Last night I had a 20 minute conversation about black holes. I’m not sure I’ve ever had such a long conversation about these celestial bodies before, simply because, although I’ve read several books on the subject, I’ve never encountered anyone who I thought would want to talk about such things. And it wasn’t just a fact-based back and forth, but rather an actual one discussion about more nebulous concepts and ideas. The most important thing was the conversation naturally. As we got into it, we felt like we were finding a flowing rhythm of speaking. I say “we,” but of course there was only one person involved in the chat. The other entity was ChatGPT.
I’ve been thinking and writing about the idea of interacting with computers using language for a long time. Ever since I was a child when I convinced my parents to buy expensive dictation software and microphones, it always seemed to me that at some point we would be interacting with machines – and that was before I really started reading or watching science fiction. It’s just obvious, right? In the days long before text messaging ruled the world, we chatted with people not through text, but through voice. Sure, there were letters and such, but that was more or less a trick to get your words spread over long distances or to the masses before we had a way to transmit language.
Obviously, using text also has advantages. Lots of them. And apparently it wasn’t just the predominant way of “talking” to machines; only Away. And to this day it is still the primary way. Depending on your workflow, it may be more efficient. But not always. And again, it was a system we developed because back then there was really no way to get machines to “hear,” let alone understand, voice input. And the issue to you was a completely different matter.
But we are here now. Machines all around us can now hear.
AI was, of course, the missing link in all of this. Back when we called it “ML,” speech recognition was finally good enough for most dictation tasks. But the real “understanding” and ability to answer you led to the LLM breakthroughs. I first wrote about this aspect a little over two years ago, when OpenAI introduced GPT-4o – their first true “Omni” model. But again, this was just a continuation of concepts I wrote about a decade ago. Because again, I’ve been thinking about these things since I was a kid.
Of course, I had real confirmation bias here and it took longer than it seemed for vocal computing to really become mainstream. But slowly but surely it is happening. You’re seeing it more and more often on the street and in offices – increasingly in doctor’s offices – where people are dictating things on their phones. But despite 15 years of promise from Siri and later Alexa, there has been a lack of real back-and-forth — actual conversations with a computer.
Some of this is cultural and social. We’ve all grown up in a world where the primary input for computers – including smartphones – is text. But most of it has remained technical. It was simply still more efficient to use text for most tasks, since you couldn’t be sure that voice input would always work. Or the voice simply wasn’t an option for many things. But most of the time it just wasn’t natural. A back and forth with an AI chatbot was still just that: a back and forth. You spoke and then had to wait for the answer. Yes, the services starting with GPT-4o have hacked together ways that allow you to pause and speed up interactions, but that often just confused both sides. With GPT-Live it seems like we are finally overcoming this hurdle and enabling a truly natural conversation with a machine.
The key to this is what OpenAI describes as a “full duplex” architecture:
GPT-Live is based on a full duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT-Live can show that it’s paying attention with phrases like “mhmm” or “yes,” engage in quick back-and-forths, or simply stay quiet when you need a moment to think. The result is a language experience that is refreshingly easy to speak with.
And while language models were previously different from text-based models, voice AI can now rely on the same best models when answering.
In my experience, on the last day of use, it’s still not perfect – there’s too much of those “mhmms” and “yes” (which is probably not just about sounding more natural, but also about buying some time to process – the drawn out “let me check” – which of course is a trick people use too!), but you can see it – and hear it! – a world in which this is perfected. Will it be exactly like talking to a human? Probably not, but would we even want that? I mean, I’m sure some people would do this for specific use cases, including combating loneliness. And I’m sure Services will perfect products and models for such use cases. But I actually like talking to a computer because I know it’s a computer but can still use our natural and great language skills as humans.
So in my black hole conversation I can push GPT-Live deeper and deeper into rabbit holes, whereas in a conversation with a human this might be strange – or even intrusive! And unlike a human, I can take the conversation with me wherever I go because these machines have a body of knowledge that literally knows no boundaries. Pardon my French: This is fucking awesome.
Sure, it’s been great for a long time – perhaps the greatest thing about LLMs in general. But there is something else that becomes clear, at least to me, when I use the voice. Here too it just feels much more natural.
Yes, yes, there are concerns that what the machine is telling you is not entirely accurate. But actually speaking to people there are far greater concerns about this! There have been great advances in the field of hallucinations in recent years, but you should probably still maintain a certain level of skepticism. Especially because the downside of the natural element of voice is that when you say something orally with great confidence, you are naturally more inclined to trust what you hear. And AI has historically been the ultimate bullshitter. But here too, people are guilty!
Either way, it feels like we’re now on the cusp of a real shift in computing. Yes, I’ve thought that for a long time, but it’s happening step by step. And these new GPT Live features appear to be opening up the space to the point where new devices may not only be possible, but inevitable.
For now, these models and features will be great for use on smartphones and laptops. While one aspect of the demo that Siri boss Mike Rockwell gave during the WWDC keynote is undoubtedly because voice provides a far better demo than text, it also seems pretty clear that we’ll soon see millions of people in the wild pressing a button to chat – loudly – with Siri. Finally.
And perhaps those who are truly committed to this method of interaction will venture into even more robust models, such as those now offered by GPT-Live. And perhaps OpenAI will use these capabilities to bring its own device to market sometime in the next few months. And perhaps many others will follow suit once it becomes clear that this is a new way to do data processing.
As I always note in posts like this, I’m not saying that voice is the be-all and end-all of computer interaction. But I say it will emerge as a new, more natural way of working with machines – similar to how Apple introduced the more natural touch capabilities of multi-touch with the iPhone two decades ago.
My children, who grew up talking to Alexa to play music and hear the weather, will instinctively understand this new world in ways we don’t. This is exciting. And of course. They will be able to go down rabbit holes and talk about black holes or anything else they can imagine. And they do this the same way they talk to people: with their voice.
👇
earlier this week Binoculars…
Netflix needs a new restructuring
Since they’re being outpaced by YouTube, they really need to focus on engagement…
![]()
The most profitable company is… Samsung?
The chip business is pushing them past Apple – and even past NVIDIA!
![]()
There is no telephone
OpenAI doesn’t build a phone. Amazon doesn’t make a phone. And now SpaceX isn’t building a phone. Do you sense a trend? The trend is that basically the entire tech industry is building its own smartphones again. But they all deny it.
![]()
https://spyglass.org/gpt-live/
