Skip to content
HN On Hacker News ↗

The Rise of Audio AR

▲ 34 points • 23 comments • by dbreunig • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,643
PEAK AI % 0% · §1
Analyzed
Sep 26
backend: pangram/v3.3
Segments scanned
1 windows
avg 1643 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,643 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

The Right Information, At The Right Time, Without Looking Photo by Justin Lynham For a bit there, virtual reality (VR) and augmented reality (AR) were being heralded as the next big thing. Meta went all-in, spending billions on Oculus R&D annually and eventually rebranding the company to signal their focus. Microsoft’s Hololens had plenty of buzz, Magic Leap was the darling of the tech press and investors, Sony shipped millions of PSVR units, and Apple was rumored to be working on something big. But the hype never materialized. Oculus has yet to find its killer app and no one’s quite sure what to do with Apple’s Vision Pro. ChatGPT arrived and AI quickly became the next big thing, bumping VR/AR to the back of the line. Even Meta seems to talk more about LLaMA than AR these days. But quietly the pieces have been coming together for a different kind of AR: Audio Augmented Reality (AAR). Thanks to the rise of smart headphones, improved voice recognition and generation, AI language models, and better contextual data, AAR is set to slowly but surely seep into our daily lives. AAR will deliver the right information, at the right time without requiring us to change our focus. It’s a monumental shift in how we interact with technology, and it’s coming sooner than you think. The Pioneers of Audio AR The concept of Audio AR isn’t entirely new. Microsoft’s Soundscape and FourSquare’s MarsBot were two early attempts at delivering an Audio AR experience. Microsoft’s Soundscape was developed primarily as an accessibility tool for the visually impaired. As users walked through a city, spatial audio pings would notify them of surrounding points of interest, from laundromats to restaurants. It demonstrated the potential of using spatial audio and location-based alerts to augment reality. However, in dense urban environments, it created an overwhelming cacophony of sounds. And in suburban areas, it was too sparse to be very useful. FourSquare’s MarsBot took a more personalized, data-driven approach. It leveraged FourSquare’s extensive database of places, their ratings, and reviews to predict and recommend locations a user might find interesting. But MarsBot only thrived in a high-density environment like New York City. It was an experiment – one with great thinking about what it means to be an audio-first AR app – that never really landed. It struggled to be truly passive, demanding too much user attention to filter through irrelevant alerts. These apps, and experiments like them, broke important ground and showed Audio AR could be used for more than simple turn-by-turn navigation. Soundscape showed the power of spatial audio and location-based notifications for wayfinding. MarsBot pioneered using personalized data and experimented with more humane notifications. But it seems only a very specific set of nerds like myself took notice. Their ideas lay dormant for a few years, waiting for the right enabling technologies to arrive. The Arrival of Enabling Technologies Several key technologies have emerged in recent years, enabling a new generation of Audio AR applications: Smart wireless headphones like Apple’s AirPods have seen massive adoption, putting an always-available, hands-free audio interface in many people’s ears. They deliver high-quality sound and give developers access to a range of controls and sensor data, from volume to head movement to biometrics. Built-in microphones and simple controls make it easy to trigger assistants and audio apps. Voice recognition has made huge leaps thanks to improved noise cancellation, language models, and edge computing. Talking to virtual assistants now feels almost as natural as conversing with a person. You can speak to them casually without having to modify your speech. Speech-to-text is now fast and accurate, without the need for a network connection. Text-to-speech engines can now generate human-like voices with realistic inflection and intonation. The latest models are hard to distinguish from a real human, especially in short exchanges. This allows for the fluid generation of dynamic content. Large language models can engage in open-ended conversation, answer follow-up questions, and even take actions on the user’s behalf. They translate imperfect and inconsistent voice commands into programmatic actions enabling a more complex interface without a screen. Further, LLMs can be used as decision engines for what content to surface and when. Rich location and context data allow Audio AR apps to deeply understand the user’s environment and current situation. This includes detailed place data, real-time weather and traffic, calendar and messaging data, and more. It’s impossible to serve the right bit of information at precisely the right time without this data. Putting it all together, we now can deliver highly contextual and personalized audio content and interactions to users as they go about their day. The challenge now moves from building enabling technologies to building the UX. The Problem to Be Solved The core UX challenge for Audio AR applications is delivering the right information at the right time. Too much irrelevant information becomes overwhelming. Too little and the experience isn’t very useful. The sweet spot is frustratingly narrow. To illustrate this challenge, let’s first establish a starting point with an Audio AR app and use case that works well: audio tours. A few weeks ago, I traveled to Ghent, Belgium for an Overture Maps Foundation meeting (speaking of improved access to location data…). Having some time to explore the first morning, I checked out an audio tour of the city on VoiceMap. I have no idea how good other tours are, but this tour was perfection. So good, it suggested the potential of Audio AR. Voice Map’s tours work by using your phone’s GPS to trigger audio descriptions tied to specific waypoints. As you walk, the narrator gently directs you where to go next, providing landmarks where the next segment of the tour will pick up. Voice Map helps creators time their content to specific route segments, (in this case) achieving a relaxed yet seamless experience as you stroll. If you pause to duck into a cafe to sit for a moment, the next cue isn’t triggered and the audio pauses. Walk outside and the content picks back up. The interface was my headphones and my location, that’s it. But, to borrow a term from game design, I was on rails. The content and notifications were perfect because I’d purchased and started the tour. Voice App knew where I was going and the information I wanted. For a small slice of Ghent, it was the perfect Audio AR experience. I started wondering what it would take to cover the entire city in content and cues, enabling me to walk wherever I wanted. But even then, I’d turn the tour off if I was taking a path I’d already walked or if I just didn’t want to be disturbed for a while. Even with content coverage, cues would get tricky quickly and require more than just my current location. Time of day, my calendar, my past location history – all of this would figure into how the content should be surfaced. Creating an always-on, ambient experience that proactively surfaces relevant information is quite the challenge. After the tour, I looked up MarsBot to see if it was still available. It’s not, but Denis Crowley, FourSquare’s founder and MarBot’s creator, is taking another bite at the apple with a new company called BeeBot. BeeBot leverages LLMs to help with content creation and determining when to push an audio notification. Crowley was recently interviewed on Alex Kantrowitz’s Big Technology Podcast. Like MarsBot and SoundScape, Crowley’s approach with BeeBot is to “push” content to the user. This is undoubtedly the hard mode of Audio AR. Not only do you have to figure out what content to push, but you have to figure out when to push it. For their entry into Audio AR, Meta has taken the opposite approach. Photo by Phil Nickinson for Digital Trends A Glimpse of the Future Meta’s Wayfarer sunglasses are perhaps the best Audio AR product currently on the market. They combine the enabling technologies – smart headphones, voice recognition and synthesis, LLMs, and context data – into a single, cohesive product. Despite this, they’re kinda flying under the radar. Before I had my pair, I knew one person who used them…and they work at Meta. If you, like me, aren’t familiar with Meta’s Wayfarers here’s a quick rundown. They look nearly identical to Ray Ban’s iconic Wayfarer glasses, with a bit more heft on their arms. They have built-in speakers that do a great job of delivering spatial audio so only you can hear it; there’s no bass, but it’s great for spoken word. They have cameras that let you take pictures and short videos of your POV. And if you join the beta program, you can access Meta’s voice multi-model AI. If you want to know more about what you’re looking at, just say “Hey Meta,” and ask. To be sure, there are rough edges. The assistant refuses to answer many types of questions and is occasionally wrong. The interface is invisible, forcing users to result to trial and error. Despite this, Meta’s Wayfarers succeed because they use a “pull” UX model: users have to explicitly ask for information. The guesswork of knowing when to push content is eliminated; a welcome decision for a product whose rough edges are still very apparent (hence the beta program). Contextual data is still used to inform the assistant’s response – both with the user’s location and the camera’s view – but the user controls the timing. And when everything clicks, Meta’s Wayfarers are a glimpse of the future. Photo by Trusted Reviews Focusing on a Specific Context Sessions There’s a hybrid approach to the UX worth considering: focusing on a specific context in a session that a user proactively engages. During this period the app can push content with confidence that it’s relavent. This design pattern is put to excellent use by Apple’s Fitness app.