Forget iPhone Duo, I’m more interested in Audio Intelligence – privacy nightmare or accessibility genius?

0
2
Forget iPhone Duo, I’m more interested in Audio Intelligence – privacy nightmare or accessibility genius?


Kerry Wan/ZDNET

ZDNET’s key takeaways

  • AI-powered Audio Intelligence offers new features.
  • These features offer users an accessibility win.
  • Privacy-minded folks should consider implications.

Apple’s September hardware event featured one of its most anticipated devices, the iPhone Duo, two new Apple Watches, an upgraded pair of base-model AirPods, and several technological upgrades to under-the-hood mobile processors. However, one of the more interesting AI-powered upgrades is Apple’s Audio Intelligence, an umbrella term encompassing new audio-related accessibility features for Apple Watch 12 wearers.

Also: How to preorder the new iPhone Duo, iPhone 18 Pro, and everything else Apple announced

Audio Intelligence can alert Apple Watch 12 owners to important noises, such as a doorbell or a baby crying; rewind live conversations to catch missed words; briefly summarize conversations; and, most importantly, Shazam when you hear a good song at the grocery store.

These features represent an accessibility win for those who are hard of hearing or have memory issues, but do others want another device listening in on their every word? Here’s what Apple says to privacy-minded folks.

What are the Audio Intelligence features?

  • Sound Recognition: Alerts the Apple Watch when it hears urgent sounds, such as a doorbell, siren, or alarm. The wearer then receives a notification. Apple says this feature will benefit hard-of-hearing users, though I can also see it benefiting work-from-home employees who often wear headphones.
  • Live Rewind: Allows users to double-press the Digital Crown to retrieve the last 15 seconds of an active conversation. I can see this feature being useful for hard-of-hearing people and those who speak a foreign language, but still miss some things here and there.
  • Siri Recap: Creates summaries of a real-time conversation, with a title and conversation notes automatically saved to the Siri app.

What does Apple say about privacy?

Privacy has been at the core of Apple’s messaging for years, with the company deploying web-tracking blockers, localized data processing, and end-to-end encryption across its ecosystem.

According to Apple’s press release, Live Rewind snippets and Siri Recaps are end-to-end encrypted, and not even Apple can access them. Apple says the Watch 12’s S11 chip features Secure Exclave, a portion of the chip that processes audio separately from the rest of the system and immediately discards it.

Also: This LG TV jailbreak turns smart devices into spies, say researchers

The same press release states that speaker attributions are not recorded in Live Rewind or Siri Recap, so the speakers involved in the conversation are not identifiable. Additionally, Apple says Siri Recaps are brief summaries of a conversation, not a recorded transcript, and that the software is designed to omit personal information, such as financial details or “government-assigned identifiers.”

Apple Watch 12 wearers can toggle Siri Recap off at any time in the Control Center.

Live Rewind, which relays the previous 15 seconds of a conversation, emits a chime when it’s enabled, even when the Apple Watch is in Silent Mode. A full-screen animation will appear on the watch’s display when this feature is on, so people nearby can see it.

What about two-party consent?

Apple’s press release states that “… raw audio used for processing is completely inaccessible to the operating systems, apps, the user, or Apple.” This statement leads me to believe that the sound that reaches the Apple Watch 12’s microphone, whether it’s speech or a siren, enters the Secure Exclave for privacy, and that the Live Rewind snippet or Siri Recap summary is information processed by a speech-recognition AI model.

This approach means the output of these features could differ from a standard voice recording. The output could be inaccurate if the AI model processes background noises, overlapping voices, or fast or slurred speech.

Also: iPhone 18 Pro hands-on: I came for the Burgundy and stayed for the dual-aperture cameras

For example, a standard voice recording may be: “So, um, I was telling my coworker last week that uh … I wanted to ask for an extension on the, um, what’s it called, that project we discussed like, three weeks ago because, you know, things aren’t going well.” Every filler word, stutter, or mid-sentence pivot is recorded verbatim.

And from this bit of conversation, the Siri Recap could be: “Told coworker project needs an extension; it’s not going well.”

In summary, Apple’s Audio Intelligence features aren’t storing recordings of your voice — the raw audio is immediately deleted, and which speaker said which words is unidentifiable. However, the watch did reproduce remnants of what another speaker said, even though an uninvolved third party couldn’t identify the speaker from the summaries alone.

Also: Wait, no base iPhone 18? Why Apple likely skipped its fan favorite this year

So, to answer the question of ‘What happened to two-party consent?’ My answer is that, as with many instances surrounding fast-moving AI-powered technologies, the laws haven’t caught up yet. Two-party consent means all parties in a conversation, whether in person or over the phone, must be aware it is being recorded.

Apple Watch 12 isn’t recording the conversation; it’s processing the audio, deleting it, and spitting out a brief AI-generated summary. Therefore, the tougher question is: ‘Do machine-processed summaries of raw audio equal a deceptive recording of someone?’

Alas, I’m not a lawyer, so I can’t answer that question. In the meantime, if you’re concerned about these kinds of things, inspect your friends’ and family’s wrists before you tell them what you did last summer.