Sonic Digest #003
Agentic Ears, Sonic Evidence, and Listening Through the World
Welcome to the third Sonic Digest, a slower dispatch from Sonic Field for those who want to receive some of our recent signals in a more condensed form.
This past month, the pieces we published kept returning to one simple but expansive question: what does listening do? Listening appeared as evidence, as interface, as bodily knowledge, as ecological relation, as memory, as software protocol, as studio infrastructure, and as a way of being with others. So this issue is grouped by a few currents rather than presented as a straight list. There is a lot here, but the constellation is quite clear: listening is becoming operational on many fronts at once.
1. AI audio is becoming conversational, agentic, and situated
One of the strongest threads this month was the rapid movement of AI audio away from isolated generation and toward systems that hear, act, remember, and operate inside computational workflows.
In our article on AI Agents Enter the Studio and The MCP Turn in Audio, we looked at how MCP connectors are beginning to bring AI agents into production environments such as Ableton, Splice, Max, REAPER, and other creative tools. The important shift is not simply “AI makes music,” but something more infrastructural: manuals, sample catalogs, DAW controls, audio processes, and creative decisions are becoming addressable through language. The studio starts to appear as an agentic network.
That same shift appears in VOID Link Audio Connectors, the Networked Audio Between Creative Tools created by Julien Bayle, although operating from another angle, and focusing on interoperability: Max/MSP, Pure Data, TouchDesigner, VCV Rack, openFrameworks, and other creative systems exchanging beat-synced audio over a local network. It is a small but meaningful example of the studio becoming less a closed setup and more a field of connected processes.
On the model side, MOSS-Audio, An Open-Source Audio Understanding Model points toward another important layer: audio understanding beyond transcription. Instead of treating sound only as speech to be converted into text, MOSS-Audio is designed to work with multiple acoustic dimensions: timing, background texture, musical cues, speaker conditions, and environmental information. It matters because open audio models are becoming more capable of dealing with sound as sound, not merely as linguistic content, so they can process audio data in new ways, thus arising important questions specially in regards to the soundscape this unfolds.
Inworld Realtime TTS-2, A Voice Model for Conversational Presence shows that synthetic speech is no longer framed as static narration, but as a real-time layer of interaction. Voice direction, conversational awareness, emotional context, multilingual continuity, latency, and timing all become central. Then we reported that AI Voice is leaving the turn-based era and how this widens the frame: recent speech systems are moving beyond the old pipeline of speech-to-text → language model → text-to-speech. The new object is the realtime speech agent: something that listens while the conversation is happening, handles interruptions, reasons through tasks, switches languages, calls tools, and responds with timing and expression that seems to aim to simulate human interaction.
These new conditions and challenges in the public sonic space also call for new ways of listening where the depth and detail reached by the generation process, is also backed up with an equivalent interpretation and processing capabilities, so we can ask about intentions, agencies, forces and powers behind this new voice landscape. This connects directly with Sonic Field Labs’ own experiment which we made public in the past month: AKOÚŌ: Operational Ears for Agentic Listening, a framework created for asking how AI agents could listen, not only what they should classify, control and transfer. The system proposes several listening modes in the form of agent skills aimed to separate signal inspection, embodied listening, forensic listening, ecological listening, political interpretation, aesthetic dimensions and speculative routes, so that machine listening does not collapse measurement, inference, and imagination into the same flat claim. This feels important: as agents begin to hear and act, we need better ways of describing what kind of listening they are performing.
All of this suggests that AI audio is becoming less about isolated outputs and more about situated interaction. The question is no longer only “can a model generate sound?” but “how does it listen, when does it speak, what can it touch, and what kind of agency does it gain inside the sonic workflow?”
2. Listening as evidence, investigation, and public knowledge
All that also lead us to another important thread in Sonic Field this past month, that is listening as evidence, expansion, revelation; listening as knowing what happened.
We covered the impressive work of Earshot, a nonprofit using sound in human and environmental rights investigations. Their work with audio ballistics, sonic profiling, sound authentication, earwitness testimony, and eco-acoustics expands what can count as evidence. A gunshot, drone, blast, engine, voice, acoustic reflection, or barely audible frequency can become part of a reconstruction when cameras are absent, obstructed, or insufficient.
This is not a naive faith in sound. The most interesting part is the care around interpretation. Hearing is unstable, embodied, contextual, and shaped by traumas of power. That makes acoustic evidence both powerful and delicate.
A more artistic and speculative version of this appeared in Sonic Detection: Listening as Investigation, our note on Rebecca Collins and Johanna Linsley’s open-access book Sonic Detection: Necessary Notes for Arts and Performance. The book treats listening as clue, method, atmosphere, fieldwork, fiction, archive, and collective practice. It imagines the listener as a kind of detective, as sound lead us into traces, residues, margins, and unstable forms of testimony. Together, Earshot and Sonic Detection show two sides of an important contemporary problem. Sound can be proof, but also atmosphere. It can support legal and human rights work, but it can also carry uncertainty and simulacra.
3. Listening through bodies, places and shared attention
Several posts this month moved away from audio as object and toward listening as a relational practice: something that happens through bodies, places, communities, and conditions of attention.
Soundcamp / Reveil 2026: Following the Dawn returned with one of the most beautiful yearly rituals in sound culture: a 24+1 hour planetary broadcast following daybreak around the Earth through live environmental streams. Reveil always reminds us that listening can be a distributed act of care. Microphones, habitats, birds, cities, wetlands, forests, radios, and local communities become part of a fragile acoustic commons.
That sense of listening as situated relation also appears in AudioSpaces and the Geography of Sonic Memory, a project that lets people record sounds and attach them to places. The simple gesture of pinning audio to a map opens bigger questions: whose memory gets mapped, what sounds are worth preserving, and what happens when listening requires movement through space? Sound becomes local, embodied, revisitable, also algorithmically connected and reinterpreted, socially listened and shared, an expanded geological ear.
The Listening Biennial: Listening as practice, care, and shared research expanded this further. The project treats listening as an active practice of attention toward places, ecologies, communities, and forms of life often pushed aside. Across exhibitions, workshops, commissioned works, publications, and collective study, it frames listening as a shared method for care and repair.
The body entered this thread directly through Dance and Silence: Listening Through the Body, a book that approaches silence through dance, ecology, architecture, neuroscience, signing, performance, and embodied knowledge. This is especially useful for sound studies because it asks us to think beyond sound itself and how silence can be a field of movement and bodily intelligence.
And in Auditory Perception in Twentieth-Century Self-Narratives, Claudia Cerulo’s work on the “oto-bio-graphical” subject brings listening into autobiography and literary self-formation. Acoustic memory as a pre-verbal experience becomes a force that shape how the self is written.
4. Instruments, books, and material listening
A final cluster this month returned to tools and instruments with posts that ask how devices, books, and modular systems organize ways of thinking with sound.
MAGMA: Listening Through Contact and Immersion introduced Jorge Barco’s handcrafted contact hydrophone, released through VIC NIC. MAGMA works as both hydrophone and contact microphone, inviting listening through matter: water, metal, soil, you name it. The devices is itself a proposition that listening may begin where the body touches the world.
Modular Synthesizers Book pointed to Heiner Kruse’s wide-ranging guide to modular synthesis past and future, covering both the tech and the culture. Modular here means an open-ended instrument, a method of composition, and a way of learning signal flow through practice.
What ties this month together
Looking across this batch, I keep returning to the same feeling: listening has never been easy to contain. It appears in AI agents and realtime voices, in forensic investigations and sonic fiction, in dance and silence, in sound maps and collective broadcasts, in hydrophones, modular systems, networked studios, literary memory, open-source models, and listening-based research communities. This is probably why sound feels so urgent and vast right now. It is being automated, modeled archived, weaponized, protected, mapped, performed, translated, and reimagined at the same time. But the one of the most interesting questions remains human and situated: how do we listen without reducing what is heard?
The answer is not one single method. Perhaps it is a plurality of ears in an open acoulogical era in which technical ears, bodily ears, ecological ears, forensic ears, literary ears, speculative ears, communal ears, agentic ears, all converge, contrast, dialog, clash, mutate together. Each one hears differently, and each one carries its own risks.
Take your time with this batch. There are lots of tools, books and projects here, but also a shared invitation: to treat listening not as passive reception, but as a way of entering relation with the world. As always, we would love to hear what resonates, what bothers you, and what kind of sonic work, research, or listening practice you have been carrying lately.
Our ears are open ✿


