Agents are only as capable as the tools they can reach. An agent that can browse the web can research; one that can run code can compute. But ask most agents to actually do something with a piece of audio (separate the vocal, transcribe the speech, find the key and tempo) and they're stuck, because they have no ear.
We built the ear as something any AI can plug into. Scarleta ships a live MCP server: an endpoint that outside AIs connect to with a key and use Scarleta's audio tools natively.
An AI other AIs can use
The Model Context Protocol is how modern agents discover and call tools. Expose your capabilities over it and any MCP-capable AI (Claude, ChatGPT, a custom agent you built) can call them as first-class tool-calls, not hand-rolled request plumbing.
That's exactly what the MCP server does. It's a developer offering alongside the REST and batch API, exposing the same audio capabilities: separate stems, remove or isolate an element, transcribe speech with timestamps and speakers, detect key, tempo, and structure, clean up noise, and more. The agent doesn't need to know how any of it works. It just sees an audio toolbelt and reaches for the right tool.
How it fits together
- Connect with a key. Your agent authenticates with a customer key, the same way it would reach any other service.
- Call tools natively. The capabilities show up as tools the model can invoke directly. No bespoke HTTP client, no polling loop to babysit.
- Billed per use. It runs on the same per-second-of-audio meter as the rest of the platform. Same prepaid tokens, nothing new to learn.
Why this is the natural shape of "the AI that listens"
If listening is a modality (an ear for machines), the most useful place for that ear to sit is right next to every other AI, ready to be called. Not a separate app someone has to leave their agent to go use, but a tool that lives inside the agent's own reach.
Give any AI an audio toolbelt. It connects, it calls, it understands sound, and you pay only for what it actually runs.
The MCP server is a live developer surface: connect with a customer key, call the audio capabilities natively, billed per use on the standard per-second meter. See Developers.