1 TB free audio storage

Give any AI a full audio toolbelt.

Connect your agent to Scarleta's MCP server with a key, and it can separate stems, transcribe speech, and detect key and tempo, as native tool calls, not hand-rolled REST plumbing. Billed per use, same prepaid tokens.

How it works

No setup, no plugins. Just ask, and get the file back.

Step 1

Grab a key

Self-serve a customer key at signup. No sales call. The same key that authenticates the REST and batch API connects your AI to the MCP server.

Step 2

Point your AI at the MCP server

Add Scarleta's MCP server to your agent's MCP config with your key. Claude, ChatGPT, or a custom agent, anything that speaks the Model Context Protocol.

Step 3

Your AI calls the audio tools

The tools show up natively in your agent. It can separate stems, transcribe, or detect key and tempo on demand. You pay per second of audio it processes.

Built-in storage

Every result is saved to your library, 1 TB free.

The files you bring and the audio tool calls from your agent results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.

1 TB free Share with a link Ask across your collection
Saved to your library1 TB free
Audio tool calls from your agent result
Ready to play, share, or reuse
Play Share Versions

Recipes

Make audio tool calls from your agent one step in a pipeline you run by name.

Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.

Agent transcription pipeline
Connect via MCPTranscribe audioLabel speakers
Music analysis agent
Connect via MCPDetect key and tempoExport results
Stem separation workflow
Connect via MCPSeparate stemsDownload stem files

See it in action

Give any AI a full audio toolbelt via MCP.

Simple, transparent pricing

Only pay for the audio you actually process.

$0.10/ minute

1 token = 1 second of audio, minimum 1 token per job.

Prepaid, no subscription 300 free tokens to start Top up from $10

More than one trick

Scarleta does a whole lot more.

The same account handles all of it. Here are a few very different things you can do with your audio.

For developers

Connect your agent

Add Scarleta's MCP server to your agent with your customer key, the same key and same per-second billing as the REST and batch API. The exact endpoint URL and the full tool list live in the developer docs.

{
  "mcpServers": {
    "scarleta": {
      "url": "https://<see-developer-docs>/mcp",
      "headers": {
        "Authorization": "Bearer scar_live_your_key"
      }
    }
  }
}

Give your agent an audio toolbelt today

Grab a key and point your AI at the MCP server. It starts on your free tokens, billed per second of audio, no subscription. The full setup is in the developer docs.

Frequently asked questions

Which AIs can connect?

Any AI that speaks the Model Context Protocol: Claude, ChatGPT, or a custom agent you build. It connects to Scarleta's MCP server with your customer key.

What can my AI do once it's connected?

The same audio capabilities as the rest of the platform, exposed as native tool calls: separate stems, transcribe speech, detect key and tempo, and more. See the docs for the full tool list.

How is it billed?

Per use, on the same meter as everything else: 1 token = 1 second of audio, $0.10 per minute, from prepaid tokens. New accounts start with 300 free tokens, and top-ups begin at $10. No subscription.

Do I need a different key from the API?

No. Your customer key authenticates the MCP server the same way it authenticates the REST and batch API. Grab a key once and use it everywhere.

Is the MCP server actually live?

Yes. It is a live developer offering alongside the REST and batch API. Head to the developer docs for the connection details.

Give Any AI a Full Audio Toolbelt via Scarleta MCP — Scarleta | Scarleta