1 TB free audio storage

Drop a conversational audio agent into your product.

One endpoint. Your users, or your own agent, describe the audio job in plain language, and the Scarleta agent picks the right capabilities, runs them, and returns the results. Async standard tier for background jobs, a streaming tier for live UIs. Authenticate with your key and ship.

How it works

No setup, no plugins. Just ask, and get the file back.

Step 1

Get a key

Self-serve a key: free (sk_free_*) to start, paid (scar_live_*) for production. No sales call.

Step 2

POST a turn to /v1/chat

Send the user's message (and any audio) to the Chat / Agent API. The agent interprets intent, selects capabilities, runs them, and returns the result, so you don't wire up individual tools.

Step 3

Poll or stream the result

Poll GET /v1/chat/:job_id on the async standard tier, or use the streaming tier (POST /api/chat) for a live UI. Metered per second of audio processed.

Built-in storage

Every result is saved to your library, 1 TB free.

The files you bring and the conversational audio agent results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.

1 TB free Share with a link Ask across your collection
Saved to your library1 TB free
Conversational audio agent result
Ready to play, share, or reuse
Play Share Versions

Recipes

Make a conversational audio agent one step in a pipeline you run by name.

Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.

Voice meeting notes
Transcribe audioLabel speakersExport JSON
Catalog tagger
Detect key and tempoSeparate stemsTag and file
Clip cleaner
Remove noiseEven out loudnessTrim silence

See it in action

Drop a conversational audio agent into your product.

Simple, transparent pricing

Only pay for the audio you actually process.

$0.10/ minute

1 token = 1 second of audio, minimum 1 token per job.

Prepaid, no subscription 300 free tokens to start Top up from $10

More than one trick

Scarleta does a whole lot more.

The same account handles all of it. Here are a few very different things you can do with your audio.

For developers

Send a turn to the agent

POST a message to the async standard tier, then poll for the result. Swap to the streaming tier (POST /api/chat) when you want a live UI. Full request/response shapes and error codes live in the docs.

# Async standard tier: submit a turn
curl -X POST https://api.scarleta.ai/v1/chat \
  -H "Authorization: Bearer scar_live_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "message": "Remove the vocals from this track",
    "audio_url": "https://example.com/song.mp3"
  }'
# → 202 { "job_id": "…", "status": "IN_QUEUE" }

# Poll for the result
curl https://api.scarleta.ai/v1/chat/JOB_ID \
  -H "Authorization: Bearer scar_live_your_key"
# → { "status": "COMPLETED", "result": { … } }

Ship a conversational audio agent this week

Self-serve a key, POST your first turn to /v1/chat, and start on 300 free tokens. Move to a production key when you go live. Prepaid, no subscription.

Frequently asked questions

What is the Chat / Agent API?

A conversational endpoint: you POST a turn (a user message, optionally with audio) and the Scarleta agent decides which audio capabilities to run, runs them, and returns the results. You integrate one endpoint instead of wiring up each tool yourself.

Async or streaming?

Both. The async standard tier (POST /v1/chat) returns a job you poll with GET /v1/chat/:job_id. Good for background work. The streaming tier (POST /api/chat) streams the turn for live, chat-style UIs.

How do I authenticate?

With your own API key. Free keys (sk_free_*) let you start without a sales call; paid keys (scar_live_*) are for production. Grab one on the developer pages.

How is it priced?

Per second of audio the agent processes, $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription.

What can the agent actually do?

It reaches Scarleta's audio capabilities: splitting and isolating audio, musical and speech analysis, and more. See the developer docs for the full capability list.

Conversational Audio Agent API — Scarleta | Scarleta