Drop a conversational audio agent into your product.
One endpoint. Your users, or your own agent, describe the audio job in plain language, and the Scarleta agent picks the right capabilities, runs them, and returns the results. Async standard tier for background jobs, a streaming tier for live UIs. Authenticate with your key and ship.
How it works
No setup, no plugins. Just ask, and get the file back.
Get a key
Self-serve a key: free (sk_free_*) to start, paid (scar_live_*) for production. No sales call.
POST a turn to /v1/chat
Send the user's message (and any audio) to the Chat / Agent API. The agent interprets intent, selects capabilities, runs them, and returns the result, so you don't wire up individual tools.
Poll or stream the result
Poll GET /v1/chat/:job_id on the async standard tier, or use the streaming tier (POST /api/chat) for a live UI. Metered per second of audio processed.
Built-in storage
Every result is saved to your library, 1 TB free.
The files you bring and the conversational audio agent results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.
Recipes
Make a conversational audio agent one step in a pipeline you run by name.
Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.
See it in action
Drop a conversational audio agent into your product.
Simple, transparent pricing
Only pay for the audio you actually process.
1 token = 1 second of audio, minimum 1 token per job.
More than one trick
Scarleta does a whole lot more.
The same account handles all of it. Here are a few very different things you can do with your audio.
For developers
Send a turn to the agent
POST a message to the async standard tier, then poll for the result. Swap to the streaming tier (POST /api/chat) when you want a live UI. Full request/response shapes and error codes live in the docs.
# Async standard tier: submit a turn
curl -X POST https://api.scarleta.ai/v1/chat \
-H "Authorization: Bearer scar_live_your_key" \
-H "Content-Type: application/json" \
-d '{
"message": "Remove the vocals from this track",
"audio_url": "https://example.com/song.mp3"
}'
# → 202 { "job_id": "…", "status": "IN_QUEUE" }
# Poll for the result
curl https://api.scarleta.ai/v1/chat/JOB_ID \
-H "Authorization: Bearer scar_live_your_key"
# → { "status": "COMPLETED", "result": { … } }Ship a conversational audio agent this week
Self-serve a key, POST your first turn to /v1/chat, and start on 300 free tokens. Move to a production key when you go live. Prepaid, no subscription.
Frequently asked questions
What is the Chat / Agent API?
A conversational endpoint: you POST a turn (a user message, optionally with audio) and the Scarleta agent decides which audio capabilities to run, runs them, and returns the results. You integrate one endpoint instead of wiring up each tool yourself.
Async or streaming?
Both. The async standard tier (POST /v1/chat) returns a job you poll with GET /v1/chat/:job_id. Good for background work. The streaming tier (POST /api/chat) streams the turn for live, chat-style UIs.
How do I authenticate?
With your own API key. Free keys (sk_free_*) let you start without a sales call; paid keys (scar_live_*) are for production. Grab one on the developer pages.
How is it priced?
Per second of audio the agent processes, $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription.
What can the agent actually do?
It reaches Scarleta's audio capabilities: splitting and isolating audio, musical and speech analysis, and more. See the developer docs for the full capability list.