Ship pro audio features without training a model.
One REST API for the whole audio toolbox. Separate stems, transcribe speech, detect key and tempo, tag sounds, and more. Send a URL, get results back. No GPUs to run, no models to train.
How it works
No setup, no plugins. Just ask, and get the file back.
Grab a key
Self-serve a scar_live_ API key in minutes. 300 free tokens get you started, no sales call to try it.
POST your audio
Send one request with your audio URLs and the analysis you want. Up to 500 files in a single batch call.
We call you back
Don't sit and poll. We POST a completion webhook the moment the batch is done, then you pull the results.
Built-in storage
Every result is saved to your library, 1 TB free.
The files you bring and the processed audio results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.
Recipes
Make processed audio one step in a pipeline you run by name.
Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.
See it in action
One REST API for the whole audio toolbox.
Simple, transparent pricing
Only pay for the audio you actually process.
1 token = 1 second of audio, minimum 1 token per job.
More than one trick
Scarleta does a whole lot more.
The same account handles all of it. Here are a few very different things you can do with your audio.
For developers
One call. Any capability. Bulk or single.
Submit a batch of audio URLs against a saved recipe; we return a batch_id and run it async. Poll for status, or let the completion webhook call you. Honest, coded errors and a hard 500-files-per-batch cap. Full reference in the docs.
curl -X POST https://api.scarleta.ai/v1/batch \
-H "Authorization: Bearer scar_live_your_key" \
-H "Content-Type: application/json" \
-d '{
"recipe_id": "rec_0001_transcribe_and_tag",
"audio_urls": ["https://example.com/episode.mp3"],
"webhook_url": "https://your-app.com/hooks/scarleta"
}'
# → 202 { "batch_id": "…", "status": "IN_QUEUE" }
# Then GET /v1/batch/{batch_id}/results, or let the
# completion webhook call you when the batch is done.Add pro audio to your app today
Grab a key and ship on your free tokens. No models to train, no GPUs to run. Send a URL, get results back. Top up only when you need more.
Frequently asked questions
What can the API do?
The whole Scarleta audio toolbox from one REST API: stem separation, speech transcription with speakers and captions, key and chord detection, tempo and beat, sound tagging, noise reduction, effects, and more. You compose the steps you want into a recipe and run it by ID.
Bulk or one file at a time?
Both. Send up to 500 files in a single batch call (the GA-validated bulk surface) and we run them async, or run a single recipe over one file. Either way you get honest, coded errors and a completion webhook. Don’t poll, we call you.
How much does it cost?
You pay per second of audio, $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription. We don’t re-analyze or re-charge you for the same audio twice within your own account.
Do I have to train or host anything?
No. There are no models to train and no GPUs to run. Send a URL, get results back. Grab a key and ship today, starting on 300 free tokens.
How do I get called back?
Pass a webhook_url in your request and we POST a completion event when the batch finishes (3 delivery attempts, SSRF-guarded, idempotent). Put a secret token in your own webhook_url if you want to verify the call is ours. Webhook delivery is live; there’s no dashboard to register endpoints. You pass the URL in the API call.