Your words. Your labels. Scarleta scores the match.
Write the labels you care about in plain words: "crowd noise," "acoustic guitar," "angry tone," anything. Scarleta scores how well each one matches the audio. No fixed tag list to pick from, and nothing to train.
How it works
No setup, no plugins. Just ask, and get the file back.
Upload your audio
Drop a clip into the Scarleta chat: a recording, a field capture, or a sample from your catalog. No plugins to install.
Write your own labels
Give it the words you care about in plain language: "dog barking," "distorted," "calm voice." You choose the labels. There is no preset list to fit into and no model to train.
Read the match scores
Get a score for how well each label matches the audio, laid out as a table or chart right in the conversation so you can rank, threshold, or triage.
Built-in storage
Every result is saved to your library, 1 TB free.
The files you bring and the label match scores results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.
Recipes
Make label match scores one step in a pipeline you run by name.
Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.
See it in action
Score any audio against your own labels
Simple, transparent pricing
Only pay for the audio you actually process.
1 token = 1 second of audio, minimum 1 token per job.
More than one trick
Scarleta does a whole lot more.
The same account handles all of it. Here are a few very different things you can do with your audio.
For developers
Do this at scale with the API
Running a whole library? Send a batch of files to a custom-label recipe and let a webhook call you back when each one is scored. Your labels are defined in the recipe.
curl https://api.scarleta.ai/v1/batch \
-H "Authorization: Bearer $SCARLETA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"recipe_id": "rec_0001_custom_audio_labels",
"audio_urls": [
"https://example.com/clip.wav"
]
}'
# 202 + batch_id. Poll GET /v1/batch/:id/results,
# or let the completion webhook call you when scoring is done.Score your first clip on your free tokens
Upload a clip, write the labels you care about, and read the match scores. No fixed list and nothing to train. Top up only when you need more.
Frequently asked questions
What do I get back?
A match score for each label you wrote, showing how well it fits the audio. You can see the scores as a table or chart right in the chat, and export them as CSV or JSON for your own tools.
How is this different from 'identify the sounds'?
Sound identification tells you what is in a clip from open-ended detection: you do not pick the words. This is the opposite: you write your own labels and it scores the audio against exactly those. Bring your own vocabulary.
Do I have to train a model or pick from a preset list?
No. There is no training and no fixed list. You type whatever labels you want and it scores them straight away.
How much does it cost?
You pay per second of audio: $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription.
Can I run this on a lot of files from my own app?
Yes. Developers can send many files at once through our API and get called back by webhook when every one is done. See the docs.