1 TB free audio storage

Upload a track and ask it anything.

Drop your own audio into the Scarleta chat and talk to it in plain language. The AI figures out what to do. No menus, no settings, no code, and it plays and charts the results right in the conversation.

How it works

No setup, no plugins. Just ask, and get the file back.

Step 1

Upload your audio

Drop a track into the Scarleta chat: a full song, a stem, a voice memo. It also lands in your library, so it is there to ask about again later.

Step 2

Ask in plain language

Type what you want to know: "what's going on in this mix?", "what key is it?", "pull the vocals out". The AI reads your audio and picks the right tool, so you do not choose one.

Step 3

See it answer inline

Results come back right in the conversation: audio players you can hear, charts and tables you can read, not a wall of raw numbers.

Built-in storage

Every result is saved to your library, 1 TB free.

The files you bring and the audio analysis results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.

1 TB free Share with a link Ask across your collection
Saved to your library1 TB free
Audio analysis result
Ready to play, share, or reuse
Play Share Versions

Recipes

Make your audio analysis one step in a pipeline you run by name.

Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.

Quick track breakdown
Upload trackAsk for key and tempoExport results
Voice memo to notes
Transcribe speechLabel speakersAdd timestamps
Mix check
Upload mixCheck loudnessAsk for structure

See it in action

Upload your audio and just ask

Simple, transparent pricing

Only pay for the audio you actually process.

$0.10/ minute

1 token = 1 second of audio, minimum 1 token per job.

Prepaid, no subscription 300 free tokens to start Top up from $10

More than one trick

Scarleta does a whole lot more.

The same account handles all of it. Here are a few very different things you can do with your audio.

For developers

For developers: the Chat / Agent API

Drop the same conversational audio agent into your own app. Submit a turn, then poll for the result.

curl -X POST https://api.scarleta.ai/v1/chat \
  -H "Authorization: Bearer scar_live_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "message": "What key and tempo is this track?",
    "audio_urls": ["https://example.com/my-song.mp3"]
  }'
# → 202 { "job_id": "…", "status": "IN_QUEUE" }
# Then poll GET /v1/chat/{job_id} for the agent's reply + inline results.

Upload a track and start asking

Start on your free tokens. Drop in your own audio and ask the AI anything. Top up only when you need more. No subscription.

Frequently asked questions

What can I ask it?

Anything the platform can do to audio, from "what key and tempo is this?" to "remove the vocals" to "transcribe this". You describe the outcome in plain language and the AI picks the tool. Ask follow-ups in the same conversation.

Does it remember my track?

Yes. Audio you upload lands in your own library and stays there until you delete it, so you can come back and ask about the same track another day.

Do I get to see and hear the results?

Yes. Results appear inline in the chat: audio players you can listen to, plus interactive charts and tables, not just text.

How much does it cost?

You pay per second of audio processed, $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens; top-ups start at $10. No subscription.

Do I need to install anything?

No. It runs in the browser. Upload a track in the Scarleta chat and start asking. Developers can also call the same conversational agent from their own app through our API.

Talk to the AI About Your Own Audio — Scarleta | Scarleta