1 TB free audio storage

250+ audio features, ready for your model.

Pull deep spectral and perceptual feature data out of any track: MFCC, chroma, HPSS, ZCR, tonnetz, chroma CENS, LUFS loudness, and more, as numeric feature blocks over the batch API. Built for ML pipelines and audio research.

How it works

No setup, no plugins. Just ask, and get the file back.

Step 1

Submit your audio to the batch API

POST your file URLs to /v1/batch with a feature-extraction recipe and your API key. One call handles a whole set of files, up to 500 per batch.

Step 2

Poll or receive a webhook

Track status with GET /v1/batch/{id}, or pass a webhook URL and we call you back when the batch completes. No polling loop needed.

Step 3

Read the numeric feature blocks

Pull per-item results and consume the feature blocks: MFCC, chroma, HPSS, ZCR, tonnetz, CENS, LUFS, and more, straight into your model or notebook. Export as JSON or CSV.

Built-in storage

Every result is saved to your library, 1 TB free.

The files you bring and the audio feature data results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.

1 TB free Share with a link Ask across your collection
Saved to your library1 TB free
Audio feature data result
Ready to play, share, or reuse
Play Share Versions

Recipes

Make audio feature data one step in a pipeline you run by name.

Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.

ML training prep
Extract featuresExport as CSVFeed into model
Catalog tagger
Batch submit tracksRead loudness and tempoTag and label
Research pipeline
Submit audio URLsPull feature blocksExport as JSON

See it in action

250+ audio features, ready for your model

Simple, transparent pricing

Only pay for the audio you actually process.

$0.10/ minute

1 token = 1 second of audio, minimum 1 token per job.

Prepaid, no subscription 300 free tokens to start Top up from $10

More than one trick

Scarleta does a whole lot more.

The same account handles all of it. Here are a few very different things you can do with your audio.

For developers

Extract features over a batch

Submit many files at once, then poll for results or receive a completion webhook. Full request/response shapes and every error code live in the docs.

curl -X POST https://api.scarleta.ai/v1/batch \
  -H "Authorization: Bearer $SCARLETA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "recipe_id": "rec_0004_audio_features",
    "audio_urls": [
      "https://example.com/track.wav"
    ]
  }'

# 202 { "batch_id": "...", "status": "IN_QUEUE" }
# Then poll GET /v1/batch/:id/results, or let the completion webhook call you.

Feature data for your ML pipeline, on tap

Start on your free tokens: submit a batch, pull the feature blocks, and feed them straight into your model. Top up only when you need more. No subscription.

Frequently asked questions

What exactly do I get back?

Numeric feature blocks. From spectral analysis: 250-300+ timbre and spectral features including MFCC, chroma, and HPSS. From signal analysis: mathematical and perceptual features like ZCR, tonnetz, chroma CENS, and LUFS loudness. You can export the analysis data as JSON or CSV.

Is this a bulk / batch API?

Yes. The batch API takes up to 500 audio URLs per request. Submit once, then poll GET /v1/batch/:id/results or receive a completion webhook. See the docs for the full contract and error codes.

How much does it cost?

You pay per second of audio: $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups begin at $10. No subscription. Batch is a paid feature.

Can I run one file instead of a batch?

The batch API handles a single URL just as well as 500. Single-recipe runs are available too. The developer docs cover both.

Where's the full API reference?

On the developer pages. This page is the quick pitch for feature extraction. /developers and /docs have every endpoint, parameter, limit, and error code.

Audio Feature Extraction API: 250+ Features — Scarleta | Scarleta