Skip to content

Media: Processing & Generation โ€‹

Inclavate handles media in two directions, and they are at very different stages:

DirectionWhat it meansStatus
Processing (media in)Send an image, audio clip, or video frames; the model answers in textโœ… Shipped
Generation (media out)Ask for an image, speech, or video and get media back๐Ÿšง Speech works; image and video not yet

This page covers both, plus the artifact storage that generation is built on and that you can already use today.


Media processing โ€‹

Send media to a model and it answers in text โ€” description, transcription, analysis, or OCR. Nothing is generated; the media is read.

A model exposes a modality when three gates pass:

  1. its text decoder graph is supported,
  2. a compatible projector (mmproj) GGUF loads, and
  3. every node in the chain advertises the decoder metadata that projector needs.

There is no model-family allowlist. Capabilities come from inspecting the projector, so a model whose decoder and projector are both already supported works with no Inclavate code change.

Supported today: Gemma 4 ยท Gemma 3 ยท Qwen2-VL ยท Qwen2.5-VL (3Bโ€“72B) ยท Qwen3-VL (2Bโ€“32B, 30B-A3B, 235B-A22B) ยท GLM-4.6V-Flash ยท DeepSeek-OCR ยท Ultravox ยท Qwen3-ASR. See Models for the full matrix.

Raw media never reaches decoder nodes โ€” it is projected to embeddings inside the orchestrator, and only embeddings cross the network.

Speech to text โ€‹

There are two ways to turn speech into an answer, and which one a deployment uses is decided by the chat model, not by a setting.

If the chat model has an audio projector โ€” Gemma 4's E-series does โ€” the clip goes to it whole. One model hears and answers: two model hops instead of three, one resident model instead of two, and no transcript in between to be wrong about.

If the chat model is text-only, a transcriber goes in front of it. Any audio-capable model can be one; qwen3-asr:1.7b is the dedicated choice. Pull its projector too โ€” the decoder alone has no ears:

bash
dppan pull hf:ggml-org/Qwen3-ASR-1.7B-GGUF:Q8_0
dppan pull hf:ggml-org/Qwen3-ASR-1.7B-GGUF:Q8_0 --projector-only

A transcript on its own comes from POST /v1/audio/transcriptions, which is OpenAI-compatible โ€” see REST API. One field behaves differently there and it is worth knowing why: language is a filter, not a hint. These models take no instruction about language, but they do report what they believe they heard, and handed near-silence they return a fluent sentence rather than nothing โ€” often in another language entirely. A transcript whose reported language disagrees with the caller's comes back empty, which is the honest answer to "what was said" when nothing was.

In the dashboard, the microphone attaches one clip to a message; the voice control beside it opens a hands-free conversation that listens continuously, decides for itself when you have finished, and reads the reply back aloud when a speech model is configured.

Video: two ways to send it โ€‹

input_video takes either frames or file, never both.

framesfile
What you sendimages you extracted, with timestampsa whole container
Who decodesyouthe server
Needs on the hostnothingffmpeg โ€” see below
Containersn/amp4, mov, webm, mkv
Ceiling16 frames per part240 frames, 8 MiB encoded

These are complementary, not one replacing the other. Sampling frames covers a long clip thinly and lets you choose which moments matter โ€” a ten-minute video works. file covers a short clip densely, up to 240 frames, and the server does the extraction.

The dashboard and dppan run --video-frames use the frames path, so they work on every deployment with no extra setup.

Optional: ffmpeg for server-side video decode โ€‹

Sending a whole video file requires ffmpeg and ffprobe on the host's PATH. Everything else โ€” images, audio, and video frames โ€” needs nothing.

DPPAN does not ship or download ffmpeg. Install it yourself if you want this feature:

bash
# macOS
brew install ffmpeg

# Debian / Ubuntu
sudo apt install ffmpeg

# Fedora / RHEL
sudo dnf install ffmpeg

# verify both are reachable
ffmpeg -version && ffprobe -version

Both binaries must be on the PATH of the process running the orchestrator โ€” not just your shell. Under systemd that usually means setting Environment=PATH= in the unit file, or installing to a standard location such as /usr/bin.

Check before sending. GET /api/models reports:

json
{ "supports_video_files": true }

It describes the host, not the model: a deployment without ffmpeg cannot decode video files however its models are configured. A request sent to a host without it is refused with that reason, so the fix is obvious from the error.

How much video fits. The decode is held whole in memory while the model consumes it, so the ceiling is frames ร— width ร— height, not file size:

ResolutionFramesAt the default 4 fps
512ร—51285~21 s
720ร—48064~16 s
1280ร—72024~6 s
1920ร—108010~2 s

For anything longer, downscale first or extract frames yourself โ€” a sampled ten-minute clip is entirely workable, a decoded one is not.

As a job, not a chat turn โ€‹

Everything above sends media through a chat completion, which streams its answer and needs a connection held open for as long as the model takes. The same reading can be queued instead, as an understanding job โ€” modality: "text" on the ordinary generation API:

bash
curl -X POST http://localhost:8001/v1/generations \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3vl:4b","modality":"text",
       "input_artifacts":["66815206-..."],
       "params":{"prompt":"What is the total on this receipt?"}}'

The answer arrives as a text/plain artifact rather than a stream. That is the whole trade: you give up watching it appear, and you get a job that survives a restart, can be polled or subscribed to later, and does not tie up a connection. For a batch of receipts, or a clip that takes a minute to read, it is the right shape. For a conversation, it is not โ€” a job takes one prompt and answers once.

It is the same queue and the same worker as speech and music, so a deployment generating a song will not start reading your receipt until the song is done.

Which models can do this, what each one can read, and every job error code are in REST API โ†’ Understanding.


Media generation โ€‹

Generating images, speech, or video needs three things: somewhere for produced bytes to live, asynchronous jobs to produce them, and the generator runtimes themselves.

All three have shipped for speech and for music โ€” artifacts, generation jobs, and two generator runtimes over them. Image and video generators are not built, so requesting those modalities answers 422. The OpenAI-compatible Images and Audio endpoints are thin adapters that will sit over the same job API, not a second implementation.

Speech and music both produce audio, and that is where the resemblance stops. A song has no voice, no language and no reference clip; an utterance has no lyrics, no length target and no solver steps. They are therefore separate routes with separate request bodies, and a generator says which kind it is โ€” see kind โ€” so a client renders the right form instead of showing every control to both.

Which kind of generator โ€‹

GET /v1/generations/capabilities describes every generator this deployment serves. Each carries a modality and a kind:

json
{ "modality": "audio", "kind": "music", "model": "minimax",
  "limits": { "max_duration_seconds": 10.0 } }

modality stopped being enough the moment a second audio generator existed. Speech and music both produce audio and share no parameters at all, so a client keying off modality alone would show every control to both and send fields the server silently ignores. kind is what a UI should branch on.

An older server omits kind; read that as speech, which is the only thing those servers served.

Speech โ€‹

Available models โ€‹

ModelCodecRateSizeBuilt-in voicesRegistries

Fetched live from dl.inclavate.io/models/tts-catalog.json โ€” the same catalog the CLI reads for dppan models --tts. Speech models are listed apart from text models because a speech model is a pair of files that always run together and never shards, so the layer count that table is organised around does not apply.

Music models have a catalog of their own again โ€” a music model is a bundle of four or five components with a quant each, so it shares neither schema. Three files, three shapes, three pull commands, and the wrong command for a shape does not work.

A model is defined by its codec, and the two differ in ways a caller can see:

qwen3ttsorpheusf5
Voice comes froma reference clip, requireda built-in name, or nothinga reference clip and its transcript, both required
Languagesten, selected per requestfixed by the checkpoint, not the requestfixed by the checkpoint โ€” eleven Indian languages for IndicF5
Decoder shipsinside the backbone's repoin a separate repoinside the same single file
Frame rate12.5 Hz11.72 Hz93.75 Hz

f5 is the odd one out: one file, not a backbone and a codec.

You do not choose a codec; you choose a model, and everything above follows from it โ€” including how you pull it. Ask the running deployment which applies rather than inferring it: GET /v1/generations/capabilities reports each generator's voices, languages and limits.

Getting a model โ€‹

dppan pull --speech fetches both halves โ€” the backbone and its codec decoder โ€” and records that they belong together:

bash
# decoder ships in the backbone's own repo (Qwen3-TTS)
dppan pull hf:ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF --speech

Some families publish the decoder separately, shared by every checkpoint of the family โ€” an Orpheus backbone is a plain Llama, and the 25 MB SNAC decoder it needs lives in its own repo. Name it with --speech-decoder and the pair is recorded exactly as if the two had shipped together:

bash
dppan pull hf:PkmX/orpheus-3b-0.1-ft-Q8_0-GGUF --speech \
  --speech-decoder hf:cstr/snac-24khz-GGUF:24khz

Without --speech-decoder, that pull fails: there is no speech model without both halves, so a success that quietly produced one would give you half a thing and no reason to look. The orpheus codec always needs the flag; qwen3tts never does.

Pulling a second model of the same family needs one more flag. Every Orpheus checkpoint shares the same SNAC decoder, and the file recording which backbone a decoder belongs to lives beside that decoder and names exactly one. Pull a second Orpheus model to the default location and it overwrites the first pairing โ€” the first model then vanishes from dppan models --tts even though its bytes are still on disk. Give each pair its own decoder path:

bash
dppan pull hf:lex-au/Orpheus-3b-Hindi-FT-Q8_0.gguf --speech \
  --speech-decoder hf:cstr/snac-24khz-GGUF:24khz \
  --projector-out ~/.dppan/models/projectors/orpheus-hindi/snac-24khz.gguf

# Veena โ€” Hindi and English from one checkpoint, including code-mixed input
dppan pull hf:Mungert/Veena-GGUF:Q8_0 --speech \
  --speech-decoder hf:cstr/snac-24khz-GGUF:24khz \
  --projector-out ~/.dppan/models/projectors/veena/snac-24khz.gguf

--projector-out names the decoder file, and it must sit in its own directory under projectors/ โ€” that directory is the unit discovery scans.

The 25 MB decoder is duplicated per model, which is the price of each pair being independently removable. This does not arise for qwen3tts, whose decoder ships per-checkpoint.

Some checkpoints are converted rather than downloaded. Qwen publishes the instruction-tuned Qwen3-TTS variants as PyTorch checkpoints, and no GGUF of them exists that this pipeline can load. --speech converts one in place โ€” there is nothing extra to install and no Python involved:

bash
dppan pull hf:Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --speech
Converting Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice: staging 4.21 GB of checkpoint into โ€ฆ
  โ€ฆ eight files, each with the same download bar the rest of `pull` uses โ€ฆ
  talker โ€” 155k-token vocabulary, embedding fold, Q8_0
  folding the text projection    [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘             ] 76673/151936 19s
  quantising tensors             [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘          ] 193/311 2s
  writing                        [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘         ] 202/311 0s
โœ“ talker: โ€ฆ/hf-Qwen-โ€ฆ-CustomVoice-full.gguf (155008 tokens, stop token 154086, 1.72 GB, 43s)
  projector โ€” RVQ codebooks, snake folds, code2wav decoder
  converting tensors             [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘            ] 168/302 0s
โœ“ projector: โ€ฆ/mmproj-โ€ฆ.gguf (302 tensors, 296M parameters, 0.60 GB, 1s)
โœ“ reclaimed 4.21 GB โ€” the checkpoint is not needed again

Every stage is counted, because the slow one takes most of a minute in a single call and a silent command is indistinguishable from a hung one. Redirect the output and the bars disappear (they are drawn only to a terminal) while the phase and โœ“ lines remain, so a logged pull still shows what happened.

What it needs while it runs, both peaking well above what the finished pair occupies:

Disk~8.4 GB freethe 4.21 GB checkpoint and the 2.3 GB pair are present together โ€” the checkpoint is not reclaimed until the conversion has succeeded
Memory~4.8 GB peakthe embedding fold holds the source table, the folded copy and the accumulating Q8_0 output at overlapping times

Both are checked before anything downloads: too little disk is refused outright, too little RAM is a warning, since that figure is one measurement on one checkpoint and refusing on it could block a machine that would have finished.

What lands is an ordinary pulled pair: same directory, same sidecar, same discovery, served on the next start like any other. One field in the sidecar records the difference โ€” source_digest_verified is false, because the bytes on disk are not bytes the publisher serves. The weights it was built from are verified against their published sha256 before any conversion starts, and each half is written under a temporary name and renamed only once complete, so an interrupted conversion leaves nothing that looks like a finished model.

IndicF5 is also a converted checkpoint, and becomes one file in ~/.dppan/models/f5/. Its repository is gated: accept the terms on the model page and set HF_TOKEN before pulling.

bash
dppan pull hf:ai4bharat/IndicF5 --speech

It needs about 2.3 GB free while it converts and takes 0.93 GB once done.

F5 models come in two generations, v0 and v1, with identical weights that compute differently. The wrong one produces fluent, wrong audio instead of an error, so a pull never guesses. IndicF5 is known to this build. Any other F5 repository is refused until you name its generation: --f5-variant v0 for fine-tunes of F5TTS_Base, --f5-variant v1 for fine-tunes of F5TTS_v1_Base. You give it once, at pull time.

The orchestrator discovers speech generators from what has been pulled, so a model fetched either way is served on the next start with no configuration at all. List what is installed against what the catalog knows about:

bash
dppan models --tts
MODEL                      PIPELINE       RATE      SIZE  SOURCES          LOCAL
qwen3-tts-1.7b             qwen3tts    24000Hz      2.3G  hf               ready (hf-ggml-org-Qwen3-TTS-12Hz-1.7B-Base-GGUF-Q8_0-full)
qwen3-tts-1.7b-customvoice qwen3tts    24000Hz      2.5G  hf ยท converts    ready (hf-Qwen-Qwen3-TTS-12Hz-1.7B-CustomVoice-full)
orpheus-3b-en              orpheus     24000Hz      4.0G  hf               ready (hf-PkmX-orpheus-3b-0.1-ft-Q8_0-GGUF-Q8_0-full)
orpheus-3b-hi              orpheus     24000Hz      3.5G  hf               ready (hf-lex-au-Orpheus-3b-Hindi-FT-Q8_0.gguf-Q8_0-full)
veena-3b                   orpheus     24000Hz      4.0G  hf               ready (hf-Mungert-Veena-GGUF-Q8_0-full)
indicf5                    f5          24000Hz      0.9G  hf ๐Ÿ”’ ยท converts  ready (ai4bharat-IndicF5)

converts marks the rows whose repository publishes a PyTorch checkpoint: same pull command, but it converts instead of downloading. ๐Ÿ”’ marks a gated repository โ€” accept its terms on Hugging Face and set HF_TOKEN before pulling.

LOCAL is answered by the same discovery the orchestrator runs at startup, so ready means it will actually be served โ€” not merely that some bytes exist. The name in brackets is the id callers use, derived from the backbone's file name.

To serve a model from outside the models directory, name both halves explicitly. This takes precedence over discovery:

bash
export DPPAN_SPEECH_MODEL=/path/to/backbone.gguf
export DPPAN_SPEECH_MMPROJ=/path/to/mmproj.gguf

Setting only one of the two is ignored with a warning, not half-applied.

Generating โ€‹

bash
curl -X POST http://localhost:8001/v1/generations \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3-tts","modality":"audio",
       "params":{"text":"Hello world","lang":"en"}}'

Poll the job; when state is completed, output_artifacts[0] is a WAV you fetch from /v1/artifacts/:id. For live progress instead of polling, open GET /v1/generations/:id/events โ€” stage and fraction stream as the generator works, and disconnecting does not cancel the job.

If you already have an OpenAI client, point it here instead โ€” POST /v1/audio/speech is a synchronous adapter over the same queue and returns the audio directly. See REST API โ†’ Audio for where compatibility stops.

Request parameters โ€‹

text is always required. Whether a speaker source (voice or speaker_artifact) is also required depends on the model, and it varies within one family: Qwen3-TTS Base refuses to speak without one, Qwen3-TTS CustomVoice cannot take a clip at all and offers nine named timbres instead, and Orpheus has no clip conditioning either. Unknown fields are ignored, not rejected.

Field
textWhat to speak.
langLanguage hint โ€” see below for what a model accepts.
voiceA built-in voice name, or one of the deployment's configured clips.
speaker_artifactID of an uploaded clip to clone the voice from.
reference_textExactly what the speaker_artifact clip says โ€” required by f5 models, refused (400) by every other.
temperature, top_k, top_p, repetition_penaltySampling. Absent means the deployment's default.
instructA sentence of style direction, e.g. "Very happy.". Accepted by every speech generator; only the instruction-tuned checkpoints act on it.
max_duration_secondsCeiling on the finished audio, however long the text is.
speedPace โ€” divides the duration, so 2.0 is twice as fast. Honoured by generators that report a speed_range (IndicF5: 0.3โ€“2.0, 400 outside it); ignored by the rest.

voice and speaker_artifact are two routes to the same input, so sending both is refused; neither silently wins.

Rather than hardcoding which models need what, GET /v1/generations/capabilities answers it per generator โ€” requires_speaker is exactly the flag a client should gate its form on, and requires_reference_text says when the form also needs a transcript box.

An f5 model continues the reference clip's speech with your text, so it needs the clip's exact transcript:

bash
curl -X POST http://localhost:8001/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{"model": "ai4bharat-IndicF5",
       "input": "เคจเคฎเคธเฅเคคเฅ‡! เคธเค‚เค—เฅ€เคค เค•เฅ€ เคคเคฐเคน เคœเฅ€เคตเคจ เคญเฅ€ เค–เฅ‚เคฌเคธเฅ‚เคฐเคค เคนเฅ‹เคคเคพ เคนเฅˆเฅค",
       "speaker_artifact": "66815206-...",
       "reference_text": "exactly what my-voice.wav says"}' \
  -o hindi.wav
  • The clip: 5โ€“10 seconds of clear speech. Keep it under 15 s: a longer clip is cut at a pause but the transcript isn't, and the output then suffers or comes back silent.
  • The transcript: word for word what the clip says, punctuation included. It isn't checked against the audio; a wrong one costs quality.
  • The text: its speaking time is estimated from its length against the transcript's. A word or two is slowed down so it still gets said, but a single character usually comes out silent, so write เค…เฅค rather than เค…. Otherwise speed (0.3โ€“2.0) sets the pace.
  • Errors say what to fix: characters the model has no token for (emoji, Chinese) are named, and a result that comes out silent fails with the likely cause (a 400) instead of returning an empty file.

Ask the deployment what it will honour: GET /v1/generations/capabilities reports each generator's voices, languages, and limits, including the longest utterance its frame budget allows.

Voices โ€‹

There are three ways a request gets a voice, and which are available is a property of the model rather than of the deployment. All three arrive through the same voice or speaker_artifact field, so a client offers one control; the server knows which mechanism a name belongs to.

Built into the checkpoint. Some models carry named speakers โ€” Orpheus's English checkpoint has eight (tara, leah, jess, leo, dan, mia, zac, zoe). Pass one as voice:

bash
curl -X POST http://localhost:8001/v1/generations \
  -H 'Content-Type: application/json' \
  -d '{"model":"hf-PkmX-orpheus-3b-0.1-ft-Q8_0-GGUF-Q8_0-full","modality":"audio",
       "params":{"text":"Hello world","voice":"tara"}}'

These are a property of the weights: no operator can add one and no configuration can remove one. They appear in capabilities as builtin_voices, listed separately from configured clips because the two cannot be merged โ€” only one of them is yours to change. A model with an empty list either has none or has none we can vouch for, and the catalog says which.

Cloned, per request. Upload a clip and pass its ID as speaker_artifact. Needs no configuration and no restart, so any caller can do it:

bash
curl -X POST http://localhost:8001/v1/files \
  -F purpose=user_data -F [email protected]
# {"id":"66815206-...","filename":"my-voice.wav",...}

Named clips, per deployment. Register clips under names an operator chooses. These appear as a voice dropdown in the dashboard and as an enum on the chat tool:

bash
export DPPAN_SPEECH_VOICES="warm=/srv/voices/warm.wav,clear=/srv/voices/clear.wav"

Paths are checked at startup, so a typo appears in the boot log, not as one caller's failed generation. f5 models cannot use them: a registered clip has no transcript, so they take a speaker_artifact and reference_text per request instead. An unknown voice is refused by name and the available ones listed; it does not fall back to the default, because answering in a voice you did not ask for is harder to notice than an error.

The last two need a speaker encoder, which only the qwen3tts codec has. A few seconds of clean single-speaker audio is what that encoder wants; more does not help much. wav, mp3 and flac are accepted, and sample rate and channel count do not matter โ€” both are converted for you. Sending a clip to an orpheus model is refused, not ignored: it has no clip conditioning, so silently dropping the clip would answer in the wrong voice.

Chat needs one of the first or third โ€‹

Chat has no way to attach a reference clip to a call the model generates, so generate_speech is offered there only when the model can speak without one: either it has built-in voices, or the deployment registered named clips.

That makes one combination a dead end, and it is the most obvious one to reach for. A qwen3tts model takes its voice only from a clip, so pulled on its own with no DPPAN_SPEECH_VOICES it is fully usable from /v1/generations and the dashboard's Generate page โ€” which can send a speaker_artifact โ€” while the chat speech control stays greyed out. That is deliberate, not a bug: advertising the tool would mean the model calls it and the call fails every time.

If chat speech is disabled and the Generate page works, this is why. Either register a clip:

bash
export DPPAN_SPEECH_VOICES="narrator=/srv/voices/narrator.wav"

or serve an orpheus model, which needs no clip at all and so is chat-ready with no configuration.

The path is checked when the orchestrator starts, so a wrong one appears in the boot log as ignoring speech voice โ€ฆ, so you are not left with a greyed-out control and no reason for it.

Languages โ€‹

Some models pick a language through a special token, so the languages they support are a property of their vocabulary. DPPAN reads them from the model at startup and reports them in capabilities; the dashboard offers only those, and the chat tool's language argument is constrained to them.

Qwen3-TTS carries ten: zh, en, de, it, pt, es, ja, ko, fr, ru. Anything else fails the job, so read the list rather than assuming it.

An empty languages list is a different statement, and worth reading correctly: it means the model has no language tokens to select between, not that it speaks nothing. Orpheus is like this โ€” each checkpoint is fine-tuned for one language, chosen when you pull it rather than per request, so there is nothing for a caller to set. The dashboard shows no language control for such a model, which is why the dropdown appears for Qwen3-TTS and not for Orpheus.

Long text โ€‹

There is no length limit to work around. Text longer than one generation can hold is split at sentence boundaries and spoken in pieces, joined into a single WAV before the job completes โ€” you submit one request and get one artifact, nothing truncated and nothing to stitch yourself.

Two things are worth knowing. On a long passage the voice may shift slightly where pieces meet, because each is generated independently. And max_duration_seconds caps the finished audio, not each piece, so it is a reliable ceiling on what you get back.

Several models at once โ€‹

Every pulled pair is served on /v1/generations. The chat tool stays a single generate_speech with a model argument rather than one tool per model, so adding a model extends an enum instead of lengthening the tool list. The dashboard shows a picker only when there is more than one to choose between.

The chat enum can be shorter than the served list, and Voices says why: a model that can only be voiced by a reference clip is reachable everywhere except chat until clips are registered for it.

Models are not all held in memory at once. Speech generators share a byte budget, and loading one that does not fit unloads the least recently used until it does; an unloaded model reloads on its next request. So the cost of serving more models than fit is latency on a switch, not failure, and a deployment does not end up holding every model it has ever served. DPPAN_SPEECH_MEMORY_BUDGET_MB sets the budget; unset, it is half of total RAM, floored at one model's worth so a small machine keeps exactly one resident rather than thrashing.

Configuration โ€‹

VariableDefault
DPPAN_SPEECH_MODELunsetBackbone GGUF. Overrides discovery.
DPPAN_SPEECH_MMPROJunsetCodec-decoder GGUF, required alongside it
DPPAN_SPEECH_MODEL_IDqwen3-ttsThe id callers name, for the explicit pair
DPPAN_SPEECH_MAX_FRAMES512Frame ceiling per piece, not per job โ€” see Long text. The frame rate is the codec's, so 512 is ~41 s on qwen3tts and ~44 s on orpheus
DPPAN_SPEECH_MEMORY_BUDGET_MBhalf of RAMBytes all resident speech models may occupy together; below it, least-recently-used models are unloaded
DPPAN_SPEECH_VOICESunsetname=/path/clip.wav, comma-separated
DPPAN_SPEECH_TEMPERATURE0.9Default sampling temperature
DPPAN_SPEECH_TOP_K50<= 0 disables the stage
DPPAN_SPEECH_TOP_P1.0>= 1.0 disables the stage
DPPAN_SPEECH_REPETITION_PENALTY1.05Semantic-token repetition penalty; 1.0 disables it
DPPAN_SPEECH_MAX_ATTEMPTS1Retries for a model that fails to stop โ€” see below
DPPAN_SPEECH_SYNC_TIMEOUT_SECS120How long /v1/audio/speech waits before handing back a job ID to poll
DPPAN_MODELS_DIR~/.dppan/modelsWhere discovery looks

Speech registers only when a model and artifact storage are both available โ€” a generator with nowhere to publish would accept work it could never deliver.

When a model will not stop โ€‹

Some checkpoints terminate unreliably: the model never emits its end token, generation runs to the frame ceiling, and the audio is the utterance on repeat.

DPPAN fails such a job with did_not_terminate rather than publishing the loop โ€” a caller cannot tell a loop from a real answer, and would reasonably conclude the service is broken rather than the model.

The failure is stochastic, so re-running with a fresh seed usually succeeds. That is off by default, because a failed run consumed the entire frame budget where a good one uses a fraction โ€” each retry costs several normal generations, and a model that never terminates would pay that repeatedly to reach the same answer. Raise DPPAN_SPEECH_MAX_ATTEMPTS where a particular checkpoint is worth paying for; replacing it with one that stops reliably is the better fix.

Cancelling a running generation stops it within about one frame (~30 ms) and publishes nothing.


Music โ€‹

Three pipelines sit behind one endpoint. Which one runs is decided by the model you name, and they differ in ways a client should read rather than assume โ€” sample rate, channel count, and whether lyrics are required.

Available music models โ€‹

ModelPipelineOutputSizePartsRegistries

Fetched live from the same catalog dppan models --music reads, which also reports which bundles are already on this machine and, for a partial one, which components it still needs.

outputlyricsseparate tracks
minimax44.1 kHz stereooptionalno
acestep48 kHz stereooptionalno
yue16 kHz monorequiredvocal / instrumental

Read the sample rate and channel count from the response rather than assuming โ€” they are properties of the model, and all three differ. YuE's mono is not a limitation: its reference mixes the two stems into one channel, and that mix is the song.

Getting a music model โ€‹

A music model is not one file โ€” it is four or five, and they belong in one directory. You name the directory; the engine works out which file is which.

dppan models --music lists what exists and what is already on this machine. Here is a complete one, MiniMax-Music3 โ€” five commands into one directory:

bash
BUNDLE=~/.dppan/models/bundles/minimax-music3

# The language model comes from a DIFFERENT repository โ€” only that export
# carries a tokenizer.
dppan pull hf:Serveurperso/MiniMax-Music3-GGUF:Q8_0 --component language_model    --out $BUNDLE

# The other four.  --f16 halves the transformer; see below.
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF:BF16 --component transformer --f16    --out $BUNDLE
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF:Q8_0 --component rvq_depth_decoder    --out $BUNDLE
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF      --component condition_encoder    --out $BUNDLE
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF      --component vocoder              --out $BUNDLE

dppan-orchestrator

About 20 GB as published, or 15 GB with the --f16 above, and the orchestrator serves it on the next start.

Component names are the publisher's, not ours โ€” they are read from the files. Here they use underscores, so --component condition_encoder works and condition-encoder does not. Get one wrong and the error lists every name the repository actually offers, which is the fastest way to find the right one:

no component matches "condition-encoder"; repo has:
  CONDITION_ENCODER, LANGUAGE_MODEL, RVQ_DEPTH_DECODER, TRANSFORMER, VOCODER

Where several quants of one component exist, name the quant on the reference (:Q8_0) as above. A component published in only one precision needs none.

Pass --out, and pass the same one every time. Without it each pull lands in a directory named after its own repository โ€” and MiniMax spans two repositories while YuE spans three, so the components would scatter into places where none of them is a complete model.

Pull it there and nothing else is needed: the orchestrator scans ~/.dppan/models/bundles at startup and serves what it finds. To keep bundles elsewhere, name them instead โ€” comma-separated, each checked independently:

bash
DPPAN_MUSIC_BUNDLE=/vol/a,/vol/b dppan-orchestrator

Startup tells you what it found, and when it skips something, why โ€” an incomplete directory names the components it is missing. A model is refused at startup rather than accepted and failed later, so if it appears in the dashboard it can produce a song.

Two things worth knowing when assembling by hand:

  • YuE needs all three of its components to be recognised at all. Two of them alone are not identifiable as YuE and will be reported as no model.
  • ACE-Step's text encoder is published twice. pull writes a loadable copy beside the original; keep what it puts there.

Smaller and faster: --f16 โ€‹

Some components are published unquantised โ€” MiniMax's transformer is 9.1 GiB, about half its bundle. Converting it to F16 halves the file and roughly halves the time a song takes. Easiest at pull time:

bash
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF:BF16 --component transformer --f16 \
     --out ~/.dppan/models/bundles/minimax-music3

--f16 is opt-in, because it is lossy. Components that are already quantised are skipped, so passing it where it buys nothing costs nothing. A conversion that could not be represented in F16 is refused, not silently turned into garbage.

If a component is already on disk unconverted, delete it and re-pull with --f16:

bash
rm ~/.dppan/models/bundles/minimax-music3/transformer_bf16.gguf
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF:BF16 --component transformer --f16 \
     --out ~/.dppan/models/bundles/minimax-music3

Delete the original. Two files that both look like the same component leave the engine to pick one by sort order, which is not a choice you want made for you.

It is worth doing for more than disk: a smaller model is one that fits on the GPU, and that is the difference between a song taking minutes and taking tens of minutes.

Generating โ€‹

bash
curl http://localhost:8080/v1/audio/music \
  -H 'content-type: application/json' \
  -d '{"model":"minimax-music3","input":"warm lo-fi hip hop, dusty piano, vinyl crackle",
       "lyrics":"[verse]\nlate night on an empty street\n",
       "duration_seconds":30}' --output song.wav
field
modelthe model, named by its directory
inputthe style prompt โ€” genre, instruments, mood, tempo. Detail helps
lyricsoptional for minimax and acestep; required for yue, which needs [verse] / [chorus] headers to structure the song. Leave it out for an instrumental
duration_secondsclamped to the deployment's ceiling rather than refused
stepssolver steps. The default is the count the model was tuned for โ€” lower is faster and smears the words
stemsmix (default), vocal, or instrumental
response_formatwav only

Expect the request to time out, and treat that as normal. A song is minutes of compute, so the wait usually runs out and hands back a job id to poll on /v1/generations/{id}. For anything longer than a short clip, use the job API directly instead of holding a connection open.

Vocals or backing track alone โ€‹

YuE renders vocals and instrumentals separately and sums them, so asking for one costs nothing extra:

bash
curl http://localhost:8080/v1/audio/music \
  -H 'content-type: application/json' \
  -d '{"model":"yue-s1-s2-en-cot","input":"slow blues, male vocal","lyrics":"[verse]\n...","stems":"vocal"}'

minimax and acestep produce one mixed track and refuse anything but mix rather than quietly returning a full mix to someone who asked for vocals. Ask the capabilities endpoint, not the model name: supports_stems says which is which, and the dashboard shows the control only where it applies.

Progress is a phase, not a percentage โ€‹

A song writes, then solves, then decodes, and those cost very different amounts โ€” so a single percentage would sit near the end for most of the wall clock. The job reports which phase it is in:

phases
minimaxwriting โ†’ solving โ†’ decoding
acestepsolving โ†’ decoding
yuewriting โ†’ refining โ†’ decoding

Show the phase name. A fraction is there too and moves, but its total can shrink โ€” the model may stop early, which shortens everything after it.

Length costs compute โ€‹

Cost scales with both length and steps. On an M-series laptop a ten-second clip is a minute or two end to end; a minute of music is a different order of commitment.

That is why the ceiling is low by default โ€” 250 frames, ten seconds. Raise it deliberately, in frames at 25 Hz:

bash
DPPAN_MUSIC_MAX_FRAMES=1500 dppan-orchestrator   # one minute

Configuration โ€‹

VariableDefault
DPPAN_MUSIC_BUNDLEunsetBundle directory, comma-separated for several. Unset disables music generation entirely
DPPAN_MUSIC_MAX_FRAMES250 (10 s)Ceiling on song length, in autoregressive frames at 25 Hz
DPPAN_MUSIC_STEPS0 (pipeline's own)Flow-solver steps โ€” the quality/cost knob. See below
DPPAN_MUSIC_THREADS0 (ask the machine)Threads for the music CPU graphs
DPPAN_MUSIC_SYNC_TIMEOUT_SECS900How long /v1/audio/music waits before handing back a job ID to poll

Steps โ€‹

steps is the flow solver's iteration count, and the main lever on quality against time. 0 means "whatever the pipeline chose", which is not one number:

PipelineDefault steps
minimax30Raising it costs time roughly linearly
acestep8Fixed. The turbo checkpoint was distilled for exactly 8; a different count is refused rather than silently honoured
yueโ€”Runs no flow solver, so it has no step count

A request may name its own steps; this variable only moves the default a request inherits by saying nothing.

Known limitations โ€‹

minimax's solver settings are this engine's best guess. The model's published files do not include them, and a wrong setting produces audio that is entirely plausible and wrong. shift and guidance are exposed as request parameters rather than buried as constants for that reason, and a running deployment reports them as unverified, never as the model's own.

acestep's timbre needs a reference. Its timbre encoder is mandatory and has no text-only path, so with no reference clip supplied a synthetic one is used โ€” which makes the timbre arbitrary rather than absent.


Artifacts โ€‹

An artifact is a blob the orchestrator keeps for longer than one request โ€” an upload you reference later by ID, or bytes a generator produced. The orchestrator has custody of artifact bytes in every deployment mode; decoder nodes never see them.

Artifacts are usable now, independently of generation.

Endpoints โ€‹

MethodPath
POST/v1/artifactsUpload. Content-Type is recorded as the media type. Returns 201.
GET/v1/artifactsList your own artifacts, newest first. ?limit= (default 100, max 1000).
GET/v1/artifacts/:idDownload. Supports Range.
GET/v1/artifacts/:id/metaMetadata only.
DELETE/v1/artifacts/:idDelete bytes and metadata. Returns 204.
bash
# upload
curl -X POST http://localhost:8001/v1/artifacts \
  -H 'Content-Type: audio/wav' \
  --data-binary @speech.wav

# {"id":"1deaf021-...","media_type":"audio/wav","size_bytes":40000,
#  "digest":"5327345f...","state":"ready","created_at":"...","expires_at":null}

# download, resume-friendly
curl -H 'Range: bytes=1000-1099' \
  http://localhost:8001/v1/artifacts/1deaf021-... -o part.bin

Bodies stream in both directions โ€” memory use does not scale with artifact size, so a multi-gigabyte object is accepted or rejected without buffering.

state is one of pending, ready, failed. Only ready artifacts are readable; a partial or failed upload never becomes readable at all.

digest is the SHA-256 of the full object, computed while streaming.

Ranges are single-range only. Multi-range (bytes=0-9,20-29) and suffix (bytes=-500) requests are refused rather than partially honoured.

Who owns an artifact โ€‹

Custody and ownership are different questions. The orchestrator holds the bytes either way; ownership decides who can read or delete them.

  • Self-hosted (no database): there are no user accounts, so there is one principal โ€” the operator. Every artifact belongs to it.
  • Hosted (database configured): the owner is the logged-in account, resolved from the session cookie exactly as /auth/me is. An unauthenticated request is rejected with 401; it never falls back to a shared principal.

A missing artifact and one belonging to somebody else both return 404, so a guessed ID cannot confirm that another owner's artifact exists.

Where the bytes live โ€‹

Bytes are stored on the local filesystem. Every backend addresses objects through the same key:

<prefix>/<aa>/<bb>/<artifact-id>

aa/bb are the first two hex byte-pairs of the ID, which keeps any single prefix small. The path under the artifact root is the object key, so the root and an S3 bucket hold the same namespace โ€” a deployment can be copied from one to the other with no re-keying:

bash
aws s3 sync ~/.dppan/artifacts/ s3://your-bucket/

S3-compatible storage is not implemented yet. The configuration is in place and the store is chosen at runtime, so it will be a drop-in when it lands. Setting DPPAN_ARTIFACT_S3_BUCKET today refuses to enable artifact storage rather than silently writing to local disk.

Metadata (owner, size, digest, state, expiry) is stored separately from bytes: embedded SQLite under the artifact root when no database is configured, PostgreSQL when one is. That follows the same signal as the rest of the orchestrator, so artifact ownership lands in the same store as the accounts it refers to.

Configuration โ€‹

VariableDefault
DPPAN_ARTIFACT_ROOT~/.dppan/artifactsWhere bytes and (self-hosted) metadata live
DPPAN_ARTIFACT_PREFIXartifactsObject-key prefix; lets one bucket host several deployments
DPPAN_ARTIFACT_MAX_BYTES0 (unlimited)Per-owner byte ceiling
DPPAN_ARTIFACT_MAX_OBJECTS0 (unlimited)Per-owner object ceiling
DPPAN_ARTIFACT_TTL_HOURS0 (keep forever)Default lifetime for new artifacts
DPPAN_ARTIFACT_S3_BUCKETunsetSelects the S3 backend โ€” not implemented yet
DPPAN_ARTIFACT_S3_ENDPOINTunsetFor R2, MinIO, and other S3-compatible services
DPPAN_ARTIFACT_S3_REGIONunsetCredentials come from the standard AWS environment

A single upload is capped at 512 MiB regardless of quota.

Hosted deployments need the schema applied out-of-band โ€” the orchestrator connects, it never migrates. Use one mechanism consistently:

bash
# either: sqlx, which records what it applied
sqlx migrate run --source rust/migrations

# or: the consolidated snapshot, for a database sqlx will never touch
psql "$DATABASE_URL" -v ON_ERROR_STOP=1 -f rust/db/schema.sql

Do not mix them. Applying a single migration with psql -f leaves _sqlx_migrations un-stamped, so a later sqlx migrate run re-applies it and fails on CREATE TYPE โ€” the migrations are deliberately not idempotent, so a double application is loud rather than silent.

If the artifact root cannot be prepared, artifact endpoints return 503 and the orchestrator still starts and serves inference โ€” nothing on the inference path depends on artifacts.

Current limits โ€‹

  • No usage metrics are exported yet.
  • The S3 backend is not built.

Generation jobs โ€‹

Producing media is asynchronous: you submit a job, poll it, and download the artifacts it produced. Speech works today โ€” see Speech for the configuration. Image and video have no generator, so requesting those modalities answers 422.

MethodPath
POST/v1/generationsSubmit. Returns 202 Accepted with a queued job.
GET/v1/generationsList your own jobs, newest first.
GET/v1/generations/:idPoll.
DELETE/v1/generations/:idCancel; returns the state it settled in.
bash
curl -X POST http://localhost:8001/v1/generations \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3-tts","modality":"audio","params":{"text":"Hello world"}}'

state is queued, running, completed, failed, cancelled, or expired. Only completed populates output_artifacts, which are ordinary artifact IDs โ€” download them through /v1/artifacts/:id.

Two refusals are worth telling apart. 429 means the queue is full: retry. 422 means no generator in this deployment produces that modality with that model: retrying will never help.

Surviving a restart โ€‹

An orchestrator records its own identity on every job it claims, and matches on it at startup to resolve the ones it left running when it died.

VariableDefault
DPPAN_WORKER_IDderived and persisted under DPPAN_ARTIFACT_ROOTIdentity this instance claims jobs under

The default is right for almost every deployment. Set it when the artifact root is not durable โ€” a container with no mounted volume โ€” because an identity that changes on restart never matches the jobs the previous run left behind, and those jobs stay running forever. It must be stable for this instance and distinct from any peer sharing the same database.

Cancellation is cooperative and reaches a running job โ€” generators stop between frames rather than being killed, so nothing is left half-written. A job that outlives its deadline is stopped even if its generator ignores the request.

VariableDefault
DPPAN_JOB_MAX_QUEUED64Waiting jobs allowed; running jobs are not counted
DPPAN_JOB_DEADLINE_MINUTES10Default deadline; 0 for none. A request may ask for less, never more than 24 h
DPPAN_CLEANUP_INTERVAL_SECS300Sweeps expired artifacts and jobs; 0 disables both

The job schema is applied out-of-band with the same mechanism as artifacts โ€” see Where the bytes live above; sqlx migrate run or rust/db/schema.sql covers both tables.

Free to run ยท Proprietary