Models
dppan runs GGUF models. The table below is the curated catalog — models validated end-to-end on the network. It is not a registry mirror: any GGUF from Ollama or Hugging Face can still be tried (dppan pull hf:<owner>/<repo>), the catalog just marks what we vouch for. Check it from the CLI anytime:
dppan models # table with local availability
dppan models --json # for scripting
dppan models --tts # speech models, catalogued separately
dppan models --music # music models, catalogued separatelyLooking for speech or music?
They are catalogued apart from the chat models below, because they are assembled from several files rather than sharded by layer. See Media: processing & generation for speech and Music for songs.
Supported models
| Model | Arch | Layers | Size | Registries | Manifest |
|---|
The table is fetched live from dl.inclavate.io/models/catalog.json — the same catalog the CLI uses (dppan models caches it for 24 hours and works offline from a built-in copy). Models larger than one machine use sharded download.
Speech models are catalogued separately — see Media → Speech. They are not listed here because a speech model is a pair of files that always run together and never shards, so the layer count this table is organised around does not apply to one.
Music models are not here either, and for a stronger version of the same reason — see Media → Music. A music model is not one file or even a pair: it is four or five heterogeneous GGUFs pulled as a bundle into one directory, and which file plays which role is read from the files rather than declared. There is no single layer count to organise such a thing around, and no single quant either: quant is a per-component property, which is why dppan pull --component records one per file.
Coverage & roadmap
The engine supports 45 architecture families — enough to run 132 of the top-200 trending text-generation models on Hugging Face (audited 2026-09-02 by reading each model's GGUF architecture, not its name). Where the rest stands:
| Tier | Covers | Status |
|---|---|---|
| Dense transformers | Llama 3.x · Qwen 2.5/3 · Gemma 2/3/4 · Phi-3/4 · GLM-4 · SmolLM · Muse Glimmer 30B — plus their distills and finetunes | ✅ Supported |
| Dense classics | GPT-2 · Pythia · BLOOM · Falcon · Phi-2 · Nemotron · Command R7B | ✅ Supported |
| Gemma 4 E-series | E2B · E4B — image and audio input; splits across at most two nodes | ✅ Supported |
| Mixture-of-Experts | Qwen3-MoE (A3B/235B) · gpt-oss · Gemma 4 A4B · OLMoE · Qwen1.5-MoE · GLM-4.5-Air and 355B-A32B · Mixtral 8x7B/8x22B · Llama 4 Scout & Maverick | ✅ Supported |
| MLA attention | DeepSeek-V2 · V2-Lite · Coder-V2-Lite · GigaChat3 | ✅ Supported |
| State-space & hybrid | Mamba (1 & 2) · FalconMamba · Codestral Mamba · Granite 4.0-H · Falcon-H1 (0.5B–34B) | ✅ Supported |
| Hybrid — single mixer per layer | Nemotron Nano 2 · Nemotron-H 47B Reasoning · Nemotron-3 MoE (30B-A3B – 550B-A55B) · LFM2 (350M–2.6B) · LFM2.5-2.6B · LFM2-8B-A1B · Jamba (900M · Mini 1.7 · Large 1.7) | ✅ Supported |
| Granite dense & MoE | Granite 3.x/4.x dense (2B/8B) · Granite 4.2 (3B/8B/30B) · Granite-MoE a400m/a800m | ✅ Supported |
| Gated delta-net | Qwen3.5 (0.8B–27B) · Qwen3.6 · Qwen3.8 (27B + 2B/4B/9B distills) · Ornith 1.5 (9B · 35B-A3B) · Qwen3.5-MoE (A3B/A10B/A17B) · Qwen3-Next-80B · Qwen3-Coder-Next · Ling 3.0 | ✅ Supported |
| MLA — frontier | DeepSeek-V3 · R1 · V3.2 · Kimi K2 · Mistral-Large-3 · GLM-5/5.2 — 671B–1T | 📋 In Validation |
| Multimodal input | Image, audio and video frames on any model with a matching projector — Gemma 4 · Gemma 3 · Qwen2-VL · Qwen2.5-VL · Qwen3-VL · GLM-4.6V-Flash · DeepSeek-OCR · Muse Glimmer · Ultravox · Qwen3-ASR. Input only — the model answers in text | ✅ Supported |
| Speech to text | OpenAI-compatible /v1/audio/transcriptions, plus a hands-free voice conversation in the dashboard. A chat model with an audio projector hears the clip directly; a text-only one gets a transcriber in front of it. See Media | ✅ Supported |
| Document input | PDF and text/Markdown files, optional OCR, native server-side video decode | 📋 Planned |
| Media generation | Speech and music ship. Speech via Qwen3-TTS (voice cloning, or nine named voices) and Orpheus/SNAC; music via MiniMax-Music3, ACE-Step 1.5 and YuE. See Media and the job queue. Image and video generation are not built | 🚧 In Progress |
No dates are promised — items ship when they pass the same validation gate as everything in the catalog. Missing a model you need? Write to [email protected] — requests directly shape this list.
Model sources
dppan pull and dppan download resolve models from three places:
| Source | Ref format | Auth |
|---|---|---|
| Ollama registry | llama3.3:70b | none |
| Hugging Face (GGUF repos) | hf:bartowski/Llama-3.3-70B-Instruct-GGUF:Q6_K | none for public repos; your own HF_TOKEN for gated ones |
| Inclavate-hosted | inclavate:<slug>[:<quant>] | none |
- Hugging Face: only repos that ship GGUF files (bartowski, unsloth, lmstudio-community, …). The quant selector defaults to
Q4_K_M, but any GGUF quantisation runs — pick another withhf:<owner>/<repo>:<QUANT>(e.g.:Q6_K,:Q8_0,:IQ4_XS); the engine keeps tensors in their native type and the right kernels are dispatched automatically. The quant shown in the catalog is simply the configuration we validated. Split quants (-00001-of-0000N.gguf) are merged automatically. For gated repos, accept the license on huggingface.co and setHF_TOKEN(or log in withhuggingface-cli) — tokens are always your own. - Inclavate-hosted is for models with no community GGUF: converted and quantized once, centrally, then served from
dl.inclavate.io— so nodes never have to quantize anything themselves.
Models published as a set of files
Most models are one GGUF. A few are not: ACE-Step 1.5 ships a diffusion transformer, a 5 Hz language model, a text encoder and an autoencoder in a single repository, and MiniMax-Music3 ships five components. Every file carries the same quant suffixes, so a plain dppan pull hf:…:Q8_0 cannot tell them apart — it refuses with the list of candidates rather than guessing at which of four models you meant.
--component names the one you want. ROLE is what the file is in your pipeline and is what gets recorded; SELECTOR finds it in the repository and defaults to ROLE, so you only write both when the publisher does not name its files after their job:
# ACE-Step names its files after the checkpoint, so role and selector differ
dppan pull hf:Serveurperso/ACE-Step-1.5-GGUF:Q8_0 --component dit=acestep-v15-turbo
dppan pull hf:Serveurperso/ACE-Step-1.5-GGUF --component vae
# MiniMax names them after the component, so the role alone is enough
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF:Q4_0 --component transformer
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF --component vocoderOne component per command. Each lands in ~/.dppan/models/bundles/<repo>/ alongside a bundle.dppan.json, which every pull extends — so a four-file model is assembled over four commands and the directory always says what it holds:
✓ component vocoder written: …/bundles/hf-audio-cpp-MiniMax-Music3-GGUF/vocoder.gguf
bundle now holds 4: condition_encoder, rvq_depth_decoder (Q8_0), transformer (Q4_0), vocoderQuant is per component, not per model — and it is read from the file that was selected, not from the one you asked for. That is not a nicety. These repositories genuinely mix precisions: MiniMax publishes its transformer at Q4_0 and its depth decoder at Q8_0, and ACE-Step publishes its autoencoder in BF16 only. Asking for :Q8_0 in the ACE-Step repo therefore gets you a Q8_0 diffusion transformer and a BF16 autoencoder, and the manifest records each as what it actually is rather than as what was requested. A component published in a single precision ignores the quant selector, because there is nothing to choose between — which is also why vocoder above needs no quant at all.
Some publishers export GGUFs that name their architecture in a way llama.cpp does not read. Where that is true and the fix is metadata rather than weights, the pull repairs the header on the way in and records that it did — the tensor payload is copied through byte for byte, so nothing is requantised.
--component works on Hugging Face repositories, requires the source to publish a sha256, and cannot be combined with --layers, --orch, --projector-only or --speech: a component is a whole file, not a layer slice. Point --out somewhere else to keep a bundle outside the model store.
Music models are assembled exactly this way, and they are served — see Music for a complete, copy-pasteable set of commands for one of them.
Integrity
A GGUF has no per-tensor checksums, so a partially-downloaded shard can't be checked against the source's whole-file hash. dppan's answer is a per-tensor manifest:
dppan manifest llama3.3:70b -o 70b.manifest.json # once, on a machine with the full file
dppan pull llama3.3:70b --layers 27-53 --verify --manifest 70b.manifest.jsonWith --verify, every downloaded tensor is hashed and compared; any mismatch names the tensor and deletes the shard. Without a manifest, shards are still structurally validated and fetched over TLS — fine for machines you control. Manifests published at dl.inclavate.io/manifests/<sha256>.json are picked up automatically, no --manifest flag needed.