Vision & video
Models that see. Send an image — or a whole video — alongside your text and ask questions about it. Frames are sampled natively by the model server, so a twenty-minute clip is a single request, not a pile of screenshots.
Qwen 3.8 · 27B● live now
lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL — Qwen3.8-27B dense hybrid-GDN, UNCENSORED via a Heretic fork (Arbitrary-Rank Ablation, layers 26-56 of 64; vendor-reported KL 0.0535 vs ba
visiontoolsreasoning128K context
c1-qwen38-27b-heretic-ara-nvfp4-vllmGemma4 26B A4Bavailable
Gemma-4-26B-A4B UNCENSORED MoE (4B active, AEON-7 NVFP4 compressed-tensors, 15.29 GiB), Gemma4ForConditionalGeneration arch, vLLM v0.21.0 TP=2 on 2x 5060 Ti (sm_120). Multimodal (v
vision
Qwen36 27B Hereticavailable
GENUINE Qwen3.6-27B flagship DENSE (base Qwen/Qwen3.6-27B, NOT Qwopus), HERETIC abliterated/UNCENSORED, PrismaQuant mixed-precision FP8+NVFP4 (compressed-tensors, ~5.25-bit, rdtand
visiontoolsreasoning
Qwen36 27B Prismascoutavailable
GENUINE Qwen3.6-27B flagship DENSE (base Qwen/Qwen3.6-27B, NOT Qwopus), PrismaSCOUT mixed-precision NVFP4 (compressed-tensors, ~5.31 bpp, 20.17 GiB, rdtand), vLLM v0.21.0 TP=2 on 2
visiontoolsreasoning
Qwen36 35B A3Bavailable
Qwen3.6-35B-A3B flagship MoE (35B/3B active, hybrid GDN, multimodal), AWQ-4bit (cyankiwi, 24.3 GiB), vLLM v0.21.0 TP=2 on 2x 5060 Ti (sm_120). SAME model as the NVFP4 winner — stag
visiontoolsreasoning
Qwen36 35B A3Bavailable
GENUINE Qwen3.6-35B-A3B flagship MoE (35B/~3B active, hybrid GDN), NVFP4 compressed-tensors (unsloth, Blackwell-native), vLLM v0.21.0 TP=2 on 2x 5060 Ti (sm_120). Cleaner-quant A/B
visiontoolsreasoning
Qwen36 35B A3Bavailable
Qwen3.6-35B-A3B UNCENSORED MoE (HauhauCS-Aggressive abliteration), NVFP4 compressed-tensors (lyf re-quant, 21.75 GiB), vLLM v0.21.0 TP=2 on 2x 5060 Ti (sm_120). Same arch/quant/siz
visiontoolsreasoning
Qwen38 27Bavailable
unsloth/Qwen3.8-27B-NVFP4 — dense 27B hybrid-GDN (48 DeltaNet + 16 full attn), 262K native, MTP head. Served on c3 Ada via vLLM Marlin, TP=4, NO DCP (4 KV heads land 1/GPU), fp8 KV
visiontoolsreasoning
Qwopus 36 27Bavailable
Qwopus3.6-27B-v2 DENSE (RL-tuned), NVFP4 (modelopt/TensorRT-MO, ~18 GiB, Blackwell-native FP4), vLLM v0.21.0 TP=2 on 2x 5060 Ti (sm_120). Native-FP4 speed lane for the dense qualit
visiontoolsreasoning
Qwopus 36 35Bavailable
Qwopus3.6-35B-A3B-v1 MoE (35B/~3B active, hybrid GDN), NVFP4 compressed-tensors (cyburn, ~4.75-bit, Blackwell-native), vLLM v0.21.0 TP=2 on 2x 5060 Ti (sm_120). Fast-lane sibling t
visiontoolsreasoning
Chat & reasoning
General-purpose language models. Most are reasoning models: they think through a problem in a hidden scratchpad before answering, which is why you should always allow at least 2048 max_tokens or you will get an empty reply. Many support tool-calling, so agents and editors can drive them directly.
Muse Glimmer · 30B● live now
Muse-Glimmer-30B (Meta muse_glimmer / OpenAI-Harmony arch, dense ~29.6B) NVFP4 modelopt_fp4, vLLM (authors' muse-glimmer image) on c3 Ada TP=2, 131K ctx (full native, util 0.95, ~1
128K context
c3-muse-glimmer-30b-nvfp4-vllmQwen 3.6 · 35B● live now
lyf/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4 — the SAME weights hermes runs on c1, served on 2 of c3's 4 Ada cards via vLLM Marlin. Stood up 2026-07-26 as a SECOND work
toolsreasoning256K context
c3-qwen36-35b-a3b-uncensored-nvfp4-tp2-vllmGlm 45 Airavailable
GLM-4.5-Air REAP 82B-A12B MoE (12B active), compressed-tensors AWQ-int4, vLLM TP=4 on c3 Ada (glm47-patched2). 131K ctx, fp8 KV, KV pool capped at 131K (num-gpu-blocks-override) to
tools
Glm 47 Flashavailable
zai-org GLM-4.7-Flash 30B-A3B MoE (glm4_moe_lite, MLA, text), AWQ-4bit (cyankiwi), vLLM TP=2 on 2x RTX 4060 Ti (Ada sm_89) on catamaran-3. GLM's MLA kernel needs datacenter shared
tools
Glm 47 Flashavailable
30B-A3B MoE (3B active); HauhauCS-Aggressive uncensored finetune of GLM-4.7-Flash; 202K ctx; mutex with qwen35-27b and Heretic on c2
Glm 47 Flashavailable
HauhauCS abliterated GLM-4.7-Flash 30B-A3B MoE Q4_K_M, native llama-server (llama.cpp b9464) on c3 Ada TP=2 row-split. MLA-native => full 202K ctx (vs vLLM ~20K MLA-shmem wall), un
Gemma4 26B A4Bavailable
Gemma-4-26B-A4B UNCENSORED MoE (4B active, AEON-7 NVFP4 compressed-tensors), SGLang v0.5.12 TP=2 on catamaran-2's 2x RTX 5060 Ti (sm_120). Native-FP4 speed attempt vs the vLLM emul
Gemma4 26B A4Bavailable
Gemma-4-26B-A4B UNCENSORED MoE (4B active, AEON-7 NVFP4 compressed-tensors), Gemma4ForConditionalGeneration, vLLM TP=2 on catamaran-2's 2x RTX 5060 Ti (sm_120). 256K context (max_p
Gemma4 26B A4Bavailable
HauhauCS-Balanced uncensored Gemma-4-26B-A4B MoE (4B active, GQA, 256K ctx), Q5_K_M GGUF (public) on llama.cpp (digest-pinned b9464), c2 Blackwell TP=2 row-split, q8_0 KV on GPU. F
Gemma4 31B Uncensoredavailable
Gemma-4-31B DENSE (30.7B, Gemma4ForConditionalGeneration, huihui abliterated-v2 UNCENSORED), MXFP4-pack compressed-tensors 4-bit, vLLM TP=2 on catamaran-2's 2x RTX 5060 Ti (sm_120,
tools
Gemma4 31B Hauhauavailable
HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP — dense 30.7B gemma4, QAT-native Q4_K_M GGUF, abliterated/uncensored, agentic/reasoning. 192K context (freeze-capped on 2x1
Gemma4 26B Uncensoredavailable
HauhauCS abliterated Gemma4-26B-A4B (Balanced) MoE, Q5_K_M GGUF on llamafile 0.10.2 (stock llama.cpp:server-cuda image, dropped-in binary), c3 Ada TP=2 row-split. Uncensored, visio
Lfm25 8B A1Bavailable
LiquidAI LFM2.5-8B-A1B (8.3B MoE / 1.5B active, 18 conv + 6 GQA layers) as cyankiwi AWQ-INT4, 5 GiB. Agentic/tool-use/instruction specialist, very fast, 128K ctx, text-only, NOT un
Laguna Xs 21available
poolside/Laguna-XS-2.1-NVFP4 — 33B-total/3B-active coding MoE (256+1 experts, 10 global + 30 SWA/512 attn), 262K ctx. NVFP4 (compressed-tensors) on c3 Ada via Marlin. One of 2× TP=
tools
Laguna Xs 21available
poolside/Laguna-XS-2.1-NVFP4 — 33B-total/3B-active coding MoE (256+1 experts, 10 global + 30 SWA/512 attn), 262K ctx. NVFP4 (compressed-tensors) on c3 Ada via Marlin. One of 2× TP=
toolsreasoning
Minicpm V45 Tp2available
Vision-language model, GPT-4o level, on c1 GPUs 0+1
Negentropy Claude Opusavailable
Claude Opus 4.7 distill on Qwen3.5-9B; SGLang v0.5.12 TP=2 across 2x 5060 Ti; BF16 weights + runtime --quantization fp8; 262K ctx; native multi-slot continuous batching (--max-runn
Qwable 36 35Bavailable
MoE 35B (~3B active), Qwen3.6-35B-A3B-class reasoning fine-tune, 262K ctx (c2 sibling, mutex with c2 27/35B-class)
Agentworld 35B A3Bavailable
Frosty40/Qwen-AgentWorld-35B-A3B-NVFP4 — world model (not executor): predicts the next environment state across 7 agent domains (MCP/Search/Terminal/SWE/Android/Web/OS). Orchestrat
Qwen3 Next 80Bavailable
Qwen3-Next-80B-A3B-Instruct hybrid (Gated-DeltaNet+attn, 3B active, 512 experts/10 active), AWQ-4bit, SGLang TP=4 on c3 Ada. NEXTN/MTP spec-decode + tiny KV -> full 262K ctx. STAGE
Qwen3 Vl 32Bavailable
AWQ 4-bit VLM for hooper_booper vision pipeline
Qwen35 08Bavailable
Smallest 0.8B model on RTX 3050
Qwen35 27Bavailable
Dense 27B, 262K context, row-split on c2
Qwen35 2Bavailable
Tiny 2B model on RTX 3050
Qwen35 2B Bf16available
BF16 2B model on RTX 3050
Qwen35 35B A3Bavailable
MoE 35B via vLLM AWQ, 32K context
Qwen35 4Bavailable
Small 4B model on RTX 3050
Qwen35 9Bavailable
AWQ 4-bit dense 9B, 32K context
Qwen35 9B Bavailable
Second instance for parallel inference
Qwen35 9B Uncensoredavailable
9B sibling of the live c1 Qwen3.6-35B-A3B UNCENSORED (same HauhauCS-Aggressive lineage; no Qwen3.6-9B base exists upstream). Dense 9B hybrid GDN, multimodal. GitMylo BF16 safetenso
Qwen36 27B Hereticavailable
DavidAU NEO-CODE Di-IMatrix Q4_K_M; uncensored Qwen3.6-27B finetune; 262K ctx; mutex with qwen35-27b on c2
Qwen36 27B Hauhauavailable
Second c3 instance of HauhauCS abliterated Qwen3.6-27B dense Q4_K_P, native llama-server (llama.cpp b9464) on c3 Ada TP=2 layer-split. Identical to #1 — same model, same args. ~17G
Qwen36 27B Hauhauavailable
HauhauCS abliterated Qwen3.6-27B dense Q4_K_P, native llama-server (llama.cpp b9464) on c3 Ada TP=2 layer-split. Hybrid GDN (48 linear + 16 full attn layers, only 16/64 carry KV),
Qwen36 27B Uncensoredavailable
HauhauCS-Aggressive uncensored Qwen3.6-27B; Q5_K_P custom imatrix; 0/465 refusals; 262K ctx; mutex with c2 27B-class siblings
Qwen36 35B A3Bavailable
GENUINE Qwen3.6-35B-A3B flagship MoE (35B/~3B active, hybrid GDN), NVFP4 compressed-tensors (unsloth). SGLang v0.5.12-cu129 TP=2 on 2x 5060 Ti (sm_120). The ONLY candidate that cle
Qwen36 35B A3Bavailable
lyf/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4 — the SAME weights hermes runs on c1 (trusted), served on c3 Ada via vLLM Marlin + DCP=2 (#40996 backport). One TP=4 instan
toolsreasoning
Qwopus 27B V35available
Reasoning fine-tune of Qwen3.5-27B, 262K context (c2 sibling, mutex with c2 27B-class)
Qwopus 36 27Bavailable
3.6 reasoning fine-tune of Qwen3.6-27B, 262K context (c2, mutex with c2 27B-class)
Qwopus 36 27Bavailable
Qwopus3.6-27B-v2 (Jackrong reasoning SFT on Qwen3.6-27B), GPTQ-Pro 4-bit (XReyRobert), vLLM v0.21.0 gptq_marlin TP=2 across 2x 5060 Ti (sm_120), 200K ctx (CUDA graphs, ~25 tok/s; 2
toolsreasoning
Qwopus 36 35Bavailable
MoE 35B (3B active), Qwen3.6 base, Jackrong reasoning fine-tune, 262K ctx (c2 sibling, mutex with c2 27/35B-class)
Qwythos 9Bavailable
9B reasoning LLM, 1M context, full-param fine-tune on Qwen3.5-9B. Claude Mythos + Fable traces. Thinking mode, function calling, uncensored.
Qwythos 9Bavailable
Qwen3.5-based 9B, 131K context, reasoning, c1
Supergemma4 26B Modeloptavailable
AEON-7 SuperGemma4-26B abliterated/uncensored gemma4 MoE (3.8B active), NVIDIA ModelOpt NVFP4, vLLM TP=2 on c2 2x RTX 5060 Ti (sm_120) with FP8-Marlin MoE — the fast+correct path t
Thinkingcap 27B Abliteratedavailable
huihui-ai abliterated fork of ThinkingCap-Qwen3.6-27B (UNCENSORED), sakamakismile NVFP4 W4A4 on c2 Blackwell native FP4 (NEVER on Ada — quantized GDN in_proj gates hit the Marlin w
toolsreasoning
Thinkingcap 27Bavailable
Same cyankiwi AWQ-INT4 artifact as the c3 TP=4 sibling, on c2's 2x RTX 5060 Ti (Marlin on sm_120). Desktop-margin-bound: ~32K ctx initial (ladder at first boot; the 262K deployment
toolsreasoning
Thinkingcap 27Bavailable
bottlecapai ThinkingCap (LoRA reasoning fine-tune of dense Qwen3.6-27B hybrid, -46% thinking tokens at ~base quality; CENSORED by design), cyankiwi AWQ-INT4 W4A16 Marlin, vLLM TP=2
toolsreasoning
Thinkingcap 27B Fp8available
bottlecapai/ThinkingCap-Qwen3.6-27B-FP8 — Qwen3.6-27B RL-finetuned for ~50% fewer thinking tokens, native block-FP8 (e4m3). Dense hybrid-GDN (64 layers, 16 full-attn), 262K ctx. TP
toolsreasoning
Image & video generation
Text-to-image and text-to-video generation, plus a synchronized audio track. These run as a job queue rather than a chat: you submit a prompt, poll for completion, and fetch the finished media.
Wan Artifacts● live now
S3 artifact store backing wan-api. Public clients re-request the /view 302 Location path+query via /proxy/wan-artifacts/... (presigned signature survives the hop; the proxy connect
wan-artifactsWan Api● live now
REST orchestrator for the wan2gp studio workers (text-to-image/video). POST /jobs {runtime:wan2gp, instance:c3-wan2gp-studio-a, model:flux-1-schnell-t2i, prompt, params:{kind:image
wan-apiWan Executor● live now
THE media submission endpoint. POST /create {prompt, kind:image|video|auto, resolution?, seconds?, quality:fast|quality} -> 202 {job_id, poll, view}. It creates+dispatches the job
wan-executorWanfleet Studio Aavailable
WAN-FLEET multi-modal studio (distinct from the public c3-wan2gp-studio-a/b workers): T2I (Flux/Qwen-Image), T2V/I2V (Wan 2.1/2.2, Hunyuan, LTX-2), MMAudio/TTS. Non-root, ephemeral
Wanfleet Studio Bavailable
WAN-FLEET multi-modal studio (distinct from the public c3-wan2gp-studio-b/b workers): T2I (Flux/Qwen-Image), T2V/I2V (Wan 2.1/2.2, Hunyuan, LTX-2), MMAudio/TTS. Non-root, ephemeral
Wanfleet Studio Cavailable
WAN-FLEET multi-modal studio (distinct from the public c3-wan2gp-studio-c/b workers): T2I (Flux/Qwen-Image), T2V/I2V (Wan 2.1/2.2, Hunyuan, LTX-2), MMAudio/TTS. Non-root, ephemeral
Wanfleet Studio Davailable
WAN-FLEET multi-modal studio (distinct from the public c3-wan2gp-studio-d/b workers): T2I (Flux/Qwen-Image), T2V/I2V (Wan 2.1/2.2, Hunyuan, LTX-2), MMAudio/TTS. Non-root, ephemeral
Wan2Gp Studio 5060available
Wan2.2 TI2V-5B FastWan worker (720p + MMAudio), profile 3 + sage2 — parity with the c3 studios. Ephemeral, FIXED to c1 GPU 0 (5060 Ti sm_120); SHARES that GPU with the planned embe
Wan2Gp Studio Aavailable
Full multi-modal studio (worker A of 2): text-to-image (Flux/Qwen-Image/Z-Image/HiDream), text+image-to-video (Wan 2.1/2.2, Hunyuan, LTX-2), video-to-audio + TTS (MMAudio, Qwen3-TT
Wan2Gp Studio Aavailable
Wan fleet multi-modal studio on 1x RTX 5060 Ti (Blackwell sm_120): T2I (Flux/Qwen-Image), T2V/I2V (Wan 2.1/2.2, Hunyuan, LTX-2), MMAudio/TTS. Model chosen per request. Ephemeral DR
Wan2Gp Studio Bavailable
Wan2.2 TI2V-5B FastWan video worker — native 720p (1280x704, 5s@24fps, ~180-250s) + MMAudio soundtrack. profile 3 + sage2, model SSD-resident. IDENTICAL to studio-a (parity twin, i
Wan2Gp Studio Bavailable
Wan fleet multi-modal studio on 1x RTX 5060 Ti (Blackwell sm_120): T2I (Flux/Qwen-Image), T2V/I2V (Wan 2.1/2.2, Hunyuan, LTX-2), MMAudio/TTS. Model chosen per request. Ephemeral DR
Wan2Gp T2V 1available
v2.0-studio on the 3050 (8GB, sm_86), gen-clip driver. Lean: Wan2.1 t2v 1.3B + small models (bigger models need a 16GB worker). Ephemeral; c1 host-RAM is tight.
Quantum simulation
State-vector quantum circuit simulators running on GPU. Not language models at all — they execute quantum circuits, which is a genuinely different kind of workload that happens to want the same hardware.
Qsim Pennylaneavailable
Qubit-capable circuit execution endpoint, RTX 3050 8 GB Ampere — PennyLane + cuQuantum, ceiling ~28 qubits
Qsim Pennylane 5060available
Qubit-capable circuit execution endpoint, RTX 5060 Ti 16 GB Blackwell sm_120 — PennyLane + cuQuantum, ceiling ~29 qubits. Sibling of c1-qsim-pennylane (3050) + c3-qsim-pennylane (4
Qsim Pennylane 5060available
Qubit-capable circuit execution endpoint, RTX 5060 Ti 16 GB Blackwell sm_120. BLITZ-ready: replicas=0 by default; activated when c1 GPU 1 (normally sglang/negentropy) is freed for
Qsim Pennylane 5060available
Qubit-capable circuit execution endpoint, RTX 5060 Ti 16 GB Blackwell sm_120. BLITZ-ready: replicas=0 by default; activated when c1 GPU 2 (normally sglang/negentropy TP=2 partner)
Qsim Pennylane 5060available
Qubit-capable circuit execution endpoint, RTX 5060 Ti 16 GB Blackwell sm_120. BLITZ-ready: replicas=0 by default; c2 GPU 0 normally hosts a TP=2 llama.cpp model. Safe for HB's <14-
Qsim Pennylane 5060available
Qubit-capable circuit execution endpoint, RTX 5060 Ti 16 GB Blackwell sm_120. BLITZ-ready: replicas=0 by default; c2 GPU 1 is the second half of the c2 TP=2 LLM. Desktop-freeze cav
Qsim Pennylane 4060available
Qubit-capable circuit execution endpoint, RTX 4060 Ti 16 GB Ada sm_89 — PennyLane + cuQuantum, ceiling ~29 qubits
Qsim Pennylane 4060available
Qubit-capable circuit execution endpoint, RTX 4060 Ti 16 GB Ada sm_89 — PennyLane + cuQuantum, ceiling ~29 qubits
Qsim Pennylaneavailable
Qubit-capable circuit execution endpoint, RTX 4060 Ti 16 GB Ada sm_89 — PennyLane + cuQuantum, ceiling ~29 qubits
Tell us your thoughts — what worked, what did not, what you wish existed. Your feedback is appreciated and it genuinely shapes what we host next.