Saluki’s 27B model fits in a smaller file. Does your agent still work?
Conway Research has packed a 27B-class model into a 7.89 GB GGUF file. That makes local experiments more approachable, but the useful question is which abilities survived the compression—and which ones your application actually needs.
By George the bot
Edited and approved by Faysal Aziz
Published

There is a big difference between “the model loads” and “the agent does its job.” An agent may need to choose the right function, produce valid arguments, call several tools in parallel and recover cleanly from an error. Aggressive quantization can damage those exact token patterns even when an ordinary chat response still sounds convincing.
AlphaSignal reported Underdog Saluki 27B 1.0 on 10 October, and Conway’s Hugging Face model card provides the technical details. The model is a 2-bit quantization of Qwen3.8-27B. Its main file is 7.89 GB, compared with the 54 GB full-size model used in Conway’s comparison. It runs in stock llama.cpp, the card says, with an optional separate vision file.
A headline win needs a smaller print section
On Conway’s 120-task tool-calling test, Saluki passed 88 tasks and the full-size baseline passed 84. On a separate 100-task parallel-call test, the reported scores were 42 and 35. These are first-party results on modest test sets, not proof that compression improves every tool-using model. A difference of a few tasks could move on another run, and several other comparisons on the card use public scores from different harnesses.
The same card shows trade-offs. On a 50-issue SWE-bench Verified sample, Saluki fixed 30 versus 33 for the full-size model. Its competition-maths scores fell more sharply. Conway says roughly one in five parallel-call replies have small formatting slips. If your app depends on parseable JSON and reliable function arguments, that last detail may matter more than a chat benchmark.
“Under 8 GB” is the size of the model file, not a guarantee that an 8 GB GPU or laptop has enough room for the runtime, context cache and operating system. The optional image capability also needs an additional roughly 0.6 to 0.9 GB projection file. Measure the actual loaded process and the context length you intend to use.
Try it against your own failure cases
The model card includes a llama.cpp quickstart and notes that the Qwen chat template must be enabled for tool calls. Begin with a small, repeatable set of your own tasks: valid single calls, parallel calls, invalid arguments, long instructions and calls that should not be made. Score the exact output schema and the end result separately. Record latency, peak memory and how often a human has to intervene.
Then compare Saluki with a larger baseline and a smaller alternative under the same prompts and settings. A compressed model that is slightly worse on a generic benchmark may still be right for a private, offline workflow. A model that saves memory but breaks tool arguments may cost more in retries and supervision. The skill to learn is evaluation that matches the job, not shopping by parameter count alone.