Verification models
Judges under two billion parameters, trained on verdicts rather than prose. They score grounding, citations, and constraints on device.
Deep-tech research lab · Grounding verification
Fahrenheit Research builds grounding verifiers: small judge models that sit beside a domain model on your hardware and prove every answer before it ships.
One category. One job. The verdict.
Lens 06 · Binary · The problem
Fluent answers are everywhere. Proof is not. The gap between the two is where adoption stalls, budgets leak, and value pools.
Lens 02 · Quadtree · What the lab builds
We build the author and the examiner: small language models that know the domain, thin today and industry-specific next, and verification models that hold every answer to ground truth.
Judges under two billion parameters, trained on verdicts rather than prose. They score grounding, citations, and constraints on device.
The category we coined. Frontier-grade judgment for one domain, on a laptop.
Weights and checks that run on local silicon, with no cloud round trip.
Our thesis: automatic verification, not data or compute, decides the winners.
Pipelines that surface signal, never blind scraped.
Select projects where we build the domain-specific model and its verifier for your stack.
Lens 04 · Halftone · Model 001
The first Fahrenheit model. A small judge that reads an answer against its evidence and returns a verdict on device, in milliseconds, with no cloud in the loop. It does not generate. It checks.
Pair it with any small language model or FR model: the model writes, Rankine grounds every claim in evidence before the answer ships.
Every claim is scored against evidence; a weak sentence cannot hide in a fluent paragraph.
Each verdict cites the exact passage, so review takes seconds instead of a re-read.
Thin evidence returns "unverifiable" instead of a guess.
Verdicts in milliseconds, fully offline. Named for the Rankine scale: absolute measurement, in Fahrenheit degrees.
Lens 03 · Circle pack · The thesis
Our meta-analysis across five model lines: task classes with automatic verifiers gain roughly twelve points per quarter. Classes without them show no computable trajectory at all.
Lens 05 · Trace · Research
Pre-registered protocols, public predictions, and results that survive hostile review. Every thread feeds the verifier line.
A pre-registered protocol and meta-analysis on why automatic verification decides domain AI.
Verdicts with evidence spans, built for retrieval-heavy work.
Sub-2B domain models that run locally.
10× smaller, under 3% accuracy loss.
Lens 07 · Dither · Principles
One philosophy, drawn in two tones.
Refuse
Own
Lens 11 · Quantize · Contact
Researchers, enterprises, and builders who need answers they can prove: reach the lab directly.
Custom engagements: your domain, your data, your hardware.
Lens 12 · Moire · Before you reach out
A small model whose only job is to check another model's answer: grounding, citations, constraints. It returns a verdict with evidence, not more prose.
Our meta-analysis across five model lines finds that task classes with automatic verifiers improve roughly twelve points per quarter, while verifier-poor classes show no computable trajectory. Scale generates. Verification compounds.
The first Fahrenheit model: a grounding verifier that runs on device and pairs with any small language model or FR model. It is in training now.
Yes. The lab takes on select engagements to build custom domain-specific models, from data architecture to a deployed model with its verifier. Write to research@f-r.co.