Deep-tech research lab · Grounding verification

Intelligence
you can verify.

Fahrenheit Research builds grounding verifiers: small judge models that sit beside a domain model on your hardware and prove every answer before it ships.

One category. One job. The verdict.

Pass -- / 12 · one subject Sweep
Live inspection
Click to run the next check

Lens 06 · Binary · The problem

Every domain can generate. Almost none can check.

Fluent answers are everywhere. Proof is not. The gap between the two is where adoption stalls, budgets leak, and value pools.

fr-lab · trust ledgerper deployment
001Fluent answers shipped without proof of groundingrisk
002Hallucinations found by customers, not by systemslate
003Every check routed to a hosted frontier model∞ $
004Data that cannot leave the device for reviewcrit
005Human review as the only quality gateslow
006No local model whose whole job is the verdictgap
Observed across the fieldVerifier-poor domains show no measurable progress

Lens 02 · Quadtree · What the lab builds

One lab. Two model lines.

We build the author and the examiner: small language models that know the domain, thin today and industry-specific next, and verification models that hold every answer to ground truth.

001 /

Verification models

Judges under two billion parameters, trained on verdicts rather than prose. They score grounding, citations, and constraints on device.

002 /

Thin Language Models

The category we coined. Frontier-grade judgment for one domain, on a laptop.

003 /

Edge-native runtime

Weights and checks that run on local silicon, with no cloud round trip.

004 /

The Verifier Boundary

Our thesis: automatic verification, not data or compute, decides the winners.

005 /

Deep data architecture

Pipelines that surface signal, never blind scraped.

006 /

Custom model engagements

Select projects where we build the domain-specific model and its verifier for your stack.

Lens 04 · Halftone · Model 001

Rankine, the grounding verifier

The first Fahrenheit model. A small judge that reads an answer against its evidence and returns a verdict on device, in milliseconds, with no cloud in the loop. It does not generate. It checks.

In training · lab hardware

FR-Rankine

Pair it with any small language model or FR model: the model writes, Rankine grounds every claim in evidence before the answer ships.

Class
Verifier
Job
Grounding
Run
On-device
Pairs with
Any SLM or FR model
001

Claim-level verdicts

Every claim is scored against evidence; a weak sentence cannot hide in a fluent paragraph.

002

Evidence spans

Each verdict cites the exact passage, so review takes seconds instead of a re-read.

003

Calibrated abstention

Thin evidence returns "unverifiable" instead of a guess.

Verdicts in milliseconds, fully offline. Named for the Rankine scale: absolute measurement, in Fahrenheit degrees.

Lens 03 · Circle pack · The thesis

Progress follows verification.

Our meta-analysis across five model lines: task classes with automatic verifiers gain roughly twelve points per quarter. Classes without them show no computable trajectory at all.

Lens 05 · Trace · Research

The lab publishes first, then ships what survives review.

Pre-registered protocols, public predictions, and results that survive hostile review. Every thread feeds the verifier line.

001

The Verifier Boundary

A pre-registered protocol and meta-analysis on why automatic verification decides domain AI.

002

Grounding verification

Verdicts with evidence spans, built for retrieval-heavy work.

003

Thin Language Models

Sub-2B domain models that run locally.

004

Compressed intelligence

10× smaller, under 3% accuracy loss.

Lens 07 · Dither · Principles

What we refuse. What we own.

One philosophy, drawn in two tones.

Refuse

  • We do not license breakthroughs.
  • We do not outsource thinking.
  • We do not rent intelligence.

Own

  • Every model we specialize.
  • Every dataset we build.
  • Every product we ship.

Lens 11 · Quantize · Contact

Work with the lab.

Researchers, enterprises, and builders who need answers they can prove: reach the lab directly.

Custom engagements: your domain, your data, your hardware.

Lens 12 · Moire · Before you reach out

Questions we get asked

A small model whose only job is to check another model's answer: grounding, citations, constraints. It returns a verdict with evidence, not more prose.

Our meta-analysis across five model lines finds that task classes with automatic verifiers improve roughly twelve points per quarter, while verifier-poor classes show no computable trajectory. Scale generates. Verification compounds.

The first Fahrenheit model: a grounding verifier that runs on device and pairs with any small language model or FR model. It is in training now.

Yes. The lab takes on select engagements to build custom domain-specific models, from data architecture to a deployed model with its verifier. Write to research@f-r.co.