0
All projects

Counterfactual Fact Verification

How well small local LLMs check FEVER claims with and without evidence, and how easily hand-written counterfactuals fool them.

Role
Wrote the initial FEVER preprocessing (claim extraction, evidence resolution) and the counterfactual template
When
Feb 2026 - Apr 2026
Status
Complete
Context
CS421, team of 4
Stack
  • Python
  • Transformers
  • bitsandbytes
  • Ollama
  • Phi-3 Mini
  • Mistral 7B
  • Llama 3.1 8B

The problem

Small models running locally are cheap enough to fact-check at scale, but how much do they actually know without evidence in front of them, and how easily are they misled? We tested that on FEVER, comparing zero-shot answers with evidence-backed ones, and then fed the models deliberately misleading counterfactual claims.

This was a four-person course project. I wrote the first FEVER preprocessing (claim extraction, picking the first sample, evidence resolution and the counterfactual template) and later rewrote one of the counterfactual tools. My teammates ran the tiering, the tier analysis and the Ollama and Llama experiments.

How it works

  1. 1

    Extract

    Pulls the 80,035 SUPPORTS claims out of FEVER's roughly 145,000.

  2. 2

    Tier

    Scores each claim on tokens, evidence sets and Wikipedia pages to sort it into low, medium or high structural complexity.

  3. 3

    Tier analysis

    Runs 300 claims per tier zero-shot and again with the gold evidence, on 4-bit quantized models.

  4. 4

    Cherry-pick

    Selects 500 claims across tiers and resolves their evidence from the FEVER Wikipedia dump.

  5. 5

    Counterfactuals

    The team hand-wrote counterfactual versions of claims in small GUI tools against local models.

  6. 6

    Batch evaluation

    Runs every counterfactual through each model and records whether it was fooled.

Decisions and tradeoffs

  • "Structural complexity" instead of "difficulty". The course TA pointed out the tiers had no empirical basis as a difficulty measure, so we renamed them to what they actually count.
  • Strategy-based counterfactuals instead of simple negation. At the midterm the models caught simple negations 93.3% of the time, so negation alone wasn't a real test.

What broke

The same Phi-3 Mini scored 14.1% zero-shot through Ollama's 4-bit build and 65.9% through Hugging Face NF4. Same model, different serving stack, a 50-point swing. We documented it as serving-stack sensitivity and moved the counterfactual work to Mistral and Llama.

Results

  • 65.9% vs 96.4%

    Phi-3 Mini (4-bit NF4) accuracy zero-shot vs. with evidence, 900 claims

    tier-analysis validation results, Apr 6 2026

  • 76.6% vs 96.9%

    Mistral 7B (4-bit NF4) accuracy zero-shot vs. with evidence, 900 claims

    tier-analysis results, Apr 9 2026

  • 90.3%

    Counterfactuals that fooled Mistral 7B (400 of 443)

    data/results/eval_final_mistral.json

  • 48.7%

    Counterfactuals that fooled Llama 3.1 8B (269 of 552)

    data/results/eval_final_v3_llama.json

  • 14.1% vs 65.9%

    The same Phi-3 Mini, zero-shot, served through Ollama's 4-bit build vs. Hugging Face NF4

    tier-analysis runs, Apr 11 and Apr 6 2026

Evidence closes most of the gap: both models land around 96% with it. Without evidence they're much weaker, and Mistral, the stronger zero-shot model, was the easier one to fool with counterfactuals.

Limits

Every claim in the tier analysis is a SUPPORTS claim, so accuracy there is really the rate of answering "supports". The runs also mix backends and model versions, so the cross-model comparisons are rough.

Next project

DevSecOps on AWS EKS

My build-scan-deploy pipeline and AWS setup for running a community three-tier app on EKS, with Terraform, Jenkins, SonarQube, Trivy, ECR and an ALB ingress.