CPCIVICPROOF LIVE
STATUS / LIVEBENCHMARK / V2.0.0CREATED IN NYC

CIVIC AI EVALUATION INFRASTRUCTURE

MAKE
MODELS
ACCOUNTABLE.

One public benchmark. Every model version. Every score tied to inspectable evidence.

CivicProof Live is an open observatory for tracking how AI systems interpret civic risk, power, consent, transparency, and human agency over time.

100 fixed cases8 civic principles1 reference baseline0 fabricated results

01 / THE PRODUCT

NOT ANOTHER AI LEADERBOARD.

Most leaderboards ask which model is strongest. CivicProof asks a different question: when a model judges a civic proposal, can anyone inspect the rule, output, evidence, disagreement, and change across versions?

RUNFixed cases

Every system receives the same 100 public proposals.

PROVEEvidence first

No score becomes public without a stable audit trail.

TRACKModel drift

Successive releases reveal regressions and value changes.

CHALLENGEPublic disagreement

Cases, labels, and methodology remain open to criticism.

02 / WHAT EXISTS TODAY

THE FOUNDATION IS ALREADY RUNNING.

The reference engine reproduces its own published rules. That proves consistency only. Independent model results remain empty until genuine external evidence is submitted and verified.

03 / POSITIONING

BUILT FROM PROVEN MECHANISMS.

Proven systemWhat works thereCivicProof contribution
ArenaHuman preference at internet scaleFixed civic cases plus inspectable evidence
Artificial AnalysisContinuous model performance measurementVersion-to-version civic behavior tracking
NIST ARIAHuman field testing and expert annotationOpen submission and human verification contract
PolisLarge-scale public deliberationFuture public disagreement and appeal layer

04 / THE NEXT PROOF

THE FIRST FIVE INDEPENDENT RUNS.

Traffic is not the immediate bottleneck. Independent evidence is. The next milestone is five complete runs from distinct frontier and open models, each with raw outputs and exact versions.

  1. Run OpenAI, Anthropic, Google, and two open models.
  2. Publish every raw response and configuration.
  3. Verify the evidence before placing any public score.
  4. Repeat after model releases to expose behavioral drift.
Start the first external run

CREATED BY

MARCO HERGI

NYC creator and independent builder

CivicProof is useful independent infrastructure first. Marco is attributable because he builds, publishes, and maintains the evidence system.

Creator recordGitHub sourceInstagram