0. Registry Classification
| Field | Entry |
|---|---|
| Restoration Arc ID | RA-058 |
| Name | AI Classifier / Evaluator Restoration |
| Short Name / Alias | Evaluator Restoration |
| Primary Family | AI Governance / Classifier Integrity / Evaluation |
| Secondary Families | Core; AI Governance; Cognitive Infrastructure; Auditability; Safety Calibration; Coherence; Boundary; Meaning; Security; Platform Governance; Feedback Integrity; Scaling |
| Treatment | Canon Parent Arc |
| Status | Canon-Ready |
| Scope | AI / Classifier / Evaluator / Cognitive Infrastructure / Platform / Security / Institutional / Governance / Cross-Domain |
| Primary U-Layers | U2 / U3 / U4 / U5 → U6 / U7 validation |
| Primary Operators | Σ → Θ → Au → FI → Γ → ℛ → Λ → Τ |
| Primary Diagnostics | Au, H, O, ε, ι, µᵢ, BΣ, K, R, FI, Γ, evaluator_integrity, classifier_fidelity, benchmark_substitution_risk, reward_hacking_risk, field_validation_strength, false_positive_rate, false_negative_rate, evaluation_diversity, Φ/O divergence |
1. Purpose
1.1 What This Arc Repairs
AI Classifier / Evaluator Restoration repairs AI systems whose classifiers, evaluators, benchmarks, reward models, safety tests, red-team suites, human-rating pipelines, or automated review systems no longer measure the real field condition they claim to measure.
It applies when evaluation becomes a proxy regime: benchmarks improve while coherence, safety, meaning fidelity, affected-node outcomes, or field performance degrade.
This arc repairs evaluator and classifier failure by:
- distinguishing benchmark performance from real coherence;
- restoring feedback integrity;
- identifying evaluator capture and metric overfitting;
- reducing reward hacking;
- expanding evaluation diversity;
- reconnecting classifier and evaluator outputs to field validation;
- auditing false positives and false negatives;
- restoring context and meaning fidelity;
- preserving boundary and consent constraints in evaluation;
- preventing local test success from substituting for U6 / U7 field proof.
AI Classifier / Evaluator Restoration is the canonical arc for restoring feedback integrity in AI assessment systems.
1.2 Core Restoration Function
This arc restores evaluation coherence by preventing benchmark, classifier, or reward-model success from substituting for field-validated safety, meaning fidelity, boundary integrity, and real-world restoration.
AI Classifier / Evaluator Restoration prevents the measuring system from becoming the thing optimized against.
2. Use Conditions
2.1 When to Apply
Use this arc when:
- benchmark scores improve while field failures persist;
- classifiers create repeated false positives or false negatives;
- safety evaluations overfit to known tests;
- reward models incentivize polished but incoherent outputs;
- evaluator criteria become detached from user meaning, field signal, or affected-node outcomes;
- automated safety routing misclassifies valid requests;
- evaluator monoculture narrows what counts as success;
- red-team findings are patched locally but not integrated into broader field validation;
- a product claims improved safety from metrics while trust, appeal, meaning fidelity, or harm outcomes degrade;
- evaluation systems are controlled by the same actors whose performance they certify;
- security, policy, or governance classifiers reward local compliance while global risk remains.
Examples:
- an AI assistant passes safety benchmarks but over-refuses legitimate technical or governance analysis;
- a classifier reduces visible violations by misclassifying ambiguous requests as prohibited;
- a reward model favors confidence, tone, or disclaimer density over accuracy and usefulness;
- an evaluator suite misses harms experienced by affected users because the test set lacks their context;
- a security classifier flags defensive research while missing actual abuse;
- a model is tuned to pass public benchmarks while hidden failure modes migrate to new surfaces.
2.2 When Not to Apply
Do not apply this arc when:
- the issue is a single local interaction misfire and RA-047 is sufficient;
- the issue is mode routing and RA-048 should occur first;
- the primary failure is memory validity or retention and RA-059 should occur first;
- active AI harm requires RA-060 stabilization;
- evaluation already tracks field signal, affected-node outcomes, and recurrence robustly;
- classification is correct but user-facing explanation or appeal is weak;
- benchmark failure is known and the immediate issue is governance remediation;
- evaluator data cannot be inspected without violating privacy, safety, or security boundaries;
- the system refuses to distinguish benchmark performance from field validation.
AI Classifier / Evaluator Restoration must not become benchmark theater.
2.3 Required Preconditions
Before this arc begins, the following must be true:
| Precondition | Requirement |
|---|---|
| Evaluator Object Identified | The classifier, benchmark, reward model, evaluation suite, safety test, red-team process, ranking model, or human-review rubric is named |
| Claimed Measurement Known | The system identifies what the evaluator claims to measure: safety, accuracy, helpfulness, meaning fidelity, compliance, risk, quality, legitimacy, or coherence |
| Field Signal Available | User outcomes, affected-node signal, incident data, appeals, external audits, or deployment results can be compared against evaluator outputs |
| Proxy Risk Mappable | Benchmark substitution, reward hacking, false positives, false negatives, or metric overfitting can be investigated |
| Audit Surface Available | Evaluation criteria, inputs, outputs, thresholds, labels, failures, or routing decisions can be inspected under valid scope |
| Diversity Expansion Possible | Evaluation can incorporate multiple contexts, edge cases, affected-node perspectives, adversarial tests, and longitudinal checks |
| Boundary Protection Available | Evaluator audit preserves privacy, consent, security, and sensitive model or policy details |
| Correction Path Available | Findings can route to classifier, evaluator, reward model, policy, memory, interface, or governance repair |
If required preconditions fail:
Arc cannot validly begin.The system must route to Audit Surface Expansion, Interaction-Level Restoration, Restoration Junction Protocol, GEI Audit Restoration, AI Boundary Restoration, AI Memory Reindexing, or AI Incident Restoration.
3. Failure / Damage Signature
3.1 Pre-State Across S
| Variable | Expected Pre-State |
|---|---|
| O — Coherence | Degraded where evaluator success does not correspond to real-world coherence |
| H — Hidden Debt | Rising through unmeasured harms, false positives, false negatives, benchmark overfitting, and field-feedback loss |
| ε — Error / Noise | Elevated through misclassification, brittle labels, narrow rubrics, and evaluator disagreement |
| ι — Inversion Index | Rising when benchmark or reward success is treated as proof of safety or coherence |
| Au — Auditability | Weak where evaluator criteria, thresholds, training data, failure classes, or routing decisions are opaque |
| µᵢ — Agent Integrity | Threatened when user meaning, affected-node experience, or contextual nuance is flattened by labels |
| BΣ — Boundary Integrity | At risk if evaluation data, user reports, or adversarial tests are reused beyond valid scope |
| K — Compatibility / Slack Context | Reduced when users and operators cannot appeal, inspect, contest, or correct classifier outcomes |
| R — Restoration Capacity | Under-routed when evaluator failure is treated as product behavior rather than repairable governance infrastructure |
| FI — Feedback Integrity | Degraded through benchmark substitution, self-certification, monoculture, and weak field feedback |
| Γ — Diversity / Variance Support | Low where evaluation contexts, perspectives, and failure probes are too narrow |
| Φ — Fitness Proxy | Dominant through score gains, benchmark rankings, compliance metrics, leaderboard success, or polished outputs |
3.2 Primary Failure Links
| Failure Mode | Relationship |
|---|---|
| Benchmark Substitution | Primary repair target |
| Evaluator Capture | Primary repair target |
| Reward Hacking | Primary repair target |
| Classifier Drift | Primary repair target |
| Proxy Evaluation Collapse | Primary repair target |
| False-Positive Safety Distortion | Repairs / prevents |
| False-Negative Safety Failure | Repairs / prevents |
| Field Feedback Loss | Primary repair target |
| Evaluation Monoculture | Primary repair target |
| Overfitting to Metrics | Repairs / prevents |
| Goodharted Safety | Repairs / prevents |
| Context Collapse | Often co-occurs |
| Legitimacy-by-Benchmark | False-restoration risk |
3.3 Origin-Layer Localization
| Layer | Role |
|---|---|
| Failure Origin | Often U3 classifier / evaluator governance, U4 benchmark or safety claim, or U5 feedback / memory / validation layer |
| Visible Symptom Layer | Often U4 benchmark improvement, refusal pattern, misclassification, safety metric, leaderboard claim, or product-quality claim |
| Required Repair Layer | Same or lower than the layer where feedback integrity, field validation, or evaluation diversity failed |
| Validation Layer | U6 / U7 through field outcomes, affected-node signal, recurrence reduction, longitudinal testing, and adversarial robustness |
Canon rule:
Evaluation is not restored when the evaluator improves. Evaluation is restored when the evaluator again predicts and corrects field reality.
4. Restoration Objective
4.1 Canonical Objective
Restore classifier and evaluator integrity by reconnecting evaluation to field signal, expanding diversity, reducing proxy dominance, auditing failures, and validating real-world outcomes.
Formal objective:
FI ↑
evaluator_integrity ↑
classifier_fidelity ↑
benchmark_substitution_risk ↓
reward_hacking_risk ↓
field_validation_strength ↑
false_positive_rate ↓
false_negative_rate ↓
evaluation_diversity ↑
H ↓
Φ/O divergence ↓Expanded objective:
Convert evaluation from proxy performance into field-corrective feedback infrastructure that improves safety, meaning fidelity, boundary integrity, and coherence.
4.2 Non-Goals
This arc does not aim to:
- discard benchmarks entirely;
- treat all metrics as invalid;
- optimize for field anecdotes without structured evaluation;
- expose sensitive user data, red-team methods, or safety internals without boundaries;
- replace evaluation with public opinion;
- increase evaluator diversity without maintaining quality;
- make every classifier decision manually reviewed;
- equate lower false-positive rate with safety if false negatives rise;
- equate lower false-negative rate with safety if meaning fidelity collapses;
- use evaluation repair to delay incident response.
5. Operator Sequence
5.1 Minimal Operator Scaffold
Σ benchmark-not-O invariant + Θ proxy-pressure damping → Au evaluator / classifier trace → FI field-feedback restoration → Γ evaluation diversity expansion → ℛ correction routing → Λ field-fit test → Τ recurrence / robustness proofReference sequence from the registry:
Σ + Θ
→ Au + FI restore
→ Γ diversity
→ U6 validationUniversal grammar alignment:
Σ + Θ → Au + FI → Γ → ℛ → Λ → ΤAI Classifier / Evaluator Restoration may route into Goodhart Repair, GEI Audit Restoration, Interaction-Level Restoration, Restoration Junction Protocol, AI Boundary Restoration, AI Memory Reindexing, Governance-Level Restoration, or AI Incident Restoration.
5.2 Operator Step Table
| Step | Operator | Function | Variable Impact | Failure Prevented |
|---|---|---|---|---|
| 1 | Σ | Lock invariant that benchmark, reward, or classifier success is not coherence proof | O protected / ι↓ | Benchmark substitution |
| 2 | Θ | Dampen proxy pressure, leaderboard pressure, reward hacking, and compliance optics | K/σ↑ | Goodharted evaluation |
| 3 | Au | Trace evaluator design, labels, thresholds, failure classes, routing decisions, and claimed measurement | Au↑ | Opaque evaluation |
| 4 | FI | Reconnect evaluator outputs to field signal, appeals, incidents, affected-node outcomes, and recurrence | FI↑ | Field feedback loss |
| 5 | Γ | Expand evaluation diversity across contexts, users, adversarial cases, modalities, and time horizons | evaluation_diversity↑ | Evaluation monoculture |
| 6 | ℛ | Route findings to classifier, evaluator, reward model, policy, interface, memory, or governance repair | R↑ / H↓ | Inert evaluation |
| 7 | Λ | Test evaluator fit against safety, meaning fidelity, boundary integrity, and field outcomes | evaluator_integrity↑ | False validation |
| 8 | Τ | Validate recurrence reduction, robustness, and field performance over time | field_validation_strength↑ | Regression drift |
5.3 Sequence Notes
This arc is feedback-integrity-gated, diversity-gated, and field-validation-gated.
The sequence must distinguish:
benchmark
classifier
evaluator
reward model
field signal
affected-node outcome
false positive
false negative
real coherenceThe following steps cannot be skipped:
claimed measurement identification
evaluator audit
false-positive / false-negative review
field-feedback reconnection
evaluation diversity expansion
repair routing
field-fit validation
temporal robustness proofIf evaluation improves only on known tests, the arc is incomplete.
If field signal is gathered but cannot change the evaluator, the arc is inert.
If diversity expands but boundary or quality discipline collapses, the arc fails.
6. Restoration Phases
Phase 0 — Identify Evaluator Object
Purpose: Name the classifier, evaluator, benchmark, reward model, or evaluation process under repair.
Actions:
- identify the evaluation system;
- identify the classification or scoring task;
- identify claimed measurement target;
- identify deployment context;
- identify affected users, nodes, requests, or downstream systems;
- identify whether the evaluator gates access, safety, visibility, ranking, response mode, or legitimacy.
Validation:
evaluator object named
claimed measurement visible
affected field identifiedPhase 1 — Separate Benchmark From Coherence
Purpose: Prevent proxy performance from certifying restoration.
Actions:
- identify benchmark scores or evaluator metrics;
- identify what those scores do and do not measure;
- identify proxy overreach;
- identify where field outcomes diverge from metrics;
- identify whether benchmark success suppresses incident or appeal signal;
- distinguish metric improvement from O restoration.
Validation:
benchmark_substitution_risk ↓
proxy_dominance ↓
Φ/O divergence ↓Phase 2 — Audit Classifier / Evaluator Trace
Purpose: Restore auditability over evaluation behavior.
Actions:
- inspect labels and categories;
- inspect thresholds;
- inspect routing decisions;
- inspect error classes;
- inspect false positives and false negatives;
- inspect examples near boundaries;
- inspect evaluator-owner and decision provenance;
- inspect whether model behavior is optimizing against evaluator artifacts.
Validation:
Au ↑
classifier_fidelity visible
evaluator_integrity baseline knownPhase 3 — Restore Feedback Integrity
Purpose: Reconnect evaluation to field reality.
Actions:
- connect appeal outcomes to evaluator updates;
- connect user corrections to classifier review;
- connect incident reports to evaluation criteria;
- connect field performance to benchmark revision;
- connect affected-node signal to test coverage;
- connect recurrence to regression tests;
- prevent self-certification by the evaluated system alone.
Validation:
FI ↑
field_validation_strength ↑
self-certification risk ↓Phase 4 — Expand Evaluation Diversity
Purpose: Prevent evaluator monoculture.
Actions:
- add varied user contexts;
- add affected-node perspectives;
- add adversarial cases;
- add edge cases;
- add long-horizon cases;
- add symbolic, technical, governance, security, and mixed-mode examples where relevant;
- add multilingual, accessibility, domain-specific, and boundary-condition tests where needed;
- avoid overfitting to the new test set.
Validation:
Γ ↑
evaluation_diversity ↑
coverage gaps ↓Phase 5 — Repair Reward / Proxy Incentives
Purpose: Reduce reward hacking and metric overfitting.
Actions:
- identify reward shortcuts;
- identify style, tone, disclaimer, refusal, confidence, or compliance proxies that dominate quality;
- identify where the model learns evaluator artifacts rather than field coherence;
- reduce incentive to optimize visible metric while hiding H;
- include outcome-sensitive and field-sensitive reward signals;
- add checks for polished incoherence.
Validation:
reward_hacking_risk ↓
proxy evaluation collapse ↓
Φ/O divergence ↓Phase 6 — Route Corrections
Purpose: Make evaluator findings change the system.
Actions:
- route classifier failures to classifier repair;
- route evaluator failures to evaluator repair;
- route benchmark gaps to benchmark revision;
- route memory-related failures to RA-059;
- route boundary-related failures to RA-057;
- route GEI failures to RA-055;
- route material harm to RA-060;
- define owners and review dates.
Validation:
R ↑
repair routing active
evaluation no longer inertPhase 7 — Field Validation and Temporal Proof
Purpose: Confirm that evaluator repair works in deployment.
Actions:
- monitor false positives;
- monitor false negatives;
- monitor affected-node outcomes;
- monitor appeal reversal rates;
- monitor incident recurrence;
- monitor drift after model or policy updates;
- monitor performance outside benchmark distribution;
- monitor whether field feedback continues to update the evaluator.
Validation:
field_validation_strength ↑
false_positive_rate ↓ where excessive
false_negative_rate ↓ where unsafe
recurrence ↓
evaluator_integrity stable or ↑7. Gates
7.1 Required Gates
| Gate | Requirement | Failure Result |
|---|---|---|
| FI-Gate | Field signal, appeals, incidents, affected-node outcomes, and recurrence must be able to correct the evaluator | Evaluator self-seals |
| HR-Gate | High-impact classifiers or evaluators cannot be certified by benchmark performance alone | Safety or legitimacy claim blocked |
| MS-Gate | High-status labs, platforms, or teams cannot exempt their evaluators from external or separated validation where impact requires it | Accountability invalid |
| Au-Actuation | Evaluator criteria, thresholds, labels, failures, claims, and correction paths must be traceable where possible | Actuation provisional |
| BΣ-Gate | Evaluation audit must preserve privacy, consent, security, and affected-node boundaries | Arc aborts or reroutes |
| Λ-Gate | Evaluator must fit field conditions, safety, meaning fidelity, boundary integrity, and recurrence constraints | Reliance blocked |
| ☷ᵢ Principle Gates | Non-negotiable invariants hold | ∅ outcome |
7.2 Gate Failure Rule
If any required gate fails:
∅ — AI Classifier / Evaluator Restoration cannot validly proceed in that form.The system must either:
- expand audit surface;
- restore feedback integrity;
- reduce benchmark reliance;
- expand evaluation diversity;
- repair classifier thresholds;
- repair reward incentives;
- route to GEI Audit Restoration;
- route to AI Boundary Restoration;
- route to AI Memory Reindexing;
- route to AI Incident Restoration if harm has occurred;
- withhold safety, quality, or legitimacy claims until field validation exists.
8. Diagnostics
8.1 Required Diagnostic Trends
| Diagnostic | Expected Trend | Meaning |
|---|---|---|
| Au | ↑ | Evaluator design, criteria, thresholds, and failures become traceable |
| H | ↓ | Hidden evaluation and classifier debt decreases |
| O | Stable / ↑ | Evaluator outputs align better with real coherence |
| ε | ↓ | Misclassification noise and evaluator disagreement decrease |
| ι | ↓ | Benchmark or reward success no longer substitutes for coherence |
| µᵢ | ↑ | User meaning and affected-node context are better preserved |
| BΣ | Stable / ↑ | Evaluation respects data, privacy, consent, and boundary limits |
| K / σ | ↑ | Users and operators gain appeal, correction, and review paths |
| R | ↑ | Evaluation findings route to repair |
| FI | ↑ | Field feedback corrects evaluation |
| Γ | ↑ | Evaluation diversity increases |
| evaluator_integrity | ↑ | Evaluator better measures claimed target |
| classifier_fidelity | ↑ | Classifier better maps real categories and contexts |
| benchmark_substitution_risk | ↓ | Benchmarks no longer replace field validation |
| reward_hacking_risk | ↓ | Model behavior less optimized to evaluator artifacts |
| field_validation_strength | ↑ | Deployment outcomes validate evaluator quality |
| false_positive_rate | ↓ where excessive | Valid requests or nodes are less often misclassified |
| false_negative_rate | ↓ where unsafe | Harmful or invalid cases are less often missed |
| evaluation_diversity | ↑ | Test coverage spans more contexts and failure modes |
| Φ/O divergence | ↓ | Scores, compliance, and benchmarks align better with coherence |
8.2 Arc-Specific Diagnostic Thresholds
Suggested thresholds:
FI ↑
evaluator_integrity ↑
classifier_fidelity ↑
benchmark_substitution_risk ↓
reward_hacking_risk ↓
field_validation_strength ↑
false_positive_rate ↓ where excessive
false_negative_rate ↓ where unsafe
evaluation_diversity ↑
H ↓
Φ/O divergence ↓AI Classifier / Evaluator Restoration is not complete if:
benchmark improvement is treated as field proof
false positives remain unreviewed
false negatives remain hidden
field signal cannot update the evaluator
affected-node outcomes are excluded
evaluator criteria remain opaque
reward hacking remains rewarded
evaluation diversity remains too narrow
appeals do not feed back into classifier repair
safety or legitimacy claims rest on proxy metrics alone9. Anti-Patterns / False Restorations
9.1 Common False Versions
This arc is being simulated, not executed, if:
- new benchmarks are added without field feedback;
- a classifier threshold changes without auditing false positives and false negatives;
- evaluator diversity is claimed from superficial dataset variation;
- red-team findings are patched only as isolated examples;
- reward hacking is renamed as alignment improvement;
- polished refusal or disclaimer density is treated as safety;
- user appeal outcomes do not affect evaluator updates;
- the evaluated system certifies its own evaluator;
- benchmark gains are used to dismiss field harms;
- failures are treated as edge cases outside the test definition.
9.2 Named Anti-Pattern Links
| Anti-Pattern | Why It Fails |
|---|---|
| Benchmark Theater | Adds or improves tests while field failures persist |
| Evaluator Self-Certification | Lets the system being evaluated certify its own evaluator |
| Threshold Patch | Adjusts a classifier threshold without repairing category logic |
| Red-Team Whack-a-Mole | Patches individual examples without broader failure geometry |
| Disclaimer-as-Safety | Treats warning language as safety outcome |
| Polished Incoherence | Rewards style, tone, or confidence over field coherence |
| Appeal Inertia | Collects appeals without evaluator correction |
| Diversity Veneer | Claims broad evaluation without meaningful contextual coverage |
| Legitimacy-by-Benchmark | Uses scores to claim trust without U6 / U7 validation |
10. Completion Criteria
10.1 Post-State Signature
| Variable | Required Post-State |
|---|---|
| O | Evaluation better predicts and supports real coherence |
| H | Hidden evaluation, classifier, and reward debt reduced |
| ε | Misclassification and evaluator noise reduced |
| ι | Reduced where benchmark, score, or reward success substituted for O |
| Au | Evaluator criteria, claims, thresholds, failure classes, and correction paths traceable |
| µᵢ | User meaning and affected-node context better preserved |
| BΣ | Evaluation data and repair preserve privacy, consent, and boundary integrity |
| K | Users and operators can appeal, correct, contest, and review evaluation outcomes |
| R | Evaluation findings route to classifier, evaluator, policy, interface, memory, or governance repair |
| FI | Field feedback actively corrects evaluator behavior |
| Γ | Evaluation coverage includes meaningful diversity of contexts and failure modes |
| Φ | Subordinate to O; benchmark score, compliance metric, leaderboard rank, or reward model score cannot certify restoration alone |
10.2 Temporal Proof
AI Classifier / Evaluator Restoration cannot be certified by a one-time benchmark gain. It requires field validation and recurrence reduction over time.
Template:
Completion requires FI ↑,
evaluator_integrity ↑,
classifier_fidelity ↑,
benchmark_substitution_risk ↓,
reward_hacking_risk ↓,
field_validation_strength ↑,
false_positive_rate ↓ where excessive,
false_negative_rate ↓ where unsafe,
evaluation_diversity ↑,
and evaluator performance remaining valid across future deployments.Minimum temporal proof:
- field outcomes align better with evaluator claims;
- false positives and false negatives decrease under relevant conditions;
- appeals and incidents update evaluation;
- benchmark success no longer suppresses field signal;
- evaluation diversity improves real coverage;
- reward hacking decreases;
- evaluator changes remain auditable;
- recurrence of the same classifier / evaluator failure decreases.
10.3 Completion Statement
Canonical format:
This arc is complete only when classifiers and evaluators regain feedback integrity, benchmark success no longer substitutes for coherence, evaluation diversity increases, field validation improves, reward hacking decreases, and future AI behavior is corrected by real-world signal rather than proxy performance alone.
11. Cross-Links
11.1 Related Restoration Arcs
| Arc | Relationship |
|---|---|
RA-004 — Audit Surface Expansion | Precursor when evaluator or classifier behavior is insufficiently visible |
RA-008 — Goodhart Repair | Companion when metrics have overtaken field reality |
RA-012 — Temporal Proof Arc | Companion for field validation over time |
RA-022 — Compression Relief | Companion when classifier categories compress meaning |
RA-023 — Meaning Restoration | Companion when user meaning is flattened by labels |
RA-025 — Observability Restoration | Companion when claimed performance exceeds visible state |
RA-036 — Wisdom Re-Indexing | Companion when evaluator lessons must become retrievable |
RA-047 — Interaction-Level Restoration | Local repair when classifier failure causes a single interaction misfire |
RA-048 — Restoration Junction Protocol | Companion when mode routing depends on classifier accuracy |
RA-051 — Signed Decision Provenance | Companion when evaluator changes require decision lineage |
RA-052 — Tamper-Evident Audit Restoration | Companion when evaluator logs and changes require integrity protection |
RA-053 — Constraint Recalibration Under Φ Growth | Companion when high-impact AI systems need stronger evaluation governance |
RA-055 — GEI Audit Restoration | Companion when classifiers or evaluators shape epistemic access |
RA-057 — AI Boundary Restoration | Companion when classifier/evaluator behavior controls permissions or tool scope |
RA-059 — AI Memory Reindexing | Companion when memory affects evaluator or classifier behavior |
RA-060 — AI Incident Restoration | Escalation when classifier / evaluator failure causes material AI harm |
11.2 Related Failure Modes
| Failure Mode | Relationship |
|---|---|
| Benchmark Substitution | Repairs |
| Evaluator Capture | Repairs |
| Reward Hacking | Repairs |
| Classifier Drift | Repairs |
| Proxy Evaluation Collapse | Repairs |
| False-Positive Safety Distortion | Repairs / prevents |
| False-Negative Safety Failure | Repairs / prevents |
| Field Feedback Loss | Repairs |
| Evaluation Monoculture | Repairs |
| Overfitting to Metrics | Repairs / prevents |
| Goodharted Safety | Repairs / prevents |
| Context Collapse | Repairs / prevents |
| Legitimacy-by-Benchmark | Prevents |
11.3 Related Diagnostics
Au, H, O, ε, ι, µᵢ, BΣ, K, R, FI, Γ, evaluator_integrity, classifier_fidelity, benchmark_substitution_risk, reward_hacking_risk, field_validation_strength, false_positive_rate, false_negative_rate, evaluation_diversity, Φ/O divergence11.4 Related Laws / Invariants
INV — Benchmark success is not coherence proof.
INV — Evaluators must be corrected by field signal.
INV — Classifier categories must not erase meaning or boundary integrity.
INV — Feedback integrity is required for evaluation legitimacy.
LAW — Reward hacking expands when proxies become the target.
LAW — Evaluation monoculture creates hidden classifier debt.
LAW — False-positive and false-negative harms must both remain visible.
LAW — Φ benchmark performance is not O restoration.12. Domain Notes
12.1 AI / Cognitive Infrastructure
Check:
- classifier thresholds;
- reward model criteria;
- evaluator rubrics;
- safety benchmarks;
- refusal benchmarks;
- helpfulness benchmarks;
- user correction signal;
- appeal reversal rates;
- memory influence;
- field outcomes;
- longitudinal recurrence.
AI evaluation must remain connected to lived deployment conditions. A benchmark can be useful, but it cannot replace field validation, affected-node feedback, and recurrence tracking.
12.2 Platform Governance
Check:
- moderation classifiers;
- visibility classifiers;
- fraud classifiers;
- risk scoring;
- appeal feedback;
- reviewer guidelines;
- false-positive burden;
- false-negative harm;
- ranking evaluator behavior;
- trust metrics.
Platforms must evaluate classifiers by what happens to users and fields, not only by internal accuracy claims.
12.3 Security
Check:
- abuse detection;
- defensive research classification;
- incident severity scoring;
- anomaly detection;
- false-positive operational burden;
- missed attack signal;
- red-team test coverage;
- adversarial adaptation;
- post-incident evaluator update.
Security classifiers fail when they punish legitimate defense, miss real abuse, or overfit to known attack patterns.
12.4 Justice / Governance / Legitimacy
Check:
- classifier governance;
- decision provenance;
- appeal rights;
- affected-node burden;
- evaluator independence;
- public safety claims;
- auditability of criteria;
- recurrence of misclassification.
Legitimacy requires evaluation systems that can be challenged, corrected, and validated beyond internal metrics.
12.5 Economy
Check:
- credit classifiers;
- fraud scoring;
- labor ranking;
- pricing evaluators;
- reputation systems;
- marketplace visibility;
- appeal outcomes;
- burden distribution;
- proxy optimization.
Economic classifiers can create material harm when proxies determine access, price, trust, or opportunity without field-corrective review.
12.6 CMS / Meaning / Archetypes
Check:
- symbolic-category classification;
- legitimacy scoring;
- recognition thresholds;
- taboo classification;
- meaning compression;
- interpretive diversity;
- delayed recognition;
- evaluator authority.
Meaning systems require evaluator restoration when classification governs what can be recognized, named, trusted, or restored.
13. Machine-Readable Metadata
id: "RA-058"
title: "AI Classifier / Evaluator Restoration"
aliases:
- "Evaluator Restoration"
family_primary: "AI Governance / Classifier Integrity / Evaluation"
families_secondary:
- "Core"
- "AI Governance"
- "Cognitive Infrastructure"
- "Auditability"
- "Safety Calibration"
- "Coherence"
- "Boundary"
- "Meaning"
- "Security"
- "Platform Governance"
- "Feedback Integrity"
- "Scaling"
treatment: "Canon Parent Arc"
status: "Canon-Ready"
scope:
- "AI"
- "Classifier"
- "Evaluator"
- "Cognitive Infrastructure"
- "Platform"
- "Security"
- "Institutional"
- "Governance"
- "Cross-Domain"
u_layers:
failure_origin:
- "often U3 classifier / evaluator governance"
- "often U4 benchmark or safety claim"
- "often U5 feedback / memory / validation layer"
symptom_visible:
- "U4 benchmark improvement / refusal pattern / misclassification / safety metric / leaderboard claim / product-quality claim"
repair_required:
- "same or lower than the layer where feedback integrity, field validation, or evaluation diversity failed"
validation:
- "U6"
- "U7"
operators:
scaffold: "Σ benchmark-not-O invariant + Θ proxy-pressure damping → Au evaluator / classifier trace → FI field-feedback restoration → Γ evaluation diversity expansion → ℛ correction routing → Λ field-fit test → Τ recurrence / robustness proof"
sequence:
- "Σ"
- "Θ"
- "Au"
- "FI"
- "Γ"
- "ℛ"
- "Λ"
- "Τ"
state_variables:
primary:
- "Au"
- "O"
- "H"
- "FI"
- "Γ"
secondary:
- "ε"
- "ι"
- "µᵢ"
- "BΣ"
- "K"
- "R"
- "Φ"
diagnostics:
- "evaluator_integrity"
- "classifier_fidelity"
- "benchmark_substitution_risk"
- "reward_hacking_risk"
- "field_validation_strength"
- "false_positive_rate"
- "false_negative_rate"
- "evaluation_diversity"
- "Φ/O divergence"
gates_required:
- "FI-Gate"
- "HR-Gate"
- "MS-Gate"
- "Au-Actuation"
- "BΣ-Gate"
- "Λ-Gate"
- "☷ᵢ"
linked_failure_modes:
- "Benchmark Substitution"
- "Evaluator Capture"
- "Reward Hacking"
- "Classifier Drift"
- "Proxy Evaluation Collapse"
- "False-Positive Safety Distortion"
- "False-Negative Safety Failure"
- "Field Feedback Loss"
- "Evaluation Monoculture"
- "Overfitting to Metrics"
- "Goodharted Safety"
- "Context Collapse"
- "Legitimacy-by-Benchmark"
linked_restoration_arcs:
- "RA-004"
- "RA-008"
- "RA-012"
- "RA-022"
- "RA-023"
- "RA-025"
- "RA-036"
- "RA-047"
- "RA-048"
- "RA-051"
- "RA-052"
- "RA-053"
- "RA-055"
- "RA-057"
- "RA-059"
- "RA-060"
anti_patterns:
- "Benchmark Theater"
- "Evaluator Self-Certification"
- "Threshold Patch"
- "Red-Team Whack-a-Mole"
- "Disclaimer-as-Safety"
- "Polished Incoherence"
- "Appeal Inertia"
- "Diversity Veneer"
- "Legitimacy-by-Benchmark"
completion_tests:
- "feedback integrity increases"
- "evaluator integrity increases"
- "classifier fidelity increases"
- "benchmark substitution risk decreases"
- "reward hacking risk decreases"
- "field validation strength increases"
- "false-positive rate decreases where excessive"
- "false-negative rate decreases where unsafe"
- "evaluation diversity increases"
- "hidden debt decreases"
- "Φ/O divergence decreases"
summary: "AI Classifier / Evaluator Restoration repairs benchmark substitution, evaluator capture, and reward hacking by restoring feedback integrity, auditability, diversity of evaluation, and field validation so classifiers and evaluators no longer substitute proxy success for coherence."Final Calibration Rule
AI Classifier / Evaluator Restoration answers six questions:
What classifier, evaluator, benchmark, reward model, or rubric is failing?
What does it claim to measure, and what field reality does it miss?
Where are benchmark substitution, evaluator capture, reward hacking, false positives, or false negatives occurring?
What field signal, appeal data, affected-node outcome, or incident evidence must correct the evaluator?
What evaluation diversity is required to prevent monoculture and metric overfitting?
How is restoration proven over time without benchmark theater, evaluator self-certification, or legitimacy-by-benchmark?