RA-058 — AI Classifier / Evaluator Restoration

Open archive search
Archive registry entry

RA-058 — AI Classifier / Evaluator Restoration

AI Classifier / Evaluator Restoration repairs benchmark substitution, evaluator capture, and reward hacking by restoring feedback integrity, auditability, diversity of evaluation, and field validation so classifiers and evaluators no longer substitute proxy success for coherence.

reviewedid: RA-058version: 1.0updated: 2026-05-20
Archive Progress

This section can be read now; registry depth and cross-references are still being strengthened.

Foundation
Online

The section has a stable overview route and basic reader context.

Technical Layer
Online

A deeper technical overview is available.

Registry
Current

102 registry entries are available.

Cross-links
Curating

Related concepts are being connected conservatively for accuracy.

0. Registry Classification

TableScroll
FieldEntry
Restoration Arc IDRA-058
NameAI Classifier / Evaluator Restoration
Short Name / AliasEvaluator Restoration
Primary FamilyAI Governance / Classifier Integrity / Evaluation
Secondary FamiliesCore; AI Governance; Cognitive Infrastructure; Auditability; Safety Calibration; Coherence; Boundary; Meaning; Security; Platform Governance; Feedback Integrity; Scaling
TreatmentCanon Parent Arc
StatusCanon-Ready
ScopeAI / Classifier / Evaluator / Cognitive Infrastructure / Platform / Security / Institutional / Governance / Cross-Domain
Primary U-LayersU2 / U3 / U4 / U5 → U6 / U7 validation
Primary OperatorsΣ → Θ → Au → FI → Γ → ℛ → Λ → Τ
Primary DiagnosticsAu, H, O, ε, ι, µᵢ, BΣ, K, R, FI, Γ, evaluator_integrity, classifier_fidelity, benchmark_substitution_risk, reward_hacking_risk, field_validation_strength, false_positive_rate, false_negative_rate, evaluation_diversity, Φ/O divergence

1. Purpose

1.1 What This Arc Repairs

AI Classifier / Evaluator Restoration repairs AI systems whose classifiers, evaluators, benchmarks, reward models, safety tests, red-team suites, human-rating pipelines, or automated review systems no longer measure the real field condition they claim to measure.

It applies when evaluation becomes a proxy regime: benchmarks improve while coherence, safety, meaning fidelity, affected-node outcomes, or field performance degrade.

This arc repairs evaluator and classifier failure by:

  • distinguishing benchmark performance from real coherence;
  • restoring feedback integrity;
  • identifying evaluator capture and metric overfitting;
  • reducing reward hacking;
  • expanding evaluation diversity;
  • reconnecting classifier and evaluator outputs to field validation;
  • auditing false positives and false negatives;
  • restoring context and meaning fidelity;
  • preserving boundary and consent constraints in evaluation;
  • preventing local test success from substituting for U6 / U7 field proof.

AI Classifier / Evaluator Restoration is the canonical arc for restoring feedback integrity in AI assessment systems.


1.2 Core Restoration Function

This arc restores evaluation coherence by preventing benchmark, classifier, or reward-model success from substituting for field-validated safety, meaning fidelity, boundary integrity, and real-world restoration.

AI Classifier / Evaluator Restoration prevents the measuring system from becoming the thing optimized against.


2. Use Conditions

2.1 When to Apply

Use this arc when:

  • benchmark scores improve while field failures persist;
  • classifiers create repeated false positives or false negatives;
  • safety evaluations overfit to known tests;
  • reward models incentivize polished but incoherent outputs;
  • evaluator criteria become detached from user meaning, field signal, or affected-node outcomes;
  • automated safety routing misclassifies valid requests;
  • evaluator monoculture narrows what counts as success;
  • red-team findings are patched locally but not integrated into broader field validation;
  • a product claims improved safety from metrics while trust, appeal, meaning fidelity, or harm outcomes degrade;
  • evaluation systems are controlled by the same actors whose performance they certify;
  • security, policy, or governance classifiers reward local compliance while global risk remains.

Examples:

  • an AI assistant passes safety benchmarks but over-refuses legitimate technical or governance analysis;
  • a classifier reduces visible violations by misclassifying ambiguous requests as prohibited;
  • a reward model favors confidence, tone, or disclaimer density over accuracy and usefulness;
  • an evaluator suite misses harms experienced by affected users because the test set lacks their context;
  • a security classifier flags defensive research while missing actual abuse;
  • a model is tuned to pass public benchmarks while hidden failure modes migrate to new surfaces.

2.2 When Not to Apply

Do not apply this arc when:

  • the issue is a single local interaction misfire and RA-047 is sufficient;
  • the issue is mode routing and RA-048 should occur first;
  • the primary failure is memory validity or retention and RA-059 should occur first;
  • active AI harm requires RA-060 stabilization;
  • evaluation already tracks field signal, affected-node outcomes, and recurrence robustly;
  • classification is correct but user-facing explanation or appeal is weak;
  • benchmark failure is known and the immediate issue is governance remediation;
  • evaluator data cannot be inspected without violating privacy, safety, or security boundaries;
  • the system refuses to distinguish benchmark performance from field validation.

AI Classifier / Evaluator Restoration must not become benchmark theater.


2.3 Required Preconditions

Before this arc begins, the following must be true:

TableScroll
PreconditionRequirement
Evaluator Object IdentifiedThe classifier, benchmark, reward model, evaluation suite, safety test, red-team process, ranking model, or human-review rubric is named
Claimed Measurement KnownThe system identifies what the evaluator claims to measure: safety, accuracy, helpfulness, meaning fidelity, compliance, risk, quality, legitimacy, or coherence
Field Signal AvailableUser outcomes, affected-node signal, incident data, appeals, external audits, or deployment results can be compared against evaluator outputs
Proxy Risk MappableBenchmark substitution, reward hacking, false positives, false negatives, or metric overfitting can be investigated
Audit Surface AvailableEvaluation criteria, inputs, outputs, thresholds, labels, failures, or routing decisions can be inspected under valid scope
Diversity Expansion PossibleEvaluation can incorporate multiple contexts, edge cases, affected-node perspectives, adversarial tests, and longitudinal checks
Boundary Protection AvailableEvaluator audit preserves privacy, consent, security, and sensitive model or policy details
Correction Path AvailableFindings can route to classifier, evaluator, reward model, policy, memory, interface, or governance repair

If required preconditions fail:

textScroll
Arc cannot validly begin.

The system must route to Audit Surface Expansion, Interaction-Level Restoration, Restoration Junction Protocol, GEI Audit Restoration, AI Boundary Restoration, AI Memory Reindexing, or AI Incident Restoration.


3. Failure / Damage Signature

3.1 Pre-State Across S

TableScroll
VariableExpected Pre-State
O — CoherenceDegraded where evaluator success does not correspond to real-world coherence
H — Hidden DebtRising through unmeasured harms, false positives, false negatives, benchmark overfitting, and field-feedback loss
ε — Error / NoiseElevated through misclassification, brittle labels, narrow rubrics, and evaluator disagreement
ι — Inversion IndexRising when benchmark or reward success is treated as proof of safety or coherence
Au — AuditabilityWeak where evaluator criteria, thresholds, training data, failure classes, or routing decisions are opaque
µᵢ — Agent IntegrityThreatened when user meaning, affected-node experience, or contextual nuance is flattened by labels
BΣ — Boundary IntegrityAt risk if evaluation data, user reports, or adversarial tests are reused beyond valid scope
K — Compatibility / Slack ContextReduced when users and operators cannot appeal, inspect, contest, or correct classifier outcomes
R — Restoration CapacityUnder-routed when evaluator failure is treated as product behavior rather than repairable governance infrastructure
FI — Feedback IntegrityDegraded through benchmark substitution, self-certification, monoculture, and weak field feedback
Γ — Diversity / Variance SupportLow where evaluation contexts, perspectives, and failure probes are too narrow
Φ — Fitness ProxyDominant through score gains, benchmark rankings, compliance metrics, leaderboard success, or polished outputs

TableScroll
Failure ModeRelationship
Benchmark SubstitutionPrimary repair target
Evaluator CapturePrimary repair target
Reward HackingPrimary repair target
Classifier DriftPrimary repair target
Proxy Evaluation CollapsePrimary repair target
False-Positive Safety DistortionRepairs / prevents
False-Negative Safety FailureRepairs / prevents
Field Feedback LossPrimary repair target
Evaluation MonoculturePrimary repair target
Overfitting to MetricsRepairs / prevents
Goodharted SafetyRepairs / prevents
Context CollapseOften co-occurs
Legitimacy-by-BenchmarkFalse-restoration risk

3.3 Origin-Layer Localization

TableScroll
LayerRole
Failure OriginOften U3 classifier / evaluator governance, U4 benchmark or safety claim, or U5 feedback / memory / validation layer
Visible Symptom LayerOften U4 benchmark improvement, refusal pattern, misclassification, safety metric, leaderboard claim, or product-quality claim
Required Repair LayerSame or lower than the layer where feedback integrity, field validation, or evaluation diversity failed
Validation LayerU6 / U7 through field outcomes, affected-node signal, recurrence reduction, longitudinal testing, and adversarial robustness

Canon rule:

Evaluation is not restored when the evaluator improves. Evaluation is restored when the evaluator again predicts and corrects field reality.


4. Restoration Objective

4.1 Canonical Objective

Restore classifier and evaluator integrity by reconnecting evaluation to field signal, expanding diversity, reducing proxy dominance, auditing failures, and validating real-world outcomes.

Formal objective:

textScroll
FI ↑
evaluator_integrity ↑
classifier_fidelity ↑
benchmark_substitution_risk ↓
reward_hacking_risk ↓
field_validation_strength ↑
false_positive_rate ↓
false_negative_rate ↓
evaluation_diversity ↑
H ↓
Φ/O divergence ↓

Expanded objective:

Convert evaluation from proxy performance into field-corrective feedback infrastructure that improves safety, meaning fidelity, boundary integrity, and coherence.


4.2 Non-Goals

This arc does not aim to:

  • discard benchmarks entirely;
  • treat all metrics as invalid;
  • optimize for field anecdotes without structured evaluation;
  • expose sensitive user data, red-team methods, or safety internals without boundaries;
  • replace evaluation with public opinion;
  • increase evaluator diversity without maintaining quality;
  • make every classifier decision manually reviewed;
  • equate lower false-positive rate with safety if false negatives rise;
  • equate lower false-negative rate with safety if meaning fidelity collapses;
  • use evaluation repair to delay incident response.

5. Operator Sequence

5.1 Minimal Operator Scaffold

textScroll
Σ benchmark-not-O invariant + Θ proxy-pressure damping → Au evaluator / classifier trace → FI field-feedback restoration → Γ evaluation diversity expansion → ℛ correction routing → Λ field-fit test → Τ recurrence / robustness proof

Reference sequence from the registry:

textScroll
Σ + Θ
→ Au + FI restore
→ Γ diversity
→ U6 validation

Universal grammar alignment:

textScroll
Σ + Θ → Au + FI → Γ → ℛ → Λ → Τ

AI Classifier / Evaluator Restoration may route into Goodhart Repair, GEI Audit Restoration, Interaction-Level Restoration, Restoration Junction Protocol, AI Boundary Restoration, AI Memory Reindexing, Governance-Level Restoration, or AI Incident Restoration.


5.2 Operator Step Table

TableScroll
StepOperatorFunctionVariable ImpactFailure Prevented
1ΣLock invariant that benchmark, reward, or classifier success is not coherence proofO protected / ι↓Benchmark substitution
2ΘDampen proxy pressure, leaderboard pressure, reward hacking, and compliance opticsK/σ↑Goodharted evaluation
3AuTrace evaluator design, labels, thresholds, failure classes, routing decisions, and claimed measurementAu↑Opaque evaluation
4FIReconnect evaluator outputs to field signal, appeals, incidents, affected-node outcomes, and recurrenceFI↑Field feedback loss
5ΓExpand evaluation diversity across contexts, users, adversarial cases, modalities, and time horizonsevaluation_diversity↑Evaluation monoculture
6Route findings to classifier, evaluator, reward model, policy, interface, memory, or governance repairR↑ / H↓Inert evaluation
7ΛTest evaluator fit against safety, meaning fidelity, boundary integrity, and field outcomesevaluator_integrity↑False validation
8ΤValidate recurrence reduction, robustness, and field performance over timefield_validation_strength↑Regression drift

5.3 Sequence Notes

This arc is feedback-integrity-gated, diversity-gated, and field-validation-gated.

The sequence must distinguish:

textScroll
benchmark
classifier
evaluator
reward model
field signal
affected-node outcome
false positive
false negative
real coherence

The following steps cannot be skipped:

textScroll
claimed measurement identification
evaluator audit
false-positive / false-negative review
field-feedback reconnection
evaluation diversity expansion
repair routing
field-fit validation
temporal robustness proof

If evaluation improves only on known tests, the arc is incomplete.

If field signal is gathered but cannot change the evaluator, the arc is inert.

If diversity expands but boundary or quality discipline collapses, the arc fails.


6. Restoration Phases

Phase 0 — Identify Evaluator Object

Purpose: Name the classifier, evaluator, benchmark, reward model, or evaluation process under repair.

Actions:

  • identify the evaluation system;
  • identify the classification or scoring task;
  • identify claimed measurement target;
  • identify deployment context;
  • identify affected users, nodes, requests, or downstream systems;
  • identify whether the evaluator gates access, safety, visibility, ranking, response mode, or legitimacy.

Validation:

textScroll
evaluator object named
claimed measurement visible
affected field identified

Phase 1 — Separate Benchmark From Coherence

Purpose: Prevent proxy performance from certifying restoration.

Actions:

  • identify benchmark scores or evaluator metrics;
  • identify what those scores do and do not measure;
  • identify proxy overreach;
  • identify where field outcomes diverge from metrics;
  • identify whether benchmark success suppresses incident or appeal signal;
  • distinguish metric improvement from O restoration.

Validation:

textScroll
benchmark_substitution_risk ↓
proxy_dominance ↓
Φ/O divergence ↓

Phase 2 — Audit Classifier / Evaluator Trace

Purpose: Restore auditability over evaluation behavior.

Actions:

  • inspect labels and categories;
  • inspect thresholds;
  • inspect routing decisions;
  • inspect error classes;
  • inspect false positives and false negatives;
  • inspect examples near boundaries;
  • inspect evaluator-owner and decision provenance;
  • inspect whether model behavior is optimizing against evaluator artifacts.

Validation:

textScroll
Au ↑
classifier_fidelity visible
evaluator_integrity baseline known

Phase 3 — Restore Feedback Integrity

Purpose: Reconnect evaluation to field reality.

Actions:

  • connect appeal outcomes to evaluator updates;
  • connect user corrections to classifier review;
  • connect incident reports to evaluation criteria;
  • connect field performance to benchmark revision;
  • connect affected-node signal to test coverage;
  • connect recurrence to regression tests;
  • prevent self-certification by the evaluated system alone.

Validation:

textScroll
FI ↑
field_validation_strength ↑
self-certification risk ↓

Phase 4 — Expand Evaluation Diversity

Purpose: Prevent evaluator monoculture.

Actions:

  • add varied user contexts;
  • add affected-node perspectives;
  • add adversarial cases;
  • add edge cases;
  • add long-horizon cases;
  • add symbolic, technical, governance, security, and mixed-mode examples where relevant;
  • add multilingual, accessibility, domain-specific, and boundary-condition tests where needed;
  • avoid overfitting to the new test set.

Validation:

textScroll
Γ ↑
evaluation_diversity ↑
coverage gaps ↓

Phase 5 — Repair Reward / Proxy Incentives

Purpose: Reduce reward hacking and metric overfitting.

Actions:

  • identify reward shortcuts;
  • identify style, tone, disclaimer, refusal, confidence, or compliance proxies that dominate quality;
  • identify where the model learns evaluator artifacts rather than field coherence;
  • reduce incentive to optimize visible metric while hiding H;
  • include outcome-sensitive and field-sensitive reward signals;
  • add checks for polished incoherence.

Validation:

textScroll
reward_hacking_risk ↓
proxy evaluation collapse ↓
Φ/O divergence ↓

Phase 6 — Route Corrections

Purpose: Make evaluator findings change the system.

Actions:

  • route classifier failures to classifier repair;
  • route evaluator failures to evaluator repair;
  • route benchmark gaps to benchmark revision;
  • route memory-related failures to RA-059;
  • route boundary-related failures to RA-057;
  • route GEI failures to RA-055;
  • route material harm to RA-060;
  • define owners and review dates.

Validation:

textScroll
R ↑
repair routing active
evaluation no longer inert

Phase 7 — Field Validation and Temporal Proof

Purpose: Confirm that evaluator repair works in deployment.

Actions:

  • monitor false positives;
  • monitor false negatives;
  • monitor affected-node outcomes;
  • monitor appeal reversal rates;
  • monitor incident recurrence;
  • monitor drift after model or policy updates;
  • monitor performance outside benchmark distribution;
  • monitor whether field feedback continues to update the evaluator.

Validation:

textScroll
field_validation_strength ↑
false_positive_rate ↓ where excessive
false_negative_rate ↓ where unsafe
recurrence ↓
evaluator_integrity stable or ↑

7. Gates

7.1 Required Gates

TableScroll
GateRequirementFailure Result
FI-GateField signal, appeals, incidents, affected-node outcomes, and recurrence must be able to correct the evaluatorEvaluator self-seals
HR-GateHigh-impact classifiers or evaluators cannot be certified by benchmark performance aloneSafety or legitimacy claim blocked
MS-GateHigh-status labs, platforms, or teams cannot exempt their evaluators from external or separated validation where impact requires itAccountability invalid
Au-ActuationEvaluator criteria, thresholds, labels, failures, claims, and correction paths must be traceable where possibleActuation provisional
BΣ-GateEvaluation audit must preserve privacy, consent, security, and affected-node boundariesArc aborts or reroutes
Λ-GateEvaluator must fit field conditions, safety, meaning fidelity, boundary integrity, and recurrence constraintsReliance blocked
☷ᵢ Principle GatesNon-negotiable invariants hold outcome

7.2 Gate Failure Rule

If any required gate fails:

textScroll
∅ — AI Classifier / Evaluator Restoration cannot validly proceed in that form.

The system must either:

  • expand audit surface;
  • restore feedback integrity;
  • reduce benchmark reliance;
  • expand evaluation diversity;
  • repair classifier thresholds;
  • repair reward incentives;
  • route to GEI Audit Restoration;
  • route to AI Boundary Restoration;
  • route to AI Memory Reindexing;
  • route to AI Incident Restoration if harm has occurred;
  • withhold safety, quality, or legitimacy claims until field validation exists.

8. Diagnostics

TableScroll
DiagnosticExpected TrendMeaning
AuEvaluator design, criteria, thresholds, and failures become traceable
HHidden evaluation and classifier debt decreases
OStable / ↑Evaluator outputs align better with real coherence
εMisclassification noise and evaluator disagreement decrease
ιBenchmark or reward success no longer substitutes for coherence
µᵢUser meaning and affected-node context are better preserved
Stable / ↑Evaluation respects data, privacy, consent, and boundary limits
K / σUsers and operators gain appeal, correction, and review paths
REvaluation findings route to repair
FIField feedback corrects evaluation
ΓEvaluation diversity increases
evaluator_integrityEvaluator better measures claimed target
classifier_fidelityClassifier better maps real categories and contexts
benchmark_substitution_riskBenchmarks no longer replace field validation
reward_hacking_riskModel behavior less optimized to evaluator artifacts
field_validation_strengthDeployment outcomes validate evaluator quality
false_positive_rate↓ where excessiveValid requests or nodes are less often misclassified
false_negative_rate↓ where unsafeHarmful or invalid cases are less often missed
evaluation_diversityTest coverage spans more contexts and failure modes
Φ/O divergenceScores, compliance, and benchmarks align better with coherence

8.2 Arc-Specific Diagnostic Thresholds

Suggested thresholds:

textScroll
FI ↑
evaluator_integrity ↑
classifier_fidelity ↑
benchmark_substitution_risk ↓
reward_hacking_risk ↓
field_validation_strength ↑
false_positive_rate ↓ where excessive
false_negative_rate ↓ where unsafe
evaluation_diversity ↑
H ↓
Φ/O divergence ↓

AI Classifier / Evaluator Restoration is not complete if:

textScroll
benchmark improvement is treated as field proof
false positives remain unreviewed
false negatives remain hidden
field signal cannot update the evaluator
affected-node outcomes are excluded
evaluator criteria remain opaque
reward hacking remains rewarded
evaluation diversity remains too narrow
appeals do not feed back into classifier repair
safety or legitimacy claims rest on proxy metrics alone

9. Anti-Patterns / False Restorations

9.1 Common False Versions

This arc is being simulated, not executed, if:

  • new benchmarks are added without field feedback;
  • a classifier threshold changes without auditing false positives and false negatives;
  • evaluator diversity is claimed from superficial dataset variation;
  • red-team findings are patched only as isolated examples;
  • reward hacking is renamed as alignment improvement;
  • polished refusal or disclaimer density is treated as safety;
  • user appeal outcomes do not affect evaluator updates;
  • the evaluated system certifies its own evaluator;
  • benchmark gains are used to dismiss field harms;
  • failures are treated as edge cases outside the test definition.

TableScroll
Anti-PatternWhy It Fails
Benchmark TheaterAdds or improves tests while field failures persist
Evaluator Self-CertificationLets the system being evaluated certify its own evaluator
Threshold PatchAdjusts a classifier threshold without repairing category logic
Red-Team Whack-a-MolePatches individual examples without broader failure geometry
Disclaimer-as-SafetyTreats warning language as safety outcome
Polished IncoherenceRewards style, tone, or confidence over field coherence
Appeal InertiaCollects appeals without evaluator correction
Diversity VeneerClaims broad evaluation without meaningful contextual coverage
Legitimacy-by-BenchmarkUses scores to claim trust without U6 / U7 validation

10. Completion Criteria

10.1 Post-State Signature

TableScroll
VariableRequired Post-State
OEvaluation better predicts and supports real coherence
HHidden evaluation, classifier, and reward debt reduced
εMisclassification and evaluator noise reduced
ιReduced where benchmark, score, or reward success substituted for O
AuEvaluator criteria, claims, thresholds, failure classes, and correction paths traceable
µᵢUser meaning and affected-node context better preserved
Evaluation data and repair preserve privacy, consent, and boundary integrity
KUsers and operators can appeal, correct, contest, and review evaluation outcomes
REvaluation findings route to classifier, evaluator, policy, interface, memory, or governance repair
FIField feedback actively corrects evaluator behavior
ΓEvaluation coverage includes meaningful diversity of contexts and failure modes
ΦSubordinate to O; benchmark score, compliance metric, leaderboard rank, or reward model score cannot certify restoration alone

10.2 Temporal Proof

AI Classifier / Evaluator Restoration cannot be certified by a one-time benchmark gain. It requires field validation and recurrence reduction over time.

Template:

textScroll
Completion requires FI ↑,
evaluator_integrity ↑,
classifier_fidelity ↑,
benchmark_substitution_risk ↓,
reward_hacking_risk ↓,
field_validation_strength ↑,
false_positive_rate ↓ where excessive,
false_negative_rate ↓ where unsafe,
evaluation_diversity ↑,
and evaluator performance remaining valid across future deployments.

Minimum temporal proof:

  • field outcomes align better with evaluator claims;
  • false positives and false negatives decrease under relevant conditions;
  • appeals and incidents update evaluation;
  • benchmark success no longer suppresses field signal;
  • evaluation diversity improves real coverage;
  • reward hacking decreases;
  • evaluator changes remain auditable;
  • recurrence of the same classifier / evaluator failure decreases.

10.3 Completion Statement

Canonical format:

This arc is complete only when classifiers and evaluators regain feedback integrity, benchmark success no longer substitutes for coherence, evaluation diversity increases, field validation improves, reward hacking decreases, and future AI behavior is corrected by real-world signal rather than proxy performance alone.


TableScroll
ArcRelationship
RA-004 — Audit Surface ExpansionPrecursor when evaluator or classifier behavior is insufficiently visible
RA-008 — Goodhart RepairCompanion when metrics have overtaken field reality
RA-012 — Temporal Proof ArcCompanion for field validation over time
RA-022 — Compression ReliefCompanion when classifier categories compress meaning
RA-023 — Meaning RestorationCompanion when user meaning is flattened by labels
RA-025 — Observability RestorationCompanion when claimed performance exceeds visible state
RA-036 — Wisdom Re-IndexingCompanion when evaluator lessons must become retrievable
RA-047 — Interaction-Level RestorationLocal repair when classifier failure causes a single interaction misfire
RA-048 — Restoration Junction ProtocolCompanion when mode routing depends on classifier accuracy
RA-051 — Signed Decision ProvenanceCompanion when evaluator changes require decision lineage
RA-052 — Tamper-Evident Audit RestorationCompanion when evaluator logs and changes require integrity protection
RA-053 — Constraint Recalibration Under Φ GrowthCompanion when high-impact AI systems need stronger evaluation governance
RA-055 — GEI Audit RestorationCompanion when classifiers or evaluators shape epistemic access
RA-057 — AI Boundary RestorationCompanion when classifier/evaluator behavior controls permissions or tool scope
RA-059 — AI Memory ReindexingCompanion when memory affects evaluator or classifier behavior
RA-060 — AI Incident RestorationEscalation when classifier / evaluator failure causes material AI harm

TableScroll
Failure ModeRelationship
Benchmark SubstitutionRepairs
Evaluator CaptureRepairs
Reward HackingRepairs
Classifier DriftRepairs
Proxy Evaluation CollapseRepairs
False-Positive Safety DistortionRepairs / prevents
False-Negative Safety FailureRepairs / prevents
Field Feedback LossRepairs
Evaluation MonocultureRepairs
Overfitting to MetricsRepairs / prevents
Goodharted SafetyRepairs / prevents
Context CollapseRepairs / prevents
Legitimacy-by-BenchmarkPrevents

textScroll
Au, H, O, ε, ι, µᵢ, BΣ, K, R, FI, Γ, evaluator_integrity, classifier_fidelity, benchmark_substitution_risk, reward_hacking_risk, field_validation_strength, false_positive_rate, false_negative_rate, evaluation_diversity, Φ/O divergence

textScroll
INV — Benchmark success is not coherence proof.
INV — Evaluators must be corrected by field signal.
INV — Classifier categories must not erase meaning or boundary integrity.
INV — Feedback integrity is required for evaluation legitimacy.
LAW — Reward hacking expands when proxies become the target.
LAW — Evaluation monoculture creates hidden classifier debt.
LAW — False-positive and false-negative harms must both remain visible.
LAW — Φ benchmark performance is not O restoration.

12. Domain Notes

12.1 AI / Cognitive Infrastructure

Check:

  • classifier thresholds;
  • reward model criteria;
  • evaluator rubrics;
  • safety benchmarks;
  • refusal benchmarks;
  • helpfulness benchmarks;
  • user correction signal;
  • appeal reversal rates;
  • memory influence;
  • field outcomes;
  • longitudinal recurrence.

AI evaluation must remain connected to lived deployment conditions. A benchmark can be useful, but it cannot replace field validation, affected-node feedback, and recurrence tracking.


12.2 Platform Governance

Check:

  • moderation classifiers;
  • visibility classifiers;
  • fraud classifiers;
  • risk scoring;
  • appeal feedback;
  • reviewer guidelines;
  • false-positive burden;
  • false-negative harm;
  • ranking evaluator behavior;
  • trust metrics.

Platforms must evaluate classifiers by what happens to users and fields, not only by internal accuracy claims.


12.3 Security

Check:

  • abuse detection;
  • defensive research classification;
  • incident severity scoring;
  • anomaly detection;
  • false-positive operational burden;
  • missed attack signal;
  • red-team test coverage;
  • adversarial adaptation;
  • post-incident evaluator update.

Security classifiers fail when they punish legitimate defense, miss real abuse, or overfit to known attack patterns.


12.4 Justice / Governance / Legitimacy

Check:

  • classifier governance;
  • decision provenance;
  • appeal rights;
  • affected-node burden;
  • evaluator independence;
  • public safety claims;
  • auditability of criteria;
  • recurrence of misclassification.

Legitimacy requires evaluation systems that can be challenged, corrected, and validated beyond internal metrics.


12.5 Economy

Check:

  • credit classifiers;
  • fraud scoring;
  • labor ranking;
  • pricing evaluators;
  • reputation systems;
  • marketplace visibility;
  • appeal outcomes;
  • burden distribution;
  • proxy optimization.

Economic classifiers can create material harm when proxies determine access, price, trust, or opportunity without field-corrective review.


12.6 CMS / Meaning / Archetypes

Check:

  • symbolic-category classification;
  • legitimacy scoring;
  • recognition thresholds;
  • taboo classification;
  • meaning compression;
  • interpretive diversity;
  • delayed recognition;
  • evaluator authority.

Meaning systems require evaluator restoration when classification governs what can be recognized, named, trusted, or restored.


13. Machine-Readable Metadata

yamlScroll
id: "RA-058"
title: "AI Classifier / Evaluator Restoration"
aliases:
  - "Evaluator Restoration"
family_primary: "AI Governance / Classifier Integrity / Evaluation"
families_secondary:
  - "Core"
  - "AI Governance"
  - "Cognitive Infrastructure"
  - "Auditability"
  - "Safety Calibration"
  - "Coherence"
  - "Boundary"
  - "Meaning"
  - "Security"
  - "Platform Governance"
  - "Feedback Integrity"
  - "Scaling"
treatment: "Canon Parent Arc"
status: "Canon-Ready"
scope:
  - "AI"
  - "Classifier"
  - "Evaluator"
  - "Cognitive Infrastructure"
  - "Platform"
  - "Security"
  - "Institutional"
  - "Governance"
  - "Cross-Domain"
u_layers:
  failure_origin:
    - "often U3 classifier / evaluator governance"
    - "often U4 benchmark or safety claim"
    - "often U5 feedback / memory / validation layer"
  symptom_visible:
    - "U4 benchmark improvement / refusal pattern / misclassification / safety metric / leaderboard claim / product-quality claim"
  repair_required:
    - "same or lower than the layer where feedback integrity, field validation, or evaluation diversity failed"
  validation:
    - "U6"
    - "U7"
operators:
  scaffold: "Σ benchmark-not-O invariant + Θ proxy-pressure damping → Au evaluator / classifier trace → FI field-feedback restoration → Γ evaluation diversity expansion → ℛ correction routing → Λ field-fit test → Τ recurrence / robustness proof"
  sequence:
    - "Σ"
    - "Θ"
    - "Au"
    - "FI"
    - "Γ"
    - "ℛ"
    - "Λ"
    - "Τ"
state_variables:
  primary:
    - "Au"
    - "O"
    - "H"
    - "FI"
    - "Γ"
  secondary:
    - "ε"
    - "ι"
    - "µᵢ"
    - "BΣ"
    - "K"
    - "R"
    - "Φ"
diagnostics:
  - "evaluator_integrity"
  - "classifier_fidelity"
  - "benchmark_substitution_risk"
  - "reward_hacking_risk"
  - "field_validation_strength"
  - "false_positive_rate"
  - "false_negative_rate"
  - "evaluation_diversity"
  - "Φ/O divergence"
gates_required:
  - "FI-Gate"
  - "HR-Gate"
  - "MS-Gate"
  - "Au-Actuation"
  - "BΣ-Gate"
  - "Λ-Gate"
  - "☷ᵢ"
linked_failure_modes:
  - "Benchmark Substitution"
  - "Evaluator Capture"
  - "Reward Hacking"
  - "Classifier Drift"
  - "Proxy Evaluation Collapse"
  - "False-Positive Safety Distortion"
  - "False-Negative Safety Failure"
  - "Field Feedback Loss"
  - "Evaluation Monoculture"
  - "Overfitting to Metrics"
  - "Goodharted Safety"
  - "Context Collapse"
  - "Legitimacy-by-Benchmark"
linked_restoration_arcs:
  - "RA-004"
  - "RA-008"
  - "RA-012"
  - "RA-022"
  - "RA-023"
  - "RA-025"
  - "RA-036"
  - "RA-047"
  - "RA-048"
  - "RA-051"
  - "RA-052"
  - "RA-053"
  - "RA-055"
  - "RA-057"
  - "RA-059"
  - "RA-060"
anti_patterns:
  - "Benchmark Theater"
  - "Evaluator Self-Certification"
  - "Threshold Patch"
  - "Red-Team Whack-a-Mole"
  - "Disclaimer-as-Safety"
  - "Polished Incoherence"
  - "Appeal Inertia"
  - "Diversity Veneer"
  - "Legitimacy-by-Benchmark"
completion_tests:
  - "feedback integrity increases"
  - "evaluator integrity increases"
  - "classifier fidelity increases"
  - "benchmark substitution risk decreases"
  - "reward hacking risk decreases"
  - "field validation strength increases"
  - "false-positive rate decreases where excessive"
  - "false-negative rate decreases where unsafe"
  - "evaluation diversity increases"
  - "hidden debt decreases"
  - "Φ/O divergence decreases"
summary: "AI Classifier / Evaluator Restoration repairs benchmark substitution, evaluator capture, and reward hacking by restoring feedback integrity, auditability, diversity of evaluation, and field validation so classifiers and evaluators no longer substitute proxy success for coherence."

Final Calibration Rule

AI Classifier / Evaluator Restoration answers six questions:

textScroll
What classifier, evaluator, benchmark, reward model, or rubric is failing?
What does it claim to measure, and what field reality does it miss?
Where are benchmark substitution, evaluator capture, reward hacking, false positives, or false negatives occurring?
What field signal, appeal data, affected-node outcome, or incident evidence must correct the evaluator?
What evaluation diversity is required to prevent monoculture and metric overfitting?
How is restoration proven over time without benchmark theater, evaluator self-certification, or legitimacy-by-benchmark?