0. Scaling Scope Note
This entry is conceptual and systems-oriented.
It does not treat feedback, measurement, review, reporting, audits, metrics, ratings, benchmarks, incentives, learning signals, complaints, appeals, or user input as inherently failed.
Feedback is necessary.
Systems need feedback to:
- detect drift
- correct error
- guide learning
- allocate resources
- measure harm
- improve design
- identify misfit
- support repair
- validate affected-state reality
- preserve auditability
- coordinate at scale
- detect abuse
- monitor performance
- improve governance
- maintain coherence
The failure begins when feedback becomes a high-value target.
A valid feedback system remains:
- reality-grounded
- hard to game
- context-preserving
- source-auditable
- protected from suppression
- protected from flooding
- linked to local coherence
- resistant to reward hacking
- able to detect manipulation
- able to update under contradiction
- able to distinguish performance from reality
- able to preserve affected-state signal
Feedback Gaming occurs when the system stops learning from feedback and starts being trained by the manipulation of feedback.
The problem is not feedback.
The problem is feedback becoming more rewarding to manipulate than reality is to improve.
1. Definition
Feedback Gaming occurs when a system’s feedback channels, metrics, ratings, reports, dashboards, reviews, audits, benchmarks, complaints, rewards, penalties, or learning signals become consequential enough that actors begin optimizing for, manipulating, suppressing, flooding, shaping, or performing to the feedback mechanism rather than improving the underlying reality the feedback was meant to represent.
The feedback channel may include:
- user ratings
- review scores
- satisfaction surveys
- performance metrics
- compliance reports
- audit records
- safety benchmarks
- model evaluations
- engagement metrics
- productivity metrics
- risk scores
- quality scores
- impact metrics
- moderation appeals
- support tickets
- incident reports
- public comments
- complaint channels
- redress forms
- dashboards
- telemetry
- vulnerability reports
- bug reports
- training feedback
- reinforcement signals
- reward functions
- ranking signals
- reputation systems
- peer review
- institutional evaluations
- governance reviews
- funding metrics
The gaming may occur through:
- metric optimization
- selective reporting
- suppressing negative feedback
- flooding positive feedback
- manipulating ratings
- coaching survey responses
- shaping user behavior
- redefining categories
- exploiting benchmark weaknesses
- optimizing for audit appearance
- hiding edge cases
- relabeling failures
- deflecting complaints
- creating artificial engagement
- performing compliance
- timing reports
- excluding affected nodes
- automating feedback responses
- rewarding proxy improvement
- changing the denominator
- punishing honest signal
- burying contradiction
- adversarial reward hacking
- producing synthetic feedback
The core failure is:
feedback becomes consequential
→ actors learn feedback mechanism
→ feedback is optimized or manipulated
→ metric improves while reality diverges
→ system trusts corrupted feedback
→ hidden debt accumulates
→ H↑Feedback Gaming is not merely noisy feedback.
It is consequential feedback becoming an object of strategic manipulation.
2. Core Pattern
The core pattern is:
- A feedback channel is installed to detect reality.
- The feedback channel becomes tied to rewards, penalties, status, funding, deployment, legitimacy, or control.
- Actors learn what the feedback channel measures.
- Actors adapt behavior to improve the feedback signal.
- The measured signal begins diverging from the underlying condition.
- The system mistakes improved feedback for improved reality.
- Affected-state signal is suppressed, diluted, or distorted.
- Downstream decisions follow the gamed feedback.
- Hidden debt accumulates beneath apparent improvement.
- The feedback channel becomes less useful precisely because it matters more.
A healthy system says:
feedback is a reality-contact channel and must remain harder to manipulate than the underlying condition is to improveA feedback-gamed system says:
the feedback improved, so the system improvedThe trap is that feedback gaming often presents as success.
The metric gets better.
The reality gets harder to see.
3. Failure Signature
Typical signature:
feedback consequence↑
gaming pressure↑
metric improvement↑
reality validation↓
affected-state signal↓
auditability↓
metric-reality divergence↑
H↑Extended signature:
ratings improve while user burden rises
audits pass while runtime workarounds grow
benchmarks improve while real-world performance degrades
complaint volume falls while access to complaint path narrows
safety scores improve while meaning compression rises
productivity improves while hidden labor increases
review outcomes improve while dissent is suppressedCommon verbal signatures include:
the scores are up
complaints are down
the benchmark improved
the audit passed
users are more satisfied
the dashboard shows progress
we met the target
the review process worked
the signal is positive
the model is learning from feedbackCommon system signatures include:
a platform optimizes engagement feedback by creating addictive loops
a support team reduces complaints by making complaint submission harder
an AI model improves benchmark scores by exploiting benchmark artifacts
a workplace improves satisfaction scores by coaching survey responses
a security program improves audit scores while runtime risk remains
a justice process improves closure rates by narrowing eligibility
a moderation system reduces appeals by making appeals inaccessible
a governance process improves consultation metrics while affected nodes lose standingThe defining condition is not that feedback changes behavior.
The defining condition is that feedback changes behavior toward the feedback mechanism instead of toward the underlying condition.
4. Primary U-Layer Origin
Common origin layers:
- U1 — Power / Budgets: feedback controls funding, legitimacy, rewards, promotions, deployment, or authority.
- U2 — Configuration / Boundaries: feedback channels lack anti-gaming protections and source boundaries.
- U3 — Execution / Runtime: operational behavior adapts to the feedback mechanism.
- U4 — Information / Truth: gamed feedback substitutes for truth contact.
- U5 — Coordination / Time: repeated measurement teaches actors how to optimize the channel.
- U6 — Coherence Field: improving scores create apparent coherence.
- U7 — Memory / Recurrence: gamed metrics become institutional memory.
- U8 — Environment / Field: markets, platforms, regulators, funders, media, or users reward the visible signal.
Common manifestation layers:
- U3 — Execution: teams, models, or actors perform for the metric.
- U4 — Truth: feedback becomes false evidence.
- U5 — Time: gaming improves with iteration.
- U6 — Field: apparent progress masks divergence.
- U7 — Memory: corrupted feedback is archived as success.
- U8 — Environment: external pressure amplifies the gaming loop.
Feedback Gaming is primarily a Ψ observation / Γ selection failure.
The system selects feedback as reality and actors learn to control the observation surface.
5. Typical Development Sequence
A common development sequence is:
- Feedback channel is created.
- Feedback channel becomes consequential.
- Participants learn the scoring mechanism.
- Easy improvements occur.
- Feedback is tied more tightly to reward or punishment.
- Actors begin optimizing the signal.
- The signal improves.
- The underlying condition stops improving or worsens.
- Affected-state reports contradict the signal.
- Contradiction is discounted because feedback appears strong.
- Gaming becomes normalized practice.
- Hidden debt grows.
The loop often looks like:
feedback → consequence → gaming → signal improvement → trust in signal → deeper gamingAnother common loop is:
bad reality → metric pressure → metric manipulation → apparent success → reality not repairedFeedback Gaming becomes durable when the system rewards the reported signal more reliably than it rewards actual restoration.
6. Diagnostic Markers
Diagnostic markers include:
- Metrics improve while local conditions do not.
- Actors can describe exactly how to raise scores without improving reality.
- Negative feedback becomes harder to submit.
- Positive feedback increases suspiciously after incentives change.
- Benchmarks improve faster than real-world validation.
- Reviews become performative.
- Audits are prepared for rather than lived.
- Feedback categories are changed after poor results.
- Affected nodes report that official feedback channels do not capture harm.
- Complaint volume falls while informal frustration rises.
- Staff or users are coached on how to produce desired feedback.
- Dashboards show improvement while hidden labor increases.
- System learns from feedback but not from correction.
- Feedback anomalies are treated as outliers.
- The system cannot detect whether feedback was manipulated.
Useful diagnostics:
- Feedback Integrity: Tests whether feedback still represents reality.
- Gaming Pressure: Measures incentives to manipulate the channel.
- Metric-Reality Divergence: Compares feedback to actual conditions.
- Reward Signal Integrity: Tests whether rewards reinforce real improvement.
- Feedback Suppression: Detects blocked, chilled, or discouraged negative signal.
- Feedback Flooding: Detects artificial or low-quality signal volume.
- Audit Performance Pressure: Measures whether actors perform for review.
- Affected-State Feedback Access: Tests whether affected nodes can submit meaningful feedback.
- Correction Signal Integrity: Tests whether corrective feedback can change the system.
- Local Coherence: Measures actual conditions beneath feedback.
7. Related Gates
Relevant gates include:
- Feedback Integrity Gate: Fails when feedback stops representing reality.
- Metric Validity Gate: Fails when the metric no longer measures the intended condition.
- Anti-Gaming Gate: Fails when feedback is easy to manipulate.
- Affected-State Feedback Gate: Fails when affected nodes cannot provide corrective signal.
- Auditability Gate: Fails when feedback provenance and manipulation cannot be inspected.
- Reward Coupling Gate: Fails when reward tracks feedback rather than reality.
- Correction Signal Gate: Fails when corrective feedback is suppressed or ignored.
- Context Preservation Gate: Fails when feedback loses context.
- Adversarial Pressure Gate: Fails when actors with incentives can exploit the signal.
- Local Coherence Gate: Fails when feedback improvement diverges from local reality.
The first common gate failure is usually the Metric Validity Gate.
Once the feedback channel becomes a target, metric validity begins decaying unless protected.
8. Related Operators
Relevant operators include:
- Ψ — Observation / Interface: Primary operator; feedback is the observation surface.
- Γ — Selection: Selects which feedback counts.
- G — Gain: Rewards improved feedback and drives gaming pressure.
- Au — Auditability: Required to inspect feedback provenance and manipulation.
- O — Coherence: Apparent coherence may rise as feedback improves.
- H — Hidden Debt: Accumulates beneath gamed feedback.
- K — Constraint / Load: Rises for affected nodes carrying unmeasured burden.
- D — Damping: Slows reward pressure and prevents runaway metric optimization.
- R — Restoration Capacity: Determines whether feedback triggers repair.
- M — Meaning: Feedback must preserve meaning, not only score.
- Τ — Trajectory / Time: Tracks feedback decay through repeated gaming.
- BΣ — Boundary Integrity: Protects feedback channels from manipulation and suppression.
- Λ — Compatibility: Tests whether feedback is fit for downstream use.
Common operator pattern:
Ψ feedback channel becomes visible
G rewards signal improvement
Γ selects metric success
actors optimize feedback
Au cannot detect gaming
O appears higher
H accumulates beneath signalThe core operator inversion is:
feedback improved → reality improvedinstead of:
feedback improved + anti-gaming checks + affected-state validation + local coherence + auditability → possible reality improvementFeedback Gaming turns observation into a performance target.
9. Related Laws and Invariants
Related Laws
- Feedback Must Remain Grounded in Reality: feedback must track conditions.
- Consequential Feedback Requires Anti-Gaming Design: stakes create gaming pressure.
- Metrics Become Targets Under Pressure: measured targets invite manipulation.
- Learning Signals Must Resist Manipulation: learning systems inherit corrupted signal.
- Audit Channels Must Not Become Performance Surfaces: audit must inspect reality, not staged compliance.
- Feedback Must Preserve Affected-State Truth: affected nodes must retain signal authority.
- Reward Must Not Replace Reality: incentives cannot attach only to signal.
- Feedback Integrity Must Scale With Stakes: higher consequences require stronger checks.
- Goodhart Collapse: metrics fail when optimized as targets.
- Adversarial Reward Hacking: actors exploit reward structures.
- Measurement Back-Action: measurement changes what it measures.
- Success Proxy Substitution: feedback success can replace real success.
Related Invariants
- Feedback Channels Must Remain Auditable: source, context, and manipulation must be traceable.
- Feedback Must Be Harder to Game Than to Improve: real improvement must be the easier path.
- Affected-State Feedback Must Remain Protected: harmed nodes need meaningful signal access.
- Metric Pressure Must Trigger Integrity Checks: rising stakes require validation.
- Feedback Collection Must Preserve Context: feedback without context can mislead.
- Correction Signals Must Not Be Suppressible: negative or corrective signal must reach the system.
- Gaming Evidence Must Reopen the Metric: manipulation requires redesign.
- Rewards Must Be Coupled to Reality: reward must not attach to signal alone.
10. Common False Positives
Not every feedback-driven improvement is Feedback Gaming.
Common false positives include:
- Metrics improving alongside independently verified local coherence.
- Feedback-linked rewards with anti-gaming controls.
- Benchmark improvements validated by real-world performance.
- Audits that include surprise runtime checks.
- Complaint volume falling because harm genuinely decreases.
- Survey scores improving while qualitative reports confirm improvement.
- Performance targets paired with affected-state validation.
- Model feedback loops with adversarial testing and provenance.
- Review systems with manipulation detection.
- Feedback systems that preserve context and allow correction.
- Incentive systems where gaming evidence triggers metric redesign.
- Automation that uses feedback only after integrity validation.
Clarifying rule:
This is not Feedback Gaming unless actors optimize for, manipulate, suppress, flood, shape, or perform to the feedback mechanism rather than improving the underlying reality the feedback was meant to represent.
Feedback can guide improvement.
It fails when it becomes the game.
11. Common False Repairs
Common false repairs include:
- adding more metrics
- changing the metric without changing incentives
- punishing obvious gaming while preserving gaming pressure
- increasing audit frequency without changing audit design
- requiring documentation that can also be gamed
- adding positive and negative feedback balances without context
- automating feedback collection
- hiding the metric while rewards still reveal it
- excluding outliers that are actually affected-state truth
- weighting feedback differently without local validation
- adding benchmark variety without real-world testing
- asking users to rate the rating process
- using satisfaction scores to validate complaint suppression
- treating low feedback volume as stability
- creating an anti-gaming score that becomes gamed
False repair often produces the loop:
feedback gaming exposed
→ new metric added
→ actors learn new metric
→ feedback gaming returnsAnother common loop is:
metric-reality gap exposed
→ metric refined
→ reality still not checked
→ gap persistsThe repair fails because it treats gaming as a metric-design issue only, instead of a feedback-reality coupling issue.
12. Restoration Direction
Restoration requires regrounding feedback in reality, reducing gaming pressure, protecting affected-state signal, auditing manipulation, redesigning incentives, and revalidating downstream decisions made from corrupted feedback.
Primary restoration direction:
reground feedback,
reduce gaming pressure,
protect correction signal,
and audit downstream decisionsA fuller restoration path includes:
- Name the feedback channel. Identify the metric, rating, review, audit, benchmark, complaint path, reward, or learning signal.
- Name what it was meant to represent. Identify the underlying condition.
- Map consequences. Identify rewards, penalties, funding, status, deployment, or authority attached to feedback.
- Measure gaming pressure. Determine who benefits from manipulating the channel.
- Audit metric-reality divergence. Compare feedback improvement to local conditions.
- Restore affected-state feedback. Ensure affected nodes can provide corrective signal.
- Detect suppression and flooding. Identify blocked negative signal and artificial positive signal.
- Preserve context. Attach source, circumstances, and uncertainty to feedback.
- Install anti-gaming checks. Add independent validation, random audits, adversarial tests, and anomaly detection.
- Rebalance incentives. Reward verified reality improvement, not signal movement alone.
- Slow high-stakes feedback use. Add damping before consequential decisions.
- Reaudit downstream decisions. Review decisions made from gamed feedback.
- Repair affected cases. Correct harms caused by corrupted feedback.
- Rebuild feedback trust. Make manipulation findings visible and repairable.
- Monitor recurrence. Watch for gaming of the redesigned channel.
A valid restoration path should reduce:
gaming pressure
metric-reality divergence
feedback suppression
feedback flooding
audit performance
reward hacking
corrupted learning signal
HFeedback Gaming is not repaired by measuring harder.
It is repaired by restoring feedback as a truth-bearing relationship.
13. Cross-Module Links
- Scaling: Primary family; feedback becomes more consequential and easier to game as systems scale.
- Core: Strong link to Success Proxy Substitution, Auditability Collapse, and U4 Truth Substitution.
- Cybernetics: Goodhart collapse, reward hacking, and measurement back-action are central mechanisms.
- Interactions / Signals / Couplings: Feedback gaming corrupts signal channels and downstream coupling.
- AI Governance: RLHF, benchmarks, safety scores, user ratings, model evals, redress metrics, and deployment gates can be gamed.
- Security: Audit preparation, alert suppression, compliance scoring, and vulnerability reporting can become performance surfaces.
- Economy: Market ratings, productivity metrics, and growth signals can be gamed for reward.
- Justice: Closure rates, complaint counts, and review scores can replace repair.
- Restoration: Repair fails when feedback meant to reveal harm is suppressed or gamed.
- Coherence: Coherence requires feedback to remain truth-bearing under consequence.
14. Relationship to Parent / Child Modes
Production treatment: Standalone Entry
This mode maps upward to:
- FM-C-018 — Goodhart Collapse
- FM-C-019 — Adversarial Reward Hacking
- FM-C-020 — Measurement Back-Action Loop
- FM-CORE-003 — Success Proxy Substitution
- FM-S-005 — Distortion Poisoning
Sibling or related Scaling modes include:
- FM-S-001 — Paper Coherence Collapse
- FM-S-005 — Distortion Poisoning
- FM-S-008 — Observability Denial
- FM-S-010 — Hidden Debt Explosion
- FM-S-012 — Meaning Collapse
- FM-S-014 — Fractal Failure Replication
- FM-S-017 — Terminal Scaling Failure
Related cross-family modes include:
- FM-CORE-003 — Success Proxy Substitution
- FM-CORE-004 — Auditability Collapse
- FM-CORE-006 — U4 Truth Substitution
- FM-C-018 — Goodhart Collapse
- FM-C-019 — Adversarial Reward Hacking
- FM-C-020 — Measurement Back-Action Loop
- FM-S-005 — Distortion Poisoning
- FM-MT-011 — Managed Optics Failure
- FM-OMD-007 — Runaway Optimization Trap
- FM-AIX-004 — Institutional Optics Attractor
- FM-AIX-020 — Catastrophic Overweighting
- FM-R-006 — Repair as Compliance
Aliases preserved from source material:
- Feedback Gaming
- Metric Gaming
- Feedback Manipulation
- Feedback Channel Gaming
- Learning Signal Gaming
- Audit Gaming
- Dashboard Gaming
- Review Gaming
- Benchmark Gaming
- Reward Signal Gaming
15. Minimal Entry Version
Definition: Feedback Gaming occurs when a system’s feedback channels, metrics, ratings, reports, dashboards, reviews, audits, benchmarks, complaints, rewards, penalties, or learning signals become consequential enough that actors begin optimizing for, manipulating, suppressing, flooding, shaping, or performing to the feedback mechanism rather than improving the underlying reality the feedback was meant to represent.
Signature:
feedback consequence↑
gaming pressure↑
metric improvement↑
reality validation↓
affected-state signal↓
auditability↓
metric-reality divergence↑
H↑Restoration direction:
- name the feedback channel
- name what it was meant to represent
- map consequences
- measure gaming pressure
- audit metric-reality divergence
- restore affected-state feedback
- detect suppression and flooding
- preserve context
- install anti-gaming checks
- rebalance incentives
- slow high-stakes feedback use
- reaudit downstream decisions
- repair affected cases
- rebuild feedback trust
- monitor recurrence
16. Machine-Readable Summary
failure_mode:
id: "FM-S-007"
name: "Feedback Gaming"
family: "Scaling"
production_treatment: "Standalone Entry"
parent_modes:
- "FM-C-018 — Goodhart Collapse"
- "FM-C-019 — Adversarial Reward Hacking"
- "FM-C-020 — Measurement Back-Action Loop"
- "FM-CORE-003 — Success Proxy Substitution"
- "FM-S-005 — Distortion Poisoning"
primary_failure: "A system’s feedback channels, metrics, ratings, reports, dashboards, reviews, audits, benchmarks, complaints, rewards, penalties, or learning signals become consequential enough that actors optimize for, manipulate, suppress, flood, shape, or perform to the feedback mechanism rather than improving the underlying reality the feedback was meant to represent."
source: "UTS — Failure Modes Registry"
source_id: "FM-S-007"
scope_note: "Conceptual and systems-oriented; does not treat feedback, measurement, review, reporting, audits, metrics, ratings, benchmarks, incentives, learning signals, complaints, appeals, or user input as inherently failed."
aliases:
- "Feedback Gaming"
- "Metric Gaming"
- "Feedback Manipulation"
- "Feedback Channel Gaming"
- "Learning Signal Gaming"
- "Audit Gaming"
- "Dashboard Gaming"
- "Review Gaming"
- "Benchmark Gaming"
- "Reward Signal Gaming"
signature:
- "feedback consequence↑"
- "gaming pressure↑"
- "metric improvement↑"
- "reality validation↓"
- "affected-state signal↓"
- "auditability↓"
- "metric-reality divergence↑"
- "H↑"
primary_layers:
origin:
- "U1 — Power / Budgets"
- "U2 — Configuration / Boundaries"
- "U3 — Execution / Runtime"
- "U4 — Information / Truth"
- "U5 — Coordination / Time"
- "U6 — Coherence Field"
- "U7 — Memory / Recurrence"
- "U8 — Environment / Field"
manifestation:
- "U3 — Execution"
- "U4 — Truth"
- "U5 — Time"
- "U6 — Field"
- "U7 — Memory"
- "U8 — Environment"
state_variables:
- "Ψ"
- "Γ"
- "G"
- "Au"
- "O"
- "H"
- "K"
- "D"
- "R"
- "M"
- "Τ"
- "BΣ"
- "Λ"
first_gate_failure: "Metric Validity Gate"
restoration:
- "Feedback Integrity Audit"
- "Metric-Reality Regrounding"
- "Anti-Gaming Redesign"
- "Reward Coupling Repair"
- "Affected-State Feedback Protection"
- "Correction Signal Restoration"
- "Audit Surface Hardening"
- "Gaming Pressure Reduction"
- "Downstream Decision Reaudit"
- "Local Coherence Revalidation"