“Human in the loop” has become one of the most reassuring phrases in technology.

The model will not act alone. The operator can take over. An expert will check the answer. A reviewer must approve the decision. A pilot, clinician, analyst, editor, safety driver, or supervisor remains responsible.

Then the human loop fails.

Sometimes the person is distracted after hours of uneventful monitoring. Sometimes automation hands back control at the worst possible moment. Sometimes an interface presents dozens of alarms without making the underlying problem intelligible. Sometimes the reviewer trusts an authoritative answer, signs it, and moves on. Sometimes the person sees the problem clearly but lacks the authority or organizational protection to stop it.

The presence of a human proves almost nothing about the quality of oversight.

That is the opening for Human Runtime—but also its boundary. HRT can make the operation of the human node observable. It cannot turn bad governance, defective automation, suppressed evidence, or an impossible handover into a safe system.

The recurring failure pattern

The classic automation literature predicted this problem decades ago.

Lisanne Bainbridge's 1983 paper “Ironies of Automation” observed that automation removes people from routine control while leaving them responsible for the hardest abnormal situations. Endsley and Kiris later demonstrated the out-of-the-loop performance problem: passive supervision can reduce situation awareness and impair takeover performance when automation fails. Parasuraman and Riley's widely cited framework described automation use, misuse, disuse, and abuse, including overreliance, complacency, alarm distrust, and inappropriate allocation of responsibility.

The pattern is now visible across transport, aviation, industrial control, public administration, and generative AI:

The machine performs the ordinary work. The human inherits the exceptional case—often with degraded context, little practice, incomplete information, and very little time.

Calling that arrangement “human in the loop” does not make it meaningful human control.

Nine times the human loop broke

Not every case below was marketed with the modern HITL label. Each nevertheless used the same safety architecture: a machine, process, or institution depended on a person to detect, verify, approve, correct, or recover from a failure.

Boeing 737 MAX: the assumed recovery

The 737 MAX safety case assumed pilots would recognize and respond appropriately to unintended MCAS activation. After two crashes killed 346 people, the NTSB concluded that those assumptions did not adequately account for the effect of multiple flight-deck alerts on pilot recognition and response. A Boeing executive later acknowledged that assumptions about the human–machine interaction had been wrong.

The NTSB safety recommendation report, the congressional investigation, and Reuters' account of Boeing testimony make this one of the clearest examples of human performance being treated as a design assumption rather than an observed variable.

HRT could have been valuable before deployment: representative simulator traces could test what ordinary crews noticed, how they interpreted competing indications, and whether they acted within the assumed time. It would not replace redundant sensing, safe control laws, training, or intelligible alerts.

Air France 447: the failed handover

When temporary unreliable airspeed caused the autopilot to disconnect, the crew of Air France 447 did not correctly identify or recover from the resulting stall. All 228 people aboard died. The French BEA investigation documents the sequence; a later paper, “Learning from AF447: Human-machine interaction,” analyzes the combination of degraded automation, unclear information, human error, and suboptimal teamwork.

HRT could make this kind of handover testable in simulation: which instruments were sampled, when the pilots' mental models diverged, which alarms were understood, and whether an intervention improved the next attempt. In the live event, measurement alone would not supply the missing diagnosis or create more time.

USS Vincennes: a team saw the wrong threat

In 1988, USS Vincennes shot down Iran Air Flight 655, killing 290 people. The ship's Aegis system recorded the civilian aircraft climbing, yet the command team was told it was descending into an attack profile. A military identification signal was incorrectly associated with the aircraft, and contrary evidence failed to break the team's interpretation during a high-pressure surface engagement.

The U.S. Naval History and Heritage Command provides a detailed account of the event. The U.S. Naval Institute later summarized the lesson directly: “The Human-Machine Team Failed Vincennes.”

An HRT trace might expose narrowing information search, unchallenged team convergence, inconsistent evidence sampling, and overload. It could support a forced independent review. It could not determine whether an aircraft was hostile, and it must never become a biological targeting or intent score.

Patriot fratricides: supervising a fast machine

During the 2003 Iraq war, Patriot batteries mistakenly engaged friendly aircraft. A U.S. Department of Defense human-systems case study described the causal pattern as “undisciplined automation”: functions were automated without sufficient regard for operator performance, training, or the consequences of reducing human involvement.

This is a strong HRT research scenario in simulation. Did the operator independently verify the classification, or merely confirm the machine? What evidence was sampled? How did workload, time pressure, prior false alarms, and team communication affect the decision? The appropriate early use is evaluation and training, not live weapons control.

Uber Tempe: the absent safety driver

In 2018, an Uber developmental self-driving vehicle struck and killed a pedestrian in Tempe, Arizona. The safety driver was visually distracted. The NTSB also identified inadequate risk assessment, ineffective operator oversight, failure to address automation complacency, and an inadequate safety culture. The findings are available in the official NTSB report, with an accessible reconstruction from IEEE Spectrum.

This is one of HRT's clearest use cases. Gaze, interaction, vigilance history, and intervention timing could reveal that the nominal safety layer was no longer operational. But the correct response might be to stop the test. Another alert to an already disengaged operator is not necessarily a safety intervention.

Tesla Autopilot: monitoring hands instead of readiness

In the fatal 2018 Mountain View crash, the NTSB found that Autopilot steered toward a highway barrier and the driver did not respond because of distraction and overreliance. Ineffective monitoring of driver engagement contributed to the crash. The NTSB report and Reuters coverage connect the accident to the broader problem of automation complacency.

HRT could improve on crude proxies such as steering-wheel contact by measuring actual road monitoring and takeover readiness. It still cannot make an inadequate operational design domain safe. A system that routinely creates an unavailable backup driver is badly allocated, even if that unavailability is measured precisely.

Three Mile Island: alarms without understanding

At Three Mile Island in 1979, a relief valve became stuck open while the control-room indicator implied that it was closed. Alarms rang, warning lights flashed, other instruments gave inadequate or misleading information, and operators took actions that reduced cooling. The result was the most serious accident in U.S. commercial nuclear power history, although official reviews found no detectable public health effects.

The U.S. Nuclear Regulatory Commission's accident backgrounder describes the sequence, while the technical literature includes a human-factors evaluation of the control room and operator performance.

HRT could reconstruct alert uptake, scanning, procedural action, communication, and recovery during training and interface evaluation. It cannot repair deceptive instrumentation or replace engineered containment and automatic protection.

Challenger: the human spoke and the institution refused to listen

Before the Challenger launch, engineers objected to launching in unusually cold conditions. Management reversed the contractor's recommendation, and critical concerns did not reach senior decision-makers in a clear and protected form. Seven crew members died. The Rogers Commission found incomplete and sometimes misleading information, conflict between engineering evidence and management judgment, and a structure that allowed safety issues to bypass key leaders.

This is principally a governance failure, not a cognitive-state failure. A durable decision trace could preserve dissent, evidence provenance, approval reversals, and responsibility. Eye tracking or a workload score would not repair schedule pressure, customer pressure, or the absence of protected technical authority.

The Post Office Horizon scandal: review became institutional confirmation

The UK Post Office pursued sub-postmasters using apparent financial shortfalls generated by the Horizon computer system. Investigations were poor, material facts about system reliability were not fully disclosed, and hundreds of convictions were later quashed. The Criminal Cases Review Commission calls it the biggest single series of wrongful convictions in UK legal history. The statutory inquiry is publishing its reports and findings, while BBC reporting has documented errors and internal warnings.

HRT might show whether a reviewer inspected contrary evidence or simply accepted the system output. But this was not mainly a failure of attention. Concealment, conflicts of interest, coercive policy, weak investigation, and institutional resistance require law, governance, independence, and accountability.

LLMs create new human-loop failures

Generative AI makes the old automation problem unusually visible. LLM output is fluent, variable, difficult to verify, and frequently delivered at a speed that makes thorough human review uneconomic. The human is asked to supervise more material precisely because the model can generate more material.

The result is not merely hallucination. It is a failure of the combined production-and-review system.

The reviewer signs something that was never verified

In Mata v. Avianca, lawyers submitted nonexistent cases and fabricated quotations generated by ChatGPT. The court sanctioned the lawyers after the material was not properly verified and continued to be defended after its authenticity was challenged. The sanctions opinion is a nearly perfect human-loop record: generation, misplaced trust, failed verification, challenge, renewed reliance on the same model, and escalation of the error.

CNET likewise published AI-assisted financial explainers that it said had been edited and fact-checked, then appended substantial corrections after errors were exposed. The Washington Post raised the central question: did the authoritative machine voice cause editors to lower their guard?

HRT could measure whether the cited sources were opened, how long verification took, which claims received attention, and whether throughput pressure changed review behaviour. It could not determine legal or financial truth without authoritative external evidence.

The organization deploys a bot without owning its answers

Air Canada's chatbot incorrectly described the airline's bereavement policy. A Canadian tribunal held the airline responsible for failing to take reasonable care to ensure the accuracy of information on its website. The decision, Moffatt v. Air Canada, turned a modest refund dispute into an important lesson: an organization cannot outsource responsibility to the conversational interface it deploys.

New York City's MyCity chatbot offered businesses incorrect and sometimes unlawful guidance. The initial problem was documented by The Markup. A later New York City Comptroller audit found weaknesses in project management, contract oversight, consistency, and the handling of ungrounded answers.

These are not cases where a more attentive end user should carry the burden. HRT is relevant inside the publishing, testing, monitoring, escalation, and correction workflow—not as surveillance of the citizen receiving bad advice.

Human feedback trains the model toward the wrong thing

Human participation can fail before deployment too.

Reinforcement learning from human feedback converts judgments into training signals. But a preference is not the same as truth, safety, or usefulness. The ICLR paper “Human Feedback is not Gold Standard” found that preference judgments under-represent factuality and can be influenced by assertiveness. Anthropic's research on sycophancy found that humans and preference models sometimes favor convincingly written answers that agree with the user over correct answers. OpenAI's work on reward-model overoptimization demonstrates the deeper Goodhart problem: optimizing an imperfect proxy can eventually reduce the ground-truth performance it was meant to improve.

The original InstructGPT work is candid about another limit. The system reflects the preferences of selected labelers and researchers; those labelers were primarily English-speaking, and average preferences may be inappropriate when outputs disproportionately affect a minority group. The paper and limitations make clear that “human feedback” is not one universal human value signal.

HRT cannot solve value pluralism. It could qualify the production of feedback: reviewer expertise, task context, information inspected, disagreement, uncertainty, correction history, time pressure, and outcome. A hurried preference click and a carefully researched expert judgment should not enter the training pipeline as equivalent evidence.

The model does not know when to ask

Even if a knowledgeable person is available, an agent must recognize when escalation is necessary. Scale's 2026 HiL-Bench adds missing, contradictory, or unknowable information to software and database tasks. Models that perform strongly with complete information lose much of that performance when they must decide whether and how to ask a human for help.

This creates two coupled runtime questions:

  • Does the agent know when to escalate?
  • Is the human receiving the escalation able to provide dependable help?

Most agent benchmarks measure the first weakly and the second not at all. HRT belongs at that junction.

The hidden human cost of model safety

Humans who label toxic content are also part of the training system. A TIME investigation reported that outsourced workers in Kenya reviewed graphic descriptions of abuse, violence, and other traumatic material to help build an OpenAI toxicity detector, with workers describing psychological harm and pay below two dollars per hour. TIME's investigation and subsequent Guardian reporting expose a different human-loop failure: the model becomes safer for users by transferring cognitive and emotional cost to largely invisible workers.

HRT could support consent-aware exposure tracking, workload limits, recovery periods, task rotation, intervention evidence, and longitudinal health safeguards. It must not become another productivity instrument used to intensify harmful work.

Beyond the artifact: what existing human-loop systems miss

An established market already argues that AI needs people.

Training-data and evaluation providers including Surge AI, Appen, Sama, Labelbox, iMerit, Scale AI, Invisible Technologies, Toloka, and others sell access to generalist or domain-expert judgment. Evaluation platforms combine human ratings with deterministic checks and model-based evaluators. Agent frameworks provide pause, approval, editing, rejection, and escalation mechanisms.

Their public arguments are consistent:

  • Frontier models need increasingly expert feedback.
  • Rare cases and subjective tasks require human judgment.
  • Human review helps find hallucinations, data gaps, and regressions.
  • Quality control, adjudication, and reviewer calibration improve labels.
  • Production systems need monitoring in addition to pre-deployment evaluation.

Examples include Surge's case for domain-expert RLHF and red teaming, Appen's account of expert preference tuning for Cohere, Sama's use of reviewers to flag hallucinations and rewrite captions, and Labelbox's workflow for routing edge cases to internal experts. Humanloop describes subject-matter experts as a way to establish ground truth and spot-check production logs.

These are vendor claims and case studies, not independent proof that every workflow produces reliable oversight. They nevertheless show that the market already pays for the human contribution.

The gap is that most systems measure the output of the person—the label, edit, rating, or approval—and operational quality measures such as agreement, acceptance, and throughput. They rarely preserve a portable trace of the conditions under which that contribution was produced.

That suggests a clear boundary:

Human-feedback companies supply, organize, and quality-control the work. HRT qualifies the human episode and connects it to the machine trace and downstream outcome.

HRT should integrate with these systems, not become another annotation marketplace, approval inbox, or generic evaluation platform.

Build a living database of human-loop failures

The cases above should not remain a static article. Human Runtime could maintain a continuously updated Human Loop Failure Database: a public, evidence-linked record of incidents in which a person was expected to supervise, verify, approve, correct, teach, or recover a machine-mediated process and that safeguard failed or nearly failed.

The database would do more than collect AI scandals. It would identify the human–machine contract that broke.

Each record should answer:

  • What system, organization, domain, and date were involved?
  • What was the machine expected to do?
  • What was the human expected to detect, verify, approve, correct, or recover?
  • Was the human continuously engaged, periodically checking, or called only on escalation?
  • What information was available, presented, sampled, missed, disputed, or withheld?
  • How much time and authority did the person have?
  • What action followed, and what was the consequence?
  • Was the failure distraction, overload, automation bias, skill decay, ambiguity, bad training, weak escalation, group convergence, incentive conflict, deliberate suppression, or something else?
  • Which claims are established, disputed, inferred, or merely alleged?
  • What primary and secondary sources support the record?
  • Could HRT plausibly have helped in design, training, live operation, or after-action learning—or not at all?

The final question is essential. A database that labels every tragedy an HRT opportunity would become marketing, not evidence.

Where to look

The best discovery system would combine several source layers.

Official accident and adverse-event records

These sources are slower but usually provide the strongest causal evidence:

  • NTSB CAROL searches investigation and recommendation data across transport modes and supports JSON or CSV exports.
  • NHTSA's automated-driving crash data provides regularly updated downloadable incident reports for ADS and Level 2 driver-assistance systems.
  • NASA's Aviation Safety Reporting System contains confidential, de-identified frontline reports and near misses. NASA explicitly warns that these voluntary reports are not independently verified.
  • FDA MAUDE receives hundreds of thousands of medical-device adverse-event reports and provides downloads and an API. FDA likewise warns that reports may be incomplete, inaccurate, delayed, unverified, or biased.
  • OSHA's accident investigation search contains reviewed accident abstracts dating back decades and is updated from federal and state offices.
  • NRC event notifications and investigation records, national aviation and rail investigation bodies, maritime casualty reports, military accident boards, and cybersecurity incident disclosures can extend the same approach internationally.

Near-miss systems may be more valuable to HRT than fatality databases. They contain recovery, weak signals, and intervention opportunities before the evidence is dominated by catastrophic outcomes.

Existing AI incident monitors

The Human Loop Failure Database should ingest and enrich existing work rather than compete on raw collection.

The OECD AI Incidents Monitor tracks incidents and hazards from reputable international news. Its published methodology says it uses Event Registry, which processes more than 150,000 articles per day and clusters reports describing the same event. The OECD is also developing a common reporting framework and explicitly discusses interoperability with the AI Incident Database.

The AI Incident Database is another important seed source. Research based on its editorial process offers useful lessons about the difficult boundary between incidents, hazards, issues, and variants. HRT's contribution would be a narrower human-oversight lens: what role was assigned to the person, and why did it fail?

News and specialist reporting

For freshness, monitor global news clustering services and selected publications rather than scraping arbitrary sites indiscriminately. Event Registry already demonstrates the model. GDELT, Media Cloud, licensed news APIs, RSS feeds, and search alerts can supply discovery. High-value editorial sources include Reuters, Associated Press, BBC, the Financial Times, the Guardian, the New York Times, the Washington Post, IEEE Spectrum, MIT Technology Review, 404 Media, The Markup, and specialist safety publications.

News should create a candidate record, not a confirmed incident. The crawler should seek the underlying court filing, regulator notice, accident report, audit, correction, company disclosure, or research paper before publication.

Courts, regulators, and professional discipline

LLM failures increasingly surface through sanctions orders, employment cases, consumer disputes, professional discipline, procurement audits, and regulatory enforcement. CourtListener and RECAP, CanLII, BAILII, EU and national court portals, FTC and SEC releases, data-protection authorities, professional licensing bodies, and government audit offices are likely to yield higher-quality records than general “AI fail” searches.

The Mata, Air Canada, Horizon, and MyCity cases all become substantially more useful when the legal decision or audit is attached to the press story.

Research literature

Search Crossref, OpenAlex, Semantic Scholar, arXiv, PubMed, IEEE Xplore, ACM Digital Library, NASA NTRS, and transport-research indexes for empirical work on:

  • automation bias and complacency
  • out-of-the-loop performance
  • vigilance decrement and passive monitoring
  • takeover readiness and automation surprise
  • alert fatigue and alarm floods
  • human–AI reliance and advice taking
  • selective escalation and help-seeking
  • reviewer disagreement and annotation quality
  • RLHF, reward hacking, sycophancy, and preference bias
  • cognitive workload and physiological measurement
  • content moderation, traumatic exposure, and data-worker health

Research records should be linked to incident mechanisms, not presented as incidents themselves.

Vendor claims and customer stories

Human-feedback, evaluation, observability, autonomous-system, and safety vendors publish a constant stream of case studies. These can reveal what customers fear, which workflow patterns are commercially valuable, and how the market defines quality.

They should be stored as claims, with declared provenance and conflicts, not treated as independent validation. A useful recurring feature for the Human Runtime site would compare the vendor's proposed safeguard with what incident and research evidence says can still go wrong.

What a crawler should search for

A useful discovery query combines three kinds of language:

  1. Machine or workflow: AI, algorithm, chatbot, copilot, agent, automation, autopilot, decision support, recommendation system, safety system, alerting system, model output.
  2. Human responsibility: reviewer, operator, safety driver, moderator, annotator, evaluator, supervisor, pilot, clinician, editor, lawyer, analyst, approval, override, takeover, verification, escalation.
  3. Failure: failed to notice, failed to intervene, rubber-stamped, overreliance, complacency, distracted, overwhelmed, alert fatigue, hallucinated, fabricated, unverified, incorrect approval, delayed response, ignored warning, automation bias, inadequate oversight, no independent review.

The crawler should also look for institutional language that appears after facts have been established: “probable cause,” “contributing factor,” “human factors,” “insufficient monitoring,” “failed to verify,” “lack of oversight,” “contrary to procedure,” “sanctioned,” “consent order,” “corrective action,” and “lessons learned.”

The database needs a stronger loop than the systems it studies

Automated discovery will generate duplicates, weak allegations, causal overreach, and false matches. A single human reviewer clicking “publish” would reproduce the very problem the project is documenting.

A defensible editorial flow would be:

Discover → cluster → extract claims → retrieve primary evidence → classify the human role → independent review → publish with confidence and dispute status → update when evidence changes.

Important safeguards include:

  • Separate incident, hazard, near miss, allegation, and research finding.
  • Require a primary source or at least two credible independent reports for a confirmed record.
  • Preserve exact source passages internally while publishing copyright-safe summaries.
  • Record when an investigation remains open.
  • Distinguish an observed fact from an HRT inference or counterfactual.
  • Never infer a person's cognitive state retrospectively from an outcome alone.
  • Give named organizations and affected people a correction route.
  • Version every record rather than silently rewriting history.
  • Record crawler, model, prompt, reviewer, and editorial-decision provenance.

That last step turns the database itself into an HRT demonstration. The project can show how machine discovery and human verification interact, where reviewers disagree, what evidence they inspect, and what changes after challenge.

A fresh-content engine for Human Runtime

The database can produce several recurring formats for the Human Runtime website:

  • The Loop Failure of the Week: one verified incident and the exact oversight contract that failed.
  • New and developing: recent reports clearly labeled as unconfirmed, under investigation, or disputed.
  • Failure pattern pages: automation bias, failed handover, alert overload, rubber-stamp approval, suppressed dissent, evaluator fatigue, and escalation failure.
  • LLM field notes: legal hallucinations, agent actions, customer-service errors, evaluation failures, RLHF pathologies, and data-worker conditions.
  • What changed afterward: recalls, interface redesigns, new training, sanctions, policy changes, and whether the intervention worked.
  • Could HRT have helped? A disciplined assessment with “yes,” “partly,” “unlikely,” and “harmful if misused” outcomes.
  • Vendor claim versus field evidence: a fair comparison between promised human oversight and documented failure mechanisms.
  • Recovery library: near misses and successful interventions, not only disasters.

The recovery library matters. A database containing only failure will teach HRT what collapse looks like, but not what effective human contribution looks like. The real objective is to compare the two.

Where HRT helps—and where it does not

HRT is most credible when it measures whether oversight is functioning:

  • prolonged supervision and vigilance decay
  • automation handovers and takeover readiness
  • alert uptake and evidence sampling
  • repeated approval or annotation work
  • team coordination and independent verification
  • simulator training and intervention efficacy
  • expert escalation and routing
  • after-action reconstruction
  • testing human-performance assumptions before deployment

It is weak or inappropriate when:

  • the person lacks time, information, or authority
  • evidence is hidden or deliberately suppressed
  • incentives reward approval despite known risk
  • the underlying model, policy, or automation is invalid
  • decisions happen faster than meaningful human intervention
  • the task lacks defensible ground truth
  • signals are used to infer intent, honesty, moral worth, or permanent competence
  • monitoring becomes coercive workforce surveillance

Physiological measurement deserves particular restraint. A systematic review of 58 studies found that physiological signals can be sensitive to workload but that no single measure satisfies all requirements. A review of 109 supervisory-control studies found far less research connecting such signals to actual task performance than to general mental-state constructs. That evidence gap is precisely why HRT must combine task events, behaviour, outcomes, self-report, and optional multimodal measurement rather than claim to read the mind.

From ceremonial oversight to observable oversight

The central problem is not that humans are unreliable and machines are reliable, or the reverse.

It is that combined systems routinely assign responsibility without measuring whether the responsible node can perform the assigned role. They count an operator, reviewer, expert, or approver as present. They rarely test whether the person had the right evidence, noticed it, understood the situation, retained the relevant skill, had time to act, possessed real authority, and produced a better outcome.

Human Runtime can make those questions operational.

The Human Loop Failure Database can make them public, comparable, and cumulative.

Together, they create a sharper standard for human–AI systems:

Do not tell us that a human was in the loop. Show us that the loop worked.