Skip to Content

Humans Spiraling in the AI loop (Part 2) (Untrusted Overseeing Untrusted)

Published on September 28, 2026

Time to read: 11 minutes

Daksha Bhasker | September 28, 2026 | Part 2 of 2

People travelling through a futuristic spiral of interconnected AI systems, data and digital controls.
Humans Spiraling in the AI Loop (Part 2): Untrusted Overseeing Untrusted

In part 1 we examined how Human-in-the-Loop (HITL) and Human-on-the-Loop (HOTL) in AI systems reintroduce human weaknesses and limitations back into automated systems that the industry has spent decades engineering them out of.

This segment examines whether HITL/HOTL mitigates AI risk, and what happens when humans, “the weakest link in security”, become a trusted security control.

The control that isn’t

OWASP’s LLM Top 10 names HITL as mitigation for  Prompt Injection and Excessive Agency. [1]. The logic is that where the AI system or model cannot be trusted to act autonomously, human approval is required before high risk or high value operations are executed.

This treats HITL/HOTL as a robust control rather than a component that brings with it human failures and limitations. A mitigation that inserts a human assumes that humans are reliable adjudicators, capable of correctly evaluating AI actions, are resistant to manipulation and are working with information the AI system has not compromised. In reality, these assumptions are untrue.

The humans in the machine

AI systems hallucinate, drift and can produce inaccurate outputs that require verification. Towards this, humans are inserted into the loop as a control. Unlike traditional software, AI systems are not trusted to operate consistently and completely unsupervised.

Humans have long been known in the industry, as the weakest link in security. Humans filling the HITL/HOTL role in AI systems present a significant source of risk. Verizon’s 2026 Data Breach Investigations Report found the human element present in 62% of breaches [2]. The category includes people who fall for phishing and hand over credentials, misconfigure systems or misdirect sensitive data, insiders who misuse legitimate access, and people who make ordinary mistakes under normal working conditions.

AI’s untrustworthiness has led to insertion of human oversight even though, HITL/HOTL remain subject to a slew of attack vectors. Humans can be phished, socially engineered, or misuse the access their role grants them, among others. Neither training nor vetting removes the limitations of fatigue, attention drift, cognitive bias, and susceptibility to manipulation. AI systems deemed too unreliable to operate independently are placed under the supervision of humans, the very factor security and software systems have spent decades engineering around.

How the humans in HITL/HOTL gets staffed can make the problem worse. Mary L. Gray and Siddharth Suri’s Ghost work documented how much of the human labour underlying AI and digital platforms are organized as on-demand piecework. Workers complete fragmented tasks through APIs. They label data, moderate content, create content, make decisions, and similar work, often with little context of the tasks before or after theirs [3].

The book does not address HITL/HOTL directly, but similar conditions can exist there. A HITL/HOTL working a queue against a clock, with limited upstream context, is poorly positioned to evaluate what they are approving. Not every role is staffed this way, but in instances they are, ghost-work conditions add to limitations of HITL/HOTL.

This discussion does not argue that AI should operate without oversight. It argues that human oversight is treated as a solution to a technology problem that it does not solve.

Trust me, I’m human

Humans in HITL/HOTL depending on where they are placed in AI systems carry different risk profiles. (Figure-1)

The HITL can be external to the system or enterprise, in that the human in question could be the customer, the end user, the general public or the party submitting input, that the AI system will act on. These humans have no obligation to be honest, no accountability to the organization, and may have every incentive to test the system’s boundaries. Treating this input as a trusted signal is a design blunder.

The internal HITL/HOTL can be the employee, designated trusted party, operator, and the control on which AI operations depend on. These humans are presumed trustworthy byvirtue of employment. Employment, however, does not neutralize insider threat or immunize against social engineering. It also does not remove fatigue or attention drift. Humans can become a target of attacks and in some cases the attack vector.

Both placements of HITL/HOTL carry different risk profiles. Neither should be treated as inherently trusted. Both are untrusted signals performing validation on AI systems while requiring security controls themselves.

Figure 1. Untrusted Signals from Humans as Security Controls in HITL/HOTL

Hacking the human in the loop

OWASP’s LLM Top 10 names human-in-the-loop as recommended mitigations for Prompt Injection and Excessive Agency. However, OWASP’s own community research document attacks that defeat it for both.

The attack is catalogued as HITL Dialog Forging, also known as Lies-in-the-Loop [4]. The approval dialog the human sees is generated by the AI system. It is rendered from the same context the model processed. Model context can include untrusted sources such as web pages, documents, github code, emails, tool outputs among others. Content placed in any of those sources can influence what the dialog with the HITL/HOTL contains. The human is not receiving descriptions of verified action, rather they are exposed to untrusted content. The research demonstrated that it is possible to manipulate the HITL/HOTL dialogue to achieve remote code execution.

On the surface leveraging the HITL/HOTL recommendation appears viable but in reality it is incomplete. The mitigation recommendation treats HITL/HOTL as a control without accounting for human failures in that control. Relying on human intervention assumes the HITL/HOTL can correctly determine required action, recognize manipulation, and make decisions based on information the AI systems has not already influenced or compromised. A control operating on compromised input cannot defend against that compromise.

The Boring Software We Actually Need

AI unreliability is well established. Placing human approval steps in front of AI output does not mitigate the underlying unreliability but introduces human error into the operations.

Where security or operational requirements can be expressed as rules, deterministic software can enforce them more reliably than humans. Software applies the same rule, the same way every time, without fatigue or drift. Properly implemented controls cannot be persuaded by well-crafted inputs or prompts to disregard those rules. Security engineering already uses deterministic enforcement extensively through policy enforcement points, rate limits, allow lists and transaction thresholds among others.

Deterministic controls can also constrain what an AI system is permitted to do. They can define acceptable output ranges, establish conditions under which actions are allowed, and set limits beyond which the system cannot proceed. These controls do not depend on a human recognizing a problem at the point of execution.

HITL/HOTL has a place, but as one control in a defense in depth model (Figure -2). It fails as the only defense between and AI systems’ output and its consequences. Deterministic controls should first constrain what the AI system can do. Human review must be selective for decisions that truly require human judgement. Software can assist humans in that review by surfacing deviations, identifying information that falls outside expected parameters, providing relevant context and narrowing the decision to exactly that which requires human attention. HITL/HOTL works better when the human is given specific, constrained decision instead of being expected to identify problems or make decisions about entire ambiguous AI output.

Figure 2. Putting Humans and AI in Their Place

Putting Humans and AI in Their Place

The discussion against HITL/HOTL in current deployment models is not a case for AI autonomy over human judgment. AI autonomy and human reviews have different strengths and weaknesses, and their placement in defense in depth strategy should consider both.

Show me the numbers. A system should not go into production, or be trusted with expanded autonomy, without verifiable error rates under real operating conditions. Pre-deployment benchmarks do not account for production inputs, adversarial conditions, or changes in data and systems behaviour. Error rates require ongoing monitoring once deployed, and reassessment as systems or input parameters change. Procurement should treat black-box components as defects rather than inherent properties of AI systems. AI systems with errors that cannot be examined cannot be reliably protected by HITL/HOTL security controls or defense in depth.

Just because you can AI it…Not all tasks benefit from AI integration. Deterministic problems with known inputs and defined outputs were solved decades ago by scripts, rules engines, preprogrammed workflows and deterministic software. They provide consistency and reliability neither AI nor humans can match. Do not spend tokens and human effort trying to force AI to do work it is poorly suited for. The costs extend beyond tokens. Inappropriate use of AI introduces failures into processes that were previously stable.

60% of the time, it works every time AI systems hallucinate and drift. A system that is 80% accurate does not deliver 80% of the value, because the errors are not labelled. Finding the flawed 20% means reviewing all 100% of the output. The review costs, often borne by HITL/HOTL, scales with total volume of AI output, not with just the error rate. This is where the promised efficiency is lost. Where a single incorrect output can produce an unacceptable regulatory, safety, or financial consequence without another control preventing it, AI is likely not the right fit for that operation, regardless of how well it performs on average. Where outputs are advisory, reversible, or checked by a downstream process that addresses errors by rule, AI may be appropriate.

Put it behind a fence. Limit what AI systems are allowed to do autonomously. Require its outputs to fall within defined, measurable limits before they are eligible for release. Flag deviations for review by rule, not by asking human reviewers to catch them unassisted. Deterministic controls apply the same rules the same way every time. They do not tire, drift, or get manipulated out of a decision or output, during routine operations. Where boundaries can be applied in software, enforcing them in software is more reliable than asking a person to hold the line. A human clicking approve should not be sufficient on its own to permit actions, that would otherwise not have been entrusted to a human to perform manually in the first place. What AI can do, what it can reach, and how far a single action can extend should be defined in advance. Human approval should also operate within those limits. Not extend them. If human judgment requires extending the limits, then a separate more privileged path with its own controls must be invoked.

Trust me, I work here. Humans placed as guardians of AI output and processes require the same scrutiny or tighter oversight as the system they are overseeing. That means identity and access controls on what the HITL/HOTL can approve, limits on the scope of approvals, monitoring of the review function itself and audit trails.

Separation of duties principle applies to HITL/HOTL. Humans, especially those who can approve actions, trigger them and clear the record afterwards are a excessive privilege and collusion risk. Design and placement of HITL/HOTL must prevent humans from exploiting AI system weaknesses.

With great power comes super-admins. AI deployments frequently grant HITL/HOTL operators broad standing access to approve whatever escalates to them, and that access is rarely scoped down afterward, producing a super-admin created by convenience rather than design. Humans should not be delegated authority the industry would never have granted them before AI made that delegation seem necessary.

Help the humans out. Use traditional software to surface what changed, flag what falls outside expected parameters, and provide independent evidence relevant to the decision. Narrow the approval to the specific decision requiring human judgement rather than presenting an entire AI output for wholesale sign-off. A scoped, specific question reviewable by humans, is more effective than open-ended approval requests that become over reliant on limited human capabilities.

Trust Neither, Constrain Both

HITL/HOTL does not resolve the trust problem introduced by AI. It introduces another component with its own failure modes, attack surface, and limits. In AI systems with HITL/HOTL, neither the AI nor the human should be inherently trusted. Both require constrains and security controls.

Where requirements can be expressed as rules, enforce them deterministically. Where judgment is required, scope the decision presented to the human and provide information independent of the AI output being reviewed. Human approval should not replace authorization, access controls, separation of duties, other established security controls, nor should AI systems create authority that a human would never otherwise have been entrusted with.
The objective is to design a system that remains resilient when the AI, the human or both get it wrong.


Author’s Note: Opinions expressed in this post are the author’s and not necessarily those of her employer. Daksha is one of our Editors/Reviewers for the Pulse & Praxis Journal for Critical Infrastructure Protection, Security and Resilience.

See Alsp Part 1 – Humans Spinning in the AI Loop

Header Image and Figures are AI generated through various prompts provided by the author.

References:

[1] OWASP Foundation, “OWASP Top 10 for Large Language Model Applications,” 2025. [Online]. Available: https://genai.owasp.org/llm-top-10/. [Accessed: Sep. 12, 2026].

[2] Verizon, “2026 Data Breach Investigations Report,” 2026. [Online]. Available: https://www.verizon.com/business/resources/Tad2/reports/2026-dbir-data-breach-investigations-report.pdf. [Accessed: Sep. 12, 2026].

[3] M. L. Gray and S. Suri, Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass. Boston, MA, USA: Houghton Mifflin Harcourt, 2019.

[4] O. Ron and D. Tumarkin, “HITL dialog forging (aka Lies-in-the-Loop),” OWASP Foundation. [Online]. Available: https://community.owasp.org/attacks/Lies_in_the_Loop. [Accessed: Sep. 12, 2026].