July 29, 2026
The AI Mirage in Healthcare Billing: Why Technology Will Not Replace Human Judgment
- by Sean Weiss, Partner & VP of Strategic Litigation Services
By Sean M. Weiss
Artificial intelligence is transforming healthcare. That statement is no longer controversial. AI is already being used to summarize clinical encounters, identify documentation gaps, suggest codes, flag claims for review, predict denials, and support utilization-management decisions.
The controversy begins when AI is marketed as something it is not: an infallible replacement for the people who understand medicine, documentation, coding, billing, compliance, and the law.
The most dangerous version of that marketing is the claim that an insurer, working with an AI vendor, can “legitimately” deny 60 percent of provider claims. That assertion should be treated with extreme skepticism. It is not a self-proving measure of accuracy, medical necessity, fraud, or improper billing. At best, it is an unexplained performance statistic. At worst, it is a sales pitch dressed up as science.
“A denial is not the same thing as a correct denial. A prediction is not a finding. A statistical outlier is not fraud. And algorithmic output is not a legal conclusion.”
AI has an important place in healthcare. It can reduce administrative burdens, improve consistency, identify patterns, and help professionals focus their time where it matters most. But AI remains a tool. It requires human interaction at the front, middle, and end of the process. That will remain true for the foreseeable future.
AI will not honestly replace auditors, billers, or coders. It will change how those professionals work. It will eliminate some repetitive tasks and increase the value of others. But the accountability, reasoning, judgment, and professional skepticism that these roles require cannot be reduced to tokens, probabilities, or a denial percentage.
The “60 Percent” Denial Claim Is Not What It Pretends to Be
The first question should be simple:
Sixty percent of what?
There is a substantial difference between:
- 60 percent of all claims submitted
- 60 percent of claims selected for an unusually aggressive audit
- 60 percent of claim lines flagged for additional review
- 60 percent of claims containing a documentation discrepancy
- 60 percent of claims that an algorithm predicts will be denied
- 60 percent of claims ultimately determined, after complete human review, to be unsupported
- 60 percent of claims that are actually and lawfully denied after considering the medical record, applicable coverage policy, coding rules, authorization requirements, and provider response
Those are not interchangeable categories. Treating them as equivalent is the first fallacy.
A model can identify claims that are expensive, unusual, inconsistent with historical patterns, or different from a payer’s preferred utilization profile. None of those characteristics proves that the claim is improper.
The number also says nothing about:
- False positives: How many legitimate claims did the system flag?
- False negatives: How many improper claims did the system miss?
- Appeal outcomes: How many denials were reversed?
- Human review: Did a qualified reviewer examine the complete record?
- Data quality: Was the model working from the full clinical and billing record?
- Denial definitions: Was a technical edit treated as a substantive denial?
- Claim-line inflation: Were individual services counted as separate “claims”?
- Model drift: Does the reported performance still exist after the patient population, providers, codes, or policies change?
A denial algorithm can achieve an impressive percentage by being aggressively wrong. If the system flags nearly everything, it may produce a high number of denials while providing very little information about whether those denials are correct.
That is not efficiency. It is the industrialization of suspicion.
Why the 60 Percent Assertion Is a Fallacy
1. Historical payment patterns are not medical necessity
AI systems learn from historical data. Historical data reflects prior human decisions, payer policies, provider behavior, documentation practices, regional practice patterns, and sometimes historical bias.
If a payer historically denied a category of service, an AI model may learn that the category is “likely to be denied.” That does not establish that the service was medically unnecessary. It may only establish that the payer has a history of denying it.
A model trained to predict payer behavior can become very good at predicting payer behavior without becoming good at determining what care was appropriate.
2. Medical records are contextual, not merely transactional
Healthcare documentation is not a collection of isolated keywords. It is a narrative that develops over time.
The significance of a symptom may depend on:
- The patient’s history
- The progression of the condition
- Failed conservative treatment
- The physician’s differential diagnosis
- Examination findings
- Diagnostic testing
- Comorbidities
- Risk factors
- Response to prior treatment
- The clinical judgment exercised at the time of care
An algorithm may identify that a particular code is absent. It may not understand why the code was absent, whether another portion of the record supplies the necessary support, or whether the documentation reflects a clinically reasonable decision under the circumstances.
3. Coding is governed by rules, not just pattern recognition
Coding requires more than matching words to codes. It requires interpretation of official guidelines, payer rules, sequencing requirements, modifiers, global-period concepts, bundling edits, medical-necessity policies, and the relationship between documentation and the service reported.
A code suggestion is not a code determination.
A system may recommend a code that appears statistically likely but is inconsistent with the operative report, the level of service, the applicable coding guidelines, or the provider’s actual work. Conversely, it may fail to recognize legitimate complexity because the relevant facts are expressed in ordinary clinical language rather than in the precise terms on which the model was trained.
4. Missing data is not negative evidence
One of the most common errors in automated review is treating an absent data element as proof that the underlying fact did not exist.
The absence of a phrase from a claim form does not necessarily mean the absence of the fact from the medical record. The absence of a code does not necessarily mean the absence of a diagnosis. The absence of an authorization number does not necessarily mean that authorization was not obtained. The absence of a structured field does not necessarily mean the clinical event did not occur.
AI systems are particularly vulnerable when data is fragmented across:
- Electronic health records
- Scanned records
- Operative reports
- Laboratory systems
- Referral platforms
- Authorization portals
- Payer correspondence
- Appeals files
- Separate provider locations
A model that reviews only what is convenient to retrieve may produce a confident conclusion from an incomplete record.
5. Denial policies are not the same as clinical truth
An insurer’s coverage policy is not necessarily a complete statement of medical truth. It is a contractual, regulatory, or business rule governing payment.
A service can be clinically appropriate even when a payer disputes coverage. A service can be supported by the record even when a payer applies a narrow interpretation of its policy. A provider can comply with professional standards even when the claim requires an appeal.
The model must not be allowed to convert “the payer does not want to pay” into “the provider was wrong.”
6. AI may reproduce payer bias
If the training data reflects aggressive utilization management, the model may reproduce and amplify that conduct.
Bias can enter through:
- Incomplete representation of patient populations
- Historical underdiagnosis
- Unequal access to specialty care
- Geographic practice differences
- Provider-type assumptions
- Socioeconomic proxies
- Language differences
- Documentation-style differences
- Race, disability, age, or gender proxies
- Feedback loops created when prior denials become future training data
A patient who receives more intensive care because of complex comorbidities may look to a model like an outlier. A safety-net provider may appear inefficient because the provider serves a population with greater clinical and social needs. A specialist may appear expensive because the specialist treats difficult cases.
Those are not legitimate reasons to deny medically necessary care.
7. A model cannot independently determine causation
Healthcare billing frequently requires a determination of why a service was provided and how the service relates to the patient’s condition.
That is a causal question, not merely a predictive question.
A model may identify that a procedure often follows a particular diagnosis. It may not be able to determine whether, in this patient’s case, the procedure was performed because of that diagnosis, because of a complication, because of a failed prior intervention, or because of a clinical circumstance that does not fit the dominant pattern.
8. The model may be optimized for the wrong objective
An insurer may measure success by:
- Reduced claim payments
- Reduced utilization
- Increased denial rates
- Shorter review times
- Lower administrative costs; or
- Fewer paid claim lines
Providers, patients, and clinicians may measure success by:
- Accurate payment
- Timely access to care
- Appropriate treatment
- Patient safety
- Correct coding
- Regulatory compliance
- Resolution of legitimate disputes
Those objectives are not identical. A model optimized to reduce payments may perform exactly as designed while producing unacceptable clinical and legal outcomes.
9. Appeals expose the weakness of automated certainty
A denial that is never appealed may be counted as a successful denial even though the provider lacked the time, resources, or information to challenge it.
Prior-authorization research has documented the substantial burden that payer review places on physicians and practices. The American Medical Association’s 2024 prior-authorization survey reflects the continuing administrative burden and the need for reform. A system should not be judged solely by how many denials it produces when many providers cannot realistically appeal every incorrect decision.
The meaningful question is not how many claims the system denies. It is how many decisions survive complete, informed, human review.
10. A black-box score cannot substitute for an explanation
A provider, patient, regulator, or court should be able to understand why a claim was denied.
” The model determined that the claim was inconsistent with expected utilization” is not an explanation. It is a restatement of the conclusion.
A defensible denial should identify:
- The specific policy or rule applied
- The clinical or documentation fact considered missing
- The portion of the record reviewed
- The reasoning connecting the facts to the decision
- The identity and qualifications of the human reviewer, when required
- The steps necessary to correct or appeal the determination
Without that information, the provider is not meaningfully reviewing a denial. The provider is attempting to reverse an unexplained computer output.
The Accountability Imbalance Is Real
The practical legal and operational imbalance is difficult to ignore.
When an insurer uses AI to recommend or initiate an adverse coverage decision, the resulting harm may be treated as a utilization-management dispute, a contract issue, or an administrative appeal. The insurer may characterize the system as merely a decision-support tool. The patient and provider are then directed into an appeals process, often after care has been delayed.
Investigations have reported concerns regarding the use of algorithms by Medicare Advantage plans to limit or terminate care. Those concerns are especially serious because coverage-decision algorithms generally do not face the same regulatory framework as AI systems regulated by the Food and Drug Administration as medical devices.
That does not mean insurers are legally immune. They are not. State insurance regulators, the Centers for Medicare & Medicaid Services, contractual obligations, federal program requirements, and other legal mechanisms may apply. The National Association of Insurance Commissioners’ model guidance recognizes the need for governance, risk management, oversight, documentation, and accountability when insurers use artificial-intelligence systems.
But the immediate and visible risk often falls on the provider.
If a provider’s AI system invents a diagnosis, adds an unsupported modifier, creates a nonexistent medical-history fact, or generates documentation that does not accurately reflect the service provided, the provider may face:
- Claim denials and recoupments
- Prepayment review
- Post-payment audits
- Contractual sanctions
- Licensing or disciplinary consequences
- Civil penalties
- Exclusion concerns
- False Claims Act exposure
- Reputational damage
The use of AI does not transfer the provider’s compliance obligations to the software vendor. It does not make an inaccurate claim accurate. It does not convert an unsupported medical record into reliable documentation.
At the same time, an AI hallucination does not automatically establish fraud or a False Claims Act violation. The False Claims Act’s knowledge standard requires actual knowledge, deliberate ignorance, or reckless disregard, not mere negligence alone.31 U.S.C. § 3729(b)(1). Materiality and the circumstances surrounding submission also matter. Universal Health Services, Inc. v. Escobar, 579 U.S. 176, 181-92 (2016)
That distinction is critical. Providers should not be punished merely because a supervised tool makes an error. But providers that deploy AI without meaningful controls, fail to review its output, ignore repeated errors, or submit claims they have reason to know are inaccurate may create significant compliance risk.
The same principle should apply to insurers.
An insurer should not be permitted to hide behind a vendor when its automated system improperly denies medically necessary care. Medicare Advantage regulations, for example, restrict an organization’s ability to contract away civil liability for damage caused to an enrollee by the organization’s denial of medically necessary care.42 C.F.R. § 422.212. The entity that deploys the system must remain accountable for the system’s operation.
“The party that chooses the model, supplies the data, defines the objective, deploys the workflow, and benefits from the result cannot disclaim responsibility by pointing to the vendor.”
AI Requires Humans at the Front, Middle, and End
The proper healthcare model is not “AI replaces the professional.” It is:
Human judgment at the front. AI assistance in the middle. Human accountability at the end.
At the front: humans create the source material
The quality of an AI output depends on the quality of the input.
Clinicians must document what occurred, why it occurred, what was found, what was considered, and what was done. Staff must enter accurate patient and insurance information. Organizations must design workflows that preserve the integrity of the medical record.
AI cannot repair a fundamentally incomplete clinical encounter. It can identify a gap. It cannot truthfully fill that gap unless the provider supplies the underlying fact.
The front-end human responsibilities include:
- Accurate clinical documentation
- Complete histories and examinations
- Clear medical decision-making
- Proper identification of services
- Accurate patient and payer information
- Appropriate authorization procedures
- Clear communication with coding and billing personnel
- Protection of patient privacy and data security
In the middle: AI can accelerate professional work
This is where AI is most useful.
AI can assist with:
| Function | Appropriate AI Role | Human Responsibility |
|---|---|---|
| Documentation | Summarize encounters and identify possible omissions | Confirm that the summary is accurate and complete |
| Coding | Suggest codes, modifiers, and documentation queries | Apply coding rules and determine whether the record supports the code |
| Billing | Identify missing data, duplicate charges, or payer edits | Decide whether the claim is accurate and submit it |
| Auditing | Prioritize records for review and identify patterns | Conduct the audit and support the conclusion with evidence |
| Denial management | Categorize denials and identify appeal deadlines | Evaluate the payer’s rationale and prepare the response |
| Compliance | Detect unusual activity or recurring errors | Determine whether corrective action is required |
The middle of the process is not a license for automation without supervision. It is the place where trained professionals use technology to work more efficiently.
At the end: humans make the accountable decision
Before a claim is submitted, a professional must be able to answer:
- Does the documentation accurately describe the service?
- Does the code reflect the documented work?
- Is the modifier supported?
- Is the medical necessity rationale present?
- Does the claim comply with applicable payer and regulatory requirements?
- Did the AI invent, assume, or omit anything?
- Would the provider be prepared to defend the claim before an auditor, regulator, payer, or court?
If the answer to those questions is unknown, the claim is not ready.
The final human review is not ceremonial. It is the point at which an organization accepts responsibility for what it submits.
Why AI Will Not Replace Auditors, Billers, and Coders
AI will replace certain tasks. It will not replace the professions.
Auditors
An auditor does not merely find mismatches. An auditor evaluates evidence, understands process failure, tests controls, recognizes patterns, interviews personnel, distinguishes isolated error from systemic conduct, and explains findings in a manner that can withstand scrutiny.
AI can identify what deserves attention. It cannot independently determine the significance of every discrepancy.
The strongest auditors will use AI to expand their reach, not surrender their judgment.
Billers
Billing is not the mechanical act of transmitting a claim. It involves payer rules, authorization requirements, edits, timely filing, coordination of benefits, documentation, appeals, communication, and the practical realities of resolving disputes.
A skilled biller understands when a denial is legitimate, when it is technical, when it reflects a payer-processing error, and when the record requires clarification. That is not merely data entry. It is operational reasoning.
Coders
Coding requires disciplined interpretation. A coder must connect the record to the code set without adding facts that are not documented or overlooking facts that are.
The coder must understand that:
- A more specific code is not always a more accurate code
- A higher-paying code is not automatically supported
- A physician’s terminology may require clarification
- A modifier must reflect actual circumstances
- A diagnosis must be clinically and documentarily supported
- A code must be defensible after the claim is submitted
AI can propose. A coder must decide.
The Correct Standard: Augmentation, Not Abdication
Healthcare organizations should stop asking whether AI can eliminate people. The more responsible questions are:
- What task is AI performing?
- What data is it using?
- What can go wrong?
- Who reviews the output?
- How is the review documented?
- What happens when the AI is wrong?
- Can the decision be explained and challenged?
- Is the model being measured for accuracy, or merely for financial performance?
A responsible AI program should include:
- Pre-deployment validation
- Ongoing accuracy testing
- Bias and disparate-impact monitoring
- Version control
- Audit trails
- Defined escalation procedures
- Human override authority
- Periodic retrospective review
- Vendor accountability
- Data-security safeguards
- A process for correcting erroneous outputs before submission or adverse action
The NAIC’s model bulletin on insurers’ use of artificial-intelligence systems reflects the broader principle that AI deployment requires governance and accountability, not merely technical capability.
The same principle must govern providers.
Providers should never allow an AI system to:
- Create unsupported clinical facts
- Automatically sign documentation
- Select a final code without review
- Submit a claim solely because the model approved it
- Alter the medical record without traceability
- Generate medical necessity language disconnected from the actual encounter; or
- Conceal uncertainty behind confident wording
The responsible question is not whether AI sounds persuasive. It is whether the output is true.
The Future Belongs to Professionals Who Know How to Use AI
The future of healthcare revenue cycle management will not be human versus machine. It will be professionals who understand how to use AI versus organizations that blindly trust it.
The best billers will use AI to find claims that need attention faster. The best coders will use AI to identify potential documentation gaps while preserving independent coding judgment. The best auditors will use AI to analyze larger data sets while applying human skepticism to the results. The best compliance officers will require evidence that the technology works before allowing it to influence patient care or payment.
AI can help us process more information. It cannot bear moral responsibility. It cannot independently understand the patient. It cannot explain why a clinician made a difficult decision. It cannot replace the professional obligation to be accurate, fair, and defensible.
Human beings reason. We use logic. We evaluate context. We understand consequences. We apply process. We recognize that two records containing similar words may describe entirely different clinical realities.
We do not reduce every human judgment to a tokenized probability.
AI has a legitimate and necessary place in healthcare. But its proper role is to support professionals, not to replace them, conceal payer conduct, manufacture denials, or create a false appearance of certainty.
The 60 percent denial claim is not proof of intelligent adjudication. It is a reminder that the healthcare industry must demand better questions, better data, better oversight, and better accountability.
The future should not be automated healthcare without humans.
It should be better healthcare in which humans use automation responsibly.
That distinction is not academic. It is the difference between technology that improves the system and technology that merely makes bad decisions faster.