The C.H.R.I.S. Protocol: How to Challenge AI Root-Cause Analysis

Developer studying code and distributed system diagrams across multiple monitors

Oct 7, 2026 · @christopher Corder

An AI-generated root cause is a hypothesis, not a finding. Treat it like one. Preferably like a hypothesis from a new hire who has never seen production.

In platform support, the cost of a wrong RCA is not abstract. A customer rolls back a deployment that was never the problem. A team scales out to fix a SNAT issue that was actually a slow dependency. A stamp gets blamed for a client-side retry storm. The work moves forward on a conclusion nobody tested, and three days later we are all back on the same bridge call.

AI makes this worse in one specific way. Its wrong answers look exactly like its right ones. Same structure. Same confidence. Same clean timeline. It will tell you the SNAT ports were exhausted with the calm certainty of someone who has never opened a netstat. Fluency is not evidence, and an RCA that reads well has earned nothing yet.

The fix is not to stop using AI. I have spent 20+ years doing this work, first on global telecom implementations at Oracle and now at Azure & AI with Microsoft, and I am not going back to reading 4 GB dumps by hand for fun. It is too useful for dump triage, trace analysis, and log correlation. The fix is to cross-examine every conclusion before you act on it. I use a five-stage method for that. I call it the CHRIS Protocol.

The five stages

The first four run in the same conversation, in this order. The fifth runs somewhere else entirely.

StageThe questionWhat it exposes
Counter“What’s the strongest argument against what you just told me?”The best competing explanation
Hidden assumptions“What assumptions is this answer based on?”What was taken as true without proof
Reconcile with evidence“Is this consistent with [specific evidence]?”Whether the conclusion fits the data
Invalidate“What would you need to see to change your answer?”Whether the conclusion is falsifiable
Stress-testFull adversarial review, fresh session, different modelSelf-confirmation from the original model

C: Counter

Ask for the strongest case against the conclusion first. Models anchor on their first answer and defend it, much like an engineer who already posted the RCA in the incident channel. Asking for the opposition before anything else breaks that anchor while it is still soft.

A useful answer names the strongest conflicting evidence, a plausible alternative, a logical gap, and what the original answer skipped. A weak answer offers a token objection, waves at it politely, and then agrees with itself. When that happens, ask it to separate objections grounded in the supplied evidence from hypothetical ones.

H: Hidden assumptions

Every RCA rests on things nobody proved. In App Service work the usual suspects are:

  • The logs supplied are complete for the incident window.
  • The instance in the trace is the instance that failed.
  • An event that happened first caused the event that happened second. (Correlation has closed more tickets incorrectly than any bug I have ever met.)
  • Guidance for one SKU, OS, or storage model applies to this one.
  • Nothing changed between data collection and analysis.

Follow up with: “Which of these is least certain, and how does your answer change if it is false?” That second question is where most of the value is. It is also where the answer usually starts to wobble.

R: Reconcile with evidence

Now put the conclusion against specific artifacts. Paste the log lines. Name the counter. Give the timestamp.

“You said DNS caused the failures. Is that consistent with these entries showing successful name resolution immediately before the connection timeout?”

Precision matters. Vague evidence gets a vague yes, and AI produces vague yes at industrial scale. And “consistent with” is not “proven by.” Several explanations can fit the same trace. Ask what the evidence rules out, what it supports, and what it leaves open.

I: Invalidate

Ask what would change the answer. A conclusion that cannot name its own falsifier is an opinion with formatting. Nice headers. Still an opinion.

Watch for falsifiers that sound decisive and are not. Ask why that evidence would separate this hypothesis from the alternative. If both explanations would produce the same observation, it is not a falsifier.

S: Stress-test

The first four stages share a flaw. The model that produced the answer is reviewing it, with the whole conversation steering it toward agreement. That is the defendant sitting on its own jury.

So the last stage moves the review. Open a fresh session on a different model. Give it the conclusion and the raw evidence, not the original reasoning. Ask for the full adversarial review below. Disagreement between models is signal. Agreement from a model that did not see the first argument is worth more than agreement from one that did.

Worked example: “The database caused the outage”

The AI reviews an incident and concludes the database caused it. Of course it does. The database is the usual suspect in every outage since the invention of the database. Here is how the protocol pressures that conclusion.

  1. Counter. “What is the strongest evidence-based argument against the database being the cause? Give the most plausible alternative.” A good response might point to connection pool exhaustion in the app, outbound port pressure, or a downstream API that was slow first.
  2. Hidden assumptions. “What does this conclusion assume? Mark each one confirmed, inferred, or unknown.” Often the answer is that database latency was inferred from app-side timing, not measured on the server.
  3. Reconcile. “Is this consistent with the database metrics, the application logs, and the incident timeline? Point out contradictions.” If server-side DTU and query duration were flat while the app timed out, that is a contradiction.
  4. Invalidate. “What observation would make you abandon this? Why would it separate the database hypothesis from the alternative?” A strong answer: application failures that begin before any database call is made.
  5. Stress-test. New session, different model, raw evidence only. “Keep or revise the conclusion. State confidence and the main unresolved uncertainty.”

The conclusion may survive. That is fine. A conclusion that survives this is worth acting on. One that only survived because nobody pushed is a future incident with a ticket number you have not been assigned yet.

The stress-test prompt

Paste this into a fresh session on a different model, followed by the conclusion and the raw evidence.

What the protocol keeps catching

I have been building a case library of real AI-generated RCAs from App Service support that read well and were wrong. The same failure patterns keep showing up:

  • Absence treated as evidence. No error in the log becomes proof that a component was healthy, when the log simply did not cover it.
  • Scary-artifact anchoring. The most alarming line in a trace gets promoted to root cause because it is alarming, not because it is causal.
  • Wrong platform model. Guidance for one storage or hosting model applied to an app running on another.
  • Survivorship bias in comparisons. Percentiles compared across populations where the failed requests never made it into the sample.

None of these are exotic. All of them survive a casual read. Most of them would sail past a tired human at 2 AM. The Hidden assumptions and Reconcile stages catch most of them.

The limits

The protocol does not make AI correct. It does not verify facts on its own. A model can still miss a flaw, invent a convincing counterargument, or name a falsifier that sounds decisive and is not.

What it does is turn a fluent answer into a checkable one. You see what it assumed, what it rests on, and what would break it. That is the gap between trusting an answer and evaluating it.

It works best with specific evidence and a specific question. Generic pushback gets generic answers.

AI is the fastest junior engineer I have ever worked with. I still check its work. So should you.

Run the protocol on the next AI conclusion you were about to act on. If it breaks, I want to hear how. I am collecting the failures, and I promise yours is not the weirdest one.Chris Corder is a Senior Azure Technical Advisor on the App Service support engineering team at

Leave a comment