Breaking Down the Monolith
One overloaded prompt becomes four sequential steps with independent retries.
First published on LinkedIn.
A few weeks ago I thought I had this figured out: first normalize the input into a canonical format, then map to X12, then generate DSL.
Then I started running real payer edits through it at scale, and too many rules needed manual review or LLM refinement. It wasn't terrible, but not good enough.
The problem was what I was asking the LLM to do in a single call.
The extraction prompt was one massive instruction set telling the model to:
Classify the validation type (FOREACH vs WHEN vs AGGREGATE)
Map natural language to X12 segments and elements
Generate conditional logic with correct boolean expressions
Handle loop hierarchy and variable scoping
Select validation functions from a catalog
Write user-facing rejection messages
Format everything as valid JSON
The failures were interesting. The LLM would get the X12 mapping right but flip the boolean logic. Or nail the conditional logic but reference the wrong segment. It was trying to juggle too many constraints simultaneously.
So I broke it into four sequential steps:
Step 1: Classification & Decomposition Just understand what the edit is asking for. No X12, no DSL, no code. What makes this claim pass? What makes it fail? What's the validation trigger?
Step 2: X12 Mapping Now map the concepts to segments. Procedure code → SV101-2. CLIA number → REF*X4. Loop context. Element indices. This step uses semantic search against a vector database of X12 field mappings and loads the 10 most relevant ones per edit instead of dumping the entire X12 spec into the prompt.
Step 3: Logic Generation Convert the classification and X12 mapping into DSL logic. FOREACH loops. WHERE filters. REQUIRE expressions. This is where the validation functions catalog comes into play.
Step 4: Message Generation Generate the rejection message with proper variable interpolation. "Procedure {{procedureCode}} requires modifier {{requiredModifier}}."
Each step now has a focused prompt, independent retry logic, and clear success criteria.
The tradeoff: 4 LLM calls instead of 1. More tokens per edit. But when Step 2 fails, I don't have to re-run Steps 1, 3, and 4.
Early results look promising, but I'm still validating whether the decomposition actually improves quality or just makes debugging easier.
I spent two weeks convinced I needed specialized medical domain agents (Lab, Surgery, DME, Pharmacy). Turns out I just needed to stop asking one LLM call to do everything.
The pattern keeps repeating: when the LLM is struggling, decompose the task. Smaller prompts. Clear contracts between steps. Single responsibility.
A few weeks ago I was extracting edits with one monolithic call. Then I added the CanonicalEdit normalization layer. Last week I implemented sequential decomposition. This week I'm fine-tuning the X12 semantic search to handle edge cases.
I'm still not done. But the failures are getting much more specific, which means they're getting easier to fix.
How do you know when to break a complex LLM task into smaller steps versus trying to improve the single prompt?