services/orchestrator/src/react_agent/prompts.py.
This page covers the prompt side. The data those prompts are filled with —
.csproj/pom.xml configs, schema mappings, harness skeletons — is covered in Context Engineering. The two work together.1. One Prompt Per Pipeline Stage
UOM is a multi-stage LangGraph pipeline, and each stage has its own purpose-built system prompt rather than one giant prompt doing everything. This keeps every model call focused and short except where length is unavoidable.
The translation prompt is by far the largest and most important — and, as §5 explains, the main reason a run takes minutes.
2. The Dynamic System Prompt Builder
The translation stage callsbuild_system_prompt(state), which assembles the prompt at runtime instead of reading a fixed string:
- Persona + hard rules — “You are a Universal Object Mapping architect…” plus the non-negotiable contract.
- Injected framework configs — the literal
.csprojandpom.xmlfor the source and target (get_framework_config_content). - Curated few-shot examples — hand-verified input→output pairs for the relevant paradigm.
- Validation-entrypoint snippets — full, compilable harness skeletons for the source and target (
get_snippet_content).
2.1 The “core translation contract”
The rules are written as an explicit, numbered contract so the model cannot drift. The most consequential ones:- Rule 9 (“this code will be executed”) reframes the task from “write plausible code” to “write code that survives a compiler and a live database” — see Why You Can Trust the Translation.
- The framework rules hard-code the Template-API decision (Design Decisions §3) directly into the model’s instructions.
2.2 Why configs are injected verbatim
Rather than telling the model “use a recent Spring Data MongoDB”, UOM pastes the actual project file it will be compiled against:3. Structured Output: The Model Fills a Typed Form
The translation node does not ask for free text. It uses a strict Pydantic schema (TranslationOutput) so the model returns clearly separated fields:
translated_schema_code,translated_query_code— the deliverables.source_validation_*,target_validation_*— runnable harnesses + their declared entry-point class names.
_create_translation_output_model(state) uses Pydantic’s create_model to exclude fields that aren’t needed for the chosen translation_type:
model_validator checks that each declared entry-point class name actually appears in the generated harness code, rejecting structurally broken output before it ever reaches a sandbox.
4. Temperature 0 and Self-Repair Across Turns
4.1 Why temperature 0
Every generation and evaluation model runs attemperature=0. For creative writing you want variety; for code translation you want the opposite:
- Determinism / reproducibility — the same input produces the same translation, which is essential for a research project running evaluation experiments and for debugging.
- Fewer “creative” deviations — at higher temperatures the model is more likely to invent an alternative API or restructure a query. Here, the closest faithful translation is always the right one.
- Stable retries — when the loop feeds an error back (below), temperature 0 means the model changes its output because of the feedback, not because of random sampling noise.
4.2 How the model repairs its own mistakes
The first attempt is not always correct — a query may compile but return slightly different data, or fail to compile. UOM is designed so the model fixes itself on the next turn instead of starting from scratch. When validation or evaluation fails, the graph loops back togenerate_translation_node, and it re-invokes the model on the accumulated translation_messages rather than the original prompt:
translation_messages thread now contains the concrete failure evidence: the compiler stderr, or the DeepDiff payload showing exactly which field/count differed, or the judge’s rejection reason. So on turn 2 the model is no longer guessing — it is reading “javac error: cannot find symbol OrderItem” or “values_changed: root['unitPrice'] 48.5 → 48.0” and correcting that specific defect.
This converts the LLM from a one-shot generator into a closed-loop, self-correcting one. The retry budget is capped (MAX_TRANSLATION_LOOPS = 3); after that the run pauses for a human rather than looping forever. The isolated translation_messages thread (separate from the user-facing messages) keeps all this noisy debug back-and-forth out of the main conversation — see State & Context.
5. Why It Takes ~12 Minutes
A full translation run typically takes around 12 minutes. That feels slow next to a normal chat completion, and the reasons are worth understanding because they are mostly deliberate trade-offs in favour of correctness.5.1 generate_translation is the slowest stage — because its prompt is the largest
The translation node’s prompt is, by a wide margin, the biggest in the pipeline. It contains:
- the full persona + multi-section rule contract,
- two injected project files (source
.csproj+ targetpom.xml), - four to five few-shot examples of complete schema/query translations,
- two full validation-entrypoint skeletons (source and target), which are long, real programs.
kimi-k2.6, with thinking-capable fallbacks like deepseek-v4-pro-thinking) are large reasoning models chosen for accuracy over speed.
5.2 The other time sinks
In short: the latency buys you the correctness guarantees. UOM is not optimised to answer fast; it is optimised to only ever return a translation that compiled with the real toolchain and produced equivalent data against a real database. Most of the 12 minutes is spent proving the answer, not guessing it.
6. Related Reading
Context Engineering
What fills these prompts: configs, mappings, and harness skeletons.
Design Decisions
Why Template APIs, these frameworks, and why you can trust the output.
Architecture & LangGraph
The full state machine these prompts run inside.
prompts.py Reference
The annotated source for every prompt described here.