This dilemma captures the essence of what holds market risk reporting back: the false choice between mathematical integrity and adaptive language. The answer is not a compromise between the two, but a deliberate separation of responsibilities. Let the math be math, and let the language be language.
The problem described, thousands of trades, variable frequencies, attribution that must be precise, is not a problem for large language models. It never was. Feed raw numbers to an LLM and ask it to calculate variance attribution, and you are asking for plausible-sounding errors. The hallucination risk is real and unacceptable in financial reporting. The deterministic layer must own the numbers completely. That means Python and Polars, or any equivalent, doing the heavy lifting: calculating shifts by asset class, by region, by individual trade, and outputting structured data that is verified and repeatable.
Where the LLM earns its place is in the translation of that structured data into narrative. Once you have a clean, audited data frame of deltas and drivers, the summarization becomes a constrained mapping task, not a creative one. The model receives the numbers, the hierarchy of contributions, and a set of templates for how to express attribution. It does not invent. It formats. The risk of hallucination drops dramatically when the model is not asked to reason about the numbers, only to arrange them in human language.
The rigidity concern is real, but the answer is not to let the LLM write and execute code dynamically in a sandbox. That introduces an unacceptable attack surface in financial infrastructure. The better pattern is to build a flexible deterministic layer that can consume new attribution dimensions, asset classes, regions, counterparties, as configuration, not as hardcoded ETL. Define the attribution tree in a schema that can be updated without touching the pipeline. The LLM's context then includes that schema plus the aggregated data. If a new business scenario introduces a "green bonds" attribution, you add it to the schema, recalculate, and the LLM receives the new dimension in its context prompt. No code rewrite, no sandbox execution, no hallucination risk.
The frameworks that support this, LangChain, LlamaIndex, are useful for managing context windows and prompt templates, not for substituting the deterministic layer. Use them to structure the prompt that feeds the verified data to the model. But keep the math separate. That is the architecture that scales without drifting into unreliability. Precision is non-negotiable. Flexibility is achievable. They are not in conflict when each tool does what it does best.