AI Agents

Balancing precision and cost to feed AI agents the right data

Feeding enterprise data to token-hungry AI agents isn't just about scale; it's about precision.

3 min readInfoQ
Balancing precision and cost to feed AI agents the right data

Fabiane Nardon's presentation on architecting enterprise data for AI agents lands at a moment when most teams are still treating token costs as an afterthought. She walks through TOTVS's approach to feeding large language models from transactional systems, and the core tension she names is one every serious practitioner will recognize: deterministic logic and non-deterministic LLMs do not naturally coexist. You need precision for financial records, security for customer data, and cost control across millions of daily operations. These are not optional constraints. They are the difference between a pilot and a production system.

What stands out in Nardon's method is the emphasis on data mesh and low-latency database architectures as the foundation, not the LLM itself. That is the right instinct. Most teams start with prompt engineering and then wonder why their context windows balloon and their bills spike. Nardon flips the sequence. She builds the semantic ontology first, then uses dynamic MCP tool selection to pull only the data the agent actually needs. That is a quiet but significant shift in mindset. We have been telling readers for months that the Scale AWS Server Deployments Effortlessly with Stateless Model Context Protocol is where the protocol is heading, and Nardon's work reinforces that direction. The stateless approach reduces overhead, but the real win is in how it lets you swap tools in and out based on the task, which is exactly what she describes.

The practical takeaway here is not about any single database or ontology. It is about the discipline of measuring token overhead as a first-class engineering metric. Nardon's talk implies that if you are not tracking how many tokens each agent consumes per successful transaction, you are flying blind. That is a concrete, quotable principle: "Optimize context windows before you optimize prompts." We would add that the related work on Unlock LLM Training: A Practical Guide to Distributed Algorithms shows a parallel pattern on the training side, where distributed systems force you to think about data locality and communication overhead. The inference side is no different. Nardon's semantic models are a form of data locality for tokens, making sure the model only sees what is relevant and nothing more.

The question we would push back on is whether this level of architectural rigor is realistic for smaller teams. TOTVS has resources that most enterprises do not. But that does not diminish the lesson. The principles scale down even if the infrastructure does not. Start with one domain, map its ontology, measure token usage, and iterate. The open question to watch is whether MCP tool selection becomes standardized enough that teams can adopt it without building custom middleware. If that happens, Nardon's approach stops being a case study and becomes a template. For now, the honest take is that her presentation gives you a roadmap, but the journey still requires you to build the car.

From InfoQ

Fabiane Nardon shares how TOTVS prepares enterprise data for token-hungry AI agents. She discusses balancing deterministic logic and non-deterministic LLMs across precision, security, and cost. Nardon details using data mesh, low-latency database architectures, semantic ontologies, and dynamic MCP tool selection to optimize context windows and reduce token overhead in transactional systems.

Read the original at InfoQ