Mastering Agentic Apps With Tiger Teams and Smart Evals

In this engaging podcast episode, Shane Hastie, Lead Editor for Culture & Methods, sits down with Sam Bhagwat, co-founder and CEO of Mastra, to explore the evolving landscape of AI engineering.

3 min readInfoQ
Mastering Agentic Apps With Tiger Teams and Smart Evals

The discipline of building agentic applications is maturing, and the conversation with Sam Bhagwat makes one thing clear: the organizations that succeed are not the ones with the smartest individual engineers, but the ones that rethink how teams are structured around the problem. The era of siloed development is over, and if you are still treating AI features as a solo coder's side project, you are already behind.

The core insight here is the "Tiger Team" structure, a cross-functional group that blends engineering, product, and design into a single unit focused on a specific agentic outcome. This is not a buzzword reshuffle. It is a practical response to the reality that agentic apps fail not because the model is weak, but because the context around it is fragmented. When you have a small, dedicated team that owns the entire loop from prompt design to evaluation, you eliminate the handoffs that kill momentum and introduce subtle errors. For our readers, this means a shift in how you staff projects. Stop trying to hire a "prompt engineer" and start building a unit that treats the agent as a product, not a script.

Equally important is the emphasis on evals as the backbone of this new workflow. Traditional testing checks if code breaks; evals check if the agent behaves correctly in ambiguous, real-world scenarios. Bhagwat's point is that you cannot ship reliable agents without a robust evaluation framework that is as carefully engineered as the app itself. This is where most teams stumble. They build a prototype that works on three examples and then wonder why it fails in production. The practical takeaway is to invest in your eval suite early, make it part of your continuous integration pipeline, and treat a regression in eval scores with the same urgency as a broken build. That is the only way to move from demo to dependable.

The broader implication is a direct challenge to the lone-wolf developer myth. If you are a technical lead or a CTO, your job is not to write better prompts; it is to create the conditions where cross-functional collaboration can flourish. That means rethinking job descriptions, reworking sprint cycles, and accepting that the person who understands the user's workflow is just as critical as the one who understands token probabilities. The future of AI engineering is not about more powerful models; it is about more intelligent teams that can harness them with discipline. Start by forming one Tiger Team on a single, well-scoped agentic use case, and build your eval harness before you write the first line of agent logic. That is the concrete first step.

From InfoQ

In this podcast Shane Hastie, Lead Editor for Culture & Methods spoke to Sam Bhagwat, co-founder and CEO of Mastra, about building and sustaining open source communities, the emerging discipline of AI engineering and evals, and how cross-functional Tiger Teams are key to shipping agentic applications.

Read the original at InfoQ