There is a quiet wisdom in Runway's decision to stop fighting a bug and start wrapping it in a feature. For weeks, the team watched AI-generated avatars drift off-center during real-time video generation, a stubborn flaw that resisted every backend patch. The eventual fix was not a technical breakthrough but a product decision: add a frontend toggle that automatically re-centers the user's image before generation begins. Unlock ChatGPT for Work: A Practical Guide to Getting Started and Showcase Your AI Skills: 10 Projects to Build Your Portfolio both speak to the same underlying truth: the gap between a model's raw capability and its practical usefulness is where most teams actually win or lose. Runway's head of enterprise product, Ryan Phillips, framed this as a core lesson for any company building on top of foundation models, and he is right to do so. The deeper point is about evaluation culture. Phillips walked through how Runway builds its quality sets, and the details matter more than the novelty. They run internal workshops where cross-functional teams review generated examples together, aligning on what "quality" actually means down to the pickiest visual artifacts. They track pass rates in a simple Excel spreadsheet, categorizing outputs as minor or major failures against a predetermined threshold. No magic, just discipline. This is the part that most enterprises skip. They assume evaluation is an engineering task, when it is actually a shared language problem. If your product team cannot articulate why a generation failed, you will never build a reliable system around it. And when you do build that evaluation set, you have to include the edge cases that make you uncomfortable. Runway tested with "Tooth," a non-human character with no nose and unusual teeth, specifically to see how the model behaves when pushed beyond standard human facial structures. That is the kind of rigor that separates a demo from a product. The drift bug itself is a masterclass in turning limitations into leverage. Instead of continuing to pour weeks into a model-level fix, Runway asked a simpler question: what if the user's input image is the problem? They discovered that perfectly centered inputs produced stable video. So they built "Optimize for Image Quality," a feature that silently re-centers the image before generation. The user perceives a helpful tool; the company avoids a costly engineering detour. This is not a hack. It is a philosophy. Phillips advised the audience to turn model limitations into product features, because customers do not see the internal struggle, they only see the outcome. That is a hard lesson for engineers who are trained to fix root causes. But in real-time generative systems, where non-determinism is a fact of life, the pragmatic move is often to design around the flaw rather than eliminate it. The same logic applies to the infrastructure side: when 8% of API calls dropped to 16 frames per second, the fix was not a config change but physically replacing GPUs in a single data center in us-east-1. The lesson is that model performance is inseparable from hardware health, and you need deep observability to catch those anomalies. What should a reader take from this? The most actionable takeaway is this: your evaluation set is a product artifact, not a technical chore, and your next feature might already be hiding inside a bug you are trying to kill. Runway's experience suggests that the path to shipping faster is not more compute or better models alone, but a willingness to reframe a limitation as a user benefit. That requires humility, because it means admitting the model has a flaw, and creativity, because it means finding a way to make that flaw invisible. The broader implication for enterprise teams is that the role of the creative is shifting from producing single assets to defining worlds and parameters that agents can generate from. That is a significant change, and it will not happen overnight. But the companies that start building their evaluation culture now, with the same seriousness as their infrastructure budgets, will be the ones ready for it. The detail to watch is how Runway handles the next inevitable bug.
natural language processing for spreadsheets
When a bug resists fixing, the smartest feature is a new path forward.
For weeks, Runway chased a bug that made AI avatars drift off-center during real-time video generation.
4 min readVentureBeat

Runway spent weeks trying to engineer its way out of a stubborn bug: AI-generated avatars would drift off-center during real-time video generation. The fix wasn't a back-end patch — it was a new front-end feature that just worked around the problem. That's the kind of lesson Ryan Phillips, head of enterprise product at Runway ML, walked through at VB Transform 2026, arguing that even companies not building foundation models themselves can learn from how Runway builds, evaluates, and ships them.