rows.com

A Smarter Path to Terminal UIs: A Tiny Model, Not a GPU Renderer

Terminal emulators have gotten absurdly complex, GPU glyph atlases, texture caches, custom shaders, all to draw a grid of characters really fast.

4 min readMachine Learning
A Smarter Path to Terminal UIs: A Tiny Model, Not a GPU Renderer
Instead of another GPU terminal renderer, I trained a 1.26M-param model to turn TUIs (htop, vim, emacs…) into real UI components [R]

There's a quiet revolution happening in the terminal, and it's not coming from a faster GPU or a cleverer shader. It's coming from a 5 MB model that doesn't even run most of the time. For years, the terminal world has been locked in an arms race. Alacritty, Kitty, WezTerm, Ghostty, each one pushing the boundaries of what a grid of characters can do. GPU glyph atlases, texture caches, custom shapers. It's impressive work, and it's solving a problem we've mostly accepted as permanent: that a terminal emulator must parse escape codes, maintain a cell grid, and paint pixels. Faithful, yes. But opaque. Your phone can't reflow it. A screen reader chokes on box-drawing characters. An agent has to parse visual formatting to understand which row is selected. We've spent decades optimizing a rendering pipeline that keeps us locked into a text-only view of the world. This project asks a different question. What if you didn't need a GPU to draw the terminal? What if, instead, you understood it once, server-side, with a small model, and sent the client something that actually looks like a UI? That's the bet behind Phosphene. A 1.26M-parameter model labels cells with roles, border, title, menu item, selected row, input, status bar. Deterministic code then maps those regions into A2UI components: lists, buttons, text fields. Once a layout is seen, it locks as a template. The model doesn't even run again; only changed content crosses the wire as JSON-pointer patches. When a user presses a button, the client sends the keystroke back. F10 is just a Button with an action. The numbers are honest. mIoU of 0.51 on held-out real screens is a solid first pass, not a finished product. 40% template hits mean the model often doesn't run at all. And yes, the A2UI stream is roughly 25× larger than raw VT. But that's missing the point. The win isn't bandwidth, it's that the client never has to be a terminal emulator. It's a UI renderer. This is a meaningful step toward accessible, agent-friendly terminal interfaces. The terminal has been a wall between users and their data. Phosphene tears that wall down by sending actual UI over the wire, not character streams. That's not just an incremental improvement; it's a fundamental shift in how we can interact with legacy systems. The implications extend beyond convenience. For accessibility, screen readers could finally interpret structure, not just read characters. For mobile, TUIs could become usable on touch interfaces. For AI agents, understanding a UI becomes a solved problem when the UI is already structured. We're not suggesting that the terminal is dead. Far from it. The terminal remains a powerful, universal interface. But Phosphene points to a future where the terminal's output is a suggestion, not a constraint. A future where the interface adapts to the user, not the other way around. It's a small step, but it's a step in the right direction. And it's a direction worth exploring.# Our Take: The Terminal's Quiet Evolution Isn't About Pixels Here's a confession: We've spent years watching terminal emulators get faster and more complex, and we've never once stopped to ask if they're getting smarter. The write-up on Phosphene makes that question unavoidable. It's a small project with a big provocation: what if we stopped trying to render a terminal, and started trying to understand it? The current state of terminal technology is genuinely impressive. GPU glyph atlases, texture caches, custom shaders, Alacritty, Kitty, WezTerm, and Ghostty are performing remarkable engineering feats. They draw a grid of characters at astonishing speeds. But as this project points out, that speed serves a format that remains fundamentally opaque. Your phone can't reflow it. A screen reader struggles to make sense of box-drawing characters. An AI agent has to parse visual layout to understand which row is selected. We've optimized the rendering pipeline to an extraordinary degree, yet we're still sending a stream of escape codes and hoping the client figures it out. The insight here is refreshingly simple: instead of making the terminal faster at drawing grids, why not make it smarter at understanding them? The approach, using a compact 1.

From Machine Learning

Write-up + demo: https://drksci.com/labs-phosphene

So this started as a bit of a gripe. Modern terminal renderers are seriously impressive and seriously complicated. GPU glyph atlases, texture caches, custom shaders, HarfBuzz shaping, ligatures, damage tracking, grid diffing, dirty-row uploads. Alacritty, Kitty, WezTerm and Ghostty are all doing heroic work to draw what is, at the end of the day, a grid of characters really fast.

Read the original at Machine Learning