Concurrent Image Understanding and Generation

When AI draws one path while describing another, CO₂Ju keeps them consistent

A model describes the correct maze solution while drawing a different path.

3 min readMachine Learning
When AI draws one path while describing another, CO₂Ju keeps them consistent
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

Here's a truth that will save you time: generating text and an image in the same model does not guarantee they agree. A new paper from Google, Google DeepMind, and Stony Brook University, presented at NeurIPS 2026, exposes this gap with a simple but damning test, a maze. The model can describe the correct solution while drawing a completely different path. That kind of inconsistency isn't just a research curiosity; it's a fundamental barrier to trusting AI for tasks where the words and the picture must match. This is a different problem than the one Apple is tackling with its Apple has a new way to prove your iPhone photos aren’t AI slop initiative, which focuses on verifying whether a human-made photo has been altered. Here, the challenge is internal: can the model itself ensure its own outputs are coherent? The researchers' solution, a sampler called CO₂Jump, offers a clear path forward.

The core insight is refreshingly practical. The mismatch happens because the model generates text and image in parallel, without a mechanism to correct one based on the other as the process unfolds. CO₂Jump fixes this by using the model's own text confidence scores and cross-modal attention to guide image updates during each denoising step. If a token in the text has low confidence, the sampler allows it to be masked again and regenerated, effectively letting the model revise earlier decisions. This isn't a new model that requires retraining, it's a smarter sampling method applied to existing fine-tuned models. On puzzle benchmarks like maze solving and nonograms, where both the textual answer and the generated image must be correct, CO₂Jump was the only sampler that improved monotonically across 8 to 512 sampling steps. This stands in contrast to the broader image-generation capabilities showcased in 5 ChatGPT 2.5 Features to Try Today!, where the emphasis is on creative editing rather than correctness under constraint.

What does this mean for you? If you rely on generative AI for anything beyond creative inspiration, technical documentation, instructional diagrams, data visualizations, or any workflow where the output must be verifiably correct, this paper points to a specific, actionable improvement. The fact that consistency degrades without a self-correcting mechanism means that current tools are likely less reliable than they appear. CO₂Jump is still a research artifact, but its architecture is lightweight: one model forward pass per denoising step, no extra training. That makes it a plausible candidate for integration into existing pipelines. The authors also introduced three new datasets, JEdit-1M, JMaze-200K, and JNono-200K, to evaluate this specific kind of joint accuracy, giving the field concrete benchmarks to measure against.

The open question worth watching is how far this approach scales. The paper suggests other tasks where text, image consistency can be evaluated, but the real test will be moving from puzzles to messy real-world data. If CO₂Jump can handle a diagram with a dozen labels or a multi-step instruction set as reliably as it handles a maze, we may be looking at the next standard for how multimodal models are sampled, not just trained.

From Machine Learning

Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University.

We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent.

Read the original at Machine Learning