fine tuning

From Codebase to Dataset: Fine-Tuning AI on Your Own Projects

Turning a well-built React or Next.js project into a fine-tuning dataset is a puzzle worth solving, especially when you're aiming for quality over generic code dumps. The real challenge isn't just extracting files; it's…

4 min readMachine Learning

Most developers have looked at a well-built codebase and felt the quiet pull of possibility. What if the patterns, the component logic, the careful styling decisions, could be distilled into something more useful than just an app? The question of how to turn a React or Next.js project into a fine-tuning dataset is one of the more practical and forward-looking we have seen in a while. It is not about building another wrapper or chasing a benchmark. It is about mining your own work for the raw material that makes AI models genuinely better. The user is asking about context between files, screenshots paired with code, and generating instructions that are not trivially generic. That is exactly the right set of problems to be wrestling with.

We would tell this developer what they already suspect: the tooling for this is still nascent, and the real challenge is not extraction but framing. A codebase is not a dataset. It is a dense web of decisions, and the value lies in turning those decisions into explicit instructions. For instance, a component that handles optimistic UI updates is not just a component. It is an answer to the prompt, "How do you update a list item without blocking the user?" The user is also thinking about screenshots, which is astute. Visual context is often the missing piece in code generation. This aligns with what we have seen in other explorations of real-world systems, such as the discussion on Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where the gap between training data and deployment realities is a recurring theme. The same logic applies here: a model trained only on code fragments without visual grounding will struggle with the nuance of a polished UI.

The question about benchmarking is where the user might be getting ahead of themselves, but in a good way. They mention a new architecture that could improve speed and quality while using less VRAM. That is ambitious, and it is the right instinct to want a solid dataset to test it against. But we would caution against waiting for perfection. Start with a small, curated set of instructions from one or two projects. See what the model learns from it. Iterate. The user is not likely to find a single tool that does this end-to-end. They will need to build a pipeline that combines static analysis, manual annotation, and perhaps some clever prompting to generate the instruction side of the pair. This is similar to the thinking behind Evolve Your Recommendations: Real-World Insights on Adaptive Systems, where the point is that the complexity lies not in the architecture but in how you adapt to new data. The same applies here. The workflow is the product.

The practical takeaway is this: do not wait for a perfect dataset generation tool. Take one component, write five instructions that describe what it does and why it is built that way, and test it. That will tell you more about your format than any theoretical framework. The user mentions a paper or repo that does this, but the honest answer is that they are likely to be the first on their team to build this. That is not a setback. It is an opportunity to define what a good coding dataset looks like. And if their new architecture really does deliver the speed and VRAM improvements they are after, having a custom benchmark built from their own code will be a far more convincing proof than any generic leaderboard. The next step is not to find a tool. It is to start writing instructions.

From Machine Learning

I have a few web projects with pretty good UI/UX and I’m wondering if there’s any tool or workflow that can turn an existing codebase into a dataset for fine tuning.

For example, given a React/Next.js project with components, pages, styling, etc. or a static html site, I’d like to turn it into something like:

Read the original at Machine Learning