The question of whether to fine-tune a local LLM with only 110 samples is the wrong one to start with. The real issue is that the approach itself is fighting the problem. This reader is trying to translate XQuery to SQL, a structured query translation task, and the instinct to throw more model capacity at it is understandable. But the evidence they have gathered, the parsing failures, the inconsistent prompt outputs, the missing columns on longer inputs, points to a fundamental mismatch between the tool and the task. A local LLM, no matter how well-tuned, will not reliably bridge the gap between two formal languages when the training signal is that sparse.
What the reader is experiencing is the difference between learning a pattern and following a rule. The parsing-based approach failed because it was too rigid. The prompt-engineering approach failed because it was too loose. Fine-tuning with a tiny dataset will likely produce a model that memorizes the 110 examples and then stumbles on anything new, which is exactly the sensitivity to variation they have already observed. The missing conditions and columns are not bugs in the model; they are symptoms of a solution that is trying to be a general translator when the problem is actually a specific, constrained mapping.
A more effective direction would be to treat this as a schema-driven problem, not a language-model problem. The enterprise context means the underlying data model is known. The XQueries and SQL queries are both operating against the same logical structure. Instead of asking the LLM to infer the translation from scratch, the reader could invest in a hybrid approach. Use the parsing script to extract the key components, but do not rely on it for correctness. Then feed those extracted components, along with the schema definition, to the LLM as a structured prompt that asks it to fill in the gaps, not generate the whole query. This reduces the task from open-ended translation to constrained generation, which is far more forgiving of small datasets.
The reader should also reconsider the fine-tuning plan. QLoRA on a 7B model is a reasonable choice, but with 110 samples, the risk of overfitting is severe. A better use of that dataset would be to generate synthetic variations programmatically. Change the order of conditions, rename variables, alter the depth of the XPath expressions. This expands the dataset without needing new human annotations. The goal is not to make the model smarter; it is to make the model more robust to the exact variations that are currently breaking it. The reader is not wrong to use an LLM, but they are wrong to expect it to do the heavy lifting alone. The practical takeaway is this: stop trying to make the model a translator. Make it a tool that works within a well-defined boundary, and let the structured parts of the problem do the heavy lifting.