A data engineer at a manufacturing company builds engines, spends significant time on ML and deep learning tasks, and wants to publish a paper. The gap between shipping a working model in production and writing a research paper that passes peer review is wide, and the engineer's question about how to bridge it is the exact kind of practical problem our field needs more people to address. Too often, industry practitioners assume their work lacks academic merit because it wasn't done in a lab, while academics struggle to validate their models on real-world constraints like latency, cost, and maintenance. This is a mismatch that hurts both sides.
The engineer's first instinct should be to look for the constraint that made the project hard. In manufacturing, that constraint is almost never "we needed a better architecture." It is usually something like: we had to detect anomalies with only ten labeled examples per engine type, or we needed inference to run on a microcontroller with 256 KB of RAM, or we had to update the model weekly without causing downtime. Those constraints are research questions. They are exactly the kind of novelty reviewers look for, because novelty does not always mean inventing a new attention mechanism. It can mean demonstrating how to adapt existing methods to a previously unstudied domain with hard operational limits. The engineer should write down every rule of thumb and workaround they developed to make the model work in production. That list is the draft of a methodology section. For a deeper look at how the field sometimes chases architectural novelty while ignoring practical constraints, see our piece on Neural architecture search promised progress, but transformers emerged elsewhere. For guidance on what to expect when your work enters the public discussion before final acceptance, read Navigating ICLR's open review timeline: when your work goes public.
The second step is to reframe the engineering deliverable as a research question. The engineer likely built a system that predicts engine failure, optimizes fuel consumption, or classifies component defects. In a paper, that system is not the contribution. The contribution is the insight about *why* the system works, or the trade-off that had to be made. A good question to ask is: what did we learn about the problem that was not obvious from reading the existing literature? If the answer is "we got 94% accuracy," that is not a paper. If the answer is "we found that synthetic data from a physics simulator generalizes better to new engine models than real data from old ones," that is a paper. The difference is that the second statement is a claim about the nature of data and transfer learning in a specific industrial context, and it can be tested and reproduced by others. The engineer should also consider that the bar for publication at applied venues like the Journal of Machine Learning Research or workshops at NeurIPS is often lower for novelty and higher for reproducibility and real-world impact than pure theory venues. For a practical guide on the infrastructure side of running experiments at scale, see Unlock LLM Training: A Practical Guide to Distributed Algorithms.
Here is the specific takeaway: the engineer should not start by writing a paper. They should start by writing a technical blog post or a short report for their team that explains one specific design decision and why it worked. That document, when shared with a colleague who asks "why did you do it that way instead of the standard approach?", is the seed of a publishable paper. The academic writing and formatting come last, not first. If that document cannot answer the "why" question with something beyond "it performed better," then the project is not ready for publication yet, and that is fine. The work itself still matters.