machine learning

Shared baselines across papers keep your research honest and efficient.

Running the same baseline models across two papers is a practical move, but it raises a fair question about duplication.

4 min readMachine Learning

There's a quiet kind of pressure that comes with preparing two papers at once, especially when you're trying to keep the science honest while also getting the work out the door. The question from Jealous_Key_4030 is simple, but it cuts to the heart of how we think about originality in research: if you run the same baseline models once and report those numbers in two different papers, are you committing plagiarism? The short answer is no, but the longer, more interesting answer is about how we frame the work we reuse. Plagiarism is about taking someone else's ideas or words without credit. Baselines are not your contribution, they are the measuring stick you use to show your own model's value. If you ran the tree and neural network models, you own those results. Reusing them across papers is not theft; it is efficient, transparent, and scientifically sound. The real issue is not the numbers themselves, but whether you are transparent about where those numbers came from and how you are presenting them to each journal.

What makes this situation more nuanced is the unspoken rule that papers should stand alone. When a reviewer sees a nearly identical RMSE table in two submissions, their first instinct might be to question the novelty of the work. But that instinct is often misguided. The baseline is not the contribution. Your proposed model is. If you clearly state in each paper that the baselines were run once and are shared across related work, you are being honest. The problem only arises if you try to hide that connection or if you imply that the baselines were run specifically for that paper. So, our advice is direct: do not run the baselines twice. That would be a waste of time and compute. Instead, write a brief note in your methodology section, something like, "The baseline results are reproduced from our prior work [reference], where the experimental setup is described in full." This is standard practice, and it protects you from any accusation of double submission or self-plagiarism because you are not reusing prose or analysis, only the raw output of a shared experiment.

This debate connects to a broader tension we see across the field. We recently explored how Cloudflare's Blog Finds Performance Gains with EmDash, Its New CMS and how teams often mistake tooling for transformation. The same logic applies here: the spreadsheet of numbers is not the science, the interpretation is. Similarly, when we looked at the Forrester Function as a tool for machine learning, we saw that the value is in how you apply the function, not the function itself. And in exploring real-world computer vision deployments, we saw that the same model can be described differently depending on the constraints of the deployment. The pattern is consistent: the field rewards clarity about what you actually did, not the illusion of doing everything fresh each time.

So, to the researcher who is worried about downvotes and accusations: you are not cheating. You are being pragmatic. The scientific community needs more of that. The specific takeaway here is this: report the shared baselines, cite them clearly, and focus your energy on making the proposed model's story compelling. If a reviewer asks, you have a clean answer. If they don't, you have saved yourself days of redundant work. The only real mistake would be pretending the baselines were unique to each paper. That is where the line between efficiency and deception lives. Watch that line, and you are fine.

From Machine Learning

Suppose I create two machine learning models suppose tree and neural network for a task let's suppose regression problem, now suppose I am sending both of this paper to two different journals, now the thing is the baseline models I need to only run once because I have reported same baseline in both papers, so the RMSE tables looks exactly same except the proposed model, does it lead to any problems like palgiarism??

Edit : I don't know why I am getting downvotes

Read the original at Machine Learning