AAMAS

When Experiments Falter, Theory Can Guide Your Next Move

A second-year PhD candidate is staring down their first AAMAS submission with a familiar dread: the experiments only half-worked, and the theory that emerged feels like it was built backward.

4 min readMachine Learning

The pressure to publish in year three of a four-year contract does strange things to scientific judgment. We see it in this AAMAS submission story: a clean hypothesis, experiments that only partially cooperate, and a slow-motion realization that the theory was reverse-engineered from the results. This researcher is not alone in the HARKing trap, and the fact they are naming it publicly rather than quietly submitting anyway suggests they still care about doing this right. That matters, because the academic pressure cooker produces far more dangerous responses than honest uncertainty.

The core question here is not whether the experiments are perfect, they are not, and the hidden parameter issue in the undocumented repo is a legitimate fire. But the deeper question is whether A* venues like AAMAS actually demand complete formal theory for empirical MARL papers, or whether they reward well-scoped empirical characterizations that acknowledge their own limits. From what we observe across the field, reviewers increasingly value transparency over polish. A paper that says "here is the phenomenon, here are the boundary conditions, and here is a plausible but incomplete sketch" can survive, provided the experimental story is rigorous and the claims are scoped precisely. What kills submissions is overclaiming, not underclaiming. The researcher's instinct to pivot to a lower-tier venue may be premature; strong empirical work with honest theoretical limits often finds a home at top venues, especially when the community is actively debating reproducibility and generalization.

The HARKing recovery is the harder problem, and it deserves direct advice. You cannot unsee the results, but you can reframe the contribution. Instead of framing the paper as confirming a hypothesis, frame it as mapping the boundary conditions of robustness under perturbation, which is a legitimate and useful empirical finding. The theory then becomes a discussion section, not a core contribution. This is where the related pressure on medical students navigating their own gauntlets becomes relevant: both fields push early-career people to treat publication or matching as the only signal of worth, and both suffer for it. The same principle applies to scaling AI infrastructure, where reliability depends on admitting what you do not control. Acknowledging limits is not weakness; it is the foundation of trust.

The concrete takeaway worth quoting: "A paper that maps where a hypothesis fails is more valuable than one that pretends it never fails." The researcher should re-run everything, document the parameter corrections transparently, and submit the empirical story with a clearly labeled theoretical sketch. If the reviewers reject it, that is data, not a verdict on their worth. The timeline is salvageable, but only if they stop treating the theory as a shield and start treating the experiments as the contribution. The field needs more honest boundary mapping, not more confident overreach. Watch how they handle the resubmission; that will tell us what kind of researcher they intend to become.

From Machine Learning

2nd-year PhD candidate here staring down my first A* submission deadline (AAMAS 2027). I could really use some perspective on theory expectations, especially since I think I’ve methodologically painted myself into a corner.

My project started with a clean hypothesis: if architecture X is more robust than Y to perturbation A, and B is a strictly harder version of A, then the X > Y ordering should hold under B as well. I isolated three variables I suspected were driving the effect, ran experiments, and… got results that only partially support the hypothesis, with clear boundary conditions.

Read the original at Machine Learning