2 min readfrom Machine Learning

First A submission (AAMAS): how much theory is enough when your experiments went sideways? [D]

Our take

Navigating the complexities of empirical MARL research, particularly under A* submission deadlines like AAMAS, often demands a careful balance between experimental rigor and theoretical grounding. A 2nd-year PhD candidate currently facing this challenge highlights a common predicament: experiments yielding nuanced results and a subsequent struggle to formulate robust theory. Recognizing the potential pitfalls of HARKing and data anomalies, the post seeks advice on acceptable theory depth at A* venues and strategies for salvaging a project timeline.

This post from a second-year PhD candidate facing their first AAMAS submission deadline strikes a chord with many researchers navigating the often-turbulent waters of empirical machine learning. The core issue—the tension between rigorous theory and compelling experimental results—is a perennial challenge, particularly in the multi-agent reinforcement learning (MARL) space. It highlights a fundamental question: how much theoretical grounding is truly necessary to accompany a strong empirical contribution at a top-tier conference? The candidate’s predicament, compounded by the unfortunate discovery of parameter errors and a potential HARKing situation, underscores the pressures faced by researchers, especially those operating under strict timelines. It’s a situation many can relate to, and the desire for guidance is entirely understandable. We've seen similar anxieties reflected in discussions about leveraging tools like [Claude Code for Research Papers [R]] to streamline the writing process, demonstrating a broader trend toward optimizing research workflows under pressure.

The candidate’s honest assessment of their situation—acknowledging the potential for training-dependent theories and the pitfalls of post-hoc rationalization—is commendable. It reflects a level of self-awareness that is crucial for maintaining scientific integrity. The question of whether strong empirical characterization can compensate for a limited theoretical sketch is a valid one, and the answer likely depends on the specific community and the novelty of the findings. While A* venues generally value theoretical rigor, they also recognize the importance of demonstrating robust empirical results. A plausible, well-supported phenomenon, even without a complete theoretical explanation, can still be valuable, particularly if it opens new avenues for future research. This resonates with the challenges discussed in Kasia Trapszo's presentation on [Presentation: From DVDs to Global Streaming: How Netflix’s Commerce Architecture Actually Evolved], which illustrates how a focus on practical functionality and iterative development can sometimes precede a full theoretical understanding. The core issue is to present the empirical work honestly and transparently, acknowledging the limitations of the theoretical explanation.

The advice regarding recovering from HARKing mid-project is particularly pertinent. It’s a situation that can feel overwhelming, but it’s crucial to be upfront about the process. Documenting the shift in approach, acknowledging the initial hypotheses, and explaining how the data informed the subsequent theoretical exploration can demonstrate intellectual honesty and rigor. Reframing the narrative to emphasize the empirical contributions and the insights gained, even if the initial theoretical goals were not fully realized, can be a viable strategy. Furthermore, the mention of undocumented public repositories and hidden parameters is a critical reminder of the importance of code reproducibility and rigorous experimental controls. The work done by AWS on [AWS Introduces Specification Driven Composition for Flexible Data Workflows] highlights the value of structured, transparent systems, a principle that should be applied to research codebases as well. Addressing these issues proactively, even if it means delaying publication, can ultimately strengthen the credibility of the research.

Ultimately, this situation highlights the complex interplay between theory and empiricism in modern AI research. While a strong theoretical foundation remains highly valued, the ability to generate novel empirical findings and provide plausible explanations—even if incomplete—can still contribute significantly to the field. The key is transparency, intellectual honesty, and a willingness to adapt and reframe the narrative in light of new evidence. The question now is: will the community continue to prioritize theoretical elegance, or will we see a greater acceptance of empirically-driven discoveries, even those lacking a complete theoretical explanation, particularly as the field moves toward increasingly complex and data-rich environments?

Hi everyone,

2nd-year PhD candidate here staring down my first A* submission deadline (AAMAS 2027). I could really use some perspective on theory expectations, especially since I think I’ve methodologically painted myself into a corner.

The setup

My project started with a clean hypothesis: if architecture X is more robust than Y to perturbation A, and B is a strictly harder version of A, then the X > Y ordering should hold under B as well. I isolated three variables I suspected were driving the effect, ran experiments, and… got results that only partially support the hypothesis, with clear boundary conditions.

Where I got stuck

Trying to explain the “why” mathematically sent me down a theory rabbit hole. I ended up with two bad options:

  1. Claims tied to specific training outputs rather than structural/architectural properties, or
  2. Weak, hand-wavy speculations that feel like post-hoc rationalizations.

I’m pretty sure I fell into HARKing.. I started building theory after seeing the results instead of deriving predictions beforehand.

Furthermore, my codebase is built on an undocumented public repo, and I recently found a bunch of hidden parameters set to wrong values for my setting. I’m currently re-running everything, which is why I’m being vague about specifics. My “insights” from the first round are probably garbage.

My actual questions

  • For those who’ve reviewed for or published at AAMAS (or similar A* venues): how much formal theory is actually expected for an empirical MARL paper? Is “here’s the phenomenon, here’s the controlled experiments, here’s a plausible but incomplete theoretical sketch” a death sentence?
  • If the theory ends up being training-dependent rather than structural, is that a sign I should pivot to a lower-tier venue, or can strong empirical characterization + limited theory still fly at A*?
  • How do you recover from HARKing mid-project when you’re under pressure to publish in year 3/4 of a 4-year contract?

Any advice on how to salvage the timeline or reframe the narrative would be hugely appreciated.

submitted by /u/ham_bam0
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article