The question lands at the heart of every comparative study: are we comparing architectures, or are we comparing our patience in tuning them? When you notice that optimal hyperparameters shift for every architecture and scenario pair in the VMAS library, you are not facing a nuisance to be eliminated. You are facing the data. The honest answer to whether you must unify is both yes and no, and the distinction matters more than the answer itself.
Methodologically, unifying hyperparameters is the only way to isolate architecture as the variable. If you tune each model to its personal best, you are no longer comparing PPO variants; you are comparing the results of a hyperparameter search. That is a different study, and it does not answer the question you started with. But your note that unification sometimes leads to non-converging models is not a flaw in your process. It is a finding. It tells you that some architectures are more sensitive to initialization and optimization conditions than others. That sensitivity is a property of the model, not an error in your experimental design. This is where the conversation connects to broader engineering decisions, much like how Scale AWS Server Deployments Effortlessly with Stateless Model Context Protocol reframes infrastructure choices as protocol-level decisions rather than configuration tweaks. The choice of what to hold fixed is itself the architecture.
What should you actually do? Run two tracks. The first track uses unified hyperparameters, and you report which models fail to converge. That is a legitimate result, and it is informative for practitioners who need to know which methods are robust enough to deploy without heavy per-task tuning. The second track can allow per-model tuning, but you must be transparent about the compute and search budget you gave each method. A model that needs three times more tuning effort to reach parity is not equivalent to one that works out of the box. This mirrors the insight from Bridging Retrieval and Action: A New Approach to AI Tasks, where the explicit connection between components revealed that the integration method, not the individual parts, determined performance. Your frozen-model adversarial attack objective makes this even more critical. If you unify hyperparameters and a model underperforms under attack, you need to know whether that fragility stems from the architecture or from the fact that you gave it suboptimal training settings. Your robustness conclusions will be meaningless if they are confounded by tuning artifacts.
The practical takeaway is this: do not unify for the sake of a clean table. Unify for the sake of a clear claim. Report both, and let the reader see the cost of each choice. And for your adversarial robustness tests, prioritize the models that converge under unified settings, because those are the ones worth attacking. A model that cannot be trained reliably under standard conditions is not a model you should be stress-testing for deployment. Watch for the gap between convergence and performance. That gap is where the real comparison lives.