The terminology confusion here is understandable, but the answer matters more than most researchers realize. When you train a model exclusively on normal data and then calibrate a decision threshold using labeled validation examples, you are running a semi-supervised anomaly detection pipeline. The representation learning is unsupervised, yes, but the threshold selection injects label information into the final system. Calling the whole setup unsupervised risks understating how much supervision actually shapes the deployed model.
The practical implication is straightforward: your paper should be precise about where labels enter the process. Say what you did, step by step. The model learns from one class alone. The threshold is chosen to maximize F1 on a labeled validation set. That is not a minor detail, it is a defining characteristic of your approach. Reviewers and readers will appreciate the clarity, and you avoid the trap of overclaiming unsupervised status when your decision boundary is explicitly tuned with ground truth.
There is also a deeper point worth making. The distinction between unsupervised and semi-supervised here is not just academic pedantry. It changes how your results should be interpreted. If your threshold is calibrated on labeled normal and anomalous examples, then your evaluation already assumes access to both classes at some stage. That is a different problem setting than pure one-class learning, where the model never sees an anomaly until deployment. Acknowledging this does not weaken your contribution, it strengthens it by being honest about the assumptions.
For the paper, describe the representation learning as one-class or unsupervised, then state that threshold calibration uses labeled validation data. If you want a single term, semi-supervised is the more accurate umbrella, because some supervision is present in the overall workflow. But do not stop at the label. Explain the split, the metric, and why this choice makes sense for your domain. The field needs less hand-waving about learning paradigms and more precise reporting of where labels actually influence the system. Your setup is common, and it deserves a name that reflects that reality.