Claims about model reliability depend on the experiments used to support them. In out-of-distribution detection, differences in datasets, score conventions, random seeds, hyperparameters, and implementations can make results difficult to compare or reproduce.
My work addresses this problem from both methodological and infrastructural
directions. It studies uncertainty and randomness in evaluation protocols,
documents practical reproducibility challenges, and develops
pytorch-ood as shared infrastructure for consistent
implementations and benchmarks.