COMPLETED2023-09 → 2023-12
CompSem -- NLI
How do different sentence encoders trade off NLI accuracy and transfer performance?
NLPPythonPyTorchDeep Learning
I wanted a clean comparison of sentence encoders beyond single benchmark scores. I built a modular pipeline to train multiple encoders on SNLI and evaluate transfer with SentEval, including reusable checkpoints. The key finding was that higher NLI accuracy did not guarantee better transfer. I handled modeling, training, evaluation, and the write-up. The repository is here: Repo.