
Sharing our experience at making LLM-based Source Code classifier

Sharing our experience at making LLM-based Source Code classifier

What to do when we train on X, inference on Y, and measure on Z

What to do what only unlabeled data is available

Yet another NLP? No No and No

3 possible pivots when a classification model training is not possible

Sub populations to take into account when sampling GitHub

From interview through task definition to evaluation

3 lean validations to make sure you're on the right path
