Step 1: Define your use case
- Use case title: Up to 82 characters.
- Who can join this use case?: Private (invite-only) or Public (visible to all users).
- What task should contributors work on?: The task, for example Object Detection. It filters the training datasets offered in step 2. See the full list of supported data types and tasks. If your use case is not yet supported, reach out to us at [email protected].
Step 2: Data & Evaluation
- Select your training dataset: The training datasets you ingested in the Prepare Data step, for the task you picked.
- Select your test dataset: The test and score lists appear after you pick the training dataset. The test list shows only test sets compatible with it. When the task has classes, the test set’s classes must be a subset of the training set’s classes. A test set missing from the list is not compatible.
- Select your score: The metric that ranks submissions on the leaderboard. The list holds the metrics in the tables below for your task. If your metric is not supported, reach out to us at [email protected].
- Upload EDA File (optional): A
.ipynbor.jsonfile, up to 25 MB, to help peers understand the data. Explore template use cases for inspiration.
- Description: The title. Edit it on the use case’s Overview tab with Edit as admin.
- Compute budget: 1 PF per peer. Change it on the Admin tab under Compute usage with Adjust limit.
- End date: None, so the use case runs until you end it. Set one on the Admin tab under Duration.
Supported metrics per data type and task
The tables below list the metrics you can pick as the score for each task today. Each metric uses either higher is better or lower is better sort order on the leaderboard. Some tasks report more metrics during training than you can pick as the score; the notes under those tables name them.Image Classification
Text Classification
Sentence Pair Classification
Training also reports precision, recall, F1 score, MCC and other classification metrics, but Accuracy is the only score you can pick today.
Token Classification
Token-level scores are computed over labelled tokens only: padding, special tokens and sub-word continuations are excluded.
Training also reports token accuracy (same value as Accuracy) and precision, recall and F1 score, macro-averaged across tag classes. Accuracy is the only score you can pick today.
Masked Language Modeling
Causal Language Modeling
Causal language modeling is next-token prediction, so its scores are token-level. Each validation sequence is scored at every position that has a training target: the model’s most likely next token is compared with the actual next token. Padding is excluded, and so is the prompt part of aprompt<TAB>completion row, so only completion tokens count.
Seq2Seq
Seq2Seq scores are token-level over the target (decoder) sequence, with padding excluded.
Training also reports token accuracy (same value as Accuracy) and perplexity, but Accuracy is the only score you can pick today.
Embeddings
Embeddings are trained with a contrastive objective on anchor/positive pairs, so the scores measure the embedding space rather than class labels. An embeddings run reports these metrics during training, but you cannot pick them as the score today:Tabular Classification
Object Detection
For ranking detectors, pick Mean Average Precision @ IoU 0.50: unlike Accuracy and IoU, it penalizes both missed objects and false boxes.
Semantic Segmentation
Keypoint Detection
Tabular Regression
Time Series Forecasting
By default, when the data has feature columns besidestimestamp and the target, the engine scales the target to the 0–1 range (MinMaxScaler) before training, so error metrics such as MAE and RMSE are in scaled units, not in your data’s units.
Training also reports direction accuracy (how often the model predicts the direction of change correctly), but you cannot pick it as the score today.
Time Series Classification
Time series classification offers the same metrics as Tabular Classification, computed per sequence instead of per row.Time-to-Event Prediction
Next Steps
Once your use case is published, invite peers to train models on it. In the use case view, monitor- total resource consumption
- daily submits and user activity
- overall leaderboard and submissions
Need Help?
- Email us at [email protected]