Skip to main content
tracebloc is continually expanding supported data types and tasks to enable your use cases. In case your use case is not yet supported, please reach out to us at [email protected]. Before you can create a use case on the tracebloc website, make sure the following requirements are met:
  • You are registered as a user on the tracebloc platform
  • Your dataset is cleaned and preprocessed
  • You are familiar with the supported data types and tasks
  • To ingest the dataset, you have a secure environment running, locally or in the cloud. Preparing the files doesn’t need one.

Your path as a new user

  1. Sign up for a tracebloc account.
  2. Install your secure environment with the Quick Start. It also installs the tracebloc CLI.
  3. Prepare your data against the dataset template for your task. You can do this before step 2; it doesn’t need a secure environment.
  4. Ingest it with the tracebloc CLI, once as training data and once as test data.
  5. Define the use case, including its evaluation metric.
  6. Peers train models on it.
  7. Evaluate the models on the leaderboard.

Supported Data Types and Tasks

The exact folder layout, CSV columns, ingest.yaml and validation rules for each task are on the dataset template pages. Where this overview and a template page differ, the template page is authoritative — it is derived from the ingestor’s own checks.

Image Data

Requirements for all image data tasks: Every image in a dataset has the same width and height and the same file type, for example all images as 256×256 RGB .jpg files. Any aspect ratio works; images don’t have to be square. tracebloc doesn’t resize or convert images when it ingests them, so convert, crop or resize them before you ingest. At training time, images are resized to the input size of the model a peer trains. The ingestor doesn’t check color mode or bit depth. If you combine datasets, declare color_mode (RGB or grayscale) and bit_depth (8 or 16) in the template so they can be checked for consistency.
Filenames in the label CSV: For all image and text tasks, the filename column may include the file extension or not: cat01 and cat01.jpg both resolve to images/cat01.jpg when the configured extension is .jpg. The extension is configured once for the whole dataset (file_options.extension in the template).
All images are validated before ingestion by the data ingestor. The ingestion process only starts when every file meets the requirements. Fix or remove any invalid images, then retry. One tracebloc data ingest takes at most 1 GiB in total and 500 MiB per file; see Size limit for larger datasets.

Image Classification

The filename column may include the file extension or leave it out.

Image Keypoint Detection

The number of keypoints per image must be fixed: every row names the same keypoints, exactly as many as you declare for the dataset. You cannot mix classes with different keypoint counts (e.g., 16 for person and 32 for car) or annotate some images with fewer keypoints.
Each row is one image. Annotation maps each keypoint name to its [x, y] coordinates, and Visibility uses the same keys with 1 (visible) or 0 (occluded or outside the image). Both are JSON objects inside a quoted CSV cell, with the inner quotes doubled:
The image_label column needs at least two distinct values; the ingestor rejects a dataset with a single class. See the keypoint detection template for every rule the ingestor checks.

Image Object Detection

Images and their Pascal VOC XML annotations pair by file stem: images/street01.jpg belongs to annotations/street01.xml. There is no labels CSV; the image list and the classes are read from the XML files. See the object detection template for the full annotation format.
Every element shown is required; the ingestor rejects a file that leaves one out.

Image Semantic Segmentation

Each image has one mask: a PNG named <image stem>_mask.png, with exactly the same width and height as the image, whose pixel values are class indices 0 to N-1. N is the number of distinct image_label values in the dataset, so if 0 is background, background must appear as an image_label value too. The labels CSV has one row per image: it links the image to its mask and names a class present in the image. Training reads each row as one sample, so list each image once. See the semantic segmentation template for the full rules.

Tabular Data

Requirements for all tabular data tasks: Each dataset must be provided as a single CSV file with a header row. Every column must contain uniform data types, for example numeric values for features and a categorical or alphanumeric column for labels. Use UTF-8 encoding with comma separators and validate that your schema matches the expected types. Invalid rows are skipped by the ingestor.

Tabular Classification

Include a header row with clear column names, using a dedicated column for the labels. An id column is recommended but not required.

Tabular Regression

Same structure as Tabular Classification, but the label column holds a continuous numeric target (not a class).

Time Series Forecasting

Provide a single CSV with a timestamp column, one or more numeric feature columns, and the numeric target column you want to forecast. Rows must be ordered by time and use a consistent timestamp format.

Time to Event Prediction

Provide a single CSV with feature columns, a time column (duration to event or censoring), and a binary event column (1 = event occurred, 0 = censored).

Text Data

Text Classification

The filename column may include the file extension or leave it out. The extension is set once for the whole dataset (file_options.extension in the template).
file example

Next Steps


Need Help?