Overview
Training and evaluation read your data from your secure environment’s storage. Raw data never leaves your infrastructure. Getting it there takes two steps:- Prepare the files. Lay your data out to match the dataset template for your task (tabular, image, text, time series). You don’t need a secure environment for this step.
- Ingest them with
tracebloc data ingest. The data ingestor validates the whole dataset, copies it into your secure environment’s storage and registers only metadata with the platform. This step needs a running secure environment; set one up with the Quick Start.
1. Prepare the files
Open the template for your task and lay out one folder per split the same way. For image classification:- Tabular and time-series datasets are a single CSV file.
- Image and text datasets are a folder:
labels.csvplus the per-sample subfolder from the template (images/,texts/, …). Object detection has nolabels.csv; its records come from the Pascal VOC files inannotations/. - Every image in one dataset has the same file type and the same width and height. tracebloc never resizes or converts images when it ingests them, so do this before you ingest.
- Clean the data first. Peers can’t view, clean or fix raw data, so model quality depends on the data you provide.
Don’t augment images offline as well. For image tasks (classification, object detection, keypoint detection, semantic segmentation), peers can turn on augmentation such as rotation, shifts, zoom and brightness when they train. It runs on the training split only, at training time, and is off unless a peer turns it on. See data augmentation. If you also add augmented copies to your dataset, the two stack.
Where to get sample data
To try tracebloc before you use your own data, start from a public dataset whose license allows your use, and convert it to your task’s template layout. Large downloads from public dataset hubs can be slow or stall when you’re not signed in. Sign in to the hub (on Hugging Face, with an access token) and re-run the download. If a download hangs, stop it completely before you retry.2. Ingest with the tracebloc CLI
The tracebloc CLI is installed with your secure environment.tracebloc data ingest finds your secure environment, checks the dataset on your machine, copies the files into its storage, runs the data ingestor and shows progress until it finishes. You don’t run Helm, write ingest.yaml or use kubectl.
- Leave out the flags and the CLI asks for each value. Add
--dry-runto check the dataset without ingesting anything. --taskis thecategoryon your template page.--nameis how you pick the dataset later. It starts with a letter or underscore, followed by letters, digits and underscores. Use different names for training and test data.- Some tasks need one more flag:
--number-of-keypoints(keypoint detection),--label-policy(tabular regression, time-series forecasting and survival analysis; defaultbucket) and--time-column(survival analysis). Object detection takes no--label-column. - Image tasks read the image size from your first image and reject any image that differs. Pass
--target-size WxHto state the size yourself.
tracebloc data ingest --help or see tracebloc CLI.
Size limit
Onetracebloc data ingest takes at most 1 GiB in total and 500 MiB per file. The CLI checks this before it copies anything. If your dataset is larger:
- Save images as JPEG instead of PNG. For photos, JPEG files are usually several times smaller; a quality of 90 to 95 keeps them close to the original.
- Lower the resolution. Resize every image to one smaller size that still suits the models peers will train.
- Use fewer samples.
- Ingest with the Helm chart instead. That path doesn’t have the CLI’s size limit; the dataset only has to fit in your secure environment’s storage.
Dataset management interface
View your datasets at ai.tracebloc.io/data after they are ingested. Interface displays:- Dataset name, ID, and record count
- Data type (Tabular, Image, Text) and purpose (Training/Testing)
- Namespace and GPU requirements
Advanced: ingest with the Helm chart
Use the Helm chart when your data is already on the cluster, when you run a GitOps workflow, or when a dataset is over the CLI’s size limit. You describe the dataset in a shortingest.yaml and helm install it. The official ingestor image (published as ghcr.io/tracebloc/ingestor) runs it. No Dockerfile, no Python script.
Before you run any commands in this section: if you installed with the one-line installer (Append
curl -fsSL https://tracebloc.io/i.sh | bash), every later helm upgrade <namespace> tracebloc/client … (the installer names the Helm release after the namespace) must include --reset-then-reuse-values, otherwise the upgrade drops the values the installer applied and breaks the secure environment:--version <version-number> to pin a specific chart version. This caveat only affects upgrades of the parent tracebloc/client chart, not the helm install tracebloc/ingestor runs below.1. Add the chart repo (one-time)
tracebloc/client parent chart bootstraps the cluster (jobs-manager, MySQL, RBAC). The tracebloc/ingestor subchart submits per-dataset ingestion runs against it.
2. Stage your data on the cluster’s shared PVC
The chart doesn’t transport data into the cluster. It reads data already on the cluster’s shared PVC (client-pvc by default, mounted at /data/shared/ inside the ingestor Pod). Before installing, put your raw files in a per-dataset folder there, /data/shared/<prefix>/. The <prefix> becomes the path you reference in ingest.yaml.
Where that PVC lives depends on the storage mode your secure environment was installed with. On macOS and Linux the one-line installer prints it at the end: Data in-node (k3s local-path) means node-local.
-
Node-local (the installer’s default). Datasets live inside the cluster node (k3s local-path storage), not in a folder on your machine, and they are deleted with the cluster. There is no host folder to copy into: copy the files in through a throwaway Pod that mounts
client-pvc(kubectl cp). The client ingestor README has the Pod manifest and commands. -
Host path (installed with
TB_STORAGE_MODE=hostpath). The PVC is a folder on the machine where the secure environment runs:~/.tracebloc/<namespace>/data/, where~/.traceblocis the data directory (unless you setTRACEBLOC_HOST_DATA_DIR) and<namespace>is your secure environment’s namespace (tracebloc cluster infoshows it). Copy your files into it:
/data/shared/<prefix>/.... That’s what you put in ingest.yaml below.
For multi-node or EKS deployments, where the PVC isn’t on a single machine, use a throwaway
kubectl cp Pod or a cloud-storage init container. See the client ingestor README for those recipes.3. Write your ingest.yaml
The example below is for image_classification. Other tasks require different fields — e.g. tabular_classification has no images: and instead needs a typed schema: block. Don’t copy this one blindly; open the dataset template for your task — one page per task with the folder layout, the CSV columns, a ready-to-edit ingest.yaml and the checks the ingestor runs — and edit from there.
apiVersion, kind, category, table, intent, label) is the same for every category; the category field picks the validator set, file-extension defaults, and column conventions. The data-source fields (csv:, images:, schema:, …) vary per category. The paths are paths inside the ingestor Pod, which is the PVC mount you populated in step 2.
4. Install once per dataset
The ingestor runs once: validates your data, copies files into the destination directory on the PVC, inserts rows into MySQL, sends metadata to the tracebloc backend, then exits. Run it twice per dataset — once withintent: train, once with intent: test — using distinct table: names. The example below shows both releases:
helm install is a separate release (the first argument is the release name), so the two runs don’t collide. The ingestor Pod picks up CLIENT_ID / CLIENT_PASSWORD automatically from the Kubernetes Secret the parent tracebloc/client chart created in <namespace> at install time — you don’t pass credentials on the helm install command.
Full chart docs (data-staging recipe, schema, every category, update model, verification, override knobs) → client ingestor README.
Advanced: custom Python script
Use this flow when the declarative schema can’t express what your data needs — typically when you have non-trivial preprocessing logic, a custom validator, or aBaseProcessor subclass. The sections below — Quick Setup and Detailed Setup — both describe this advanced path.
Quick Setup
Use this quick setup if you already have an ingestor configured and just want to switch datasets or toggle between training and testing. If you are setting up for the first time, go to the next section for the detailed walkthrough.Steps
- Edit your ingestion script (the one you wrote in Configure a script below)
- Update csv options and data_path
- Only for tabular data: Update schema
- Set
schemaandCSVIngestor()parameters like category, intent, label_column, etc. to match data type, task and train/test purpose
- Build and push docker image:
- Edit ingestor-job.yaml:
metadata.name: Unique job name (e.g. ingestor-job-train and ingestor-job-test)image: The tag you built and pushedTRACEBLOC_LABEL_FILE: Path inside the pod to the labels CSV, under the PVC mount (e.g./data/shared/labels.csv). For tabular data, this is the same file that contains both labels and features.TRACEBLOC_TABLE_NAME: Unique table name (no spaces, one per dataset). Title is optionalTRACEBLOC_SRC_PATH: Root of the mounted dataset directory inside the pod (/data/shared, the shared PVC)
TRACEBLOC_ prefix still work; see Configure Kubernetes.
- Deploy to Kubernetes
Detailed Setup
1. Configure a script
This section walks you through the step-by-step setup of a data ingestor. You will install the ingestor package, lay your data out against the dataset template for your task, and write a short ingestion script that matches it. Follow this guide if you are setting up an ingestor for the first time or need full control beyond the quick setup.Install the ingestor package
The ingestion library is published on PyPI astracebloc-ingestor (import name tracebloc_ingestor). It is the same code the official ingestor image runs, so a script written against it behaves exactly like the declarative path. It needs Python 3.11 or newer:
Pick the dataset template for your task
Open the dataset template for your task. It gives you the folder layout and the labels CSV the ingestor expects — lay your data out the same way, then set the matchingcategory and data_format in your script. The rows below cover the most common tasks; every other task on the templates index works the same way, and its TaskCategory constant is the upper-cased category identifier from its page (for example keypoint_detection → TaskCategory.KEYPOINT_DETECTION).
High Level Script Structure
Every ingestion script follows the same structure — save it asingestor.py next to your Dockerfile:
ingestor_job.yaml.
config.LABEL_FILE: Path to local csv label fileconfig.BATCH_SIZE: Batch size used during ingestion
Customize the script
The structure above is a starting point, but every dataset has its own format and labels. In this step you adapt the script to your data by tuning CSV ingestion options and setting the ingestor parameters (category, label column, intent, data path and schema). The following example shows how to ingest a tabular dataset, but the setup works the same way for image or text data.Needed for Tabular Data: Define Schema
Define the dataset schema as a Python dictionary, mapping each column to its SQL type and constraints. Do not include IDs or the label column into the schema.Needed for Image Classification Data: Define Image Options
Define image size and file extension.Needed for Object Detection Data: Define Image Options
Define image size and file extension.Needed for Text Data: Define File Extension
Define file extensions.Set CSV ingestion options
Customize parsing, memory handling, and data cleaning with the csv_options dictionary:Set Up the Ingestor
Define the Ingestor instance with the required configuration. See the tabular data example below:category, choose the ML task type (TABULAR_CLASSIFICATION, IMAGE_CLASSIFICATION, OBJECT_DETECTION)label_column, target column or class labelsintent, set as TRAIN or TEST depending on dataset purpose- include
file_optionsorschemadepending on the data type
category, data_format and options for your task from the table above.
2. Build Docker Image
With your script configured, the next step is to package it into a Docker image so it can run inside the Kubernetes cluster.Docker Hub Setup (first-time users)
The cluster pulls your ingestor image from a public Docker registry, so you need an account before you can push. If you already have one, skip to Write the Dockerfile.- Create a Docker Hub account at hub.docker.com/signup and verify your email.
-
Log in from your terminal so the
docker pushcommand can authenticate: -
Push the data ingestor image to your account using the build/push commands in the next section. The image name takes the form
<your-docker-username>/<image-name>:<tag>— the username segment must match the account you just created. -
Make the image public so the cluster can pull it without credentials:
- Go to hub.docker.com/repositories, open the repository you just pushed.
- Click Settings → Visibility settings → Make public.
imagePullSecretnamedregcredin the client namespace (theingestor-job.yamlalready references it).
Place data files on the shared PVC
Datasets are not baked into the Docker image. They live on the shared PVC of your secure environment (client-pvc) and are mounted into the ingestor pod at /data/shared.
Stage your files there the same way as for the Helm chart: see Stage your data on the cluster’s shared PVC. On a node-local install (the installer’s default) you copy them in through a throwaway Pod; on a host path install (TB_STORAGE_MODE=hostpath) you copy them into ~/.tracebloc/<namespace>/data/.
If you place images/ and labels.csv at the top of the PVC, they appear inside the ingestor pod as /data/shared/images/... and /data/shared/labels.csv. Set TRACEBLOC_SRC_PATH and TRACEBLOC_LABEL_FILE in ingestor-job.yaml to point at those in-pod paths (see Configure Kubernetes below). For tabular data, stage the single CSV with features and labels.
Write the Dockerfile
The Dockerfile only needs to install the ingestor package and copy in your script — the dataset is mounted at runtime, so do notCOPY data into the image:
restricted Pod Security Standard (see Run as non-root below), also add a non-root user to the Dockerfile, before the # Set the entrypoint line:
Build Docker Image
You need a docker user and password to proceed with the next step. Cloud platforms run a mix of x86 and ARM nodes (e.g. AWS Graviton, Azure Ampere, GCP Tau T2A). Building a multi-arch image with--platform linux/amd64,linux/arm64 guarantees the image runs on either, particularly if you build on Apple Silicon (M1/M2) or other ARM-based systems. Build and push the image with a single command:
3. Configure Kubernetes
With the image generated and pushed to the registry, editingestor-job.yaml with your settings:
JOBNAME, to distinguish between train and test data jobs.NAMESPACE, use the same as your client.image, your Docker image (imagePullPolicy: Always for DockerHub, IfNotPresent for local)CLIENT_ID,CLIENT_PASSWORD: read from the Kubernetes Secret<namespace>-secretsthat the installer created in your namespace, so you don’t paste credentials into the fileTRACEBLOC_TABLE_NAME, unique per dataset, train and test use different names, no spaces. Different names for train and test data is mandatoryTRACEBLOC_LABEL_FILE, path inside the ingestor pod (under/data/shared) to the CSV with file paths and labels — must match where you staged the file on the shared PVCTRACEBLOC_SRC_PATH, root inside the pod where the dataset directory is mounted (/data/shared)TRACEBLOC_BATCH_SIZEis the number of entries sent to the server per request. Optional — defaults to 4000. Keep it consistent across data types. It depends on available CPU memory, not for example image size. Too large can exhaust memory. It was tested up to 10,000, but 5,000 is a safe default for most systems.TRACEBLOC_LOG_LEVEL, “WARNING” for all warnings and errors, “INFO” for all logs, “ERROR” for errors only
Older setting names still work. Each setting above was first named without the
TRACEBLOC_ prefix: SRC_PATH, LABEL_FILE, TABLE_NAME, TITLE, BATCH_SIZE and LOG_LEVEL. An existing ingestor-job.yaml that uses those names keeps working, and if you set both spellings, the TRACEBLOC_ name wins. TRACEBLOC_TITLE, TRACEBLOC_BATCH_SIZE and TRACEBLOC_LOG_LEVEL are read from tracebloc-ingestor 0.8.69 on; if your image pins an older version, set the older names for those three (or both, with the same value).4. Deploy
Run the ingestor as a Kubernetes Job:JOBNAME and TABLE_NAME values for each run (e.g. ingestor-job-train / ingestor-job-test), and set intent to TRAIN or TEST accordingly in your ingestion script.
Run as non-root
If the namespace enforces therestricted Pod Security Standard, kubectl apply will be admitted but the pod will be rejected with a warning like:
securityContext block to the container in ingestor-job.yaml (already shown in the YAML above):
# Set the entrypoint line so the image ships with a UID that satisfies runAsNonRoot: true:
Verify Deployment
Verify if jobs and pods are deployed successfully and running:Best practices
- Use consistent, descriptive dataset names (for example
insurance_claims_trainandinsurance_claims_test) - Check the dataset before you ingest it (
tracebloc data ingest --dry-run, ortracebloc data validate ingest.yamlfor aningest.yaml) to catch errors early - Clean data before ingestion. Peers can’t view, clean or fix raw data, so model performance depends entirely on the quality of data you provide
- Helm chart and custom script: run the training and test jobs in parallel under different job names
Troubleshooting the Helm chart and custom script
Recommended for debugging: Use k9s, a terminal-based Kubernetes dashboard, to monitor jobs, pods, and logs in real time. Runk9s -n <namespace> to get a live view of resources, switch between them instantly, and inspect logs or events with a few keystrokes. Compared to kubectl, it is faster and more convenient.
Stale Kubernetes Job preventing new Job execution:
Next Steps
- Define and publish your use case: Define Use Case
Need Help?
- Email us at [email protected]