Skip to main content
Everything the installer does, explained. This page walks through the full setup process — what each step does, what to expect, and how to verify it worked.
The installer deploys a single-node secure environment on one machine — a local Kubernetes cluster running inside Docker on that host. For multi-node or production deployments, see the EKS deployment guide.

Requirements

The installer runs on any modern machine (one host per secure environment) and checks these before it starts. Below the minimum it stops on Linux with a clear message; on macOS and Windows it handles Docker’s memory as described below the table. “To run” keeps a secure environment online; training on the same machine needs the headroom — even the smallest experiment (4 GiB) only fits beside the platform’s own services in an 11 GB Docker budget. On macOS and Windows, Docker runs in a VM with its own memory limit, which is often 8 GB by default.
  • macOS: the installer offers to raise Docker Desktop’s memory for you, which restarts Docker Desktop. If you decline, set it in Docker Desktop → Settings → Resources → Memory.
  • Windows (WSL 2): the installer warns and tells you how much to give Docker. Set memory= under [wsl2] in %UserProfile%\.wslconfig, then run wsl --shutdown.
  • Below 5 GB for Docker, the installer stops on macOS and Linux.
On macOS the installer can’t see Docker Desktop’s disk, so check the Disk usage limit under Docker Desktop → Settings → Resources yourself: at least 20 GB, and 50 GB to train. Admin rights on macOS: use an administrator account. The installer asks for your password once and stops with “Administrator rights required” on a standard account, unless Docker is already installed and running. Then it needs no administrator rights and installs kubectl, k3d and helm into ~/.local/bin. Supported platforms: macOS (Intel & Apple Silicon) · Linux (x86_64 & arm64) · Windows (x86_64 & arm64; on arm64 the tracebloc CLI needs Windows 11, where it runs the x86_64 build) GPU (optional): training on a GPU needs an NVIDIA GPU of the Turing generation or newer, a driver that runs CUDA 12, and an x86_64 machine running Linux or Windows. See GPU support for supported GPUs, drivers and platforms, and what isn’t supported. Outbound access needed: The installer pulls container images, the install scripts, and the Helm chart, then connects to the tracebloc platform. The host list, split into install-time and run-time, is on Security & data handling — use it as the single source when you write the firewall rules, since a partial allowlist fails in ways that read like a broken image rather than a blocked host.
Storage must be on a local disk. The data directory (TRACEBLOC_HOST_DATA_DIR, default ~/.tracebloc; the older HOST_DATA_DIR still works) holds the secure environment’s database, which corrupts on network filesystems — the installer fails fast if it is on NFS/CIFS/SMB. To keep large datasets on a network mount, set TRACEBLOC_HOST_DATASET_DIR to it together with TRACEBLOC_STORAGE_MODE=hostpath (with the default node-local storage the installer stops): only the dataset volume moves there, while the database stays local. See Configuration.

GPU support

A GPU is optional: every secure environment can train on CPU. To train on a GPU, the machine needs an NVIDIA GPU of the Turing generation or newer, an NVIDIA driver that runs CUDA 12, and an x86_64 (amd64) CPU. Anything else trains on CPU.

At a glance

NVIDIA GPU generations

The training images carry kernels for the generations below. The list is the architecture set of the PyTorch build the images ship, torch==2.13.0 from the cu129 wheel index, as of 2026-09-21. The generation decides support, not the product line: a consumer GeForce card trains the same way as a data-centre card of the same generation. What differs is memory (see GPU memory). Grace Hopper (GH200) and Grace Blackwell (GB200) systems have a supported GPU but an arm64 CPU. The GPU training images are built for x86_64 only, so these machines train on CPU. From the installer release after v1.9.200, the installer leaves an NVIDIA GPU on an arm64 machine out: it installs no driver, container toolkit or device plugin, installs CPU-only and prints why. Re-running it on an arm64 machine an older installer set up with the GPU, on Linux or Windows, switches that machine to CPU. The secure environment’s page in the web app then shows that the GPU isn’t supported. Older GPUs install CPU-only. A card older than Turing (compute capability below 7.5, for example GTX 10xx, GT 710, V100, P100) has no kernels in the training images. The installer leaves the GPU out, installs the secure environment CPU-only and prints why. The secure environment’s page in the web app then shows that the GPU isn’t supported and that training runs on CPU.

NVIDIA driver and CUDA

The training images bring their own CUDA libraries: PyTorch 2.13.0 built for CUDA 12.9, with the CUDA 12.9 runtime, cuDNN 9 and NCCL from PyTorch’s own packages. The host only needs the NVIDIA driver. Don’t install the CUDA toolkit or cuDNN on the host for tracebloc. A driver runs these images when it supports CUDA 12. Through CUDA minor version compatibility, that starts at 525.60.13 on Linux and 528.33 on Windows (NVIDIA’s CUDA Toolkit release notes, the minor version compatibility table). Newer driver series, including those made for CUDA 13, run the images too. A Blackwell card needs a driver series that supports Blackwell, whatever the minimum says. The installer reads the driver version from nvidia-smi before it sets up the GPU:

Platforms

On Linux, the installer also installs the NVIDIA Container Toolkit and the Kubernetes device plugin. After it installs a driver, reboot and re-run the installer. For the Kubernetes settings, see Configuration. macOS trains on CPU by design: training happens in Linux containers, which on macOS run in a virtual machine with no access to the GPU.
AMD GPUs can’t train today. The training images are built for CUDA, and there are no ROCm images. From the installer release after v1.9.200, the installer leaves an AMD GPU out on every platform: it installs no ROCm and no device plugin, installs the secure environment CPU-only and prints why. Re-run the installer on a machine set up by an older installer, and it switches the secure environment to CPU; ROCm stays installed, and you can remove it yourself. With an older installer, run tracebloc resources set --gpus 0 to train on CPU.

Multi-GPU

Each experiment uses one GPU by default. On a machine with several GPUs, give each experiment more of them:
The setting applies from the next experiment. A PyTorch experiment spreads across those GPUs with DistributedDataParallel. Test and inference runs use one GPU. Not supported:
  • One experiment across several machines. An experiment’s GPUs all come from one machine (one Kubernetes node).
  • MIG partitions. Training pods request whole GPUs (nvidia.com/gpu), never a MIG profile.

GPU memory

There is no minimum GPU memory: what fits depends on the model and its batch size. When a PyTorch run hits a CUDA out-of-memory error, it retries the failing batch before it gives up: first with a paged optimizer, then with half-size micro-batches (the effective batch size stays the same), then with activation checkpointing. If the batch still doesn’t fit, the run fails with the out-of-memory error. Use a smaller model or batch size, or a GPU with more memory.

Check your machine

This prints each GPU’s name, driver version, compute capability and memory. Compare them with the tables above.
  • At the end of the install, the installer prints the secure environment’s mode: GPU or CPU. If it installs CPU-only on a GPU machine, it prints why.
  • tracebloc resources shows how many GPUs the machine has and how many each experiment may use.
  • tracebloc doctor warns when experiments ask for a GPU that no node can give them.
  • If a GPU can’t train, the experiment fails at start with a message such as GPU not usable: <GPU> (compute capability <X.Y>, driver <version>); supported: Turing (7.5) or newer.

1. Create an Account

Sign up at ai.tracebloc.io. Free to get started — no credit card required.

2. Deploy

One command sets up your entire secure environment on any machine — macOS, Linux, or Windows. The installer is idempotent: it detects what’s already installed and skips it, so it’s safe to re-run at any time.
What it changes on your machine: Docker (installed if missing), a local cluster named tracebloc, kubectl, k3d and helm, the tracebloc CLI (plus a PATH line in your shell profile, if needed), and ~/.tracebloc (config and logs). It also switches kubectl’s current context to k3d-tracebloc.

What the Installer Does

The installer runs six labelled steps: a) Checking your machine Checks CPU, memory, disk and network against the requirements above, and detects GPU hardware (falls back to CPU mode if none). b) Installing what tracebloc needs Installs whatever is missing: Docker, kubectl, k3d, helm, and the tracebloc CLI. On macOS, Docker Desktop is downloaded from docker.com and checked against its published checksum; a headless Mac gets Colima through Homebrew instead. kubectl, k3d and helm are checksum-verified direct downloads, not Homebrew: k3d and helm at pinned versions, kubectl at the current stable release. On Windows, Helm comes through winget when winget is available. On a Mac without Docker, Docker Desktop opens for the first time. In the Docker window, approve the privileged-helper prompt (your admin password) and accept the license agreement. The installer then continues by itself. c) Creating your secure environment Provisions a lightweight local Kubernetes cluster inside Docker. First run takes 1–2 minutes to download components. d) Registering this machine Signs you in and registers this machine on your account — there’s nothing to create on the clients page beforehand:
  1. The installer prints a link and a short code. Open the link on any device (this computer or your phone), sign in the way you already do — password, Google, or GitHub — and approve the code. The code is valid for 10 minutes; if it lapses, press Enter for a fresh one.
  2. It asks you to name your secure environment — the name shown on your dashboard. Set TRACEBLOC_CLIENT_NAME before the install command to skip the prompt. The location comes from your system timezone; set TRACEBLOC_CLIENT_LOCATION to a country code such as DE to choose it.
e) Installing tracebloc Installs the tracebloc Helm chart into the cluster with the credential minted in step d. You never see or type it. f) Connecting to the tracebloc network Waits until the services are up and connected, then prints a summary with your secure environment’s name, version and mode (CPU or GPU), and what to do next. Install logs are kept in ~/.tracebloc/ if you need to debug anything.
Browser sign-in needs someone at a terminal. For an unattended install (CI, device management), first create the secure environment on the clients page: you set its password there, and you copy its ID from the list. Then set TRACEBLOC_CLIENT_ID and TRACEBLOC_CLIENT_PASSWORD before the install command. The installer skips sign-in and connects with those credentials. To install into a namespace other than tracebloc, also set TRACEBLOC_NAMESPACE.

Status reference

Your secure environment moves through these states on your dashboard:
A one-liner install keeps its services current through the hourly auto-upgrade, and tracebloc upgrade brings the CLI to the latest release. To upgrade the services right away, see Operations → Upgrade, which covers the Helm route (helm upgrade <namespace> tracebloc/client -n <namespace> --reset-then-reuse-values; --reset-then-reuse-values keeps the values the installer applied).

GPU Support

The installer detects your GPU and configures the cluster:
  • Linux x86_64 (NVIDIA) — drivers, container toolkit, and Kubernetes device plugin are installed automatically. A reboot may be required after driver installation. AMD GPUs, and NVIDIA GPUs on arm64, install CPU-only.
  • macOS — CPU-only. For GPU workloads, deploy on a Linux machine or use AWS (EKS).
  • Windows — pre-install NVIDIA drivers before running the installer. The installer detects them via nvidia-smi.
See Configuration > GPU for detailed platform-specific behavior.

3. Verify

After the installer finishes, open a new terminal (so the tracebloc command is on your PATH) and run:
tracebloc client status --wait waits until it’s online. Then open ai.tracebloc.io and check that your secure environment shows Online. This confirms it has a secure connection to the tracebloc platform.
Stuck on Offline? Run tracebloc doctor, then see the Troubleshooting page.

What’s Next

Your secure environment is running. Here’s where to go from here: Data owners: Ingest your dataset, then create a use case — pick from prepared datasets, define evaluation metrics, and invite peers to submit models. Data scientists: Join a use case — connect to a use case, train models on real data, and submit results. Advanced configuration: Configuration — customize installer options, manage the cluster, configure GPU settings, or deploy manually with Helm.