k9s -n <namespace> to get a live view of pods, logs, and events.
Your namespace is the Environment shown in the summary at the end of the install (or run kubectl get ns). Use it wherever this page says <namespace>.
Stuck? Generate a support bundle. Re-run the installer with It writes a redacted
--diagnose:~/.tracebloc/tracebloc-diagnose-<timestamp>.tgz — logs, pod status, and versions with credentials removed — that you can send to support. The first line of output shows your installer version.Quick Checks
Error Messages
General
These errors typically indicate storage pressure on your Kubernetes nodes:Local
Issues specific to local (k3d) deployments:Debugging Commands
When the quick checks don’t resolve the issue, use these commands to dig deeper.Pod status and logs
Resource usage
See if your nodes or pods are running out of CPU or memory:Storage
Check that persistent volume claims are bound and have enough capacity:Image pull credentials
The default images onghcr.io are public and need no pull secret. If you set registry credentials (dockerRegistry.create: true) and pods fail with ErrImagePull, verify that the chart created the secret <release>-regcred (the installer names the release after the namespace):
dockerRegistry.existingSecret, check that name instead.
CPU and Memory Optimization
Hitting resource limits during training? Two levers:- Reduce data size — smaller batches, lower resolution, shorter sequences
- Smaller models — fewer layers, smaller hidden dimensions
Memory Consumption (RAM / GPU VRAM)
Understanding what drives memory usage helps you right-size your resource limits:- Batch size — memory scales roughly linearly with batch size
- Model size — more parameters = more memory for weights, activations, and gradients
- Precision — FP16/BF16 uses half the memory of FP32. Mixed-precision training helps significantly
- Optimizer — Adam requires ~2-3x the memory of SGD (stores running averages)
- Input dimensions — transformer attention grows quadratically with sequence length; CNN memory grows quadratically with image resolution
Compute Consumption (CPU / GPU)
If training is slow, these are the factors to look at:- Batch size — larger batches increase GPU utilization up to memory saturation
- Model complexity — transformer attention is O(seq_len² x hidden_dim); CNNs scale with kernel x feature map x filters
- Precision — FP16/BF16 can speed up training 2-3x on modern GPUs
- Data pipeline — slow CPU preprocessing (augmentation, tokenization) can bottleneck training
- Parallelization — data parallelism splits batches across GPUs; model parallelism splits the model itself