Overview
Use EKS for production — multi-node, autoscaling, or shared GPU clusters on AWS. For a single machine (a laptop or one server), the local installer is simpler and faster.
Prerequisites
Before proceeding with the EKS infrastructure setup, make sure the following requirements are met:AWS Setup
1. Create or use an AWS Account
- If you don’t already have one, create an AWS account.
- Your account must have permissions to create and manage EKS resources, VPCs, EC2 instances, and IAM roles.
2. Install and configure AWS CLI
- Install the AWS CLI on your local machine.
- Configure your credentials with:
eu-central-1), and output format.
3. Verify your AWS CLI configuration
Check if your credentials and region are set correctly:Required Permissions
Requires permissions for:- Amazon EKS cluster management
- VPC and networking resources
- EC2 instances and security groups
- IAM roles and policies
Required Tooling & Tracebloc Account
- Helm 3.14 or later (Helm 4 works): the upgrade command in Operations uses
--reset-then-reuse-values, which needs 3.14. Install Helm on your local machine. Installation Guide - kubectl: Install kubectl to interact with your EKS cluster. Installation Guide
- Tracebloc Account: You will need your Client ID and Client Password (from the tracebloc client view).
- Docker registry credentials (optional): The tracebloc images are public, so a default install pulls them without credentials. You need credentials only to avoid Docker Hub rate limits or to pull from a private mirror.
Recommended for Monitoring: Use k9s
You can use k9s, a terminal-based Kubernetes dashboard, to monitor jobs, pods, and logs in real time. Runk9s -n <NAMESPACE> to get a live view of resources, switch between them instantly, and inspect logs or events with a few keystrokes. Compared to kubectl, it is faster and more convenient.
Components
The EKS infrastructure setup consists of two main parts: Core Infrastructure and Client Deployment. Together, they give you a scalable Kubernetes cluster that runs your ML workloads inside the tracebloc secure environment.Core Infrastructure
VPC with isolated networking Creates a dedicated, secure network environment with subnets across multiple Availability Zones for high availability. EKS Cluster A managed Kubernetes control plane with auto-scaling nodegroups to run your ML workloads. IAM Roles and Policies Granular roles for the cluster and worker nodes following least-privilege principles. Security Groups Network-level access controls for pods, nodes, and storage. EFS Storage System Shared persistent storage for training datasets and model artifacts, accessible from all nodes.Client Deployment
tracebloc secure environment Deployed into the EKS cluster using Helm, configured with your account credentials and storage. Monitoring and Verification Tools for validating that your cluster, workloads, and secure environment are running correctly. By combining these components, you get a production-ready Kubernetes environment tailored for secure, high-performance machine learning workloads.Process Flow
- Infrastructure Setup - Create VPC, networking, EKS cluster, and storage
- Client Deployment - Configure and deploy the tracebloc application
- Verification - Validate everything is running correctly
Security
Security is a core consideration when setting up and running ML workloads on EKS. This setup follows AWS and Kubernetes best practices to ensure data protection, secure communication, and controlled access at every layer.Data Protection
- Data Locality: Training data always remains within your infrastructure. Only pre-defined metrics and logs are shared externally.
- Model Encryption: Model weights are encrypted and never leave your environment.
- Secure Communication: All platform communication is protected with TLS.
tracebloc Backend Security
- Code Analysis: Submitted external models are scanned with the Bandit library to detect vulnerabilities or malicious code.
- Input Validation: Training scripts and model code undergo strict validation before execution.
- Sandboxed Execution: Training workloads run in isolated environments with restricted system access.
- Namespace Isolation: Kubernetes namespaces ensure logical separation from other applications.
Identity & Access Management
AWS IAM Roles:- Cluster Role: Minimal permissions (
AmazonEKSClusterPolicy) for cluster management. - Nodegroup Role: Least-privilege permissions for worker nodes, networking, container registry, and storage access.
- Cluster roles provide limited permissions for jobs, pods, and deployments.
- Service accounts are bound to roles within their namespace, preventing cross-namespace access.
Network Security
- VPC: This guide builds a dedicated VPC with public subnets and an Internet Gateway. Nodes get public IPs and unrestricted outbound access. For production, move the nodes to private subnets behind a NAT gateway and restrict outbound traffic with your own security groups or firewall.
- Security Groups: The EFS mount targets use the EKS cluster security group, so only cluster nodes can reach the file system.
- Training pod egress: The tracebloc chart blocks direct outbound HTTPS from training pods by default and routes it through an in-cluster egress gateway allowlist. This needs a CNI that enforces NetworkPolicy (see NetworkPolicy for training pods).
Detailed Setup
This section walks through a step-by-step build with AWS CLI and kubectl. You choose your own resource settings (# of CPUs, Memory, VPC CIDRs, instance types, namespace, Helm release name, StorageClass, etc.). Expect about 1–2 hours end-to-end.What you’ll do (Steps 1–6):
- VPC & Network Configuration — Create an isolated Virtual Private Cloud (VPC) with subnets distributed across three availability zones for high availability, an Internet Gateway for outbound connectivity, and routing tables with proper associations. This provides network isolation and fault tolerance for your cluster infrastructure.
- EKS Cluster Setup — Create the cluster service role and provision the EKS control plane which manages the Kubernetes API server.
- EKS Nodegroup Setup — Create a node service role and provision two managed nodegroups: A system nodegroup for Kubernetes system components and a training nodegroup for ML workloads. Each nodegroup can be customized with different instance types, scaling parameters, and capacity types depending on the work loads and data types.
- Storage — Create an Amazon EFS file system for shared persistent storage and create mount targets in each availability zone, using the EKS cluster security group so the nodes can reach EFS over NFS. Install the EFS CSI (Container Storage Interface) driver, which lets Kubernetes create and mount EFS volumes, and metrics-server.
- Client Configuration — Add the tracebloc Helm repository and configure your deployment values (authentication credentials, storage settings, resource limits).
- Client Deployment — Install the chart into your chosen namespace and verify that all pods are running and persistent volume claims are bound.
tracebloc/client chart with EKS-specific values.
1. VPC and Network Configuration
Your AWS EKS cluster must run in a secure and isolated network with subnets across availability zones, DNS resolution, and internet access. This is provided by an AWS VPC (Virtual Private Cloud), which is a logically isolated section of AWS where you define your own IP address range (CIDR block), subnets, routing, and gateways. This section sets up the foundation.VPC Creation
VPC provides network isolation, so your workloads are not exposed to the public internet by default.10.0.0.0/16 to define the IP address range for all resources in the VPC and tags it with the specified owner name. Replace <OWNER_NAME> with your preferred identifier. Keep the output and the VPC_IDfor the next steps.
Enable DNS Hostnames
DNS (Domain Name System) support and hostnames let Kubernetes pods and services find each other by name, which is essential for service discovery in Kubernetes, which relies on DNS resolution.VPC_IDfrom the previous step or list VPCs using aws ec2 describe-vpcs.
Create Subnets Across Availability Zones
Subnets (smaller ranges of IP addresses inside the VPC) are spread across multiple AZs (Availability Zones) to ensure high availability and fault tolerance. EKS requires subnets in at least 2 availability zones for high availability - using 3 zones provides better fault tolerance and load distribution.10.0.1.0/24, 10.0.2.0/24, 10.0.3.0/24) distributed across different availability zones. Use or create tags if needed. Register the generated SUBNET_IDs for the following steps.
Enable Public IPs on Subnets
EKS worker nodes need public IPs to communicate with the EKS control plane, download container images from Dockerhub, and allow external traffic to reach applications.Create and Attach an Internet Gateway
An Internet Gateway (IGW) connects your VPC to the internet so nodes can pull container images and reach AWS APIs.Create Default Route Table to Internet Gateway
A route table defines how network traffic is directed within the VPC and out through the IGW.2. EKS Cluster Setup
Amazon EKS provides the managed Kubernetes control plane. It needs permissions through AWS IAM (Identity and Access Management) to create and manage resources such as load balancers, security groups, and network interfaces. The cluster itself is the logical container for the API server, linking your VPC networking with the control plane and IAM role. To interact with it you must update your local kubeconfig so kubectl can send commands. Setting this up gives you a working API server that can manage workloads once worker nodes are added in the next step.Create EKS Cluster Role
EKS needs an IAM role to manage your cluster’s control plane and AWS resources:--role-name: Name for your EKS service role--assume-role-policy-document: Trust policy allowing EKS to use this role--tags: Optional resource tags for organization
Attach EKS Cluster Policy to Role
To give the EKS control plane the necessary permissions, you must attach the AmazonEKSClusterPolicy.Create EKS Cluster
Create the EKS cluster to provision the managed Kubernetes control plane and connect it to your VPC, subnets, and IAM role.K8S_VERSION to a version in standard support (see the EKS Kubernetes version calendar); older versions cost extra under extended support. Set an appropriate CLUSTER_NAME, expect this to take around 10 minutes. If you do not have your accound ID ready, run aws sts get-caller-identity.
Configure kubectl
Your local kubectl must know the cluster’s API endpoint and credentials, otherwise you cannot interact with the cluster you just created:Manage Multiple Clusters (Optional)
If you work with several clusters, make sure your context points to the right one. List all contexts:CONTEXT_NAME, then apply it:
Create Application Namespace
Namespaces let you isolate your workloads from system components and other applications running in the cluster. Define an appropriate namespace for your setup3. EKS Nodegroup Setup
Nodegroups are sets of EC2 instances that act as worker nodes, running your system and training workloads. Each nodegroup needs an IAM role so the nodes can join the cluster and interact with AWS services such as pulling images from registries, attaching EFS storage, and sending logs. It is best practice to create at least two groups: a small system nodegroup to host cluster services and a training nodegroup sized for your ML workloads. The training nodegroup can be CPU-based for general jobs or GPU-based for deep learning, ensuring that heavy jobs do not interfere with core Kubernetes functions.Create Nodegroup Role
Worker nodes need an IAM role to join the cluster and access AWS services.NODEGROUP_ROLE_NAME.
Attach Required Policies
Attach minimum required AWS-managed policies so the nodes have permissions:- AmazonEKSWorkerNodePolicy: lets nodes communicate with the cluster control plane
- AmazonEC2ContainerRegistryReadOnly: allows pulling images from ECR
- AmazonEKS_CNI_Policy: required for pod networking via the Container Network Interface (CNI) plugin
Create Managed Nodegroups
Now let`s create two nodegroups: A system nodegroup that runs critical Kubernetes components (CoreDNS, kube-proxy, Container Network Interface (CNI), Storage Interface drivers (CSI), metrics-server, etc.) and a training nodegroup that runs your machine learning workloads. This separation ensures cluster stability even under heavy training load.System Nodegroup
It is recommended to use AWS EC2 instance type t3.medium for system pods as they are relatively cheap and sufficient. Spread nodes across three Availability Zones for resilience, set ON_DEMAND for reliability and label as “system”:t3.medium instances (2 vCPUs, 4 GiB memory) spread across three AZs. The group scales between 2 and 5 nodes (EC2s), each node is labeled for system workloads. Set an appropriate SYSTEM_NODEGROUP_NAME. Use an x86_64 AMI type for system nodes so that the data ingestor runs on them, too. AWS publishes no Amazon Linux 2 (AL2_*) EKS AMIs for Kubernetes 1.33 and later, so use the AL2023_* types.
Training Nodegroup
This group runs your ML training workloads — size it for your dataset, model type, number of parallel workloads, and whether you need GPUs. Refer to the EC2 instance types list and EKS managed nodegroups docs for guidance.TRAINING_NODEGROUP_NAME. Depending on the compute type (CPU vs. GPU), set variables as follows:
The first example provisions a training nodegroup with t3.xlarge instances (4 vCPUs, 16 GiB memory). It starts with 2 nodes, can scale down to 0 when idle, and grows to 5 under load.
You can limit the compute per peer or team: Refer to the creating a use case section for details.
Optional for GPU Training Nodegroup
Install the NVIDIA device plugin so Kubernetes can automatically detect and schedule GPUs on any new node:4. Storage
Your workloads need persistent storage across nodes to store:- train and test datasets
- model checkpoints and weights during training
- secure environment logs
Create EFS File System
Let`s create an EFS file system:FILE_SYSTEM_ID for later use.
Identify Security Groups
List all security groups in the VPC so you can allow NFS traffic from your worker nodes to the EFS mount targets. NFS traffic is the file system protocol that EFS uses to let your nodes read and write shared storage.SECURITY_GROUP_ID in the next step. If you pick a different security group, add an inbound rule that allows TCP 2049 from your nodes:
Create EFS Mount Targets
Mount targets act as network endpoints in each Availability Zone, letting your worker nodes connect to the EFS file system using NFS.Attach EFS CSI Driver Policy to Node Role
The EFS CSI (Container Storage Interface) driver is the Kubernetes plugin that makes Amazon EFS usable inside your cluster. It translates Kubernetes PersistentVolumeClaims into actual EFS mounts. For the driver to do this, the worker nodes need IAM permissions to create, mount, and manage EFS volumes.Add EFS CSI Driver Helm Repository
The EFS CSI driver is packaged as a Helm chart, which makes it easy to install and upgrade in your cluster.Configure EFS CSI Driver Placement
The EFS CSI driver consists of a controller (manages provisioning of storage) and node pods (handle actual EFS mounts on each worker). By default, the controller could be scheduled onto any node, including your training nodes. Since training jobs consume heavy CPU and memory, this risks starving the controller and blocking storage operations. Lets manually create a driver configuration fileefs-csi-driver.yaml:
type=system in step 3) instead of training nodes.
Install EFS CSI Driver with Helm
Now, the EFS CSI driver needs to be installed into the cluster (namespace kube-system by convention) using the configuration file.Install Metrics Server
The Metrics Server is the cluster-wide aggregator for resource usage data. It collects CPU and memory metrics for autoscaling and live resource monitoring. The chart’s resource monitor requires it, and EKS does not install it by default. Install it once per cluster:5. Client Configuration
You can now prepare the Helm chart that deploys the tracebloc secure environment itself: it is the component that runs your model evaluation and training jobs inside the cluster. To install it, you use a Helm chart, which bundles all required Kubernetes manifests into a single, configurable package. A values.yaml file controls this deployment, where you specify credentials (to authenticate with tracebloc) and storage settings (to mount your EFS volumes). This configuration ensures the secure environment can securely connect, schedule workloads, and store results.Add Helm Repository
The tracebloc secure environment is delivered as a single unified Helm chart (tracebloc/client) that supports AKS, EKS, bare-metal, and OpenShift. Source: tracebloc/client.
Configure your Deployment Settings
Export the chart’s default configuration into a local file that you can edit:values.yaml:
Tracebloc Authentication
Provide your Client ID and password from the tracebloc client view:Docker Registry Configuration (Optional)
The tracebloc images are public, and by default the chart pulls them without credentials. Add adockerRegistry block only to avoid Docker Hub rate limits or to pull from a private mirror. With create: true, the chart creates a pull secret named {{ .Release.Name }}-regcred and uses it for every pod; without create: true, no secret is created and every pull stays anonymous:
DOCKER_USERNAME: Docker Hub usernameDOCKER_PASSWORD/DOCKER_TOKEN: Password, or an access token (preferred — required if 2FA is enabled)DOCKER_EMAIL: Email linked to your Docker account
Storage (EKS / EFS)
The chart provisions an EFS-backed storage class via the EFS CSI driver. Use EFS Access Points (provisioningMode: efs-ap) for proper UID/GID enforcement on shared filesystems:
FILE_SYSTEM_ID from your EFS setup (step 4). Adjust PVC sizes as needed.
StorageClass options:
create: true— chart creates a release-unique storage class (e.g.<release>-storage-class). Each release gets its own.create: false— reuse an existing class;namemust match.
Set Resource Limits for Training Jobs
Each training Job inherits these resource requests/limits. Make sure they fit within your EC2 instance capacity. IMPORTANT: For GPU jobs,requests and limits must be equal — Kubernetes rejects configs where they differ.
- One EC2 equals one node; one node runs many pods.
- One pod contains one or more containers.
- One experiment equals one Job, which in most cases creates one pod by default.
Proxy Settings (Optional)
Required only if your EKS worker nodes reach the internet through a corporate proxy:NetworkPolicy for Training Pods
The chart applies a NetworkPolicy to training pods that denies all ingress and restricts egress — arbitrary pod-to-pod traffic and the Kubernetes API are blocked, while the in-cluster MySQL that serves the training data and the proxy that reports results stay reachable. Direct outbound HTTPS is blocked by default: external HTTPS goes through the in-cluster egress gateway allowlist. SetallowExternalHttps: true only to opt out.
enableNetworkPolicy=true), or run Calico or Cilium as an add-on. The plain AWS VPC CNI without the agent does not enforce NetworkPolicy. If your cluster has no enforcing CNI, set enabled: false so you do not rely on protection you do not have.
Resource Monitor (metrics-server)
The chart’s resource-monitor DaemonSet polls/apis/metrics.k8s.io/v1beta1 and requires metrics-server, which you installed in step 4. If you cannot install metrics-server, disable the DaemonSet:
6. Client Deployment
With the configuration ready, deploy the tracebloc secure environment into your EKS cluster using Helm.Deploy the client with Helm
Install the chart into your namespace using your customized values file:- A MySQL pod:
mysql-client-... - The jobs manager pod, which runs your training jobs:
<release>-jobs-manager-... - A requests proxy pod:
<release>-requests-proxy-... - An egress gateway pod:
<release>-egress-proxy-... - A
<release>-resource-monitorDaemonSet in thetracebloc-node-agentsnamespace - A
<release>-auto-upgradeCronJob for hourly chart upgrades (see Configuration → Auto-upgrade) - Supporting resources: Service, ConfigMap, Secrets, PVCs, PDBs, and a PriorityClass only if you set
priorityClass.create: true
Running.
Verification and Maintenance
After deployment, confirm the secure environment is running correctly and learn how to maintain it.Verify Deployment
Check that all pods are running in your namespace:Check pod status:
mysql-client-..., <release>-jobs-manager-..., <release>-requests-proxy-... and <release>-egress-proxy-... pods in Running state. If not, inspect the logs for troubleshooting.
Check services:
mysql-client.
Check persistent volumes:
Maintenance
The chart’s auto-upgrade CronJob handles routine version bumps every hour. To upgrade manually:Update your values:
Upgrade the deployment:
Uninstall
Remove the Helm release:
Clean up persistent resources (optional):
helm uninstall leaves some resources in place on purpose:
- The
tracebloc-node-agentsnamespace: it is shared by every release on the cluster. Delete it withkubectl delete namespace tracebloc-node-agentsonce no tracebloc release remains. - The PriorityClass, if you set
priorityClass.create: true. Delete it withkubectl delete priorityclass <NAME>. - Your data on EFS: the storage class uses
reclaimPolicy: Retain, so deleting the PVCs keeps the PersistentVolumes and their EFS data. Delete the released PersistentVolumes (kubectl get pv), and delete the EFS file system when you no longer need it.
Next Steps
Create a New Use Case- Prepare your dataset - Ingest and configure your dataset
- Create an AI use case - Set up a new use case Join an Existing Use Case
- Explore and join available use cases - Browse ongoing AI projects
- Start training models - Begin training on shared datasets
Need Help?
- Email: [email protected]
- Docs: tracebloc Documentation Portal