Using Automated ML workflows on dedicated servers provides a reliable and high-performance environment for training and deploying machine learning models. Dedicated servers are the best choice for ML workflows because they ensure consistent computing power, which avoids slowdowns. They also offer better scalability for large datasets and complex models.
In this guide, we will use Ubuntu 22.04 or Ubuntu 24.04 on a reliable and high-performance dedicated server, and use Docker and Docker Compose to set up an automated ML workflow.
System Prerequisites for Automated ML Workflows on Dedicated Servers
The automated ML workflow we want to set up in this guide has several key components, including:
- MinIO (S3) for storing datasets and artifacts.
- PostgreSQL for metadata management.
- MLflow for experiment tracking and model registry.
- Apache Airflow for automation.
- DVC for data versioning.
- Optionally, NVIDIA Triton can be added for model serving.
- Prometheus, Grafana, and cAdvisor handle monitoring.
Before setting up this workflow, you need to prepare your operating system by configuring firewall rules, installing Docker, and other necessary steps.
Install Required Packages and Configure Firewall Rules
Run the system upgrade and install the required packages on your system with the following commands:
sudo apt update && sudo apt upgrade -ysudo apt install curl ufw ca-certificates gnupg lsb-release jq unzip git -y
Allow the necessary ports through your firewall and enable the firewall with the commands below:
sudo ufw allow 9000,9001,5000,8080,5432,9090,3000,5555/tcpsudo ufw enable
Note: In a production environment, allow only within your LAN or VPN.
Set up Docker and Docker Compose
Use the following commands to install Docker and Docker Compose on your system:
sudo install -m 0755 -d /etc/apt/keyringscurl -fsSL https://download.docker.com/linux/ubuntu/gpg \ | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpgecho \ "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] \ https://download.docker.com/linux/ubuntu $(lsb_release -cs) stable" \ | sudo tee /etc/apt/sources.list.d/docker.list > /dev/null sudo apt updatesudo apt install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin -y sudo usermod -aG docker $USER
Install NVIDIA Driver and Container Toolkit for GPU Servers
If you have servers with GPUs, you must install the NVIDIA driver with the container toolkit. Install the proper driver on your system with the commands below:
sudo apt install ubuntu-drivers-common -ysudo ubuntu-drivers autoinstall
Reboot your system, and after that, check for NVIDIA drivers with the command below:
To set up the container toolkit, you can run the following commands:
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \ | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpgcurl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \ | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#' \ | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list sudo apt update && sudo apt install nvidia-container-toolkit -y
Configure Docker to recognize and use NVIDIA GPUs through the NVIDIA Container Toolkit with the command below:
sudo nvidia-ctk runtime configure --runtime=docker
Restart Docker to apply the changes:
sudo systemctl restart docker
You can run a quick GPU test with the command below:
docker run --rm --gpus all nvidia/cuda:12.3.2-base-ubuntu22.04 nvidia-smi
Tip: For users who prefer ready-to-deploy GPU infrastructure rather than configuring drivers manually, consider using PerLod's High-performance GPU Dedicated Servers.
Create Main Services With Docker Compose for ML Workflow
You need to create an env file that sets environment variables to configure connections between the services, including PostgreSQL, MLflow, and MinIO. This allows MLflow to store experiments in PostgreSQL and artifacts in MinIO.
Create a working directory with the command below:
sudo mkdir -p ml-platform
Create the env file in your working directory by using the command below:
sudo nano ml-platform/.env
Add the following variables to the file:
# .envPOSTGRES_USER=platformPOSTGRES_PASSWORD=platformpassPOSTGRES_DB=platformdb MLFLOW_BACKEND=postgresql://platform:platformpass@postgres:5432/platformdbMLFLOW_S3_ENDPOINT_URL=http://minio:9000AWS_ACCESS_KEY_ID=minioadminAWS_SECRET_ACCESS_KEY=minioadminsecretMLFLOW_ARTIFACT_BUCKET=mlflow MINIO_ROOT_USER=minioadminMINIO_ROOT_PASSWORD=minioadminsecret
Then, create a Docker Compose file for PostgreSQL, MinIO, and MLflow services with the command below:
sudo nano ml-platform/compose/docker-compose.core.yml
Add the following config to the file:
1version: "3.9"2services:3 postgres:4 image: postgres:165 restart: unless-stopped6 environment:7 POSTGRES_USER: ${POSTGRES_USER}8 POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}9 POSTGRES_DB: ${POSTGRES_DB}10 volumes:11 - pgdata:/var/lib/postgresql/data12 networks: [core]13 ports:14 - "5432:5432"15 16 minio:17 image: minio/minio:latest18 command: server /data --console-address ":9001"19 environment:20 MINIO_ROOT_USER: ${MINIO_ROOT_USER}21 MINIO_ROOT_PASSWORD: ${MINIO_ROOT_PASSWORD}22 volumes:23 - minio:/data24 networks: [core]25 ports:26 - "9000:9000"27 - "9001:9001"28 29 mlflow:30 image: python:3.11-slim31 restart: unless-stopped32 depends_on: [postgres, minio]33 working_dir: /app34 environment:35 MLFLOW_BACKEND_STORE_URI: ${MLFLOW_BACKEND}36 MLFLOW_S3_ENDPOINT_URL: ${MLFLOW_S3_ENDPOINT_URL}37 AWS_ACCESS_KEY_ID: ${AWS_ACCESS_KEY_ID}38 AWS_SECRET_ACCESS_KEY: ${AWS_SECRET_ACCESS_KEY}39 BACKEND_BUCKET: ${MLFLOW_ARTIFACT_BUCKET}40 command: >41 sh -c "pip install --no-cache-dir mlflow[boto3]==2.14.1 psycopg2-binary==2.9.9 &&42 mlflow server43 --backend-store-uri ${MLFLOW_BACKEND_STORE_URI}44 --default-artifact-root s3://${BACKEND_BUCKET}45 --host 0.0.0.0 --port 5000"46 networks: [core]47 ports:48 - "5000:5000"49 50volumes:51 pgdata:52 minio:53 54networks:55 core:
This Docker Compose file creates a complete ML platform with experiment tracking and artifact storage. You can check the official Docker Compose documentation for detailed configuration syntax, service options, and advanced networking setups.
Switch to the compose directory and run Docker Compose with the following commands:
cd ml-platform/composedocker compose -f docker-compose.core.yml up -d
Initialize MLflow Artifact Storage in MinIO
At this point, you must configure MinIO object storage to store MLflow artifacts. It creates a dedicated bucket where MLflow will save experiment results, models, and other machine learning artifacts for tracking and versioning.
Download and install the MinIO client system-wide on your server with the following commands:
sudo wget https://dl.min.io/client/mc/release/linux-amd64/mc -O mcsudo chmod +x mc && sudo mv mc /usr/local/bin/
Configure the connection to the MinIO server with the following command:
sudo mc alias set local http://127.0.0.1:9000 ${MINIO_ROOT_USER} ${MINIO_ROOT_PASSWORD}
Create the bucket by using the command below:
mc mb local/${MLFLOW_ARTIFACT_BUCKET}
You can list the bucket with the following command:
You can access the MinIO console by navigating to the URL below:
Also, access the MLflow Web UI with:
Set up DVC with MinIO for Data Versioning
Now you must configure DVC (Data Version Control) to use MinIO as remote storage for datasets and models. This will enable version control for large data files while storing them efficiently in S3-compatible storage.
Run the following commands from your repo root to set up DVC and configure MinIO as remote storage for DVC files:
git initpipx install dvc 2>/dev/null || python3 -m pip install --user dvc[s3]dvc init dvc remote add -d s3remote s3://${MLFLOW_ARTIFACT_BUCKET}/dvcdvc remote modify s3remote endpointurl http://127.0.0.1:9000dvc remote modify s3remote access_key_id ${MINIO_ROOT_USER}dvc remote modify s3remote secret_access_key ${MINIO_ROOT_PASSWORD}dvc remote modify s3remote use_ssl falsegit add .dvc .gitignoregit commit -m "Init DVC with MinIO remote"
Then, track and version a dataset with DVC with the following commands:
mkdir -p data/raw && cp /path/to/your.csv data/raw/dvc add data/raw/your.csvgit add data/raw/your.csv.dvcgit commit -m "Track dataset with DVC"dvc push
ML Model Training with MLflow Experiment Tracking
In this step, we will show you machine learning training with comprehensive MLflow logging. It includes a complete workflow for training a classifier while automatically tracking parameters, metrics, and models to the MLflow server for experiment management and reproducibility.
Create the requirements file with the following command:
sudo nano ml-platform/ml/requirements.txt
Add the following variables to the file:
numpypandasscikit-learnmlflow==2.14.1boto3
Then, create the train file with the following command:
sudo nano ml-platform/ml/train.py
Add the following config to the file:
1import os2import mlflow3import mlflow.sklearn4import pandas as pd5from sklearn.model_selection import train_test_split6from sklearn.linear_model import LogisticRegression7from sklearn.metrics import f1_score8 9MLFLOW_TRACKING_URI = os.environ.get("MLFLOW_TRACKING_URI", "http://localhost:5000")10mlflow.set_tracking_uri(MLFLOW_TRACKING_URI)11mlflow.set_experiment("demo-exp")12 13def main():14 15 df = pd.read_csv("data/raw/your.csv")16 X = df.drop(columns=["label"])17 y = df["label"]18 19 Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)20 21 with mlflow.start_run():22 params = {"C": 1.0, "max_iter": 200}23 model = LogisticRegression(**params).fit(Xtr, ytr)24 preds = model.predict(Xte)25 f1 = f1_score(yte, preds, average="macro")26 27 mlflow.log_params(params)28 mlflow.log_metric("f1_macro", f1)29 mlflow.sklearn.log_model(model, artifact_path="model")30 31 print(f"F1 (macro): {f1:.4f}")32 33if __name__ == "__main__":34 main()
Now you can run the training model locally in a Python environment shell:
python3 -m venv .venv && . .venv/bin/activatepip install -r ml/requirements.txtexport MLFLOW_TRACKING_URI=http://127.0.0.1:5000python ml/train.py
From the MLflow Web UI, you can verify the model that appears under the run’s artifacts.
ML Pipeline Automation with Apache Airflow
At this point, you can automate machine learning workflows using Apache Airflow. This setup will create a complete automation system that can pull data with DVC, run training jobs, and log experiments to MLflow, all managed through scheduled pipelines with monitoring and retry capabilities.
Create the Airflow requirements file with the command below:
sudo nano ml-platform/airflow/requirements.txt
Add the following variables to the file:
apache-airflow-providers-httpapache-airflow-providers-cncf-kubernetesboto3mlflow==2.14.1dvc[s3]
Then, create the Airflow Docker Compose file with the command below:
sudo nano ml-platform/compose/docker-compose.airflow.yml
Add the following configuration to the file:
1version: "3.9"2x-airflow-common: &airflow-common3 image: apache/airflow:2.9.34 environment:5 AIRFLOW__CORE__LOAD_EXAMPLES: "False"6 AIRFLOW__CORE__EXECUTOR: LocalExecutor7 AIRFLOW__CORE__FERNET_KEY: "generate_a_fernet_key_and_put_here"8 AIRFLOW__DATABASE__SQL_ALCHEMY_CONN: postgresql+psycopg2://platform:platformpass@postgres:5432/airflow9 _PIP_ADDITIONAL_REQUIREMENTS: "apache-airflow-providers-http apache-airflow-providers-cncf-kubernetes boto3 mlflow==2.14.1 dvc[s3]"10 11 MLFLOW_TRACKING_URI: "http://mlflow:5000"12 AWS_ACCESS_KEY_ID: ${AWS_ACCESS_KEY_ID}13 AWS_SECRET_ACCESS_KEY: ${AWS_SECRET_ACCESS_KEY}14 MLFLOW_S3_ENDPOINT_URL: ${MLFLOW_S3_ENDPOINT_URL}15 volumes:16 - ../airflow/dags:/opt/airflow/dags17 - ../ml:/opt/airflow/ml18 depends_on:19 - postgres20 - minio21 - mlflow22 networks: [core]23services:24 airflow-webserver:25 <<: *airflow-common26 command: webserver27 ports: ["8080:8080"]28 airflow-scheduler:29 <<: *airflow-common30 command: scheduler31 airflow-init:32 <<: *airflow-common33 entrypoint: /bin/bash34 command: -c "airflow db migrate && airflow users create --username admin --password admin --firstname Admin --lastname User --role Admin --email admin@example.com"35 36networks:37 core:38 external: true
Once you are done, navigate to the compose directory and run the Airflow Compose file:
cd ml-platform/composedocker compose -f docker-compose.airflow.yml run --rm airflow-initdocker compose -f docker-compose.airflow.yml up -d
You can access the Airflow Web UI and change the admin credentials from the init command:
Now, create the Airflow DAG (Directed Acyclic Graph) that defines a machine learning pipeline with the command below:
sudo nano ml-platform/airflow/dags/train_example.py
Add the following ML pipeline to the file:
1from datetime import datetime2from airflow import DAG3from airflow.operators.bash import BashOperator4 5default_args = {"owner": "you", "retries": 0}6 7with DAG(8 dag_id="train_example",9 default_args=default_args,10 start_date=datetime(2025, 1, 1),11 schedule_interval=None,12 catchup=False13) as dag:14 15 16 dvc_pull = BashOperator(17 task_id="dvc_pull",18 bash_command="""19 cd /opt/airflow && \20 dvc pull -v21 """22 )23 24 25 pip_install = BashOperator(26 task_id="pip_install",27 bash_command="pip install -r /opt/airflow/ml/requirements.txt --no-cache-dir"28 )29 30 31 train = BashOperator(32 task_id="train",33 env={34 "MLFLOW_TRACKING_URI": "http://mlflow:5000",35 "AWS_ACCESS_KEY_ID": "{{ var.value.AWS_ACCESS_KEY_ID if var.value.AWS_ACCESS_KEY_ID else '' }}",36 "AWS_SECRET_ACCESS_KEY": "{{ var.value.AWS_SECRET_ACCESS_KEY if var.value.AWS_SECRET_ACCESS_KEY else '' }}",37 "MLFLOW_S3_ENDPOINT_URL": "http://minio:9000",38 },39 bash_command="""40 cd /opt/airflow && \41 python ml/train.py42 """43 )44 45 dvc_pull >> pip_install >> train
You can run the Airflow DAG using the Airflow REST API:
curl -X POST "http://localhost:8080/api/v1/dags/train_example/dagRuns" \-u admin:admin \-H "Content-Type: application/json" \-d '{"conf": {"run_id": "manual_1"}}'
Tip: To expand automation beyond ML workflows and include infrastructure and hosting operations, check out the AI Automated Hosting Operations tutorial.
CI/CD Pipeline for ML Platform with GitHub Actions
You can also set up a GitHub Actions workflow, which automates testing, containerization, and deployment of machine learning code. It runs tests, builds Docker images, pushes to a container registry, and can optionally trigger Airflow pipelines.
Create the GitHub Actions file with the following command:
sudo nano .github/workflows/ci.yml
Add the following config to the file:
1name: ci2on:3 push:4 branches: [ "main" ]5jobs:6 test-build:7 runs-on: ubuntu-latest8 steps:9 - uses: actions/checkout@v410 11 - uses: actions/setup-python@v512 with: { python-version: "3.11" }13 - run: python -m pip install -r ml/requirements.txt14 - run: python -m pip install pytest15 - run: pytest -q || true 16 17 18 - name: Log in to GHCR19 run: echo "${{ secrets.GHCR_TOKEN }}" | docker login ghcr.io -u ${{ github.actor }} --password-stdin20 - name: Build21 run: docker build -t ghcr.io/${{ github.repository }}/ml-app:latest -f Dockerfile .22 - name: Push23 run: docker push ghcr.io/${{ github.repository }}/ml-app:latest24 25 26 - name: Trigger Airflow27 run: |28 curl -X POST "http://YOUR_AIRFLOW_URL/api/v1/dags/train_example/dagRuns" \29 -u "${{ secrets.AIRFLOW_USER }}:${{ secrets.AIRFLOW_PASS }}" \30 -H "Content-Type: application/json" \31 -d '{"conf":{"source":"ci"}}'
If you prefer containerized tasks, you can create a simple Dockerfile to package your training code:
FROM python:3.11-slimWORKDIR /appCOPY ml/requirements.txt .RUN pip install --no-cache-dir -r requirements.txtCOPY ml/ /app/ml/CMD ["python", "ml/train.py"]
ML Model Serving with NVIDIA Triton Inference Server
You can deploy machine learning models for high-performance inference using NVIDIA Triton. It provides a production-ready serving solution with support for multiple frameworks such as TensorFlow, PyTorch, ONNX, and optimized GPU inference through a scalable inference server.
You can start an NVIDIA Triton Inference Server with GPU support by using the following commands:
mkdir -p serving/models docker run -d --name triton \ --gpus all \ -p 8000:8000 -p 8001:8001 -p 8002:8002 \ -v $PWD/serving/models:/models \ nvcr.io/nvidia/tritonserver:24.05-py3 \ tritonserver --model-repository=/models
There is an alternative method with a simple model served with FastAPI and Docker. It creates RESTful APIs for inference that can be containerized and deployed behind NGINX, which offers a straightforward solution for production deployment without complex model conversion.
Tip: For a deeper guide on configuring and optimizing PyTorch inference environments on virtual private servers, check PyTorch Model Inference Setup on VPS.
ML Platform Monitoring with Prometheus and Grafana
Monitoring the ML workflow is an essential task. This setup collects metrics from containers, hosts, and services using Prometheus, visualizes them in Grafana dashboards, and monitors resource usage, which gives you full visibility into system performance and model serving metrics.
Create the monitoring Compose file with the command below:
sudo nano ml-platform/compose/docker-compose.monitoring.yml
Add the following configuration to the file:
1version: "3.9"2services:3 prometheus:4 image: prom/prometheus:latest5 volumes:6 - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro7 networks: [core]8 ports: ["9090:9090"]9 10 cadvisor:11 image: gcr.io/cadvisor/cadvisor:latest12 privileged: true13 networks: [core]14 ports: ["5555:8080"]15 volumes:16 - /:/rootfs:ro17 - /var/run:/var/run:rw18 - /sys:/sys:ro19 - /var/lib/docker/:/var/lib/docker:ro20 21 grafana:22 image: grafana/grafana:latest23 networks: [core]24 ports: ["3000:3000"]25 volumes:26 - grafana:/var/lib/grafana27 28volumes:29 grafana:30 31networks:32 core:33 external: true
Then, create the Prometheus file with the command below:
sudo nano ml-platform/compose/prometheus.yml
Add the configuration to the file:
1global:2 scrape_interval: 15s3scrape_configs:4 - job_name: 'prometheus'5 static_configs: [{ targets: ['prometheus:9090'] }]6 - job_name: 'cadvisor'7 static_configs: [{ targets: ['cadvisor:8080'] }]8 - job_name: 'triton'9 static_configs: [{ targets: ['HOST_IP_OR_TRITON:8002'] }]
Navigate to the compose directory and bring up the monitoring:
cd ml-platform/composedocker compose -f docker-compose.monitoring.yml up -d
Access the Grafana dashboard with the following URL and change the default admin credentials:
You can connect Prometheus as your data source, and import pre-built dashboards to monitor Docker containers, host resources, and system performance in real time.
For additional insights on maximizing performance and resource utilization, you can refer to Optimizing VPS Resource Allocation with AI.
That's it, you are done with setting up Automated ML workflows on dedicated servers.
Conclusion
Building automated ML workflows on dedicated servers gives you the power to manage the entire ML platform from data collection to production deployment, with full visibility, reproducibility, and scalability. By using tools like Docker, Airflow, MLflow, DVC, and Prometheus, you create a powerful foundation for continuous experimentation and model delivery.
This tutorial was tested and deployed on PerLod Hosting, which offered optimized VPS and dedicated GPU servers for machine learning, AI automation, and high-performance computing.
We hope you enjoy this guide. Subscribe to our X and Facebook channels to get the latest updates and articles in machine learning.