In this tutorial, we want to discover how to set up an AIOps stack for server monitoring. An AIOps stack is a combination of tools that use AI and automation to monitor, analyze, and manage servers. It is a great way to detect issues, predict failures, and automate responses. In this guide, we will use the following tools for the AIOps stack:
- Metrics: Prometheus, Node Exporter, and Blackbox Exporter.
- Logs: Loki and Promtail.
- Visualization and alerting UI: Grafana.
- Alert routing: Alertmanager.
- AIOps: A small Python service that fetches Prometheus metrics, detects anomalies, and triggers Alertmanager via webhook.
In this guide, we use Ubuntu 24.04 server from PerLod Hosting, where you can find support for server monitoring with AIOps.
Requirements for AIOps Stack for Server Monitoring
To complete the guide steps, you need a fresh Ubuntu 24.04 VM with 4 vCPU, 8 GB RAM, and 100 GB disk. Also, you must open the required firewall ports, including:
Prometheus: 9090Alertmanager: 9093Loki: 3100Grafana: 3000Port 22And 80/443, if you use a reverse proxy.
Remember to set the correct timezone on your server:
sudo timedatectl set-timezone Asia/Dubai
Once you are done, proceed to the next step to install Docker and Docker Compose, which are used for setting up an AIOps stack.
Install Docker and Docker Compose For Setting up an AIOps Stack
Run the system update and install the required packages with the commands below:
sudo apt updatesudo apt install ca-certificates curl gnupg -y
Add Docker GPG key and repository to your server with the following commands:
sudo install -m 0755 -d /etc/apt/keyrings curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpgecho \ "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] \ https://download.docker.com/linux/ubuntu noble stable" \| sudo tee /etc/apt/sources.list.d/docker.list >/dev/null
Again, run the system update and install Docker and Docker Compose with:
sudo apt updatesudo apt install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin -y
You can add your user to the Docker group with the command below:
sudo usermod -aG docker $USER
Log out and log in again to apply the changes.
Create AIOps Stack Directory
You must create the AIOps stack directory and the required tools path in it. To do this, you can run the command below:
sudo mkdir -p /opt/aiops/{prometheus,alertmanager,grafana-provisioning/{datasources,dashboards},loki,promtail,blackbox,anomaly-detector}
Set the correct ownership for the AIOps stack directory with:
sudo chown -R $USER:$USER /opt/aiops
Switch to the AIOps stack directory:
Prometheus Configuration for AIOps Stack
Prometheus is the main monitoring engine. The first step is to configure Prometheus server monitoring by creating the Prometheus YAML file with the following command:
sudo nano /opt/aiops/prometheus/prometheus.yml
Add the following configuration to the file:
1global:2 scrape_interval: 15s3 evaluation_interval: 15s4 5alerting:6 alertmanagers:7 - static_configs:8 - targets: ["alertmanager:9093"]9 10rule_files:11 - /etc/prometheus/rules/*.yml12 13scrape_configs:14 - job_name: "prometheus"15 static_configs:16 - targets: ["prometheus:9090"]17 18 - job_name: "node-exporters"19 static_configs:20 21 - targets: ["10.0.0.11:9100","10.0.0.12:9100"]22 23 - job_name: "blackbox-http"24 metrics_path: /probe25 params:26 module: [http_2xx]27 static_configs:28 - targets:29 - https://example.com30 - https://your-api.example31 relabel_configs:32 - source_labels: [__address__]33 target_label: __param_target34 - source_labels: [__param_target]35 target_label: instance36 - target_label: __address__37 replacement: blackbox:9115
Also, you can create a simple alert rule in the Prometheus rule file:
sudo nano /opt/aiops/prometheus/rules/infra.yml
Add the following alert rule to the file:
1groups:2- name: infra3 rules:4 - alert: NodeDown5 expr: up{job="node-exporters"} == 06 for: 2m7 labels: {severity: critical}8 annotations:9 summary: "Node down: {{ $labels.instance }}"10 description: "No scrape data for 2m."11 12 - alert: HighCPU13 expr: 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 8514 for: 5m15 labels: {severity: warning}16 annotations:17 summary: "High CPU on {{ $labels.instance }}"18 description: "CPU usage >85% for 5m."19 20 - alert: DiskFilling21 expr: (node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} - node_filesystem_free_bytes{fstype!~"tmpfs|overlay"}) 22 / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} > 0.923 for: 10m24 labels: {severity: warning}25 annotations:26 summary: "Disk >90% on {{ $labels.instance }}"27 description: "Filesystem filling up."
Tip: To learn more about Prometheus Linux server monitoring, you can check this guide on Monitoring a Linux host using Prometheus.
Alertmanager Configuration for AIOps Stack
The next step is to set up Alertmanager for server monitoring, which handles notifications from Prometheus. Create the Alertmanager YAML file with:
sudo nano /opt/aiops/alertmanager/alertmanager.yml
Add the following configuration to the file. You can replace SMTP with real credentials or remove the email if not needed yet:
1route:2 receiver: default3 group_by: ["alertname", "instance"]4 group_wait: 30s5 group_interval: 5m6 repeat_interval: 3h7 8receivers:9- name: default10 email_configs:11 - to: "alerts@example.com"12 from: "aiops@example.com"13 smarthost: "smtp.example.com:587"14 auth_username: "aiops@example.com"15 auth_identity: "aiops@example.com"16 auth_password: "CHANGE_ME"17 webhook_configs:18 - url: "http://anomaly-detector:8080/alertmanager"19 send_resolved: true
Loki and Promtail Log Configuration for AIOps Stack
Loki is a log aggregation system, and Promtail is its log collector. Promtail gathers logs from your server and sends them to Loki, which makes it easy to search logs in Grafana alongside metrics.
Create the Loki YAML file with the following command:
sudo nano /opt/aiops/loki/config.yml
Add the following configuration to it:
1auth_enabled: false2server:3 http_listen_port: 31004common:5 path_prefix: /loki6 ring:7 instance_addr: 127.0.0.18 kvstore:9 store: inmemory10schema_config:11 configs:12 - from: 2024-01-0113 store: boltdb-shipper14 object_store: filesystem15 schema: v1316 index:17 prefix: index_18 period: 24h19storage_config:20 boltdb_shipper:21 active_index_directory: /loki/index22 cache_location: /loki/boltdb-cache23 filesystem:24 directory: /loki/chunks25limits_config:26 ingestion_burst_size_mb: 6427 ingestion_rate_mb: 3228 max_cache_freshness_per_query: 10m29chunk_store_config:30 max_look_back_period: 720h31table_manager:32 retention_deletes_enabled: true33 retention_period: 720h
Also, create the Promtail file with the command below:
sudo nano /opt/aiops/promtail/config.yml
Add the following config to the file:
1server:2 http_listen_port: 90803positions:4 filename: /positions.yml5clients:6 - url: http://loki:3100/loki/api/v1/push7scrape_configs:8 - job_name: system-logs9 static_configs:10 - targets: [localhost]11 labels:12 job: varlogs13 host: ${HOSTNAME}14 __path__: /var/log/*.log
Configure Blackbox Exporter for AIOps Stack
Blackbox Exporter checks external endpoints like APIs or websites using HTTP probes, which helps you monitor the availability and response of your web services.
At this point, you can create the Blackbox Exporter YAML file with the following command:
sudo nano /opt/aiops/blackbox/blackbox.yml
Add the following configuration to the file:
1modules:2 http_2xx:3 prober: http4 http:5 preferred_ip_protocol: ip46 no_follow_redirects: false7 fail_if_not_ssl: false8 fail_if_body_not_matches_regexp:9 - ".+"
Grafana Configuration for AIOps Stack
You can connect Grafana to Prometheus and Loki as data sources, and load dashboards to visualize server performance, logs, and alerts in one place.
To create the Grafana data sources YAML file, run the command below:
sudo nano /opt/aiops/grafana-provisioning/datasources/datasource.yml
Add the following configuration to the file:
1apiVersion: 12datasources:3 - name: Prometheus4 type: prometheus5 url: http://prometheus:90906 access: proxy7 isDefault: true8 - name: Loki9 type: loki10 url: http://loki:310011 access: proxy12 jsonData:13 maxLines: 1000
Optional Note: You can drop JSON dashboards in grafana-provisioning/dashboards/ and add a dashboards.yml provisioning file.
Create the Grafana dashboards YAML file with:
sudo nano /opt/aiops/grafana-provisioning/dashboards/dashboards.yml
Add the following config to the file:
1apiVersion: 12providers:3 - name: 'Default'4 orgId: 15 folder: 'AIOps'6 type: file7 options:8 path: /etc/grafana/dashboards
AIOps Microservice: Anomaly Detector Configuration
This custom Python microservice analyzes Prometheus metrics in real time using STL decomposition and z-score detection to identify anomalies. When it detects unusual behavior, it sends a synthetic alert to Alertmanager automatically.
Create the requirements file with the following command:
sudo nano /opt/aiops/anomaly-detector/requirements.txt
Add the following requirements to the file:
flask==3.0.3requests==2.32.3pandas==2.2.3numpy==2.1.1statsmodels==0.14.3
Then, create the anomaly detector script file with:
sudo nano /opt/aiops/anomaly-detector/app.py
Add the following script to the file:
1from flask import Flask, request, jsonify2import os, time, requests, json3import pandas as pd4import numpy as np5from datetime import datetime, timedelta6from statsmodels.tsa.seasonal import STL7 8PROM = os.getenv("PROM_URL", "http://prometheus:9090")9ALERTMAN = os.getenv("ALERTMAN_URL", "http://alertmanager:9093/api/v2/alerts")10QUERY = os.getenv("PROM_QUERY", '100 - (avg by (instance)(rate(node_cpu_seconds_total{mode="idle"}[5m]))*100)')11 12app = Flask(__name__)13 14def fetch_series(minutes=120, step="30s"):15 end = int(time.time())16 start = end - minutes*6017 url = f"{PROM}/api/v1/query_range"18 r = requests.get(url, params={"query": QUERY, "start": start, "end": end, "step": step}, timeout=30)19 r.raise_for_status()20 return r.json()21 22def stl_anomaly(values):23 if len(values) < 60:24 return None25 ts = pd.Series([float(v[1]) for v in values])26 stl = STL(ts, period=60, robust=True).fit()27 resid = stl.resid28 z = (resid - resid.mean()) / (resid.std() + 1e-9)29 if abs(z.iloc[-1]) > 3.5:30 return float(ts.iloc[-1]), float(z.iloc[-1])31 return None32 33def push_alert(instance, value, zscore):34 payload = [{35 "labels": {36 "alertname": "AIOpsDetectedAnomaly",37 "severity": "warning",38 "instance": instance39 },40 "annotations": {41 "summary": f"AIOps anomaly on {instance}",42 "description": f"Value={value:.2f}, z-score={zscore:.2f} on query: {QUERY}"43 }44 }]45 rr = requests.post(ALERTMAN, data=json.dumps(payload), headers={"Content-Type":"application/json"}, timeout=10)46 rr.raise_for_status()47 48@app.route("/run", methods=["POST","GET"])49def run():50 data = fetch_series()51 if data.get("status") != "success":52 return jsonify({"status":"error","msg":data}), 50053 result = data["data"]["result"]54 anomalies = []55 for series in result:56 metric = series.get("metric", {})57 inst = metric.get("instance","unknown")58 values = series.get("values", [])59 res = stl_anomaly(values)60 if res:61 val, z = res62 push_alert(inst, val, z)63 anomalies.append({"instance":inst,"value":val,"z":z})64 return jsonify({"status":"ok","anomalies":anomalies})65 66@app.route("/alertmanager", methods=["POST"])67def inbound():68 _ = request.json69 return jsonify({"status":"received"}), 20070 71if __name__ == "__main__":72 app.run(host="0.0.0.0", port=8080)
Now you must create the Dockerfile in the anomaly detector:
sudo nano /opt/aiops/anomaly-detector/Dockerfile
Add the following config to the file:
FROM python:3.11-slimWORKDIR /appCOPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txtCOPY app.py .ENV PROM_URL=http://prometheus:9090ENV ALERTMAN_URL=http://alertmanager:9093/api/v2/alertsENV PROM_QUERY=100 - (avg by (instance)(rate(node_cpu_seconds_total{mode="idle"}[5m]))*100)EXPOSE 8080CMD ["python","/app/app.py"]
Set up AIOps Stack Docker Compose File
At this point, you can easily create the whole stack in a Docker Compose file. Create the file with the command below:
sudo nano /opt/aiops/docker-compose.yml
Add the following config to the file:
1services:2 prometheus:3 image: prom/prometheus:v2.55.14 command:5 - "--config.file=/etc/prometheus/prometheus.yml"6 - "--storage.tsdb.path=/prometheus"7 - "--web.enable-lifecycle"8 volumes:9 - ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro10 - ./prometheus/rules:/etc/prometheus/rules:ro11 - prom_data:/prometheus12 ports: ["9090:9090"]13 restart: unless-stopped14 15 alertmanager:16 image: prom/alertmanager:v0.27.017 volumes:18 - ./alertmanager/alertmanager.yml:/etc/alertmanager/alertmanager.yml:ro19 ports: ["9093:9093"]20 restart: unless-stopped21 22 loki:23 image: grafana/loki:2.9.824 command: ["-config.file=/etc/loki/config.yml"]25 volumes:26 - ./loki/config.yml:/etc/loki/config.yml:ro27 - loki_data:/loki28 ports: ["3100:3100"]29 restart: unless-stopped30 31 promtail:32 image: grafana/promtail:2.9.833 command: ["-config.file=/etc/promtail/config.yml"]34 volumes:35 - ./promtail/config.yml:/etc/promtail/config.yml:ro36 - /var/log:/var/log37 - promtail_positions:/positions.yml38 restart: unless-stopped39 40 blackbox:41 image: prom/blackbox-exporter:v0.25.042 command: ["--config.file=/etc/blackbox/blackbox.yml"]43 volumes:44 - ./blackbox/blackbox.yml:/etc/blackbox/blackbox.yml:ro45 ports: ["9115:9115"]46 restart: unless-stopped47 48 grafana:49 image: grafana/grafana:11.2.050 environment:51 - GF_SECURITY_ADMIN_PASSWORD=ChangeMe!52 - GF_USERS_ALLOW_SIGN_UP=false53 volumes:54 - grafana_data:/var/lib/grafana55 - ./grafana-provisioning/datasources:/etc/grafana/provisioning/datasources:ro56 - ./grafana-provisioning/dashboards:/etc/grafana/dashboards57 - ./grafana-provisioning/dashboards/dashboards.yml:/etc/grafana/provisioning/dashboards/dashboards.yml:ro58 ports: ["3000:3000"]59 restart: unless-stopped60 61 anomaly-detector:62 build: ./anomaly-detector63 environment:64 - PROM_URL=http://prometheus:909065 - ALERTMAN_URL=http://alertmanager:9093/api/v2/alerts66 - PROM_QUERY=100 - (avg by (instance)(rate(node_cpu_seconds_total{mode="idle"}[5m]))*100)67 ports: ["8080:8080"]68 restart: unless-stopped69 70volumes:71 prom_data:72 grafana_data:73 loki_data:74 promtail_positions:
Once you are done, you can run the anomaly detector with the following commands:
docker compose build anomaly-detectordocker compose up -d
Check if the service is up and running with the command below:
Install Node Exporter on Each Server To Monitor
At this point, you must set up Node Exporter on each server target you want to monitor. It collects CPU, memory, disk, and network metrics and exposes them on port 9100, and then Prometheus scrapes these metrics for analysis and alerting.
To do this, you can run the following commands:
NODE_EXPORTER_VERSION=1.8.2cd /tmp curl -LO https://github.com/prometheus/node_exporter/releases/download/v${NODE_EXPORTER_VERSION}/node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz tar xzf node_exporter-*.tar.gzsudo mv node_exporter-*/node_exporter /usr/local/bin/sudo useradd --no-create-home --shell /usr/sbin/nologin nodeexp || true
Then, create the systemd service for Node Exporter with the command below:
sudo tee /etc/systemd/system/node_exporter.service >/dev/null <<'EOF'[Unit]Description=Prometheus Node ExporterAfter=network.target[Service]User=nodeexpGroup=nodeexpType=simpleExecStart=/usr/local/bin/node_exporterRestart=on-failure[Install]WantedBy=multi-user.targetEOF
Enable and start the service with the command below:
sudo systemctl daemon-reloadsudo systemctl enable --now node_exporter
sudo ss -lntp | grep 9100
You must be sure the Prometheus server can reach server_ip:9100.
Verifying AIOps Stack and Creating Dashboards
It is recommended to confirm that every component of the AIOps stack is running correctly. To verify your setup, open the Prometheus Web UI by navigating to the URL below:
http://your-server-ip:9090
From there, go to Status and then Targets. Make sure all targets show as UP. Also, check the Blackbox Exporter to confirm that your target pages are listed and also show as UP.
For Loki, navigate to the following URL:
http://your-server-ip:3100/ready
It must return ready, which means the service is active.
Then, navigate to the Grafana UI and log in with the default credentials:
http://your-server-ip:3000
In Grafana’s Explore section, select Prometheus and run the query up to confirm metrics are available, then switch to Loki and run the query {job="varlogs"} to verify logs are being collected correctly.
To quickly visualize your metrics, you can create a dashboard in Grafana.
Open Grafana, go to Dashboards and then Import, and enter the dashboard ID 1860 from Grafana.com to import Node Exporter Full and get a detailed view of your servers’ performance.
Alternatively, you can place your own JSON dashboard files in the /opt/aiops/grafana-provisioning/dashboards/ directory, and they’ll automatically appear under the AIOps folder when Grafana starts.
Test Alerts and Validate AIOps Anomalies
After setting up the AIOps stack, it’s important to test that alerts and anomaly detection are working correctly. You can simulate different scenarios to confirm that Prometheus, Alertmanager, and the AIOps anomaly detector respond as expected.
For example, stop the Node Exporter service on a monitored server to test a NodeDown alert after about two minutes.
You can also manually test the AIOps anomaly detection by running the command below:
curl -X POST http://your-server-ip:8080/run
If the last metric point is an outlier, it sends an AIOpsDetectedAnomaly alert to Alertmanager via webhook.
You can automate this process with a cron job, sidecar container, or Grafana alert rule to call /run periodically and keep continuous anomaly checks active.
That's it, you are done with setting up the AIOps stack for server monitoring.
Conclusion
Setting up an AIOps stack for server monitoring provides proactive insights, automatic anomaly detection, and clear visibility across your entire infrastructure. By combining Prometheus, Grafana, Loki, Alertmanager, and a lightweight Python anomaly detector, you can monitor system health and performance in real time and receive smart alerts.
We hope you enjoy this guide. Subscribe to our X and Facebook channels to get the latest server monitoring guides and tips.
For optimal performance and reliability of your AIOps monitoring stack, it is recommended to use Reliable Dedicated Servers or Flexible VPS hosting.
For further reading:
Learn Server Memory Disaggregation
Deploy PyTorch Model on a VPS Environment
File Integrity Monitoring for Secure Data Integrity