OperateObservability Early access
Metrics, logs and traces
for everything you run.
GPU, InfiniBand, node and storage metrics are collected from the moment a machine starts. Read them in the console, in Grafana dashboards, or from any tool that speaks PromQL. Send your own metrics, logs and traces to the same place.
Sources on the left, three stores in the middle and four ways to read on the right. GPUs, InfiniBand ports, machines, volumes, shared filesystems and buckets are collected automatically and feed the metrics store; machines also feed the logs store. 19 NVIDIA DCGM fields are read from every GPU. Your applications send metrics, logs and traces. Metrics are kept for 14 days, and one query returns up to 3,000 series. The console and alerts read metrics, dashboards read metrics and logs, and the query API reads all three. One route is lit: GPU metrics reach the console with no agent to install, where gpu utilization, power draw, gpu temperature and the other signals are listed. Connections: GPUs to Metrics: no agent; InfiniBand ports to Metrics; Machines to Metrics; Volumes to Metrics; Shared filesystems to Metrics; Buckets to Metrics; Machines to Logs; Your applications to Metrics; Your applications to Logs; Your applications to Traces; Metrics to Console; Metrics to Dashboards; Metrics to Query API; Metrics to Alerts; Logs to Dashboards; Logs to Query API; Traces to Query API.
- fields
- 19 per GPU
- names
- DCGM_FI_*
- agent
- none
Metrics, logs and traces
- in
- Prometheus remote-write
- out
- PromQL
- Kept
- 14days
- Series per query
- 3,000
- in
- OpenTelemetry Collector
- out
- LogQL
- in
- OTLP over gRPC
- out
- TraceQL
- GPU utilization
- %
- Power draw
- W
- GPU temperature
- °C
- GPUs
- Volumes
- Filesystems
- Buckets
- PromQL
- LogQL
- TraceQL
- Pending
- Firing
- fields
- 19 per GPU
- names
- DCGM_FI_*
Metrics, logs and traces
- in
- Prometheus remote-write
- out
- PromQL
- Kept
- 14days
- Per query
- 3,000
- in
- OpenTelemetry Collector
- out
- LogQL
- in
- OTLP over gRPC
- out
- TraceQL
- GPU utilization
- %
- Power draw
- W
- GPU temperature
- °C
Prebuilt, on Grafana
All three stores
Thresholds, by email
- GPUs
- InfiniBand
- Machines
- Volumes
- Filesystems
- Buckets
- to
- Logs, from machines
- in
- Prometheus remote-write
- out
- PromQL
- Kept
- 14days
- Series per query
- 3,000
- to
- Dashboards · Query API · Alerts
- GPU utilization
- %
- Power draw
- W
- GPU temperature
- °C
Metrics, logs and traces
- to
- Metrics · Traces
- in
- OpenTelemetry Collector
- out
- LogQL
- to
- Dashboards · Query API
- in
- OTLP over gRPC
- out
- TraceQL
- to
- Query API
Prebuilt, on Grafana
PromQL · LogQL · TraceQL
Thresholds, by email
- GPU metrics, no agent to install
The five parts of Observability
- Metrics PromQL Collected automatically. Query through a Prometheus-compatible API.
- Logs LogQL Service logs and your own. Export to JSON or Parquet.
- Traces OpenTelemetry Send OTLP. Search with TraceQL filters.
- Alerts Thresholds On machine and bucket metrics, delivered by email.
- Dashboards Prebuilt For GPUs, volumes, shared filesystems and buckets.
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
series = fx.metrics.query(
'avg by (node) (gpu_utilization{cluster="pretrain-a"})',
start="-6h", step="1m",
)
idle = [s.labels["node"] for s in series if s.mean() < 20]
print(idle) Output · illustrative
['node-3', 'node-7'] # prometheus.yml: send your own metrics
remote_write:
- url: ${FANTASTI_METRICS_WRITE_URL}
authorization:
credentials: ${FANTASTI_API_KEY} # traces from an application that already uses OpenTelemetry
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
export OTEL_EXPORTER_OTLP_ENDPOINT="$FANTASTI_TRACES_ENDPOINT"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer $FANTASTI_API_KEY" - FANTASTI_METRICS_URL
- PromQL queries
- FANTASTI_METRICS_WRITE_URL
- Prometheus remote-write
- FANTASTI_TRACES_ENDPOINT
- OTLP over gRPC
Open standards in,
open standards out.
Prometheus remote-write, LogQL and OTLP on the way in. PromQL, LogQL and TraceQL on the way out. Nothing here needs a Fantasti-specific agent in your code.
What we collect from every GPU.
| Signal | NVIDIA DCGM field | Unit | Fields |
|---|---|---|---|
| GPU utilization | DCGM_FI_DEV_GPU_UTIL | % | 1 |
| Memory utilization | DCGM_FI_DEV_MEM_COPY_UTIL | % | 1 |
| Frame buffer used, free, total, reserved | DCGM_FI_DEV_FB_USED, _FB_FREE, _FB_TOTAL, _FB_RESERVED | MB | 4 |
| Power draw | DCGM_FI_DEV_POWER_USAGE | W | 1 |
| Power limit | DCGM_FI_DEV_POWER_MGMT_LIMIT | W | 1 |
| Energy since driver load | DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION | mJ | 1 |
| GPU temperature | DCGM_FI_DEV_GPU_TEMP | °C | 1 |
| Memory temperature | DCGM_FI_DEV_MEMORY_TEMP | °C | 1 |
| Slowdown threshold | DCGM_FI_DEV_SLOWDOWN_TEMP | °C | 1 |
| Throttle reasons | DCGM_FI_DEV_CLOCK_THROTTLE_REASONS | bitmask | 1 |
| SM clock, memory clock | DCGM_FI_DEV_SM_CLOCK, DCGM_FI_DEV_MEM_CLOCK | frequency | 2 |
| PCIe traffic, receive and transmit | DCGM_FI_PROF_PCIE_RX_BYTES, _TX_BYTES | bytes | 2 |
| NVLink traffic, receive and transmit | DCGM_FI_PROF_NVLINK_RX_BYTES, _TX_BYTES | bytes/s | 2 |
| Per GPU | NVIDIA DCGM fields collected | 19 |
| Source | Signals |
|---|---|
| InfiniBand ports (clusters) | Link downed, link error recovery, port data received and transmitted, transmit discards, receive errors, packets received and transmitted, transfer rate |
| Machine | CPU utilization, memory total and used, disk bytes and operations, network bytes and packets |
| Volumes | Read and write latency by quantile, throttling latency, read and write IOPS, read and write throughput, share of the IOPS limit in use |
| Shared filesystems | Read and write latency by quantile, read and write IOPS and throughput, read and write errors, index operations and errors |
| Buckets | Traffic, total size, read and modify requests, space and object counts by storage class, API errors, lifecycle transition failures |
- Sheet
- 01 / 02
- Title
- Signals collected
- Field names
- NVIDIA DCGM
- Reviewed
- 2026-10-09
19 signals per GPU, from NVIDIA DCGM.
- GPU utilization: DCGM_FI_DEV_GPU_UTIL (%)
- Memory utilization: DCGM_FI_DEV_MEM_COPY_UTIL (%)
- Frame buffer used, free, total, reserved: DCGM_FI_DEV_FB_USED, DCGM_FI_DEV_FB_FREE, DCGM_FI_DEV_FB_TOTAL, DCGM_FI_DEV_FB_RESERVED (MB)
- Power draw: DCGM_FI_DEV_POWER_USAGE (W)
- Power limit: DCGM_FI_DEV_POWER_MGMT_LIMIT (W)
- Energy since driver load: DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION (mJ)
- GPU temperature: DCGM_FI_DEV_GPU_TEMP (°C)
- Memory temperature: DCGM_FI_DEV_MEMORY_TEMP (°C)
- Slowdown threshold: DCGM_FI_DEV_SLOWDOWN_TEMP (°C)
- Throttle reasons: DCGM_FI_DEV_CLOCK_THROTTLE_REASONS (bitmask)
- SM clock, memory clock: DCGM_FI_DEV_SM_CLOCK, DCGM_FI_DEV_MEM_CLOCK (frequency)
- PCIe traffic, receive and transmit: DCGM_FI_PROF_PCIE_RX_BYTES, DCGM_FI_PROF_PCIE_TX_BYTES (bytes)
- NVLink traffic, receive and transmit: DCGM_FI_PROF_NVLINK_RX_BYTES, DCGM_FI_PROF_NVLINK_TX_BYTES (bytes/s)
Dashboards, already wired.
Open prebuilt dashboards for GPUs, volumes, shared filesystems and buckets with the data sources already connected. They run on Grafana, so you build your own next to them with PromQL and LogQL.
The dashboards run as an application on your own cluster.
Already run Grafana? Add Fantasti as a Prometheus data source for metrics and a Loki data source for logs, and keep your dashboards.
# provisioning/datasources/fantasti.yaml
apiVersion: 1
datasources:
- name: Fantasti metrics
type: prometheus
url: ${FANTASTI_METRICS_URL}
jsonData:
httpHeaderName1: Authorization
secureJsonData:
httpHeaderValue1: >-
Bearer ${FANTASTI_API_KEY} Two data sources are connected before you open anything: Metrics, added as a Prometheus data source and queried with PromQL, and Logs, added as a Loki data source and queried with LogQL. The dashboard application runs on your own cluster and holds 4 prebuilt dashboards: GPUs, Volumes, Shared filesystems and Buckets. Each reads Metrics. GPUs: utilization, frame buffer, power, temperature, clocks, nvlink. Volumes: latency, iops, throughput, throttling, iops limit. Shared filesystems: latency, iops, throughput, errors, index operations. Buckets: traffic, size, requests, objects, api errors, lifecycle. Dashboards you build yourself sit next to them and read Metrics and Logs. Connections: Metrics to GPUs; Metrics to Volumes; Metrics to Shared filesystems; Metrics to Buckets; Metrics to Your own; Logs to Your own.
- type
- Prometheus
- query
- PromQL
- type
- Loki
- query
- LogQL
- Utilization
- Frame buffer
- Power
- Temperature
- Clocks
- NVLink
- Latency
- IOPS
- Throughput
- Throttling
- IOPS limit
- Latency
- IOPS
- Throughput
- Errors
- Index operations
- Traffic
- Size
- Requests
- Objects
- API errors
- Lifecycle
- type
- Loki
- query
- LogQL
- to
- Your own
- type
- Prometheus
- query
- PromQL
- to
- Volumes · Shared filesystems ·
- Buckets · Your own
- Utilization
- Power
- Temperature
- NVLink
- Latency
- IOPS
- Throughput
- IOPS limit
- Latency
- IOPS
- Errors
- Index ops
- Traffic
- Size
- Requests
- API errors
- PromQL
- LogQL
- Connected when the application opens
- Yours to build
Bring your own telemetry.
Send your own metrics, logs and traces to the same place.
-
Metrics
Send Prometheus remote-write from a Prometheus server, the Prometheus Operator or Vector.
- Prometheus server
- Prometheus Operator
- Vector
-
Cluster logs
Collect pod and system logs with a Helm chart we provide, or with the OpenTelemetry Collector.
- Helm chart
- OpenTelemetry Collector
-
Machine logs
Collect journald logs from your machines.
- journald
-
Traces
Export OTLP over gRPC from your application, or route it through the cluster agent.
- OTLP over gRPC
- Cluster agent
Know when a number crosses a line.
Pick a metric, a condition, a threshold and how long it must hold. An alert moves from Pending to Firing and notifies the email addresses you choose.
- PendingThe condition is true and the hold time is running.
- FiringIt held for the whole time. An email goes to the addresses on the rule.
In this preview, alerts cover machine and bucket metrics.
- Metric
- cpu_utilization · Machine
- Condition
- above
- Threshold
- 90 %
- Must hold for
- 5 minutes
An example rule: the machine metric cpu_utilization, above 90 %, must hold for 5 minutes. A made-up series stays under the threshold, then crosses it. The alert is Pending while the condition holds, and Firing once it has held for the time the rule names. An email goes to the addresses you chose.
Limits.
| Item | Value |
|---|---|
| Metric retention | 14 days |
| Log retention | 14 days; longer by request |
| Series per query | 3,000 |
| Points per series | 10,000 |
| Log entry size | 256 KB |
| Log batch | 1,000 entries or 4 MB |
| Trace search | TraceQL, basic filters |
- Sheet
- 02 / 02
- Title
- Observability limits
- Reviewed
- 2026-10-09