OperateObservability Early access

Metrics, logs and traces
for everything you run.

GPU, InfiniBand, node and storage metrics are collected from the moment a machine starts. Read them in the console, in Grafana dashboards, or from any tool that speaks PromQL. Send your own metrics, logs and traces to the same place.

Opens with the first cohort. Nothing to install for GPU and system metrics.
Telemetry path. Schematic · not to scale.

Sources on the left, three stores in the middle and four ways to read on the right. GPUs, InfiniBand ports, machines, volumes, shared filesystems and buckets are collected automatically and feed the metrics store; machines also feed the logs store. 19 NVIDIA DCGM fields are read from every GPU. Your applications send metrics, logs and traces. Metrics are kept for 14 days, and one query returns up to 3,000 series. The console and alerts read metrics, dashboards read metrics and logs, and the query API reads all three. One route is lit: GPU metrics reach the console with no agent to install, where gpu utilization, power draw, gpu temperature and the other signals are listed. Connections: GPUs to Metrics: no agent; InfiniBand ports to Metrics; Machines to Metrics; Volumes to Metrics; Shared filesystems to Metrics; Buckets to Metrics; Machines to Logs; Your applications to Metrics; Your applications to Logs; Your applications to Traces; Metrics to Console; Metrics to Dashboards; Metrics to Query API; Metrics to Alerts; Logs to Dashboards; Logs to Query API; Traces to Query API.

Collected automatically
  • GPUsNVIDIA DCGM
    fields
    19 per GPU
    names
    DCGM_FI_*
    agent
    none
  • InfiniBand ports
  • Machines
  • Volumes
  • Shared filesystems
  • Buckets
  • Your applicationsYou send

    Metrics, logs and traces

  • Metrics (highlighted)
    in
    Prometheus remote-write
    out
    PromQL
    Kept
    14days
    Series per query
    3,000
  • Logs
    in
    OpenTelemetry Collector
    out
    LogQL
  • Traces
    in
    OTLP over gRPC
    out
    TraceQL
  • ConsoleNo setup
    GPU utilization
    %
    Power draw
    W
    GPU temperature
    °C
  • DashboardsOn Grafana
    • GPUs
    • Volumes
    • Filesystems
    • Buckets
  • Query API
    • PromQL
    • LogQL
    • TraceQL
  • AlertsBy email
    • Pending
    • Firing
    • Email
StoresReadOpen standards inOpen standards out
Automatic
  • GPUsNVIDIA DCGM
    fields
    19 per GPU
    names
    DCGM_FI_*
  • InfiniBand ports
  • Machines
  • Volumes
  • Shared filesystems
  • Buckets
  • Your applications

    Metrics, logs and traces

  • Metrics (highlighted)
    in
    Prometheus remote-write
    out
    PromQL
    Kept
    14days
    Per query
    3,000
  • Logs
    in
    OpenTelemetry Collector
    out
    LogQL
  • Traces
    in
    OTLP over gRPC
    out
    TraceQL
  • ConsoleNo setup
    GPU utilization
    %
    Power draw
    W
    GPU temperature
    °C
  • Dashboards

    Prebuilt, on Grafana

  • Query API

    All three stores

  • Alerts

    Thresholds, by email

StoresRead
  • Collected automaticallyNo agent
    • GPUs
    • InfiniBand
    • Machines
    • Volumes
    • Filesystems
    • Buckets
    to
    Logs, from machines
  • Metrics (highlighted)
    in
    Prometheus remote-write
    out
    PromQL
    Kept
    14days
    Series per query
    3,000
    to
    Dashboards · Query API · Alerts
  • ConsoleNo setup
    GPU utilization
    %
    Power draw
    W
    GPU temperature
    °C
  • Your applicationsYou send

    Metrics, logs and traces

    to
    Metrics · Traces
  • Logs
    in
    OpenTelemetry Collector
    out
    LogQL
    to
    Dashboards · Query API
  • Traces
    in
    OTLP over gRPC
    out
    TraceQL
    to
    Query API
  • Dashboards

    Prebuilt, on Grafana

  • Query API

    PromQL · LogQL · TraceQL

  • Alerts

    Thresholds, by email

  • GPU metrics, no agent to install
§02 Code Preview API
Preview API · names illustrative
import fantasti

fx = fantasti.Client()                      # reads FANTASTI_API_KEY

series = fx.metrics.query(
    'avg by (node) (gpu_utilization{cluster="pretrain-a"})',
    start="-6h", step="1m",
)
idle = [s.labels["node"] for s in series if s.mean() < 20]
print(idle)

Output · illustrative

['node-3', 'node-7']
# prometheus.yml: send your own metrics
remote_write:
  - url: ${FANTASTI_METRICS_WRITE_URL}
    authorization:
      credentials: ${FANTASTI_API_KEY}
# traces from an application that already uses OpenTelemetry
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
export OTEL_EXPORTER_OTLP_ENDPOINT="$FANTASTI_TRACES_ENDPOINT"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer $FANTASTI_API_KEY"
FANTASTI_METRICS_URL
PromQL queries
FANTASTI_METRICS_WRITE_URL
Prometheus remote-write
FANTASTI_TRACES_ENDPOINT
OTLP over gRPC

Open standards in,
open standards out.

Prometheus remote-write, LogQL and OTLP on the way in. PromQL, LogQL and TraceQL on the way out. Nothing here needs a Fantasti-specific agent in your code.

Read the API

§03 GPU signals 19 fields per GPU

What we collect from every GPU.

TAB 01GPU signals Source · NVIDIA DCGM field identifiers Reviewed 2026-10-09
GPU signals: signal, NVIDIA DCGM field, unit and number of fields
Signal NVIDIA DCGM field Unit Fields
GPU utilization DCGM_FI_DEV_GPU_UTIL % 1
Memory utilization DCGM_FI_DEV_MEM_COPY_UTIL % 1
Frame buffer used, free, total, reserved DCGM_FI_DEV_FB_USED, _FB_FREE, _FB_TOTAL, _FB_RESERVED MB 4
Power draw DCGM_FI_DEV_POWER_USAGE W 1
Power limit DCGM_FI_DEV_POWER_MGMT_LIMIT W 1
Energy since driver load DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION mJ 1
GPU temperature DCGM_FI_DEV_GPU_TEMP °C 1
Memory temperature DCGM_FI_DEV_MEMORY_TEMP °C 1
Slowdown threshold DCGM_FI_DEV_SLOWDOWN_TEMP °C 1
Throttle reasons DCGM_FI_DEV_CLOCK_THROTTLE_REASONS bitmask 1
SM clock, memory clock DCGM_FI_DEV_SM_CLOCK, DCGM_FI_DEV_MEM_CLOCK frequency 2
PCIe traffic, receive and transmit DCGM_FI_PROF_PCIE_RX_BYTES, _TX_BYTES bytes 2
NVLink traffic, receive and transmit DCGM_FI_PROF_NVLINK_RX_BYTES, _TX_BYTES bytes/s 2
Per GPU NVIDIA DCGM fields collected 19
TAB 02Other signals Source · Fantasti Reviewed 2026-10-09
Other signals
SourceSignals
InfiniBand ports (clusters)Link downed, link error recovery, port data received and transmitted, transmit discards, receive errors, packets received and transmitted, transfer rate
MachineCPU utilization, memory total and used, disk bytes and operations, network bytes and packets
VolumesRead and write latency by quantile, throttling latency, read and write IOPS, read and write throughput, share of the IOPS limit in use
Shared filesystemsRead and write latency by quantile, read and write IOPS and throughput, read and write errors, index operations and errors
BucketsTraffic, total size, read and modify requests, space and object counts by storage class, API errors, lifecycle transition failures
Sheet
01 / 02
Title
Signals collected
Field names
NVIDIA DCGM
Reviewed
2026-10-09
EQ Signals per GPU Counted from TAB 01

19 signals per GPU, from NVIDIA DCGM.

One GPU · 19 NVIDIA DCGM fields
EQ 01 Source · TAB 01 · GPU signals Reviewed 2026-10-09
TAB 03Field register, one GPU Source · NVIDIA DCGM field identifiers Reviewed 2026-10-09
  1. GPU utilization: DCGM_FI_DEV_GPU_UTIL (%)
  2. Memory utilization: DCGM_FI_DEV_MEM_COPY_UTIL (%)
  3. Frame buffer used, free, total, reserved: DCGM_FI_DEV_FB_USED, DCGM_FI_DEV_FB_FREE, DCGM_FI_DEV_FB_TOTAL, DCGM_FI_DEV_FB_RESERVED (MB)
  4. Power draw: DCGM_FI_DEV_POWER_USAGE (W)
  5. Power limit: DCGM_FI_DEV_POWER_MGMT_LIMIT (W)
  6. Energy since driver load: DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION (mJ)
  7. GPU temperature: DCGM_FI_DEV_GPU_TEMP (°C)
  8. Memory temperature: DCGM_FI_DEV_MEMORY_TEMP (°C)
  9. Slowdown threshold: DCGM_FI_DEV_SLOWDOWN_TEMP (°C)
  10. Throttle reasons: DCGM_FI_DEV_CLOCK_THROTTLE_REASONS (bitmask)
  11. SM clock, memory clock: DCGM_FI_DEV_SM_CLOCK, DCGM_FI_DEV_MEM_CLOCK (frequency)
  12. PCIe traffic, receive and transmit: DCGM_FI_PROF_PCIE_RX_BYTES, DCGM_FI_PROF_PCIE_TX_BYTES (bytes)
  13. NVLink traffic, receive and transmit: DCGM_FI_PROF_NVLINK_RX_BYTES, DCGM_FI_PROF_NVLINK_TX_BYTES (bytes/s)
§04 Dashboards 4 prebuilt

Dashboards, already wired.

Open prebuilt dashboards for GPUs, volumes, shared filesystems and buckets with the data sources already connected. They run on Grafana, so you build your own next to them with PromQL and LogQL.

The dashboards run as an application on your own cluster.

Already run Grafana? Add Fantasti as a Prometheus data source for metrics and a Loki data source for logs, and keep your dashboards.

YAML
Preview API · names illustrative
# provisioning/datasources/fantasti.yaml
apiVersion: 1
datasources:
- name: Fantasti metrics
  type: prometheus
  url: ${FANTASTI_METRICS_URL}
  jsonData:
    httpHeaderName1: Authorization
  secureJsonData:
    httpHeaderValue1: >-
      Bearer ${FANTASTI_API_KEY}
Dashboards Data sources connected
Dashboards and their data sources. Schematic · not to scale.

Two data sources are connected before you open anything: Metrics, added as a Prometheus data source and queried with PromQL, and Logs, added as a Loki data source and queried with LogQL. The dashboard application runs on your own cluster and holds 4 prebuilt dashboards: GPUs, Volumes, Shared filesystems and Buckets. Each reads Metrics. GPUs: utilization, frame buffer, power, temperature, clocks, nvlink. Volumes: latency, iops, throughput, throttling, iops limit. Shared filesystems: latency, iops, throughput, errors, index operations. Buckets: traffic, size, requests, objects, api errors, lifecycle. Dashboards you build yourself sit next to them and read Metrics and Logs. Connections: Metrics to GPUs; Metrics to Volumes; Metrics to Shared filesystems; Metrics to Buckets; Metrics to Your own; Logs to Your own.

Dashboard applicationOn your own cluster
  • MetricsConnected
    type
    Prometheus
    query
    PromQL
  • LogsConnected
    type
    Loki
    query
    LogQL
  • GPUs19 fields per GPU
    • Utilization
    • Frame buffer
    • Power
    • Temperature
    • Clocks
    • NVLink
  • VolumesPrebuilt
    • Latency
    • IOPS
    • Throughput
    • Throttling
    • IOPS limit
  • Shared filesystemsPrebuilt
    • Latency
    • IOPS
    • Throughput
    • Errors
    • Index operations
  • BucketsPrebuilt
    • Traffic
    • Size
    • Requests
    • Objects
    • API errors
    • Lifecycle
  • Your ownPromQL · LogQL
  • LogsConnected
    type
    Loki
    query
    LogQL
    to
    Your own
  • MetricsConnected
    type
    Prometheus
    query
    PromQL
    to
    Volumes · Shared filesystems ·
    Buckets · Your own
  • GPUs19 fields
    • Utilization
    • Power
    • Temperature
    • NVLink
  • Volumes
    • Latency
    • IOPS
    • Throughput
    • IOPS limit
  • Shared filesystems
    • Latency
    • IOPS
    • Errors
    • Index ops
  • Buckets
    • Traffic
    • Size
    • Requests
    • API errors
  • Your ownYou build
    • PromQL
    • LogQL
Application · On your cluster
  • Connected when the application opens
  • Yours to build
§05 Your telemetry 4 ways in

Bring your own telemetry.

Send your own metrics, logs and traces to the same place.

  1. Metrics

    Send Prometheus remote-write from a Prometheus server, the Prometheus Operator or Vector.

    • Prometheus server
    • Prometheus Operator
    • Vector

    Lands in the Metrics store, read back with PromQL

  2. Cluster logs

    Collect pod and system logs with a Helm chart we provide, or with the OpenTelemetry Collector.

    • Helm chart
    • OpenTelemetry Collector

    Lands in the Logs store, read back with LogQL

  3. Machine logs

    Collect journald logs from your machines.

    • journald

    Lands in the Logs store, read back with LogQL

  4. Traces

    Export OTLP over gRPC from your application, or route it through the cluster agent.

    • OTLP over gRPC
    • Cluster agent

    Lands in the Traces store, read back with TraceQL

§06 Alerts Machine and bucket metrics

Know when a number crosses a line.

Pick a metric, a condition, a threshold and how long it must hold. An alert moves from Pending to Firing and notifies the email addresses you choose.

  • PendingThe condition is true and the hold time is running.
  • FiringIt held for the whole time. An email goes to the addresses on the rule.

In this preview, alerts cover machine and bucket metrics.

Alerts Delivered by email
PLT 01One alert rule illustrative
Metric
cpu_utilization · Machine
Condition
above
Threshold
90 %
Must hold for
5 minutes

An example rule: the machine metric cpu_utilization, above 90 %, must hold for 5 minutes. A made-up series stays under the threshold, then crosses it. The alert is Pending while the condition holds, and Firing once it has held for the time the rule names. An email goes to the addresses you chose.

§07 Limits Early access

Limits.

TAB 04Observability limits Source · Fantasti Reviewed 2026-10-09
Observability limits
Item Value
Metric retention 14 days
Log retention 14 days; longer by request
Series per query 3,000
Points per series 10,000
Log entry size 256 KB
Log batch 1,000 entries or 4 MB
Trace search TraceQL, basic filters
Sheet
02 / 02
Title
Observability limits
Reviewed
2026-10-09
§08 Questions 6 answers

Questions about Observability.

Do I need to install an agent?

Not for system and GPU metrics. New machines and cluster nodes report them automatically. Collecting application logs from a Kubernetes cluster uses a Helm chart.

Can I use my own Grafana?

Yes. Connect a Prometheus data source for metrics and a Loki data source for logs.

How long is data kept?

14 days for metrics and logs. Ask us if you need longer log retention.

Is there a service level?

Not during early access.

What does it cost?

Pricing is not published during early access.

Is Fantasti affiliated with Grafana Labs?

No. Grafana is open-source software from Grafana Labs.