Profiling

Stack: torch-spyre (new, Inductor-based).

Scope: performance — why is it slow? For correctness questions (why is the result wrong?) see Debugging.

Torch-Spyre provides tooling to measure the performance of PyTorch workloads running on the Spyre accelerator. The full design of the planned toolkit is in RFC 0601 — Spyre Profiling Toolkit.

The in-tree torch_spyre.profiler package is importable and exports FFDC retrieval (get_diagnostic_report). Device presence is torch.spyre.is_available(). Day-to-day performance work still goes through torch.profiler plus the integrations on this page (upstream Kineto with libaiupti, aiu-smi, aiu-trace-analyzer). Broader in-tree profiling APIs will land with RFC 0601.

What can be profiled today

Capability

Status

Where

Compiler pipeline logs

Available

Environment variables

FFDC diagnostic reports on Spyre compile/runtime/unimplemented failures

Available (TORCH_SPYRE_FFDC=1)

FFDC user guide · API: get_diagnostic_report · Environment variables

CPU-side timing with torch.profiler

Available

PyTorch Profiler

Device telemetry (power, temperature, bandwidth)

Available — PF and VF mode (IBM-internal distribution; public release tracked in #1335)

Device monitoring

Device-side kernel timing via ProfilerActivity.PrivateUse1

Merged and built by default (#1856) via upstream Kineto and libaiupti

PyTorch Profiler

Trace post-processing (aiu-trace-analyzer)

Available, known gaps

Trace analysis

torch.spyre.memory.memory_allocated() / max_memory_allocated()

Available — delegates to torch.accelerator.memory (PR #770)

Quick example

Scratchpad utilization metrics

Planned

RFC 0601

IR-instrumentation-based fine-grained profiler

Planned

RFC 0601

FFDC quick example

When TORCH_SPYRE_FFDC=1, frontend-compile, backend-compile, kernel-launch, and unimplemented-operation failures write a JSON diagnostic report. Retrieve the newest valid report with:

import torch
import torch_spyre

report = torch.spyre.get_diagnostic_report()
if report is not None:
    print(report["failure"]["category"], report["failure"]["message"])
    print(report["failure"]["file"], report["failure"]["lineno"])
    print(report["_report_path"])

See the FFDC user guide for failure categories, report locations, pod/CI workflow, and JSON triage. The API reference documents the function contract.

Memory API quick example

torch.spyre.memory re-exports torch.accelerator.memory, so the same memory-query calls used on CUDA apply to Spyre. The example below allocates a tensor, frees it, and reads the current and peak allocations:

import torch

# Reset the peak counter so max_memory_allocated() starts from zero.
torch.spyre.memory.reset_peak_memory_stats()

# Allocate on the device; memory_allocated() reflects the new total.
x = torch.rand((64, 64), dtype=torch.float16, device="spyre")
print(torch.spyre.memory.memory_allocated())     # bytes currently allocated

# Free the tensor. memory_allocated() drops back, but the peak persists.
del x
print(torch.spyre.memory.memory_allocated())     # current allocation
print(torch.spyre.memory.max_memory_allocated()) # peak since reset

The module also exposes reset_accumulated_memory_stats() and memory_stats().

Toolkit layers

Layer

Tool

Granularity

Application / PyTorch

torch.profiler + upstream Kineto

Kernel-level

Compiler frontend

Inductor logging

Pass-level

Compiler backend

IR instrumentation (planned)

Intra-kernel

Runtime

libaiupti kernel + memory events

Kernel + memory

Device / HW

aiu-smi

Device-level telemetry

Post-processing

aiu-trace-analyzer

Derived metrics

Profiling topics

See also

Work in Progress

Some subsystems above are labelled Planned and are under active development as part of RFC 0601. The APIs reflect planned design and may change.