Universal performance engineering for Python.
PerfX is a measurement-first performance platform for Python applications, web APIs, and AI/ML workloads. It provides timing, CPU and memory telemetry, optional GPU/TPU/NPU backends, statistically valid benchmarking, empirical complexity analysis, evidence-based bottleneck classification, and regression detection — all behind a single, hardware-agnostic API.
PerfX reports measurements, not claims. Every metric it cannot verify is reported as unavailable rather than estimated or fabricated.
- Design Principles
- Requirements
- Installation
- Quick Start
- Core API
- Benchmarking
- Empirical Complexity Analysis
- Bottleneck Classification
- Regression Detection
- Web Framework Integration
- Command Line Interface
- Configuration
- Reporting
- Architecture
- Testing
- Contributing
- Limitations
- License
- Author
| Principle | Description |
|---|---|
| Measurement over inference | Metrics are only reported when directly observable. |
| No fabricated hardware data | Unsupported metrics return unavailable, never a guessed value. |
| Dependency-light core | The core package has zero mandatory third-party dependencies. |
| Scientific honesty | Complexity results are labeled empirical estimates, not proofs. |
| Exception safety | Instrumentation never alters or swallows application exceptions. |
| Extensibility | New hardware vendors and frameworks integrate via a stable plugin protocol. |
- Python 3.10, 3.11, or 3.12
- No mandatory third-party packages
Optional extras enable additional telemetry and integrations:
| Extra | Enables |
|---|---|
psutil |
Process CPU percentage, RSS, context switches, core counts |
cuda |
NVIDIA NVML device discovery, utilization, power |
pytorch |
CUDA event kernel timing, VRAM, transfer profiling, MPS discovery |
tensorflow |
TensorFlow graph execution integration |
jax |
TPU device discovery, XLA compilation/execution separation |
numpy |
NumPy array input-size resolution |
pandas |
DataFrame / Series input-size resolution |
fastapi |
FastAPI/Starlette request-level instrumentation |
flask |
Flask request-level instrumentation |
django |
Django/DRF request-level instrumentation |
dev |
pytest, coverage, ruff, mypy, build |
full |
psutil, pynvml, torch, numpy, pandas |
pip install perfxWith optional extras:
pip install perfx[psutil]
pip install perfx[cuda]
pip install perfx[pytorch]
pip install perfx[fastapi]
pip install perfx[full]From source:
git clone https://github.com/SyntaxilitY/PerfX.git
cd PerfX
python -m venv .venv
.venv\Scripts\activate # Windows
source .venv/bin/activate # Linux / macOS
pip install -e .[dev]from perfx import performance
@performance(cpu=True, memory=True, gpu="auto")
def process(items):
return sorted(items)
process([3, 1, 2])from perfx.core.results_store import get_results
from perfx.reporting.console import render_console
result = get_results("__main__.process")[-1]
print(render_console(result))from perfx import performance
@performance(cpu=True, memory=True, gpu="auto", accelerator="auto")
def train_step(batch):
...Function metadata, signature, return values, and exceptions are preserved unchanged. Instrumentation failures never mask the original exception raised by the wrapped function.
from perfx import aperformance
@aperformance()
async def fetch(url):
...from perfx import performance_block
with performance_block("serialization") as block:
serialize(payload)
print(block.result.timing.wall_time_ns)from perfx import performance_class
@performance_class(include=["process", "transform"])
class Pipeline:
def process(self, data): ...
def transform(self, data): ...
print(Pipeline.performance_summary())A single execution is never treated as a valid benchmark. benchmark()
performs warmup iterations, runs a fixed number of timed repetitions, and
computes descriptive statistics including percentiles and outlier detection.
from perfx import benchmark
result = benchmark(sorted, args=([3, 1, 2],), warmup=5, iterations=30)
print(result.mean_ns)
print(result.p95_ns)
print(result.outliers_ns)Complexity analysis requires an explicit, safe workload generator. Functions
marked repeatable=False are refused, preventing accidental repeated
execution of non-idempotent operations such as database writes or payments.
from perfx import complexity
@complexity(workload=lambda n: list(range(n)), sizes=[100, 1000, 10000, 100000])
def sort_data(data):
return sorted(data)
report = sort_data.analyze()
print(report.best_model)
print(report.confidence)
print(report.note)All results are explicitly labeled as an Empirical Complexity Estimate, not a formal algorithmic proof.
from perfx import classify_bottleneck, recommend
from perfx.core.results_store import get_results
result = get_results("train_step")[-1]
bottleneck = classify_bottleneck(result)
recommendations = recommend(result, bottleneck)
print(bottleneck.classification)
print(bottleneck.evidence)
for r in recommendations:
print(r)Classification is derived from multiple measured signals and is never inferred from a single metric in isolation.
perfx baseline mypackage.train_step --output baseline.json
perfx baseline mypackage.train_step --output current.json
perfx compare baseline.json current.jsonExit codes:
| Code | Meaning |
|---|---|
| 0 | No regression |
| 1 | Performance regression detected |
| 2 | Configuration error |
| 3 | Execution error |
PerfX distinguishes server-side instrumentation (work done inside your process) from client-side load testing (end-to-end network latency).
from fastapi import FastAPI
from perfx.integrations.fastapi_middleware import install
app = FastAPI()
install(app)from flask import Flask
from perfx.integrations.flask_middleware import install
app = Flask(__name__)
install(app)MIDDLEWARE = [
"perfx.integrations.django_middleware.PerfXMiddleware",
]from perfx.benchmark.http import run_http_load_test, evaluate_http_result
from perfx.integrations.web import WebPerformanceThresholds
result = run_http_load_test("http://127.0.0.1:8000/endpoint", total_requests=200, concurrency=20)
verdict = evaluate_http_result(result, WebPerformanceThresholds(max_p95_ms=200))
print(verdict.needs_improvement)
print(verdict.reasons)perfx devices # Discover CPU / GPU / TPU / NPU hardware
perfx benchmark module:function # Run a statistical benchmark
perfx baseline module.function # Record a performance baseline
perfx compare baseline.json current.json
perfx report module.function # Print recorded results as JSON
perfx overhead module:function # Measure PerfX's own instrumentation cost
perfx --version
perfx --help[tool.perfx]
cpu = true
memory = true
gpu = "auto"
accelerator = "auto"
[tool.perfx.benchmark]
warmup = 5
iterations = 30
[tool.perfx.regression]
runtime_threshold = 10
memory_threshold = 15
gpu_threshold = 10| Format | Module |
|---|---|
| Console | perfx.reporting.console |
| JSON | perfx.reporting.json_reporter |
| Markdown | perfx.reporting.markdown |
| HTML | perfx.reporting.html |
from perfx.reporting.html import render_performance_html
from perfx.core.results_store import get_results
result = get_results("train_step")[-1]
html = render_performance_html(result)
with open("report.html", "w", encoding="utf-8") as f:
f.write(html)Application
|
Public API (performance, benchmark, complexity)
|
Instrumentation Layer
|
Measurement Engine
|
Hardware Abstraction Layer
|-- CPU Backend
|-- Memory Backend
|-- NVIDIA CUDA Backend
|-- AMD ROCm Backend (plugin extension point)
|-- Apple Metal / MPS Backend
|-- Google TPU Backend
|-- NPU Backend (plugin extension point)
|
Analysis Engine
|-- Statistics
|-- Complexity Analysis
|-- Bottleneck Classification
|-- Regression Detection
|
Reporting Engine
|-- Console / JSON / Markdown / HTML
|
CLI / Pytest / Web Framework / CI Integrations
The core engine has no direct dependency on any accelerator SDK, web
framework, or ML framework. All hardware-specific and framework-specific
behavior is implemented behind the AcceleratorBackend protocol and
discovered through the perfx.plugins entry-point group.
pip install -e .[dev]
pytest -v
pytest -m "not slow"
pytest --cov=perfx --cov-report=term-missingBrowser-based reports:
pip install pytest-html
pytest --html=report.html --self-contained-html
pytest --cov=perfx --cov-report=htmlIssues and pull requests are welcome at https://github.com/SyntaxilitY/PerfX.
Before submitting a change:
ruff check perfx
mypy perfx
pytest -m "not slow"New hardware or framework backends should implement the
AcceleratorBackend protocol and register through the perfx.plugins
entry-point group rather than modifying perfx.core.
Measurements are affected by CPU frequency scaling, thermal throttling, OS scheduling, cache and branch prediction behavior, garbage collection, GPU driver and allocator behavior, and general system load. Empirical complexity results are statistical inferences from measured samples, not formal mathematical proofs. Results are only meaningfully comparable across runs captured on identical or explicitly documented hardware and software environments.
Metrics that cannot be verified through an available hardware or framework API are never estimated. They are reported as unavailable.
Apache License 2.0
Tariq Mehmood
- GitHub: SyntaxilitY/PerfX
- PyPI: perfx