Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
-
Updated
Sep 21, 2026 - Python
Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
Cross-platform .NET performance engineering skill for coding agents, covering CPU, memory, GC, benchmarking, concurrency, startup, native profiling, GPU rendering, and production diagnostics on macOS, Windows, and Linux.
Profine automatically profiles and optimizes PyTorch training jobs on real GPUs, delivering measurable speedups and lower GPU costs before teams waste days tuning configs by hand.
Evidence-first tools to measure, verify, and optimize LLM inference across local GPUs and OpenAI-compatible APIs.
Automated GPU profiling analysis for Adreno — turns Snapdragon Profiler captures into actionable insights with LLM
NAV extracts and analyzes GPU performance traces from NVIDIA Nsight™ Systems (NSYS), enabling comparative analysis and visualization for efficient performance profiling and regression testing.
Per-precision XMX (matrix engine) profiling for Intel Arc GPUs via Level Zero metric streamers — observes any workload without wrapping it
Communication cost modeling for tensor parallel LLM inference with TP vs PP vs hybrid comparison, VRAM analysis, pipeline bubble modeling, regime detection, and cost-efficiency. Shows TP dominates on NVLink, PP has 47% bubble at 8 GPUs, and LLaMA-70B needs 8× A100 or 2× H100 for VRAM.
Reproducible MiniMax-H3 benchmarks on NVIDIA GB300: SGLang scaling, host eBPF CUDA tracing, Grace-GPU PMU traffic, energy, and Nsight comparison.
NAV extracts and analyzes GPU performance traces from NVIDIA Nsight™ Systems (NSYS), enabling comparative analysis and visualization for efficient performance profiling and regression testing.
Collection of examples and links that uses different profiling tools to show memory usage and timings.
JAX benchmarking, profiling and evaluation metrics for Flax NNX: a registry of pure-function metrics (regression, classification, calibration, uncertainty, forecasting, generative, image, text, audio, graph, fairness), XLA FLOP counting, roofline analysis, GPU and energy monitoring, regression detection, publication exports, W&B and MLflow.
Reproducible study of how tensor layout, reduction length, and pointer alignment drive vendor GEMM dispatch and latency cliffs.
Agent Skills and an MCP server for GPU performance profiling, benchmarking, optimization, and reporting, with an inference focus.
Attention backend benchmark on Turing GPUs comparing Vanilla, SDPA Math, SDPA Efficient, and a custom Triton FlashAttention implementation. SDPA efficient achieves 130× memory reduction and 10× speedup; Triton FA achieves O(n) memory but is 64× slower than SDPA efficient on RTX 2070.
Live 3D GPU visualizer synced to real PyTorch training telemetry — SM activity, memory bandwidth, kernel execution, rendered in real time.
Kernel-only profiling workflow for CUDA and Triton kernels with Nsight Compute, standardized reports, visual analysis, and vendor-portable adapters.
Profiling and Triton-based KV-cache optimization for protein language model inference on consumer GPUs.
Long-context benchmark pushing Qwen2-0.5B from 4K to 32K tokens on RTX 2070 using SDPA + chunked prefill. Shows 40x speedup at 8K, FP16 beating INT4 at long context, and that quantization is NOT a long-context solution — KV-cache is the real bottleneck.
Experimental macOS GPU benchmarking, RenderDoc capture/replay, Mesa llvmpipe counters, RDC slicing, and analysis.
To associate your repository with the gpu-profiling topic, visit your repo's landing page and select "manage topics."