Platform Security and AI Infrastructure Engineer. TS/SCI cleared, USAF cyber defense operations background.
I build the security and platform layer around model serving. My work covers authenticated inference paths, policy and audit controls, release-gate tooling, and the observability that makes serving behavior measurable. I like deterministic systems that can be tested under pressure and explained line by line.
- secure-gpu-inference-gateway: inference gateway demo showing the security control plane around a mock backend: access control, spending limits, audit logs, and supply-chain checks.
- deterministic-inference-scheduler: deterministic continuous-batching and paged KV-cache scheduler with replayable traces and release gates.
- triton-kernel-lab: fused Triton RMSNorm and SwiGLU kernels with FP32 oracles and measured latency comparisons.
- triton-inference-benchmark: load-generation toolkit for Triton and OpenAI-compatible serving. Its public CI includes a verified TLS 1.2+ path with explicit CA trust: two authenticated HTTPS agents, wrong-bearer and wrong-CA rejection before target work, and artifact privacy checks. The published
benchmark_report.pycompares ordered saved JSON runs with explicit p95-latency, throughput, success-rate, and retry-amplification gates without copying endpoints or prompts.lifecycle_qualification.pystarts an explicit service command without a shell, waits for loopback HTTP-200 readiness, then runs the benchmark and emits a privacy-safe startup report. Existing fixtures also recover one completed shard across an agent restart and both coordinator shards after response loss without duplicate target requests. Its publicqualification_manifest.pyre-derives a saved trend report and binds it to exact ordered input bytes with SHA-256 digests while keeping paths, prompts, endpoints, credentials, outputs, and raw telemetry out of the manifest. This is single-host synthetic protocol evidence; the lifecycle value is process-launch to selected health readiness, not model cold-start time, partial-workflow continuation, interrupted-child recovery, multi-host behavior, or a production-network measurement. - heterocore-compiler: compiler and cost model for mixed analog-digital inference accelerators.
- market-microstructure-engine: limit-order-book matching engine with a C++20 core and Python parity checks.
- readiness-control-tower: mission readiness decision tool built on synthetic operational data.
USAF cyber defense operations.
- Supported cyber defense across 26 wings.
- Cut threat resolution time from 5 days to 2.
- Hardened 30,000 assets.
- Named A6 Airman of the Year, 2025.
- Portfolio: https://wafflebits.github.io/WaffleBits/
- LinkedIn: https://www.linkedin.com/in/adnanberik/


