Skip to content
View WaffleBits's full-sized avatar

Highlights

  • Pro

Block or report WaffleBits

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
WaffleBits/README.md

Adnan Berik

Platform Security and AI Infrastructure Engineer. TS/SCI cleared, USAF cyber defense operations background.

I build the security and platform layer around model serving. My work covers authenticated inference paths, policy and audit controls, release-gate tooling, and the observability that makes serving behavior measurable. I like deterministic systems that can be tested under pressure and explained line by line.

Selected work

  • secure-gpu-inference-gateway: inference gateway demo showing the security control plane around a mock backend: access control, spending limits, audit logs, and supply-chain checks.
  • deterministic-inference-scheduler: deterministic continuous-batching and paged KV-cache scheduler with replayable traces and release gates.
  • triton-kernel-lab: fused Triton RMSNorm and SwiGLU kernels with FP32 oracles and measured latency comparisons.
  • triton-inference-benchmark: load-generation toolkit for Triton and OpenAI-compatible serving. Its public CI includes a verified TLS 1.2+ path with explicit CA trust: two authenticated HTTPS agents, wrong-bearer and wrong-CA rejection before target work, and artifact privacy checks. The published benchmark_report.py compares ordered saved JSON runs with explicit p95-latency, throughput, success-rate, and retry-amplification gates without copying endpoints or prompts. lifecycle_qualification.py starts an explicit service command without a shell, waits for loopback HTTP-200 readiness, then runs the benchmark and emits a privacy-safe startup report. Existing fixtures also recover one completed shard across an agent restart and both coordinator shards after response loss without duplicate target requests. Its public qualification_manifest.py re-derives a saved trend report and binds it to exact ordered input bytes with SHA-256 digests while keeping paths, prompts, endpoints, credentials, outputs, and raw telemetry out of the manifest. This is single-host synthetic protocol evidence; the lifecycle value is process-launch to selected health readiness, not model cold-start time, partial-workflow continuation, interrupted-child recovery, multi-host behavior, or a production-network measurement.
  • heterocore-compiler: compiler and cost model for mixed analog-digital inference accelerators.
  • market-microstructure-engine: limit-order-book matching engine with a C++20 core and Python parity checks.
  • readiness-control-tower: mission readiness decision tool built on synthetic operational data.

Background

USAF cyber defense operations.

  • Supported cyber defense across 26 wings.
  • Cut threat resolution time from 5 days to 2.
  • Hardened 30,000 assets.
  • Named A6 Airman of the Year, 2025.

Links

Pinned Loading

  1. market-microstructure-engine market-microstructure-engine Public

    Deterministic Python and C++20 limit-order-book matching engine with latency benchmarks

    C++

  2. readiness-control-tower readiness-control-tower Public

    Mission readiness dashboard: root-cause scoring and what-if analysis over synthetic operational data. Live demo on Pages

    TypeScript 1

  3. secure-gpu-inference-gateway secure-gpu-inference-gateway Public

    Inference gateway demo: RBAC, token budgets, audit logging, Prometheus/OTLP observability

    Python 1

  4. triton-kernel-lab triton-kernel-lab Public

    Fused RMSNorm and SwiGLU GPU kernels in OpenAI Triton, validated against FP32 oracles with measured RTX 5070 Ti benchmarks

    Python

  5. deterministic-inference-scheduler deterministic-inference-scheduler Public

    Deterministic continuous-batching and paged KV-cache scheduler in Rust: replayable traces, promote/hold/rollback release gates

    Rust

  6. heterocore-compiler heterocore-compiler Public

    HeteroCore hub: ONNX compiler and analytical cost model for mixed analog-digital AI inference

    Python