Senior AI Systems Engineer. I optimize LLM inference at the metal — CUDA kernels, ARM NEON intrinsics, Vulkan compute shaders, quantization pipelines. From 300% speedups on NVIDIA Jetson to building a Vulkan LLM runtime from scratch.
A framework-free LLM agent that diagnoses automotive ECU faults from live OBD-II signals. 62 tools across 9 namespaces, data-driven tool registry with scored search, explicit context compaction with scratchpad persistence, and structurally-isolated typed subagents. Graded 5-dimension eval rubric with deterministic CI replay.
Built a from-scratch Vulkan 1.1 compute inference engine for on-device LLM inference. 9 GGUF quantization shaders (Q4_0 through IQ4_XS), thermal governor, paged KV cache, and Kotlin SDK. Discovered Qualcomm's 6ms Vulkan fence overhead, pivoted to NEON SDOT achieving 11.2 tok/s.
Integrated the Q4_HQQ quantization format deep into llama.cpp with custom AVX2 and NEON SIMD kernels. Built a fast Python GGUF converter and enabled KV cache quantization. 3.3x model footprint reduction with near-lossless quality.
Redesigned ORB-SLAM3's memory architecture and execution pipeline for TI TDA4VM DSP hardware. Achieved ~3x rendering and localization speedup, unlocking real-time visual SLAM on constrained embedded devices.
Production RAG system with physically-isolated role-based access control, hybrid dense + sparse retrieval via Reciprocal Rank Fusion, grounded answer generation with citation tracking, and a full evaluation pipeline (RAGAs + retrieval benchmarking). Five roles, four access levels, 167 tests.
Integrated BEVDet and BEVFormer perception models into Autoware Universe. Real-time multi-sensor fusion with deterministic latency on NVIDIA Jetson Orin. ~3x pipeline throughput increase for production AV stacks.
I'm looking for my next role in AI systems engineering, inference optimization, or on-device ML. If you're building something that needs to run fast on real hardware, I'd love to hear about it.