Designed and delivered the company's first self-hosted SLM inference platform end-to-end, from architecture review to production: vLLM based, drop-in OpenAI Chat Completions API, LoRA adapter serving, FP8 quantization, schema-constrained structured outputs, per-use-case isolation, and KEDA autoscaling.
Migrated production workloads off external LLM vendors onto the platform, saving over $7M annually, including an 87% reduction (~$5.8M/year) on a 60M request/day generation workload and a 93% reduction (~$1.34M/year) on a fine-tuned ranking model ramped to 100% of traffic.
Delivered inline query embeddings on the live job-search path at ~20ms p95, a latency budget third-party embedding APIs (~500ms p90) could not meet.
Productionized a fine-tuned SLM reranker at flat online latency, making a 44% NDCG@10 gain and 63% fewer bad results shippable.
Led the platform's full lifecycle from an accepted architecture design review through production deployment, standing up 2 new repositories, service-to-service authentication, and 3 Production Readiness Reviews that standardized onboarding for new use cases.
Adopted genai-bench for inference benchmarking and contributed a bug fix back to the open source project.
Kept the ranking rollout healthy through A/B testing, resolving production issues in real time, implementing warmup orchestration to cut cold-start latency, and tuning KEDA downscaling to absorb cron-driven traffic spikes.
Contributed key design input to the LLM proxy integration, shaping the routing layer through which client teams consume the inference platform.
Spearheaded cost-saving initiatives delivering over $3.3M in annual savings, including sunsetting a legacy email alerts product (~$1.7M/year), cutting the team's AWS account costs (~$1.44M/year), and removing unused data builders (~$180k/year).
Traced a GPU cost-attribution bug misreporting spend across every GPU deployment in the company, driving the cross-team investigation and proposing the fix.
Led the migration of bidding models onto the shared model-serving platform as the single point of contact for the initiative, establishing SLO monitoring and shadow testing for safe rollout, and consolidated bidding onto a global cluster, simplifying deployments and improving CPU utilization.
Drove the creation of comprehensive monitoring dashboards and SLOs for critical services, enhancing system observability and reliability.
Owned and resolved an 11-month availability failure in a business-critical data pipeline, coordinating 3 engineers and an intern across 5 partner teams, migrating 3 of its indexes to AWS, and restoring the model refresh cadence to daily.
Built the data-comparison tooling that validated those AWS index migrations within a 1% deviation, adopted by every team performing the migration.
Led the migration of SERP second-phase scoring off the index service onto the shared model-serving platform, from design through implementation, across three partner teams.
Led the handover of an entire product area to a new team, writing the transition documentation, running knowledge-sharing sessions, and onboarding their engineers into the on-call rotation.
Identified and drove company-wide storage savings, including disabling versioning on outbox S3 buckets (2.2PB reclaimed, ~$29k/month) and applying lifecycle policies to analytics buckets.
Pioneered research into Java 21 with Generational ZGC, documenting the findings and a benchmarking plan for adoption across the platform's services.
Stepped up as team lead, mentored and onboarded numerous engineers and an intern, and supported the growth of 4+ engineers through regular 1:1s.
Co-created a new pair programming interview, building the questions and defining the scoring signals, removing 2 interviews from each candidate's loop.
Set the team's AI-assisted engineering practice, running demos and knowledge-sharing sessions on a spec-driven, agent-assisted workflow that other engineers adopted as their own.
Corrected the team's testing strategy: wrote the proposal that became a team objective and guided a QA engineer in raising unit test coverage on a core service from under 10% to over 90%.
Presented the team's technical direction and business impact to senior engineering leadership, and authored the runbooks, load-testing tools, and wiki documentation the team onboards from.
Acted as a key technical partner and point of contact for numerous cross-functional teams, unblocking critical initiatives and ensuring successful project delivery.
Improved developer productivity by enabling caching and incremental builds on a Gradle task used in every Java build company-wide, cutting time spent on it by 30%.