Lead Machine Learning Engineer, Inference & Performance

Egen

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

About the Company

Egen is a fast-growing and entrepreneurial company with a data-first mindset. We bring together the best engineering talent working with the most advanced technology platforms, including Google Cloud and Salesforce, to help clients drive action and impact through data and insights.

Responsibilities

  • Optimize inference by building and tuning production LLM serving with vLLM and SGLang
  • Profile and accelerate training runs to resolve bottlenecks using FlashAttention
  • Engineer for hardware by applying understanding of GPU architecture and attention internals
  • Deploy and operate multiple models within shared GPU clusters on GKE with autoscaling
  • Drive efficiency by owning GPU utilization and improving throughput-per-dollar
  • Collaborate and consult with clients to translate business needs into high-performance AI architectures

Requirements

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field
  • 5+ years of experience in ML/AI engineering focused on performance, infrastructure, or systems
  • Proven track record of deploying and optimizing models in production environments
  • Experience profiling and improving GPU utilization for training and/or inference
  • Mastery of Python and shell scripting
  • Hands-on experience with vLLM, SGLang, or comparable high-performance serving stacks
  • Strong Kubernetes (GKE) experience for deploying models on Google Cloud
  • Knowledge of Data Engineering and SQL

Preferred Qualifications

  • Experience with Classic Machine Learning (neural nets, training, tuning)
  • Ability to reason about lower-level (CUDA-adjacent) performance code

Benefits

  • Comprehensive Health Insurance
  • Paid Leave (Vacation/PTO)
  • Paid Holidays
  • Sick Leave
  • Parental Leave
  • Bereavement Leave
  • 401(k) Employer Match
  • Employee Referral Bonuses

Skills & tools

PythonKubernetesLLM

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 015+ years ML/AI engineering experience
  2. 02Python mastery
  3. 03vLLM or SGLang experience
  4. 04Kubernetes (GKE) experience
  5. 05GPU architecture knowledge
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches