Machine Learning Engineer

Featherlessai

Completely RemoteFull TimeInformation Technology
Posted 4 days ago

Job description

Responsibilities

  • Optimize inference latency, throughput, and cost for large-scale ML models in production
  • Profile and bottleneck GPU/CPU inference pipelines including memory, kernels, batching, and IO
  • Implement and tune techniques such as quantization, KV-cache optimization, speculative decoding, and model pruning
  • Collaborate with research engineers to productionize new model architectures
  • Build and maintain inference-serving systems using Triton, custom runtimes, or bespoke stacks
  • Benchmark performance across NVIDIA/AMD GPUs, CPUs, and cloud setups
  • Improve system reliability, observability, and cost efficiency under real workloads

Requirements

  • Strong experience in ML inference optimization or high-performance ML systems
  • Solid understanding of deep learning internals such as attention, memory layout, and compute graphs
  • Hands-on experience with PyTorch or similar frameworks and model deployment
  • Familiarity with GPU performance tuning including CUDA, ROCm, Triton, or kernel-level optimizations
  • Experience scaling inference for real users
  • Ability to work in fast-moving startup environments

Preferred Qualifications

  • Experience with LLM or long-context model inference
  • Knowledge of inference frameworks like TensorRT, ONNX Runtime, vLLM, or Triton
  • Experience optimizing across different hardware vendors
  • Open-source contributions in ML systems or inference tooling
  • Background in distributed systems or low-latency services

About the Company

We are a Series A startup focused on building high-performance, reliable, and cost-efficient machine learning systems that serve real users.

Skills & tools

PyTorchCUDATritonLLMGPU Optimization

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 01ML inference optimization experience
  2. 02Deep learning internals knowledge
  3. 03PyTorch experience
  4. 04GPU performance tuning
  5. 05Scaling inference for real users
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches