AI Researcher — Inference Optimization

Featherlessai

Completely RemoteFull TimeInformation Technology
Posted 4 days ago

Job description

Responsibilities

  • Research and develop techniques to optimize inference performance for large neural networks
  • Improve latency, throughput, memory efficiency, and cost per inference
  • Design and evaluate model-level optimizations such as quantization, pruning, and KV-cache optimization
  • Implement systems-level optimizations including dynamic batching, kernel fusion, and multi-GPU inference
  • Benchmark inference workloads across hardware accelerators
  • Collaborate with engineering teams to deploy optimized inference pipelines
  • Translate research insights into production-ready improvements

Requirements

  • Strong background in machine learning, deep learning, or AI systems
  • Hands-on experience optimizing inference for large-scale models
  • Proficiency in Python and modern ML frameworks like PyTorch
  • Experience with inference tooling such as Triton, TensorRT, vLLM, or ONNX Runtime
  • Ability to design experiments and communicate results clearly

Preferred Qualifications

  • Experience deploying production inference systems at scale
  • Familiarity with distributed and multi-GPU inference
  • Experience contributing to open-source ML or inference frameworks
  • Authorship or co-authorship of peer-reviewed research papers
  • Experience working close to hardware using CUDA, ROCm, or profiling tools

Skills & tools

PythonPyTorchTritonTensorRTvLLM

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 01Machine learning or deep learning background
  2. 02Inference optimization experience
  3. 03Python proficiency
  4. 04PyTorch experience
  5. 05Inference tooling experience
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches