Technical Support Engineer

Together AI

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

Responsibilities

  • Resolve complex technical challenges involving Kubernetes GPU clusters
  • Act as a customer-facing SRE to ensure cluster health and stability
  • Monitor GPU cluster health and communicate hardware issues like thermal throttling or NVLink degradation
  • Maintain production infrastructure including Slurm cluster maintenance and node repair
  • Investigate storage and networking issues involving Weka filesystem or InfiniBand
  • Collaborate with Engineering and Product teams to drive the product roadmap
  • Maintain detailed documentation of system configurations and troubleshooting guides

Requirements

  • 3+ years in a customer-facing technical role
  • 1+ year supporting AI services or mission-critical SaaS APIs
  • Experience as an SRE or DevOps engineer with Kubernetes
  • Knowledge of AI, ML, and GPU technologies in HPC environments
  • Proficiency with infrastructure services like SLURM and Ansible
  • Experience with high-speed networking such as InfiniBand and RDMA
  • Experience with distributed storage systems like Weka or NFS
  • Strong troubleshooting skills for compute clusters and I/O issues

About the Company

Together AI is a research-driven artificial intelligence company on a mission to lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.

Skills & tools

KubernetesGPUSRE

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 013+ years customer-facing technical experience
  2. 02Kubernetes SRE or DevOps experience
  3. 03Knowledge of AI/ML and GPU technologies
  4. 04Experience with SLURM and InfiniBand
  5. 05Distributed storage troubleshooting skills
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches