Senior Solutions Engineer

TensorWave

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

About the Company

TensorWave's mission is to deliver seamless, secure, reliable, and resilient AI compute at scale. They have built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation.

Responsibilities

  • Resolve complex escalations acting as the final authority on issues exceeding GOC scope
  • Partner with customer technical leads to diagnose production issues and ensure rapid resolution
  • Develop diagnostic scripts and workarounds to maintain customer operations
  • Drive end-to-end P1 root cause analysis and deliver actionable post-incident analysis
  • Convert recurring customer pain points into evidence-based feature requests for the product roadmap
  • Document platform behaviors and refine GOC runbooks to build scalable knowledge

Requirements

  • 5–9 years in Infrastructure Engineering, Platform Engineering, or SRE
  • Deep experience in Kubernetes cluster administration and scheduler internals
  • Proficiency in orchestrating GPU workloads and diagnosing training job failures using ROCm or CUDA
  • Skilled in RDMA/RoCEv2, SRIOV, and BGP networking
  • Expert knowledge of Linux kernel networking, hugepages, and cgroups
  • Proficient in Python and Ansible for automation and diagnostic tool development
  • Strong technical communication skills for presenting findings to executive leadership

Preferred Qualifications

  • Prior experience in customer-facing engineering roles such as Solutions Engineering or Technical Support Engineering
  • Experience working in high-uptime environments requiring 24/7/365 availability

Benefits

  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance
  • 401(k)
  • Flexible PTO and Paid Holidays
  • Parental Leave

Skills & tools

KubernetesPythonAnsibleCUDALinux

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 015-9 years Infrastructure/SRE experience
  2. 02Kubernetes expertise
  3. 03GPU infrastructure proficiency (ROCm/CUDA)
  4. 04Networking skills (RDMA/RoCEv2/BGP)
  5. 05Linux kernel expertise
  6. 06Python and Ansible proficiency
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches