Lead AI Infrastructure & Distributed Systems Engineer

Supervity

  • Posted: 1 month ago
  • Openings: 10
  • Applicants: 0

Job Description

Lead AI Infrastructure & Distributed Systems Engineer


Role Overview:

We are seeking a Lead AI Infrastructure Engineer to design, scale, and maintain our local, high-speed GPU training rings. Instead of using unlimited cloud clusters, you will be responsible for orchestrating a cost-efficient cluster of basic/consumer GPUs (e.g., RTX 4090s/5090s, PCIe server nodes) to execute full-parameter distillation of 26B-parameter standalone models. You will eliminate hardware bottlenecks by slicing model architectures across our physical network topology.

Core Responsibilities:

* Design and implement distributed training topologies using PyTorch FSDP, DeepSpeed (ZeRO-3), and Megatron-LM.

* Fragment and assign the 4864 transformer layers of a dense 26B student model sequentially across available hardware nodes using Pipeline Parallelism (PP) and Tensor Parallelism (TP).

* Minimize data-transfer idle time by optimizing host-CPU RAM offloading, gradient checkpointing, and Linux system-level communication links (10GbE/100GbE networking, NCCL tuning).

* Manage hardware cluster health, profiling thermal thresholds, and maximizing PCIe lane utilization on mixed or consumer-grade GPU arrays.


Required Technical Skills:

* Languages: Deep proficiency in Python and low-level C/C++ configurations.

* Distributed ML Architecture: 3+ years of experience setting up multi-node training routines (FSDP, DeepSpeed, or Horovod).

* Systems Networking: Mastery of Linux systems performance tuning, PCIe Gen4/Gen5 routing, and NCCL multi-GPU communication optimization.

* Hiring Signal: Experience working in university research labs, decentralized computing projects (e.g., Exo, Petals), or custom mining/rendering farm setups is a significant plus.


More Info

Full Time
o
Software Product
Not Disclosed
English
Not Disclosed

Education

Any Graduate
Not Disclosed

Required Skills

ZeRO3 GPU Communication Optimization Megatron PyTorch Distributed Training python RTX 4090/5090 Linux Performance Tuning Model Distillation

Contact Details

Supervity
+91 987654567
sales@supervity.ai
  • Experience3 years
  • Salary Above 10 LAKHS ANNUALLY
  • Location for Hiring Mumbai
  • Apply Now
Latest Job

Similar Jobs

HR Recruiter
Transient HR
  • 1 years
  • Mumbai
  • 18 Hours
Senior Human Resource Recruiter
Eclat Health Solutions
  • 2 years
  • Mumbai
  • 18 Hours
Recruitment Executive -HR
be3 Human Resource Management
  • Fresher
  • Mumbai
  • 18 Hours
Jr Recruiter -Ghansoli
PES Hr Services
  • Fresher
  • Mumbai
  • 18 Hours
HR Recruiter
Csml Group
  • 1 years
  • Mumbai
  • 18 Hours
Hr Recruiter Work From Home
Swati 7011890554
  • Fresher
  • Kolkata
  • 18 Hours