Machine Learning Engineer - Inference & Fine-Tuning

Insaito Software Pvt. Ltd.

  • Posted: 1 year ago
  • Openings: 10
  • Applicants: 0

Job Description

Get hired by a US-based company focused on the US, UK, and European markets. This would include travel to the US office located in California. What we're looking for We are looking for a talented Machine Learning Engineer to lead the development of an Inference Service built on open-source technologies. You will design, deploy, and fine-tune machine learning models in the cloud using frameworks like Hugging Face, TensorFlow, PyTorch, and other open-source tools. The ideal candidate has hands-on experience building scalable inference services, fine-tuning pre-trained models, and working with cloud-native infrastructure. In this role, you'll work on exciting machine learning applications spanning NLP, computer vision, and custom tasksall with a focus on efficient deployment and real-time inference. If you're passionate about open-source technologies and cloud-based ML infrastructure, this is the perfect role for you! Responsibilities Build and optimize scalable inference pipeline using popular open-source frameworks (e.g. TensorFlow serving, ONNX). Design real-time API endpoints for model serving and integration using frameworks like FastAPI or Flask. Implement and optimize batch processing and streaming data pipelines to handle large-scale workloads. Fine-tune pre-trained models (e.g. Llama, GPT, YOLO), leverage open source tools like RayTune, and Hugging Face Trainer. Deploy inference services on cloud platforms (GCP, AWS, Azure) using a containerized environment with Docker and Kubernetes. Leverage cloud-native tools like Kubernetes, and Kubeflow to manage and scale model deployment. Implement robust monitoring and logging systems using Prometheus, Grafana, or ELK stack. Continuously optimize models for latency and throughput, troubleshooting bottlenecks. Qualifications/ Required Skills and Experiences Educational Background: Bachelors or Master’s degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience. Experience: 3+ years of hands-on experience deploying machine learning models into production, with a focus on inference services and fine-tuning. Proficiency in Python or C/CC+. Proven track record in creating high-performance libraries and tools. Proficiency in model serving frameworks such as TensorFlow Serving, TorchServe, and ONNX Runtime. Familiarity with hyperparameter optimization libraries like Ray Tune, Optuna, and Hugging Face Trainer. Experience with containerization tools (Docker) and orchestration platforms (Kubernetes, Helm). Strong grasp of low-level OS concepts, including multi-threading, memory management, networking, storage, performance, and scalability. Preferred: Knowledge of AI inference techniques like speculative decoding. Preferred: Experience with CUDA/Triton programming, Rust, Cython, and compiler technologies. Soft Skills: Excellent communication skills with the ability to explain complex ML concepts to both technical and non-technical stakeholders. Strong problem-solving and debugging skills, with the ability to analyze and resolve performance issues in production environments. Comfortable working in an Agile environment and collaborating across teams to deliver results.

More Info

Full Time
Software Product
Not Disclosed
English
Not Disclosed

Education

Any Graduate
Not Disclosed

Required Skills

Artificial Intelligence Computer science Machine Learning Machine Science Intelligence Computer Engineering

Contact Details

Not Disclosed
Not Disclosed
Not Disclosed
  • Experience1 years
  • Salary Not Disclosed
  • Location for Hiring Remote
  • Apply Now
Latest Job

Similar Jobs

  • 6+ years
  • Pune
  • 1 Year
Machine Learning Engineer
Buzzworks Business Services
  • 6+ years
  • Pune
  • 1 Year
  • 6+ years
  • Pune
  • 1 Year
  • 2 years
  • Pune
  • 1 Year
  • 3 years
  • Gurugram
  • 1 Year