MLOps Engineer

Guddge

  • Posted: 8 hours ago
  • Openings: 10
  • Applicants: 0

Job Description

About the Role

We're looking for an MLOps Engineer to support and evolve production machine-learning platform that produces daily risk scores used by operational leaders to intervene before incidents occur. You'll own the reliability, deployment, and monitoring of a production ML pipeline that spans AWS (S3, Glue, SageMaker, Step Functions), and a fully automated Terraform + GitHub Actions delivery pipeline across multi environments.

This is a hands-on infrastructure-and-operations role. The model itself is trained offline; your focus is keeping the scoring pipeline dependable, observable, secure, and easy to change.


What You'll Do

  • Operate and improve the end-to-end inference pipeline (SAP HANA --> AWS Glue --> AWS SageMaker --> AWS Step Functions --> SAP HANA), keeping scheduled daily runs reliable.
  • Own the infrastructure-as-code: maintain Terraform modules and configuration, manage remote state in Terraform Cloud, and promote changes through dev --> qa --> prd.
  • Maintain and extend the CI/CD workflows in GitHub Actions
  • Keep the pipeline observable: structured JSON logging, CloudWatch Logs Insights queries, and Step Functions execution monitoring.
  • Support model monitoring (drift detection) and data-quality checks; tune alert thresholds and triage failure alerts.
  • Troubleshoot production incidents using established runbooks (database connectivity, SageMaker job failures, Glue write-back, scheduler issues) and drive root-cause fixes.
  • Enforce security and compliance controls: encryption, least-privilege IAM, secrets handling, and no-PII-in-logs discipline.
  • Collaborate with data scientists to deploy retrained models through the SageMaker Model Registry and cross-account promotion process.

Required Qualifications

  • 3+ years in MLOps supporting production systems.
  • Strong AWS experience, especially the data/ML services.
  • Production Terraform experience with a remote backend and multi-environment promotion.
  • Solid Python for data processing and operational scripting.
  • Experience operating CI/CD pipelines and diagnosing production failures from logs and metrics.

Technical Skillset Breakdown

The percentages reflect the relative weight of each area for day-to-day success in this role.


1. AWS Cloud & ML Services: ~35%

The core of the platform runs on AWS.

  • AWS Glue - Spark/PySpark ETL jobs, Glue connections (VPC/JDBC), Data Catalog, crawlers, job bookmarks
  • Amazon SageMaker - Processing Jobs, Model Registry, cross-account model packages, execution roles
  • AWS Step Functions - orchestration, Amazon States Language (ASL), retry/error handling
  • Amazon EventBridge Scheduler - cron scheduling, timezone handling
  • Amazon S3 - bucket policies, versioning, server-side encryption, lifecycle
  • Amazon CloudWatch - Logs, Logs Insights, monitoring
  • AWS Lake Formation - fine-grained catalog access control
  • VPC / Security Groups / Prefix Lists network connectivity to on-prem

2. Infrastructure as Code (Terraform): ~20%

Nearly all infrastructure is managed as code.

  • Terraform module authoring and reuse
  • Terraform Cloud (remote state, workspaces, locking)
  • Multi-environment configuration (.tfvars, per-env config files)
  • Provider configuration (AWS, AWSCC)
  • Managing drift, plan/apply safety, state troubleshooting

3. CI/CD & Automation: ~15%

Delivery is fully automated.

  • GitHub Actions: workflow authoring, reusable composite actions, environment approvals
  • Sequential environment promotion (dev qa prd)
  • Git and branch/PR workflows, GitHub CLI / AWS Kiro
  • Make for build automation

4. Python & ML Tooling: ~10%

Supporting the model code and tests.

  • Python 3.8+ for inference/ETL scripts and operational tooling
  • pandas, pyarrow, scipy for data processing
  • boto3 for AWS automation
  • SHAP (model explainability), imbalanced-learn (SMOTE) familiarity to support the model
  • Scikit-learn model artifacts (Random Forest, MinMaxScaler) and pickle handling
  • pytest + coverage (50% gate), Black, flake8, isort

5. Docker & Containerization: ~10%

Containers underpin local testing and reproducible builds.

  • Building and running Docker images for development and CI
  • Containerized test execution (e.g., PySpark unit tests that don't run natively on Windows)
  • Managing dependencies and reproducible environments across local and pipeline runs
  • Familiarity with container-based execution in AWS (Glue/SageMaker managed containers)

6. Observability, Reliability & Security: ~10%

Keeping it production-grade.

  • Structured JSON logging and a typed error hierarchy (config/data/model/inference/downstream)
  • Model drift and data-quality monitoring; alert threshold tuning
  • Incident triage using runbooks; alerting via SNS/SES dispatcher integration
  • Security controls: encryption in transit/at rest, least-privilege IAM, no PII in logs
  • Data-quality validation

Working Environment

  • This is a offshore role, you will be working from your base location. No travel required
  • Ability to work US Pacific Standard Time Zone (i.e. 8 pm IST to 5 am IST)

More Info

Full Time
o
Not Disclosed
English
Not Disclosed

Education

Any Graduate
Not Disclosed

Required Skills

Terraform github aws sagemaker Aws Glue python docker Glue AWS

Contact Details

Guddge
+91 987654567
info@guddge.com
  • Experience5 years
  • Salary Above 10 LAKHS ANNUALLY
  • Location for Hiring Remote
  • Apply Now
Latest Job

Similar Jobs

Senior Project Scheduler
Core Energy Systems
  • 5 years
  • Mumbai
  • 8 Hours
Technical Recruiter
Procreator
  • 4 years
  • Mumbai
  • 8 Hours
Project Technician
Esbee Dynamed
  • 1 years
  • Mumbai
  • 8 Hours
Coordinator
MSB Educational Institute
  • 2 years
  • Mumbai
  • 8 Hours
Autosys Scheduler / Workload Automation Engineer
CMMI Level 5 IT Company - MNC
  • 6+ years
  • Mumbai
  • 8 Hours