Job Description
## About the Role
Crunchyroll is hiring a Senior MLOps Engineer in Hyderabad to build and operate the machine learning infrastructure required to move models reliably from research into production. The role sits within the AI/ML organization and focuses on ML pipelines, model deployment, observability, automation, cloud infrastructure, and production ML lifecycle management.
## Responsibilities
- Design, build, and maintain end-to-end ML infrastructure and production ML pipelines.
- Develop CI/CD pipelines for automated and reliable ML model delivery.
- Implement model registries, experiment tracking, versioning, and lifecycle management.
- Build monitoring and observability systems for model drift, degradation, and production anomalies.
- Partner with data scientists and ML engineers to productionize models.
- Optimize ML training and inference workflows for scalability, performance, and cost.
- Operate ML workloads using AWS SageMaker, Databricks, Kinesis, Lambda, EKS, and Docker.
- Build reproducible training, validation, deployment, monitoring, and retraining workflows.
- Integrate ML services with large-scale distributed systems.
- Establish best practices for ML security, governance, compliance, testing, and automation.
## Requirements
- 8+ years of experience in MLOps, ML infrastructure, or DevOps for AI/ML systems.
- Strong Python and scripting skills.
- Hands-on experience with MLflow, SageMaker, or Databricks ML.
- Experience with Airflow or AWS Step Functions for workflow orchestration.
- Strong CI/CD and automation experience using GitHub Actions, Terraform, CloudFormation, or similar tools.
- Hands-on experience with Docker and Kubernetes/EKS.
- Strong AWS experience, particularly with SageMaker, Lambda, S3, Kinesis, and Step Functions.
- Understanding of ML model monitoring, observability, security, governance, and compliance.
- Experience working closely with data scientists and ML engineering teams.
## Nice to Have
- Experience with PyTorch, TensorFlow, or Scikit-learn production deployments.
- Experience implementing model drift and data drift detection.
- Experience with LLM production infrastructure.
- Experience building highly automated ML platforms.
- Experience with Infrastructure as Code at scale.
- Knowledge of distributed systems and cloud-native architecture.