Lead Machine Learning Engineer
Maya- Posted 9 hours ago
- Be among the first 10 applicants
Job Description
Core Profile
We are looking for a Lead Machine Learning Engineer to spearhead the machine learning engineering function for Maya's Marketing, Growth, and Personalization teams. In this role, you will lead, mentor, and elevate a team of machine learning engineers while driving close collaboration with data science, product, and cross-functional engineering partners.
You operate as a technical authority who designs highly resilient, large-scale distributed ML systems—including next-generation agentic AI solutions and autonomous workflows—and masters compute optimization, balancing cloud economics, hardware constraints, and latency-throughput trade-offs. You will own complex mesh-facing service development, translate deep infrastructure trade-offs into business ROI, and ensure our marketing intelligence platforms scale efficiently and reliably.
Nature of Work & Key Responsibilities
- Team Leadership & Engineering Growth: Lead and mentor Machine Learning Engineers across the Marketing, Growth, and Personalization domains. Guide staff and senior engineers on architectural blind spots, systemic debugging, and advanced software engineering paradigms.
- Agentic Solutions & Autonomous Workflows: Architect, deploy, and scale production-grade agentic AI systems(multi-agent frameworks, tool-using LLMs, autonomous decision-making loops, and goal-driven orchestration engines) to power hyper-personalized, real-time marketing journeys and automated growth campaigns.
- Cross-Functional Collaboration & Stakeholder Alignment: Partner closely with Product Owners, Data Scientists, and external engineering teams. Act as the bridge who translates complex technical infrastructure and architectural trade-offs to business leaders, aligning engineering costs directly with business ROI.
- Distributed ML Systems & Resilient Architecture: Architect and design highly resilient, large-scale distributed machine learning systems that anticipate, isolate, and mitigate catastrophic failure modes. Oversee complete system design and integration decisions.
- Mesh-Facing Service Development: Take complete ownership of mesh-facing service development to seamlessly integrate ML capabilities (such as real-time personalization, propensity scoring, and recommendation feeds) across disparate legacy and modern enterprise systems.
- Compute Optimization & Cloud Economics: Master deep compute optimization, managing cloud cost models, token economics for LLMs/agents, hardware-level constraints (GPUs/TPUs), and precise latency/throughput trade-offs. Make definitive strategic calls on leveraging managed cloud AI services versus building bespoke, in-house infrastructure.
- Full-Lifecycle MLOps & Production Standards: Set the standard for rigorous code reviews (platform, tool, utility, and model-level), enforce team coding standards, drive automated model monitoring deployment, and manage production deployment cycles from staging to high-availability production environments.
Displayed Skill Mastery & Foundational Competencies
Advanced Leadership & Architectural Mastery (Lead Level)
- Resilient Distributed Systems: Expert in designing fault-tolerant, large-scale distributed ML systems capable of handling high-volume traffic without degradation.
- Mesh & Service Integration: Extensive experience building microservices, RESTful APIs (FastAPI/Flask), and mesh-facing services that expose model outputs reliably.
- Cost & Hardware Optimization: Deep expertise in cloud economics, resource optimization for model training and inferencing, and hardware constraint tuning (GPUs/TPUs).
- Strategic Technology Judgment: Definitive ownership of end-to-end MLOps solutions, deciding when to adopt managed cloud AI services versus building custom components.
- Agentic & Generative AI Architecture: Deep expertise in building production-grade agentic workflows, multi-agent coordination frameworks (e.g., LangGraph, AutoGen, CrewAI), vector databases, retrieval-augmented generation (RAG), and tool-augmented LLM orchestration.
Comprehensive Technical Foundation (Levels 1–4 Execution)
- Model Lifecycle & MLOps: Mastery of the complete model development lifecycle—from translating experimental data science models into production code, handling model staging/deployment, to implementing automated health checks, monitoring, and retraining triggers.
- Code Quality & Review Standards: Proficiency in reviewing complex DS and MLE code, establishing team-wide coding standards, and leveraging AI-assisted tools for code refactoring and baseline documentation.
- Core Stack Proficiency: Strong command of Python, SQL, PySpark/Spark, Bash, Docker, Kubernetes, Terraform/CloudFormation, and CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins).
- Frameworks & Infrastructure: Familiarity with modern data and ML frameworks (Pandas, Scikit-learn, TensorFlow, PyTorch, MLflow) and cloud ecosystems (AWS, Databricks).
Expected Results
- High-Performing Engineering Team: A mentored, highly productive team of ML engineers adhering to rigorous software engineering and MLOps standards.
- Resilient & Scalable Platforms: Mission-critical marketing, growth, and personalization ML systems that maintain high availability, low latency, and robust fault tolerance.
- Optimized Infrastructure Spend: A tightly managed balance between cloud computing costs (LLM Tokens, GPUs/TPUs, infrastructure overhead) and actual business value/ROI delivered.
- Seamless System Integration: Flawless mesh-facing service development that integrates advanced ML models cleanly into Maya's broader technical ecosystem.
Required Qualifications
- Education: Bachelor's or Master's degree in Computer Science, Computer Engineering, Information Technology, or a quantitative discipline.
- Experience: 6+ years of hands-on experience building, deploying, and operating data pipelines and production machine learning systems.
- Leadership Track Record: 3+ years of experience leading technical initiatives, projects, or engineering teams, with proven success mentoring engineers and aligning cross-functional stakeholders.
- Infrastructure & MLOps Depth: 3+ years of experience designing CI/CD pipelines, distributed systems, microservices, and cloud-native infrastructure (AWS, GCP, or Azure).
- Agentic & GenAI Expertise: Demonstrated experience architecting and deploying LLM-based applications, multi-agent frameworks, vector databases, or automated decision-making workflows into production.
- Domain Context: Experience supporting AI/ML use cases in fintech, digital banking, e-commerce, or high-growth technology environments.
More Info
Key Skills
Scikit-learn
MLflow
GitLab CI
GitHub Actions
Cloud Economics
