Search by job, company or skills

Lead GenAI Engineer

Lead GenAI Engineer

Digital Impetus
Fresher
Not Disclosed
  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

Roles & Responsibilities

We are looking for an AI Platform Senior Lead to join our engineering team and take ownership of the infrastructure, pipelines, and platform capabilities that power our AI and LLM-based solutions. This is a hands-on engineering role for someone who thrives at the intersection of cloud infrastructure, data engineering, and AI systems — not someone who builds models, but someone who builds the platforms that make AI systems reliable, observable, and scalable in production.

You will work closely with application teams, data scientists, and product stakeholders to ensure AI workloads run efficiently, cost-effectively, and at scale.

What You Will Do

AI Platform & Infrastructure

  • Design, build, and maintain scalable platforms that support LLM inference, embedding pipelines, RAG systems, guardrails, and multi-tenant AI services
  • Architect and manage APIs and microservices that abstract AI capabilities for consumption by application teams
  • Own platform reliability — uptime, latency SLAs, cost per request, and performance benchmarks
  • Build and maintain multi-tenant infrastructure with strong isolation, quota management, and usage tracking across teams and clients

Data Engineering & Processing

  • Design and implement data ingestion, transformation, and processing pipelines that feed AI systems
  • Build and manage vector databases, document stores, and retrieval infrastructure for RAG and semantic search use cases
  • Ensure data quality, lineage, and governance across AI data pipelines
  • Optimize data pipelines for throughput, latency, and cost at scale

Cloud & DevOps

  • Own cloud infrastructure on AWS — including ECS/EKS, Lambda, API Gateway, S3, RDS, SQS, and AI/ML services like Bedrock and SageMaker
  • Implement Infrastructure as Code using Terraform or CDK
  • Build CI/CD pipelines for AI workloads including model serving, evaluation, and deployment automation
  • Manage Kubernetes clusters and containerised workloads for AI services

Observability & Operations

  • Implement comprehensive observability for AI systems — logging, tracing, metrics, cost dashboards, and alerting
  • Build evaluation pipelines to monitor LLM output quality, hallucination rates, latency, and token consumption in production
  • Proactively identify and resolve performance bottlenecks, reliability gaps, and cost inefficiencies

Collaboration & Technical Leadership

  • Partner with product, architecture, and application teams to translate requirements into scalable platform capabilities
  • Define and enforce platform engineering standards, patterns, and best practices
  • Mentor junior engineers and contribute to technical design reviews

More Info

Job Type:
Industry:
Employment Type:

Key Skills

RAG systems

vector databases

alerting

EKS

SageMaker

document stores

data ingestion

observability

About Company

Similar Jobs

Bengaluru, India
Skills:
amazon dynamodb , Aws Lambda, Amazon S3, Typescript, Python, Agentic AI architectures, AIDLC, AWS Step Functions, AgentCore, Knowledge Bases, Amazon API Gateway, AWS CDK, Foundation Models, Serverless AWS, GitHub Actions, RAG, Strands Agents SDK, AWS Storage, Amazon Bedrock