Senior Software Engineer, AI Operations
Summary
Oversees AI system stability, performance, and maintenance for public-sector clients, ensuring operational excellence by managing SLAs, model drift, and incident governance while bridging engineering and client governance.
As a Senior Software Engineer AI Ops at Scale AI you will own the long-term technical health performance and stability of AI solutions deployed across our strategic public sector partners While our Delivery Teams build and launch new use cases you are the technical steward ensuring these deployments operate with Operational Excellence You will bridge software engineering MLOps and client governance managing tiered SLAs tracking model drift and executing maintenance protocols that protect both system integrity and operational margins
Key Responsibilities- Handover Gate amp Onboarding Act as the technical gatekeeper during the formal transition from Delivery to Maintenance Conduct deep-dive reviews to ensure baseline code prompts and architecture meet strict maintainability and documentation standards before sign-off
- Tiered SLA amp Incident Management Own technical response and resolution targets across multi-tiered service models from Business-Hours Essential to 24 7 Mission-Critical
- Lead Incident Governance
- Root Cause Analysis RCA and P1 P2 mitigations within strict active support windows
- AI Lifecycle Governance Monitor production model performance latency and data drift Manage prompt configuration repositories to maintain behavioral consistency and perform regression testing when LLM providers update underlying endpoints
- Request Classification amp Technical Scope Operationalize the boundary between Routine Maintenance In-Scope and System Evolution Out-of-Scope Assess incoming client requests and run comparative benchmarking on new AI models
- Automation amp Reliability Engineering Eliminate operational toil by engineering self-healing data pipelines automated RAG indexing syncs and telemetry tooling Influence upstream Delivery teams to adopt architectural patterns that simplify ongoing maintenance
- Client Technical Interface Serve as the senior technical point of contact for government and enterprise IT leads Translate technical AI concepts data drift prompt versioning API deprecation into clear business impacts for non-technical stakeholders
Ideally you'd have Background: 5+ years in Software Engineering, MLOps, SRE, or Forward Deployed Engineering in heavy data or production AI environments.
Technical Stack:- Advanced proficiency in Python, SQL, REST/gRPC APIs, and cloud architecture (AWS, Azure, or GCP).
- Hands-on experience with MLOps tooling, vector databases, and LLM orchestration frameworks (e.g., LangChain, LlamaIndex).
Practical understanding of prompt version control, model benchmarking against evaluation datasets, RAG pipeline mechanics, and data drift detection.
Engineering Mindset:A drive to build systematic, automated fixes rather than applying temporary patches.
Strong grasp of CI/CD for machine learning pipelines.
Client Acumen & Boundary Control:Strong technical communication skills with the ability to manage client expectations, defend operational boundaries (Maintenance vs. Evolution), and advise on long-term system roadmaps.