AI Platform Engineer Senior/Mid
Summary
Build and own the AI-native platform that powers an insurance claims operation, focusing on reliability, observability, and security for Python-based agentic services and Kubernetes infrastructure.
We're building an AI-native claims operation. The models can already do remarkable work — what makes them reliable enough to hand real operational decisions to is the platform around them. That platform is what you'll own. We're an agile, senior team automating a complex, communication-heavy insurance operation, starting with residential property claims. Our agents run live cases today. Your job is to make the infrastructure they run on fast, observable, secure and boringly reliable — and to build it on solid, portable foundations rather than locking us into any one vendor.
Your tasks
The Python services that orchestrate our agentic pipeline (eg. FastAPI, LangChain), and the platform they run on
Our production runtime — deployment, scaling, resilience, and keeping services healthy under real load
Release management — a proper branching and release strategy in Git, environment promotion, semantic versioning, and safe rollouts: canary/progressive delivery, feature flags, and clean rollback when something's wrong
CI/CD pipelines and infrastructure-as-code — fast, safe, repeatable delivery from commit to production
The async processing backbone — queued, scheduled and event-driven claim handling
The data layer — choosing and running the right persistence and building it to scale, not inheriting whatever's easiest
Observability and ops: tracing, logging, metrics, alerting
Security and data protection — how we handle and isolate sensitive claim data in a GDPR-heavy domain
Your profile
Have 5+ years in backend/platform/DevOps engineering, and think in systems, not just features
Have real, hands-on experience building and operating production services on Kubernetes
Have own a release process end to end — you know how to structure Git for a team, promote code across environments safely, and ship to production frequently
Are strong in Python and comfortable owning cloud infrastructure end to end — cloud-agnostic by instinct, wary of lock-in, fluent in the primitives rather than any one provider's product catalogue
Live in Docker, CI/CD
Have solid database depth — you've diagnosed a slow query under load and fixed the root cause, and you can pick and run the right datastore
Treat reliability, observability and security as first-class concerns
Are comfortable being the platform owner in a small team — high autonomy, high ownership
Nice to have
Experience with LLM-serving infrastructure, queues, or agentic/async workloads
GitOps / progressive-delivery tooling (Argo, Flux, or similar)
Java experience alongside Python
Insurance, fintech, or another regulated domain; German (helpful, not required — we work in English)