freehire launches on Product Hunt on 26 August.

Follow →

Staff AI Platform & Reliability Engineer

Location: Santa Ana, California — on-site
Employment type: Full-time
Salary: $172,000 – $220,000

About us

Created.ai is an AI-powered creative platform combining text, image and video generation into production tools for commercial creative work. The platform is operated by eJam, a multi-brand consumer products and technology company.

We are expanding our engineering team with two appointments covering our AI platform infrastructure and our product engineering surface. Both roles carry substantial technical ownership and report directly into engineering leadership.

This position is AI-native by design. Our engineers work with agentic coding tools as a primary part of their daily workflow, which allows a small team to operate at a scope that would conventionally require a much larger one. We are looking for engineers whose judgement, architectural thinking and review discipline scale that leverage rather than being replaced by it.

Role summary

We are seeking a Staff Engineer to take ownership of the AI generation services, provider integrations, billing integrity, tenant security and overall platform reliability underpinning created.ai. The role is Python and GCP-first, with sufficient TypeScript proficiency to trace and modify cross-service contracts.

This is a senior individual contributor position with architectural authority over the platform layer.

Key responsibilities

  • Own the design, delivery and operation of AI generation services and all third-party provider integrations
  • Ensure billing and usage-metering integrity across the platform, including reconciliation and credit accounting
  • Maintain tenant isolation and platform security controls
  • Lead reliability engineering: failure handling, degradation strategy, capacity and incident response
  • Strengthen deployment safeguards, observability and cost attribution
  • Set technical standards for the Python services and mentor engineers working within them

Requirements

  • Python 3.12 with FastAPI, Pydantic, asyncio and strict typing in production
  • Google Cloud Platform: Cloud Run, Pub/Sub, Cloud Tasks, GCS, Firestore, Cloud SQL/Postgres
  • Demonstrated experience with webhooks, queues, retries, idempotency, dead-letter queues and durable asynchronous jobs
  • SQLAlchemy and Alembic, together with billing or usage-metering experience
  • Infrastructure and delivery tooling: Terraform, IAM/OIDC, Secret Manager, CI/CD
  • Production integrations with image, video or LLM providers
  • Working proficiency in NestJS/TypeScript sufficient to modify cross-service contracts
  • Strong background in observability, cost tracking and incident debugging
  • Fluency with agentic coding tools (Claude Code, Codex, Cursor or comparable), including the ability to scope work for them, review their output critically and maintain architectural coherence across AI-assisted changes

Benefits

  • Health, Dental, Vision
  • 401k Plan
  • PTO Plan
  • 14 observed local holidays
  • Stock options
  • Amazing, pet-friendly office environment
  • Equipment budget and learning allowance
  • Real technical ownership within a small engineering team, and visible impact on a commercially active AI product with real users and real scale constraints
  • Direct involvement in product decisions — the team is small enough that there is nowhere to hide, in both directions

What this application asks

workable

First name, Last name, Email, Phone, Address, Education, Resume

  • Are you able to work full-time, Monday–Friday, on-site at our Santa Ana, CA office? yes / no
  • Are you legally authorized to work in the United States without employer sponsorship? yes / no
  • Do you have a minimum of 7 years building and operating production backend services in Python? yes / no
  • Within the LAST 12 MONTHS, have you personally built and operated production services on Google Cloud Platform (Cloud Run, Pub/Sub, Cloud Tasks, Firestore, Cloud SQL) — not just AWS or Azure? yes / no
  • Name the most recent system where you personally owned billing or usage-metering integrity. Include: what was being metered, the stack, how you handled reconciliation, and rough timeframe. written answer
  • Within the LAST 12 MONTHS, have you personally integrated a production service with a third-party image, video or LLM provider (OpenAI, Anthropic, Replicate, fal, Stability, ElevenLabs, or similar)? Name the provider(s). yes / no
  • What is your target annual base salary? Format: $XXX,XXX (e.g., $195,000).

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available