freehire launches on Product Hunt on 26 August.

Follow →

Platform Engineer - AI Agent Infrastructure

Open 20d

Summary

Build and operate a secure, multi-tenant platform that runs production AI agents, handling runtimes, security, data governance, and cloud infrastructure for business workflows.

How to apply

Email hiring@sunnystep.com with subject: [Position]-[Your Full Name]. Include your CV, GitHub/portfolio, and 2-3 examples of production systems you built or shipped. Applications that do not follow this format may not be reviewed.

About the role

We are building Maxify, an AI-native operating platform where agents execute real business workflows through company systems while remaining accountable to source data, permissions, evaluations and measurable outcomes.

We are looking for a hands-on Platform Engineer to build the reusable, multi-tenant infrastructure that makes production AI agents secure, reliable and economical to operate. You will own the shared platform beneath business-specific agents while still shipping practical capabilities end-to-end.

This is not a pure infrastructure, prompt-engineering or research role. You will work across agent runtimes, backend services, data, integrations, deployment and developer tooling.

What you will own

  • Own the multi-tenant agent runtime, orchestration, state and governed memory.
  • Build secure tool execution, authentication, authorization and customer isolation.
  • Design approvals, audit trails, privacy controls and secrets management.
  • Create reusable connectors for APIs, databases and business systems.
  • Build evaluation infrastructure, testing environments and safe release gates.
  • Implement observability for quality, failures, latency, reliability and cost.
  • Establish deployment, versioning, backup, rollback and incident-response mechanisms.
  • Build platform APIs, operator interfaces and developer tooling.
  • Own cloud architecture, infrastructure automation, capacity, availability and disaster recovery.
  • Turn repeated workflow patterns into legible, reusable platform capabilities.

What we are looking for

  • Production experience building and operating multi-tenant backend or platform systems.
  • Strong backend and distributed-systems engineering judgment.
  • Production proficiency in Python and/or TypeScript.
  • Strong PostgreSQL and cloud-hosted data-infrastructure experience.
  • Experience with APIs, queues, asynchronous jobs and event-driven systems.
  • Strong understanding of authentication, authorization, secrets management and data isolation.
  • Experience with CI/CD, infrastructure automation, monitoring and incident response.
  • Experience operating production workloads on AWS, GCP or Azure.
  • Practical knowledge of containers, networking, backups, rollback and disaster recovery.
  • Practical LLM experience including tool use, structured outputs, context management and evaluation.
  • Ability to debug across application, infrastructure and external-service layers.
  • Clear written communication and sound architectural judgment.

Strong signals

  • You have shipped and operated a production agentic or automation platform, not only a demo.
  • You have built multi-tenant systems with explicit isolation and authorization controls.
  • You have designed evaluation, observability, approval or audit systems for AI workflows.
  • You have owned reliability and cloud infrastructure while remaining product-oriented.
  • You use AI development tools extensively while validating outputs and retaining technical ownership.
  • You can explain a system you built end-to-end, what failed and how you prevented recurrence.

Nice to have

  • Experience with OpenAI Agents SDK, LangGraph, OpenClaw or comparable systems.
  • Experience with Supabase or PostgreSQL at production scale.
  • Experience with Lark/Feishu, Shopify, CRM, accounting or enterprise SaaS integrations.
  • Familiarity with knowledge graphs, ontology design or governed memory.
  • Early-stage startup or founding-engineer experience.

Current stack

OpenClaw, Python, TypeScript, Supabase/PostgreSQL, Lark/Feishu APIs and CLI, Shopify and other business-system APIs, Git-based development, automated testing, deployment controls, monitoring and rollback.

What this role is not

  • Prompt writing without software ownership.
  • Pure machine-learning research or model training.
  • Pure DevOps, Kubernetes or cloud administration.
  • Frontend-only or backend-only feature delivery.
  • Building abstractions for hypothetical scale before real workflows require them.

Success in the first 30 days

  • Complete an architecture, security and reliability review and identify the major risks and delivery bottlenecks.
  • Operate the current agent runtime and deployment environment independently.
  • Ship at least one meaningful production platform improvement.
  • Deliver one reusable platform capability with tests, telemetry and rollback.
  • Establish minimum quality, security and observability gates.
  • Document the reusable pattern for onboarding the next customer or workflow.
  • Produce a prioritized 60-day platform roadmap based on evidence from the first month.

These outcomes are subject to timely access and platform readiness.

How we work

  • We move quickly and build for permanence.
  • We prefer small, reversible decisions over speculative architecture.
  • Documentation, testing, evaluation and observability are part of the product.
  • AI accelerates the work; it does not remove engineering accountability.
  • We measure success through customer outcomes, revenue impact, cost reduction and lower operational attention.

Interview process

Shortlisted candidates will complete an onsite prototyping test based on a practical platform problem. We evaluate working output, engineering judgment, effective use of AI tools, debugging ability and clarity of explanation. Test spec: Live Prototyping Test v2 — Zero-Handoff (Platform Engineer & Software Engineer)

See also