freehire launches on Product Hunt on 26 August.

Follow →

Sr. SRE Platform Architect

Open 33d

You will lead the design, development, and evolution of Bitdeer's next-generation public cloud platform, owning the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services. You'll be the single point of architectural accountability for the NeoCloud SRE platform, a multi-region GPU rental fleet spanning self-built and OEM-rented data centers, with dozens of bounded contexts, frameworks, and three operational tiers. You'll write and defend the design under review, and shepherd it through the engineering squads that build it. You'll collaborate with cross-functional teams and global partners to define the cloud technology roadmap, optimize multi-region deployments, and deliver world-class infrastructure and platform solutions that power large-scale AI and enterprise workloads.

Responsibilities

  • Write and maintain the platform architecture document, keeping the design coherent across all sections, frameworks, and tiers
  • Review every framework-level change including new bounded contexts, plugin kinds, tier-deployment shifts, schema changes, naming changes, and cross-context contract changes
  • Set design invariants such as residency rules, Tier 2 self-sufficiency budget, survival-uplink contracts, naming conventions, SLO catalogues, and redaction-at-boundary rules
  • Run the plugin framework and author and evolve the uniform contract for every extension
  • Decide tier placement for Edge DC versus Regional Controller versus Global Hub, making data-residency, compliance, and availability tradeoffs explicit
  • Coordinate with cloud-service teams and tenants who author plugins, SDKs, dashboards, and agent recipes riding the platform
  • Coordinate with Security on joint ownership of vulnerability management, exposure management, and joint operations
  • Pre-flight roadmap items by producing one-page designs that fit the existing layered model, tier topology, naming conventions, and extension contracts before implementation
  • Defend the design under review, rejecting scope creep and one-off integrations while approving genuinely needed new plugin kinds

Requirements

  • 10+ years of production SRE, platform-engineering, or infra-architecture experience, including 3+ years at architect level
  • Hands-on experience with GPU / AI-compute infrastructure including NVIDIA GPU ops (DCGM, MIG, vGPU, NVLink/NVSwitch, XID semantics, NCCL), InfiniBand or RoCE fabrics, and HPC storage (Lustre, NetApp/Pure/DDN/VAST, NVMe-oF)
  • Multi-region observability at scale including metrics, logs, traces, profiles, analytics-lake substrate, recording rules, MWMBR burn-rate alerting, and SLI/SLO discipline
  • Experience with cluster platforms including Kubernetes (control plane, GPU Operator, topology-aware scheduling) and at least one of Slurm, Volcano, Kueue, Ray, or KubeRay
  • Data-center operations experience including ZTP, BMC/IPMI/Redfish, BIOS/firmware lifecycle, RMA, and multi-vendor OEM management
  • Strong DDD instincts including bounded contexts, public contracts, and one-context-one-repo discipline
  • Experience designing a plugin framework with a uniform manifest and lifecycle
  • Writing fluency to author and maintain a multi-thousand-line architecture document as well as executive one-pagers
  • Cross-team operating tempo including design reviews, runbook authorship, on-call shadowing, and post-mortem facilitation
  • Hyperscale or NeoCloud experience
  • BS/MS in Computer Science or similar

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available