freehire launches on Product Hunt on 26 August.

Follow →

Senior Platform & System Engineer

Summary

Senior engineer owns hybrid cloud-edge infrastructure for AI-powered QSR operations, including on-prem edge servers, observability stacks, and automation to keep thousands of stores running smoothly.

Berry AI builds AI-powered operations platforms for QSR restaurants — drive-thru analytics, loss prevention, and store management tooling deployed at thousands of locations across the US — and growing. Everything we ship runs on hybrid cloud-edge infrastructure — in-store edge servers and cloud services — and the system team keeps that infrastructure healthy, observable, and efficient to run as it scales. We're hiring a Senior Platform & System Engineer to go deep on that infrastructure as a hands-on individual contributor.

What you'll work on

  • Own the on-prem edge fleet — thousands of in-store servers and cameras across customer networks, reached over resilient VPN/mesh paths (OpenVPN, Tailscale).

  • Build the internal tools, scripts, and APIs that let the CS team and on-site technicians install and troubleshoot store servers, networks, and cameras on their own — turning cross-team escalations into self-serve fixes.

  • Own our self-hosted observability platform — a Prometheus-ecosystem stack (Grafana, Mimir, Loki, VMAgent, Vector) across edge and cloud, with SLO dashboards and alerting that feeds automated ticketing and remediation.

  • Drive the automation that keeps operational load flat as the fleet grows — event-driven auto-ticketing and auto-remediation across Ansible/AWX and serverless AWS (SNS/SQS/Lambda).

  • Run the core infrastructure and services at the Taipei headquarters — servers and office network, virtualization, Kubernetes, storage, device monitoring, and self-hosted services (container registry, auth, reverse proxy).

  • Shape the roadmap so systems work enables product velocity — partnering with ML engineers, product managers, customer success, and customer-side IT.

  • Take your turn in the on-call rotation — own incident response on your shifts, and turn recurring issues into durable fixes and runbooks.

  • Own security and compliance operations — the ISO 27001 ISMS, vulnerability management, code security scanning, and disaster-recovery drills.

You're a strong fit if you have

  • Skilled in systems, infrastructure, or DevOps engineering, with strong Linux (Ubuntu) administration and real experience operating fleets at scale.

  • Strong networking fundamentals — TCP/IP, DNS, VLANs, VPN, and firewalls.

  • Deep infrastructure-as-code, config-management, and orchestration experience across on-prem and cloud — e.g. Terraform, Pulumi.

  • Observability fluency — running metrics, logs, and alerting at scale, defining SLOs, and balancing monitoring detail against the storage and cost it drives.

  • Solid automation and scripting skills — Bash and Python especially, and ideally some Go — you reach for code to remove repetitive operational work.

  • Excellent communication — clear docs and runbooks, a calm head in incidents, and coordination across product, customer success, customers, and third-party IT.

  • Fluent in Mandarin and English — Mandarin is the team's working language; fluent English is required for daily work with US teams, customers, and IT vendors.

  • A pragmatic sense of ownership — balancing reliability against delivery velocity, and knowing which problems are worth solving now.

Bonus points

  • Familiarity with IP cameras/NVRs and the ONVIF protocol, plus real-time video streaming (RTSP, H.264/H.265, MediaMTX).

  • Storage and NAS operations at scale (ZFS/TrueNAS, Synology, QNAP, RAID/HA).

  • Hands-on vulnerability management and SAST — e.g. OpenVAS, SonarQube.

  • Workflow-automation tooling (n8n or similar) and a track record of building low-toil operational systems.

Our engineering culture

Small team, high ownership, fast feedback from customers — and the operational rigor to make that velocity sustainable. Modern AI tooling — LLMs, coding agents, agent-driven workflows — is a normal part of how we work, and you're encouraged to push on what these tools can do.

============================

Interview Process

  1. Online (Google Meet)

    1. Team Lead Interview (0.5 - 1 hr)

  2. Onsite

    1. Technical Interview (1.5 hrs)

    2. CEO & VP Interview (1.5 hrs)

    3. HR Interview (0.5 hr)

What this application asks

ashby

Name, Email, Resume

  • Phone optional

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available