Senior Backend Engineer — Distributed Systems
Summary
Senior backend engineer who designs, builds, and operates high-throughput distributed backend services and APIs on Kubernetes for a client's large-scale AI model-training data platform. Day-to-day work centers on microservices, Kubernetes/Docker, CI/CD with GitOps, performance, reliability, and observability, working U.S. EST hours.
This role is open to candidates based in LATAM, Africa, and Eastern Europe. Please note that as this role supports U.S.-based clients, candidates must be available to work during U.S. business hours aligned with the client’s time zone.
Our client is a well-funded, early-stage AI company building a large-scale data platform used to train, fine-tune, and evaluate the next generation of AI models. As the platform continues to scale, they are seeking a deeply technical backend engineering specialist to build reliable, high-throughput distributed systems.
Role Overview
The Senior Backend Engineer — Distributed Systems will design, build, and operate the backend services that power a large-scale, high-throughput data platform.
This is a specialist engineering role focused on backend development and distributed systems rather than generalist software development or people management. The Senior Backend Engineer will work directly with engineering leadership on system architecture, reliability, scalability, observability, containerized production environments, and deployment infrastructure.
The ideal candidate brings exceptional computer science fundamentals, deep distributed systems expertise, and significant backend engineering experience. This person will also contribute to the broader engineering team by sharing technical context, supporting less-experienced colleagues, and delegating work effectively.
Location
Fully Remote | 9:00 AM - 5:00 PM EST
Key Responsibilities
Backend Services & APIs
Design, develop, and maintain scalable backend services and APIs.
Build backend systems using a microservices architecture.
Support the backend infrastructure required for a large-scale, high-throughput data platform.
Distributed Systems Architecture
Apply distributed systems principles to the design and operation of production systems.
Design systems with consideration for consistency, partitioning, replication, consensus, and failure handling.
Build fault-tolerant, highly available, and scalable distributed systems.
Kubernetes & Container Orchestration
Build and operate production workloads on Kubernetes.
Manage deployments, services, ingress, autoscaling, and Helm.
Work with managed Kubernetes environments such as EKS.
Support containerized production environments using Docker.
Performance & Reliability
Optimize application performance across distributed and containerized environments.
Improve the reliability of backend systems and services.
Identify opportunities to strengthen system scalability and availability.
CI/CD & Deployment
Design and implement robust CI/CD pipelines.
Support GitOps-based deployment practices.
Implement Kubernetes-native deployment workflows.
Observability
Improve observability across distributed backend services.
Support effective system monitoring and alerting.
Strengthen visibility into the health and performance of production systems.
Architecture & Technical Collaboration
Partner directly with engineering leadership on architectural decisions.
Share technical knowledge and context with less-experienced colleagues.
Delegate work appropriately to support effective team execution.
Contribute collaboratively to technical decisions and overall engineering quality.
Responsible AI Use
Use AI tools as part of daily engineering workflows to improve productivity.
Select appropriate AI models for planning and architectural work.
Orchestrate AI-assisted execution efficiently while critically reviewing outputs.
Follow established AI governance practices.
Ensure access flows through approved API endpoints and MCPs rather than direct database access or other shortcuts.
Qualifications
Experience
10+ years of professional software engineering experience with a strong backend focus.
Deep hands-on experience designing and working with distributed systems.
3+ years of hands-on Kubernetes experience running production workloads.
Experience working with managed Kubernetes environments such as EKS, GKE, or AKS.
Experience designing and developing microservices architectures.
Experience working with relational and/or NoSQL databases.
Skills
Exceptional computer science fundamentals with a deeply architectural and systems-minded approach.
Strong understanding of distributed systems concepts, including consistency models, partitioning, replication, consensus, CAP theorem, and failure handling.
Proficiency in at least one modern backend programming language, with Go or Python strongly preferred and C++ or Rust also applicable.
Strong hands-on Kubernetes capabilities across deployments, services, ingress, autoscaling, Helm, and production workload management.
Experience working with Docker and containerized applications.
Strong understanding of microservices architecture and inter-service communication patterns.
Familiarity with infrastructure-as-code, with Terraform preferred.
Ability to use AI tools effectively in daily engineering work while critically reviewing their output and following defined governance requirements.
Exposure to Elasticsearch or other search technologies is a plus.
Experience with GitOps tools such as Argo CD or Flux is a plus.
Knowledge of service mesh technologies such as Istio or Linkerd is a plus.
Experience with workflow orchestration tools such as Temporal or Airflow is a plus.
Experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, or ELK is a plus.
Experience building or operating search-heavy systems is a plus.
Contributions to open-source projects related to Kubernetes or distributed systems are a plus.
What Success Looks Like
Scalable backend services and APIs reliably support the company's high-throughput data platform.
Distributed systems are designed with strong fault tolerance, availability, and scalability.
Kubernetes production workloads operate reliably and efficiently.
Application performance and reliability continuously improve across distributed and containerized environments.
CI/CD and GitOps processes support reliable Kubernetes-native deployments.
Monitoring, alerting, and observability provide clear visibility into distributed production systems.
Architectural decisions reflect strong computer science and distributed systems fundamentals.
Technical knowledge and context are effectively shared across the engineering team.
AI tools increase engineering productivity while remaining within established governance requirements.
Opportunity
This role offers the opportunity to apply deep backend and distributed systems expertise to a large-scale data platform supporting the training, fine-tuning, and evaluation of AI models. The Senior Backend Engineer — Distributed Systems will work directly with engineering leadership on architecture, scalability, reliability, Kubernetes, deployment infrastructure, and observability while contributing technical knowledge across a collaborative engineering team.
Application Process:
To be considered for this role these steps need to be followed:
Fill in the application form
Record a video showcasing your skill sets
Skills
As published by ashby · 17 questions · 6 written answers
Basics
Full Name:, Email:, Resume:, What country will you be working from?
Short answers (3)
- Phone Number:
- What's your expected monthly compensation in USD for a 40h/week (full-time) role?
- What's your expected monthly compensation in USD for a 20h/week (part-time) role? optional
Pick from a list (8)
- What is your level of English proficiency?
- Are you available to work 9 AM - 5 PM EST?
- Do you have any friends, family members, or personal connections currently working at Scale Army or involved in any of our projects?
- How many years of professional software engineering experience do you have with a strong focus on backend development?
- Which of the following technologies have you worked with hands-on?
- Did you submit your video using the same email you are entering in this application form?
- I’d like to stay in touch! Please send me job opportunities, updates, and relevant news from Scale Army Careers. I can unsubscribe anytime.
- Are you comfortable with us using your video for marketing purposes (ex: social, web, email, etc.) to help more hiring managers see your application?
Written answers (6)
- Were you referred by a current Scale Army contractor? If yes, please add their full name. optional
- Please share your GitHub profile or links to public repositories that demonstrate your backend engineering, distributed systems, Kubernetes, or infrastructure experience. optional
- Describe your hands-on experience designing and operating distributed systems. Please include examples of how you have addressed concepts such as consistency, partitioning, replication, consensus, high availability, or failure handling in production.
- Describe your experience running production workloads on Kubernetes. What types of systems have you operated, and how have you worked with areas such as deployments, services, ingress, autoscaling, Helm, and managed Kubernetes environments?
- Tell us about your experience designing, building, and scaling backend systems using Kubernetes. What types of systems have you worked on, and what were your main responsibilities?
- How do you use AI in your work?