Sr. Principal Software Engineer
Summary
Leads the design and development of a distributed edge compute platform, focusing on hyperscaler cloud technologies (AWS/Azure/GCP), Kubernetes, and high-throughput L7 traffic routing to ensure scalability, fault tolerance, and API-driven interactions with customers and partners.
- Design and implement a massively distributed edge compute platform, ensuring high availability and fault tolerance.
- Develop external-facing APIs, considering design, versioning, and interactions with customers and partners.
- Optimize high-throughput L7 traffic routing for efficient data processing.
- Utilize Kubernetes internals for resource management and container deployment.
- Collaborate with a team of engineers to build and operate the platform, sharing knowledge and best practices.
- Translate distributed systems complexity into clear architectural tradeoffs
- Communicate infrastructure risks and mitigation strategies to executive stakeholders
- Manage technical requirements from customers’ needs.
- Collaborate effectively with AI/ML systems engineers, Hardware architects, Network engineering teams, Security and compliance teams
- Present scalability models and reliability metrics in measurable, defensible terms
- Mentor engineering teams on distributed systems best practices
- Anticipate internal and external business challenges and / or regulatory issues and drive process, product or service improvements that create competitive advantage.
- Stay updated with the latest trends and technologies in distributed systems and edge computing.
- Establish design reviews and architectural governance standards
- Ensure the platform's security and data privacy, adhering to industry standards and regulations.
- Conduct code reviews and provide constructive feedback to maintain code quality.
- Extensive experience (5+ years) with hyperscaler technologies (AWS, Azure, GCP) and managed services.
Have worked at or intimately with AWS, Azure, GCP, or similar-scale platforms - built internal cloud, compute, or infrastructure services
- Proven track record in designing and operating large-scale distributed systems.
- Strong fundamentals in distributed systems, including consistency models and fault tolerance.
- Proficiency in API design and development, with an understanding of versioning and SDK foundations.
- Experience with Kubernetes and containerization technologies for resource management.
- Solid understanding of high-throughput traffic routing and network protocols.
- Ability to work independently and manage complex technical tasks.
- Excellent communication and collaboration skills, with a willingness to share knowledge.
- Familiarity with security best practices and data privacy regulations.
- Master's degree in Computer Science, Engineering, or a related field. PhD preferred.