Principal Software Engineer
About the Role
We are seeking a highly experienced and hands-on Principal Software Engineer to help define and build the next generation of our enterprise Observability and Operational Intelligence Platform. This role combines deep full-stack engineering expertise, platform architecture leadership, product strategy, and modern observability practices.
The ideal candidate will have experience owning product direction, platform architecture, and technology strategy for enterprise-scale engineering solutions. Prior experience building observability, monitoring, operational intelligence, AIOps, or telemetry-driven platforms is highly desirable.
Key Responsibilities
- Own and drive the technical product vision, platform architecture, and technology strategy for a cloud-native observability and intelligence platform.
- Lead architecture, design, and development of scalable, resilient, and highly available enterprise solutions.
- Build modern user experiences using React, TypeScript, and contemporary front-end frameworks.
- Design and develop backend services, APIs, microservices, telemetry pipelines, and intelligence services.
- Define platform standards, technology selections, and engineering best practices aligned with long-term product strategy.
- Design and optimize large-scale telemetry ingestion, aggregation, correlation, and analytics pipelines supporting real-time operational intelligence.
- Drive innovation across Monitoring, Observability, Event Correlation, AIOps, Operational Analytics, Predictive Intelligence, and Self-Healing Automation.
- Evaluate emerging technologies, market trends, and customer needs to continuously evolve product capabilities and roadmap direction.
- Partner with engineering, operations, and business stakeholders to translate enterprise challenges into scalable platform solutions.
- Establish engineering excellence across reliability, performance, automation, security, scalability, and developer experience.
- Mentor engineers, influence technical direction across teams, and serve as a trusted architecture leader.
Qualifications
- 10+ years of software engineering experience with strong expertise in full-stack product development.
- Proven experience owning and delivering technical product direction, platform architecture, and enterprise technology strategy.
- Expert-level proficiency in React, TypeScript, JavaScript, HTML/CSS, and modern UI architectures.
- Strong backend engineering experience using .NET, Java, Node.js, or Python.
- Demonstrated experience designing enterprise-scale architectures, distributed systems, microservices, event-driven architectures, and cloud-native platforms.
- Strong experience with Azure (preferred), AWS, or GCP, Kubernetes, CI/CD, DevOps, Infrastructure as Code, and platform engineering practices.
- Experience designing and optimizing large-scale telemetry ingestion, processing, event correlation, and analytics architectures supporting real-time operational intelligence.
- Deep understanding of observability concepts including Metrics, Logs, Traces, Dashboards, Alerting, SLO/SLI Management, Incident Management, Operational Analytics, and Platform Reliability Engineering.
- Experience with leading observability platforms such as Datadog, Dynatrace, Grafana, Splunk, New Relic, AppDynamics, Elastic, OpenTelemetry, or similar technologies.
- Experience designing AI-powered intelligence layers leveraging:
- Machine Learning and Predictive Analytics
- Agentic AI Workflows
- LLM-based Operational Assistants
- Graph-based Correlation and Dependency Mapping
- Automated Incident Detection and Triage
- Root Cause Analysis and Remediation Recommendations
- Experience translating customer feedback, operational insights, and industry trends into platform capabilities and product roadmap investments.
- Strong product mindset with a focus on customer value, business outcomes, innovation, and engineering excellence.
Preferred Experience
- Experience building enterprise Observability, AIOps, Operational Intelligence, Reliability Engineering, IT Operations, Monitoring, or Platform Engineering products.
- Experience designing systems that process and analyze high-volume telemetry data at scale.
- Experience building intelligent platforms that combine telemetry, AI/ML, automation, and operational workflows.
Exposure to SRE practices, self-healing automation, anomaly detection, capacity prediction, and operational decision intelligence.
Skills
- Agentic AI
- AI
- Analytics
- Anomaly Detection
- API
- Automation
- AWS
- Azure
- CI/CD
- Cloud
- Cloud Native
- CSS
- Datadog
- Developer Experience
- DevOps
- Distributed Systems
- .NET
- Dynatrace
- Event Driven Architecture
- GCP
- Grafana
- HTML
- Infrastructure as Code
- Java
- JavaScript
- Kubernetes
- LLM
- Machine Learning
- Microservices
- New Relic
- Node.js
- Observability
- OpenTelemetry
- Predictive Analytics
- Python
- React
- Splunk
- TypeScript