Principal Platform Software Engineer - JOIN OCI NASHVILLE!
PRINCIPAL PLATFORM SOFTWARE ENGINEER (IC4)
OCI Log Analytics - Query Language and Distributed Execution
Location: Nashville, Tennessee
Work Arrangement: On-site Monday through Friday
On-Call Requirement: Participation in the team’s on-call rotation is required
Relocation: Relocation assistance is available for qualified candidates
POSITION OVERVIEW
Oracle Cloud Infrastructure (OCI) Log Analytics is a cloud-scale observability platform that enables customers to search, analyze, and derive actionable insights from massive volumes of machine-generated data. Its query language supports advanced capabilities, including clustering, anomaly detection, natural language processing, statistical analysis, and time-series analysis across terabytes of log data.
We are seeking a Principal Platform Software Engineer to help shape the architecture, query language, and distributed execution platform behind these capabilities. In this role, you will solve complex engineering problems involving query processing, distributed systems, machine learning, and generative AI.
You will also have opportunities to develop AI agents and agentic workflows that make large-scale log investigation and operational troubleshooting faster, more intuitive, and more effective.
This is a production-facing engineering role based in Nashville, Tennessee. The selected candidate must work on-site Monday through Friday and participate in the team’s on-call rotation. Candidates who are interested in relocating to Nashville are encouraged to apply. Relocation assistance is available for qualified candidates.
RESPONSIBILITIES
• Lead the architecture, design, and development of OCI Log Analytics query-language capabilities, including parsing, planning, optimization, and distributed execution.
• Build scalable systems that execute advanced analytics and machine-learning algorithms across terabytes of data while meeting demanding latency, availability, and cost-efficiency requirements.
• Design and develop AI agents and agentic workflows that translate user intent into effective queries, investigate complex operational issues, and surface actionable insights.
• Architect secure, resilient, multi-tenant cloud services with appropriate fault tolerance, replication, throttling, load shedding, and automated recovery.
• Identify and eliminate performance bottlenecks across query execution, indexing, storage, retrieval, and analytics pipelines.
• Define performance, scalability, correctness, availability, and reliability requirements.
• Develop load, stress, fault-injection, recovery, and performance tests that validate system behavior at scale.
• Establish service-level objectives, telemetry, dashboards, and alerts that support operational excellence.
• Lead the investigation and root-cause analysis of complex production issues and drive corrective and preventive improvements.
• Participate in the team’s on-call rotation, respond to production incidents, restore service health, and help maintain the reliability and availability of OCI services.
• Collaborate with product managers, architects, software engineers, site reliability engineers, security teams, and adjacent OCI Observability organizations.
• Provide technical leadership, mentor engineers, promote engineering best practices, and contribute to hiring and talent development.
MINIMUM QUALIFICATIONS
• Bachelor’s, master’s, or doctoral degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
• Six or more years of software engineering experience, including experience designing, building, and operating large-scale distributed systems.
• Strong programming expertise in Java, Python, or both.
• Experience building cloud-scale, multi-tenant products or services that support high-throughput, concurrent processing.
• Strong understanding of distributed computing, large-scale data processing, system reliability, performance optimization, and the complete service development and operations lifecycle.
• Demonstrated ability to design systems for scalability, elasticity, availability, durability, security, and cost efficiency.
• Familiarity with modern AI-assisted software development and a practical understanding of large language models, AI agents, tool use, and interoperability protocols such as Model Context Protocol (MCP) and Agent-to-Agent (A2A).
• Strong analytical, problem-solving, collaboration, and written and verbal communication skills.
• Demonstrated ability to provide technical leadership and deliver results in a fast-moving environment with evolving requirements.
• Ability and willingness to participate in the team’s on-call rotation and respond to production incidents.
• Ability to work on-site in Nashville, Tennessee, Monday through Friday.
PREFERRED QUALIFICATIONS
• Experience building observability, logging, search, analytics, or query-processing platforms.
• Experience scaling machine-learning, statistical, or analytical algorithms across very large datasets.
• Experience with technologies such as Kafka, Lucene, Solr, Spark, Parquet, Kubernetes, or Terraform.
• Experience designing or building query languages, compilers, query planners, distributed execution engines, indexing systems, or information-retrieval platforms.
• Proven success improving the performance, scalability, reliability, resiliency, and recovery characteristics of large-scale production systems.
• Experience developing AI agents, generative AI applications, retrieval systems, or natural-language interfaces for technical products.
LOCATION, RELOCATION, AND ON-CALL EXPECTATIONS
This position requires working on-site at Oracle’s Nashville, Tennessee, location Monday through Friday. This is not a remote or hybrid position.
Candidates who do not currently live in the Nashville area must be willing to relocate. Relocation assistance is available for qualified candidates.
Participation in the team’s on-call rotation is required. The engineer will respond to production incidents, troubleshoot complex service issues, restore service health, and implement improvements that reduce the likelihood and impact of future incidents.
Disclaimer:Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $104,500 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC4