Principal Software Engineer, AI & Data Platform
Summary
Architect and build scalable AI and data pipelines for scientific applications, focusing on clean, reusable datasets and end-to-end model training workflows using Python and cloud technologies.
Responsibilities
- Architect the data foundation for large scale scientific and engineering output, keeping results clean, queryable, reusable, and ready for model training.
- Model domain specific scientific data so the same datasets can support interactive analysis, automation, and downstream machine learning workflows.
- Build scalable data processing patterns across object storage, analytical stores, and training optimized formats.
- Create machine learning data pipelines for curation, deduplication, formatting, evaluation sets, and regression tracking.
- Build and operate training and fine tuning pipelines for models used in scientific and workflow driven products.
- Develop intelligent workflow interfaces that connect user intent, structured platform capabilities and executable workflows without exposing unnecessary complexity to users.
- Own model evaluation, benchmarking, automated scoring, and quality tracking so each iteration is measurable.
- Set data and AI engineering standards for the team and turn them into code, documentation, and reusable patterns.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related engineering field, with 10 plus years building and shipping production software.
- Expert Python and a strong record of shipping systems end to end.
- Deep experience with large scale data systems, including object storage, analytical processing, training optimized formats, and production data pipelines.
- Hands on experience building data pipelines for model training, fine tuning, evaluation, and continuous improvement.
- Direct experience training or fine tuning models for structured outputs, tool use, workflow automation, or domain specific applications.
- Strong understanding of relational, document, and columnar data models, with judgment about where each belongs.
- Comfort operating in cloud, enterprise, and technical compute environments, including distributed training or large scale batch processing.
- Ability to set technical direction in ambiguous early stage environments and carry it through implementation.
- Nice to have: Experience applying machine learning to scientific data, such as property prediction, generative models, graph based methods, or simulation data.
- Experience with atomistic, materials, chemistry, or engineering data systems.
- Experience with retrieval over structured data, knowledge graphs, or hybrid search systems.
- Experience designing APIs or tool interfaces that intelligent systems can call reliably.
- Experience building complex data and machine learning workflows on production orchestrators.
- Contributions to open source machine learning, data infrastructure, or scientific computing tools.
Core Competencies
Demonstrates expertise in building and optimizing data pipelines for machine learning and scientific applications, with a strong focus on data quality, model training, and workflow automation. Proficient in Python and experienced in large scale data systems, ensuring effective data management and analysis.