Sr. Storage Design Engineer
Mission:
Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. Our infrastructure teams design and operate the systems that provide the compute, storage, networking, and platform capabilities behind Groq's rapidly growing AI infrastructure.
We are looking for a Senior Storage Design Engineer to architect, design, validate, and evolve the storage infrastructure supporting Groq's large-scale AI and compute environments.
As a Senior Storage Design Engineer, you will own storage architecture and design across high-performance data center and AI infrastructure environments. You will translate workload requirements—including capacity, throughput, latency, durability, availability, and cost—into scalable storage architectures that can be deployed and operated consistently.
This role requires deep expertise in distributed storage, high-performance storage systems, data center infrastructure, Linux, and automation. You will work closely with compute, network, platform, data center, security, and infrastructure operations teams to take storage designs from requirements and benchmarking through production deployment.
You will also help establish the engineering standards, reference architectures, validation methodologies, and technical direction that allow Groq's storage infrastructure to scale.
Responsibilities & opportunities in this role:
- Architect highly available, high-performance storage platforms supporting large-scale AI and compute infrastructure.
- Work closely with software teams building inference and training stacks to align storage topology with stack architecture.
- Develop storage architectures, high-level designs (HLDs), low-level designs (LLDs), implementation standards, and reference architectures.
- Design storage solutions optimized for high-throughput, low-latency, and highly parallel AI workloads.
- Architect file, object, block, and local storage solutions based on workload and application requirements.
- Evaluate distributed storage technologies and platforms such as Ceph, Lustre, Weka, VAST Data, Pure Storage, NetApp, Dell, IBM, or equivalent technologies.
- Design scalable storage systems using technologies such as NVMe, NVMe-oF, SSD, high-capacity HDD, object storage, and distributed file systems.
- Develop storage strategies for AI models, datasets, inference workloads, application data, logs, backups, and other infrastructure requirements.
- Analyze application I/O patterns and translate workload characteristics into storage performance and capacity requirements.
- Perform capacity planning and develop forecasting models for storage growth, utilization, performance, and lifecycle management.
- Define storage availability, durability, replication, erasure coding, data protection, backup, and disaster recovery strategies.
- Design storage architectures spanning multiple data centers or infrastructure environments where appropriate.
- Partner closely with network engineering to optimize storage traffic, including high-bandwidth east-west connectivity and technologies such as RDMA and RoCEv2.
- Evaluate storage servers, controllers, drives, network interfaces, and other hardware components for performance, reliability, density, and cost.
- Build lab environments, proofs of concept, benchmarks, and design-validation frameworks before introducing new technologies into production.
- Develop representative workload tests to measure throughput, IOPS, latency, metadata performance, scalability, and failure behavior.
- Define and test failure scenarios involving drives, storage nodes, network connectivity, controllers, and entire failure domains.
- Drive storage automation and infrastructure-as-code practices using Python, Ansible, APIs, Git, and CI/CD pipelines.
- Establish storage observability requirements, including telemetry, performance metrics, health monitoring, capacity visibility, and alerting.
- Troubleshoot complex performance and reliability issues spanning storage, network, compute, operating systems, and applications.
- Lead technical design reviews and provide guidance on complex storage architecture decisions.
- Partner with operations teams to ensure storage designs are maintainable, observable, upgradeable, and safe to operate at scale.
- Mentor engineers and raise the technical bar for storage architecture, benchmarking, automation, documentation, and engineering practices.
Ideal candidates have/are:
- 8+ years of experience in storage engineering, systems engineering, infrastructure engineering, or architecture, with significant experience designing large-scale storage environments.
- Deep understanding of distributed storage architecture and high-performance storage systems.
- Strong knowledge of file, object, and block storage technologies and their respective design tradeoffs.
- Experience designing storage for data-intensive, highly parallel, or high-performance computing environments.
- Strong understanding of NVMe, NVMe over Fabrics, SSD technologies, storage networking, and modern storage server architectures.
- Experience with distributed file systems, object storage platforms, or software-defined storage technologies.
- Strong understanding of storage performance characteristics, including IOPS, throughput, latency, queue depth, caching, metadata performance, and read/write patterns.
- Experience with storage resiliency concepts including replication, erasure coding, failure domains, snapshots, backup, and disaster recovery.
- Strong Linux systems knowledge, including filesystems, storage devices, multipathing, kernel I/O, and performance analysis.
- Experience with storage benchmarking and performance tools and methodologies.
- Experience automating infrastructure using Python, APIs, configuration management, or infrastructure-as-code tooling.
- Familiarity with Git-based workflows, automated testing, and CI/CD engineering practices.
- Experience with monitoring, telemetry, capacity management, and storage performance analysis.
- Demonstrated ability to develop architecture documents, design standards, implementation plans, and technical specifications.
- Ability to make architectural decisions involving performance, reliability, scalability, operational complexity, and cost.
- Strong communication skills and the ability to influence technical decisions across engineering organizations.
- Experience designing storage for AI/ML infrastructure, HPC environments, distributed computing systems, or large accelerator clusters.
- Experience supporting storage environments with extremely high aggregate throughput and highly concurrent access patterns.
- Knowledge of RDMA, RoCEv2, GPUDirect Storage, NVMe-oF, and high-performance Ethernet storage fabrics.
- Experience with parallel filesystems such as Lustre, IBM Storage Scale/GPFS, WekaFS, or equivalent technologies.
Compensation
Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is TBD, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.
US Job Posting
This position may require access to technology and/or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet certain citizenship or residency criteria. Specifically, they must qualify as U.S. Persons for export control purposes (i.e., U.S. citizen, U.S. lawful permanent resident (Green Card holder), or a protected individual under 8 U.S.C. § 1324b(a)(3) such as a refugee or asylee), or otherwise be eligible for an applicable export license.