Tech jobs
Job listings
Software Engineer, Observability
Build and maintain scalable observability infrastructure for CoreWeave’s AI cloud platform, focusing on metrics, logging, tracing, and telemetry pipelines to support thousands of GPUs and petabyte-scale data.
Staff Software Engineer, Observability
Leads the design, scaling, and reliability of CoreWeave’s global observability stack—logging, tracing, and metrics—using tools like Prometheus, Grafana, and Kubernetes, while mentoring engineers and managing production clusters.
Senior Software Engineer II, Applied Training
Builds and maintains Kubernetes-native AI research infrastructure for CoreWeave’s customers, focusing on job orchestration, distributed training, and developer tooling to accelerate AI model development.
Principal Software Engineer, Developer Experience
Principal engineer defines and drives the technical strategy for developer tools, build/test systems, and CI/CD at a large-scale GPU cloud, using Kubernetes and Go.
Senior Manager, Production Engineering
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading…
Solutions Architect
Design and deploy secure AI workloads for U.S. government agencies, ensuring compliance with FedRAMP and DoD standards while optimizing GPU-powered cloud infrastructure for mission-critical applications.
Senior Software and AI Engineer
Builds AI-powered agents and internal tools using frameworks like LangChain and React, integrating with enterprise systems (ERP, HRIS) and deploying on cloud infrastructure.
IT Operations Specialist
Supports and scales CoreWeave's internal IT environment, handling end-user support, identity management, and automation while ensuring security and reliability for a growing workforce.
Staff Software Engineer, Network Development
A senior engineer designs and builds software to automate and operate CoreWeave’s global GPU cloud network, using Python and Go to provision, monitor, and heal infrastructure while collaborating with platform teams.
Senior Software Engineer, Box Office Platform
Senior Software Engineer builds and maintains Go-based services and APIs to automate hardware lifecycle management for CoreWeave's global AI cloud fleet, including RMA, repairs, and vendor integrations.
Senior Software Engineer, Fleet Monitoring Analysis
Builds and maintains observability systems for CoreWeave’s global fleet of GPU servers, using Prometheus, Grafana, and Kubernetes to automate monitoring and alerting.
Senior Software Engineer, Server Fleet Infrastructure
Senior engineer builds and maintains automation for CoreWeave’s global bare-metal server fleet using Go, gRPC, Kubernetes CRDs, and data-center tooling.
Senior Software Engineer, Box Office Platform
Senior Software Engineer builds and maintains Go-based services and APIs that automate hardware lifecycle management for CoreWeave’s global AI cloud fleet, integrating data centers, vendors, and internal platforms.
Security Operations Engineer II
Triage and respond to security alerts in a 24/7 SOC, using SIEM/EDR and Linux/MacOS/Kubernetes tooling to investigate incidents and escalate as needed.
Senior Software Engineer - GPU Kernel Authoring & Optimization
Write, profile, and optimize CUDA kernels for LLM inference to maximize throughput and minimize latency on NVIDIA GPUs, using DSLs like Triton or Mojo and benchmarking with MLPerf.
Senior Software Engineer
Senior Software Engineer builds and operates CoreWeave’s Kubernetes-based data platform, ensuring reliability, scalability, and security for AI workloads and mission-critical services.
Security Programs Senior TPM - Enterprise Security & IAM
Leads enterprise security and Identity & Access Management (IAM) programs at a cloud AI infrastructure company, ensuring secure, least-privileged access across systems while managing cross-functional initiatives.
Staff Software Engineer
Designs and maintains CoreWeave’s Kubernetes-based data platform, ensuring high availability, security, and performance for AI workloads and data pipelines using tools like Spark, Kafka, and Prometheus.