Senior DevOps Engineer
Summary
Build and maintain secure, scalable cloud infrastructure and CI/CD pipelines using AWS, Azure, OpenTofu, and Datadog to support product teams and accelerate software delivery.
The Platform Engineering team works across the stack to give product teams paved, secure, cost-efficient paths to build, ship, and run software with minimal cognitive load. We own the "how," so product teams can focus on the "what." You will join the team responsible for running the core infrastructure that supports Shelf products. This role is primarily based in our European offices in Wroclaw, Poland and Lviv, Ukraine.
You will develop reusable components, improve system performance, and create scalable abstractions that accelerate product development across the organization.
You will maintain high standards for reliability and security in your work and in the systems used by other teams.
This is a high-ownership, hands‑on engineering role. You will manage everything from Terraform/OpenTofu modules and CI/CD pipelines to SSO permissions and observability tools, with a mandate to build infrastructure that works and keeps working.
You will work with AWS, Datadog, OpenTofu, Snowflake, GitHub, Azure, various LLMs, and many other tools and services.
Responsibilities
- Write and maintain infrastructure as code in OpenTofu, making modules more reusable and robust so that more engineers can ship infrastructure safely on their own.
- Write clear runbooks and playbooks that explain how things work and what to do when they break; present your work in a clean, structured way, favor a good doc once to enable self‑serve.
- Care deeply about the health of our infrastructure by keeping databases, LLMs, and third‑party self‑hosted services on current, supported versions, standardizing them across environments, and actively hunting down and removing outdated components.
- Participate in on‑call rotations and incident response, and write clear post‑mortems with concrete action items. Turn every incident into an opportunity to improve, define and refine SLOs and error budgets, and follow through on the work that prevents repeats.
- Treat CI/CD pipelines as a critical product. Own and improve hundreds of pipelines by making them faster, more reliable, easier to roll back, and more standardized to reduce manual toil and mental overhead for developers.
- Become a Datadog and observability expert, tuning logging, metrics, tracing, dashboards, and alerts to squeeze out as much useful signal as possible. Build simple defaults, automation, and clear docs so developers can self‑serve, contribute to observability, and rely on a solid platform.
- Make thoughtful build‑vs‑buy decisions and work directly with vendors and cloud support (AWS, Azure, GCP, and others) to solve infrastructure problems, plan upgrades, and find cost savings while asking good questions and pulling in expertise.
- Implement and enforce SOC 2‑aligned policies for infrastructure and deployments, including disaster recovery, business continuity, change management, and security policies, ensuring they are practical, documented, and followed in day‑to‑day work.
Qualifications
- Take pride in building and operating scalable, reliable, secure systems and are not comfortable bypassing protocols or cutting corners.
- Took full ownership of work, handle ambiguity and rapid change, and proactively remove obstacles to deliver results.
- Comfortable diving into any part of the stack, from infrastructure and backend services to product frontends, when that is what it takes.
- Use Python to automate repetitive work and improve workflows.
- Read AWS re:Invent announcements and us‑east‑1 post‑mortems for fun.
Sample Projects
- Provision the entire product infrastructure and applications in a new cloud or region.
- Design and implement a live database migration from us‑east‑1 to us‑east‑2.
- Maintain a 100 % score on the AWS CIS Benchmark in our environments.
- Centralize audit trail logs from AWS, GCP, and Azure into a single place.
- Write a clear runbook describing how you conducted a disaster recovery test of a system component.
- Change the SSO provider and reconfigure services to use the new provider.
- Optimize infrastructure costs by improving configuration, identifying abandoned resources, and applying reserved or committed compute purchases where appropriate.
Benefits
- Premier AI development environment: GitHub Copilot, Claude Code, OpenAI, TypingMind, v0, MCP Servers, plus credits to experiment with emerging AI tools.