SRE (BigData)

Summary

Maintains and scales large Apache Hadoop, Hive, and Kafka clusters on Linux, ensuring uptime and performance while supporting East Coast U.S. hours.

What you will be doing

  • Deploying, configuring, monitoring, and maintaining multiple big data stores across multiple data centers.
  • Performing planning, configuration, deployment, and maintenance work relevant to the environment.
  • Managing the large-scale Linux infrastructure to ensure maximum uptime.
  • Developing and documenting system configuration standards and procedures.
  • Performance and reliability testing. This may include reviewing configuration, software choices/versions, hardware specs, etc.
  • Advancing our technology stack with innovative ideas and new creative solutions.

What you will need

  • Multi-faceted Apache Hadoop + Hive (GitOps) understanding.
  • Experience managing Kafka clusters on Linux.
  • Thorough understanding of Linux (we use Rocky Linux in production).
  • Any scripting language (Python/Ruby/Shell, etc.).
  • Understanding of basic networking concepts (TCP/IP stack, DNS, CDN, load balancing).
  • Willing and able to work East Coast U.S. hours 9 am-6 pm EST.

Bonus, but not required

  • Experience administering Percona XtraDB Cluster.
  • Experience using Kerberos for data storage and Trino for data retrieval.
  • Experience with Security-related best practices.
  • Puppet configuration management tool.
  • Experience with infrastructure monitoring solutions: Icinga, Prometheus, Graphite, Grafana, and ELK.
  • Experience working with Kubernetes.
  • Cassandra cluster installation, troubleshooting, and maintenance.
  • Train/mentor junior-level staff.
  • Experience in AdTech or High-Frequency Trading.

We offer

  • Remote work. Relocation to EU/UK/US is negotiable (depends on your current location and legal status).
  • Salary: 7-9k USD / month, higher figures may be negotiated.
  • US holiday schedule.
  • 21 days of vacation.