SRE (BigData)
Summary
Maintains and scales large Apache Hadoop, Hive, and Kafka clusters on Linux, ensuring uptime and performance while supporting East Coast U.S. hours.
What you will be doing
- Deploying, configuring, monitoring, and maintaining multiple big data stores across multiple data centers.
- Performing planning, configuration, deployment, and maintenance work relevant to the environment.
- Managing the large-scale Linux infrastructure to ensure maximum uptime.
- Developing and documenting system configuration standards and procedures.
- Performance and reliability testing. This may include reviewing configuration, software choices/versions, hardware specs, etc.
- Advancing our technology stack with innovative ideas and new creative solutions.
What you will need
- Multi-faceted Apache Hadoop + Hive (GitOps) understanding.
- Experience managing Kafka clusters on Linux.
- Thorough understanding of Linux (we use Rocky Linux in production).
- Any scripting language (Python/Ruby/Shell, etc.).
- Understanding of basic networking concepts (TCP/IP stack, DNS, CDN, load balancing).
- Willing and able to work East Coast U.S. hours 9 am-6 pm EST.
Bonus, but not required
- Experience administering Percona XtraDB Cluster.
- Experience using Kerberos for data storage and Trino for data retrieval.
- Experience with Security-related best practices.
- Puppet configuration management tool.
- Experience with infrastructure monitoring solutions: Icinga, Prometheus, Graphite, Grafana, and ELK.
- Experience working with Kubernetes.
- Cassandra cluster installation, troubleshooting, and maintenance.
- Train/mentor junior-level staff.
- Experience in AdTech or High-Frequency Trading.
We offer
- Remote work. Relocation to EU/UK/US is negotiable (depends on your current location and legal status).
- Salary: 7-9k USD / month, higher figures may be negotiated.
- US holiday schedule.
- 21 days of vacation.