Sr. Platform Engineer
Summary
Sr. Platform Engineer at Carbon Arc, a remote-first, NYC-headquartered startup, building the foundation of its data intelligence platform: terabyte-scale ETL/ELT pipelines, Spark performance tuning, Apache Iceberg lakehouse tables queried via Trino/StarRocks, Airflow orchestration, and internal tooling, primarily in Python on AWS.
Compensation: $150k – $210k
# Platform Engineer ## About the Role Carbon Arc is looking for a talented Platform Engineer to help build the foundation of our data platform. You'll design, build, and maintain the systems that power our data intelligence products, spanning large-scale data pipelines, workflow orchestration, and the internal tooling that keeps the platform reliable and scalable. You'll thrive here if you enjoy a fast-paced startup environment where your decisions carry real technical and product weight. If you care about building reliable data systems and writing clean, well-tested code, this role offers a wide range of challenging, high-impact problems. ## What You'll Do - Design and build performant ETL/ELT pipelines for massive data sources, including terabytes of structured and unstructured data - Tune Spark workloads for scale: diagnose and fix data skew (salting, record sharding, bucketed joins), manage executor/driver heap, and benchmark changes to prove out performance wins - Work deep in the lakehouse: manage Apache Iceberg tables, partitions, and snapshots; query through Trino and StarRocks; and keep catalogs (e.g. Polaris) consistent and safe - Develop and maintain workflow orchestration using Airflow or similar DAG-based systems, including config-driven DAG frameworks that make pipeline onboarding repeatable - Write clean, well-tested Python (and, where it counts, Spark/Scala internals such as native Catalyst expressions and UDFs) to solve data and infrastructure challenges - Build internal tooling, runbooks, and automation to improve developer productivity and pipeline reliability - Implement data quality monitoring with traceability back to source systems, reading from live table state rather than stale caches - Define schemas, partitions, and indexes aligned to usage patterns and performance needs - Create, optimize, and maintain queries against analytical databases - Collaborate on platform architecture decisions and help establish engineering best practices - Debug complex issues across distributed data systems and services, including remote/classic Spark session conflicts and memory pressure - Document systems, pipelines, and operational procedures to support long-term scalability ## What You'll Bring - 3–5+ years of software engineering experience focused on data systems or platform engineering - Strong hands-on experience with Spark, including performance tuning, skew mitigation, and memory/heap management on terabyte-scale workloads - Strong proficiency in Python, with an emphasis on testing and code quality - Deep experience building and operating ETL pipelines, data quality controls, and workflow orchestration systems (Airflow, Dagster, Prefect) - Hands-on experience with data lake architectures and Apache Iceberg, including partition and snapshot management and catalog awareness - Strong SQL skills against analytical engines such as Trino, StarRocks, or similar, including query optimization and partitioning strategies - Hands-on experience with AWS services such as S3, IAM, Glue, EMR, and RDS - Comfort reasoning about benchmarking and measuring changes, not just shipping them - Comfort integrating AI-assisted development tools into your daily workflow - Strong written and verbal technical communication skills, including clear runbooks and operational docs - A proactive mindset with the ability to adapt quickly and drive change in evolving systems ## Nice to Have - JVM-level Spark work: writing or optimizing native Catalyst expressions, Scala UDFs, or codegen paths - Experience with data index performance tuning - An active GitHub profile showcasing personal or open-source projects in data or platform engineering - Experience with MLOps, feature stores, or ML pipeline orchestration - Kubernetes/EKS and container orchestration experience - Infrastructure-as-code tools such as Terraform - GitOps workflows (e.g., ArgoCD) - Additional analytical databases such as ClickHouse or Snowflake - Experience ingesting regulatory or third-party datasets (e.g. FERC/XBRL, PUDL) with reliable streaming sinks - Exposure to data security, compliance, and access management frameworks (e.g., SOC 2) ## What We Offer - Competitive compensation - Fully paid healthcare benefits (medical, dental, vision) - Remote work options - 401(k) plan with employer contributions - Paid time off - Generous parental and family leave policies - Challenging problems to tackle with a supportive, entrepreneurially driven team - A strong work culture centered on integrity, excellence, truth, trust, and transparency ## Location Carbon Arc is headquartered in New York City and operates as a remote-first company. Team members travel for in-person onsite gatherings at least one week per quarter to collaborate, plan, and connect as a team. --- *Carbon Arc is an Equal Opportunity Employer committed to fair and equitable hiring. All candidates are considered without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, genetic information, marital status, pregnancy, veteran or military status, citizenship status, or any other characteristic protected by applicable law.* *Need an accommodation to participate in the application or interview process? Reach out to careers@carbonarc.co.* *Note to recruiters and placement agencies: Carbon Arc does not accept unsolicited resumes. Any resume submitted without a prior written agreement will be deemed the property of Carbon Arc, and no fee will be paid in the event of a hire.*