Cybersecurity Data Engineer
Summary
Build and maintain secure data pipelines and ETL workflows in Azure to ingest, clean, and deliver cybersecurity datasets for real-time analytics and SIEM/SOAR use cases.
Whitehall Resources are looking for a Cybersecurity Data Engineer. This role is hybrid working with 3 days per week onsite in South Yorkshire, and the remainder remote working, on an initial 12-month contract.
***Inside IR35***
Job Description:
- Ingestion and provisioning of raw datasets, enriched tables, and/or curated, re-usable data assets to enable Cybersecurity use cases.
- Driving improvements in the reliability and frequency of data ingestion including increasing real-time coverage.
- Support and enhancement of data ingestion infrastructure and pipelines.
- Designing and implementing data pipelines that will collect data from disparate sources across the enterprise, and from external sources, transport said data, and deliver it to our data platform.
- Extract Translate and Load (ETL) workflows, using both advanced data manipulation tools and programmatically manipulating data throughout our data flows, ensuring data is available at each stage in the data flow, and in the form needed for each system, service, and customer along said data flow.
- Identifying and onboarding data sources using existing schemas and, where required, conducting exploratory data analysis to investigate and determine new schemas.
Required Skills:
The successful candidate will be a student of the Google Site Reliability Engineering (SRE) philosophy as applied to managing large-scale cloud infrastructure, possess skills and experience within one or more of the following areas, and demonstrate a willingness to learn additional skills via certification and/or on-the-job learning where required.
Programming, Software & Network Principles:
- Experience with SRE and Azure DevOps
- Ability to script (Bash/PowerShell, Azure CLI), code (Python, C#, Java), query (SQL, Kusto query language) coupled with experience with software versioning control systems (e.g., GitHub) and CI/CD systems.
- Programming experience in the following languages: PowerShell, Terraform, Python Windows command prompt and object orientated programming languages.
- Demonstrable experience of Linux administration and scripting (preferably Red Hat Systems)
- Understanding of hardware and software principles and storage technologies (SSD, HDD, NVMe), CPU architectures, and Memory & Operating system principles (especially network stack fundamentals)
- Understanding of network protocols and network design
- Data Acquisition, Cloud-based Data Pipelines (Azure preferred)
- Data Transport and Data Cleaning
- Data Engineering pipeline automation, productionisation, and optimisation
- Designing, building, and maintaining data pipelines and ETL workflows across disparate datasets
- Dataset and Data Asset Curation
- Data Modelling and Cataloguing
- Database Architecture and Design
- Data Warehousing and Data Integration
- Real-Time Analytics Deployment for Large-Scale Datasets
- Applying data engineering methods to the cyber security domain
- Technical knowledge and breadth of Azure technology services (Identity, Networking, Compute, Storage, Web, Containers, Databases)
- Experience with server, operating system, and infrastructure technologies such as Nginx/Apache, CosmosDB, Linux, Bash, PowerShell, Prometheus, Grafana, Elasticsearch)
- Experience with Infrastructure-as-Code and Automation tools such as Terraform, Chef, Ansible, CloudFormation/Azure Resource Manager (ARM)
- Streaming platforms such as Azure Event Hubs or Kafka, and stream processing services such as Spark streaming
- Experience with Security Information & Event Management (SIEM) and Security Orchestration, Automation & Response (SOAR) technologies, especially cloud based, is a significant asset