Data Engineer
Senior lvl.
We are looking for a Data Engineer to join our Detection-Coverage Data Platform team, working on high-volume, production data pipelines. This is not a BI, dashboarding, or pure analytics role. You will own the platform that extracts, normalizes and models large-scale customer SIEM data from systems such as Splunk, Microsoft Sentinel, and Google SecOps into graph, relational, and lakehouse models that power our product and data science teams. The role offers strong ownership, complex technical challenges, and the opportunity to design durable, fault-tolerant data systems that operate reliably under strict performance, rate, concurrency and scale constraints.

Responsibilities
>/ Own and develop durable, back-pressured, fault-tolerant ETL pipelines operating across thousands of data streams per tenant.
>/ Design and evolve graph, relational, and lakehouse data models, choosing the right storage layer for each access pattern while maintaining a reliable source of truth.
>/ Diagnose and resolve database concurrency and performance issues, including lock contention, deadlocks, hotspots, inefficient queries, and concurrent write challenges.
>/ Optimize data processing through batching, partitioning, idempotent upserts, and appropriately sized worker fleets.
>/ Build end-to-end observability and reliability into pipelines through metrics, structured logging, distributed tracing, data-quality validation, and pipeline SLIs.
>/ Write production-grade, async-first, typed, and tested Python within structured workflow orchestration.
>/ Containerize and deploy data services using Docker, Kubernetes, and cloud-native infrastructure.
>/ Collaborate closely with data scientists and product engineers to build the platform supporting statistical and ML workflows and maintain clear data contracts.

Requirements
>/ Proven experience building and operating data-intensive backend systems or large-scale data infrastructure.
>/ Strong experience with production-grade async Python, including typing, data validation, testing, and CI.
>/ Experience with workflow orchestration such as Temporal, Airflow, Dagster, or Step Functions.
>/ Strong PostgreSQL experience, including relational schema design, query tuning, concurrency, and performance troubleshooting.
>/ Experience modeling data with graph databases (Neo4j / Cypher) or open lakehouse technologies (Apache Iceberg / Delta / Parquet on S3).
>/ Strong understanding of distributed data processing, concurrent writes, batching, partitioning, and idempotent processing.
>/ Hands-on experience with S3 / object storage, columnar formats such as Parquet, and query engines such as Athena, Trino, or Spark.
>/ Strong understanding of observability and data quality, including metrics, distributed tracing, schema validation, pipeline SLIs and ideally OpenTelemetry.
>/ Experience with Docker, Kubernetes / EKS, Infrastructure as Code and cloud-native deployment practices.

Nice to Have
>/ Hands-on experience with SIEM or cybersecurity platforms, such as Splunk, Microsoft Sentinel, Google SecOps, or CrowdStrike.
>/ Experience extracting data from rate-limited, credential-scoped customer environments.
>/ Experience with security data normalization standards, such as OCSF.
>/ Understanding of streaming vs.
>/ batch architectures and technologies such as Kafka, Kinesis, or Firehose.
>/ Experience with Helm and ArgoCD / GitOps deployment workflows.
>/ Experience working closely with data science / ML teams on production data platforms.