At Sword, we’re building AI to heal billions and unlock humanity’s full potential.
Requirements
Strong experience building batch pipelines and analytics datasets at scale.
Proficiency with Python and SQL - SQL especially; this role leans analytics.
Deep data modeling skills: dimensional, Data Vault, or similar; semantic layer design; metric definitions.
dbt in production as a first-class skill - you know how to structure a large dbt project, design for testability, and keep it maintainable as it grows.
Hands-on experience with at least one modern warehouse or lakehouse engine - Snowflake, BigQuery, Databricks, or Trino/Starburst.
Production experience with a workflow orchestrator (Airflow, Dagster, or similar).
Clear communicator: you can talk to PMs, analysts, and clinicians without drowning them in jargon.
Pragmatic: you ship the 80% solution and iterate.
Ownership: you don’t hand off broken pipelines.
Familiarity with lakehouse table formats (Iceberg, Delta, Hudi).
Understanding of Kafka and event-driven sources, enough to consume them into batch layers sensibly.
Streaming exposure - Flink or Spark Structured Streaming - for when batch isn’t enough.
Experience with reverse-ETL, metrics stores, or semantic layer tooling (Cube, LookML, MetricFlow).
Experience in healthcare, HIPAA, or FedRAMP environments.
PySpark or Spark SQL at scale.
Benefits & Perks
A stimulating, fast-paced environment with lots of room for creativity.
A bright future at a promising high-tech startup company.
Career development and growth, with a competitive salary.
The opportunity to work with a talented team and to add real value to an innovative solution with the potential to change the future of healthcare.
A flexible environment where you can control your hours (remotely) with unlimited vacation.
Access to our health and well-being program (digital therapist sessions).
Remote or Hybrid work policy.
To get to know more about our Tech Stack, check here .