claudeers.
// Claude Skills

awesome-data-engineering-skills

36 data-engineering skills for AI coding agents — dbt, Airflow, Dagster, Prefect, Spark, Snowflake, BigQuery, Databricks, Kafka, Iceberg. Portable SKILL.md A…

Actively maintained
90/100
last commit about 1 month ago
last release none
releases 0
open issues 0
// star history+2 this week

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up awesome-data-engineering-skills (claude-plugin project) into my current project.
Found on https://claudeers.com/awesome-data-engineering-skills
Repo: https://github.com/Unknown-333/awesome-data-engineering-skills
Homepage/docs: —
Detected install method: claude-plugin → /plugin install awesome-data-engineering-skills@Unknown-333/awesome-data-engineering-skills
Category: skills. Platforms: cli.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or install directly (claude-plugin)
/plugin marketplace add Unknown-333/awesome-data-engineering-skills
/plugin install awesome-data-engineering-skills@Unknown-333/awesome-data-engineering-skills
// or clone
git clone https://github.com/Unknown-333/awesome-data-engineering-skills

// compatibility

Platformscli
Operating systems—
AI compatibilityclaude
LicenseNOASSERTION
Pricingopen-source
LanguagePython

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

Awesome Data Engineering Skills — data-engineering skills for your AI coding agent

Awesome Data Engineering Skills

A curated, portable collection of Agent Skills that make AI coding agents genuinely useful for data engineering work — dbt, Airflow, Dagster, Spark, Snowflake, BigQuery, Databricks, data quality, idempotency, backfills, and pipeline operations.

Each skill is a folder with a SKILL.md the agent loads only when relevant, so your context stays lean until you need the expertise. Skills follow the open Agent Skills standard and work across Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI, and other compatible agents.

Built from real, recurring data-engineering pain points (idempotency & backfills, silent schema changes, data quality incidents, warehouse cost, testing pipelines) — not generic advice.

Why skills for data engineering?

Data bugs are expensive and quiet: a non-idempotent retry duplicates a fact table, a backfill rewrites yesterday's revenue, a SELECT * becomes a recurring bill. These skills encode the correct default so the agent gets it right the first time.

Skill catalog

SQL & data modeling

SkillUse it when
optimizing-sql-queriesA query is slow/expensive, scans too much, or spills — across Postgres, Snowflake, BigQuery, Spark, Redshift
modeling-dimensional-dataDesigning star/snowflake schemas, choosing grain, or tracking history with SCDs
writing-idempotent-transformationsA retry could duplicate data, or a load must be backfill-safe

dbt

SkillUse it when
building-dbt-modelsCreating/refactoring models, choosing materialization, writing incremental logic
testing-dbt-projectsAdding tests, catching data quality regressions, source freshness
debugging-dbt-runsdbt run/build fails, incremental is stale/duplicated, CI differs from local
documenting-dbt-modelsAdding descriptions, exposures, and generated docs/lineage

Orchestration

SkillUse it when
authoring-airflow-dagsWriting DAGs, scheduling, retries/backfills, fixing non-idempotent tasks
debugging-airflow-pipelinesA task fails or is stuck queued, scheduler won't run, zombie tasks
building-dagster-assetsBuilding software-defined assets, partitions, resources, asset checks
orchestrating-prefect-flowsWriting Prefect flows/tasks, retries, caching, deployments/schedules

Distributed processing

SkillUse it when
optimizing-pyspark-jobsA Spark job is slow, spills, OOMs, or has skew/large shuffles
engineering-databricks-pipelinesBuilding Delta/DLT pipelines, Auto Loader, Unity Catalog

Lakehouse & storage

SkillUse it when
designing-medallion-architectureOrganizing a lakehouse into bronze/silver/gold layers
building-iceberg-tablesCreating/maintaining Iceberg tables, partitioning, compaction, time travel
optimizing-parquet-storageSlow/costly Parquet reads, small-files problem, compression/layout

Warehouses & cost

SkillUse it when
optimizing-snowflake-workloadsSnowflake is slow/expensive, warehouses spill/queue, sizing decisions
optimizing-bigquery-queriesBigQuery bytes billed are high or a query full-scans

Data quality & contracts

SkillUse it when
implementing-data-quality-checksCatching bad data before consumers — freshness, volume, schema, integrity
designing-data-contractsA producer change could silently break downstream pipelines
handling-schema-evolutionAdding/renaming/retyping columns without breaking readers

Ingestion & streaming

SkillUse it when
building-ingestion-pipelinesExtracting from DBs/APIs/files — incremental, CDC, watermarks, pagination
processing-streaming-dataKafka/Spark/Flink streaming — delivery semantics, windowing, late data
building-kafka-consumersWriting Kafka consumers/producers, offsets, consumer groups, dead-letter
implementing-cdc-with-debeziumReplicating an OLTP DB with CDC, capturing deletes, applying change streams

Pipeline engineering & ops

SkillUse it when
debugging-data-pipelinesNumbers look wrong, data missing/duplicated, a dashboard is stale
designing-backfills-and-replaysBackfilling history or reprocessing after a fix — safely
implementing-pipeline-observabilityFailures are found by stakeholders, not alerts; setting SLAs/monitoring
reviewing-data-pipeline-codeReviewing a data PR for idempotency, grain, cost, tests, PII
implementing-data-cicdAdding CI (lint/compile/test), dbt Slim CI, environment promotion
managing-data-lineage-openlineageCross-tool lineage, impact analysis, scoping incidents/backfills
migrating-legacy-etlMigrating SSIS/Informatica/stored procs/on-prem to dbt/Spark/cloud
generating-synthetic-test-dataRealistic seeded test data, edge cases, referential integrity

Platform, governance & ML

SkillUse it when
terraform-for-data-infraProvisioning warehouses/buckets/IAM/orchestration as code
masking-pii-dataMasking/tokenizing PII, dynamic masking, GDPR/CCPA/HIPAA, deletion
building-feature-pipelinesML features, point-in-time joins, offline/online parity, feature stores

Install

Skills live in this repo under skills/. Point your agent at them by copying or symlinking into the tool's skills directory. .agents/skills/ is recognized by the widest set of tools.

Quick install (macOS/Linux)

./scripts/install.sh claude     # -> ./.claude/skills/   (Claude Code)
./scripts/install.sh cursor     # -> ./.cursor/skills/   (Cursor)
./scripts/install.sh codex      # -> ./.codex/skills/    (Codex)
./scripts/install.sh copilot    # -> ./.github/skills/   (GitHub Copilot)
./scripts/install.sh agents     # -> ./.agents/skills/   (broadest support)
# add --user for a global install, --copy to copy instead of symlink

Per-tool skills directories

ToolProject pathPersonal (global) path
Claude Code.claude/skills/ or .agents/skills/~/.claude/skills/
Cursor.cursor/skills/ or .agents/skills/ (also reads .claude/, .codex/)~/.cursor/skills/, ~/.agents/skills/
Codex.codex/skills/~/.codex/skills/
GitHub Copilot.github/skills/, .claude/skills/, or .agents/skills/~/.copilot/skills/, ~/.agents/skills/

Claude Code plugin (optional)

Install the whole set as a plugin marketplace:

/plugin marketplace add <your-org>/awesome-data-engineering-skills
/plugin install data-engineering-skills@awesome-data-engineering-skills

How skills load

  1. Discovery — the agent preloads each skill's name + description (~tiny).
  2. Activation — when your task matches, it reads the full SKILL.md.
  3. Resources — deeper references/*.md load only when needed.

You can also invoke a skill directly by typing / in chat and picking it by name.

Portability

Phase-1 skills use only the portable spec frontmatter (name, description) so they behave identically across tools. Tool-specific extensions (dynamic command injection, forked subagents, paths) are intentionally avoided. See CONTRIBUTING.md.

Contributing

New skills and improvements welcome — see CONTRIBUTING.md. Validate before opening a PR:

python scripts/validate_skills.py   # frontmatter, naming, references
python scripts/check_evals.py       # every skill has trigger/non-trigger evals

Each skill has description-tuning eval cases in evals/triggering.json (prompts that should and should not activate it) to keep discovery accurate.

Roadmap (backlog)

The initial catalog above covers the core plus the first expansion set. Candidate future skills: writing-great-expectations-suites · tuning-warehouse-costs · building-realtime-analytics · managing-data-catalogs · orchestrating-dbt-airflow.

License

Apache-2.0

// faq

What is awesome-data-engineering-skills?

36 data-engineering skills for AI coding agents — dbt, Airflow, Dagster, Prefect, Spark, Snowflake, BigQuery, Databricks, Kafka, Iceberg. Portable SKILL.md Agent Skills (the format Anthropic launched as Claude Skills), with idempotency, backfill, and data-quality know-how. Runs in Claude Code, Cursor, Codex, Copilot, Gemini CLI and more.. It is open-source on GitHub.

Is awesome-data-engineering-skills free to use?

awesome-data-engineering-skills is open-source under the NOASSERTION license, so it is free to use.

What category does awesome-data-engineering-skills belong to?

awesome-data-engineering-skills is listed under skills in the Claudeers registry of Claude-compatible tools.

7 views
★ 21 stars
unclaimed
updated about 1 month ago

// embed badge

awesome-data-engineering-skills on Claudeers
[![Claudeers](https://claudeers.com/api/badge/awesome-data-engineering-skills.svg)](https://claudeers.com/awesome-data-engineering-skills)

// retro hit counter

awesome-data-engineering-skills hit counter
[![Hits](https://claudeers.com/api/counter/awesome-data-engineering-skills.svg)](https://claudeers.com/awesome-data-engineering-skills)

// reviews

// guestbook

0/500

// related in Claude Skills

🔓

An agentic skills framework & software development methodology that works.

// skillsobra/⟨Shell⟩★ 292,190◷ MIT[ claude ]
🔓

Public repository for Agent Skills

// skillsanthropics/⟨Python⟩★ 178,324[ claude ]
🔓

💫 Toolkit to help you get started with Spec-Driven Development

// skillsgithub/⟨Python⟩★ 138,919◷ MIT[ claude ]
🔓

AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs,…

// skillsGraphify-Labs/⟨Python⟩★ 123,800◷ MIT[ claude ]
→ see how awesome-data-engineering-skills connects across the ecosystem