DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
Status
| State | DraftAccepted |
| Discussion Thread | |
| Vote Threadtbd. | [LAZY,CONSENSUS] Example dags |
| Vote Result Threadtbd. | Re: [LAZY,CONSENSUS] Example dags |
| Progress Tracking (PR/GitHub Project/Issue Label) | Github Project "Airflow Examples Refurbish" |
| Date Created |
|
| Version Released | tbd. |
| Authors |
Motivation
Airflow Examples have been grown in number and focus over the past years. They purpose multiple things:
...
- The number of examples should be reduced to 20-30
- If possible examples from docs should be represented in examples. Some code examples which are stand-alone in code should be moved into examples if possible.
- Otherwise example content not referenced in documentation might be questioned if beneficial
- If possible examples should be directly usable by copy&paste (consider imports for integrated snippets)
- Examples should be arranged along a story-line if possible which might represent a virtual company and support real-life use cases. Might be good if we can structure it as growth of use-cases bringing in need for more features. i.e. starting off basic transformation → needing more complex task so using groups → using sensors, assets... showing progression of usage with project/company maturity.
- Examples should follow best practices in coding
- Existing examples should be reviewed which DAGs are just used for testing. Testing DAGs should be separated and not pollute the example collection
- Some examples are specific for providers. They should be moved to provider packages
- A mechanism is to be created that uses DAG bundle loading mechanism to load example DAGs from providers w/o need to copy them to global examples.
- Same like today if loading of example DAGs is enabled also needed plugins e.g. timetables should be usable out-of-the-box
...
- "Tailwind" - A virtual / non existing wind park energy company that powers a farm of win-mills to produce clean energy. The company has a strong demand to ETL sensor data from the windmills as well as need to act on data events when base data changes or contracts with customers renew. The company values also the DEI rules and has sustainable targets for clean energy and CO2 reduction.
Event driven case: New data is dropped on a file system (S3 would be great but can not be executed w/o S3, can be added to AWS provider as extension) that triggers data processing for wind energy. Data is loaded and Asset events are generated (see Github Issue 52481)
- Asset driven pipeline around reporting, data is split up per city of wind turbine and reports about production are distributed to shareholders. Branching can be used to check if a notification is sent via email or a custom notifier. This can also use branch labels. A third notification channel might be broken and as these shareholders are important we inform the admin in case of any task fails (trigger rule) and start a recovery task.
- Scheduled nightly use case, example reporting is written to file system (where event pipeline is picking up!). As it would be too easy some tasks migth fail and then a custom weight rule is used for retries.
- Manual correction trigger: Correction wind production counters can be submitted which then also are written to file system
- Timetable example for maintenance schedule (selected calendar dates) where maintenance notifications are sent (see Github Issue 52479)
- Scheduled hourly check for wind turbines state. This requires som some special infrastructure to start and stop, using setup+teardown to open a VPN tunnel to the remote machines. This is using a generator pattern and produces the same logic for 3 counties. (see Github Issue 52480)
As for demos and examples a lot of functionality is needed in both decorator as well as classic Dag implementation it would be good to have two similar use cases. Or alternatively describe that the Tailwind south branch prefers to implement all in Pythonic manner whereas the Tailwind North branch data engineers like the classic implementation?
...
-
--load-example-dagsmust load examples from standard provider at least in Airflow 3.1 (same like in breeze hack today) (see Github Issue 52469) - Testing DAGs must be loadable (at least in breeze) to be able to remove them from example tree (See Github Issue 52474)
Likely we should have a dedicated "test_dags" folder/bundle that should contain dags used for testing only. We can automatically add such dags to be used in unit test via auto-fixture in the sharedconftest.pyor similar - this way it will work in both breeze and local venv. I think we should aim to have breeze == local env - DAG Bundles must be extended depending on installed/available providers to extend examples - allowing to move examples from core to providers (e.g.
example_kubernetes_executor.py→ cncf.kubernetes)
Likely we can add "examples" section inprovider.yamland move the example dags from "tests" to separate "examples" folder that will be also embedded in the.whlfile (so that you can also conditionally enable examples from a given provider). This means that "system" tests will not be a separate "system" folder, but something separate. -
We should figure out a way how to show examples from a provider - possibly "load_core_examples=true/false" and "load_provider_examples=[list of providers]" would be a nice way how to do it
-
The example have to be reviewed with "security" point of view. We often have security reports that are pointing to RCE , lack of sanitizations etc. in our examples - which is very important as those examples can be used by others and bad practices propagated to production code.
- Add a review checklist into the repo to remember the qulity gates we defined for future reviews and extensions after the examples have been cleaned-up. (see Github Issue 52476)
- Add a check that links in example dags/doc_md are valid to currest RST (See Github Issue 52477)
Proposed Technical Excellence in Example DAGs
- All code has documentation (pydoc)
- All DAGs have DAG MD docs and task MD docs
The MD must include a lightweight Dag Preview - Let's Include a short summary(1–2 sentences) in the Dag’s description, explaining what it demonstrates/features and why it’s useful. So that users don't have to really go through the story to understand that one feature.
The MD should include links to the official Airflow documentation. This would provide users with instant access to deeper reference material for operators, hooks, or features being demonstrated, without them needing to search separately. - All DAGs and Tasks use Typing
- The Examples use tags mapping to use cases of storyline.
- All examples carry the the tag "
Example" - If the DAG is serving as a tutorial, it is having the tag "
Tutorial" - Current examples that we move off for testing (or which are dual-use) get a tag "
Testing" Feature Tags for Easier Discoverability - The new storyline-based example Dags are great for understanding end-to-end workflows. That said, they can be a bit overwhelming if you’re just looking to learn a specific concept (e.g., Dynamic Task Mapping, Assets, Sensors, XCom). It helps to add feature tags to each Dag. So users can more easily find Dags showcasing a specific capability. We should maintain a specific set of tags.
- Examples for a provider (other than standard) will get a tag with the name of the provider (e.g.
"Edge3")
- All examples carry the the tag "
- Ruff + Mypy checks are enabled and examples follow code quality guidelines
- DAGs and Tasks have nice display names
- Examples do not use deprecated functions
- Examples do not carry top-level code
- Code in the examples follow best practices with regards to security (sanitizatio, RCE protection, no exposure of sensitive information and the like)
- All examples are running in the standard setup out-of-the-box. They need to run out-of-the-box (e.g. no connections, SW packages need to be created prior run)
Note: This will not be possible for some provider examples - for example Google - they need at least configuration of Google connection and some of them need a separate setup. Possibly in those cases we should just document prerequisites for them. Or we need to find other creative options like checking for existence and short-circuit if per-requisites are not existing (e.g. a snowflake back-end needed). If can be set-up automatically then setup- and teardown tasks might be an option. - Examples should integrate into the story-line and not just stand-alone to showcase a technical feature
- All xamples should be (if possible) available as Taskflow and Non-Taskflow (we call it Classic in the examples)
...
