Page History

...

New committers [Julien]
Release overview (0.6.0-0.6.1) [Michael R.]
Process for blog posts [Ross]
Retrospective: Spark integration [Willy et al.]
Open discussion

Meeting:

Widget Connector

url	http://youtube.com/watch?v=XaSboAHNk-k

Notes:

New committers [Julien]
- 4 new committers were voted in last week
- We had fallen behind
- Congratulations to all
Release overview (0.6.0-0.6.1) [Michael R.]
- Added
  - Extract source code of PythonOperator code similar to SQL facet @mobuchowski (0.6.0)
  - Airflow: extract source code from BashOperator @mobuchowski (0.6.0)
    - These first two additions are similar to SQL facet
    - Offer the ability to see top-level code
  - Add DatasetLifecycleStateDatasetFacet to spec @pawel-big-lebowski (0.6.0)
    - Captures when someone is conducting dataset operations (overwrite, create, etc.)
  - Add generic facet to collect environmental properties (EnvironmentFacet) @harishsune (0.6.0)
    - Collects environment variables
    - Depends on Databricks runtime but can be reused in other environments
  - OpenLineage sensor for OpenLineage-Dagster integration @dalinkim (0.6.0)
    - The first iteration of the Dagster integration to get lineage from Dagster
  - Java-client: make generator generate enums as well @pawel-big-lebowski (0.6.0)
    - Small addition to Java client feat. better types; was string
- Fixed
  - Airflow: increase import timeout in tests, fix exit from integration @mobuchowski (0.6.0)
    - The former was a particular issue with the Great Expectations integration
- - Reduce logging level for import errors to info @rossturk (0.6.0)
    - Airflow users were seeing warnings about missing packages if they weren't using a part of an integration
    - This fix reduced the level to Info
  - Remove AWS secret keys and extraneous Snowflake parameters from connection URI @collado-mike (0.6.0)
    - Parses Snowflake connection URIs to exclude some parameters that broke lineage or posed security concerns (e.g., login data)
    - Some keys are Snowflake-specific, but more can be added from other data sources
  - Convert to LifecycleStateChangeDatasetFacet @pawel-big-lebowski (0.6.0)
    - Mandates the LifecycleStateChange facet from the global spec rather than the custom tableStateChange facet used in the past
  - Catch possible failures when emitting events and log them @mobuchowski (0.6.1)
    - Previously when an OL event failed to emit, this could break an integration
    - This fix catches possible failures and logs them
Process for blog posts [Ross]
- Moving the process to Github Issues
- Follow release tracker there
- Go to https://github.com/OpenLineage/website/tree/main/contents/blog to create posts
- No one will have a monopoly
- Proposals for blog posts also welcome and we can support your efforts with outlines, feedback
- Throw your ideas on the issue tracker on Github
Retrospective: Spark integration [Willy et al.]
- Willy: originally this part of Marquez – the inspiration behind OL
  - OL was prototyped in Marquez with a few integrations, one of which was Spark (other: Airflow)
  - Donated the integration to OL
- Srikanth: #559 very helpful to Azure
- Pawel: is anything missing from the Spark integration? E.g., column-level lineage?
- Will: yes to column-level; also, delta tables are an issue due to complexity; Spark 3.2 support also welcome
- Maciej: should be more active about tracking projects we have integrations with; add to test matrix
- Julien: let’s open some issues to address these
Open Discussion
- Flink updates? [Julien]
  - Maciej: initial exploration is done
    - challenge: Flink has 4 APIs
    - prioritizing Kafka lineage currently because most jobs are writing to/from Kafka
    - track this on Github milestones, contribute, ask questions there
  - Will: can you share thoughts on the data model? How would this show up in MZ? How often are you emitting lineage?
  - Maciej: trying to model entire Flink run as one event
  - Srikanth: proposed two separate streams, one for data updates and one for metadata
  - Julien: do we have an issue on this topic in the repo?
  - Michael C.: only a general proposal doc, not one on the overall strategy; this worth a proposal doc
  - Julien: see notes for ticket number; MC will create the ticket
    - https://github.com/OpenLineage/OpenLineage/issues/596
  - Srikanth: we can collaborate offline

...

Page tree

Versions Compared

Old Version 84

New Version 85

Key