-
Notifications
You must be signed in to change notification settings - Fork 2.9k
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
fix(ingest/bigquery): fix partition and median queries for profiling #8778
fix(ingest/bigquery): fix partition and median queries for profiling #8778
Conversation
column_profile.median = str( | ||
self.dataset.engine.execute( | ||
sa.select( | ||
sa.text(f"approx_quantiles(`{column}`, 2) [OFFSET (1)]") |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
we already have a patch for bq quantiles here https://github.com/datahub-project/datahub/blob/master/metadata-ingestion/src/datahub/ingestion/source/ge_data_profiler.py#L150 - is this also required?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
yes, this is required. As of version 0.15.50 of GX, median calculation does not use this bigquery patched method.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
could you add a comment explaining that
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
can do it in a follow-up too
* tag 'v0.11.0': (188 commits) fix(spark-test): upgrade gradle and fix spark smoke test (datahub-project#8777) fix(gms): Fixed Recently Viewed section for users with '@' in the URN. (datahub-project#8754) feat: add feedback widget (datahub-project#8732) fix(custom-search): fix custom search to be able to use unquoted query (datahub-project#8805) docs(db-retention): update with default setting (datahub-project#8797) feat(openapi): entity endpoints & analytics raw (datahub-project#8537) feat(search): Also de-duplicate the field queries based on field names (datahub-project#8788) fix(ingest): drop `wrap_aspect_as_workunit` method (datahub-project#8766) feat(ingest): drop sql_metadata parser (datahub-project#8765) docs: minor fix on versioning navbar and dropdown (datahub-project#8790) chore(ingest): upgrade sqlglot fork (datahub-project#8775) docs: add datahub source to integrations page (datahub-project#8787) fix(ingest/bigquery): fix partition and median queries for profiling (datahub-project#8778) fix(ingest/tableau): fix tableau native CLL for snowflake, add type annotations (datahub-project#8779) refactor(ingest): Add support for group-owners in dataflow entities (datahub-project#8154) feat(systemMetadata): Adding a lastRunId field system metadata (datahub-project#8672) feat(airflow-plugin): add package type information (datahub-project#8795) fix(ingest/datahub): Support postgres; build(postgres): Modernize postgres docker setup (datahub-project#8762) docs(session): add documentation for session token duration and fix default (datahub-project#8791) chore(analytics): bump version (datahub-project#8786) ...
Checklist