查询引擎
Google BigQuery
登录后可跨设备保存划线和私人笔记登录
Iceberg 表
要读取 Apache XTable™ (Incubating) 同步的 Iceberg 表(来自 BigQuery),你有两种选择:
使用 Iceberg JSON 元数据文件创建 Iceberg BigLake 表:
Apache XTable™ (Incubating) 会为 Iceberg 目标格式的同步输出元数据文件,BigQuery 可以利用这些文件读取 BigLake 表。
sql
CREATE EXTERNAL TABLE xtable_synced_iceberg_table
WITH CONNECTION `myproject.mylocation.myconnection`
OPTIONS (
format = 'ICEBERG',
uris = ["gs://mybucket/mydata/mytable/metadata/iceberg.metadata.json"]
)注意:
此方法需要你在表发生更新时手动更新最新元数据,因此 Google 建议使用 BigLake Metastore 来创建 Iceberg BigLake 表。请参阅 同步到 BigLake Metastore 指南了解具体步骤。
重要:适用于 Hudi 源格式到 Iceberg 目标格式的用例
- Hudi 扩展提供了在使用 Hudi 写入时向 Parquet schema 添加字段 ID 的能力。这是某些引擎(如 BigQuery 和 Snowflake)读取 Iceberg 表时的要求。如果你不打算使用 Iceberg,则无需在 Hudi 写入器中添加这些内容。
- 为了避免插入操作走 row writer,我们需要手动禁用它。对 row writer 的支持即将推出。
为 Hudi 写入器添加额外配置的步骤:
将扩展 jar 包(
xtable-hudi-extensions-0.5.0-SNAPSHOT-bundled.jar)添加到你的类路径中
例如,如果你正在使用 Hudi 针对 Spark 的快速入门指南,只需在命令末尾添加--jars xtable-hudi-extensions-0.5.0-SNAPSHOT-bundled.jar即可。在写入器选项中设置以下配置:
shell
hoodie.avro.write.support.class: org.apache.xtable.hudi.extensions.HoodieAvroWriteSupportWithFieldIds hoodie.client.init.callback.classes: org.apache.xtable.hudi.extensions.AddFieldIdsClientInitCallback hoodie.datasource.write.row.writer.enable : false运行你现有的使用 Hudi 写入器的代码
使用 BigLake Metastore 创建 Iceberg BigLake 表:
你可以通过两种方式将 Apache XTable™ (Incubating) 同步的 Iceberg 表注册到 BigLake Metastore:
- 要直接将 Apache XTable™ (Incubating) 同步的 Iceberg 表注册到 BigLake Metastore,请遵循 Apache XTable™ 与 BigLake Metastore 集成指南
- 在 BigQuery 上使用 Spark 存储过程 将表注册到 BigLake Metastore,并从 BigQuery 查询这些表。
Hudi 和 Delta 表
本文档 介绍了如何通过 manifest 文件查询 Hudi 和 Delta 表格式。
评论
登录后参与评论
正在加载评论…
KnowForge