数据目录

Google BigQuery

师成师成· 更新于 2026-09-29· 阅读 12 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

Hudi 表可以作为外部表通过 Google Cloud BigQuery 进行查询。目前,Hudi 与 BigQuery 的集成仅支持 hive 风格分区的写时复制(Copy-On-Write)表和读优化(Read-Optimized)的写时合并(Merge-On-Read)表。

同步模式

Manifest 文件

从 0.14.0 版本开始,BigQuerySyncTool 支持通过 manifest 将表同步到 BigQuery。首次运行时,该工具会根据提供的配置创建一个表示表中当前基础文件的 manifest 文件,以及 BigQuery 中的一张表。此后每次运行都会生成新的 manifest 文件,并且当 Hudi 表的 schema 发生变化时,会更新 BigQuery 中表的 schema。

使用新 manifest 方式的优势:

  1. 只会扫描 manifest 中列出的文件,从而降低查询成本并提升查询性能
  2. schema 现在从 Hudi 提交元数据同步,支持正常的 schema 演进
  3. 在 BigQuery 中查询时,list 类型不再产生不必要的嵌套,因为默认启用了 list 推断
  4. 由于新的 schema 处理改进,分区列不再需要从文件中剔除

要启用此功能,请将 hoodie.gcp.bigquery.sync.use_bq_manifest_file 设置为 true。

基于文件的视图(旧版)

这是当前的默认行为,用于在用户升级到 0.14.0 及更高版本时保持兼容性。
运行后,同步工具会在 BigQuery 的目标数据集中创建 2 张表和 1 个视图。这些表和视图共享相同的名称前缀,该前缀取自 Hudi 表名。查询该视图可获得与查询写时复制 Hudi 表相同的结果。
注意: 该视图会扫描表基础路径下的所有 parquet 文件,因此建议升级到基于 manifest 的方式,以获得更好的成本和性能表现。

配置

Hudi 使用 org.apache.hudi.gcp.bigquery.BigQuerySyncTool 来同步表。通过设置同步工具类,它可以与 Hudi Streamer 配合使用。需要一些 BigQuery 专用的配置项。

配置说明
hoodie.gcp.bigquery.sync.project_id目标 Google Cloud 项目
hoodie.gcp.bigquery.sync.dataset_nameBigQuery 数据集名称;请在运行同步工具前创建
hoodie.gcp.bigquery.sync.dataset_location数据集的区域信息;与存储 Hudi 表的 GCS 存储桶相同
hoodie.gcp.bigquery.sync.source_uri指向第一级分区的通配符路径模式;可显式指定分区键,也可自动推断。仅分区表需要该配置
hoodie.gcp.bigquery.sync.source_uri_prefixsource_uri 的公共前缀,通常是 Hudi 表的路径,结尾斜杠有无均可。
hoodie.gcp.bigquery.sync.base_pathHudi 表通常使用的 basepath 配置。
hoodie.gcp.bigquery.sync.use_bq_manifest_file设置为 true 以启用基于清单文件的同步
hoodie.gcp.bigquery.sync.require_partition_filter该配置在 Hudi 0.14.1 版本引入,接受 BOOLEAN 值,默认为 false。启用后(设置为 true),所有查询都必须创建分区过滤条件(WHERE 子句),且针对分区表的分区列。缺少此类过滤条件的查询将导致错误。

请参阅 org.apache.hudi.gcp.bigquery.BigQuerySyncConfig 以获取完整的配置列表。

分区处理

除了 BigQuery 专有的配置之外,你还需要使用 hive 风格的分区方式,以便在 BigQuery 中实现分区裁剪。除此之外,分区路径中的值将成为查询时该字段返回的值。例如,如果按一个以毫秒为单位的时间字段 time 进行分区,且输出格式为 time=yyyy-MM-dd,那么查询返回的 time 值将是天级别的粒度,而不是原始的毫秒值,因此在建表时请务必注意这一点。

hoodie.datasource.write.hive_style_partitioning = 'true'

对于基于视图的同步,你还必须指定以下配置:

hoodie.datasource.write.drop.partition.columns  = 'true'
hoodie.partition.metafile.use.base.format       = 'true'

示例

下面展示了结合 Hudi Streamer 运行 BigQuerySyncTool 的一个示例。

spark-submit --master yarn \
--packages com.google.cloud:google-cloud-bigquery:2.10.4 \
--jars "/opt/hudi-gcp-bundle-1.2.0.jar,/opt/hudi-utilities-slim-bundle_2.12-1.2.0.jar,/opt/hudi-spark3.5-bundle_2.12-1.2.0.jar" \
--class org.apache.hudi.utilities.streamer.HoodieStreamer \
/opt/hudi-utilities-slim-bundle_2.12-1.2.0.jar \
--target-base-path gs://my-hoodie-table/path \
--target-table mytable \
--table-type COPY_ON_WRITE \
--base-file-format PARQUET \
# ... other Hudi Streamer options
--enable-sync \
--sync-tool-classes org.apache.hudi.gcp.bigquery.BigQuerySyncTool \
--hoodie-conf hoodie.streamer.source.dfs.root=gs://my-source-data/path \
--hoodie-conf hoodie.gcp.bigquery.sync.project_id=hudi-bq \
--hoodie-conf hoodie.gcp.bigquery.sync.dataset_name=rxusandbox \
--hoodie-conf hoodie.gcp.bigquery.sync.dataset_location=asia-southeast1 \
--hoodie-conf hoodie.gcp.bigquery.sync.table_name=mytable \
--hoodie-conf hoodie.gcp.bigquery.sync.base_path=gs://rxusandbox/testcases/stocks/data/target/${NOW} \
--hoodie-conf hoodie.gcp.bigquery.sync.partition_fields=year,month,day \
--hoodie-conf hoodie.gcp.bigquery.sync.source_uri=gs://my-hoodie-table/path/year=* \
--hoodie-conf hoodie.gcp.bigquery.sync.source_uri_prefix=gs://my-hoodie-table/path/ \
--hoodie-conf hoodie.gcp.bigquery.sync.use_file_listing_from_metadata=true \
--hoodie-conf hoodie.gcp.bigquery.sync.assume_date_partitioning=false \
--hoodie-conf hoodie.datasource.hive_sync.mode=jdbc \
--hoodie-conf hoodie.datasource.hive_sync.jdbcurl=jdbc:hive2://localhost:10000 \
--hoodie-conf hoodie.datasource.hive_sync.skip_ro_suffix=true \
--hoodie-conf hoodie.datasource.hive_sync.ignore_exceptions=false \
--hoodie-conf hoodie.datasource.hive_sync.database=mydataset \
--hoodie-conf hoodie.datasource.hive_sync.table=mytable \
--hoodie-conf hoodie.datasource.write.recordkey.field=mykey \
--hoodie-conf hoodie.datasource.write.partitionpath.field=year,month,day \
--hoodie-conf hoodie.table.ordering.fields=ts \
--hoodie-conf hoodie.datasource.write.keygenerator.type=COMPLEX \
--hoodie-conf hoodie.datasource.write.hive_style_partitioning=true \
--hoodie-conf hoodie.datasource.write.drop.partition.columns=true \
--hoodie-conf hoodie.partition.metafile.use.base.format=true \

评论

登录后参与评论

正在加载评论…