目录服务

Hive Metastore

师成师成· 更新于 2026-09-28· 阅读 6 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

本文档将逐步介绍如何在 Hive Metastore (HMS) 上注册一个由 Apache XTable™ (Incubating) 同步的表。

前提条件

  1. 源表(Hudi/Delta/Iceberg)已写入本地存储或 S3/GCS/ADLS 等外部存储位置。如果源表尚未写入,可以按照本教程中的步骤进行设置。
  2. 一台可以运行 Apache Spark 的计算实例。可以是本地机器、Docker,也可以是 Amazon EMR、Google Cloud Dataproc、Azure HDInsight 等分布式系统。这是使用 Spark 客户端在 HMS 中注册表的必要步骤。
  3. 克隆 XTable™ (Incubating)代码仓库,并按照安装页面上的步骤生成 xtable-utilities_2.12-0.5.0-SNAPSHOT-bundled.jar。
  4. 本指南还假设您已在本地或 EMR/Dataproc/HDInsight 上配置好 Hive Metastore,并且其正在运行。

步骤

运行同步

在克隆的 Apache XTable™ (Incubating) 目录中创建 my_config.yaml。

  • targetFormat: HUDI
  • targetFormat: DELTA
  • targetFormat: ICEBERG

yaml

sourceFormat: DELTA|ICEBERG # choose only one
targetFormats:
  - HUDI
datasets:
  -
    tableBasePath: file:///path/to/source/data
    tableName: table_name

注意:

  1. 请将 sourceFormat、tableBasePath 和 tableName 字段替换为适当的值。

  2. 如果你的源表位于 S3/GCS/ADLS,请将 file:///path/to/source/data 替换为相应的源数据路径,即:

    • S3 - s3://path/to/source/data
    • GCS - gs://path/to/source/data,或
    • ADLS - abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<path-to-data>

在克隆的 Apache XTable™ (Incubating) 目录下打开终端,运行以下命令执行同步过程。

shell

java -jar xtable-utilities/target/xtable-utilities_2.12-0.5.0-SNAPSHOT-bundled.jar --datasetConfig my_config.yaml

注意:

此时,如果你查看存储桶路径,将会看到 .hoodie 或 _delta_log 或 metadata 目录,其中包含相关的元数据文件,这些文件可帮助查询引擎将数据识别为 Hudi/Delta/Iceberg 表。

在 Hive Metastore 中注册目标表

现在,你需要在 Hive Metastore 中注册通过 Apache XTable™ (Incubating) 同步的目标表。

  • targetFormat: HUDI
  • targetFormat: DELTA
  • targetFormat: ICEBERG

Hudi 表可以直接通过 Hive Sync Tool 同步到 Hive Metastore,随后由不同的查询引擎进行查询。有关 Hive Sync Tool 的更多信息,请参阅 Hudi Hive Metastore 文档。

shell

cd $HUDI_HOME/hudi-sync/hudi-hive-sync

./run_sync_tool.sh  \
--jdbc-url <jdbc_url> \
--user <username> \
--pass <password> \
--partitioned-by <partition_field> \
--base-path <'/path/to/synced/hudi/table'> \
--database <database_name> \
--table <tableName>

注意:

如果您的源数据位于 S3/GCS/ADLS,请将 file:///path/to/source/data 替换为相应的源数据路径,即

  • S3 - s3://path/to/source/data
  • GCS - gs://path/to/source/data 或
  • ADLS - abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<path-to-data>

现在,您可以在同一个 spark 会话中直接将创建的表作为 Hudi 表进行查询,也可以使用 Presto 和/或 Trino 等查询引擎进行查询。有关更多信息,请查阅 Presto 或 Trino 查询引擎的 Apache XTable™ (Incubating) 同步表查询指南。

sql

SELECT * FROM <database_name>.<table_name>;

结论

在本指南中,我们介绍了如何:

  1. 使用 Apache XTable™ (Incubating) 同步源表,为目标表格式生成元数据
  2. 在 Hive Metastore 中将目标表格式的数据登记到数据目录
  3. 使用 Spark 查询目标表

评论

登录后参与评论

正在加载评论…