Hive Metastore
本文档将逐步介绍如何在 Hive Metastore (HMS) 上注册一个由 Apache XTable™ (Incubating) 同步的表。
前提条件
- 源表(Hudi/Delta/Iceberg)已写入本地存储或 S3/GCS/ADLS 等外部存储位置。如果源表尚未写入,可以按照本教程中的步骤进行设置。
- 一台可以运行 Apache Spark 的计算实例。可以是本地机器、Docker,也可以是 Amazon EMR、Google Cloud Dataproc、Azure HDInsight 等分布式系统。这是使用 Spark 客户端在 HMS 中注册表的必要步骤。
- 克隆 XTable™ (Incubating)代码仓库,并按照安装页面上的步骤生成
xtable-utilities_2.12-0.5.0-SNAPSHOT-bundled.jar。 - 本指南还假设您已在本地或 EMR/Dataproc/HDInsight 上配置好 Hive Metastore,并且其正在运行。
步骤
运行同步
在克隆的 Apache XTable™ (Incubating) 目录中创建 my_config.yaml。
- targetFormat: HUDI
- targetFormat: DELTA
- targetFormat: ICEBERG
yaml
sourceFormat: DELTA|ICEBERG # choose only one
targetFormats:
- HUDI
datasets:
-
tableBasePath: file:///path/to/source/data
tableName: table_name注意:
请将
sourceFormat、tableBasePath和tableName字段替换为适当的值。如果你的源表位于 S3/GCS/ADLS,请将
file:///path/to/source/data替换为相应的源数据路径,即:- S3 -
s3://path/to/source/data - GCS -
gs://path/to/source/data,或 - ADLS -
abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<path-to-data>
- S3 -
在克隆的 Apache XTable™ (Incubating) 目录下打开终端,运行以下命令执行同步过程。
shell
java -jar xtable-utilities/target/xtable-utilities_2.12-0.5.0-SNAPSHOT-bundled.jar --datasetConfig my_config.yaml注意:
此时,如果你查看存储桶路径,将会看到 .hoodie 或 _delta_log 或 metadata 目录,其中包含相关的元数据文件,这些文件可帮助查询引擎将数据识别为 Hudi/Delta/Iceberg 表。
在 Hive Metastore 中注册目标表
现在,你需要在 Hive Metastore 中注册通过 Apache XTable™ (Incubating) 同步的目标表。
- targetFormat: HUDI
- targetFormat: DELTA
- targetFormat: ICEBERG
Hudi 表可以直接通过 Hive Sync Tool 同步到 Hive Metastore,随后由不同的查询引擎进行查询。有关 Hive Sync Tool 的更多信息,请参阅 Hudi Hive Metastore 文档。
shell
cd $HUDI_HOME/hudi-sync/hudi-hive-sync
./run_sync_tool.sh \
--jdbc-url <jdbc_url> \
--user <username> \
--pass <password> \
--partitioned-by <partition_field> \
--base-path <'/path/to/synced/hudi/table'> \
--database <database_name> \
--table <tableName>注意:
如果您的源数据位于 S3/GCS/ADLS,请将 file:///path/to/source/data 替换为相应的源数据路径,即
- S3 -
s3://path/to/source/data - GCS -
gs://path/to/source/data或 - ADLS -
abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<path-to-data>
现在,您可以在同一个 spark 会话中直接将创建的表作为 Hudi 表进行查询,也可以使用 Presto 和/或 Trino 等查询引擎进行查询。有关更多信息,请查阅 Presto 或 Trino 查询引擎的 Apache XTable™ (Incubating) 同步表查询指南。
sql
SELECT * FROM <database_name>.<table_name>;结论
在本指南中,我们介绍了如何:
- 使用 Apache XTable™ (Incubating) 同步源表,为目标表格式生成元数据
- 在 Hive Metastore 中将目标表格式的数据登记到数据目录
- 使用 Spark 查询目标表
评论
登录后参与评论
KnowForge