目录服务

Glue Data Catalog

师成师成· 更新于 2026-09-28· 阅读 6 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

本文档将引导你在 AWS 上的 Glue Data Catalog 中注册一个 Apache XTable™ (Incubating) 同步表的步骤。

前提条件

  1. 源表(Hudi/Delta/Iceberg)已经写入 Amazon S3。如果源表尚未写入 S3,可以按照本教程中的步骤进行设置
  2. 设置从命令行访问 AWS API 的权限。如果尚未安装 AWSCLIv2,可以按照 AWS 文档中列出的步骤进行安装,并按照此处的步骤设置访问凭证
  3. 克隆 Apache XTable™ (Incubating)代码仓库,并按照安装页面上的步骤生成 xtable-utilities_2.12-0.5.0-SNAPSHOT-bundled.jar

步骤

运行同步

在克隆的 Apache XTable™ (Incubating) 目录中创建 my_config.yaml。

  • targetFormat: HUDI
  • targetFormat: DELTA
  • targetFormat: ICEBERG

yaml

sourceFormat: DELTA|ICEBERG # choose only one
targetFormats:
  - HUDI
datasets:
  -
    tableBasePath: s3://path/to/source/data
    tableName: table_name

::: 注意

请将 sourceFormat、tableBasePath 和 tableName 字段替换为合适的值。

:::

在终端中进入克隆的 xtable 目录,使用以下命令运行同步过程。

shell

java -jar xtable-utilities/target/xtable-utilities_2.12-0.5.0-SNAPSHOT-bundled.jar --datasetConfig my_config.yaml

注意:

此时,如果你检查存储桶路径,就能看到 .hoodie 或 _delta_log 或 metadata 目录,其中包含元数据文件,这些文件所包含的信息有助于查询引擎将数据作为目标表进行解析。

在 Glue Data Catalog 中注册目标表

在终端中,创建一个 glue 数据库。

shell

aws glue create-database --database-input "{\"Name\":\"xtable_synced_db\"}"

在终端中创建一个 Glue 爬取器(crawler)。请将 <yourAccountId>、<yourRoleName> 和 <path/to/your/data> 修改为适当的值。

shell

export accountId=<yourAccountId>
export roleName=<yourRoleName>
export s3DataPath=s3://<path/to/source/data>
  • targetFormat: HUDI
  • targetFormat: DELTA
  • targetFormat: ICEBERG

shell

aws glue create-crawler --name xtable_crawler --role arn:aws:iam::${accountId}:role/service-role/${roleName} --database xtable_synced_db --targets "{\"HudiTargets\":[{\"Paths\":[\"${s3DataPath}\"]}]}"

在终端中,运行 Glue 爬虫。

 aws glue start-crawler --name xtable_crawler

爬虫成功后,你就可以从 Athena、EMR 和/或 Redshift 查询引擎中查询这个 Iceberg 表。

  • targetFormat:HUDI
  • targetFormat:DELTA
  • targetFormat:ICEBERG

Hudi 目标格式的限制:

要验证 Hudi targetFormat 表的查询结果,你需要确保所使用的查询引擎支持 Hudi 0.14.0 版本,详见此处。

结语

本指南介绍了如何:

  1. 使用 Apache XTable™ (Incubating) 将源表同步为目标表格式的元数据;
  2. 在 Glue Data Catalog 中为目标表格式的数据建立目录;
  3. 使用 Amazon Athena 查询目标表。

评论

登录后参与评论

正在加载评论…