工具
Trino 提取器
登录后可跨设备保存划线和私人笔记登录
Trino 提取器
概述
Trino 提取器是一个全面的元数据提取工具,用于将 Apache Atlas 与 Trino 集成。它提供对 Trino 元数据(包括目录、模式、表和列)的发现、提取和同步功能,将其导入 Apache Atlas,以增强数据治理和元数据管理能力。
主要特性
元数据提取
- 全面发现:自动发现并提取 Trino 目录、模式、表和列的元数据
- 基于 JDBC 的连接:使用标准的 Trino JDBC 驱动程序,确保连接可靠
- 选择性提取:支持针对特定目录、模式或表名进行提取
Atlas 集成
- 实体管理:为 Trino 元数据对象创建和更新 Atlas 实体
- 关系映射:建立目录、模式、表和列之间的正确层级关系
- 同步:通过移除 Trino 中已不存在的 Atlas 实体来保持一致性
- 连接器支持:对 Atlas 通过各自 Hook(如 Hive、Iceberg)捕获元数据的 Trino 连接器进行专门处理
调度与自动化
- 基于 Cron 的调度:支持使用 cron 表达式进行自动化的周期性提取
- 一次性执行:可以作为单次提取任务运行
- 错误处理:具备健壮的错误处理机制和详细的日志记录
架构
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Trino Cluster │ │ Trino Extractor │ │ Apache Atlas │
│ │ │ │ │ │
│ ┌─────────────┐ │ │ ┌─────────────┐ │ │ ┌─────────────┐ │
│ │ Catalogs │◄┼────┼►│ JDBC Client │ │ │ │ Entities │ │
│ └─────────────┘ │ │ └─────────────┘ │ │ └─────────────┘ │
│ ┌─────────────┐ │ │ ┌─────────────┐ │ │ ┌─────────────┐ │
│ │ Schemas │ │ │ │ Extraction │◄┼────┼►│Relationships│ │
│ └─────────────┘ │ │ │ Service │ │ │ └─────────────┘ │
│ ┌─────────────┐ │ │ └─────────────┘ │ │ ┌─────────────┐ │
│ │ Tables │ │ │ ┌─────────────┐ │ │ │ Lineage │ │
│ └─────────────┘ │ │ │Atlas Client │◄┼────┼►│ Data │ │
│ ┌─────────────┐ │ │ └─────────────┘ │ │ └─────────────┘ │
│ │ Columns │ │ │ │ │ │
│ └─────────────┘ │ └─────────────────┘ └─────────────────┘
└─────────────────┘ 快速开始
1. 配置设置
配置 atlas-trino-extractor.properties 文件:
# Atlas connection
atlas.rest.address=http://localhost:21000/
# Trino connection
atlas.trino.jdbc.address=jdbc:trino://localhost:8080/
atlas.trino.jdbc.user=your-username
# Catalogs to extract
atlas.trino.catalogs.registered=hive_catalog,iceberg_catalog2. 基本执行
# Extract all registered catalogs
./bin/run-trino-extractor.sh
# Extract specific catalog
./bin/run-trino-extractor.sh -c my_catalog
# Schedule periodic extraction (every 6 hours)
./bin/run-trino-extractor.sh -cx "0 0 */6 * * ?"配置属性
| 属性 | 说明 | 默认值 | 示例 |
|---|---|---|---|
atlas.rest.address | Atlas REST API 端点 | http://localhost:21000/ | https://atlas.company.com:21443/ |
atlas.trino.jdbc.address | Trino JDBC URL | - | jdbc:trino://trino-server:8080/ |
atlas.trino.jdbc.user | Trino 用户名 | - | admin |
atlas.trino.jdbc.password | Trino 密码 | "" | password123 |
atlas.trino.namespace | Trino 实例命名空间 | cm | production-cluster |
atlas.trino.catalogs.registered | 要抽取的 Catalog | - | hive,iceberg,mysql |
atlas.trino.catalog.hook.enabled.<catalog-name> | 是否为该 Catalog 在 Atlas 中启用 Hook? | false | true |
atlas.trino.catalog.hook.enabled.<catalog-name>.namespace | 该 Hook 在 Atlas 下的命名空间 | cm | cm |
atlas.trino.extractor.schedule | Cron 表达式 | - | 0 0 2 * * ? |
命令行用法
可用选项
| 选项 | 长格式 | 说明 | 示例 |
|---|---|---|---|
-c | --catalog | 抽取指定的 catalog | -c hive_catalog |
-s | --schema | 抽取指定的 schema | -s sales_data |
-t | --table | 抽取指定的表 | -t customer_orders |
-cx | --cronExpression | 使用 cron 表达式进行调度 | -cx "0 0 2 * * ?" |
-h | --help | 显示帮助信息 | -h |
特定连接器的处理
例如:Hive 连接器集成
# Enable Hive hook integration
atlas.trino.catalog.hook.enabled.hive_catalog=true
atlas.trino.catalog.hook.enabled.hive_catalog.namespace=cm优势:
- 将 Trino 实体与现有的 Hive 实体关联起来
- 保持 Hive 与 Trino 元数据之间的一致性
- 支持启用了 Atlas Hive Hook 的环境
常见问题
故障排查
问:为什么有些实体没有出现在 Atlas 中?
答:请检查目录(catalog)注册情况、权限以及网络连通性,并查看日志以获取具体的错误信息。
问:如何处理拥有数千张表的大型集群?
答:可以使用选择性抽取、增加内存分配、在非高峰时段进行调度,并逐个处理各个目录(catalog)。
文档
评论
登录后参与评论
正在加载评论…
KnowForge