Apache Paimon(孵化中)
Apache Paimon(孵化中)是一个流式数据湖平台,支持高速数据摄入、变更数据追踪以及高效的实时分析。
提示
本文假定你已经掌握了 Apache Paimon(孵化中) 的基础知识和操作。关于本文中未提及的 Apache Paimon(孵化中)相关知识,你可以查阅其官方文档。
通过使用 Kyuubi,我们可以对 Apache Paimon(孵化中)运行 SQL 查询,这比直接使用 Trino 操作 Apache Paimon(孵化中)更方便、更易理解,也更易于扩展。
Apache Paimon(孵化中)集成
要启用 Kyuubi Trino SQL 引擎与 Apache Paimon(孵化中)的集成,你需要:
依赖项
支持 Apache Paimon(孵化中)的 Kyuubi Trino SQL 引擎的 classpath 由以下部分组成:
kyuubi-trino-sql-engine-1.9.1_2.12.jar,随 Kyuubi 发行版部署的引擎 jar- 一份 Trino 发行版的拷贝
paimon-trino-<version>.jar(例如:paimon-trino-0.2.jar),其源码可在源代码中找到flink-shaded-hadoop-2-uber-<version>.jar,其源码可在预打包 Hadoop 中找到
为了使 Apache Paimon(孵化中)的包对引擎的运行时 classpath 可见,你需要:
- 参照 Apache Paimon(孵化中)Trino README 构建
paimon-trino-<version>.jar - 将
paimon-trino-<version>.jar和flink-shaded-hadoop-2-uber-<version>.jar包直接放入$TRINO_SERVER_HOME/plugin/tablestore目录
警告
请注意不同 Apache Paimon(孵化中)与 Trino 版本之间的兼容性,可以在 Apache Paimon(孵化中)多引擎支持页面上确认。
配置
要启用 Apache Paimon(孵化中)的功能,我们可以设置以下配置:
Catalog 通过在 $TRINO_SERVER_HOME/etc/catalog 目录中创建 catalog 属性文件来注册。例如,创建 $TRINO_SERVER_HOME/etc/catalog/tablestore.properties 并写入以下内容,即可将 tablestore 连接器挂载为 tablestore catalog:
connector.name=tablestore
warehouse=file:///tmp/warehouseApache Paimon(孵化中)操作
Apache Paimon(孵化中)支持通过 Trino 读取表存储(table store)中的表。一个常见的场景是使用 Spark 或 Flink 写入数据,然后使用 Trino 读取数据。您可以按照文档 Apache Paimon(孵化中)引擎 Flink 快速入门 向表存储表中写入数据,然后使用 kyuubi trino sql engine 通过以下 SELECT SQL 语句查询该表。
SELECT * FROM tablestore.default.t1评论
登录后参与评论
KnowForge