Iceberg
Apache Iceberg 是一个面向海量分析型数据集的开放表格式。它旨在改进 Hive 等传统表格式的局限,并提供模式演进、隐藏分区、时间旅行等特性。
Iceberg 使用 Ozone 存储层来构建可扩展的数据湖仓,作为表数据与元数据的持久化记录系统。Ozone 原生的原子重命名能力满足了 Iceberg 的原子提交需求,无需外部锁定服务即可为数据管理提供强一致性。Ozone 处理海量对象的能力以及其强一致性模型(通过 Ratis 实现),使其成为 Iceberg 事务化、基于快照的结构的合适且可靠的后端。
关键集成细节
- 存储与元数据管理: Iceberg 将数据文件和元数据文件(manifest、快照)直接存储在 Ozone 上。
- 原子操作: Ozone 支持 Iceberg 提交流程所需的原子操作,确保并发写入时的数据一致性。
- 性能: 二者的结合可实现 PB 级分析与快速查询规划,克服传统 HDFS NameNode 的可扩展性瓶颈。
- 兼容性: Iceberg 通过 S3 兼容 API 或 Hadoop FileSystem 接口与 Ozone 交互,实现无缝集成。
快速上手
本教程介绍如何借助 S3 Gateway 和 Docker Compose,开始在 Apache Iceberg 中使用 Apache Ozone。
快速上手环境
非安全模式的 Ozone 与 Iceberg 集群。
Ozone S3G 启用了基于子域名
s3.ozone的虚拟主机风格寻址。- 该子域名以及包含桶名的子域名
warehouse.s3.ozone均映射到 S3 Gateway。
- 该子域名以及包含桶名的子域名
Iceberg 通过 S3 Gateway 访问 Ozone。
第 1 步 — 创建 Ozone 服务的 docker-compose.yaml
创建一个 docker-compose.yaml 文件,内容如下,用于:
启动一个单 Datanode 的 Ozone 集群
启动 S3 Gateway,并带上 Iceberg 所需的配置
- 启动前先等待 OM 就绪
- 启动时创建
warehouse桶 - 只有在桶创建完成后才将 S3 Gateway 标记为健康状态,因为这是 Iceberg 容器的前置条件。
- 定义并将桶子域名映射到 S3 Gateway。
# Licensed to the Apache Software Foundation (ASF) under one
# or more contributor license agreements. See the NOTICE file
# distributed with this work for additional information
# regarding copyright ownership. The ASF licenses this file
# to you under the Apache License, Version 2.0 (the
# "License"); you may not use this file except in compliance
# with the License. You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
x-image:
&image
image: ${OZONE_IMAGE:-apache/ozone}:${OZONE_IMAGE_VERSION:-2.1.0}${OZONE_IMAGE_FLAVOR:-}
x-common-config:
&common-config
OZONE-SITE.XML_hdds.datanode.dir: "/data/hdds"
OZONE-SITE.XML_ozone.metadata.dirs: "/data/metadata"
OZONE-SITE.XML_ozone.om.address: "om"
OZONE-SITE.XML_ozone.om.http-address: "om:9874"
OZONE-SITE.XML_ozone.recon.address: "recon:9891"
OZONE-SITE.XML_ozone.recon.db.dir: "/data/metadata/recon"
OZONE-SITE.XML_ozone.replication: "1"
OZONE-SITE.XML_ozone.scm.block.client.address: "scm"
OZONE-SITE.XML_ozone.scm.client.address: "scm"
OZONE-SITE.XML_ozone.scm.datanode.id.dir: "/data/metadata"
OZONE-SITE.XML_ozone.scm.names: "scm"
no_proxy: "om,recon,scm,s3g,localhost,127.0.0.1"
OZONE-SITE.XML_hdds.scm.safemode.min.datanode: "1"
OZONE-SITE.XML_hdds.scm.safemode.healthy.pipeline.pct: "0"
OZONE-SITE.XML_ozone.s3g.domain.name: "s3.ozone"
version: "3"
services:
datanode:
<<: *image
ports:
- 9864
command: ["ozone","datanode"]
environment:
<<: *common-config
networks:
iceberg_net:
om:
<<: *image
ports:
- 9874:9874
environment:
<<: *common-config
CORE-SITE.XML_hadoop.proxyuser.hadoop.hosts: "*"
CORE-SITE.XML_hadoop.proxyuser.hadoop.groups: "*"
ENSURE_OM_INITIALIZED: /data/metadata/om/current/VERSION
WAITFOR: scm:9876
command: ["ozone","om"]
networks:
iceberg_net:
scm:
<<: *image
ports:
- 9876:9876
environment:
<<: *common-config
ENSURE_SCM_INITIALIZED: /data/metadata/scm/current/VERSION
command: ["ozone","scm"]
networks:
iceberg_net:
recon:
<<: *image
ports:
- 9888:9888
environment:
<<: *common-config
command: ["ozone","recon"]
networks:
iceberg_net:
s3g:
<<: *image
ports:
- 9878:9878
environment:
<<: *common-config
WAITFOR: om:9874
command:
- sh
- -c
- |
set -e
ozone s3g &
s3g_pid=$$!
until ozone sh volume list >/dev/null 2>&1; do echo '...waiting...' && sleep 1; done;
ozone sh bucket delete /s3v/warehouse || true
ozone sh bucket create /s3v/warehouse
wait "$$s3g_pid"
healthcheck:
test: [ "CMD", "ozone", "sh", "bucket", "info", "/s3v/warehouse" ]
interval: 5s
timeout: 3s
retries: 10
start_period: 30s
networks:
iceberg_net:
aliases:
- s3.ozone
- warehouse.s3.ozone第 2 步 — 为 Iceberg 服务创建 iceberg-spark.yml
services:
spark-iceberg:
image: tabulario/spark-iceberg
container_name: spark-iceberg
build: spark/
networks:
iceberg_net:
depends_on:
rest:
condition: service_started
s3g:
condition: service_healthy
volumes:
- ./warehouse:/home/iceberg/warehouse
environment:
- AWS_ACCESS_KEY_ID=admin
- AWS_SECRET_ACCESS_KEY=password
- AWS_REGION=us-east-1
ports:
- 8888:8888
- 8080:8080
- 10000:10000
- 10001:10001
rest:
image: apache/iceberg-rest-fixture
container_name: iceberg-rest
networks:
iceberg_net:
ports:
- 8181:8181
environment:
- AWS_ACCESS_KEY_ID=admin
- AWS_SECRET_ACCESS_KEY=password
- AWS_REGION=us-east-1
- CATALOG_WAREHOUSE=s3://warehouse/
- CATALOG_IO__IMPL=org.apache.iceberg.aws.s3.S3FileIO
- CATALOG_S3_ENDPOINT=http://s3.ozone:9878
networks:
iceberg_net:第 3 步 —— 同时启动 Iceberg 与 Ozone
将 docker-compose.yaml(用于 Ozone)和 docker-compose-flink.yml(用于 Flink)放在同一目录下后,你就可以使用以下命令同时启动这两个服务,并让它们共享同一个网络:
export COMPOSE_FILE=docker-compose.yaml:iceberg-spark.yml
docker compose up -d验证容器正在运行:
docker ps步骤 4 — 启动 Spark SQL 客户端
docker exec -it spark-iceberg spark-sql你现在应该位于:
spark-sql ()>第 5 步 —— 创建并查询由 Ozone S3 支撑的表
在 Ozone S3 中创建一个 Iceberg 表:
CREATE NAMESPACE IF NOT EXISTS demo.nyc;
CREATE TABLE demo.nyc.taxis
(
vendor_id bigint,
trip_id bigint,
trip_distance float,
fare_amount double,
store_and_fwd_flag string
)
PARTITIONED BY (vendor_id);向表中插入数据:
INSERT INTO demo.nyc.taxis
VALUES (1, 1000371, 1.8, 15.32, 'N'), (2, 1000372, 2.5, 22.15, 'N'), (2, 1000373, 0.9, 9.01, 'N'), (1, 1000374, 8.4, 42.13, 'Y');查询该表:
SELECT * FROM demo.nyc.taxis;验证数据文件已存储在 Ozone S3 中:
docker compose exec -it s3g ozone fs -ls -R ofs://om/s3v/warehouseSpark UI 的访问地址为 http://localhost:8080,你可以在此监控 Spark 作业。
评论
登录后参与评论
KnowForge