从 Python 访问 Apache Ozone
Apache Ozone 项目本身并不提供 Python 客户端库。不过,有多个第三方开源库可用于构建应用程序,通过不同接口访问 Ozone 集群:ofs 文件系统、Ozone HttpFS REST API 以及 Ozone S3。
本文档概述了这些方法,并提供了简明的配置说明和经过验证的代码示例。
安装配置与前置条件
开始之前,请确保满足以下条件:
已安装 Python(推荐 3.x 版本)
Apache Ozone 已配置并可访问
使用 PyArrow 与 libhdfs 时:
- PyArrow 库(
pip install pyarrow) - 已配置 Hadoop 原生库,并指定了 Ozone classpath(详见下文)
- PyArrow 库(
使用 S3 访问时:
- Boto3(
pip install boto3) - Ozone S3 Gateway 端点和存储桶名称
- 访问凭证(类似 AWS 的密钥和访问密钥)
- Boto3(
使用 HttpFS 访问时:
- Requests(
pip install requests)或 fsspec(pip install fsspec)
- Requests(
方法一:通过 PyArrow 和 libhdfs 访问 Ozone
该方法利用 PyArrow 的 HadoopFileSystem API,它依赖于 libhdfs.so 原生库。libhdfs.so 并未包含在 PyArrow 中,必须单独从 Hadoop 获取。
配置
请确保 Ozone 配置文件(core-site.xml 和 ozone-site.xml)可用,并且已设置 OZONE_CONF_DIR。同时确保 ARROW_LIBHDFS_DIR 和 CLASSPATH 已正确设置。
例如,
export ARROW_LIBHDFS_DIR=hadoop-3.4.0/lib/native/
export CLASSPATH=$(ozone classpath ozone-tools)代码示例
import pyarrow.fs as pafs
# Connect to Ozone using HadoopFileSystem
# "default" tells PyArrow to use the fs.defaultFS property from core-site.xml
fs = pafs.HadoopFileSystem("default")
# Create a directory inside the bucket
fs.create_dir("volume/bucket/aaa")
# Write data to a file
path = "volume/bucket/file1"
with fs.open_output_stream(path) as stream:
stream.write(b'data')note
在 core-site.xml 中将 fs.defaultFS 配置为指向 Ozone 集群。例如,
<configuration>
<property>
<name>fs.defaultFS</name>
<value>ofs://om:9862</value>
<description>Ozone Manager endpoint</description>
</property>
</configuration>自己试试!查看 PyArrow 教程,快速上手使用 Ozone 的 Docker 镜像。
方法二:通过 Boto3 和 S3 Gateway 访问 Ozone
配置
- 确定你的 Ozone S3 Gateway 端点(例如
http://s3g:9878)。 - 使用 AWS 兼容的凭据(来自 Ozone)。
代码示例
import boto3
# Create a local file to upload
with open("localfile.txt", "w") as f:
f.write("Hello from Ozone via Boto3!\n")
# Configure Boto3 client
s3 = boto3.client(
's3',
endpoint_url='http://s3g:9878',
aws_access_key_id='ozone-access-key',
aws_secret_access_key='ozone-secret-key'
)
# List buckets
response = s3.list_buckets()
print("Buckets:", response['Buckets'])
# Upload the file
s3.upload_file('localfile.txt', 'bucket', 'file.txt')
print("Uploaded 'localfile.txt' to 'bucket/file.txt'")
# Download the file back
s3.download_file('bucket', 'file.txt', 'downloaded.txt')
print("Downloaded 'file.txt' as 'downloaded.txt'")::: note
请将端点 URL、凭据和存储桶名称替换为你的实际配置。
:::
自己动手试试!
快速上手请参阅 Boto3 教程,使用 Ozone 的 Docker 镜像进行体验。
方法 3:通过 HttpFS REST API 访问 Ozone
首先,安装 requests Python 模块:
pip install requests配置
- 使用 Ozone 的 HttpFS 端点(例如
http://httpfs:14000)。
代码示例(requests)
#!/usr/bin/python
import requests
# Ozone HTTPFS endpoint and file path
host = "http://httpfs:14000"
volume = "vol1"
bucket = "bucket1"
filename = "hello.txt"
path = f"/webhdfs/v1/{volume}/{bucket}/{filename}"
user = "ozone" # can be any value in simple auth mode
# Step 1: Initiate file creation (responds with 307 redirect)
params_create = {
"op": "CREATE",
"overwrite": "true",
"user.name": user
}
print("Creating file...")
resp_create = requests.put(host + path, params=params_create, allow_redirects=False)
if resp_create.status_code != 307:
print(f"Unexpected response: {resp_create.status_code}")
print(resp_create.text)
exit(1)
redirect_url = resp_create.headers['Location']
print(f"Redirected to: {redirect_url}")
# Step 2: Write data to the redirected location with correct headers
headers = {"Content-Type": "application/octet-stream"}
content = b"Hello from Ozone HTTPFS!\n"
resp_upload = requests.put(redirect_url, data=content, headers=headers)
if resp_upload.status_code != 201:
print(f"Upload failed: {resp_upload.status_code}")
print(resp_upload.text)
exit(1)
print("File created successfully.")
# Step 3: Read the file back
params_open = {
"op": "OPEN",
"user.name": user
}
print("Reading file...")
resp_read = requests.get(host + path, params=params_open, allow_redirects=True)
if resp_read.ok:
print("File contents:")
print(resp_read.text)
else:
print(f"Read failed: {resp_read.status_code}")
print(resp_read.text)自己动手试试!查看使用 HttpFS REST API 访问 Ozone 教程,快速上手 Ozone 的 Docker 镜像。
代码示例(webhdfs)
首先,安装 fsspec Python 模块:
pip install fsspecfrom fsspec.implementations.webhdfs import WebHDFS
fs = WebHDFS(host='httpfs', port=14000, user='ozone')
# Read a file from /vol1/bucket1/hello.txt
file_path = "/vol1/bucket1/hello.txt"
with fs.open(file_path, mode='rb') as f:
content = f.read()
print("File contents:")
print(content.decode('utf-8'))::: note
请根据你的环境替换 host、port 和 path。
:::
故障排查提示
- 认证错误:请检查凭据以及 Kerberos 令牌(如使用)是否正确。
- 连接问题:请检查端点 URL、端口和防火墙规则。
- 文件系统错误:请确保 Ozone 配置正确且拥有相应权限。
- 缺少依赖:请安装所需的 Python 包(
pip install pyarrow boto3 requests fsspec)。
参考资料与延伸资源
评论
登录后参与评论
KnowForge