客户端接口

从 Python 访问 Apache Ozone

师成师成· 更新于 2026-09-28· 阅读 12 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

Apache Ozone 项目本身并不提供 Python 客户端库。不过,有多个第三方开源库可用于构建应用程序,通过不同接口访问 Ozone 集群:ofs 文件系统、Ozone HttpFS REST API 以及 Ozone S3。

本文档概述了这些方法,并提供了简明的配置说明和经过验证的代码示例。

安装配置与前置条件

开始之前,请确保满足以下条件:

  • 已安装 Python(推荐 3.x 版本)

  • Apache Ozone 已配置并可访问

  • 使用 PyArrow 与 libhdfs 时:

    • PyArrow 库(pip install pyarrow)
    • 已配置 Hadoop 原生库,并指定了 Ozone classpath(详见下文)
  • 使用 S3 访问时:

    • Boto3(pip install boto3)
    • Ozone S3 Gateway 端点和存储桶名称
    • 访问凭证(类似 AWS 的密钥和访问密钥)
  • 使用 HttpFS 访问时:

    • Requests(pip install requests)或 fsspec(pip install fsspec)

方法一:通过 PyArrow 和 libhdfs 访问 Ozone

该方法利用 PyArrow 的 HadoopFileSystem API,它依赖于 libhdfs.so 原生库。libhdfs.so 并未包含在 PyArrow 中,必须单独从 Hadoop 获取。

配置

请确保 Ozone 配置文件(core-site.xml 和 ozone-site.xml)可用,并且已设置 OZONE_CONF_DIR。同时确保 ARROW_LIBHDFS_DIR 和 CLASSPATH 已正确设置。

例如,

export ARROW_LIBHDFS_DIR=hadoop-3.4.0/lib/native/
export CLASSPATH=$(ozone classpath ozone-tools)

代码示例

import pyarrow.fs as pafs

# Connect to Ozone using HadoopFileSystem
# "default" tells PyArrow to use the fs.defaultFS property from core-site.xml
fs = pafs.HadoopFileSystem("default")

# Create a directory inside the bucket
fs.create_dir("volume/bucket/aaa")

# Write data to a file
path = "volume/bucket/file1"
with fs.open_output_stream(path) as stream:
  stream.write(b'data')

note

在 core-site.xml 中将 fs.defaultFS 配置为指向 Ozone 集群。例如,

<configuration>
  <property>
    <name>fs.defaultFS</name>
    <value>ofs://om:9862</value>
    <description>Ozone Manager endpoint</description>
  </property>
</configuration>

自己试试!查看 PyArrow 教程,快速上手使用 Ozone 的 Docker 镜像。

方法二:通过 Boto3 和 S3 Gateway 访问 Ozone

配置

  • 确定你的 Ozone S3 Gateway 端点(例如 http://s3g:9878)。
  • 使用 AWS 兼容的凭据(来自 Ozone)。

代码示例

import boto3

# Create a local file to upload
with open("localfile.txt", "w") as f:
  f.write("Hello from Ozone via Boto3!\n")

# Configure Boto3 client
s3 = boto3.client(
  's3',
  endpoint_url='http://s3g:9878',
  aws_access_key_id='ozone-access-key',
  aws_secret_access_key='ozone-secret-key'
)

# List buckets
response = s3.list_buckets()
print("Buckets:", response['Buckets'])

# Upload the file
s3.upload_file('localfile.txt', 'bucket', 'file.txt')
print("Uploaded 'localfile.txt' to 'bucket/file.txt'")

# Download the file back
s3.download_file('bucket', 'file.txt', 'downloaded.txt')
print("Downloaded 'file.txt' as 'downloaded.txt'")

::: note
请将端点 URL、凭据和存储桶名称替换为你的实际配置。
:::

自己动手试试!

快速上手请参阅 Boto3 教程,使用 Ozone 的 Docker 镜像进行体验。

方法 3:通过 HttpFS REST API 访问 Ozone

首先,安装 requests Python 模块:

pip install requests

配置

  • 使用 Ozone 的 HttpFS 端点(例如 http://httpfs:14000)。

代码示例(requests)

#!/usr/bin/python
import requests

# Ozone HTTPFS endpoint and file path
host = "http://httpfs:14000"
volume = "vol1"
bucket = "bucket1"
filename = "hello.txt"
path = f"/webhdfs/v1/{volume}/{bucket}/{filename}"
user = "ozone"  # can be any value in simple auth mode

# Step 1: Initiate file creation (responds with 307 redirect)
params_create = {
  "op": "CREATE",
  "overwrite": "true",
  "user.name": user
}

print("Creating file...")
resp_create = requests.put(host + path, params=params_create, allow_redirects=False)

if resp_create.status_code != 307:
  print(f"Unexpected response: {resp_create.status_code}")
  print(resp_create.text)
  exit(1)

redirect_url = resp_create.headers['Location']
print(f"Redirected to: {redirect_url}")

# Step 2: Write data to the redirected location with correct headers
headers = {"Content-Type": "application/octet-stream"}
content = b"Hello from Ozone HTTPFS!\n"

resp_upload = requests.put(redirect_url, data=content, headers=headers)
if resp_upload.status_code != 201:
  print(f"Upload failed: {resp_upload.status_code}")
  print(resp_upload.text)
  exit(1)
print("File created successfully.")

# Step 3: Read the file back
params_open = {
  "op": "OPEN",
  "user.name": user
}

print("Reading file...")
resp_read = requests.get(host + path, params=params_open, allow_redirects=True)
if resp_read.ok:
  print("File contents:")
  print(resp_read.text)
else:
  print(f"Read failed: {resp_read.status_code}")
  print(resp_read.text)

自己动手试试!查看使用 HttpFS REST API 访问 Ozone 教程,快速上手 Ozone 的 Docker 镜像。

代码示例(webhdfs)

首先,安装 fsspec Python 模块:

pip install fsspec
from fsspec.implementations.webhdfs import WebHDFS

fs = WebHDFS(host='httpfs', port=14000, user='ozone')

# Read a file from /vol1/bucket1/hello.txt
file_path = "/vol1/bucket1/hello.txt"

with fs.open(file_path, mode='rb') as f:
  content = f.read()
  print("File contents:")
  print(content.decode('utf-8'))

::: note
请根据你的环境替换 host、port 和 path。
:::

故障排查提示

  • 认证错误:请检查凭据以及 Kerberos 令牌(如使用)是否正确。
  • 连接问题:请检查端点 URL、端口和防火墙规则。
  • 文件系统错误:请确保 Ozone 配置正确且拥有相应权限。
  • 缺少依赖:请安装所需的 Python 包(pip install pyarrow boto3 requests fsspec)。

参考资料与延伸资源

评论

登录后参与评论

正在加载评论…