Bolt Parquet Write Configuration

qianmoQqianmoQ· 更新于 2026-10-02· 阅读 24 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

Bolt Parquet Write Configuration

This document outlines the Parquet write-time configuration options supported in Bolt. The Bolt Parquet writer is based on the Arrow C++ Parquet library and is primarily exposed through the Hive connector.

In this document, “Support” indicates that a configuration is available and effective when set through Bolt’s public configuration pathways (e.g., Hive session/table properties, QueryConfig).

Spark-Level Configurations (Reference)

These standard Spark properties are not directly used by Bolt’s writer. Instead, Bolt provides equivalent functionality through its own configuration mechanisms.

PropertySupportBolt DefaultConfiguration PathNotes
spark.sql.parquet.outputTimestampTypePartialINT96 disabled; Coercion to SECONDHiveConfig arrowBridgeTimestampUnit, WriterOptions writeInt96AsTimestampBolt manages timestamp precision via arrowBridgeTimestampUnit (SECOND, MILLI, MICRO, NANO). The internal writeInt96AsTimestamp option provides an override for legacy INT96 format and takes precedence. See bolt/connectors/hive/HiveDataSink.cpp and bolt/dwio/parquet/writer/Writer.h.
spark.sql.parquet.writeLegacyFormatNoN/ANot applicableBolt does not use this flag. The format version is controlled by the parquet.writer.version property.

Arrow Writer Properties in Bolt

These properties correspond to settings within the underlying parquet::arrow::WriterProperties and parquet::arrow::ArrowWriterProperties. Bolt exposes some of these directly or maps them from other configurations.

PropertySupportBolt DefaultConfiguration PathNotes
write_batch_sizeInternal1024Not exposedDefault from Arrow writer properties. Bolt uses its own heuristics (writeBatchBytes, minBatchSize) for batching. See bolt/dwio/parquet/arrow/Properties.h.
max_row_group_lengthYes1,048,576 rowsFlush PolicyThe default flush policy triggers at ~1M rows or ~128MiB. Configurable via a custom flush policy factory in WriterOptions. See bolt/dwio/parquet/writer/Writer.h.
parquet_block_sizeYes128 MiB (when enabled)WriterOptions.parquet_block_sizeThis is only effective when enableFlushBasedOnBlockSize is true, which overrides the default row/byte flush policy. See bolt/dwio/parquet/writer/Writer.cpp.
data_page_versionYesV1WriterOptions.dataPageVersionCan be set to V1 or V2. See bolt/dwio/parquet/writer/Writer.h.
writer.versionYesPARQUET_2_6WriterOptions.parquetVersionControls the Parquet format version. See bolt/dwio/parquet/arrow/Properties.h.
compressionYesUNCOMPRESSEDHive INSERT compressionKind propertySupported codecs: SNAPPY, GZIP, ZSTD, LZ4, UNCOMPRESSED. Mapped in bolt/connectors/hive/HiveDataSink.cpp.
compression_level & codec_optionsInternalVaries by codecWriterOptions.codecOptionsSupported internally by the Arrow writer but not exposed through Hive session/table properties. Can be configured per-column. See bolt/dwio/parquet/arrow/Properties.h.
dictionary_enabledYestrueWriterOptions.enableDictionary (global), columnEnableDictionaryMap (per-column)Dictionary encoding can be controlled globally or for specific columns. See bolt/dwio/parquet/writer/Writer.h.
data_page_sizeYes1 MiBWriterOptions.dataPageSize (global), columnDataPageSizeMap (per-column)See bolt/dwio/parquet/writer/Writer.h.
dictionary_page_size_limitYes1 MiBWriterOptions.dictionaryPageSizeLimit (global), columnDictionaryPageSizeLimitMap (per-column)See bolt/dwio/parquet/writer/Writer.h.
page_index_enabledInternalfalseNot exposedThe Arrow writer supports page index generation, but it is not currently configurable in Bolt. See bolt/dwio/parquet/arrow/Properties.h.
statistics_enabledInternaltrueNot exposedStatistics are enabled by default with a max size of 4096 bytes per column. Not externally configurable. See bolt/dwio/parquet/arrow/Properties.h.
store_decimal_as_integerYestrueWriterOptions.storeDecimalAsIntegerBolt defaults to storing compatible decimals as int32/int64 for efficiency. See bolt/dwio/parquet/writer/Writer.h.
created_byInternal"parquet-cpp-bolt"HardcodedThe created_by string in the file metadata is set internally. See bolt/dwio/parquet/arrow/Properties.h.
sorting_columnsInternalNot setNot exposedThe Arrow writer can store sorting metadata, but Bolt does not expose a pathway to set this. See bolt/dwio/parquet/arrow/Properties.h.
page_checksum_enabledYesfalseparquet.page.write-checksum.enabled table propertyMapped internally to the Arrow writer’s enable_page_checksum() builder method. See bolt/dwio/parquet/arrow/Properties.h.
encryptionInternalDisabledWriterOptions.encryptionOptionsThe writer supports AES_GCM_V1 and AES_GCM_CTR_V1 encryption if properties are provided, but this is not exposed through the Hive connector. See bolt/dwio/parquet/writer/Writer.h.
threading (use_threads)Yesfalse (0 threads)WriterOptions.threadPoolSizeIf threadPoolSize > 0, threading is enabled for parallel column writing using a global static thread pool. See bolt/dwio/parquet/writer/Writer.cpp.
compliant_nested_typesInternaltrueNot exposedThe writer follows the Parquet specification for nested list element naming (“element”). See bolt/dwio/parquet/arrow/Properties.h.

Parquet-MR Style Settings

This table indicates whether classic parquet-mr Hadoop configurations are effectively supported by Bolt’s writer, typically by mapping to an equivalent WriterOptions or WriterProperties setting.

PropertyEffective in BoltBolt DefaultConfiguration PathNotes
parquet.block.sizeYes128 MiBWriterOptions.parquet_block_sizeUsed when enableFlushBasedOnBlockSize is true.
parquet.page.sizeYes1 MiBWriterOptions.dataPageSize
parquet.compressionYesUNCOMPRESSEDHive compressionKind table propertyMaps to WriterOptions.compression.
parquet.enable.dictionaryYestrueWriterOptions.enableDictionary
parquet.dictionary.page.sizeYes1 MiBWriterOptions.dictionaryPageSizeLimit
parquet.writer.versionYesPARQUET_2_6WriterOptions.parquetVersion
parquet.compression.codec.zstd.levelYes3WriterOptions.codecOptionsExposed via session/table properties and mapped internally. Default is from zstd library. See bolt/dwio/parquet/arrow/util/CompressionZstd.cpp.
parquet.page.write-checksum.enabledYesfalseWriterProperties::BuilderConfigurable via table properties.
parquet.enable.summary-metadataNoN/ANot implementedBolt does not create a separate summary file.
parquet.bloom.filter.enabledNoN/ANot implementedBloom filter writing is not supported.
parquet.crypto.factory.classNoN/ANot implementedEncryption is handled internally via WriterOptions.
parquet.compression.codec.zstd.workersNoN/ANot implementedBolt’s Parquet writer threading is controlled by WriterOptions.threadPoolSize.
parquet.validationNoN/ANot implemented

Bolt-Specific Writer Options and Behaviors

These options and behaviors are specific to Bolt’s implementation and provide more granular control over the writing process.

  • Default Flush Policy: By default, the Parquet writer flushes a row group when it reaches approximately 1,048,576 rows or its estimated size exceeds 128 MiB. This is defined in DefaultFlushPolicy in bolt/dwio/parquet/writer/Writer.h.
  • enableFlushBasedOnBlockSize : A bool in WriterOptions that changes the flush behavior from the default row/byte count policy to a purely block-size-based policy. When true, it uses WriterOptions.parquet_block_size (defaults to 128 MiB) to control row group size and sets max_row_group_length to a very large value to prevent it from triggering first. The Arrow writer’s NewBufferedRowGroup() is used. Found in bolt/dwio/parquet/writer/Writer.cpp.
  • enableRowGroupAlignedWrite : A bool in WriterOptions used for specialized data retention scenarios. It works with expectedRowsInEachBlock to create row groups with an exact number of rows. See bolt/dwio/parquet/writer/Writer.h.
  • writeBatchBytes / minBatchSize : Internal heuristics within WriterOptions to manage memory and performance when converting large Bolt VectorPtr batches to Arrow RecordBatch. The writer may split large batches into smaller ones based on these thresholds (defaults: 40 MiB and 512 rows). See bolt/dwio/parquet/writer/Writer.cpp.
  • Filename Extension: For Hive INSERT operations, HiveDataSink automatically appends the .parquet extension to output files when the storage format is PARQUET. This is handled in bolt/connectors/hive/HiveDataSink.cpp.
  • Timestamp Bridge Unit: The precision of timestamps written to Parquet is controlled by the arrow_bridge_timestamp_unit session property (via HiveConfig), which can be set to SECOND, MILLI, MICRO, or NANO. For legacy compatibility, the internal WriterOptions.writeInt96AsTimestamp (bool) can be set to true to force the deprecated INT96 format, and this setting takes precedence over the bridge unit. See bolt/connectors/hive/HiveConfig.cpp and bolt/dwio/parquet/writer/Writer.h.

评论

登录后参与评论

正在加载评论…