Basic Configurations
This page covers the basic configurations you may use to write/read Hudi tables. This page only features a subset of the most frequently used configurations. For a full list of all configs, please visit the All Configurations page.
- Hudi Table Config: Basic Hudi Table configuration parameters.
- Spark Datasource Configs: These configs control the Hudi Spark Datasource, providing ability to define keys/partitioning, pick out the write operation, specify how to merge records or choosing query type to read.
- Flink Sql Configs: These configs control the Hudi Flink SQL source/sink connectors, providing ability to define record keys, pick out the write operation, specify how to merge records, enable/disable asynchronous compaction or choosing query type to read.
- Write Client Configs: Internally, the Hudi datasource uses a RDD based HoodieWriteClient API to actually perform writes to storage. These configs provide deep control over lower level aspects like file sizing, compression, parallelism, compaction, write schema, cleaning etc. Although Hudi provides sane defaults, from time-time these configs may need to be tweaked to optimize for specific workloads.
- Metastore and Catalog Sync Configs: Configurations used by the Hudi to sync metadata to external metastores and catalogs.
- Metrics Configs: These set of configs are used to enable monitoring and reporting of key Hudi stats and metrics.
- Kafka Connect Configs: These set of configs are used for Kafka Connect Sink Connector for writing Hudi Tables
- Hudi Streamer Configs: These set of configs are used for Hudi Streamer utility which provides the way to ingest from different sources such as DFS or Kafka.
note
In the tables below (N/A) means there is no default value set
Hudi Table Config
Basic Hudi Table configuration parameters.
Hudi Table Basic Configs
Configurations of the Hudi Table like type of ingestion, storage formats, hive table name etc. Configurations are loaded from hoodie.properties, these properties are usually set during initializing a path as hoodie base path and never changes during the lifetime of a hoodie table.
Config NameDefaultDescription
hoodie.bootstrap.base.path(N/A)Base path of the dataset that needs to be bootstrapped as a Hudi tableConfig Param: BOOTSTRAP_BASE_PATH
hoodie.compaction.payload.class(N/A)Payload class to use for performing merges, compactions, i.e merge delta logs with current base file and then produce a new base file.Config Param: PAYLOAD_CLASS_NAME
hoodie.database.name(N/A)Database name. If different databases have the same table name during incremental query, we can set it to limit the table name under a specific databaseConfig Param: DATABASE_NAME
hoodie.record.merge.mode(N/A)org.apache.hudi.common.config.RecordMergeMode: Determines the logic of merging updates COMMIT_TIME_ORDERING: Using transaction time to merge records, i.e., the record from later transaction overwrites the earlier record with the same key. EVENT_TIME_ORDERING: Using event time as the ordering to merge records, i.e., the record with the larger event time overwrites the record with the smaller event time on the same key, regardless of transaction time. The event time or ordering fields need to be specified by the user. CUSTOM: Using custom merging logic specified by the user.Config Param: RECORD_MERGE_MODESince Version: 1.0.0
hoodie.record.merge.strategy.id(N/A)Id of merger strategy. Hudi will pick HoodieRecordMerger implementations in hoodie.write.record.merge.custom.implementation.classes which has the same merger strategy idConfig Param: RECORD_MERGE_STRATEGY_IDSince Version: 0.13.0
hoodie.table.checksum(N/A)Table checksum is used to guard against partial writes in HDFS. It is added as the last entry in hoodie.properties and then used to validate while reading table config.Config Param: TABLE_CHECKSUMSince Version: 0.11.0
hoodie.table.create.schema(N/A)Schema used when creating the tableConfig Param: CREATE_SCHEMA
hoodie.table.index.defs.path(N/A)Relative path to table base path where the index definitions are storedConfig Param: RELATIVE_INDEX_DEFINITION_PATHSince Version: 1.0.0
hoodie.table.keygenerator.class(N/A)Key Generator class property for the hoodie tableConfig Param: KEY_GENERATOR_CLASS_NAME
hoodie.table.keygenerator.type(N/A)Key Generator type to determine key generator classConfig Param: KEY_GENERATOR_TYPESince Version: 1.0.0
hoodie.table.legacy.payload.class(N/A)Payload class to indicate the payload class that is used to create the table and is not used anymore.Config Param: LEGACY_PAYLOAD_CLASS_NAMESince Version: 1.1.0
hoodie.table.metadata.partitions(N/A)Comma-separated list of metadata partitions that have been completely built and in-sync with data table. These partitions are ready for use by the readersConfig Param: TABLE_METADATA_PARTITIONSSince Version: 0.11.0
hoodie.table.metadata.partitions.inflight(N/A)Comma-separated list of metadata partitions whose building is in progress. These partitions are not yet ready for use by the readers.Config Param: TABLE_METADATA_PARTITIONS_INFLIGHTSince Version: 0.11.0
hoodie.table.name(N/A)Table name that will be used for registering with Hive. Needs to be same across runs.Config Param: NAME
hoodie.table.ordering.fields(N/A)Comma separated fields used in records merging comparison. By default, when two records have the same key value, the largest value for the ordering field determined by Object.compareTo(..), is picked. If there are multiple fields configured, comparison is made on the first field. If the first field values are same, comparison is made on the second field and so on.Config Param: ORDERING_FIELDS
hoodie.table.partial.update.mode(N/A)This property when set, will define how two versions of the record will be merged together when records are partially formedConfig Param: PARTIAL_UPDATE_MODESince Version: 1.1.0
hoodie.table.partition.fields(N/A)Comma separated field names used to partition the table. These field names also include the partition type which is used by custom key generatorsConfig Param: PARTITION_FIELDS
hoodie.table.precombine.field(N/A)Comma separated fields used in preCombining before actual write. By default, when two records have the same key value, the largest value for the precombine field determined by Object.compareTo(..), is picked. If there are multiple fields configured, comparison is made on the first field. If the first field values are same, comparison is made on the second field and so on.Config Param: PRECOMBINE_FIELD
hoodie.table.recordkey.fields(N/A)Columns used to uniquely identify the table. Concatenated values of these fields are used as the record key component of HoodieKey.Config Param: RECORDKEY_FIELDS
hoodie.table.secondary.indexes.metadata(N/A)The metadata of secondary indexesConfig Param: SECONDARY_INDEXES_METADATASince Version: 0.13.0
hoodie.timeline.layout.version(N/A)Version of timeline used, by the table.Config Param: TIMELINE_LAYOUT_VERSION
hoodie.archivelog.folderarchivedpath under the meta folder, to store archived timeline instants at.Config Param: ARCHIVELOG_FOLDER
hoodie.bootstrap.index.classorg.apache.hudi.common.bootstrap.index.hfile.HFileBootstrapIndexImplementation to use, for mapping base files to bootstrap base file, that contain actual data.Config Param: BOOTSTRAP_INDEX_CLASS_NAME
hoodie.bootstrap.index.enabletrueWhether or not, this is a bootstrapped table, with bootstrap base data and an mapping index defined, default true.Config Param: BOOTSTRAP_INDEX_ENABLE
hoodie.bootstrap.index.typeHFILEBootstrap index type determines which implementation to use, for mapping base files to bootstrap base file, that contain actual data.Config Param: BOOTSTRAP_INDEX_TYPESince Version: 1.0.0
hoodie.datasource.write.hive_style_partitioningfalseFlag to indicate whether to use Hive style partitioning. If set true, the names of partition folders follow <partition_column_name>=<partition_value> format. By default false (the names of partition folders are only partition values)Config Param: HIVE_STYLE_PARTITIONING_ENABLE
hoodie.partition.metafile.use.base.formatfalseIf true, partition metafiles are saved in the same format as base-files for this dataset (e.g. Parquet / ORC). If false (default) partition metafiles are saved as properties files.Config Param: PARTITION_METAFILE_USE_BASE_FORMAT
hoodie.populate.meta.fieldstrueWhen enabled, populates all meta fields. When disabled, no meta fields are populated and incremental queries will not be functional. This is only meant to be used for append only/immutable data for batch processingConfig Param: POPULATE_META_FIELDS
hoodie.table.base.file.formatPARQUETBase file format to store all the base file data.Config Param: BASE_FILE_FORMAT
hoodie.table.cdc.enabledfalseWhen enable, persist the change data if necessary, and can be queried as a CDC query mode.Config Param: CDC_ENABLEDSince Version: 0.13.0
hoodie.table.cdc.supplemental.logging.modeDATA_BEFORE_AFTERorg.apache.hudi.common.table.cdc.HoodieCDCSupplementalLoggingMode: Change log capture supplemental logging mode. The supplemental log is used for accelerating the generation of change log details. OP_KEY_ONLY: Only keeping record keys in the supplemental logs, so the reader needs to figure out the update before image and after image. DATA_BEFORE: Keeping the before images in the supplemental logs, so the reader needs to figure out the update after images. DATA_BEFORE_AFTER(default): Keeping the before and after images in the supplemental logs, so the reader can generate the details directly from the logs.Config Param: CDC_SUPPLEMENTAL_LOGGING_MODESince Version: 0.13.0
hoodie.table.formatnativeTable format name used when writing to the table.Config Param: TABLE_FORMAT
hoodie.table.initial.versionNINEInitial Version of table when the table was created. Used for upgrade/downgrade to identify what upgrade/downgrade paths happened on the table. This is only configured when the table is initially setup.Config Param: INITIAL_VERSIONSince Version: 1.0.0
hoodie.table.log.file.formatHOODIE_LOGLog format used for the delta logs.Config Param: LOG_FILE_FORMAT
hoodie.table.multiple.base.file.formats.enablefalseWhen set to true, the table can support reading and writing multiple base file formats.Config Param: MULTIPLE_BASE_FILE_FORMATS_ENABLESince Version: 1.0.0
hoodie.table.timeline.timezoneLOCALUser can set hoodie commit timeline timezone, such as utc, local and so on. local is defaultConfig Param: TIMELINE_TIMEZONE
hoodie.table.typeCOPY_ON_WRITEThe table type for the underlying data.Config Param: TYPE
hoodie.table.versionNINEVersion of table, used for running upgrade/downgrade steps between releases with potentially breaking/backwards compatible changes.Config Param: VERSION
hoodie.timeline.history.pathhistorypath under the meta folder, to store timeline history at.Config Param: TIMELINE_HISTORY_PATH
hoodie.timeline.pathtimelinepath under the meta folder, to store timeline instants at.Config Param: TIMELINE_PATH
Spark Datasource Configs
These configs control the Hudi Spark Datasource, providing ability to define keys/partitioning, pick out the write operation, specify how to merge records or choosing query type to read.
Read Options
Options useful for reading tables via read.format.option(...)
Config NameDefaultDescription
hoodie.datasource.read.begin.instanttime(N/A)Required when hoodie.datasource.query.type is set to incremental. Represents the completion time to start incrementally pulling data from. The completion time here need not necessarily correspond to an instant on the timeline. New data written with completion_time >= START_COMMIT are fetched out. For e.g: ‘20170901080000’ will get all new data written on or after Sep 1, 2017 08:00AM.Config Param: START_COMMITSince Version: 0.9.0
hoodie.datasource.read.end.instanttime(N/A)Used when hoodie.datasource.query.type is set to incremental. Represents the completion time to limit incrementally fetched data to. When not specified latest commit completion time from timeline is assumed by default. When specified, new data written with completion_time <= END_COMMIT are fetched out. Point in time type queries make more sense with begin and end completion times specified.Config Param: END_COMMITSince Version: 0.9.0
hoodie.datasource.read.incr.table.version(N/A)The table version assumed for incremental readConfig Param: INCREMENTAL_READ_TABLE_VERSIONSince Version: 1.0.0
hoodie.datasource.read.streaming.table.version(N/A)The table version assumed for streaming readConfig Param: STREAMING_READ_TABLE_VERSIONSince Version: 1.0.0
hoodie.datasource.write.precombine.field(N/A)Comma separated list of fields used in preCombining before actual write. When two records have the same key value, we will pick the one with the largest value for the precombine field, determined by Object.compareTo(..). For multiple fields if first key comparison is same, second key comparison is made and so on. This config is used for combining records within the same batch and also for merging using event time merge modeConfig Param: READ_PRE_COMBINE_FIELD
hoodie.datasource.query.typesnapshotWhether data needs to be read, in incremental mode (new data since an instantTime) (or) read_optimized mode (obtain latest view, based on base files) (or) snapshot mode (obtain latest view, by merging base and (if any) log files)Config Param: QUERY_TYPESince Version: 0.9.0
Write Options
You can pass down any of the WriteClient level configs directly using options() or option(k,v) methods.
inputDF.write()
.format("org.apache.hudi")
.options(clientOpts) // any of the Hudi client opts can be passed in as well
.option(DataSourceWriteOptions.RECORDKEY_FIELD_OPT_KEY(), "_row_key")
.option(DataSourceWriteOptions.PARTITIONPATH_FIELD_OPT_KEY(), "partition")
.option(HoodieTableConfig.ORDERING_FIELDS(), "timestamp")
.option(HoodieWriteConfig.TABLE_NAME, tableName)
.mode(SaveMode.Append)
.save(basePath);Options useful for writing tables via write.format.option(...)
Config NameDefaultDescription
hoodie.datasource.hive_sync.mode(N/A)Mode to choose for Hive ops. Valid values are hms, jdbc and hiveql.Config Param: HIVE_SYNC_MODE
hoodie.datasource.write.partitionpath.field(N/A)Partition path field. Value to be used at the partitionPath component of HoodieKey. Actual value obtained by invoking .toString()Config Param: PARTITIONPATH_FIELD
hoodie.datasource.write.precombine.field(N/A)Comma separated list of fields used in preCombining before actual write. When two records have the same key value, we will pick the one with the largest value for the precombine field, determined by Object.compareTo(..). For multiple fields if first key comparison is same, second key comparison is made and so on. This config is used for combining records within the same batch and also for merging using event time merge modeConfig Param: ORDERING_FIELDS
hoodie.datasource.write.precombine.field(N/A)Comma separated list of fields used in preCombining before actual write. When two records have the same key value, we will pick the one with the largest value for the precombine field, determined by Object.compareTo(..). For multiple fields if first key comparison is same, second key comparison is made and so on. This config is used for combining records within the same batch and also for merging using event time merge modeConfig Param: PRECOMBINE_FIELD
hoodie.datasource.write.recordkey.field(N/A)Record key field. Value to be used as the recordKey component of HoodieKey. Actual value will be obtained by invoking .toString() on the field value. Nested fields can be specified using the dot notation eg: a.b.cConfig Param: RECORDKEY_FIELD
hoodie.datasource.write.secondarykey.column(N/A)Columns that constitute the secondary key component. Actual value will be obtained by invoking .toString() on the field value. Nested fields can be specified using the dot notation eg: a.b.cConfig Param: SECONDARYKEY_COLUMN_NAME
hoodie.write.record.merge.mode(N/A)org.apache.hudi.common.config.RecordMergeMode: Determines the logic of merging updates COMMIT_TIME_ORDERING: Using transaction time to merge records, i.e., the record from later transaction overwrites the earlier record with the same key. EVENT_TIME_ORDERING: Using event time as the ordering to merge records, i.e., the record with the larger event time overwrites the record with the smaller event time on the same key, regardless of transaction time. The event time or ordering fields need to be specified by the user. CUSTOM: Using custom merging logic specified by the user.Config Param: RECORD_MERGE_MODESince Version: 1.0.0
hoodie.clustering.async.enabledfalseEnable running of clustering service, asynchronously as inserts happen on the table.Config Param: ASYNC_CLUSTERING_ENABLESince Version: 0.7.0
hoodie.clustering.inlinefalseTurn on inline clustering - clustering will be run after each write operation is completeConfig Param: INLINE_CLUSTERING_ENABLESince Version: 0.7.0
hoodie.datasource.hive_sync.enablefalseWhen set to true, register/sync the table to Apache Hive metastore.Config Param: HIVE_SYNC_ENABLED
hoodie.datasource.hive_sync.jdbcurljdbc:hive2://localhost:10000Hive metastore urlConfig Param: HIVE_URL
hoodie.datasource.hive_sync.metastore.uristhrift://localhost:9083Hive metastore urlConfig Param: METASTORE_URIS
hoodie.datasource.meta.sync.enablefalseEnable Syncing the Hudi Table with an external meta store or data catalog.Config Param: META_SYNC_ENABLED
hoodie.datasource.write.hive_style_partitioningfalseFlag to indicate whether to use Hive style partitioning. If set true, the names of partition folders follow <partition_column_name>=<partition_value> format. By default false (the names of partition folders are only partition values)Config Param: HIVE_STYLE_PARTITIONING
hoodie.datasource.write.operationupsertWhether to do upsert, insert or bulk_insert for the write operation. Use bulk_insert to load new data into a table, and there on use upsert/insert. bulk insert uses a disk based write path to scale to load large inputs without need to cache it.Config Param: OPERATIONSince Version: 0.9.0
hoodie.datasource.write.table.typeCOPY_ON_WRITEThe table type for the underlying data, for this write. This can’t change between writes.Config Param: TABLE_TYPESince Version: 0.9.0
Flink Sql Configs
These configs control the Hudi Flink SQL source/sink connectors, providing ability to define record keys, pick out the write operation, specify how to merge records, enable/disable asynchronous compaction or choosing query type to read.
Flink Options
Flink jobs using the SQL can be configured through the options in WITH clause. The actual datasource level configs are listed below.
Config NameDefaultDescription
hoodie.database.name(N/A)Database name to register to Hive metastoreConfig Param: DATABASE_NAME
hoodie.datasource.write.recordkey.field(N/A)Record key field. Value to be used as the recordKey component of HoodieKey. Actual value will be obtained by invoking .toString() on the field value. Nested fields can be specified using the dot notation eg: a.b.cConfig Param: RECORD_KEY_FIELD
hoodie.table.name(N/A)Table name to register to Hive metastoreConfig Param: TABLE_NAME
ordering.fields(N/A)Comma separated list of fields used in records merging. When two records have the same key value, we will pick the one with the largest value for the ordering field, determined by Object.compareTo(..). For multiple fields if first key comparison is same, second key comparison is made and so on. Config precombine.field is now deprecated, please use ordering.fields instead.Config Param: ORDERING_FIELDS
path(N/A)Base path for the target hoodie table. The path would be created if it does not exist, otherwise a Hoodie table expects to be initialized successfullyConfig Param: PATH
read.commits.limit(N/A)The maximum number of commits allowed to read in each instant check, if it is streaming read, the avg read instants number per-second would be 'read.commits.limit'/'read.streaming.check-interval', by default no limitConfig Param: READ_COMMITS_LIMIT
read.end-commit(N/A)End commit instant for reading, the commit time format should be 'yyyyMMddHHmmss'Config Param: READ_END_COMMIT
read.start-commit(N/A)Start commit instant for reading, the commit time format should be 'yyyyMMddHHmmss', by default reading from the latest instant for streaming readConfig Param: READ_START_COMMIT
archive.max_commits50Max number of commits to keep before archiving older commits into a sequential log, default 50Config Param: ARCHIVE_MAX_COMMITS
archive.min_commits40Min number of commits to keep before archiving older commits into a sequential log, default 40Config Param: ARCHIVE_MIN_COMMITS
cdc.enabledfalseWhen enable, persist the change data if necessary, and can be queried as a CDC query modeConfig Param: CDC_ENABLED
cdc.supplemental.logging.modeDATA_BEFORE_AFTERSetting 'op_key_only' persists the 'op' and the record key only, setting 'data_before' persists the additional 'before' image, and setting 'data_before_after' persists the additional 'before' and 'after' images.Config Param: SUPPLEMENTAL_LOGGING_MODE
changelog.enabledfalseWhether to keep all the intermediate changes, we try to keep all the changes of a record when enabled: 1). The sink accept the UPDATE_BEFORE message; 2). The source try to emit every changes of a record. The semantics is best effort because the compaction job would finally merge all changes of a record into one. default false to have UPSERT semanticsConfig Param: CHANGELOG_ENABLED
clean.async.enabledtrueWhether to cleanup the old commits immediately on new commits, enabled by defaultConfig Param: CLEAN_ASYNC_ENABLED
clean.retain_commits30Number of commits to retain. So data will be retained for num_of_commits * time_between_commits (scheduled). This also directly translates into how much you can incrementally pull on this table, default 30Config Param: CLEAN_RETAIN_COMMITS
clustering.async.enabledfalseAsync Clustering, default falseConfig Param: CLUSTERING_ASYNC_ENABLED
clustering.plan.strategy.small.file.limit600Files smaller than the size specified here are candidates for clustering, default 600 MBConfig Param: CLUSTERING_PLAN_STRATEGY_SMALL_FILE_LIMIT
clustering.plan.strategy.target.file.max.bytes1073741824Each group can produce 'N' (CLUSTERING_MAX_GROUP_SIZE/CLUSTERING_TARGET_FILE_SIZE) output file groups, default 1 GBConfig Param: CLUSTERING_PLAN_STRATEGY_TARGET_FILE_MAX_BYTES
compaction.async.enabledtrueAsync Compaction, enabled by default for MORConfig Param: COMPACTION_ASYNC_ENABLED
compaction.delta_commits5Max delta commits needed to trigger compaction, default 5 commitsConfig Param: COMPACTION_DELTA_COMMITS
hive_sync.enabledfalseAsynchronously sync Hive meta to HMS, default falseConfig Param: HIVE_SYNC_ENABLED
hive_sync.jdbc_urljdbc:hive2://localhost:10000Jdbc URL for hive sync, default 'jdbc:hive2://localhost:10000'Config Param: HIVE_SYNC_JDBC_URL
hive_sync.metastore.urisMetastore uris for hive sync, default ''Config Param: HIVE_SYNC_METASTORE_URIS
hive_sync.modeHMSMode to choose for Hive ops. Valid values are hms, jdbc and hiveql, default 'hms'Config Param: HIVE_SYNC_MODE
hoodie.datasource.query.typesnapshotDecides how data files need to be read, in 1) Snapshot mode (obtain latest view, based on row & columnar data); 2) incremental mode (new data since an instantTime); 3) Read Optimized mode (obtain latest view, based on columnar data) .Default: snapshotConfig Param: QUERY_TYPE
hoodie.datasource.write.hive_style_partitioningfalseWhether to use Hive style partitioning. If set true, the names of partition folders follow <partition_column_name>=<partition_value> format. By default false (the names of partition folders are only partition values)Config Param: HIVE_STYLE_PARTITIONING
hoodie.datasource.write.partitionpath.fieldPartition path field. Value to be used at the partitionPath component of HoodieKey. Actual value obtained by invoking .toString(), default ''Config Param: PARTITION_PATH_FIELD
hoodie.table.services.enabledtrueMaster control to disable all table services including archive, clean, compact, cluster, etc.Config Param: TABLE_SERVICES_ENABLED
index.typeFLINK_STATEIndex type of Flink write job, default is using state backed index.Config Param: INDEX_TYPE
lookup.asyncfalseWhether to enable async lookup join.Config Param: LOOKUP_ASYNC
lookup.async-thread-number16The thread number for lookup async.Config Param: LOOKUP_ASYNC_THREAD_NUMBER
lookup.join.cache.ttlPT1HThe cache TTL (e.g. 10min) for the build table in lookup join.Config Param: LOOKUP_JOIN_CACHE_TTL
lookup.join.cache.typeheapThe storage backend for the lookup join cache. Possible values: 'heap' (default) stores all dimension-table rows in JVM heap memory (may cause OutOfMemoryError for large tables); 'rocksdb' stores rows off-heap in an embedded RocksDB instance on local disk, which avoids OOM at the cost of additional serialization overhead.Config Param: LOOKUP_JOIN_CACHE_TYPE
lookup.join.rocksdb.path/var/folders/vh/zgs02hf51dn7r08pbl5m2jc00000gn/T//hudi-lookup-rocksdbLocal directory path for storing RocksDB data when 'lookup.join.cache.type' is set to 'rocksdb'. Each task manager will create a unique subdirectory under this path. The directory is cleaned up when the lookup function is closed.Config Param: LOOKUP_JOIN_ROCKSDB_PATH
metadata.compaction.async.enabledtrueWhether to enable async compaction for metadata table,if true, the compaction for metadata table will be performed in the compaction pipeline, default enabled.Config Param: METADATA_COMPACTION_ASYNC_ENABLED
metadata.compaction.delta_commits10Max delta commits for metadata table to trigger compaction, default 10Config Param: METADATA_COMPACTION_DELTA_COMMITS
metadata.enabledtrueEnable the internal metadata table which serves table metadata like level file listings, default enabledConfig Param: METADATA_ENABLED
read.source-v2.enabledfalseWhether to use Flink FLIP27 new source to consume data files.Config Param: READ_SOURCE_V2_ENABLED
read.splits.limit2147483647The maximum number of splits allowed to read in each instant check, if it is streaming read, the avg read splits number per-second would be 'read.splits.limit'/'read.streaming.check-interval', by default no limitConfig Param: READ_SPLITS_LIMIT
read.streaming.enabledfalseWhether to read as streaming source, default falseConfig Param: READ_AS_STREAMING
read.streaming.skip_insertoverwritefalseWhether to skip insert overwrite instants to avoid reading base files of insert overwrite operations for streaming read. In streaming scenarios, insert overwrite is usually used to repair data, here you can control the visibility of downstream streaming read.Config Param: READ_STREAMING_SKIP_INSERT_OVERWRITE
table.typeCOPY_ON_WRITEType of table to write. COPY_ON_WRITE (or) MERGE_ON_READConfig Param: TABLE_TYPE
write.insert.partitioner.class.nameInsert partitioner to use aiming to re-balance records and reducing small file number in the scenario of multi-level partitioning. For example dt/hour/eventIDCurrently support org.apache.hudi.sink.partitioner.GroupedInsertPartitionerConfig Param: INSERT_PARTITIONER_CLASS_NAME
write.insert.partitioner.default_parallelism_per_partition30The parallelism to use in each partition when using GroupedInsertPartitioner.Config Param: DEFAULT_PARALLELISM_PER_PARTITION
write.operationupsertThe write operation, that this write should doConfig Param: OPERATION
write.parquet.max.file.size120Target size for parquet files produced by Hudi write phases. For DFS, this needs to be aligned with the underlying filesystem block size for optimal performance.Config Param: WRITE_PARQUET_MAX_FILE_SIZE
Write Client Configs
Internally, the Hudi datasource uses a RDD based HoodieWriteClient API to actually perform writes to storage. These configs provide deep control over lower level aspects like file sizing, compression, parallelism, compaction, write schema, cleaning etc. Although Hudi provides sane defaults, from time-time these configs may need to be tweaked to optimize for specific workloads.
Common Configurations
The following set of configurations are common across Hudi.
Config NameDefaultDescription
hoodie.base.path(N/A)Base path on lake storage, under which all the table data is stored. Always prefix it explicitly with the storage scheme (e.g hdfs://, s3:// etc). Hudi stores all the main meta-data about commits, savepoints, cleaning audit logs etc in .hoodie directory under this base path directory.Config Param: BASE_PATH
Metadata Configs
Configurations used by the Hudi Metadata Table. This table maintains the metadata about a given Hudi table (e.g file listings) to avoid overhead of accessing cloud storage, during queries.
Config NameDefaultDescription
hoodie.index.name(N/A)Name of the expression index. This is also used for the partition name in the metadata table.Config Param: SECONDARY_INDEX_NAMESince Version: 1.0.0
hoodie.index.name(N/A)Name of the expression index. This is also used for the partition name in the metadata table.Config Param: EXPRESSION_INDEX_NAMESince Version: 1.0.0
hoodie.metadata.index.drop(N/A)Drop the specified index. The value should be the name of the index to delete. You can check index names using SHOW INDEXES command. The index name either starts with or matches exactly can be one of the following: files, column_stats, bloom_filters, record_index, expr_index_, secondary_index_, partition_stats, filesConfig Param: DROP_METADATA_INDEXSince Version: 1.0.1
hoodie.expression.index.typeCOLUMN_STATSType of the expression index. Default is column_stats if there are no functions and expressions in the command. Valid options could be BITMAP, COLUMN_STATS, LUCENE, etc. If index_type is not provided, and there are functions or expressions in the command then a expression index using column stats will be created.Config Param: EXPRESSION_INDEX_TYPESince Version: 1.0.0
hoodie.metadata.enabletrueEnable the internal metadata table which serves table metadata like level file listingsConfig Param: ENABLESince Version: 0.7.0
hoodie.metadata.index.bloom.filter.enablefalseEnable indexing bloom filters of user data files under metadata table. When enabled, metadata table will have a partition to store the bloom filter index and will be used during the index lookups.Config Param: ENABLE_METADATA_INDEX_BLOOM_FILTERSince Version: 0.11.0
hoodie.metadata.index.column.stats.enablefalseEnable indexing column ranges of user data files under metadata table key lookups. When enabled, metadata table will have a partition to store the column ranges and will be used for pruning files during the index lookups. For the Spark engine, this config defaults to true (enabled), overriding the base default of false. For Flink and Java engines, this remains false by default.Config Param: ENABLE_METADATA_INDEX_COLUMN_STATSSince Version: 0.11.0
hoodie.metadata.index.expression.enablefalseEnable expression index within the metadata table. When this configuration property is enabled (true), the Hudi writer automatically keeps all expression indexes consistent with the data table. When disabled (false), all expression indexes are deleted. Note that individual expression index can only be created through a CREATE INDEX and deleted through a DROP INDEX statement in Spark SQL.Config Param: EXPRESSION_INDEX_ENABLE_PROPSince Version: 1.0.0
hoodie.metadata.index.secondary.enabletrueEnable secondary index within the metadata table. When this configuration property is enabled (true), the Hudi writer automatically keeps all secondary indexes consistent with the data table. When disabled (false), all secondary indexes are deleted. Note that individual secondary index can only be created through a CREATE INDEX and deleted through a DROP INDEX statement in Spark SQL.Config Param: SECONDARY_INDEX_ENABLE_PROPSince Version: 1.0.0
Storage Configs
Configurations that control aspects around writing, sizing, reading base and log files.
Config NameDefaultDescription
hoodie.hfile.writes.allow.duplicatesfalseWhen bootstrapping RI, if the main dataset contains duplicates then it will fail the bootstrap job. TO avoid the failure and bootstrap the RI with dups this config can be set to true. One thing to note is that, there is no deterministic way to specify which among these records will be ingested into RI.Config Param: HFILE_WRITER_TO_ALLOW_DUPLICATES
hoodie.parquet.compression.codecgzipCompression Codec for parquet filesConfig Param: PARQUET_COMPRESSION_CODEC_NAME
hoodie.parquet.max.file.size125829120Target size in bytes for parquet files produced by Hudi write phases. For DFS, this needs to be aligned with the underlying filesystem block size for optimal performance.Config Param: PARQUET_MAX_FILE_SIZE
Archival Configs
Configurations that control archival.
Config NameDefaultDescription
hoodie.keep.max.commits30Archiving service moves older entries from timeline into an archived log after each write, to keep the metadata overhead constant, even as the table size grows. This config controls the maximum number of instants to retain in the active timeline.Config Param: MAX_COMMITS_TO_KEEP
hoodie.keep.min.commits20Similar to hoodie.keep.max.commits, but controls the minimum number of instants to retain in the active timeline.Config Param: MIN_COMMITS_TO_KEEP
Bootstrap Configs
Configurations that control how you want to bootstrap your existing tables for the first time into hudi. The bootstrap operation can flexibly avoid copying data over before you can use Hudi and support running the existing writers and new hudi writers in parallel, to validate the migration.
Config NameDefaultDescription
hoodie.bootstrap.base.path(N/A)Base path of the dataset that needs to be bootstrapped as a Hudi tableConfig Param: BASE_PATHSince Version: 0.6.0
Clean Configs
Cleaning (reclamation of older/unused file groups/slices).
Config NameDefaultDescription
hoodie.clean.async.enabledfalseOnly applies when hoodie.clean.automatic is turned on. When turned on runs cleaner async with writing, which can speed up overall write performance.Config Param: ASYNC_CLEAN
hoodie.clean.commits.retained10When KEEP_LATEST_COMMITS cleaning policy is used, the number of commits to retain, without cleaning. This will be retained for num_of_commits * time_between_commits (scheduled). This also directly translates into how much data retention the table supports for incremental queries.Config Param: CLEANER_COMMITS_RETAINED
hoodie.prewrite.cleaner.policyNONEorg.apache.hudi.common.model.HoodiePreWriteCleanerPolicy: If set, force attempting clean and/or failed writes rollback before starting a new ingestion write commit. This should only be set to ensure that data files do not build up on DFS if an ingestion writer is perpetually failing before completing a CLEAN. NONE(default): No pre-write clean or rollback. Default behavior. CLEAN: Force a CLEAN table service call before starting the write (also performs rollback of failed writes). ROLLBACK_FAILED_WRITES: Only perform rollback of failed writes before starting the write.Config Param: PREWRITE_CLEANER_POLICYSince Version: 1.2.0
Clustering Configs
Configurations that control the clustering table service in hudi, which optimizes the storage layout for better query performance by sorting and sizing data files.
Config NameDefaultDescription
hoodie.clustering.async.enabledfalseEnable running of clustering service, asynchronously as inserts happen on the table.Config Param: ASYNC_CLUSTERING_ENABLESince Version: 0.7.0
hoodie.clustering.inlinefalseTurn on inline clustering - clustering will be run after each write operation is completeConfig Param: INLINE_CLUSTERINGSince Version: 0.7.0
hoodie.clustering.plan.generation.use.local.engine.contextfalseWhen enabled, uses a local engine context (e.g., driver-side in Spark) instead of the distributed engine context to compute clustering groups for each partition during clustering plan generation. By default this is disabled, meaning the distributed engine context is used (e.g., with Spark, each partition's clustering groups are computed in a separate Spark task). Enable this for cases where there are guaranteed to only be a few partitions with many files in the clustering plan, and it would be more resource-efficient to compute locally on the driver rather than allocate executor resources.Config Param: PLAN_GENERATION_USE_LOCAL_ENGINE_CONTEXTSince Version: 1.2.0
hoodie.clustering.plan.strategy.small.file.limit314572800Files smaller than the size in bytes specified here are candidates for clusteringConfig Param: PLAN_STRATEGY_SMALL_FILE_LIMITSince Version: 0.7.0
hoodie.clustering.plan.strategy.target.file.max.bytes1073741824Each group can produce 'N' (CLUSTERING_MAX_GROUP_SIZE/CLUSTERING_TARGET_FILE_SIZE) output file groupsConfig Param: PLAN_STRATEGY_TARGET_FILE_MAX_BYTESSince Version: 0.7.0
Compaction Configs
Configurations that control compaction (merging of log files onto a new base files).
Config NameDefaultDescription
hoodie.compact.inlinefalseWhen set to true, compaction service is triggered after each write. While being simpler operationally, this adds extra latency on the write path.Config Param: INLINE_COMPACT
hoodie.compact.inline.max.delta.commits5Number of delta commits after the last compaction, before scheduling of a new compaction is attempted. This config takes effect only for the compaction triggering strategy based on the number of commits, i.e., NUM_COMMITS, NUM_COMMITS_AFTER_LAST_REQUEST, NUM_AND_TIME, and NUM_OR_TIME.Config Param: INLINE_COMPACT_NUM_DELTA_COMMITS
Error table Configs
Configurations that are required for Error table configs
Config NameDefaultDescription
hoodie.errortable.base.path(N/A)Base path for error table under which all error records would be stored.Config Param: ERROR_TABLE_BASE_PATH
hoodie.errortable.target.table.name(N/A)Table name to be used for the error tableConfig Param: ERROR_TARGET_TABLE
hoodie.errortable.write.class(N/A)Class which handles the error table writes. This config is used to configure a custom implementation for Error Table Writer. Specify the full class name of the custom error table writer as a value for this configConfig Param: ERROR_TABLE_WRITE_CLASS
hoodie.errortable.enablefalseConfig to enable error table. If the config is enabled, all the records with processing error in DeltaStreamer are transferred to error table.Config Param: ERROR_TABLE_ENABLED
hoodie.errortable.insert.shuffle.parallelism200Config to set insert shuffle parallelism. The config is similar to hoodie.insert.shuffle.parallelism config but applies to the error table.Config Param: ERROR_TABLE_INSERT_PARALLELISM_VALUE
hoodie.errortable.source.rdd.persistfalseEnabling this config, persists the sourceRDD to disk which helps in faster processing of data table + error table write DAGConfig Param: ERROR_TABLE_PERSIST_SOURCE_RDD
hoodie.errortable.upsert.shuffle.parallelism200Config to set upsert shuffle parallelism. The config is similar to hoodie.upsert.shuffle.parallelism config but applies to the error table.Config Param: ERROR_TABLE_UPSERT_PARALLELISM_VALUE
hoodie.errortable.validate.recordcreation.enabletrueRecords that fail to be created due to keygeneration failure or other issues will be sent to the Error TableConfig Param: ERROR_ENABLE_VALIDATE_RECORD_CREATIONSince Version: 0.15.0
hoodie.errortable.validate.targetschema.enablefalseRecords with schema mismatch with Target Schema are sent to Error Table.Config Param: ERROR_ENABLE_VALIDATE_TARGET_SCHEMA
hoodie.errortable.write.failure.strategyROLLBACK_COMMITThe config specifies the failure strategy if error table write fails. Use one of - [ROLLBACK_COMMIT (Rollback the corresponding base table write commit for which the error events were triggered) , LOG_ERROR (Error is logged but the base table write succeeds) ]Config Param: ERROR_TABLE_WRITE_FAILURE_STRATEGY
hoodie.errortable.write.union.enablefalseEnable error table union with data table when writing for improved commit performance. By default it is disabled meaning data table and error table writes are sequentialConfig Param: ENABLE_ERROR_TABLE_WRITE_UNIFICATION
Write Configurations
Configurations that control write behavior on Hudi tables. These can be directly passed down from even higher level frameworks (e.g Spark datasources, Flink sink) and utilities (e.g Hudi Streamer).
Config NameDefaultDescription
hoodie.base.path(N/A)Base path on lake storage, under which all the table data is stored. Always prefix it explicitly with the storage scheme (e.g hdfs://, s3:// etc). Hudi stores all the main meta-data about commits, savepoints, cleaning audit logs etc in .hoodie directory under this base path directory.Config Param: BASE_PATH
hoodie.datasource.write.precombine.field(N/A)Comma separated list of fields used in preCombining before actual write. When two records have the same key value, we will pick the one with the largest value for the precombine field, determined by Object.compareTo(..). For multiple fields if first key comparison is same, second key comparison is made and so on. This config is used for combining records within the same batch and also for merging using event time merge modeConfig Param: PRECOMBINE_FIELD_NAME
hoodie.table.name(N/A)Table name that will be used for registering with metastores like HMS. Needs to be same across runs.Config Param: TBL_NAME
hoodie.write.record.merge.mode(N/A)org.apache.hudi.common.config.RecordMergeMode: Determines the logic of merging updates COMMIT_TIME_ORDERING: Using transaction time to merge records, i.e., the record from later transaction overwrites the earlier record with the same key. EVENT_TIME_ORDERING: Using event time as the ordering to merge records, i.e., the record with the larger event time overwrites the record with the smaller event time on the same key, regardless of transaction time. The event time or ordering fields need to be specified by the user. CUSTOM: Using custom merging logic specified by the user.Config Param: RECORD_MERGE_MODESince Version: 1.0.0
hoodie.block.writes.on.speculative.executiontrueWhen enabled (default), throws an exception if Spark speculative execution is enabled during the client creation. This prevents potential data corruption due to duplicate writes by speculative executors. Set to false only if you understand the risks of running with Spark speculative execution enabled.Config Param: BLOCK_WRITES_ON_SPECULATIVE_EXECUTIONSince Version: 0.14.0
hoodie.fail.job.on.duplicate.data.file.detectionfalseIf config is enabled, entire job is failed on invalid file detectionConfig Param: FAIL_JOB_ON_DUPLICATE_DATA_FILE_DETECTION
hoodie.write.auto.upgradetrueIf enabled, writers automatically migrate the table to the specified write table version if the current table version is lower.Config Param: AUTO_UPGRADE_VERSIONSince Version: 1.0.0
hoodie.write.can.ignore.post.commit.failuresfalseWhen this config is true, any failures in post-commit operations are ignored and do not kill the application.Config Param: CAN_IGNORE_POST_COMMIT_FAILURES
hoodie.write.concurrency.modeSINGLE_WRITERorg.apache.hudi.common.model.WriteConcurrencyMode: Concurrency modes for write operations. SINGLE_WRITER(default): Only one active writer to the table. Maximizes throughput. OPTIMISTIC_CONCURRENCY_CONTROL: Multiple writers can operate on the table with lazy conflict resolution using locks. This means that only one writer succeeds if multiple writers write to the same file group. NON_BLOCKING_CONCURRENCY_CONTROL: Multiple writers can operate on the table with non-blocking conflict resolution. The writers can write into the same file group with the conflicts resolved automatically by the query reader and the compactor.Config Param: WRITE_CONCURRENCY_MODE
hoodie.write.ignore.failedtrueFlag to indicate whether to ignore any non exception error (e.g. write status error).By default true for backward compatibility.Config Param: IGNORE_FAILEDSince Version:
hoodie.write.table.version9The table version this writer is storing the table in. This should match the current table version.Config Param: WRITE_TABLE_VERSIONSince Version: 1.0.0
Lock Configs
Configurations that control locking mechanisms required for concurrency control between writers to a Hudi table. Concurrency between Hudi's own table services are auto managed internally.
Common Lock Configurations
Config NameDefaultDescription
hoodie.write.lock.heartbeat_interval_ms60000Heartbeat interval in ms, to send a heartbeat to indicate that hive client holding locks.Config Param: LOCK_HEARTBEAT_INTERVAL_MSSince Version: 0.15.0
Key Generator Configs
Hudi maintains keys (record key + partition path) for uniquely identifying a particular record. These configs allow developers to setup the Key generator class that extracts these out of incoming records.
Key Generator Options
Config NameDefaultDescription
hoodie.datasource.write.partitionpath.field(N/A)Partition path field. Value to be used at the partitionPath component of HoodieKey. Actual value obtained by invoking .toString()Config Param: PARTITIONPATH_FIELD_NAME
hoodie.datasource.write.recordkey.field(N/A)Record key field. Value to be used as the recordKey component of HoodieKey. Actual value will be obtained by invoking .toString() on the field value. Nested fields can be specified using the dot notation eg: a.b.cConfig Param: RECORDKEY_FIELD_NAME
hoodie.datasource.write.secondarykey.column(N/A)Columns that constitute the secondary key component. Actual value will be obtained by invoking .toString() on the field value. Nested fields can be specified using the dot notation eg: a.b.cConfig Param: SECONDARYKEY_COLUMN_NAME
hoodie.datasource.write.hive_style_partitioningfalseFlag to indicate whether to use Hive style partitioning. If set true, the names of partition folders follow <partition_column_name>=<partition_value> format. By default false (the names of partition folders are only partition values)Config Param: HIVE_STYLE_PARTITIONING_ENABLE
Index Configs
Configurations that control indexing behavior, which tags incoming records as either inserts or updates to older records.
Common Index Configs
Config NameDefaultDescription
hoodie.expression.index.function(N/A)Function to be used for building the expression index.Config Param: INDEX_FUNCTIONSince Version: 1.0.0
hoodie.index.name(N/A)Name of the expression index. This is also used for the partition name in the metadata table.Config Param: INDEX_NAMESince Version: 1.0.0
hoodie.table.checksum(N/A)Index definition checksum is used to guard against partial writes in HDFS. It is added as the last entry in index.properties and then used to validate while reading table config.Config Param: INDEX_DEFINITION_CHECKSUMSince Version: 1.0.0
hoodie.expression.index.typeCOLUMN_STATSType of the expression index. Default is column_stats if there are no functions and expressions in the command. Valid options could be BITMAP, COLUMN_STATS, LUCENE, etc. If index_type is not provided, and there are functions or expressions in the command then a expression index using column stats will be created.Config Param: INDEX_TYPESince Version: 1.0.0
Common Index Configs
Config NameDefaultDescription
hoodie.index.type(N/A)org.apache.hudi.index.HoodieIndex$IndexType: Determines how input records are indexed, i.e., looked up based on the key for the location in the existing table. Default is SIMPLE on Spark engine, and INMEMORY on Flink and Java engines. INMEMORY: Uses in-memory hashmap in Spark and Java engine and Flink in-memory state in Flink for indexing. BLOOM: Employs bloom filters built out of the record keys, optionally also pruning candidate files using record key ranges. Key uniqueness is enforced inside partitions. GLOBAL_BLOOM: Employs bloom filters built out of the record keys, optionally also pruning candidate files using record key ranges. Key uniqueness is enforced across all partitions in the table. SIMPLE: Performs a lean join of the incoming update/delete records against keys extracted from the table on storage.Key uniqueness is enforced inside partitions. GLOBAL_SIMPLE: Performs a lean join of the incoming update/delete records against keys extracted from the table on storage.Key uniqueness is enforced across all partitions in the table. BUCKET: locates the file group containing the record fast by using bucket hashing, particularly beneficial in large scale. Use hoodie.index.bucket.engine to choose bucket engine type, i.e., how buckets are generated. FLINK_STATE: Internal Config for indexing based on Flink state. RECORD_INDEX: Index which saves the record key to location mappings in the HUDI Metadata Table. Record index is a global index, enforcing key uniqueness across all partitions in the table. Supports sharding to achieve very high scale. For a table with keys that are only unique inside each partition, use RECORD_LEVEL_INDEX instead. This enum is deprecated. Use GLOBAL_RECORD_LEVEL_INDEX for global uniqueness of record keys or RECORD_LEVEL_INDEX for partition-level uniqueness of record keys. GLOBAL_RECORD_LEVEL_INDEX: Index which saves the record key to location mappings in the HUDI Metadata Table. Record index is a global index, enforcing key uniqueness across all partitions in the table. Supports sharding to achieve very high scale. For a table with keys that are only unique inside each partition, use RECORD_LEVEL_INDEX instead. RECORD_LEVEL_INDEX: Index which saves the record key to location mappings in the HUDI Metadata Table. Supports sharding to achieve very high scale. This is a non global index, where keys can be replicated across partitions, since a pair of partition path and record keys will uniquely map to a location using this index. If users expect record keys to be unique across all partitions, use GLOBAL_RECORD_LEVEL_INDEX instead.Config Param: INDEX_TYPE
hoodie.bucket.index.query.pruningtrueControl if table with bucket index use bucket query or notConfig Param: BUCKET_QUERY_INDEX
hoodie.bucket.index.remote.partitioner.enablefalseUse Remote Partitioner using centralized allocation of partition IDs to do repartition based on bucket aiming to resolve data skew. Default local hash partitionerConfig Param: BUCKET_PARTITIONER
Metastore and Catalog Sync Configs
Configurations used by the Hudi to sync metadata to external metastores and catalogs.
Common Metadata Sync Configs
Config NameDefaultDescription
hoodie.datasource.meta.sync.enablefalseEnable Syncing the Hudi Table with an external meta store or data catalog.Config Param: META_SYNC_ENABLED
Glue catalog sync based client Configurations
Configs that control Glue catalog sync based client.
Config NameDefaultDescription
hoodie.datasource.meta.sync.glue.partition_index_fieldsSpecify the partitions fields to index on aws glue. Separate the fields by semicolon. By default, when the feature is enabled, all the partition will be indexed. You can create up to three indexes, separate them by comma. Eg: col1;col2;col3,col2,col3Config Param: META_SYNC_PARTITION_INDEX_FIELDSSince Version: 0.15.0
hoodie.datasource.meta.sync.glue.partition_index_fields.enablefalseEnable aws glue partition index feature, to speedup partition based query patternConfig Param: META_SYNC_PARTITION_INDEX_FIELDS_ENABLESince Version: 0.15.0
BigQuery Sync Configs
Configurations used by the Hudi to sync metadata to Google BigQuery.
Config NameDefaultDescription
hoodie.datasource.meta.sync.enablefalseEnable Syncing the Hudi Table with an external meta store or data catalog.Config Param: META_SYNC_ENABLED
Hive Sync Configs
Configurations used by the Hudi to sync metadata to Hive Metastore.
Config NameDefaultDescription
hoodie.datasource.hive_sync.mode(N/A)Mode to choose for Hive ops. Valid values are hms, jdbc and hiveql.Config Param: HIVE_SYNC_MODE
hoodie.datasource.hive_sync.enablefalseWhen set to true, register/sync the table to Apache Hive metastore.Config Param: HIVE_SYNC_ENABLED
hoodie.datasource.hive_sync.jdbcurljdbc:hive2://localhost:10000Hive metastore urlConfig Param: HIVE_URL
hoodie.datasource.hive_sync.metastore.uristhrift://localhost:9083Hive metastore urlConfig Param: METASTORE_URIS
hoodie.datasource.meta.sync.enablefalseEnable Syncing the Hudi Table with an external meta store or data catalog.Config Param: META_SYNC_ENABLED
Global Hive Sync Configs
Global replication configurations used by the Hudi to sync metadata to Hive Metastore.
Config NameDefaultDescription
hoodie.datasource.hive_sync.mode(N/A)Mode to choose for Hive ops. Valid values are hms, jdbc and hiveql.Config Param: HIVE_SYNC_MODE
hoodie.datasource.hive_sync.enablefalseWhen set to true, register/sync the table to Apache Hive metastore.Config Param: HIVE_SYNC_ENABLED
hoodie.datasource.hive_sync.jdbcurljdbc:hive2://localhost:10000Hive metastore urlConfig Param: HIVE_URL
hoodie.datasource.hive_sync.metastore.uristhrift://localhost:9083Hive metastore urlConfig Param: METASTORE_URIS
hoodie.datasource.meta.sync.enablefalseEnable Syncing the Hudi Table with an external meta store or data catalog.Config Param: META_SYNC_ENABLED
DataHub Sync Configs
Configurations used by the Hudi to sync metadata to DataHub.
Config NameDefaultDescription
hoodie.datasource.meta.sync.enablefalseEnable Syncing the Hudi Table with an external meta store or data catalog.Config Param: META_SYNC_ENABLED
Metrics Configs
These set of configs are used to enable monitoring and reporting of key Hudi stats and metrics.
Metrics Configurations
Enables reporting on Hudi metrics. Hudi publishes metrics on every commit, clean, rollback etc. The following sections list the supported reporters.
Config NameDefaultDescription
hoodie.metrics.onfalseTurn on/off metrics reporting. off by default.Config Param: TURN_METRICS_ONSince Version: 0.5.0
hoodie.metrics.reporter.typeGRAPHITEType of metrics reporter.Config Param: METRICS_REPORTER_TYPE_VALUESince Version: 0.5.0
hoodie.metricscompaction.log.blocks.onfalseTurn on/off metrics reporting for log blocks with compaction commit. off by default.Config Param: TURN_METRICS_COMPACTION_LOG_BLOCKS_ONSince Version: 0.14.0
Metrics Configurations for M3
Enables reporting on Hudi metrics using M3. Hudi publishes metrics on every commit, clean, rollback etc.
Config NameDefaultDescription
hoodie.metrics.m3.envproductionM3 tag to label the environment (defaults to 'production'), applied to all metrics.Config Param: M3_ENVSince Version: 0.15.0
hoodie.metrics.m3.hostlocalhostM3 host to connect to.Config Param: M3_SERVER_HOST_NAMESince Version: 0.15.0
hoodie.metrics.m3.port9052M3 port to connect to.Config Param: M3_SERVER_PORT_NUMSince Version: 0.15.0
hoodie.metrics.m3.servicehoodieM3 tag to label the service name (defaults to 'hoodie'), applied to all metrics.Config Param: M3_SERVICESince Version: 0.15.0
hoodie.metrics.m3.tagsOptional M3 tags applied to all metrics.Config Param: M3_TAGSSince Version: 0.15.0
Kafka Connect Configs
These set of configs are used for Kafka Connect Sink Connector for writing Hudi Tables
Kafka Sink Connect Configurations
Configurations for Kafka Connect Sink Connector for Hudi.
Config NameDefaultDescription
bootstrap.serverslocalhost:9092The bootstrap servers for the Kafka Cluster.Config Param: KAFKA_BOOTSTRAP_SERVERS
Hudi Streamer Configs
These set of configs are used for Hudi Streamer utility which provides the way to ingest from different sources such as DFS or Kafka.
Hudi Streamer Configs
Config NameDefaultDescription
hoodie.streamer.source.kafka.topic(N/A)Kafka topic name. The config is specific to HoodieMultiTableStreamerConfig Param: KAFKA_TOPIC
Hudi Streamer SQL Transformer Configs
Configurations controlling the behavior of SQL transformer in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.transformer.sql(N/A)SQL Query to be executed during writeConfig Param: TRANSFORMER_SQL
hoodie.streamer.transformer.sql.file(N/A)File with a SQL script to be executed during writeConfig Param: TRANSFORMER_SQL_FILE
Hudi Streamer Source Configs
Configurations controlling the behavior of reading source data.
DFS Path Selector Configs
Configurations controlling the behavior of path selector for DFS source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.dfs.root(N/A)Root path of the source on DFSConfig Param: ROOT_INPUT_PATH
Hudi Incremental Source Configs
Configurations controlling the behavior of incremental pulling from a Hudi table as a source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.hoodieincr.path(N/A)Base-path for the source Hudi tableConfig Param: HOODIE_SRC_BASE_PATH
Kafka Source Configs
Configurations controlling the behavior of Kafka source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.kafka.topic(N/A)Kafka topic name.Config Param: KAFKA_TOPIC_NAME
hoodie.streamer.source.kafka.proto.value.deserializer.classorg.apache.kafka.common.serialization.ByteArrayDeserializerKafka Proto Payload Deserializer ClassConfig Param: KAFKA_PROTO_VALUE_DESERIALIZER_CLASSSince Version: 0.15.0
Kinesis Source Configs
Configurations controlling the behavior of Kinesis source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.kinesis.stream.name(N/A)Kinesis Data Streams stream name.Config Param: KINESIS_STREAM_NAMESince Version: 1.2.0
Pulsar Source Configs
Configurations controlling the behavior of Pulsar source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.pulsar.topic(N/A)Name of the target Pulsar topic to source data fromConfig Param: PULSAR_SOURCE_TOPIC_NAME
hoodie.streamer.source.pulsar.endpoint.admin.urlhttp://localhost:8080URL of the target Pulsar endpoint (of the form 'pulsar://host:port'Config Param: PULSAR_SOURCE_ADMIN_ENDPOINT_URL
hoodie.streamer.source.pulsar.endpoint.service.urlpulsar://localhost:6650URL of the target Pulsar endpoint (of the form 'pulsar://host:port'Config Param: PULSAR_SOURCE_SERVICE_ENDPOINT_URL
S3 Source Configs
Configurations controlling the behavior of S3 source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.s3.source.queue.url(N/A)Queue url for cloud object eventsConfig Param: S3_SOURCE_QUEUE_URL
File-based SQL Source Configs
Configurations controlling the behavior of File-based SQL Source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.sql.file(N/A)SQL file path containing the SQL query to read source data.Config Param: SOURCE_SQL_FILESince Version: 0.14.0
SQL Source Configs
Configurations controlling the behavior of SQL source in Hudi Streamer.
Config NameDefaultDescription
hoodie.streamer.source.sql.sql.query(N/A)SQL query for fetching source data.Config Param: SOURCE_SQL
Hudi Streamer Schema Provider Configs
Configurations that control the schema provider for Hudi Streamer.
Hudi Streamer Schema Provider Configs
Config NameDefaultDescription
hoodie.streamer.schemaprovider.registry.targetUrl(N/A)The schema of the target you are writing to e.g. https://foo:bar@schemaregistry.orgConfig Param: TARGET_SCHEMA_REGISTRY_URL
hoodie.streamer.schemaprovider.registry.url(N/A)The schema of the source you are reading from e.g. https://foo:bar@schemaregistry.orgConfig Param: SRC_SCHEMA_REGISTRY_URL
File-based Schema Provider Configs
Configurations for file-based schema provider.
Config NameDefaultDescription
hoodie.streamer.schemaprovider.source.schema.file(N/A)The schema of the source you are reading fromConfig Param: SOURCE_SCHEMA_FILE
hoodie.streamer.schemaprovider.target.schema.file(N/A)The schema of the target you are writing toConfig Param: TARGET_SCHEMA_FILE
评论
登录后参与评论
KnowForge