Spark configurations status in Gluten Bolt Backend

qianmoQqianmoQ· 更新于 2026-10-02· 阅读 81 分钟· 0 次阅读

登录后可跨设备保存划线和私人笔记登录

The file lists the if Spark configurations are hornored by Gluten Bolt backend or not. Table is from Spark4.0 configuration page. The status are:

  • H: hornored
  • P: Transparent to Gluten
  • I: ignored. Gluten doesn’t use it.
  • <blank>: unknown yet

Application Properties

Property NameDefaultSince VersionGluten Status
spark.app.name(none)0.9.0
spark.driver.cores11.3.0
spark.driver.maxResultSize1g1.2.0
spark.driver.memory1g1.1.1
spark.driver.memoryOverheaddriverMemory * spark.driver.memoryOverheadFactor, with minimum of spark.driver.minMemoryOverhead2.3.0
spark.driver.minMemoryOverhead384m4.0.0
spark.driver.memoryOverheadFactor0.103.3.0
spark.driver.resource.{resourceName}.amount03.0.0
spark.driver.resource.{resourceName}.discoveryScriptNone3.0.0
spark.driver.resource.{resourceName}.vendorNone3.0.0
spark.resources.discoveryPluginorg.apache.spark.resource.ResourceDiscoveryScriptPlugin3.0.0
spark.executor.memory1g0.7.0
spark.executor.pyspark.memoryNot set2.4.0
spark.executor.memoryOverheadexecutorMemory * spark.executor.memoryOverheadFactor, with minimum of spark.executor.minMemoryOverhead2.3.0
spark.executor.minMemoryOverhead384m4.0.0
spark.executor.memoryOverheadFactor0.103.3.0
spark.executor.resource.{resourceName}.amount03.0.0
spark.executor.resource.{resourceName}.discoveryScriptNone3.0.0
spark.executor.resource.{resourceName}.vendorNone3.0.0
spark.extraListeners(none)1.3.0
spark.local.dir/tmp0.5.0
spark.logConffalse0.9.0
spark.master(none)0.9.0
spark.submit.deployModeclient1.5.0
spark.log.callerContext(none)2.2.0
spark.log.level(none)3.5.0
spark.driver.supervisefalse1.3.0
spark.driver.timeout0min4.0.0
spark.driver.log.localDir(none)4.0.0
spark.driver.log.dfsDir(none)3.0.0
spark.driver.log.persistToDfs.enabledfalse3.0.0
spark.driver.log.layout%d{yy/MM/dd HHss.SSS} %t %p %c{1}: %m%n%ex3.0.0
spark.driver.log.allowErasureCodingfalse3.0.0
spark.decommission.enabledfalse3.1.0
spark.executor.decommission.killInterval(none)3.1.0
spark.executor.decommission.forceKillTimeout(none)3.2.0
spark.executor.decommission.signalPWR3.2.0
spark.executor.maxNumFailuresnumExecutors * 2, with minimum of 33.5.0
spark.executor.failuresValidityInterval(none)3.5.0

Runtime Environment

Property NameDefaultSince VersionGluten Status
spark.driver.extraClassPath(none)1.0.0
spark.driver.defaultJavaOptions(none)3.0.0
spark.driver.extraJavaOptions(none)1.0.0
spark.driver.extraLibraryPath(none)1.0.0
spark.driver.userClassPathFirstfalse1.3.0
spark.executor.extraClassPath(none)1.0.0
spark.executor.defaultJavaOptions(none)3.0.0
spark.executor.extraJavaOptions(none)1.0.0
spark.executor.extraLibraryPath(none)1.0.0
spark.executor.logs.rolling.maxRetainedFiles-11.1.0
spark.executor.logs.rolling.enableCompressionfalse2.0.2
spark.executor.logs.rolling.maxSize1024 * 10241.4.0
spark.executor.logs.rolling.strategy"" (disabled)1.1.0
spark.executor.logs.rolling.time.intervaldaily1.1.0
spark.executor.userClassPathFirstfalse1.3.0
spark.executorEnv.[EnvironmentVariableName](none)0.9.0
spark.redaction.regex(?i)secret\password\token\access[.]?key2.1.2
spark.redaction.string.regex(none)2.2.0
spark.python.profilefalse1.2.0
spark.python.profile.dump(none)1.2.0
spark.python.worker.memory512m1.1.0
spark.python.worker.reusetrue1.2.0
spark.files1.0.0
spark.submit.pyFiles1.0.1
spark.jars0.9.0
spark.jars.packages1.5.0
spark.jars.excludes1.5.0
spark.jars.ivy1.3.0
spark.jars.ivySettings2.2.0
spark.jars.repositories2.3.0
spark.archives3.1.0
spark.pyspark.driver.python2.1.0
spark.pyspark.python2.1.0

Shuffle Behavior

Property NameDefaultSince VersionGluten Status
spark.reducer.maxSizeInFlight48m1.4.0
spark.reducer.maxReqsInFlightInt.MaxValue2.0.0
spark.reducer.maxBlocksInFlightPerAddressInt.MaxValue2.2.1
spark.shuffle.compresstrue0.6.0
spark.shuffle.file.buffer32k1.4.0
spark.shuffle.file.merge.buffer32k4.0.0
spark.shuffle.unsafe.file.output.buffer32k2.3.0
spark.shuffle.localDisk.file.output.buffer32k4.0.0
spark.shuffle.spill.diskWriteBufferSize1024 * 10242.3.0
spark.shuffle.io.maxRetries31.2.0
spark.shuffle.io.numConnectionsPerPeer11.2.1
spark.shuffle.io.preferDirectBufstrue1.2.0
spark.shuffle.io.retryWait5s1.2.1
spark.shuffle.io.backLog-11.1.1
spark.shuffle.io.connectionTimeoutvalue of spark.network.timeout1.2.0
spark.shuffle.io.connectionCreationTimeoutvalue of spark.shuffle.io.connectionTimeout3.2.0
spark.shuffle.service.enabledfalse1.2.0
spark.shuffle.service.port73371.2.0
spark.shuffle.service.namespark_shuffle3.2.0
spark.shuffle.service.index.cache.size100m2.3.0
spark.shuffle.service.removeShuffletrue3.3.0
spark.shuffle.maxChunksBeingTransferredLong.MAX_VALUE2.3.0
spark.shuffle.sort.bypassMergeThreshold2001.1.1
spark.shuffle.sort.io.plugin.classorg.apache.spark.shuffle.sort.io.LocalDiskShuffleDataIO3.0.0
spark.shuffle.spill.compresstrue0.9.0
spark.shuffle.accurateBlockThreshold100 * 1024 * 10242.2.1
spark.shuffle.accurateBlockSkewedFactor-1.03.3.0
spark.shuffle.registration.timeout50002.3.0
spark.shuffle.registration.maxAttempts32.3.0
spark.shuffle.reduceLocality.enabledtrue1.5.0
spark.shuffle.mapOutput.minSizeForBroadcast512k2.0.0
spark.shuffle.detectCorrupttrue2.2.0
spark.shuffle.detectCorrupt.useExtraMemoryfalse3.0.0
spark.shuffle.useOldFetchProtocolfalse3.0.0
spark.shuffle.readHostLocalDisktrue3.0.0
spark.files.io.connectionTimeoutvalue of spark.network.timeout1.6.0
spark.files.io.connectionCreationTimeoutvalue of spark.files.io.connectionTimeout3.2.0
spark.shuffle.checksum.enabledtrue3.2.0
spark.shuffle.checksum.algorithmADLER323.2.0
spark.shuffle.service.fetch.rdd.enabledfalse3.0.0
spark.shuffle.service.db.enabledtrue3.0.0
spark.shuffle.service.db.backendROCKSDB3.4.0

Spark UI

Property NameDefaultSince VersionGluten Status
spark.eventLog.logBlockUpdates.enabledfalse2.3.0
spark.eventLog.longForm.enabledfalse2.4.0
spark.eventLog.compresstrue1.0.0
spark.eventLog.compression.codeczstd3.0.0
spark.eventLog.erasureCoding.enabledfalse3.0.0
spark.eventLog.dirfile:///tmp/spark-events1.0.0
spark.eventLog.enabledfalse1.0.0
spark.eventLog.overwritefalse1.0.0
spark.eventLog.buffer.kb100k1.0.0
spark.eventLog.rolling.enabledfalse3.0.0
spark.eventLog.rolling.maxFileSize128m3.0.0
spark.ui.dagGraph.retainedRootRDDsInt.MaxValue2.1.0
spark.ui.groupSQLSubExecutionEnabledtrue3.4.0
spark.ui.enabledtrue1.1.1
spark.ui.store.pathNone3.4.0
spark.ui.killEnabledtrue1.0.0
spark.ui.threadDumpsEnabledtrue1.2.0
spark.ui.threadDump.flamegraphEnabledtrue4.0.0
spark.ui.heapHistogramEnabledtrue3.5.0
spark.ui.liveUpdate.period100ms2.3.0
spark.ui.liveUpdate.minFlushPeriod1s2.4.2
spark.ui.port40400.7.0
spark.ui.retainedJobs10001.2.0
spark.ui.retainedStages10000.9.0
spark.ui.retainedTasks1000002.0.1
spark.ui.reverseProxyfalse2.1.0
spark.ui.reverseProxyUrl2.1.0
spark.ui.proxyRedirectUri3.0.0
spark.ui.showConsoleProgressfalse1.2.1
spark.ui.consoleProgress.update.interval2002.1.0
spark.ui.custom.executor.log.url(none)3.0.0
spark.ui.prometheus.enabledtrue3.0.0
spark.worker.ui.retainedExecutors10001.5.0
spark.worker.ui.retainedDrivers10001.5.0
spark.sql.ui.retainedExecutions10001.5.0
spark.streaming.ui.retainedBatches10001.0.0
spark.ui.retainedDeadExecutors1002.0.0
spark.ui.filtersNone1.0.0
spark.ui.requestHeaderSize8k2.2.3
spark.ui.timelineEnabledtrue3.4.0
spark.ui.timeline.executors.maximum2503.2.0
spark.ui.timeline.jobs.maximum5003.2.0
spark.ui.timeline.stages.maximum5003.2.0
spark.ui.timeline.tasks.maximum10001.4.0
spark.appStatusStore.diskStoreDirNone3.4.0

Compression and Serialization

Property NameDefaultSince VersionGluten Status

spark.broadcast.compress true 0.6.0

spark.checkpoint.dir (none) 4.0.0

spark.checkpoint.compress false 2.2.0

spark.io.compression.codec lz4 0.8.0

spark.io.compression.lz4.blockSize 32k 1.4.0

spark.io.compression.snappy.blockSize 32k 1.4.0

spark.io.compression.zstd.level 1 2.3.0

spark.io.compression.zstd.bufferSize 32k 2.3.0

spark.io.compression.zstd.bufferPool.enabled true 3.2.0

spark.io.compression.zstd.workers 0 4.0.0

spark.io.compression.lzf.parallel.enabled false 4.0.0

spark.kryo.classesToRegister (none) 1.2.0

spark.kryo.referenceTracking true 0.8.0

spark.kryo.registrationRequired false 1.1.0

spark.kryo.registrator (none) 0.5.0

spark.kryo.unsafe true 2.1.0

spark.kryoserializer.buffer.max 64m 1.4.0

spark.kryoserializer.buffer 64k 1.4.0

spark.rdd.compress false 0.6.0

spark.serializer org.apache.spark.serializer.
JavaSerializer 0.5.0

spark.serializer.objectStreamReset 100 1.0.0

Memory Management

Property NameDefaultSince VersionGluten Status
spark.memory.fraction0.61.6.0
spark.memory.storageFraction0.51.6.0
spark.memory.offHeap.enabledfalse1.6.0
spark.memory.offHeap.size01.6.0
spark.storage.unrollMemoryThreshold1024 * 10241.1.0
spark.storage.replication.proactivetrue2.2.0
spark.storage.localDiskByExecutors.cacheSize10003.0.0
spark.cleaner.periodicGC.interval30min1.6.0
spark.cleaner.referenceTrackingtrue1.0.0
spark.cleaner.referenceTracking.blockingtrue1.0.0
spark.cleaner.referenceTracking.blocking.shufflefalse1.1.1
spark.cleaner.referenceTracking.cleanCheckpointsfalse1.4.0

Execution Behavior

Property NameDefaultSince VersionGluten Status

spark.broadcast.blockSize 4m 0.5.0

spark.broadcast.checksum true 2.1.1

spark.broadcast.UDFCompressionThreshold 1 * 1024 * 1024 3.0.0

spark.executor.cores 1 in YARN mode, all the available cores on the worker in standalone mode. 1.0.0

spark.default.parallelism For distributed shuffle operations like reduceByKey and join, the largest number of partitions in a parent RDD. For operations like parallelize with no parent RDDs, it depends on the cluster manager:

  • Local mode: number of cores on the local machine
  • Others: total number of cores on all executor nodes or 2, whichever is larger

0.5.0

spark.executor.heartbeatInterval 10s 1.1.0

spark.files.fetchTimeout 60s 1.0.0

spark.files.useFetchCache true 1.2.2

spark.files.overwrite false 1.0.0

spark.files.ignoreCorruptFiles false 2.1.0

spark.files.ignoreMissingFiles false 2.4.0

spark.files.maxPartitionBytes 134217728 (128 MiB) 2.1.0

spark.files.openCostInBytes 4194304 (4 MiB) 2.1.0

spark.hadoop.cloneConf false 1.0.3

spark.hadoop.validateOutputSpecs true 1.0.1

spark.storage.memoryMapThreshold 2m 0.9.2

spark.storage.decommission.enabled false 3.1.0

spark.storage.decommission.shuffleBlocks.enabled true 3.1.0

spark.storage.decommission.shuffleBlocks.maxThreads 8 3.1.0

spark.storage.decommission.rddBlocks.enabled true 3.1.0

spark.storage.decommission.fallbackStorage.path (none) 3.1.0

spark.storage.decommission.fallbackStorage.cleanUp false 3.2.0

spark.storage.decommission.shuffleBlocks.maxDiskSize (none) 3.2.0

spark.hadoop.mapreduce.fileoutputcommitter.algorithm.version 1 2.2.0

Executor Metrics

These configurations are handled by Spark and do not affect Gluten’s behavior.

Networking

These configurations are handled by Spark and do not affect Gluten’s behavior.

Scheduling

Property NameDefaultSince Version
spark.cores.max(not set)0.6.0
spark.locality.wait3s0.5.0
spark.locality.wait.nodespark.locality.wait0.8.0
spark.locality.wait.processspark.locality.wait0.8.0
spark.locality.wait.rackspark.locality.wait0.8.0
spark.scheduler.maxRegisteredResourcesWaitingTime30s1.1.1
spark.scheduler.minRegisteredResourcesRatio0.8 for KUBERNETES mode; 0.8 for YARN mode; 0.0 for standalone mode1.1.1
spark.scheduler.modeFIFO0.8.0
spark.scheduler.revive.interval1s0.8.1
spark.scheduler.listenerbus.eventqueue.capacity100002.3.0
spark.scheduler.listenerbus.eventqueue.shared.capacityspark.scheduler.listenerbus.eventqueue.capacity3.0.0
spark.scheduler.listenerbus.eventqueue.appStatus.capacityspark.scheduler.listenerbus.eventqueue.capacity3.0.0
spark.scheduler.listenerbus.eventqueue.executorManagement.capacityspark.scheduler.listenerbus.eventqueue.capacity3.0.0
spark.scheduler.listenerbus.eventqueue.eventLog.capacityspark.scheduler.listenerbus.eventqueue.capacity3.0.0
spark.scheduler.listenerbus.eventqueue.streams.capacityspark.scheduler.listenerbus.eventqueue.capacity3.0.0
spark.scheduler.resource.profileMergeConflictsfalse3.1.0
spark.scheduler.excludeOnFailure.unschedulableTaskSetTimeout120s2.4.1
spark.standalone.submit.waitAppCompletionfalse3.1.0
spark.excludeOnFailure.enabledfalse2.1.0
spark.excludeOnFailure.application.enabledfalse4.0.0
spark.excludeOnFailure.taskAndStage.enabledfalse4.0.0
spark.excludeOnFailure.timeout1h2.1.0
spark.excludeOnFailure.task.maxTaskAttemptsPerExecutor12.1.0
spark.excludeOnFailure.task.maxTaskAttemptsPerNode22.1.0
spark.excludeOnFailure.stage.maxFailedTasksPerExecutor22.1.0
spark.excludeOnFailure.stage.maxFailedExecutorsPerNode22.1.0
spark.excludeOnFailure.application.maxFailedTasksPerExecutor22.2.0
spark.excludeOnFailure.application.maxFailedExecutorsPerNode22.2.0
spark.excludeOnFailure.killExcludedExecutorsfalse2.2.0
spark.excludeOnFailure.application.fetchFailure.enabledfalse2.3.0
spark.speculationfalse0.6.0
spark.speculation.interval100ms0.6.0
spark.speculation.multiplier30.6.0
spark.speculation.quantile0.90.6.0
spark.speculation.minTaskRuntime100ms3.2.0
spark.speculation.task.duration.thresholdNone3.0.0
spark.speculation.efficiency.processRateMultiplier0.753.4.0
spark.speculation.efficiency.longRunTaskFactor23.4.0
spark.speculation.efficiency.enabledtrue3.4.0
spark.task.cpus10.5.0
spark.task.resource.{resourceName}.amount13.0.0
spark.task.maxFailures40.8.0
spark.task.reaper.enabledfalse2.0.3
spark.task.reaper.pollingInterval10s2.0.3
spark.task.reaper.threadDumptrue2.0.3
spark.task.reaper.killTimeout-12.0.3
spark.stage.maxConsecutiveAttempts42.2.0
spark.stage.ignoreDecommissionFetchFailuretrue3.4.0

Barrier Execution Mode

These configurations are handled by Spark and do not affect Gluten’s behavior.

Dynamic Allocation

These configurations are handled by Spark and do not affect Gluten’s behavior.

Thread Configurations

These configurations are handled by Spark and do not affect Gluten’s behavior.

Spark Connect

Server Configuration

These configurations are handled by Spark and do not affect Gluten’s behavior.

Security

These configurations are handled by Spark and do not affect Gluten’s behavior.

Spark SQL

Runtime SQL Configuration

Property NameDefaultSince VersionGluten Status
spark.sql.adaptive.advisoryPartitionSizeInBytes(value of spark.sql.adaptive.shuffle.targetPostShuffleInputSize)3.0.0
spark.sql.adaptive.autoBroadcastJoinThreshold(none)3.2.0
spark.sql.adaptive.coalescePartitions.enabledtrue3.0.0
spark.sql.adaptive.coalescePartitions.initialPartitionNum(none)3.0.0
spark.sql.adaptive.coalescePartitions.minPartitionSize1MB3.2.0
spark.sql.adaptive.coalescePartitions.parallelismFirsttrue3.2.0
spark.sql.adaptive.customCostEvaluatorClass(none)3.2.0
spark.sql.adaptive.enabledtrue1.6.0
spark.sql.adaptive.forceOptimizeSkewedJoinfalse3.3.0
spark.sql.adaptive.localShuffleReader.enabledtrue3.0.0
spark.sql.adaptive.maxShuffledHashJoinLocalMapThreshold0b3.2.0
spark.sql.adaptive.optimizeSkewsInRebalancePartitions.enabledtrue3.2.0
spark.sql.adaptive.optimizer.excludedRules(none)3.1.0
spark.sql.adaptive.rebalancePartitionsSmallPartitionFactor0.23.3.0
spark.sql.adaptive.skewJoin.enabledtrue3.0.0
spark.sql.adaptive.skewJoin.skewedPartitionFactor5.03.0.0
spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes256MB3.0.0
spark.sql.allowNamedFunctionArgumentstrue3.5.0
spark.sql.ansi.doubleQuotedIdentifiersfalse3.4.0
spark.sql.ansi.enabledtrue3.0.0
spark.sql.ansi.enforceReservedKeywordsfalse3.3.0
spark.sql.ansi.relationPrecedencefalse3.4.0
spark.sql.autoBroadcastJoinThreshold10MB1.1.0
spark.sql.avro.compression.codecsnappy2.4.0
spark.sql.avro.deflate.level-12.4.0
spark.sql.avro.filterPushdown.enabledtrue3.1.0
spark.sql.avro.xz.level64.0.0
spark.sql.avro.zstandard.bufferPool.enabledfalse4.0.0
spark.sql.avro.zstandard.level34.0.0
spark.sql.binaryOutputStyle(none)4.0.0
spark.sql.broadcastTimeout3001.3.0
spark.sql.bucketing.coalesceBucketsInJoin.enabledfalse3.1.0
spark.sql.bucketing.coalesceBucketsInJoin.maxBucketRatio43.1.0
spark.sql.catalog.spark_catalogbuiltin3.0.0
spark.sql.cbo.enabledfalse2.2.0
spark.sql.cbo.joinReorder.dp.star.filterfalse2.2.0
spark.sql.cbo.joinReorder.dp.threshold122.2.0
spark.sql.cbo.joinReorder.enabledfalse2.2.0
spark.sql.cbo.planStats.enabledfalse3.0.0
spark.sql.cbo.starSchemaDetectionfalse2.2.0
spark.sql.charAsVarcharfalse3.3.0
spark.sql.chunkBase64String.enabledtrue3.5.2
spark.sql.cli.print.headerfalse3.2.0
spark.sql.columnNameOfCorruptRecord_corrupt_record1.2.0
spark.sql.csv.filterPushdown.enabledtrue3.0.0
spark.sql.datetime.java8API.enabledfalse3.0.0
spark.sql.debug.maxToStringFields253.0.0
spark.sql.defaultCacheStorageLevelMEMORY_AND_DISK4.0.0
spark.sql.defaultCatalogspark_catalog3.0.0
spark.sql.error.messageFormatPRETTY3.4.0
spark.sql.execution.arrow.enabledfalse2.3.0
spark.sql.execution.arrow.fallback.enabledtrue2.4.0
spark.sql.execution.arrow.localRelationThreshold48MB3.4.0
spark.sql.execution.arrow.maxRecordsPerBatch100002.3.0
spark.sql.execution.arrow.pyspark.enabled(value of spark.sql.execution.arrow.enabled)3.0.0
spark.sql.execution.arrow.pyspark.fallback.enabled(value of spark.sql.execution.arrow.fallback.enabled)3.0.0
spark.sql.execution.arrow.pyspark.selfDestruct.enabledfalse3.2.0
spark.sql.execution.arrow.sparkr.enabledfalse3.0.0
spark.sql.execution.arrow.transformWithStateInPandas.maxRecordsPerBatch100004.0.0
spark.sql.execution.arrow.useLargeVarTypesfalse3.5.0
spark.sql.execution.interruptOnCanceltrue4.0.0
spark.sql.execution.pandas.inferPandasDictAsMapfalse4.0.0
spark.sql.execution.pandas.structHandlingModelegacy3.5.0
spark.sql.execution.pandas.udf.buffer.size(value of spark.buffer.size)3.0.0
spark.sql.execution.pyspark.udf.faulthandler.enabled(value of spark.python.worker.faulthandler.enabled)4.0.0
spark.sql.execution.pyspark.udf.hideTraceback.enabledfalse4.0.0
spark.sql.execution.pyspark.udf.idleTimeoutSeconds(value of spark.python.worker.idleTimeoutSeconds)4.0.0
spark.sql.execution.pyspark.udf.simplifiedTraceback.enabledtrue3.1.0
spark.sql.execution.python.udf.buffer.size(value of spark.buffer.size)4.0.0
spark.sql.execution.python.udf.maxRecordsPerBatch1004.0.0
spark.sql.execution.pythonUDF.arrow.concurrency.level(none)4.0.0
spark.sql.execution.pythonUDF.arrow.enabledfalse3.4.0
spark.sql.execution.pythonUDTF.arrow.enabledfalse3.5.0
spark.sql.execution.topKSortFallbackThreshold21474836322.4.0
spark.sql.extendedExplainProviders(none)4.0.0
spark.sql.files.ignoreCorruptFilesfalse2.1.1
spark.sql.files.ignoreInvalidPartitionPathsfalse4.0.0
spark.sql.files.ignoreMissingFilesfalse2.3.0
spark.sql.files.maxPartitionBytes128MB2.0.0
spark.sql.files.maxPartitionNum(none)3.5.0
spark.sql.files.maxRecordsPerFile02.2.0
spark.sql.files.minPartitionNum(none)3.1.0
spark.sql.function.concatBinaryAsStringfalse2.3.0
spark.sql.function.eltOutputAsStringfalse2.3.0
spark.sql.groupByAliasestrue2.2.0
spark.sql.groupByOrdinaltrue2.0.0
spark.sql.hive.convertInsertingPartitionedTabletrue3.0.0
spark.sql.hive.convertInsertingUnpartitionedTabletrue4.0.0
spark.sql.hive.convertMetastoreCtastrue3.0.0
spark.sql.hive.convertMetastoreInsertDirtrue3.3.0
spark.sql.hive.convertMetastoreOrctrue2.0.0
spark.sql.hive.convertMetastoreParquettrue1.1.1
spark.sql.hive.convertMetastoreParquet.mergeSchemafalse1.3.1
spark.sql.hive.dropPartitionByName.enabledfalse3.4.0
spark.sql.hive.filesourcePartitionFileCacheSize2621440002.1.1
spark.sql.hive.manageFilesourcePartitionstrue2.1.1
spark.sql.hive.metastorePartitionPruningtrue1.5.0
spark.sql.hive.metastorePartitionPruningFallbackOnExceptionfalse3.3.0
spark.sql.hive.metastorePartitionPruningFastFallbackfalse3.3.0
spark.sql.hive.thriftServer.asynctrue1.5.0
spark.sql.icu.caseMappings.enabledtrue4.0.0
spark.sql.inMemoryColumnarStorage.batchSize100001.1.1
spark.sql.inMemoryColumnarStorage.compressedtrue1.0.1
spark.sql.inMemoryColumnarStorage.enableVectorizedReadertrue2.3.1
spark.sql.inMemoryColumnarStorage.hugeVectorReserveRatio1.24.0.0
spark.sql.inMemoryColumnarStorage.hugeVectorThreshold-1b4.0.0
spark.sql.json.filterPushdown.enabledtrue3.1.0
spark.sql.json.useUnsafeRowfalse4.0.0
spark.sql.jsonGenerator.ignoreNullFieldstrue3.0.0
spark.sql.leafNodeDefaultParallelism(none)3.2.0
spark.sql.mapKeyDedupPolicyEXCEPTION3.0.0
spark.sql.maven.additionalRemoteRepositorieshttps://maven-central.storage-download.googleapis.com/maven2/3.0.0
spark.sql.maxMetadataStringLength1003.1.0
spark.sql.maxPlanStringLength21474836323.0.0
spark.sql.maxSinglePartitionBytes128m3.4.0
spark.sql.operatorPipeSyntaxEnabledtrue4.0.0
spark.sql.optimizer.avoidCollapseUDFWithExpensiveExprtrue4.0.0
spark.sql.optimizer.collapseProjectAlwaysInlinefalse3.3.0
spark.sql.optimizer.dynamicPartitionPruning.enabledtrue3.0.0
spark.sql.optimizer.enableCsvExpressionOptimizationtrue3.2.0
spark.sql.optimizer.enableJsonExpressionOptimizationtrue3.1.0
spark.sql.optimizer.excludedRules(none)2.4.0
spark.sql.optimizer.runtime.bloomFilter.applicationSideScanSizeThreshold10GB3.3.0
spark.sql.optimizer.runtime.bloomFilter.creationSideThreshold10MB3.3.0
spark.sql.optimizer.runtime.bloomFilter.enabledtrue3.3.0
spark.sql.optimizer.runtime.bloomFilter.expectedNumItems10000003.3.0
spark.sql.optimizer.runtime.bloomFilter.maxNumBits671088643.3.0
spark.sql.optimizer.runtime.bloomFilter.maxNumItems40000003.3.0
spark.sql.optimizer.runtime.bloomFilter.numBits83886083.3.0
spark.sql.optimizer.runtime.rowLevelOperationGroupFilter.enabledtrue3.4.0
spark.sql.optimizer.runtimeFilter.number.threshold103.3.0
spark.sql.orc.aggregatePushdownfalse3.3.0
spark.sql.orc.columnarReaderBatchSize40962.4.0
spark.sql.orc.columnarWriterBatchSize10243.4.0
spark.sql.orc.compression.codeczstd2.3.0
spark.sql.orc.enableNestedColumnVectorizedReadertrue3.2.0
spark.sql.orc.enableVectorizedReadertrue2.3.0
spark.sql.orc.filterPushdowntrue1.4.0
spark.sql.orc.mergeSchemafalse3.0.0
spark.sql.orderByOrdinaltrue2.0.0
spark.sql.parquet.aggregatePushdownfalse3.3.0
spark.sql.parquet.binaryAsStringfalse1.1.1
spark.sql.parquet.columnarReaderBatchSize40962.4.0
spark.sql.parquet.compression.codecsnappy1.1.1
spark.sql.parquet.enableNestedColumnVectorizedReadertrue3.3.0
spark.sql.parquet.enableVectorizedReadertrue2.0.0
spark.sql.parquet.fieldId.read.enabledfalse3.3.0
spark.sql.parquet.fieldId.read.ignoreMissingfalse3.3.0
spark.sql.parquet.fieldId.write.enabledtrue3.3.0
spark.sql.parquet.filterPushdowntrue1.2.0
spark.sql.parquet.inferTimestampNTZ.enabledtrue3.4.0
spark.sql.parquet.int96AsTimestamptrue1.3.0
spark.sql.parquet.int96TimestampConversionfalse2.3.0
spark.sql.parquet.mergeSchemafalse1.5.0
spark.sql.parquet.outputTimestampTypeINT962.3.0
spark.sql.parquet.recordLevelFilter.enabledfalse2.3.0
spark.sql.parquet.respectSummaryFilesfalse1.5.0
spark.sql.parquet.writeLegacyFormatfalse1.6.0
spark.sql.parser.quotedRegexColumnNamesfalse2.3.0
spark.sql.pivotMaxValues100001.6.0
spark.sql.planner.pythonExecution.memory(none)4.0.0
spark.sql.preserveCharVarcharTypeInfofalse4.0.0
spark.sql.pyspark.inferNestedDictAsStruct.enabledfalse3.3.0
spark.sql.pyspark.jvmStacktrace.enabledfalse3.0.0
spark.sql.pyspark.plotting.max_rows10004.0.0
spark.sql.pyspark.udf.profiler(none)4.0.0
spark.sql.readSideCharPaddingtrue3.4.0
spark.sql.redaction.options.regex(?i)url2.2.2
spark.sql.redaction.string.regex(value of spark.redaction.string.regex)2.3.0
spark.sql.repl.eagerEval.enabledfalse2.4.0
spark.sql.repl.eagerEval.maxNumRows202.4.0
spark.sql.repl.eagerEval.truncate202.4.0
spark.sql.scripting.enabledfalse4.0.0
spark.sql.session.localRelationCacheThreshold671088643.5.0
spark.sql.session.timeZone(value of local timezone)2.2.0
spark.sql.shuffle.partitions2001.1.0
spark.sql.shuffleDependency.fileCleanup.enabledfalse4.0.0
spark.sql.shuffleDependency.skipMigration.enabledfalse4.0.0
spark.sql.shuffledHashJoinFactor33.3.0
spark.sql.sources.bucketing.autoBucketedScan.enabledtrue3.1.0
spark.sql.sources.bucketing.enabledtrue2.0.0
spark.sql.sources.bucketing.maxBuckets1000002.4.0
spark.sql.sources.defaultparquet1.3.0
spark.sql.sources.parallelPartitionDiscovery.threshold321.5.0
spark.sql.sources.partitionColumnTypeInference.enabledtrue1.5.0
spark.sql.sources.partitionOverwriteModeSTATIC2.3.0
spark.sql.sources.v2.bucketing.allowCompatibleTransforms.enabledfalse4.0.0
spark.sql.sources.v2.bucketing.allowJoinKeysSubsetOfPartitionKeys.enabledfalse4.0.0
spark.sql.sources.v2.bucketing.enabledfalse3.3.0
spark.sql.sources.v2.bucketing.partiallyClusteredDistribution.enabledfalse3.4.0
spark.sql.sources.v2.bucketing.partition.filter.enabledfalse4.0.0
spark.sql.sources.v2.bucketing.pushPartValues.enabledtrue3.4.0
spark.sql.sources.v2.bucketing.shuffle.enabledfalse4.0.0
spark.sql.sources.v2.bucketing.sorting.enabledfalse4.0.0
spark.sql.stackTracesInDataFrameContext14.0.0
spark.sql.statistics.fallBackToHdfsfalse2.0.0
spark.sql.statistics.histogram.enabledfalse2.3.0
spark.sql.statistics.size.autoUpdate.enabledfalse2.3.0
spark.sql.statistics.updatePartitionStatsInAnalyzeTable.enabledfalse4.0.0
spark.sql.storeAssignmentPolicyANSI3.0.0
spark.sql.streaming.checkpointLocation(none)2.0.0
spark.sql.streaming.continuous.epochBacklogQueueSize100003.0.0
spark.sql.streaming.disabledV2Writers2.3.1
spark.sql.streaming.fileSource.cleaner.numThreads13.0.0
spark.sql.streaming.forceDeleteTempCheckpointLocationfalse3.0.0
spark.sql.streaming.metricsEnabledfalse2.0.2
spark.sql.streaming.multipleWatermarkPolicymin2.4.0
spark.sql.streaming.noDataMicroBatches.enabledtrue2.4.1
spark.sql.streaming.numRecentProgressUpdates1002.1.1
spark.sql.streaming.sessionWindow.merge.sessions.in.local.partitionfalse3.2.0
spark.sql.streaming.stateStore.encodingFormatunsaferow4.0.0
spark.sql.streaming.stateStore.stateSchemaChecktrue3.1.0
spark.sql.streaming.stopActiveRunOnRestarttrue3.0.0
spark.sql.streaming.stopTimeout03.0.0
spark.sql.streaming.transformWithState.stateSchemaVersion34.0.0
spark.sql.thriftServer.interruptOnCancel(value of spark.sql.execution.interruptOnCancel)3.2.0
spark.sql.thriftServer.queryTimeout0ms3.1.0
spark.sql.thriftserver.scheduler.pool(none)1.1.1
spark.sql.thriftserver.ui.retainedSessions2001.4.0
spark.sql.thriftserver.ui.retainedStatements2001.4.0
spark.sql.timeTravelTimestampKeytimestampAsOf4.0.0
spark.sql.timeTravelVersionKeyversionAsOf4.0.0
spark.sql.timestampTypeTIMESTAMP_LTZ3.4.0
spark.sql.transposeMaxValues5004.0.0
spark.sql.tvf.allowMultipleTableArguments.enabledfalse3.5.0
spark.sql.ui.explainModeformatted3.1.0
spark.sql.variable.substitutetrue2.0.0

Static SQL Configuration

Property NameDefaultSince VersionGluten Status
spark.sql.cache.serializerorg.apache.spark.sql.execution.columnar.DefaultCachedBatchSerializer3.1.0
spark.sql.catalog.spark_catalog.defaultDatabasedefault3.4.0
spark.sql.event.truncate.length21474836473.0.0
spark.sql.extensions(none)2.2.0
spark.sql.extensions.test.loadFromCptrue
spark.sql.hive.metastore.barrierPrefixes1.4.0
spark.sql.hive.metastore.jarsbuiltin1.4.0
spark.sql.hive.metastore.jars.path3.1.0
spark.sql.hive.metastore.sharedPrefixescom.mysql.jdbc,org.postgresql,com.microsoft.sqlserver,oracle.jdbc1.4.0
spark.sql.hive.metastore.version2.3.101.4.0
spark.sql.hive.thriftServer.singleSessionfalse1.6.0
spark.sql.hive.version2.3.101.1.1
spark.sql.metadataCacheTTLSeconds-1000ms3.1.0
spark.sql.queryExecutionListeners(none)2.3.0
spark.sql.sources.disabledJdbcConnProviderList3.1.0
spark.sql.streaming.streamingQueryListeners(none)2.4.0
spark.sql.streaming.ui.enabledtrue3.0.0
spark.sql.streaming.ui.retainedProgressUpdates1003.0.0
spark.sql.streaming.ui.retainedQueries1003.0.0
spark.sql.ui.retainedExecutions10001.5.0
spark.sql.warehouse.dir(value of $PWD/spark-warehouse)2.0.0

Cluster Managers

These configurations are handled by Spark and do not affect Gluten’s behavior.

Push-based shuffle overview

External Shuffle service(server) side configuration options

Property NameDefaultSince VersionGluten Status
spark.shuffle.push.server.mergedShuffleFileManagerImplorg.apache.spark.network.shuffle. NoOpMergedShuffleFileManager3.2.0
spark.shuffle.push.server.minChunkSizeInMergedShuffleFile2m3.2.0
spark.shuffle.push.server.mergedIndexCacheSize100m3.2.0

Client side configuration options

Property NameDefaultSince VersionGluten Status
spark.shuffle.push.enabledfalse3.2.0
spark.shuffle.push.finalize.timeout10s3.2.0
spark.shuffle.push.maxRetainedMergerLocations5003.2.0
spark.shuffle.push.mergersMinThresholdRatio0.053.2.0
spark.shuffle.push.mergersMinStaticThreshold53.2.0
spark.shuffle.push.numPushThreads(none)3.2.0
spark.shuffle.push.maxBlockSizeToPush1m3.2.0
spark.shuffle.push.maxBlockBatchSize3m3.2.0
spark.shuffle.push.merge.finalizeThreads83.3.0
spark.shuffle.push.minShuffleSizeToWait500m3.3.0
spark.shuffle.push.minCompletedPushRatio1.03.3.0

评论

登录后参与评论

正在加载评论…