ConceptsqianmoQ· 更新于 2026-10-09· 阅读 35 分钟· 0 次阅读朗读举报登录后可跨设备保存划线和私人笔记登录ConceptsNow that you’re familiar with the basics and have run your own dataflow, let’s dive into the concepts that makes Apache Hamilton unique and powerful. Glossary Functions, nodes & dataflow Functions Specifying dependencies Helper function Function naming tips Nodes Anatomy of a node Dataflow How other frameworks build graphs Readability Maintainability Recap Next step Driver Define the Driver Visualize the dataflow Execute the dataflow Development tips With a Python module With a Jupyter notebook Recap Next step Visualization Available visualizations View full dataflow View executed dataflow View node dependencies Configure your visualization Custom node labels with display_name Apply custom style Materialization Different ways to write the same dataflow Without materialization Limitations With materialization Simple Materialization Static materializers Dynamic materializers Function modifiers DataLoader and DataSaver SQL metadata and lineage Fields Supported connections Schema precedence and unknown cases Call the helper yourself Function modifiers Decorators Reminder: Anatomy of a node Add metadata to a node @tag Query node by tag Customize visualization by tag @schema Validate node output @check_output* pandera support pydantic support Split node output into n nodes @unpack_fields @extract_fields @extract_columns Define one function, create n nodes @parameterize Select functions to include @config Load and save external data @load_from @save_to Builder with_modules() with_config() with_materializers() with_cache() with_adapters() enable_dynamic_execution() Caching How does it work? Cache key Observing the cache Logging Visualization Structured logs Cached result format Caching behavior Setting caching behavior via @cache via Builder().with_cache() Set a default behavior Code version Data version Hashing backend Recursion depth Support additional types Storage Setting the cache path By project Globally Separate locations Inspect storage In-memory Persist cache Load cache Roadmap Function modifiers (Advanced) Dynamic DAGs/Parallel Execution Using an Adapter Using the Parallelizable[] and Collect[] types Known Caveats Serialization Multiple Collects UI Overview Local Mode Docker/Deployed Mode Install Building the Docker Images locally Self-Hosting Running on Snowflake Get started Existing Apache Hamilton Code I need some Apache Hamilton code to run Features Dataflow versioning Assets/features catalog Browser Run tracking + telemetry SDK Configuration Changing where data is sent Changing behavior of what is captured Best Practices Function Naming It enables you to define your Apache Hamilton dataflow It drives collaboration and reuse It serves as documentation itself Migrating to Apache Hamilton Continuous Integration for Comparisons Integrate into your code base via a “custom wrapper object” Code Organization Team thinking Helps isolate what you’re working on Enables you to replace parts of your DAG easily for different contexts Common Indices Best practice: Output Immutability Best practice: Using within your ETL System Compatibility Matrix ETL Recipe Loading Data Plugging in new Data Sources Modules as Interfaces Using the Config to Decide Sources 评论登录后参与评论正在加载评论…