{
  "schemaVersion": 3,
  "dataset": {
    "version": 3,
    "date": "2026-08-13",
    "group": {
      "id": "data-infrastructure",
      "name": "Data / Messaging / Storage Infrastructure"
    },
    "repository": {
      "id": "clickhouse",
      "repo": "ClickHouse/ClickHouse",
      "name": "ClickHouse",
      "keywords": [
        "ClickHouse"
      ]
    },
    "context": {
      "repository": "ClickHouse/ClickHouse",
      "url": "https://github.com/ClickHouse/ClickHouse",
      "description": "ClickHouse® is a real-time analytics database management system",
      "homepage": "https://clickhouse.com",
      "language": "C++",
      "topics": [
        "ai",
        "analytics",
        "big-data",
        "clickhouse",
        "cloud-native",
        "cpp",
        "database",
        "dbms",
        "distributed",
        "embedded",
        "hacktoberfest",
        "lakehouse",
        "mpp",
        "olap",
        "rust",
        "self-hosted",
        "sql"
      ],
      "license": "Apache-2.0",
      "defaultBranch": "master",
      "stars": 49229,
      "forks": 8784,
      "openIssues": 6897,
      "archived": false,
      "collectedAt": "2026-08-13T18:02:06.669887+00:00"
    },
    "news": {
      "repository": "ClickHouse/ClickHouse",
      "collectedAt": "2026-08-13T18:02:06.669887+00:00",
      "latestRelease": {
        "repository": "ClickHouse/ClickHouse",
        "tag": "v26.7.3.19-stable",
        "title": "Release v26.7.3.19-stable",
        "url": "https://github.com/ClickHouse/ClickHouse/releases/tag/v26.7.3.19-stable",
        "publishedAt": "2026-08-06T07:18:47Z",
        "highlights": [],
        "prerelease": false
      },
      "upcoming": [],
      "communityDiscussions": []
    },
    "runs": [
      {
        "collectedAt": "2026-08-13T12:26:38.318Z",
        "since": "2026-08-12T12:26:38.318Z",
        "observedCount": 500,
        "changedCount": 500
      },
      {
        "collectedAt": "2026-08-13T13:48:00.446149Z",
        "since": "2026-08-12T13:48:00.446149Z",
        "observedCount": 500,
        "changedCount": 500
      },
      {
        "collectedAt": "2026-08-13T16:19:22.035158Z",
        "since": "2026-08-12T16:19:22.035158Z",
        "observedCount": 500,
        "changedCount": 210
      },
      {
        "collectedAt": "2026-08-13T17:43:20.785491Z",
        "since": "2026-08-12T17:43:20.785491Z",
        "observedCount": 500,
        "changedCount": 136
      },
      {
        "collectedAt": "2026-08-13T17:47:07.884300Z",
        "since": "2026-08-12T17:47:07.884300Z",
        "observedCount": 500,
        "changedCount": 13
      },
      {
        "collectedAt": "2026-08-13T18:01:55.420671Z",
        "since": "2026-08-12T18:01:55.420671Z",
        "observedCount": 500,
        "changedCount": 45
      }
    ],
    "signals": [
      {
        "id": "github:ClickHouse/ClickHouse:issue:101325",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Logical error: Shard number is greater than shard count: shard_num=A shard_count=B cluster=C (STID: 5066-564d)",
        "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Shard number is greater than shard count: shard_num=A shard_count=B cluster=C (STID: 5066-564d) CI report: [Stress test (arm_tsan)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=94690&sha=298cf42e6cd833562f76eef191bfd491b6937198&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBTaGFyZCBudW1iZXIgaXMgZ3JlYXRlciB0aGFuIHNoYXJkIGNvdW50OiBzaGFyZF9udW09QSBzaGFyZF9jb3VudD1CIGNsdXN0ZXI9QyAoU1RJRDogNTA2Ni01NjRkKScKICAgIC0tIEFORCBjaGVja19uYW1lID0gJ1N0cmVzcyB0ZXN0IChhcm1fdHNhbiknCiAgICBBTkQgdGVzdF9zdGF0dXMgSU4gKCdGQUlMJywgJ0VSUk9SJykKICAgIEFORCAocHVsbF9yZXF1ZXN0X251bWJlciA9IDAgT1IgYmFzZV9yZWYgSU4gKCdtYXN0ZXInKSkKICAgIEFORCB0ZXN0X2NvbnRleHRfcmF3IExJS0UgJyVFeGNlcHRpb246JScKR1JPVVAgQlkgZGF5Ck9SREVSIEJZIGRheSBERVNDCg==) Test output: ``` Log files: clickhouse-server.err.log, stderr.log Error: Logical error: 'Shard number is greater than shard count: shard_num=2 shard_count=1 cluster=parallel_replicas'. --- Stack trace: __pthread_kill_implementation @ 0x0000000000082009 raise @ 0x000000000003a83c __GI_abort @ 0x0000000000027134 contrib/llvm-project/compiler-rt/lib/tsan/rtl/tsan_interceptors_posix.cpp:0: __interceptor_abort @ 0x000000000bf6e8c4 src/Common/Exception.cpp:60:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x0000000014cf403c src/Common/Exception.cpp:93:13: DB::handle_error_code(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x0000000014cf5548 src/Common/Exception.cpp:146:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x0000000014cf5964 ./src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000c006c00 ./src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000c0063fc ./src/Common/Exception.h:189:77: DB::Exception::Exception<unsigned long&, unsigned long const&, String const&>(int, FormatStringHelperImpl<std::type_identity<unsigned long&>::type, std::type_identity<unsigned long const&>::type, std::type_identity<String const&>::type>, unsigned long&, unsigned long const&, String const&) @ 0x0000000020cd1ba8 src/Interpreters/ClusterProxy/executeQuery.cpp:565:19: DB::ClusterProxy::prepareClusterForParallelReplicas(std::shared_ptr<Poco::Logger> const&, std::shared_ptr<DB::Context const> const&) @ 0x0000000020cc7eb0 src/Interpreters/ClusterProxy/executeQuery.cpp:688:33: DB::ClusterProxy::executeQueryWithParallelReplicas(DB::QueryPlan&, DB::StorageID const&, std::shared_ptr<DB::Block const>, DB::QueryProcessingStage::Enum, boost::intrusive_ptr<DB::IAST> const&, std::shared_ptr<DB::IQueryTreeNode>, std::shared_ptr<DB::PlannerContext>, std::shared_ptr<DB::Context const>, std::shared_ptr<std::list<DB::StorageLimits, std::allocator<DB::StorageLimits>> const>, std::unique_ptr<DB::IQueryPlanStep, std::default_delete<DB::IQueryPlanStep>>) @ 0x0000000020cc5074 src/Interpreters/ClusterProxy/executeQuery.cpp:837:5: DB::ClusterProxy::executeQueryWithParallelReplicas(DB::QueryPlan&, DB::StorageID const&, DB::QueryProcessingStage::Enum, boost::intrusive_ptr<DB::IAST> const&, std::shared_ptr<DB::Context const>, std::shared_ptr<std::list<DB::StorageLimits, std::allocator<DB::StorageLimits>> const>) @ 0x0000000020ccaeb0 src/Storages/StorageMergeTree.cpp:322:9: DB::StorageMergeTree::read(DB::QueryPlan&, std::vector<String, std::allocator<String>> const&, std::shared_ptr<DB::StorageSnapshot> const&, DB::SelectQueryInfo&, std::shared_ptr<DB::Context const>, DB::QueryProcessingStage::Enum, unsigned long, unsigned long) @ 0x00000000216ced50 src/Interpreters/InterpreterSelectQuery.cpp:2840:18: DB::InterpreterSelectQuery::executeFetchColumns(DB::QueryProcessingStage::Enum, DB::QueryPlan&) @ 0x000000001d6c91d4 src/Interpreters/InterpreterSelectQuery.cpp:1797:9: DB::InterpreterSelectQuery::executeImpl(DB::QueryPlan&, std::optional<DB::Pipe>) @ 0x000000001d6bdbd4 src/Interpreters/InterpreterSelectQuery.cpp:1156:5: DB::InterpreterSelectQuery::buildQueryPlan(DB::QueryPlan&) @ 0x000000001d6bd23c src/Interpreters/InterpreterSelectWithUnionQuery.cpp:319:38: DB::InterpreterSelectWithUnionQuery::buildQueryPlan(DB::QueryPlan&) @ 0x000000001d700de4 src/Interpreters/InterpreterSelectWithUnionQuery.cpp:394:5: DB::InterpreterSelectWithUnionQuery::execute() @ 0x000000001d701f74 src/Interpreters/executeQuery.cpp:1811:40: DB::executeQueryImpl(char const*, char const*, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum, std::unique_ptr<DB::ReadBuffer, std::default_delete<DB::ReadBuffer>>&, boost::intrusive_ptr<DB::IAST>&, std::shared_ptr<DB::ImplicitTransactionControlExecutor>, std::function<void ()>, DB::QueryResultDetails&) @ 0x000000001db90dd0 src/Interpreters/executeQuery.cpp:2199:11: DB::executeQuery(String const&, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum) @ 0x000000001db8a30c src/Server/TCPHandler.cpp:823:68: DB::TCPHandler::runImpl() @ 0x00000000228469a8 src/Server/TCPHandler.cpp:2963:9: DB::TCPHandler::run() @ 0x000000002286fa38 base/poco/Net/src/TCPServerConnection.cpp:40:3: Poco::Net::TCPServerConnection::start() @ 0x000000002be0cb34 base/poco/Net/src/TCPServerDispatcher.cpp:115:42: Poco::Net::TCPServerDispatcher::run() @ 0x000000002be0d34c base/poco/Foundation/src/ThreadPool.cpp:205:14: Poco::PooledThread::run() @ 0x000000002bd70be4 base/poco/Foundation/src/Thread.cpp:45:11: Poco::(anonymous namespace)::RunnableHolder::run() @ 0x000000002bd6edc0 ./base/poco/Foundation/src/Thread_POSIX.cpp:341:27: Poco::ThreadImpl::runnableEntry(void*) @ 0x000000002bd6d344 contrib/llvm-project/compiler-rt/lib/tsan/rtl/tsan_interceptors_posix.cpp:1080:15: __tsan_thread_start_func @ 0x000000000bf663d0 start_thread @ 0x0000000000080398 thread_start @ 0x00000000000e9e9c ```",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/101325",
        "timestamp": "2026-08-12T22:50:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "testing",
          "fuzz"
        ],
        "author": "SmitaRKulkarni",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:101474",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Vector similarity: support TurboQuant quantization",
        "text": "### Company or project name _No response_ ### Use case [TurboQuant, a quantization method](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/) that perhaps could be applied to ClickHouse to improve Approximate Nearest Neighbor (ANN) searches. technical details are somewhat beyond my expertise ### Describe the solution you'd like New TurboQuant quantization option for ANN indexes, something similar to the existing quantization QBit ### Describe alternatives you've considered _No response_ ### Additional context - [Blog post](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/) - [Paper](https://arxiv.org/pdf/2504.19874)",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/101474",
        "createdAt": "2026-04-01T08:55:13Z",
        "updatedAt": "2026-08-13T04:41:17Z",
        "timestamp": "2026-08-13T04:41:17Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "feature"
        ],
        "author": "Wachynaky",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:101747",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "checkTableNameLengthUnlocked in renameDatabase gated behind dependency-checking condition, bypassed with check_table_dependencies=0",
        "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _Retrospective finding from a historical scan of [PR #79488](https://github.com/ClickHouse/ClickHouse/pull/79488) (merged 2025-04-25). Confirmed on current codebase — close with a note if already fixed._ ### Describe what's wrong RENAME DATABASE to a longer name bypasses table name length validation when check_table_dependencies=0, creating tables that cannot be dropped (filesystem error: File name too long) **Root cause:** DatabaseAtomic.cpp:725-729: checkTableNameLengthUnlocked is placed inside the `if (check_ref_deps || check_loading_deps)` block, but the name length check is independent of dependency checking and should be unconditional **Why we believe this is a bug:** DatabaseAtomic::renameDatabase (DatabaseAtomic.cpp:715) → check_ref_deps/check_loading_deps computed from settings (line 723-724) → if-block at line 725 gates BOTH dependency check AND name length check → when check_table_dependencies=0, the entire block is skipped including checkTableNameLengthUnlocked at line 729 **Affected locations:** - `src/Databases/DatabaseAtomic.cpp:729` — checkTableNameLengthUnlocked inside dependency-checking if-block **Impact:** When a user renames a database to a longer name with check_table_dependencies=0, existing tables may end up with names exceeding filesystem limits. These tables become undropable (DROP TABLE and DROP DATABASE both fail with 'File name too long' error), requiring manual intervention to recover. ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `18dfe15a9516`). ### How to reproduce ```sql DROP DATABASE IF EXISTS db_short_79488; CREATE DATABASE db_short_79488 ENGINE=Atomic; CREATE TABLE db_short_79488.aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa (x UInt8) ENGINE=MergeTree ORDER BY x; -- With check_table_dependencies=0, RENAME should STILL check table name length -- but currently bypasses it (the check is inside the dependency-checking if-block) SET check_table_dependencies=0; RENAME DATABASE db_short_79488 TO db_short_79488_with_a_much_longer_name_reduces_table_len; -- { serverError ARGUMENT_OUT_OF_BOUND } -- Cleanup: rename back if the above succeeded (bug), then drop SET check_table_dependencies=0; RENAME DATABASE db_short_79488_with_a_much_longer_name_reduces_table_len TO db_short_79488; -- { serverError UNKNOWN_DATABASE } DROP DATABASE IF EXISTS db_short_79488; ``` [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/9a5afb99-69f7-4a1c-b97e-d21efa069052) ### Expected behavior ``` No output (RENAME should fail with ARGUMENT_OUT_OF_BOUND error code 69) ``` ### Error message and/or stacktrace ``` The query succeeded but the server error '69' was expected (query: RENAME DATABASE db_short_79488 TO db_short_79488_with_a_much_longer_name_reduces_table_len; -- { serverError ARGUMENT_OUT_OF_BOUND }). ``` ### Additional context **Open risks:** - Other database engines inheriting from DatabaseOnDisk may have similar issues if they override renameDatabase **Suggested fix:** Move the checkTableNameLengthUnlocked loop outside and before the `if (check_ref_deps || check_loading_deps)` block, so the name length check runs unconditionally regardless of dependency-checking settings. **Analysis details:** Confidence HIGH | Severity P1 | Testability: `STATELESS_SQL` Found during automated review of [PR #79488](https://github.com/ClickHouse/ClickHouse/pull/79488). --- _ClickGapAI · Confidence: HIGH · Severity: P1 · Finding: `h_pr79488_001`_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/101747",
        "createdAt": "2026-04-04T04:52:38Z",
        "updatedAt": "2026-08-13T01:16:05Z",
        "timestamp": "2026-08-13T01:16:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "comp-database-engines"
        ],
        "author": "clickgapai",
        "state": "closed",
        "assignees": [
          "mstetsyuk"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:102037",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Logical error: Can't extract iceberg table state from storage snapshot for table location A (STID: 2606-4a47)",
        "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Can't extract iceberg table state from storage snapshot for table location A (STID: 2606-4a47) CI report: [Stress test (amd_msan)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=96682&sha=bc37844d817ee88bc66bf40870fb4d9c5373d488&name_0=PR&name_1=Stress%20test%20%28amd_msan%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBDYW4nJ3QgZXh0cmFjdCBpY2ViZXJnIHRhYmxlIHN0YXRlIGZyb20gc3RvcmFnZSBzbmFwc2hvdCBmb3IgdGFibGUgbG9jYXRpb24gQSAoU1RJRDogMjYwNi00YTQ3KScKICAgIC0tIEFORCBjaGVja19uYW1lID0gJ1N0cmVzcyB0ZXN0IChhbWRfbXNhbiknCiAgICBBTkQgdGVzdF9zdGF0dXMgSU4gKCdGQUlMJywgJ0VSUk9SJykKICAgIEFORCAocHVsbF9yZXF1ZXN0X251bWJlciA9IDAgT1IgYmFzZV9yZWYgSU4gKCdtYXN0ZXInKSkKICAgIEFORCB0ZXN0X2NvbnRleHRfcmF3IExJS0UgJyVFeGNlcHRpb246JScKR1JPVVAgQlkgZGF5Ck9SREVSIEJZIGRheSBERVNDCg==) Test output: ``` Log files: clickhouse-server.err.log, stderr.log Error: Logical error: 'Can't extract iceberg table state from storage snapshot for table location /var/lib/clickhouse/user_files/t_test_pxpyjno7_28122/'. --- Stack trace: __pthread_kill @ 0x00000000000969fd gsignal @ 0x0000000000042476 __lgamma_r_finite@GLIBC_2.15 @ 0x00000000000287f3 src/Common/Exception.cpp:60:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x000000002bc3f621 src/Common/Exception.cpp:93:13: DB::handle_error_code(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x000000002bc4178a src/Common/Exception.cpp:146:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x000000002bc4211e ./src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000b75d1ae ./src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000b75bb0d ./src/Common/Exception.h:189:77: DB::Exception::Exception<String const&>(int, FormatStringHelperImpl<std::type_identity<String const&>::type>, String const&) @ 0x000000000b798f4a src/Storages/ObjectStorage/DataLakes/Iceberg/IcebergMetadata.cpp:1037:15: DB::IcebergMetadata::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) const @ 0x000000003d64f43e ./src/Storages/ObjectStorage/DataLakes/DataLakeConfiguration.h:276:34: DB::DataLakeConfiguration<DB::StorageLocalConfiguration, DB::IcebergMetadata>::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) @ 0x0000000034ed9d1c src/Storages/ObjectStorage/StorageObjectStorageSource.cpp:259:36: DB::StorageObjectStorageSource::createFileIterator(std::shared_ptr<DB::StorageObjectStorageConfiguration>, DB::StorageObjectStorageQuerySettings const&, std::shared_ptr<DB::IObjectStorage>, std::shared_ptr<DB::StorageInMemoryMetadata const>, bool, std::shared_ptr<DB::Context const> const&, DB::ActionsDAG::Node const*, DB::ActionsDAG const*, DB::NamesAndTypesList const&, DB::NamesAndTypesList const&, std::vector<std::shared_ptr<DB::ObjectInfo>, std::allocator<std::shared_ptr<DB::ObjectInfo>>>*, std::function<void (DB::FileProgress)>, bool, bool, bool) @ 0x000000003d25bab6 src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:156:24: DB::ReadFromObjectStorageStep::createIterator() @ 0x0000000056ed3981 src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:85:5: DB::ReadFromObjectStorageStep::initializePipeline(DB::QueryPipelineBuilder&, DB::BuildQueryPipelineSettings const&) @ 0x0000000056ed0902 src/Processors/QueryPlan/ISourceStep.cpp:20:5: DB::ISourceStep::updatePipeline(std::vector<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>, std::allocator<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>>>, DB::BuildQueryPipelineSettings const&) @ 0x0000000056ac3ac0 src/Processors/QueryPlan/QueryPlan.cpp:211:47: DB::QueryPlan::buildQueryPipeline(DB::QueryPlanOptimizationSettings const&, DB::BuildQueryPipelineSettings const&, bool) @ 0x0000000056bf17ae src/Interpreters/MutationsInterpreter.cpp:1776:37: DB::MutationsInterpreter::addStreamsForLaterStages(std::vector<DB::MutationsInterpreter::Stage, std::allocator<DB::MutationsInterpreter::Stage>> const&, DB::QueryPlan&) const @ 0x0000000044b55e3a src/Interpreters/MutationsInterpreter.cpp:1831:20: DB::MutationsInterpreter::execute() @ 0x0000000044b62cde src/Storages/ObjectStorage/DataLakes/Iceberg/Mutations.cpp:166:72: DB::Iceberg::writeDataFiles(DB::MutationCommands const&, std::shared_ptr<DB::Context const>, std::shared_ptr<DB::StorageInMemoryMetadata const>, DB::StorageID, std::shared_ptr<DB::IObjectStorage>, String, DB::FileNamesGenerator&, DB::Iceberg::IcebergPathResolver const&, std::optional<DB::FormatSettings> const&, std::optional<DB::ChunkPartitioner>&, Poco::SharedPtr<Poco::JSON::Object, Poco::ReferenceCounter, Poco::ReleasePolicy<Poco::JSON::Object>>) @ 0x000000003d83743d src/Storages/ObjectStorage/DataLakes/Iceberg/Mutations.cpp:611:31: DB::Iceberg::mutate(DB::MutationCommands const&, std::shared_ptr<DB::Context const>, std::shared_ptr<DB::StorageInMemoryMetadata const>, DB::StorageID, std::shared_ptr<DB::IObjectStorage>, DB::DataLakeStorageSettings const&, DB::Iceberg::PersistentTableComponents const&, String const&, std::optional<DB::FormatSettings> const&, std::shared_ptr<DataLake::ICatalog>) @ 0x000000003d8309e6 src/Storages/ObjectStorage/DataLakes/Iceberg/IcebergMetadata.cpp:586:5: DB::IcebergMetadata::mutate(DB::MutationCommands const&, std::shared_ptr<DB::StorageObjectStorageConfiguration>, std::shared_ptr<DB::Context const>, DB::StorageID const&, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DataLake::ICatalog>, std::optional<DB::FormatSettings> const&) @ 0x000000003d63b34e ./src/Storages/ObjectStorage/DataLakes/DataLakeConfiguration.h:154:27: DB::DataLakeConfiguration<DB::StorageLocalConfiguration, DB::IcebergMetadata>::mutate(DB::MutationCommands const&, std::shared_ptr<DB::Context const>, DB::StorageID const&, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DataLake::ICatalog>, std::optional<DB::FormatSettings> const&) @ 0x0000000034ede9b6 src/Storages/ObjectStorage/StorageObjectStorage.cpp:781:20: DB::StorageObjectStorage::mutate(DB::MutationCommands const&, std::shared_ptr<DB::Context const>) @ 0x000000003d15f44a src/Interpreters/InterpreterAlterQuery.cpp:303: DB::(anonymous namespace)::runCommandSegments(std::vector<std::variant<DB::AlterCommands, DB::MutationCommands, std::vector<DB::PartitionCommand, std::allocator<DB::PartitionCommand>>, std::vector<DB::ASTAlterCommand const*, std::allocator<DB::ASTAlterCommand const*>>>, std::allocator<std::variant<DB::AlterCommands, DB::MutationCommands, std::vector<DB::PartitionCommand, std::allocator<DB::PartitionCommand>>, std::vector<DB::ASTAlterCommand const*, std::allocator<DB::ASTAlterCommand const*>>>>>&, std::shared_ptr<DB::IStorage> const&, std::shared_ptr<DB::Context const> const&) src/Interpreters/InterpreterAlterQuery.cpp:447:24: DB::InterpreterAlterQuery::executeToTable(DB::ASTAlterQuery const&) @ 0x000000004471b19c src/Interpreters/InterpreterAlterQuery.cpp:345:16: DB::InterpreterAlterQuery::execute() @ 0x0000000044707b25 src/Interpreters/executeQuery.cpp:1815:40: DB::executeQueryImpl(char const*, char const*, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum, std::unique_ptr<DB::ReadBuffer, std::default_delete<DB::ReadBuffer>>&, boost::intrusive_ptr<DB::IAST>&, std::shared_ptr<DB::ImplicitTransactionControlExecutor>, std::function<void ()>, DB::QueryResultDetails&) @ 0x00000000456817e5 src/Interpreters/executeQuery.cpp:2203:11: DB::executeQuery(String const&, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum) @ 0x000000004566eb09 src/Server/TCPHandler.cpp:823:68: DB::TCPHandler::runImpl() @ 0x00000000551b54b8 src/Server/TCPHandler.cpp:2974:9: DB::TCPHandler::run() @ 0x000000005522e104 base/poco/Net/src/TCPServerConnection.cpp:40:3: Poco::Net::TCPServerConnection::start() @ 0x0000000069704c60 base/poco/Net/src/TCPServerDispatcher.cpp:115:42: Poco::Net::TCPServerDispatcher::run() @ 0x0000000069705c9a base/poco/Foundation/src/ThreadPool.cpp:205:14: Poco::PooledThread::run() @ 0x00000000695a459a ./base/poco/Foundation/src/Thread_POSIX.cpp:341:27: Poco::ThreadImpl::runnableEntry(void*) @ 0x000000006959d811 start_thread @ 0x0000000000094ac3 __clone3 @ 0x00000000001268d0 ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #102033 - Merged into: `26.7.1.448` (included in `26.7` and later) - Backported to: `26.5.7.46` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/102037",
        "createdAt": "2026-04-08T10:00:31Z",
        "updatedAt": "2026-08-13T01:36:04Z",
        "timestamp": "2026-08-13T01:36:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "testing",
          "fuzz"
        ],
        "author": "KochetovNicolai",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:102845",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Documentation examples for `h3GetDestinationIndexFromUnidirectionalEdge` and `h3GetOriginIndexFromUnidirectionalEdge` use invalid edge index that throws `INCORRECT_DATA`",
        "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _Retrospective finding from a historical scan of [PR #82286](https://github.com/ClickHouse/ClickHouse/pull/82286) (merged 2025-10-17). Confirmed on current codebase — close with a note if already fixed._ ### Describe what's wrong Both `h3GetDestinationIndexFromUnidirectionalEdge` and `h3GetOriginIndexFromUnidirectionalEdge` documentation examples use edge index 1248204388774707197, which is not a valid H3 directed edge. Running the documented query throws `INCORRECT_DATA` exception instead of returning the documented values. **Root cause:** h3GetDestinationIndexFromUnidirectionalEdge.cpp:117 and h3GetOriginIndexFromUnidirectionalEdge.cpp:117: the example input 1248204388774707197 is not a valid H3 directed edge (h3UnidirectionalEdgeIsValid returns 0), and the documented output values are also wrong even if the correct edge (1248204388774707199) is used **Why we believe this is a bug:** h3GetDestinationIndexFromUnidirectionalEdge.cpp:117 and h3GetOriginIndexFromUnidirectionalEdge.cpp:117 — both examples use index 1248204388774707197 which fails `isValidDirectedEdge` check in h3Common.cpp:38. The valid edge used elsewhere in the PR is 1248204388774707199 (differs by 2). The documented output values (599686043507097597 and 599686042433355773) also differ by 2 from the actual values for the valid edge (599686043507097599 and 599686042433355775). **Affected locations:** - `src/Functions/h3GetDestinationIndexFromUnidirectionalEdge.cpp:117` — example uses invalid edge 1248204388774707197, documented output 599686043507097597 - `src/Functions/h3GetOriginIndexFromUnidirectionalEdge.cpp:117` — example uses invalid edge 1248204388774707197, documented output 599686042433355773 **Impact:** Users running the documented example query get an unexpected `INCORRECT_DATA` exception instead of the promised result. Even with `functions_h3_default_if_invalid=1`, the function returns 0 (not the documented value). ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `19cbcd782ed8`). ### How to reproduce ```sql -- Verify documented edge index is invalid SELECT h3UnidirectionalEdgeIsValid(1248204388774707197) AS invalid_edge; SELECT h3UnidirectionalEdgeIsValid(1248204388774707199) AS valid_edge; -- Actual correct results with valid edge SELECT h3GetDestinationIndexFromUnidirectionalEdge(1248204388774707199) AS destination; SELECT h3GetOriginIndexFromUnidirectionalEdge(1248204388774707199) AS origin; ``` [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/8eabd45c-b67c-477b-939f-3946f68a10a4) ### Expected behavior ``` Documentation claims h3GetDestinationIndexFromUnidirectionalEdge(1248204388774707197) returns 599686043507097597 and h3GetOriginIndexFromUnidirectionalEdge(1248204388774707197) returns 599686042433355773, but index 1248204388774707197 is invalid (throws INCORRECT_DATA) and even the correct index produces different values ``` ### Error message and/or stacktrace ``` 0 1 599686043507097599 599686042433355775 ``` ### Additional context **Open risks:** - The existing test 02292_h3_unidirectional_funcs.sql already expects INCORRECT_DATA for index 1248204388774707197 (line 3), confirming the documentation is wrong **Suggested fix:** Replace edge index 1248204388774707197 with the valid edge 1248204388774707199 in both files, and update output values: destination from 599686043507097597 to 599686043507097599, origin from 599686042433355773 to 599686042433355775 **Analysis details:** Confidence HIGH | Severity P3 | Testability: `STATELESS_SQL` Found during automated review of [PR #82286](https://github.com/ClickHouse/ClickHouse/pull/82286). --- _ClickGapAI · Confidence: HIGH · Severity: P3 · Finding: `h_pr82286_002`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #111152 - Merged into: `26.8.1.1324` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/102845",
        "createdAt": "2026-04-15T18:38:21Z",
        "updatedAt": "2026-08-13T12:50:47Z",
        "timestamp": "2026-08-13T12:50:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "comp-documentation",
          "comp-geo"
        ],
        "author": "clickgapai",
        "state": "closed",
        "assignees": [
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:106460",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "FTS index after Light-Weight Update is not working for MATERIALIZED column",
        "text": "### Company or project name ClickHouse ### Describe what's wrong During implementation of Full-Text Search framework I found the issue when using Light-Weight Updates. Particularly, when patch parts have to be applied, `MATERIALIZED` column can't be utilized in FTS index - `serverError UNKNOWN_IDENTIFIER` is thrown. However, other non-`MATERIALIZED` columns can be searched using FTS index. ### Does it reproduce on the most recent release? Yes ### How to reproduce ```sql SET allow_experimental_full_text_index = 1; SET allow_experimental_json_type = 1; SET enable_lightweight_update = 1; DROP TABLE IF EXISTS tab; CREATE TABLE tab ( id Int64, text String, data_as_json JSON, json_title String MATERIALIZED data_as_json.title::String, INDEX idx_text text TYPE text(tokenizer = 'splitByNonAlpha'), INDEX idx_json_title json_title TYPE text(tokenizer = 'splitByNonAlpha') ) ENGINE = MergeTree ORDER BY id SETTINGS enable_block_number_column = 1, enable_block_offset_column = 1; INSERT INTO tab (id, text, data_as_json) VALUES (1, 'database is fast', '{\"title\": \"performance tuning guide\"}'), (2, 'no match here', '{\"title\": \"unrelated\"}'), (3, 'database stuff', '{\"title\": \"performance test results\"}'); -- Sanity check before the lightweight UPDATE (no patch parts yet) — -- compound WHERE works. SELECT 'compound WHERE, no patch parts yet'; SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance'); -- Lightweight UPDATE on a non-materialised column → creates a patch -- part. The materialised column's source (data_as_json) is NOT -- touched. UPDATE tab SET text = 'database updated' WHERE id = 1; -- Single-column WHEREs still work after the patch part is created. SELECT 'WHERE on text only, post-UPDATE'; SELECT count() FROM tab WHERE hasToken(text, 'database'); SELECT 'WHERE on json_title only, post-UPDATE'; SELECT count() FROM tab WHERE hasToken(json_title, 'performance'); -- The compound WHERE now fails with UNKNOWN_IDENTIFIER. SELECT 'compound WHERE, post-UPDATE (BUG)'; SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance'); -- { serverError UNKNOWN_IDENTIFIER } -- The same compound WHERE with apply_patch_parts = 0 succeeds -- (workaround). SELECT 'compound WHERE, apply_patch_parts = 0 (workaround)'; SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance') SETTINGS apply_patch_parts = 0; DROP TABLE tab; ``` [04305_text_index_bug_json_materialized_patch_part.sql](https://github.com/user-attachments/files/28594469/04305_text_index_bug_json_materialized_patch_part.sql) ### Expected behavior Patch parts should be applied and `MATERIALIZED` column should work with FTS index ### Error message and/or stacktrace ``` {acd9a2a5-2c75-4ce3-a848-babcfcaad691} <Error> executeQuery: Code: 47. DB::Exception: Unknown expression or function identifier `data_as_json.title` in scope _CAST(hasToken(json_title, 'performance'), 'UInt8') AS __text_index_idx_json_title_hasToken_b115b01d1cac96a451f08eb68b7633a0, _CAST(CAST(data_as_json.title, 'String'), 'String') AS json_title: (while reading from part /home/alexey/Documents/ClickHouse/store/f2a/f2a80d0a-0b1f-44ee-ba88-640b07bc0411/all_1_1_0/ located on disk default of type local): While executing MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder). (UNKNOWN_IDENTIFIER) (version 26.6.1.374 (official build)) (from 127.0.0.1:55680) (comment: 04305_text_index_bug_json_materialized_patch_part.sql-test_u6trcoga) (query 15, line 79) (in query: SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance');), Stack trace (when copying this message, always include the lines below): 0. /ClickHouse/contrib/llvm-project/libcxx/include/__exception/exception.h:115:14: Poco::Exception::Exception(String&&, int) @ 0x0000000026eb20b3 1. /ClickHouse/src/Common/Exception.cpp:139:7: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x0000000014b73ca9 2. /ClickHouse/src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d08ac96 3. /ClickHouse/src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d08a758 4. /ClickHouse/src/Common/Exception.h:189:77: DB::Exception::Exception<char const*, String&, String, String, String>(int, FormatStringHelperImpl<std::type_identity<char const*>::type, std::type_identity<String&>::type, std::type_identity<String>::type, std::type_identity<String>::type, std::type_identity<String>::type>, char const*&&, String&, String&&, String&&, String&&) @ 0x00000000198e909a 5. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3226:27: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198cb0ce 6. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 7. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 8. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 9. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 10. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 11. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 12. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:1243:9: DB::QueryAnalyzer::tryResolveIdentifierFromAliases(DB::IdentifierLookup const&, DB::IdentifierResolveScope&, DB::IdentifierResolveContext) @ 0x00000000198d6e15 13. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:1548:34: DB::QueryAnalyzer::tryResolveIdentifier(DB::IdentifierLookup const&, DB::IdentifierResolveScope&, DB::IdentifierResolveContext) @ 0x00000000198d7f75 14. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3029:57: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c947b 15. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 16. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 17. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 18. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 19. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 20. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 21. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 22. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:285:17: DB::QueryAnalyzer::resolve(std::shared_ptr<DB::IQueryTreeNode>&, std::shared_ptr<DB::IQueryTreeNode> const&, std::shared_ptr<DB::Context const>) @ 0x00000000198c18e5 23. /ClickHouse/src/Interpreters/inplaceBlockConversions.cpp:238:14: DB::(anonymous namespace)::createExpressionsAnalyzer(DB::Block const&, boost::intrusive_ptr<DB::IAST>, bool, std::shared_ptr<DB::Context const>) @ 0x000000001a941cbe 24. /ClickHouse/src/Interpreters/inplaceBlockConversions.cpp:312:16: DB::evaluateMissingDefaults(DB::Block const&, DB::NamesAndTypesList const&, DB::ColumnsDescription const&, std::shared_ptr<DB::Context const>, bool, bool) @ 0x000000001a943559 25. /ClickHouse/src/Storages/MergeTree/IMergeTreeReader.cpp:269:20: DB::IMergeTreeReader::evaluateMissingDefaults(DB::Block, std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) const @ 0x000000001f25b0ea 26. /ClickHouse/src/Storages/MergeTree/MergeTreeReadersChain.cpp:250:28: DB::MergeTreeReadersChain::executeActionsBeforePrewhere(DB::MergeTreeRangeReader::ReadResult&, std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&, DB::MergeTreeRangeReader&, DB::Block const&, unsigned long) const @ 0x000000001f5589ea 27. /ClickHouse/src/Storages/MergeTree/MergeTreeReadersChain.cpp:120:9: DB::MergeTreeReadersChain::read(unsigned long, DB::MarkRanges&, std::vector<DB::MarkRanges, std::allocator<DB::MarkRanges>>&, std::function<void (std::vector<DB::ColumnWithTypeAndName, AllocatorWithMemoryTracking<DB::ColumnWithTypeAndName>> const&, unsigned long, std::optional<bool>&)> const&) @ 0x000000001f556d4d 28. /ClickHouse/src/Storages/MergeTree/MergeTreeReadTask.cpp:396:38: DB::MergeTreeReadTask::read() @ 0x000000001f553de3 29. /ClickHouse/src/Storages/MergeTree/MergeTreeSelectAlgorithms.h:53:103: DB::MergeTreeThreadSelectAlgorithm::readFromTask(DB::MergeTreeReadTask&) @ 0x000000001f56ce0c 30. /ClickHouse/src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:251:31: DB::MergeTreeSelectProcessor::readCurrentTask(DB::MergeTreeReadTask&, DB::IMergeTreeSelectAlgorithm&) const @ 0x000000001f561f91 31. /ClickHouse/src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:430:23: DB::MergeTreeSelectProcessor::read() @ 0x000000001f5642a5 Received exception from server (version 26.6.1): Code: 47. DB::Exception: Received from localhost:9000. DB::Exception: Unknown expression or function identifier `data_as_json.title` in scope _CAST(hasToken(json_title, 'performance'), 'UInt8') AS __text_index_idx_json_title_hasToken_b115b01d1cac96a451f08eb68b7633a0, _CAST(CAST(data_as_json.title, 'String'), 'String') AS json_title: (while reading from part /home/alexey/Documents/ClickHouse/store/f2a/f2a80d0a-0b1f-44ee-ba88-640b07bc0411/all_1_1_0/ located on disk default of type local): While executing MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder). (UNKNOWN_IDENTIFIER) ``` ### Related issues and pull requests _No response_ ### Additional context Tested on `ClickHouse local version 26.6.1.374 (official build)`",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/106460",
        "createdAt": "2026-06-04T12:05:07Z",
        "updatedAt": "2026-08-13T13:00:32Z",
        "timestamp": "2026-08-13T13:00:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "bug",
          "comp-text-index",
          "clickgap-analyzed",
          "culprit-pr-pinned"
        ],
        "author": "alexbakharew",
        "state": "open",
        "assignees": [
          "CurtizJ"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:107334",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "DeltaLake Logical error: 'No version found in table state snapshot'",
        "text": "### Describe the bug Seems not related to https://github.com/ClickHouse/ClickHouse/issues/102037 Reading a persistent DeltaLake (or Iceberg) table through a code path that does **not** call `updateExternalDynamicMetadataIfExists()` reaches `IDataLakeMetadata::iterate()` with a storage snapshot whose `datalake_table_state` is empty. `iterate()` treats this as impossible and throws `LOGICAL_ERROR`, which aborts the server: ``` <Fatal> : Logical error: 'No version found in table state snapshot'. DeltaLakeMetadataDeltaKernel::iterate (DeltaLakeMetadataDeltaKernel.cpp:331) -> DataLakeConfiguration<…, DeltaLakeMetadata>::iterate (DataLakeConfiguration.h:294) -> StorageObjectStorageSource::createFileIterator (StorageObjectStorageSource.cpp:287) -> ReadFromObjectStorageStep::createIterator / initializePipeline ``` Found by BuzzHouse on master `26.6.1.694` via `SELECT * FROM d2.t284` on a `DeltaLakeS3` table that had been through many ALTER/mutation/cluster-function operations. ## Root cause A persistent data-lake table's in-memory metadata gets its snapshot version pinned (`setDataLakeTableState`) only inside `updateExternalDynamicMetadataIfExists()` (`StorageObjectStorage.cpp:397`). The analyzer calls it on the table it resolves (`IdentifierResolver.cpp:335`), so a *direct* SELECT is fine. But several read paths build the child storage snapshot straight from `getInMemoryMetadataPtr()` **without** that call: - `StorageMerge` — `StorageMerge.cpp:730` (and `:414`) - cluster/distributed table-function reads on the executing node - (the materialized-view target path used to have this bug; it was fixed by adding the update call — `StorageMaterializedView.cpp:400`) When the table is reached first through one of these, the snapshot carries no `datalake_table_state`, and `iterate()` throws instead of falling back. The throw is also asymmetric with its sibling in the *same file*: `DeltaLakeMetadataDeltaKernel::prepareReadingFromFormat()` (`DeltaLakeMetadataDeltaKernel.cpp:505-511`) handles the identical missing-version case **gracefully** by falling back to the version from settings. `iterate()` should do the same. Iceberg is identical: `IcebergMetadata::iterate()` (`IcebergMetadata.cpp:1023`) throws `\"Can't extract iceberg table state from storage snapshot\"` under the same condition. ## Relationship to prior findings Same bug *class* as the Iceberg over-assertions filed in #107316 — a `LOGICAL_ERROR` guarding an invariant that is actually reachable from normal SQL. Here the invariant is \"the storage snapshot always carries a pinned table-state version\", which is false for read paths that skip `updateExternalDynamicMetadataIfExists`. ### How to reproduce With clickhouse binary in the PATH and the server running, run this script: ```bash set -u BIN=clickhouse # 26.6.1.694, has delta-kernel TABLE=contrib/delta-kernel-rs/kernel/tests/data/table-without-dv-small P=\"/var/lib/clickhouse/user_files\" rm -rf \"$P\" && mkdir -p \"$P\" cp -r \"$TABLE\" \"$P/delta1\" ABS=\"$P/delta1/\" echo \"=== control: direct SELECT pins state, works ===\" \"$BIN\" --client --multiquery \" SET allow_experimental_delta_kernel_rs=1; CREATE TABLE ok ENGINE = DeltaLakeLocal('$ABS'); SELECT count() FROM ok; \" echo \"=== repro: merge() is the SOLE first access -> aborts ===\" \"$BIN\" --client --multiquery \" SET allow_experimental_delta_kernel_rs=1; CREATE TABLE t ENGINE = DeltaLakeLocal('$ABS'); SELECT * FROM merge(currentDatabase(), '^t\\$') ORDER BY 1 LIMIT 5; \" echo \"exit=$? (134 = SIGABRT = reproduced)\" ``` ### Error message and/or stacktrace Stack trace: ``` <Fatal> : Logical error: 'No version found in table state snapshot'. <Fatal> : Format string: 'No version found in table state snapshot'. <Fatal> : Stack trace (when copying this message, always include the lines below): 0. /ClickHouse/contrib/llvm-project/libcxx/include/__exception/exception.h:115:14: Poco::Exception::Exception(String&&, int) @ 0x0000000027021373 1. /ClickHouse/src/Common/Exception.cpp:139:7: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x0000000014bdab29 2. /ClickHouse/src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d189cd6 3. /ClickHouse/src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d189798 4. /ClickHouse/src/Common/Exception.h:189:77: DB::Exception::Exception<>(int, FormatStringHelperImpl<>) @ 0x0000000014bd9279 5. /ClickHouse/src/Storages/ObjectStorage/DataLakes/DeltaLakeMetadataDeltaKernel.cpp:331:15: DB::DeltaLakeMetadataDeltaKernel::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) const @ 0x0000000018ab9f3d 6. /ClickHouse/src/Storages/ObjectStorage/DataLakes/DataLakeConfiguration.h:294:34: DB::DataLakeConfiguration<DB::StorageLocalConfiguration, DB::DeltaLakeMetadata>::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) @ 0x0000000016eaa901 7. /ClickHouse/src/Storages/ObjectStorage/StorageObjectStorageSource.cpp:287:36: DB::StorageObjectStorageSource::createFileIterator(std::shared_ptr<DB::StorageObjectStorageConfiguration>, DB::StorageObjectStorageQuerySettings const&, std::shared_ptr<DB::IObjectStorage>, std::shared_ptr<DB::StorageInMemoryMetadata const>, bool, std::shared_ptr<DB::Context const> const&, DB::ActionsDAG::Node const*, DB::ActionsDAG const*, DB::NamesAndTypesList const&, DB::NamesAndTypesList const&, std::vector<std::shared_ptr<DB::ObjectInfo>, std::allocator<std::shared_ptr<DB::ObjectInfo>>>*, std::function<void (DB::FileProgress)>, bool, bool, bool) @ 0x00000000189951a0 8. /ClickHouse/src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:170:24: DB::ReadFromObjectStorageStep::createIterator() @ 0x0000000020405658 9. /ClickHouse/src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:98:5: DB::ReadFromObjectStorageStep::initializePipeline(DB::QueryPipelineBuilder&, DB::BuildQueryPipelineSettings const&) @ 0x00000000204048f1 10. /ClickHouse/src/Processors/QueryPlan/ISourceStep.cpp:20:5: DB::ISourceStep::updatePipeline(std::vector<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>, AllocatorWithMemoryTracking<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>>>, DB::BuildQueryPipelineSettings const&) @ 0x00000000202cf75e 11. /ClickHouse/src/Processors/QueryPlan/QueryPlan.cpp:214:47: DB::QueryPlan::buildQueryPipeline(DB::QueryPlanOptimizationSettings const&, DB::BuildQueryPipelineSettings const&, bool) @ 0x0000000020342ec9 12. /ClickHouse/src/Storages/StorageMerge.cpp:1223:31: DB::ReadFromMerge::buildPipeline(DB::ReadFromMerge::ChildPlan&, DB::QueryProcessingStage::Enum) const @ 0x000000001eed7f9b 13. /ClickHouse/src/Storages/StorageMerge.cpp:582:32: DB::ReadFromMerge::initializePipeline(DB::QueryPipelineBuilder&, DB::BuildQueryPipelineSettings const&) @ 0x000000001eed755d 14. /ClickHouse/src/Processors/QueryPlan/ISourceStep.cpp:20:5: DB::ISourceStep::updatePipeline(std::vector<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>, AllocatorWithMemoryTracking<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>>>, DB::BuildQueryPipelineSettings const&) @ 0x00000000202cf75e 15. /ClickHouse/src/Processors/QueryPlan/QueryPlan.cpp:214:47: DB::QueryPlan::buildQueryPipeline(DB::QueryPlanOptimizationSettings const&, DB::BuildQueryPipelineSettings const&, bool) @ 0x0000000020342ec9 16. /ClickHouse/src/Interpreters/InterpreterSelectQueryAnalyzer.cpp:399:34: DB::InterpreterSelectQueryAnalyzer::buildQueryPipeline() @ 0x000000001a693cb0 17. /ClickHouse/src/Interpreters/InterpreterSelectQueryAnalyzer.cpp:364:29: DB::InterpreterSelectQueryAnalyzer::execute() @ 0x000000001a69389e 18. /ClickHouse/src/Interpreters/executeQuery.cpp:1824:40: DB::executeQueryImpl(char const*, char const*, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum, std::unique_ptr<DB::ReadBuffer, std::default_delete<DB::ReadBuffer>>&, boost::intrusive_ptr<DB::IAST>&, std::shared_ptr<DB::ImplicitTransactionControlExecutor>, std::function<void ()>, DB::QueryResultDetails&) @ 0x000000001aa21fff 19. /ClickHouse/src/Interpreters/executeQuery.cpp:2196:11: DB::executeQuery(std::basic_string_view<char, std::char_traits<char>>, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum) @ 0x000000001aa1ad8a 20. /ClickHouse/src/Server/TCPHandler.cpp:832:68: DB::TCPHandler::runImpl() @ 0x000000001fbcacb5 21. /ClickHouse/src/Server/TCPHandler.cpp:3075:9: DB::TCPHandler::run() @ 0x000000001fbe84e4 22. /ClickHouse/base/poco/Net/src/TCPServerConnection.cpp:40:3: Poco::Net::TCPServerConnection::start() @ 0x000000002708af8e 23. /ClickHouse/base/poco/Net/src/TCPServerDispatcher.cpp:115:42: Poco::Net::TCPServerDispatcher::run() @ 0x000000002708b4d2 24. /ClickHouse/base/poco/Foundation/src/ThreadPool.cpp:207:14: Poco::PooledThread::run() @ 0x000000002705443f 25. /ClickHouse/base/poco/Foundation/src/Thread_POSIX.cpp:341:27: Poco::ThreadImpl::runnableEntry(void*) @ 0x000000002705280f 26. start_thread @ 0x000000000009caa4 27. __GI___clone3 @ 0x0000000000129c6 ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #102033 - Merged into: `26.7.1.448` (included in `26.7` and later) - Backported to: `26.5.7.46` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/107334",
        "createdAt": "2026-06-12T13:57:01Z",
        "updatedAt": "2026-08-13T01:36:06Z",
        "timestamp": "2026-08-13T01:36:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "bug",
          "fuzz",
          "comp-datalake"
        ],
        "author": "PedroTadim",
        "state": "closed",
        "assignees": [
          "SmitaRKulkarni"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:109216",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Wrong DISTINCT results (and debug abort) with partial_merge JOIN + optimize_distinct_in_order — Equal values are not contiguous",
        "text": "# `Logical error: 'Equal values are not contiguous within the range assumed to be sorted'` — `DISTINCT` in order over `partial_merge` JOIN (wrong results in release) ## Summary With `join_algorithm = 'partial_merge'` (or `'prefer_partial_merge'`), the query plan assumes the join preserves the left stream's sort order and applies the pre-`DISTINCT` stage as `DistinctSortedStreamTransform` (`optimize_distinct_in_order`, default on). But `PartialMergeJoin` matches the left stream against the right side **per right block**: when the right table has more than one block/part, the left key ranges are re-emitted for each right block, so equal left-key values are no longer contiguous in the joined stream. - **Debug builds**: `chassert` in `findEqualRangeEndAssumeSorted.h:88` fires → `Logical error: 'Equal values are not contiguous within the range assumed to be sorted'` → abort. Stack: `DistinctSortedStreamTransform::transform → getEqualRangeEndAssumeSorted → checkEqualRangeEndAssumeSorted (ColumnVector<UInt32>)`. - **Release builds**: the check is compiled out, so `DistinctSortedStreamTransform` computes equal-ranges on a stream that violates its precondition — **`DISTINCT` can silently return duplicate rows** (wrong result). Found by the AST-fuzzer correctness-oracle campaign (PR #99980): two independent fleet hits (2026-06-29 `DISTINCT … FULL OUTER JOIN … QUALIFY`, 2026-06-30 `DISTINCT … INNER JOIN … ORDER BY`), minimized from the second snapshot. ## Minimal repro Flaky by nature (depends on chunk arrival order across the two right-side parts; `max_threads = 4` observed rate ≈ 40–50%, single attempt often enough on a debug build). Run a few times: ```sql CREATE TABLE t1 (x UInt32, y UInt64) ENGINE = MergeTree ORDER BY (x, y); CREATE TABLE t2 (x UInt32, y UInt64) ENGINE = MergeTree ORDER BY (x, y); INSERT INTO t1 VALUES (0,0),(1,10),(2,20),(3,30),(4,40); INSERT INTO t2 VALUES (2,21),(2,22),(4,41); INSERT INTO t2 VALUES (0,0),(4,42),(5,50); -- second part is load-bearing SET join_algorithm = 'prefer_partial_merge'; SELECT DISTINCT t1.*, t2.* FROM t1 INNER JOIN t2 ON intDiv(t2.y, 2147483647) = toUInt64(t1.x); ``` Also reproduces in `clickhouse-local` with the same script (abort, exit 134). The original fuzzed query had a baroque `ON and(if(...), key = key)` condition and an `ORDER BY … DESC NULLS LAST` — neither is needed (verified: pure-equality `ON` hits, no-`ORDER BY` hits). ## Mechanism (EXPLAIN PIPELINE) `prefer_partial_merge` (bug path): ``` DistinctTransform DistinctSortedStreamTransform × 2 <-- assumes sorted-by-left-prefix stream JoiningTransform × 2 2 → 1 <-- PartialMergeJoin MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) × 2 <-- left read in order FillingRightJoinSide MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) ``` `hash` (correct path): plain `DistinctTransform`, no in-order read, no sorted assumption. The left read is in `(x, y)` order and the plan carries that sort property through the join to justify `DistinctSortedStreamTransform`. `PartialMergeJoin` breaks the property: it processes the right side block-by-block and re-emits matching left ranges for each right block, so a left key that matches rows in two right blocks appears in two separate runs. Controls (each verified with a 10-attempt harness against the snapshot server): - `join_algorithm = 'hash'` → never reproduces (plain `DistinctTransform`). - `optimize_distinct_in_order = 0` → does not reproduce (MISS ×10). - `join_algorithm = 'partial_merge'` (strict) → reproduces, same as `prefer_partial_merge`. - `SELECT DISTINCT t1.*` only (no `t2` columns) → does not reproduce (MISS ×10); the `DISTINCT` set must include right-side columns. - Baroque original `ON and(if(...), key = key)` and the `ORDER BY … DESC NULLS LAST` → both unnecessary (pure-equality, no-`ORDER BY` variant reproduces). ## Environment - HEAD `01e08fd49182` (2026-06-28 master merge), debug build, aarch64. - Snapshot: `tmp/fuzz_lab/runs/merged7/crashes/inst-06-20260630-184002` (original), also `inst-12-20260629-115230` (same family, first sighting), `inst-05-20260701-012001` (`Sort order of blocks violated` — likely same root: a sort-order property claimed across `PartialMergeJoin` and violated downstream in a merge-sort transform instead of DISTINCT). ## Suggested fix direction `PartialMergeJoin` (and any join that rescans the left side per right block) must not report the left input's sort description as preserved on its output — the plan should either drop the sort property across it (forcing plain `DistinctTransform` / a re-sort before order-dependent consumers) or the join should be excluded from `optimize_distinct_in_order` / read-in-order propagation.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/109216",
        "createdAt": "2026-07-02T19:25:34Z",
        "updatedAt": "2026-08-13T13:05:26Z",
        "timestamp": "2026-08-13T13:05:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "fuzz",
          "sqlancer"
        ],
        "author": "qoega",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:109308",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Implement serde for `WindowStep` and support it for `makeDistributed`",
        "text": "We have basic infrastructure for `QueryPlan` serialization and deserialization: https://github.com/ClickHouse/clickhouse/blob/master/src/Processors/QueryPlan/QueryPlan.h#L101-L104. It's implemented for multiple steps, for example `GROUP BY`: https://github.com/ClickHouse/clickhouse/blob/master/src/Processors/QueryPlan/AggregatingStep.cpp#L945-L1105. However implementation for `WindowStep` is missing https://github.com/ClickHouse/clickhouse/blob/master/src/Processors/QueryPlan/WindowStep.h#L11. It makes it impossible to run distributed queries with `make_distributed_plan=1` if they contain any window functions. So the task is to support `WindowStep`. <!-- ch-version-info:start --> ### Version info - Resolved by: #109802 - Merged into: `26.8.1.615` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/109308",
        "createdAt": "2026-07-03T15:00:14Z",
        "updatedAt": "2026-08-13T13:33:28Z",
        "timestamp": "2026-08-13T13:33:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "unfinished code"
        ],
        "author": "alesapin",
        "state": "closed",
        "assignees": [
          "davenger"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:109326",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Enable `rewrite_in_to_join` for makeDistributed by default",
        "text": "Currently pre-built sets for `IN (subquery)` clauses don't work well, so as a workaround it make sense to enable `rewrite_in_to_join=1` by default for distributed queries v2.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/109326",
        "createdAt": "2026-07-03T15:59:04Z",
        "updatedAt": "2026-08-13T13:00:40Z",
        "timestamp": "2026-08-13T13:00:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "unfinished code",
          "make it worse"
        ],
        "author": "alesapin",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:109476",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Automatically change some settings if `make_distributed_plan=1` is enabled",
        "text": "* New distributed execution doesn't support some kinds of `IN (subquery)` queries. However we have amazing feature which rewrites `IN -> JOIN`: https://github.com/ClickHouse/ClickHouse/pull/83991. So `rewrite_in_to_join=1` should be enabled automatically if `make_distributed_plan=1` is specified. * Parallel replicas-style reads are not supported by `make_distributed_plan=1`. Instead it uses less efficient static distribution of work. So `enable_parallel_replicas=0` should be set automatically. * Recent optimization for correlated subqueries is not supported yet https://github.com/ClickHouse/ClickHouse/pull/91205. Should be disabled automatically with `correlated_subqueries_use_in_memory_buffer=0`. * `use_skip_indexes_on_data_read` should be disabled just in case. In case of huge table we can use distributed index analysis. * `compile_expressions=0` -- doesn't work yet, need to fix <!-- ch-version-info:start --> ### Version info - Resolved by: #112463 - Merged into: `26.8.1.506` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/109476",
        "createdAt": "2026-07-06T10:38:23Z",
        "updatedAt": "2026-08-13T13:17:21Z",
        "timestamp": "2026-08-13T13:17:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "unfinished code"
        ],
        "author": "alesapin",
        "state": "closed",
        "assignees": [
          "alesapin"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:109678",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "SIGSEGV: data race on `AsynchronousBoundedReadBuffer::prefetch_future` (ParquetV3 prefetcher)",
        "text": "### Describe the bug Having difficulty reproducing it. # The crash `26.7.1.569`, thread from the Parquet prefetcher fast pool: ``` <Fatal> BaseDaemon: Address: 0x28. Access: read. Address not mapped to object. 3. pthread_mutex_lock 5. std::mutex::lock() 6. std::__assoc_state<DB::IAsynchronousReader::Result>::move() future:636 (future::get) 7. DB::AsynchronousBoundedReadBuffer::readBigAt() AsynchronousBoundedReadBuffer.cpp:486 8. DB::Parquet::Prefetcher::readSync() Prefetcher.cpp:102 9. DB::Parquet::Prefetcher::runTask() Prefetcher.cpp:523 ... DB::ThreadPoolCallbackRunnerFast::threadFunction() (query: SELECT * FROM `d4`.`t93` LIMIT 10000 INTO OUTFILE '/tmp/file.data' TRUNCATE FORMAT Null;) ``` `d4.t93` is an S3-backed data-lake table (`engine: iceberg`) read through `ParquetV3`. ## Root cause — confirmed by code (a real ClickHouse bug, unreported) `AsynchronousBoundedReadBuffer::readBigAt()` is a positioned read intended to be safe for concurrent use, but it also *consumes an in-flight sequential prefetch* with no synchronization (`AsynchronousBoundedReadBuffer.cpp:480-489`): ```cpp if (prefetch_future.valid()) { ... result = prefetch_future.get(); // line 486 prefetch_future = {}; // line 488 ... } ``` `Prefetcher::readSync()` calls `reader->readBigAt(...)` from multiple `runTask` threads on the fast pool **without a lock in `ReadMode::RandomRead`** (`Prefetcher.cpp:102`) — note the `SeekAndRead` branch *does* take `read_mutex`. So two prefetcher threads race on the shared `prefetch_future`: ``` A: prefetch_future.valid() -> true B: prefetch_future.get(); prefetch_future = {} A: prefetch_future.get() on the moved-from future -> __state_ == nullptr -> __assoc_state::move() locks &__mut_ at offset 0x28 of nullptr -> SIGSEGV at 0x28 ``` The `0x28` fault address is exactly the offset of the `std::mutex` inside a null `__assoc_state`. Enablers in the fuzzer run: `local_filesystem_read_method='pread_fake_async'` (routes even local reads through the async buffer), `input_format_parquet_use_offset_index=1` (RandomRead → `readBigAt`), and `ThreadFuzzer` active (frame 4), which widened the tiny `valid()`-vs-`get()` window. ## Suggested fix Make the prefetch consumption in `readBigAt()` thread-safe, or take `read_mutex` in the `ReadMode::RandomRead` branch of `Prefetcher::readSync()`. Cleanest: guard the `if (prefetch_future.valid())` block in `readBigAt()` (e.g. atomically claim the future under a mutex / `std::exchange`), since `readBigAt()` is contractually concurrent. ### How to reproduce No success in reproducing it a second time. `d4`.`t93` is an IcebergS3 table. ### Error message and/or stacktrace Stack trace: ``` [node0] 2026.07.07 14:53:56.692776 [ 2113 ] <Fatal> BaseDaemon: ######################################## [node0] 2026.07.07 14:53:56.692837 [ 2113 ] <Fatal> BaseDaemon: (version 26.7.1.569 (official build), build id: 5654F8064F99BC8E4D1CB60CB74EBEE5400C905D, git hash: 20e1dcae5f06d30435d11bfc051996307d1972e7) (from thread 3980) (query_id: 2093fbc7-1d33-4441-ba60-ae33b42628cf) (query: SELECT * FROM `d4`.`t93` LIMIT 10000 INTO OUTFILE '/tmp/file.data' TRUNCATE FORMAT Null;) Received signal Segmentation fault (11) [node0] 2026.07.07 14:53:56.692853 [ 2113 ] <Fatal> BaseDaemon: Address: 0x28. Access: read. Address not mapped to object. [node0] 2026.07.07 14:53:56.692865 [ 2113 ] <Fatal> BaseDaemon: Stack trace: 0x000071b47cb33ef5 0x000058f615feec6f 0x000058f62ebed889 0x000058f619d12721 0x000058f619e564fb 0x000058f6230c1852 0x000058f6230c4082 0x000058f6230c7e33 0x000058f6177ce96d 0x000058f6160a3d89 0x000058f6160acd8b 0x000058f6160a0fbc 0x000058f6160aa14e 0x000071b47cb30ac3 0x000071b47cbc28d0 [node0] 2026.07.07 14:53:56.692930 [ 2113 ] <Fatal> BaseDaemon: 3. pthread_mutex_lock@@GLIBC_2.2.5 @ 0x0000000000097ef5 [node0] 2026.07.07 14:53:56.866200 [ 2113 ] <Fatal> BaseDaemon: 4. src/Common/ThreadFuzzer.cpp:447:1: pthread_mutex_lock @ 0x00000000155c8c6f [node0] 2026.07.07 14:53:56.928475 [ 2113 ] <Fatal> BaseDaemon: 5.0. inlined from contrib/llvm-project/libcxx/include/__thread/support/pthread.h:95: std::__libcpp_mutex_lock[abi:sqe220101](pthread_mutex_t*) [node0] 2026.07.07 14:53:56.928498 [ 2113 ] <Fatal> BaseDaemon: 5. contrib/llvm-project/libcxx/src/mutex.cpp:30:10: std::mutex::lock() @ 0x000000002e1c7889 [node0] 2026.07.07 14:53:57.044961 [ 2113 ] <Fatal> BaseDaemon: 6.0. inlined from contrib/llvm-project/libcxx/include/__mutex/unique_lock.h:44: unique_lock [node0] 2026.07.07 14:53:57.044988 [ 2113 ] <Fatal> BaseDaemon: 6. contrib/llvm-project/libcxx/include/future:636:11: std::__assoc_state<DB::IAsynchronousReader::Result>::move() @ 0x00000000192ec721 [node0] 2026.07.07 14:53:57.132500 [ 2113 ] <Fatal> BaseDaemon: 7.0. inlined from contrib/llvm-project/libcxx/include/future:986: std::future<DB::IAsynchronousReader::Result>::get() [node0] 2026.07.07 14:53:57.132527 [ 2113 ] <Fatal> BaseDaemon: 7. src/Disks/IO/AsynchronousBoundedReadBuffer.cpp:486:15: DB::AsynchronousBoundedReadBuffer::readBigAt(char*, unsigned long, unsigned long, std::function<bool (unsigned long)> const&) const @ 0x00000000194304fb [node0] 2026.07.07 14:53:57.332448 [ 2113 ] <Fatal> BaseDaemon: 8. src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:102:29: DB::Parquet::Prefetcher::readSync(char*, unsigned long, unsigned long) @ 0x000000002269b852 [node0] 2026.07.07 14:53:57.359378 [ 2113 ] <Fatal> BaseDaemon: 9. src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:523:13: DB::Parquet::Prefetcher::runTask(DB::Parquet::Prefetcher::Task*) @ 0x000000002269e082 [node0] 2026.07.07 14:53:57.382614 [ 2113 ] <Fatal> BaseDaemon: 10.0. inlined from src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:424: operator() [node0] 2026.07.07 14:53:57.382649 [ 2113 ] <Fatal> BaseDaemon: 10.1. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>::type std::__invoke[abi:sqe220101]<DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>(DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&) [node0] 2026.07.07 14:53:57.382681 [ 2113 ] <Fatal> BaseDaemon: 10.2. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>(DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&) [node0] 2026.07.07 14:53:57.382696 [ 2113 ] <Fatal> BaseDaemon: 10.3. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>(DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&) [node0] 2026.07.07 14:53:57.382707 [ 2113 ] <Fatal> BaseDaemon: 10. contrib/llvm-project/libcxx/include/__functional/function.h:443:17: ? @ 0x00000000226a1e33 [node0] 2026.07.07 14:53:57.506449 [ 2113 ] <Fatal> BaseDaemon: 11.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.506480 [ 2113 ] <Fatal> BaseDaemon: 11.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.506489 [ 2113 ] <Fatal> BaseDaemon: 11. src/Common/threadPoolCallbackRunner.cpp:225:12: DB::ThreadPoolCallbackRunnerFast::threadFunction() @ 0x0000000016da896d [node0] 2026.07.07 14:53:57.592515 [ 2113 ] <Fatal> BaseDaemon: 12.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.592537 [ 2113 ] <Fatal> BaseDaemon: 12.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.592548 [ 2113 ] <Fatal> BaseDaemon: 12. src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x000000001567dd89 [node0] 2026.07.07 14:53:57.627238 [ 2113 ] <Fatal> BaseDaemon: 13.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.627266 [ 2113 ] <Fatal> BaseDaemon: 13.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.627274 [ 2113 ] <Fatal> BaseDaemon: 13.2. inlined from src/Common/ThreadPool.cpp:1293: operator() [node0] 2026.07.07 14:53:57.627294 [ 2113 ] <Fatal> BaseDaemon: 13.3. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) [node0] 2026.07.07 14:53:57.627310 [ 2113 ] <Fatal> BaseDaemon: 13.4. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) [node0] 2026.07.07 14:53:57.627321 [ 2113 ] <Fatal> BaseDaemon: 13.5. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) [node0] 2026.07.07 14:53:57.627329 [ 2113 ] <Fatal> BaseDaemon: 13. contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x0000000015686d8b [node0] 2026.07.07 14:53:57.638085 [ 2113 ] <Fatal> BaseDaemon: 14.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.638103 [ 2113 ] <Fatal> BaseDaemon: 14.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.638130 [ 2113 ] <Fatal> BaseDaemon: 14. src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x000000001567afbc [node0] 2026.07.07 14:53:57.658980 [ 2113 ] <Fatal> BaseDaemon: 15.0. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) [node0] 2026.07.07 14:53:57.659015 [ 2113 ] <Fatal> BaseDaemon: 15.1. inlined from contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) [node0] 2026.07.07 14:53:57.659030 [ 2113 ] <Fatal> BaseDaemon: 15. contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x000000001568414e [node0] 2026.07.07 14:53:57.659079 [ 2113 ] <Fatal> BaseDaemon: 16. start_thread @ 0x0000000000094ac3 [node0] 2026.07.07 14:53:57.659105 [ 2113 ] <Fatal> BaseDaemon: 17. __GI___clone3 @ 0x00000000001268d0 [node0] 2026.07.07 14:53:58.543009 [ 2113 ] <Fatal> BaseDaemon: Integrity check of the executable successfully passed (checksum: C14EA45234712E0579DCDA202C426E57) [node0] 2026.07.07 14:53:58.704024 [ 2113 ] <Fatal> BaseDaemon: ClickHouse version 26.7.1.569 is old and should be upgraded to the latest version. [node0] 2026.07.07 14:53:58.704285 [ 2113 ] <Fatal> BaseDaemon: Changed settings: min_compress_block_size = 32, use_strict_insert_block_limits = false, max_insert_block_size_bytes = 1024, min_insert_block_size_rows = 1824, min_insert_block_size_bytes = 6127879, min_insert_block_size_rows_for_materialized_views = 4096, min_insert_block_size_bytes_for_materialized_views = 0, min_external_table_block_size_bytes = 32, max_joined_block_size_rows = 0, min_joined_block_size_rows = 1, min_joined_block_size_bytes = 6511566, joined_block_split_single_row = false, parallel_non_joined_rows_processing = false, max_final_threads = 2, max_threads_for_indexes = 1, max_threads = 32, max_parsing_threads = 2, max_read_buffer_size = 2048, max_read_buffer_size_remote_fs = 4, use_hedged_requests = true, s3_truncate_on_insert = true, azure_truncate_on_insert = true, s3_create_new_file_on_insert = false, s3_skip_empty_files = false, s3_allow_parallel_part_upload = true, s3queue_enable_logging_to_s3queue_log = true, hdfs_create_new_file_on_insert = true, dictionary_validate_primary_key_type = false, distributed_background_insert_batch = true, optimize_move_to_prewhere = true, enable_multiple_prewhere_read_steps = true, move_primary_key_columns_to_end_of_prewhere = false, allow_reorder_prewhere_conditions = false, load_balancing = 'hostname_levenshtein_distance', allow_suspicious_low_cardinality_types = true, allow_suspicious_fixed_string_types = true, allow_suspicious_indices = true, allow_suspicious_ttl_expressions = true, allow_suspicious_variant_types = true, allow_suspicious_primary_key = true, allow_suspicious_types_in_group_by = true, allow_suspicious_types_in_order_by = true, variant_throw_on_type_mismatch = false, dynamic_throw_on_type_mismatch = true, compile_expressions = false, min_count_to_compile_expression = 1, min_count_to_compile_sort_description = 0, group_by_two_level_threshold_bytes = 1024, distributed_aggregation_memory_efficient = false, aggregation_memory_efficient_merge_threads = 32, enable_positional_arguments = true, enable_positional_arguments_for_projections = false, allow_nonconst_timezone_arguments = true, enable_time_time64_type = true, function_locate_has_mysql_compatible_argument_order = false, parallel_distributed_insert_select = 1, optimize_distributed_group_by_sharding_key = true, optimize_skip_unused_shards = false, allow_nondeterministic_optimize_skip_unused_shards = true, min_chunk_bytes_for_parallel_parsing = 498990, merge_tree_min_rows_for_concurrent_read = 16384, merge_tree_min_bytes_for_concurrent_read = 4, merge_tree_min_rows_for_seek = 1, merge_tree_max_bytes_to_use_cache = 4096, enable_automatic_decision_for_merging_across_partitions_for_final = true, split_parts_ranges_into_intersecting_and_non_intersecting_final = true, split_intersecting_parts_ranges_into_layers_final = false, defer_partition_pruning_after_final = false, optimize_min_equality_disjunction_chain_length = 0, optimize_min_inequality_conjunction_chain_length = 0, min_bytes_to_use_direct_io = 4, min_bytes_to_use_mmap_io = 1, use_lightweight_primary_key_index_analysis = true, use_partition_pruning = false, use_skip_indexes_if_final_exact_mode = false, use_skip_indexes_on_data_read = true, use_statistics_for_part_pruning = false, use_top_k_dynamic_filtering_for_variable_length_types = true, query_plan_max_limit_for_top_k_optimization = 0, per_part_index_stats = false, secondary_indices_enable_bulk_filtering = true, max_streams_to_max_threads_ratio = 0.74030601978302, log_queries = false, distributed_product_mode = 'allow', deduplicate_insert_select = 'enable_when_possible', insert_quorum_parallel = false, select_sequential_consistency = 1, update_sequential_consistency = true, update_parallel_mode = 'async', table_function_remote_max_addresses = 62, enable_http_compression = false, count_distinct_implementation = 'uniqExact', send_profile_events = true, http_wait_end_of_query = false, join_output_by_rowlist_perkey_rows_threshold = 16, query_plan_join_swap_table = false, query_plan_optimize_join_order_limit = 64, query_plan_optimize_join_order_randomize = 0, enable_join_transitive_predicates = true, preferred_block_size_bytes = 16384, parts_to_throw_insert = 32, number_of_mutations_to_delay = 32, number_of_mutations_to_throw = 2, ignore_on_cluster_for_replicated_udf_queries = true, insert_allow_materialized_columns = true, http_max_fields = 512, http_make_head_request = false, use_index_for_in_with_subqueries = false, analyze_index_with_space_filling_curves = false, allow_key_condition_coalesce_rewrite = true, empty_result_for_aggregation_by_constant_keys_on_empty_set = false, allow_distributed_ddl = true, allow_suspicious_codecs = true, opentelemetry_start_keeper_trace_probability = 0.9900000095367432, max_bytes_before_external_group_by = 8, max_bytes_ratio_before_external_group_by = 0.99, prefer_external_sort_block_bytes = 0, max_bytes_ratio_before_external_sort = 0.1, max_bytes_before_remerge_sort = 10485760, remerge_sort_lowered_memory_bytes_ratio = 0.10000000149011612, max_execution_time = 60., allow_fuzz_query_functions = true, rows_before_aggregation = true, cross_to_inner_join_rewrite = 2, cross_join_min_rows_to_compress = 8, default_max_bytes_in_join = 4096, temporary_files_codec = 'none', max_rows_to_transfer = 16384, max_reverse_dictionary_lookup_cache_size_bytes = 16, log_profile_events = true, log_query_views = true, enable_optimize_predicate_expression_to_final_subquery = true, allow_push_predicate_when_subquery_contains_with = true, allow_custom_error_code_in_throwif = true, prefer_localhost_replica = true, allow_ddl = true, parallel_view_processing = false, enable_unaligned_array_join = false, read_in_order_use_virtual_row_per_block = true, optimize_aggregation_in_order = false, optimize_aggregation_in_order_limit = false, read_in_order_use_buffering = true, cancel_http_readonly_queries_on_client_close = false, allow_hyperscan = false, reject_expensive_hyperscan_regexps = true, allow_introspection_functions = true, allow_execute_multiif_columnar = true, parsedatetime_e_requires_space_padding = true, functions_h3_default_if_invalid = true, check_query_single_value_result = true, allow_drop_detached = true, allow_replace_partition_from_empty_source = true, dynamic_disk_allow_from_zk = false, max_parts_to_move = 0, max_partition_size_to_drop = 1, glob_expansion_max_elements = 8, show_table_uuid_in_table_create_query_if_not_nil = true, enable_scalar_subquery_optimization = true, optimize_trivial_count_query = false, optimize_trivial_approximate_count_query = true, optimize_trivial_group_by_limit_query = false, optimize_count_from_files = true, use_cache_for_count_from_files = true, optimize_respect_aliases = true, enable_lightweight_delete = true, lightweight_delete_mode = 'lightweight_update', optimize_normalize_count_variants = false, optimize_injective_functions_inside_uniq = false, optimize_arithmetic_operations_in_aggregate_functions = true, optimize_redundant_functions_in_order_by = true, optimize_if_transform_strings_to_enum = true, optimize_substitute_columns = false, normalize_function_names = true, allow_materialized_view_with_bad_select = true, materialized_views_squash_parallel_inserts = true, use_compact_format_in_distributed_parts_names = false, validate_polygons = true, recursive_cte_max_steps_in_type_inference = 10, allow_settings_after_format_in_insert = true, allow_nondeterministic_mutations = true, cast_keep_nullable = true, allow_non_metadata_alters = true, enable_materialized_cte = true, flatten_nested = false, optimize_skip_merged_partitions = false, optimize_use_projections = false, optimize_use_implicit_projections = false, optimize_use_projection_filtering = false, insert_null_as_default = true, enable_lightweight_update = true, apply_patch_parts = true, apply_patch_parts_join_cache_buckets = 64, mutations_execute_subqueries_on_initiator = true, mutations_max_literal_size_to_replace = 0, delta_lake_log_metadata = true, delta_lake_reload_schema_for_consistency = false, iceberg_metadata_log_level = 'metadata', iceberg_data_file_size_upper_threshold_compaction = 4, iceberg_compaction_delay_bias = 2652., iceberg_compaction_data_cleanup = 4642., use_query_cache = true, enable_writes_to_query_cache = false, enable_reads_from_query_cache = false, query_cache_for_subqueries = true, query_cache_nondeterministic_function_handling = 'ignore', query_cache_system_table_handling = 'ignore', query_cache_max_size_in_bytes = 16, query_cache_max_entries = 1024, query_cache_min_query_runs = 5481, query_cache_compress_entries = false, query_cache_squash_partial_results = true, use_query_condition_cache = false, optimize_rewrite_aggregate_function_with_if = false, optimize_rewrite_has_to_in = true, optimize_dictget_tuple_element = true, use_hash_table_stats_for_join_reordering = true, allow_experimental_kafka_offsets_storage_in_keeper = true, enable_software_prefetch_in_aggregation = true, allow_aggregate_partitions_independently = false, force_aggregate_partitions_independently = true, allow_limit_by_partitions_independently = false, min_hit_rate_to_use_consecutive_keys_optimization = 0.5, engine_file_empty_if_not_exists = false, engine_file_truncate_on_insert = true, enable_url_encoding = true, database_replicated_allow_replicated_engine_arguments = 1, database_replicated_allow_heavy_create = true, cloud_mode_engine = 0, external_storage_max_read_rows = 16, external_storage_max_read_bytes = 2048, allow_experimental_correlated_subqueries = true, max_streams_for_union_step = 32, optimize_aggregators_of_group_by_keys = false, optimize_injective_functions_in_group_by = true, legacy_column_name_of_tuple_literal = true, query_plan_enable_optimizations = true, query_plan_lift_up_array_join = false, query_plan_push_down_limit = true, query_plan_top_k_through_join = false, query_plan_split_filter = true, query_plan_merge_expressions = true, query_plan_filter_push_down = false, query_plan_convert_outer_join_to_inner_join = false, query_plan_convert_any_join_to_semi_or_anti_join = true, query_plan_merge_filter_into_join_condition = true, optimize_prewhere_after_pushdown = false, query_plan_execute_functions_after_sorting = false, query_plan_reuse_storage_ordering_for_window_functions = true, query_plan_lift_up_union = true, query_plan_read_in_order = false, query_plan_read_in_order_through_join = false, query_plan_aggregation_in_order = false, query_plan_optimize_lazy_final = false, max_rows_for_lazy_final = 4, max_bytes_for_lazy_final = 9963842, min_filtered_ratio_for_lazy_final = 0.10000000149011612, enable_lazy_columns_replication = false, enable_software_prefetch_in_join = false, correlated_subqueries_substitute_equivalent_expressions = true, correlated_subqueries_use_in_memory_buffer = true, function_range_max_elements_in_block = 0, function_base58_max_input_size = 4, local_filesystem_read_method = 'pread_fake_async', local_filesystem_read_prefetch = true, remote_filesystem_read_prefetch = true, merge_tree_min_rows_for_concurrent_read_for_remote_filesystem = 8, merge_tree_min_bytes_for_concurrent_read_for_remote_filesystem = 2, remote_read_min_bytes_for_seek = 0, merge_tree_min_bytes_per_task_for_remote_reading = 1, merge_tree_determine_task_size_by_prewhere_columns = false, merge_tree_min_read_task_size = 8192, async_insert_max_query_number = 2, cluster_function_process_archive_on_multiple_nodes = false, max_streams_for_files_processing_in_cluster_functions = 32, enable_filesystem_cache = false, filesystem_cache_name = 'fcache0', enable_filesystem_cache_on_write_operations = false, filesystem_cache_max_download_size = 3724327, throw_on_error_from_cache_on_write_operations = false, filesystem_cache_segments_batch_size = 50, filesystem_cache_boundary_alignment = 4, use_page_cache_with_distributed_cache = true, use_page_cache_for_object_storage = true, read_from_page_cache_if_exists_otherwise_bypass_cache = false, page_cache_inject_eviction = true, page_cache_block_size = 2048, page_cache_max_coalesced_bytes = 4096, allow_prefetched_read_pool_for_local_filesystem = true, prefetch_buffer_size = 0, filesystem_prefetch_step_bytes = 10485760, allow_calculating_subcolumns_sizes_for_merge_tree_reading = true, check_table_dependencies = true, check_named_collection_dependencies = false, allow_unrestricted_reads_from_keeper = true, allow_rank_dense_rank_arguments = false, schema_inference_use_cache_for_file = false, schema_inference_cache_require_modification_time_for_url = false, read_through_distributed_cache = false, distributed_cache_throw_on_error = true, distributed_cache_read_request_max_tries = 65536, distributed_cache_alignment = 8688762, distributed_cache_min_bytes_for_seek = 10485760, write_through_distributed_cache_buffer_size = 0, table_engine_read_through_distributed_cache = true, read_from_distributed_cache_if_exists_otherwise_bypass_cache = false, distributed_cache_registry_show_certificate_and_signature = false, filesystem_cache_enable_background_download_during_fetch = false, parallelize_output_from_storages = false, multiple_joins_try_to_keep_original_names = true, keeper_max_retries = 15, optimize_uniq_to_count = true, enable_order_by_all = true, allow_dynamic_type_in_join_keys = true, cast_string_to_variant_use_inference = false, enable_blob_storage_log_for_read_operations = true, allow_create_index_without_type = true, allow_named_collection_override_by_default = true, use_async_executor_for_materialized_views = true, short_circuit_function_evaluation_for_nulls_threshold = 1., allow_experimental_geo_types_in_iceberg = true, show_data_lake_catalogs_in_system_tables = true, delta_lake_throw_on_engine_predicate_error = true, delta_lake_insert_max_bytes_in_data_file = 4262856, allow_experimental_delta_lake_writes = true, use_iceberg_partition_pruning = true, extract_key_value_pairs_max_pairs_per_row = 0, allow_experimental_parallel_reading_from_replicas = 0, automatic_parallel_replicas_mode = 2, cluster_for_parallel_replicas = 'cluster0', parallel_replicas_allow_in_with_subquery = true, parallel_replicas_for_non_replicated_merge_tree = false, parallel_replicas_prefer_local_join = false, parallel_replicas_index_analysis_only_on_coordinator = false, parallel_replicas_support_projection = true, parallel_replicas_only_with_analyzer = false, parallel_replicas_allow_materialized_views = true, parallel_replicas_allow_view_over_mergetree = true, distributed_index_analysis = false, allow_experimental_database_iceberg = true, allow_experimental_database_unity_catalog = true, allow_experimental_database_glue_catalog = true, allow_experimental_analyzer = true, max_limit_for_vector_search_queries = 7447, vector_search_with_rescoring = false, use_join_disjunctions_push_down = true, shared_merge_tree_sync_parts_on_partition_operations = false, allow_general_join_planning = true, cluster_table_function_buckets_batch_size = 32, validate_enum_literals_in_operators = true, use_hive_partitioning = false, s3_uri_style = 'virtual_hosted', iceberg_insert_max_partitions = 2, min_outstreams_per_resize_after_split = 7655, enable_add_distinct_to_in_subqueries = true, jemalloc_profile_text_collapsed_use_count = true, allow_experimental_nullable_tuple_type = true, archive_adaptive_buffer_max_size_bytes = 8, max_bytes_before_external_join = 4236297, enable_join_fixed_hash_table_conversion = true, query_plan_min_columns_for_join_lazy_indexing = 5, allow_experimental_materialized_postgresql_table = true, allow_experimental_funnel_functions = true, allow_experimental_nlp_functions = true, allow_experimental_hash_functions = true, allow_experimental_time_series_table = true, allow_experimental_unique_key = true, allow_experimental_codecs = true, wait_changes_become_visible_after_commit_mode = 'wait_unknown', grace_hash_join_max_buckets = 1024, join_to_sort_minimum_perkey_rows = 2, allow_experimental_join_right_table_sorting = true, allow_experimental_json_lazy_type_hints = true, allow_statistics_optimize = true, use_statistics = false, allow_statistics = true, enable_full_text_index = true, query_plan_text_index_add_hint = true, use_text_index_like_evaluation_by_dictionary_scan = false, text_index_like_max_postings_to_read = 286, use_text_index_header_cache = false, text_index_lazy_intersection_density_threshold = 1., allow_experimental_window_view = true, allow_experimental_database_materialized_postgresql = true, allow_nullable_tuple_in_extracted_subcolumns = true, allow_experimental_database_hms_catalog = true, allow_experimental_kusto_dialect = true, allow_experimental_prql_dialect = true, allow_experimental_polyglot_dialect = true, allow_experimental_delta_kernel_rs = true, allow_insert_into_iceberg = true, allow_experimental_iceberg_compaction = true, allow_iceberg_remove_orphan_files = true, allow_experimental_expire_snapshots = true, write_full_path_in_iceberg_metadata = true, iceberg_metadata_compression_method = 'deflate', distributed_plan_force_exchange_kind = 'Streaming', distributed_plan_prefer_replicas_over_workers = true, allow_experimental_ytsaurus_table_engine = true, allow_experimental_ytsaurus_table_function = true, allow_experimental_ytsaurus_dictionary_source = true, enable_join_runtime_filters = true, join_runtime_bloom_filter_bytes = 2, join_runtime_bloom_filter_hash_functions = 1, join_runtime_filter_blocks_to_skip_before_reenabling = 16384, rewrite_in_to_join = false, allow_experimental_time_series_aggregate_functions = true, allow_experimental_paimon_storage_engine = true, use_paimon_partition_pruning = true, allow_experimental_object_storage_queue_hive_partitioning = true, query_plan_optimize_join_order_algorithm = 'dpsize', allow_experimental_database_paimon_rest_catalog = true, allow_experimental_ai_functions = true, allow_experimental_query_deduplication = true, update_insert_deduplication_token_in_dependent_materialized_views = true, allow_experimental_alias_table_engine = true, partial_merge_join_optimizations = 0, allow_not_comparable_types_in_order_by = true, allow_not_comparable_types_in_comparison_functions = true, enable_zstd_qat_codec = true, enable_deflate_qpl_codec = true, throw_if_deduplication_in_dependent_materialized_views_enabled_with_async_insert = false, async_insert_threads = 2, distributed_cache_read_alignment = 4, output_format_parallel_formatting = false, input_format_null_as_default = false, input_format_parquet_preserve_order = false, input_format_parquet_filter_push_down = false, input_format_parquet_enable_json_parsing = true, input_format_parquet_memory_low_watermark = 1024, input_format_parquet_memory_high_watermark = 1024, input_format_parquet_use_offset_index = true, input_format_parquet_local_time_as_utc = false, input_format_allow_seeks = false, input_format_orc_use_fast_decoder = true, input_format_orc_filter_push_down = true, input_format_orc_dictionary_as_low_cardinality = false, input_format_parquet_local_file_min_bytes_for_seek = 5762911, input_format_parquet_enable_row_group_prefetch = false, input_format_csv_use_best_effort_in_schema_inference = true, input_format_parquet_prefer_block_bytes = 4096, input_format_capn_proto_skip_fields_with_unsupported_types_in_schema_inference = true, schema_inference_make_columns_nullable = 3, schema_inference_make_json_columns_nullable = true, input_format_json_read_bools_as_numbers = true, input_format_json_read_bools_as_strings = false, input_format_json_try_infer_numbers_from_strings = true, input_format_json_infer_incomplete_types_as_strings = false, input_format_json_named_tuples_as_objects = true, input_format_json_throw_on_bad_escape_sequence = true, type_json_skip_duplicated_paths = true, input_format_try_infer_datetimes = true, input_format_try_infer_datetimes_only_datetime64 = true, input_format_protobuf_flatten_google_wrappers = false, input_format_tsv_skip_trailing_empty_lines = false, output_format_native_use_flattened_dynamic_and_json_serialization = false, input_format_values_accurate_types_of_literals = false, input_format_avro_null_as_default = false, input_format_binary_read_json_as_string = false, output_format_binary_write_json_as_string = true, output_format_json_quote_64bit_floats = true, output_format_json_array_of_rows = false, output_format_pretty_max_rows = 500, output_format_pretty_glue_chunks = 1, output_format_parquet_row_group_size = 8, output_format_parquet_row_group_size_bytes = 2872120, output_format_parquet_batch_size = 1571, output_format_parquet_write_bloom_filter = false, output_format_avro_codec = 'null', output_format_avro_sync_interval = 10000000, input_format_geojson_unsupported_geometry_handling = 'throw', output_format_pretty_multiline_fields = true, output_format_pretty_named_tuples_as_json = false, output_format_arrow_use_signed_indexes_for_dictionary = false, output_format_arrow_use_64_bit_indexes_for_dictionary = true, format_capn_proto_use_autogenerated_schema = true, output_format_sql_insert_include_column_names = true, input_format_bson_skip_fields_with_unsupported_types_in_schema_inference = true, format_display_secrets_in_show_and_select = true, validate_experimental_and_suspicious_types_inside_nested_types = false, show_create_query_identifier_quoting_rule = 'user_display', input_format_parquet_allow_geoparquet_parser = true ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #112573 - Merged into: `26.8.1.1240` (included in `26.8` and later) - Backported to: `26.7.4.27` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/109678",
        "createdAt": "2026-07-07T16:18:13Z",
        "updatedAt": "2026-08-13T14:31:48Z",
        "timestamp": "2026-08-13T14:31:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "crash",
          "fuzz",
          "comp-parquet-reader-v3"
        ],
        "author": "PedroTadim",
        "state": "closed",
        "assignees": [
          "grantholly-clickhouse"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:109974",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "[Bitmap] Function subBitmap execution result error",
        "text": "### Company or project name _No response_ ### Describe what's wrong ### Bug description `subBitmap` / `bitmapSubsetOffsetLimit` can return an incorrect subset when the input bitmap is stored in the small representation. In `src/AggregateFunctions/AggregateFunctionGroupBitmapData.h`, `rb_offset_limit()` : ```cpp if (isSmall()) { UInt64 count = 0; UInt64 offset_count = 0; auto it = small.begin(); for (; it != small.end() && offset_count < offset; ++it) ++offset_count; for (; it != small.end() && count < limit; ++it, ++count) r1.add(it->getValue()); return count; } ``` This assumes the `SmallSet ` is **sorted**, but that is not guaranteed. ### Reproduction ```text VM-20-3-centos :) SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res; SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res Query id: 678e3377-baeb-41a7-88c4-5bf601bd7220 ┌─res───┐ 1. │ [1,2] │ └───────┘ ``` excepted: `[3, 4]` ### Does it reproduce on the most recent release? Yes ### How to reproduce ### Reproduction ```text VM-20-3-centos :) SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res; SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res Query id: 678e3377-baeb-41a7-88c4-5bf601bd7220 ┌─res───┐ 1. │ [1,2] │ └───────┘ ``` excepted: `[3, 4]` ### Expected behavior _No response_ ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context _No response_ <!-- ch-version-info:start --> ### Version info - Resolved by: #110072 - Merged into: `26.8.1.1309` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/109974",
        "createdAt": "2026-07-10T10:16:47Z",
        "updatedAt": "2026-08-13T07:34:21Z",
        "timestamp": "2026-08-13T07:34:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "linrrzqqq",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:110281",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Join-order cardinality estimation runs before MergeTree partition/PK analysis and ignores pruned parts",
        "text": "## Describe the unexpected behaviour When column statistics are enabled (`use_statistics = 1`, the default), the join-order optimizer estimates relation cardinalities over **all** active parts of a MergeTree table, ignoring partition/PK pruning that execution will later perform. A query whose `WHERE` prunes to a 1,000-row partition of a 5,001,000-row table is planned as if it read ~2,500,500 rows. The error is identical under `use_statistics_cache = 0` and `= 1` — this is not a stale-cache problem; the cache-miss path lands on the same all-parts set. There is also an asymmetry: with statistics *disabled*, `estimateReadRowsCount` falls back to `ReadFromMergeTree::selectRangesToRead()` and gets pruning-aware `selected_rows`. Measured on the same query: | setting | fact-side estimate | best plan cost | |---|---|---| | `use_statistics = 1` (default) | 2,500,500 rows | 100000 | | `use_statistics = 0` | 1,000 rows (exact) | 1000 | So enabling column statistics currently makes the plan 100× worse by the optimizer's own cost model on partition-pruned queries. ## How to reproduce Verified on the official release `26.7.1.448` @ `cff54e151b69` (macOS arm64) and on `master` @ `5a9528b4db5` (2026-07-13, macOS arm64 debug build) — identical estimates on both. Probes were run on a freshly started server, twice per cache setting (bit-identical), with `collect_hash_table_stats_during_joins = 0` to exclude the runtime-feedback loop. ```sql CREATE TABLE fact (p UInt8, id UInt64) ENGINE = MergeTree PARTITION BY p ORDER BY id SETTINGS refresh_statistics_interval = 1; -- fast background stats refresh CREATE TABLE dim (id UInt64) ENGINE = MergeTree ORDER BY id SETTINGS refresh_statistics_interval = 1; SET materialize_statistics_on_insert = 1; INSERT INTO fact SELECT 1, number FROM numbers(5000000); -- NDV(id) = 5M INSERT INTO fact SELECT 2, number % 10 FROM numbers(1000); -- NDV(id) = 10 INSERT INTO dim SELECT number FROM numbers(100000); -- wait a few seconds for the background statistics refresh, then: SELECT count() FROM fact AS f INNER JOIN dim AS d ON f.id = d.id WHERE f.p = 2 SETTINGS use_statistics_cache = 0, collect_hash_table_stats_during_joins = 0; -- run again with use_statistics_cache = 1 (same estimates), -- and with use_statistics = 0 (pruning-aware fallback, see table above) ``` Observe the `optimizeJoin` trace (`--send_logs_level=trace`): ```text -- use_statistics_cache = 0 (identical under = 1, minus the Loading lines): a1.fact: Loading statistics optimizeJoin: estimate statistics fact: 2500500 rows, columns: [id: 2500500, p: 2] optimizeJoin: Estimated statistics for Filter f: 2500500 rows, columns: [__table1.p: 2, __table1.id: 2500500] a1.dim: Loading statistics optimizeJoin: estimate statistics dim: 100000 rows, columns: [id: 100315] JoinOrderOptimizer: Optimized join order in 0.18 ms, best plan cost: 100000, estimated cardinality: 100000 ``` The fact side is estimated at 2,500,500 rows — exactly 5,001,000 × 1/NDV(p) = 1/2, the all-parts selectivity model. A pruning-aware estimate could not exceed 1,000. NDV(id) is estimated at ~2,500,500 vs 10 actual in the surviving partition (250,000×). Estimates are bit-identical for `use_statistics_cache = 0` and `1` across repeated runs; `collect_hash_table_stats_during_joins = 0` excludes the runtime-feedback loop. Meanwhile execution does prune — `EXPLAIN indexes = 1` on the same query: ```text └──Join (JOIN FillRightFirst) │ f[2500500] ⋈ d[100000] <-- planner: all-parts estimate │ ... ├──ReadFromMergeTree (a1.fact) │ Parts: 1 | Granules: 1 <-- execution: pruned to the p=2 part │ Indexes: │ Min-Max │ Condition: (p in [2, 2]) │ Parts: 1/6 │ Granules: 1/612 │ Partition │ Condition: (p in [2, 2]) │ Parts: 1/1 ``` The same `EXPLAIN` output contains both the pruning-blind estimate on the Join node and the pruned read on the leaf. ## Plan damage (counterfactual) Same data restricted to p=2 (`CREATE TABLE fact_p2 ...; INSERT ... WHERE p=2`), same join without the `WHERE`: ```text optimizeJoin: estimate statistics fact_p2: 1000 rows, columns: [id: 10] optimizeJoin: estimate statistics dim: 100000 rows, columns: [id: 100315] JoinOrderOptimizer: Optimized join order in 0.01 ms, best plan cost: 996.86, estimated cardinality: 996 ``` versus `best plan cost: 100000, estimated cardinality: 100000` for the identical rows behind the pruned `WHERE p = 2`. The chosen build side flips from `dim` (100,000 rows) to the small fact side (1,000 rows), and the optimizer's own cost drops ~100× — i.e. today the pruning-blind estimate makes the optimizer reject the plan it would itself prefer given correctly-scoped cardinalities (the `use_statistics = 0` probe above confirms this on the original table, not just the counterfactual). On real partitioned fact tables (time-series with a sharp `WHERE` on the partition key) the same mechanism inflates every join-order decision. ## Root cause (hypothesis) At the time `optimizeJoin` requests the estimator, `ReadFromMergeTree::analyzed_result_ptr` appears to be unset, so `getParts()` falls back to `prepared_parts` (all active parts) (src/Processors/QueryPlan/ReadFromMergeTree.h, `getParts()`); `MergeTreeData::getConditionSelectivityEstimator` then folds statistics over that set. The statistics branch of `estimateReadRowsCount` (src/Processors/QueryPlan/Optimizations/optimizeJoin.cpp) returns before the `selectRangesToRead()` fallback that would have run the index analysis. `use_statistics_cache` is irrelevant to the scope: the background `refreshStatistics()` task also folds over all active parts, so cache hit and miss agree — both wrong-scoped. It is possible the current pass ordering is deliberate (index analysis is expensive, and join reordering changes which filters reach which table) — if so, this issue is a request to document that trade-off and close the gap, not a claim of an oversight. A fix presumably needs pass-ordering (make partition/PK analysis results available to join planning) plus composing statistics over surviving parts only. Note that once that happens, the table-wide `cached_estimator` *would* diverge from the pruned scope and needs a part-set-aware design (e.g. cache decoded per-part statistics and fold per query). ## Related - #55065 (column statistics umbrella) - #78441 (wrong build-side swap with `WHERE`) - #85126 (unknown `LIKE` selectivity) - #62870 (physical partition pruning through JOIN condition — different: that one is about execution, this one is about estimation scope)",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/110281",
        "createdAt": "2026-07-13T14:42:22Z",
        "updatedAt": "2026-08-13T06:16:49Z",
        "timestamp": "2026-08-13T06:16:49Z",
        "metrics": {
          "reactions": 1,
          "comments": 2
        },
        "labels": [
          "potential bug"
        ],
        "author": "skuznetsov-clickhouse",
        "state": "closed",
        "assignees": [
          "skuznetsov-clickhouse"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:110352",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Add aiFilter AI function (boolean predicate for WHERE / PREWHERE / JOIN)",
        "text": "### Company or project name _No response_ ### Use case The current AI functions are all value-producing and row-wise. There is no way to express an LLM decision as a native SQL predicate. Users who want to filter rows by a natural-language condition must wrap `aiClassify` in a comparison, which is awkward and does not compose in joins. ### Describe the solution you'd like A scalar function returning `UInt8` so it drops directly into `WHERE`, `PREWHERE`, and `JOIN ... ON`: ```sql aiFilter(text, condition[, temperature]) -- returns UInt8 ``` Example (semantic filter and semantic join): ```sql SELECT * FROM reviews WHERE aiFilter(body, 'the customer is angry about shipping'); SELECT p.id, r.id FROM products p JOIN reviews r ON aiFilter(concat(p.description,' || ',r.feedback), 'product fits customer needs'); ``` Should reuse the existing server-side provider configuration, quota, retry, and ai_function_throw_on_error infrastructure. ### Describe alternatives you've considered Wrapping `aiClassify(..., ['yes','no']) = 'yes'`. This works but reads poorly, does not express intent in joins, and prevents the planner from treating the call as a predicate. ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/110352",
        "timestamp": "2026-08-12T22:51:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "feature"
        ],
        "author": "ayakovlev-clickhouse",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:110893",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "RWLockImpl::getLock re-entrant deadlock in CREATE OR REPLACE internal DROP (STID 2043-3c5c)",
        "text": "## Summary `Logical error: RWLockImpl::getLock(): Cannot acquire exclusive lock while RWLock is already locked` fires from the internal cleanup DROP inside `CREATE OR REPLACE TABLE`. This is a fresh manifestation of the re-entrant DDL-lock LOGICAL_ERROR class previously seen in #79413 / #47023. Surfaced by CI: `Stress test (arm_debug)`, STID 2043-3c5c. Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110620&sha=ee4b5b869d5d3656964166ce79cdc044545bf273&name_0=PR&name_1=Stress%20test%20%28arm_debug%29 CIDB (30d): 1 occurrence, 0 on master. Rare, timing-dependent (needs the stress thread-fuzzer scheduling). ## Stack ``` RWLock.cpp:145 (LOGICAL_ERROR) <- IStorage::lockExclusively (IStorage.cpp:117) <- InterpreterDropQuery::executeToTableImpl (InterpreterDropQuery.cpp:326/354) <- InterpreterCreateQuery::doCreateOrReplaceTable (InterpreterCreateQuery.cpp:2502 / 2525) ``` ## Root cause `doCreateOrReplaceTable` runs its internal cleanup DROP (the post-EXCHANGE drop at InterpreterCreateQuery.cpp:2502 and the catch-block temp-cleanup drop at :2525) on a context built by `make_drop_context` = `Context::createCopy(current_context)`, which inherits the OUTER statement's `query_id`. `RWLockImpl::getLock` keys its fast-path on `query_id` (RWLock.cpp:130-146). When the same `query_id` already owns `IStorage::drop_lock` and a `Write` acquisition is requested, it throws `Cannot acquire exclusive lock while RWLock is already locked` (read->write upgrade / double-write is unsupported). Note that both `lockForShare` (Read) and `lockExclusively` (Write) operate on the SAME `drop_lock` (IStorage.cpp:79 and :115). So the inner DROP re-enters a lock that the outer `CREATE OR REPLACE` statement still holds under the shared `query_id` (e.g. a share-lock on the source/old-target storage held by an in-flight `AS SELECT` fill pipeline that has not yet released when the internal DROP requests `Write`). ## Relation to prior fixes Same LOGICAL_ERROR class as #79413 / #47023. #86751 (merged, Closes #79413) fixed the sibling `PARALLEL WITH` surface by giving each parallel branch a distinct `query_id`. The `CREATE OR REPLACE` internal-DROP path was not covered and still inherits the outer `query_id` via `createCopy`. ## Candidate fix (direction, mirrors #86751) Give the internal DROP context a distinct `query_id` (or an empty one) in the `make_drop_context` lambda (InterpreterCreateQuery.cpp:2352). An empty `query_id` sets `request_has_query_id = false`, which bypasses the RWLock fast-path entirely, so the internal DROP's exclusive acquisition can never re-enter a lock held under the outer statement's `query_id`. The pre-swap `force_drop`/size-guard semantics are unaffected (the drop context still sets `max_table_size_to_drop`/`max_partition_size_to_drop` when bypassing). ## Reproduction status Not yet deterministic locally. On a debug build with the stress thread-fuzzer env I ran: concurrent same-table `CREATE OR REPLACE` (mixed engines), forced catch-path drops via `max_table_size_to_drop=1`, forced fill-throw via `max_memory_usage`, and self-referential `CREATE OR REPLACE t AS SELECT FROM t` (all under 8-12 concurrent workers). All exercised the internal-DROP path (confirmed via error codes 359/241) but none hit the re-entrant window. The holders destruct synchronously on unwind before the internal DROP, so a foreground load does not land in the exact scheduling window; the CI stress thread-fuzzer does. Per our no-speculative-fix policy the fix will be validated against a failing reproducer (fails without, passes with) before a PR; this issue tracks the analysis and candidate fix in the meantime. <!-- ch-version-info:start --> ### Version info - Resolved by: #114420 - Merged into: `26.8.1.1308` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/110893",
        "createdAt": "2026-07-17T16:14:04Z",
        "updatedAt": "2026-08-13T07:35:51Z",
        "timestamp": "2026-08-13T07:35:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [],
        "author": "groeneai",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:110933",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "convertMySQLDataType wraps Array/Variant-based geo types in Nullable, throwing ILLEGAL_TYPE_OF_ARGUMENT on nullable MySQL spatial columns",
        "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ ### Describe what's wrong Reading a MySQL table that has a NULLABLE spatial column (`LINESTRING`/`POLYGON`/`MULTILINESTRING`/`MULTIPOLYGON`/generic `GEOMETRY`) via the `mysql()` table function, the `MySQL` table engine, or the `MySQL` database engine throws `Nested type LineString cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT)` at schema-inference time, so `CREATE`/`DESCRIBE`/`SELECT` fail entirely. Nullable is the MySQL default for a spatial column, and the `geometry` mapping is enabled by default, so this triggers out of the box. **Root cause:** src/DataTypes/convertMySQLDataType.cpp:171-172 applies `std::make_shared<DataTypeNullable>(res)` to `res` regardless of whether `res` can be inside Nullable. Before this PR the non-Point spatial types produced `String` (canBeInsideNullable=true), so `Nullable(String)` was valid; the new geo types (`Array`/`Variant` based) cannot be inside `Nullable`. **Why we believe this is a bug:** fetchTablesColumnsList (src/Databases/MySQL/FetchTablesColumnsList.cpp:116-119) calls convertMySQLDataType with is_nullable = external_table_functions_use_nulls (default true) && IS_NULLABLE='YES' -> convertMySQLDataType maps the spatial type to a geo type (e.g. `LineString` = `Array(Point)`) at src/DataTypes/convertMySQLDataType.cpp:143 -> then unconditionally wraps it at line 172 `res = std::make_shared<DataTypeNullable>(res)` -> DataTypeNullable ctor throws because DataTypeArray/DataTypeVariant::canBeInsideNullable() == false. **Affected locations:** - `src/DataTypes/convertMySQLDataType.cpp:172` — unconditional Nullable wrap of a geo-typed res - `src/DataTypes/convertMySQLDataType.cpp:143` — linestring -> LineString (Array(Point)) branch; same for polygon/multilinestring/multipolygon/geometry at 147/151/155/163 - `src/Databases/MySQL/FetchTablesColumnsList.cpp:119` — passes is_nullable=true for nullable MySQL columns when external_table_functions_use_nulls (default true) **Impact:** By default (geometry mapping on), any MySQL integration over a table with a nullable spatial column becomes unusable - schema inference throws and the table cannot be created/described/read. This is a regression: those columns previously mapped to `Nullable(String)` and worked. Affects the `mysql()` table function, `MySQL` table engine, and `MySQL` database engine. ## Assumptions The bot recorded these claims it could not verify directly from source. A maintainer ✅ confirms; ❌ flags a wrong premise (the bot should rework or close the finding). - [ ] **A plain MySQL spatial column (declared without NOT NULL) is reported as IS_NULLABLE='YES' in information_schema.columns** - *Why unverifiable:* the test harness could not connect to the MySQL container in this sandbox - *Falsifiable test:* CREATE TABLE t (g LINESTRING) in MySQL, then read information_schema.columns.IS_NULLABLE for column g (expected 'YES'); then DESCRIBE mysql(...) from ClickHouse (expected exception). ### Does it reproduce on most recent release? Likely yes — see testability note in additional context. ### How to reproduce ```python Run: pytest tests/integration/test_storage_mysql/test.py -k test_mysql_nullable_geometry -v -s (with a reachable mysql80 container). OR reproduce the invariant directly on any server: SELECT CAST(NULL AS Nullable(LineString)); -- throws ILLEGAL_TYPE_OF_ARGUMENT, whereas SELECT CAST(NULL AS Nullable(String)); succeeds. ``` ### Expected behavior ``` DESCRIBE returns rows including 'ls\\tLineString' and 'geo\\tGeometry' (schema inference succeeds). ``` ### Error message and/or stacktrace ``` Could not run end-to-end (MySQL container unreachable in sandbox). Type-invariant confirmed via SQL: `CREATE TABLE t (x Nullable(LineString)) ENGINE=Memory` -> Code: 43. DB::Exception: Nested type LineString cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT); same for Polygon/MultiLineString/MultiPolygon/Geometry. ``` ### Additional context **Open risks:** - The umbrella `Geometry` (Variant) also cannot be inside Nullable - same failure for a nullable generic GEOMETRY column. - The MySQL database engine (DatabaseMySQL) shares fetchTablesColumnsList and is affected identically. **Suggested fix:** When the mapped type cannot be inside Nullable (e.g. `!res->canBeInsideNullable()`), either skip the Nullable wrap for geo types (geo columns are inherently non-Nullable in ClickHouse) or fall back to `Nullable(String)` for nullable spatial columns. **Analysis details:** Confidence HIGH | Severity P1 | Testability: `INTEGRATION_TEST` Found during automated review of [PR #108944](https://github.com/ClickHouse/ClickHouse/pull/108944). ### CI Proof Bug confirmed by `Integration tests (amd_llvm_coverage, 2/8)` on [proof PR #110924](https://github.com/ClickHouse/ClickHouse/pull/110924). <details><summary>Reproducer test</summary> ```sql def test_mysql_nullable_geometry(started_cluster): # A nullable MySQL spatial column (MySQL's default when NOT NULL is omitted) must not # break schema inference. With the `geometry` mapping on by default, convertMySQLDataType # maps `linestring` to `LineString` (Array(Point)) and then wraps it in Nullable, which # throws ILLEGAL_TYPE_OF_ARGUMENT because Array/Variant-based geo types cannot be inside Nullable. table_name = \"test_mysql_nullable_geometry\" node1.query(f\"DROP TABLE IF EXISTS {table_name}\") conn = get_mysql_conn(started_cluster, cluster.mysql8_ip) drop_mysql_table(conn, table_name) with conn.cursor() as cursor: cursor.execute( f\"\"\" CREATE TABLE `clickhouse`.`{table_name}` ( `id` int NOT NULL, `ls` linestring, `geo` geometry, PRIMARY KEY (`id`)) ENGINE=InnoDB; \"\"\" ) conn.commit() # Expected (correct behavior): DESCRIBE succeeds and reports the geo types. # Actual (bug): raises 'Nested type LineString cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT)'. result = node1.query( f\"DESCRIBE mysql('mysql80:3306', 'clickhouse', '{table_name}', 'root', '{mysql_pass}')\" ) assert \"LineString\" in result, result drop_mysql_table(conn, table_name) conn.close() ``` </details> <details><summary>CI log excerpt</summary> ``` 2026-07-18T10:37:18.7126776Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_TESTS_SERVER_BIN_PATH to /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/clickhouse 2026-07-18T10:37:18.7127449Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_BINARY to /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/clickhouse 2026-07-18T10:37:18.7128117Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_TESTS_CLIENT_BIN_PATH to /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/clickhouse 2026-07-18T10:37:18.7128634Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_USE_OLD_ANALYZER to 0 2026-07-18T10:37:18.7129004Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_USE_DISTRIBUTED_PLAN to 0 2026-07-18T10:37:18.7129369Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_USE_DATABASE_DISK to 0 2026-07-18T10:37:18.7129977Z [2026-07-18 10:37:18] Setting environment variable PYTEST_CLEANUP_CONTAINERS to 1 2026-07-18T10:37:18.7130398Z [2026-07-18 10:37:18] Setting environment variable JAVA_PATH to /usr/lib/jvm/java-11-openjdk-amd64/bin/java 2026-07-18T10:37:18.7130811Z [2026-07-18 10:37:18] Setting environment variable COMPLIANCE_RESULT_FILE to 2026-07-18T10:37:18.7131220Z /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/promql_compliance_result.json 2026-07-18T10:37:18.7131649Z [2026-07-18 10:37:18] Setting environment variable LLVM_PROFILE_FILE to it-%4m.profraw 2026-07-18T10:37:18.7132392Z [2026-07-18 10:37:18] Run command: [pytest test_storage_mysql/test.py::test_mysql_nullable_geometry --report-log-exclude-logs-on-passed-tests --tb=short -n 1 2026-07-18T10:37:18.7133099Z --dist=loadfile --session-timeout=1200 --report-log=/home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/pytest_retries.jsonl 2026-07-18T10:37:18.7133681Z --log-file=/home/ubuntu/actions-runner/_work/ClickHouse/Cli ``` </details> --- _ClickGapAI · Confidence: HIGH · Severity: P1 · Finding: `h_pr108944_001`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #110943 - Merged into: `26.8.1.1302` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/110933",
        "createdAt": "2026-07-18T11:04:41Z",
        "updatedAt": "2026-08-13T01:36:30Z",
        "timestamp": "2026-08-13T01:36:30Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "comp-mysql"
        ],
        "author": "clickgapai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:111206",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Automatic parallel replicas (parallel_replicas_min_number_of_rows_per_replica > 0) fails RIGHT JOIN with NOT_FOUND_COLUMN_IN_BLOCK",
        "text": "**Describe what's wrong** With automatic parallel replicas — `parallel_replicas_min_number_of_rows_per_replica > 0` — a valid `RIGHT JOIN` that selects a left-table column fails with `10 NOT_FOUND_COLUMN_IN_BLOCK`: the left table's column is resolved against the **right** table. The same query works with plain parallel replicas (`parallel_replicas_min_number_of_rows_per_replica = 0`) and without parallel replicas. **Does it reproduce on the most recent release?** Reproduces on `26.7.1.408` and near-HEAD master `3090a4fc` (`26.7.1.653`). **How to reproduce** Any cluster works (below: the stateless-test 3-replica localhost cluster; `parallel_replicas_for_non_replicated_merge_tree` only because the tables are non-replicated). ```sql CREATE TABLE tl (k Int32, a Int32) ENGINE = MergeTree ORDER BY k; CREATE TABLE tr (k Int32, ver Int32) ENGINE = MergeTree ORDER BY k; INSERT INTO tl SELECT number, number FROM numbers(1000); INSERT INTO tr SELECT number, number FROM numbers(500); -- OK: 500 rows SELECT r.ver, (l.a + 2) FROM tl AS l RIGHT JOIN tr AS r USING (k) SETTINGS enable_analyzer = 1; -- Code: 10. DB::Exception: Column `a` not found in table default.tr SELECT r.ver, (l.a + 2) FROM tl AS l RIGHT JOIN tr AS r USING (k) SETTINGS enable_analyzer = 1, enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_min_number_of_rows_per_replica = 100; ``` **Characterization** - Only `RIGHT JOIN` fails; `LEFT JOIN` and `INNER JOIN` are correct under the same settings. - `ON l.k = r.k` form fails the same way as `USING (k)`; an `ON` clause with an added one-sided condition (`ON l.id = r.id AND l.v != 64`) fails identically (`Column ts not found in table ...` for whichever left column the projection references). - `parallel_replicas_min_number_of_rows_per_replica = 0` (non-automatic mode): correct. - Deterministic (3/3 runs). The row-estimation path of automatic mode appears to build/propagate a header for the wrong side of a `RIGHT JOIN`. Found by the optimizer-tester topology differential (`tests/optimizer_tester`, branch `optimizer-tester-framework`) the first time it varied the `parallel_replicas_*` sub-settings — the fixed configuration used before could not reach automatic mode. Related: https://github.com/ClickHouse/ClickHouse/commit/4531517ba0a5",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/111206",
        "createdAt": "2026-07-21T10:31:18Z",
        "updatedAt": "2026-08-13T13:08:22Z",
        "timestamp": "2026-08-13T13:08:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "bug",
          "comp-joins",
          "comp-parallel-replicas"
        ],
        "author": "zlareb1",
        "state": "closed",
        "assignees": [
          "devcrafter"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:111211",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "ORDER BY is not applied globally when reading a Distributed table through a Merge engine: per-shard-sorted streams are concatenated, LIMIT returns wrong rows",
        "text": "**Describe what's wrong** Reading a `Distributed` table through a `Merge` engine drops the global sort: `SELECT ... ORDER BY x` returns each shard's stream sorted individually and concatenated, not a globally ordered result. With `LIMIT n` this silently returns the wrong rows (the first shard's top-n instead of the global top-n). Querying the same `Distributed` table directly is correct. Aggregations over the `Merge` table (`count`, `max`) are correct — only the sorted path is wrong. **Does it reproduce on the most recent release?** Reproduces on `26.7.1.408` and near-HEAD master `3090a4fc` (`26.7.1.653`). **How to reproduce** Cluster `two_shards`: two shards on one server with `default_database` `sh0` / `sh1`. ```sql CREATE DATABASE sh0; CREATE DATABASE sh1; CREATE TABLE sh0.t (w Int64) ENGINE = MergeTree ORDER BY w; CREATE TABLE sh1.t (w Int64) ENGINE = MergeTree ORDER BY w; INSERT INTO sh0.t SELECT number * 3 FROM numbers(100); -- max 297 INSERT INTO sh1.t SELECT number * 3 + 1000 FROM numbers(100); -- max 1297 CREATE TABLE dist_t AS sh0.t ENGINE = Distributed(two_shards, '', t); CREATE TABLE merge_t AS dist_t ENGINE = Merge(currentDatabase(), '^dist_t$'); -- correct: 1297, 1294, 1291 SELECT w FROM dist_t ORDER BY w DESC LIMIT 3 SETTINGS enable_analyzer = 1; -- WRONG: 297, 294, 291 (shard 0's top-3; global maximum 1297 not returned) SELECT w FROM merge_t ORDER BY w DESC LIMIT 3 SETTINGS enable_analyzer = 1; -- also wrong WITHOUT LIMIT: 297, 294, 291, 288, ... (per-shard sorted streams -- concatenated; the result multiset is complete but the order is not global) SELECT w FROM merge_t ORDER BY w DESC SETTINGS enable_analyzer = 1; ``` **Characterization** - `max_threads = 1` (no thread-interleave explanation); deterministic. - Independent of `optimize_read_in_order` and `optimize_skip_unused_shards`. - `count()` / `max()` over `merge_t` are correct (500 rows / global max) — the row set is complete; only the final sort is missing on the `Merge`-over-`Distributed` read path, so any `ORDER BY ... LIMIT` silently returns wrong rows. Found by the optimizer-tester topology differential (`tests/optimizer_tester`, branch `optimizer-tester-framework`) the first run after a `Merge`-over-`Distributed` arm was added; detected by the order-sensitive comparison for total `ORDER BY` queries. Related: https://github.com/ClickHouse/ClickHouse/issues/111206",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/111211",
        "createdAt": "2026-07-21T10:51:36Z",
        "updatedAt": "2026-08-13T13:29:45Z",
        "timestamp": "2026-08-13T13:29:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "bug",
          "comp-distributed",
          "comp-storage-merge"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "Michicosun"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:111271",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Distributed ANY INNER JOIN returns one row per key per shard: deduplication is per-shard, not global",
        "text": "**Describe what's wrong** `ANY INNER JOIN` deduplicates per join key — but over a `Distributed` table it deduplicates on each shard independently, so a key whose rows span shards is returned once **per shard**. When the sharding key differs from the join key (the common case), the distributed result is a multiple of the correct local result. **Does it reproduce on the most recent release?** Reproduces on `26.7.1.408` and near-HEAD master `3090a4fc` (`26.7.1.653`). **How to reproduce** Cluster `two_shards`: two shards with `default_database` `sh0`/`sh1`. Sharding is by `k`; the join is on `a`, whose values appear on both shards. ```sql CREATE TABLE sh0.sha (k UInt32, a UInt32) ENGINE = MergeTree ORDER BY k; CREATE TABLE sh1.sha (k UInt32, a UInt32) ENGINE = MergeTree ORDER BY k; INSERT INTO sh0.sha SELECT number, number % 20 FROM numbers(0, 50); INSERT INTO sh1.sha SELECT number, number % 20 FROM numbers(50, 50); CREATE TABLE sha_d AS sh0.sha ENGINE = Distributed(two_shards, '', sha); CREATE TABLE shr (a UInt32) ENGINE = MergeTree ORDER BY a; INSERT INTO shr SELECT number FROM numbers(20); -- 20 (correct: one row per key) SELECT count() FROM (SELECT * FROM sh0.sha UNION ALL SELECT * FROM sh1.sha) AS l ANY INNER JOIN shr AS r ON l.a = r.a SETTINGS enable_analyzer = 1; -- 40 (WRONG: one row per key PER SHARD) SELECT count() FROM sha_d AS l ANY INNER JOIN shr AS r ON l.a = r.a SETTINGS enable_analyzer = 1, distributed_product_mode = 'global'; ``` **Characterization** - Deterministic. The count is exactly `#shards × correct` when every key spans all shards. - The single-node semantics (\"completely disables the cartesian product\" — one row per key) require the deduplication to happen globally (initiator-side or with `ANY`-awareness in the shard-merge), or the distributed rewrite of `ANY` strictness should be rejected/documented. Found by the optimizer-tester topology differential's genuinely-sharded arm (`tests/optimizer_tester`, branch `optimizer-tester-framework`) in a 10,000-query hunt. Related: https://github.com/ClickHouse/ClickHouse/issues/111195",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/111271",
        "createdAt": "2026-07-21T17:52:52Z",
        "updatedAt": "2026-08-13T12:13:18Z",
        "timestamp": "2026-08-13T12:13:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "bug",
          "comp-joins",
          "comp-distributed"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "vdimir"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:111838",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "`fsync_part_directory = 1` on an `encrypted` disk fails every `INSERT` with `FILE_DOESNT_EXIST`: the INSERT-path directory sync guard double-wraps the absolute part path",
        "text": "### Describe what's wrong Enabling `fsync_part_directory = 1` on a `MergeTree` table stored on an `encrypted` disk (with a non-empty `path`, as in the documented configuration) makes **every `INSERT` fail** with `FILE_DOESNT_EXIST` (Code 107). The durability hardening setting is therefore unusable on encrypted disks — users who want crash-safe part commits on encrypted storage have to turn it off. Root cause is a path double-wrap on the INSERT path only: - `MergeTreeDataWriter::writeTempPartImpl` requests the directory sync guard with the **absolute** part path (`new_data_part->getDataPartStorage().getFullPath()`, `src/Storages/MergeTree/MergeTreeDataWriter.cpp:855-858`). - On an encrypted disk this reaches `DiskEncryptedTransaction::wrappedPath` (`src/Disks/DiskEncryptedTransaction.h:36-42`), which unconditionally prepends `disk_path` again. With the docs' own layout (`<path>encrypted/</path>` over a local disk at `/disk/`) the guard tries to open `/disk/encrypted//disk/encrypted/store/<uuid>/tmp_insert_all_1_1_0/` → `ENOENT` → `LocalDirectorySyncGuard` constructor throws `FILE_DOESNT_EXIST` (`src/Disks/LocalDirectorySyncGuard.cpp:31-37`) → the INSERT fails. - The merge/mutate/fetch path is unaffected because `DataPartStorageOnDiskBase::getDirectorySyncGuard` passes the **relative** `root_path/part_dir` (`src/Storages/MergeTree/DataPartStorageOnDiskBase.cpp:956-958`) — so `OPTIMIZE` on the same table works while `INSERT` throws. - Plain `DiskLocal` (and an encrypted disk with an empty `path`) only work by accident: `std::filesystem::path::operator/` with an absolute right-hand side discards the left-hand side. ### Does it reproduce on the most recent release? Yes — master (26.7.1.1380) and 26.3.2.3 (longstanding). ### How to reproduce * Which ClickHouse server version to use: master; also reproduced on 26.3.2.3. Storage configuration (as in `docs/en/operations/storing-data.md`, encrypted disk with non-empty `path`): ```xml <storage_configuration> <disks> <local_disk><type>local</type><path>/data/local_disk/</path></local_disk> <enc> <type>encrypted</type> <disk>local_disk</disk> <path>enc/</path> <key>0123456789abcdef</key> </enc> </disks> <policies> <enc_policy><volumes><main><disk>enc</disk></main></volumes></enc_policy> </policies> </storage_configuration> ``` ```sql CREATE TABLE t_bug (x UInt64) ENGINE = MergeTree ORDER BY x SETTINGS storage_policy = 'enc_policy', fsync_part_directory = 1; INSERT INTO t_bug VALUES (1); -- Code: 107, FILE_DOESNT_EXIST — every time ``` Controls (all pass): - same table with `fsync_part_directory = 0` → INSERT OK; - same `fsync_part_directory = 1` on the underlying plain `local` disk → INSERT OK (and issues the directory `fdatasync`); - merge path on the encrypted disk: insert with the setting off, `ALTER TABLE ... MODIFY SETTING fsync_part_directory = 1`, `OPTIMIZE TABLE ... FINAL` → OK (relative-path caller). 3/3 on master, 1/1 on 26.3.2.3. ### Expected behavior `INSERT` with `fsync_part_directory = 1` on an encrypted disk succeeds and fsyncs the part directory of the underlying storage, same as on a plain local disk. The INSERT-path caller should pass the disk-relative part path (as the merge path does), or `wrappedPath` should not double-prepend an already-absolute path. ### Error message and/or stacktrace ``` Received exception from server (version 26.7.1): Code: 107. DB::Exception: Received from localhost:9000. DB::Exception: Cannot open file /data/local_disk/enc//data/local_disk/enc/store/b6a/b6a09e61-c268-4145-a690-ae9ae563557e/tmp_insert_all_1_1_0/: , errno: 2, strerror: No such file or directory: While executing WaitForAsyncInsert. (FILE_DOESNT_EXIST) ``` Note the double-wrapped path: `/data/local_disk/enc/` + `/data/local_disk/enc/store/...`. ### Additional context - File-data fsync on encrypted disks is intact (`WriteBufferFromEncryptedFile::sync` forwards down to `fdatasync`, `src/IO/WriteBufferFromEncryptedFile.cpp:44-50`) — only the directory-sync guard on the INSERT path is broken, and loudly. - No integration test covers fsync settings on encrypted disks (`tests/integration/test_encrypted_disk*` contain no `fsync` mentions). - Found by a crash-durability testing framework while auditing whether `DiskEncrypted` drops fsync (it does not — but the directory-fsync knob turns out to be unusable).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/111838",
        "createdAt": "2026-07-24T17:55:12Z",
        "updatedAt": "2026-08-13T12:40:00Z",
        "timestamp": "2026-08-13T12:40:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "bug",
          "comp-mergetree",
          "comp-disk-abstractions"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "CheSema"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:111879",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "pointInPolygon: throwing CAST(tuple, 'Point') hoisted into PREWHERE ahead of its NULL guard (CANNOT_INSERT_NULL), regression from #107988",
        "text": "**Describe what's wrong** A `WHERE` predicate that casts a tuple to `Point` (a non-`Nullable` type) throws `CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN` at execution, even though a preceding filter guarantees the cast never sees a NULL. The throwing `CAST(tuple(...), 'Point')` is materialized in a PREWHERE read step **before** the guarding filter is applied, so it runs on rows the guard would have removed. The query returns a result on 26.2 and throws on 26.6. ``` Code: 349. DB::Exception: Cannot convert NULL value to non-Nullable type: while executing 'FUNCTION CAST(tuple(__table1.longitude, __table1.latitude) :: 7, 'Point'_String :: 9) -> CAST(tuple(__table1.longitude, __table1.latitude), 'Point'_String) Point : 4': While executing MergeTreeSelect(pool: PrefetchedReadPool, algorithm: Thread). (CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN) ``` **Does it reproduce on the most recent release?** Yes, reproduces on 26.6. Good on 26.2. **How to reproduce** ```sql DROP TABLE IF EXISTS repro_pip SYNC; CREATE TABLE repro_pip ( id Int64, session_id Int64, payload JSON, filler String ) ENGINE = MergeTree ORDER BY id; INSERT INTO repro_pip SELECT number, number, if(number % 3 = 0, '{}', '{\"longitude\":4.35,\"latitude\":52.06}')::JSON, -- 1/3 of rows have no coords repeat('x', 200) FROM numbers(500000); -- Throws Code 349 on 26.6; returns 333333 on 26.2. SELECT count(DISTINCT session_id) FROM ( SELECT session_id, CAST(payload.latitude AS Nullable(Float64)) AS latitude, CAST(payload.longitude AS Nullable(Float64)) AS longitude FROM repro_pip WHERE payload.longitude IS NOT NULL -- guards the cast; latitude itself is never guarded ) a WHERE 1=1 -- constant-true term is the trigger AND pointInPolygon((longitude, latitude)::Point, readWKTPolygon('POLYGON ((4.3 52.0,4.4 52.0,4.4 52.1,4.3 52.1,4.3 52.0))')) = 1; ``` Essential ingredients (established by minimization; each is required): - A `JSON` column read via subcolumns (`payload.longitude` / `payload.latitude`). Plain `Nullable(Float64)` columns do **not** reproduce it — the dynamic-subcolumn read is what makes the plan materialize the cast in a PREWHERE read step. - `latitude` / `longitude` NULL together on some rows. - A subquery/CTE guard `payload.longitude IS NOT NULL` (latitude itself is never guarded). - A wide row (`filler`) so move-to-prewhere engages. - The outer `WHERE 1=1 AND pointInPolygon(...)`. The constant-true `1=1` is the trigger: it leaves a residual `Filter` in addition to the PREWHERE, and the throwing cast ends up materialized in a PREWHERE read step evaluated over the whole granule, before the guard filters the NULL rows. **Expected behavior** The query should not throw: the `payload.longitude IS NOT NULL` guard removes the NULL rows before the cast, as it did on 26.2 (returns `333333`). A potentially-throwing expression must not be evaluated in a read step that precedes the filter guarding it. Notably, adding `AND payload.latitude IS NOT NULL` to the subquery does **not** help — the cast is hoisted ahead of that guard too. So this cannot be worked around with NULL guards. **Bisect** First bad commit: `74a2218e61bf3bee47bad4851817d372bbfe4ee1`, the merge of #107988 (\"Fix performance regression for Map subcolumns with PREWHERE\"), merged to master 2026-06-24. Bisected on `origin/master` with the repro above (good `38ac1a2` 26.2-era → bad `bb9c5e5` post-26.6). The changed files match the mechanism: `MergeTreeSplitPrewhereIntoReadSteps.cpp`, `MergeTreeWhereOptimizer.cpp`, `ReadFromMergeTree.cpp`, `MergeTreeSelectProcessor.*`, and subcolumn serialization (`SerializationSparse` / `SerializationMapKeyValue` / `ISerialization`). The PR fixed a Map-subcolumn PREWHERE *performance* regression; its PREWHERE-split changes introduced this *correctness* regression. Caused by: https://github.com/ClickHouse/ClickHouse/pull/107988 **Workarounds** (both return `333333` on 26.6) - `SETTINGS optimize_move_to_prewhere = 0` (also: `allow_reorder_prewhere_conditions = 0`, `query_plan_merge_filters = 0`). - Make the cast NULL-safe: `pointInPolygon((ifNull(longitude, 0.), ifNull(latitude, 0.))::Point, readWKTPolygon('POLYGON ((...))')) = 1`.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/111879",
        "createdAt": "2026-07-25T07:49:13Z",
        "updatedAt": "2026-08-13T12:47:15Z",
        "timestamp": "2026-08-13T12:47:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug"
        ],
        "author": "fm4v",
        "state": "open",
        "assignees": [
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:111898",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "DPsub join-order reordering with `query_plan_enable_optimizations = 0` silently drops a non-equi `JOIN ON` conjunct on chained joins (post-#109638)",
        "text": "## Describe what's wrong With `query_plan_optimize_join_order_algorithm = 'dpsub'` and `query_plan_enable_optimizations = 0`, a chained join whose first `ON` clause carries a non-equi conjunct returns wrong results: the DPsub reordering still runs (it restructures the join tree even though plan optimizations are disabled), but the residual, non-equi part of the `ON` condition is silently dropped — rows that fail it come back as matched. This is the surviving corner of #109617: the fix (#109638, merged 2026-07-07) repairs the default path, but with `query_plan_enable_optimizations = 0` the same query still returns the pre-fix wrong result on current master. Settings must not change results — either the reordering should not run at all when plan optimizations are disabled, or it must preserve the residual predicate. ## How to reproduce Masters `26.8.1.102` (`eb577ff401a3`, 2026-07-24) and `26.7.1.1169` (`38334977be1e`, 2026-07-18), official builds (reduced from `tests/queries/0_stateless/02372_analyzer_join`, same data as #109617): ```sql CREATE TABLE t1 (id UInt64, value String) ENGINE = MergeTree ORDER BY tuple(); CREATE TABLE t2 (id UInt64, value String) ENGINE = MergeTree ORDER BY tuple(); CREATE TABLE t3 (id UInt64, value String) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t1 VALUES (0, 'Join_1_Value_0'), (1, 'Join_1_Value_1'), (2, 'Join_1_Value_2'); INSERT INTO t2 VALUES (0, 'Join_2_Value_0'), (1, 'Join_2_Value_1'), (3, 'Join_2_Value_3'); INSERT INTO t3 VALUES (0, 'Join_3_Value_0'), (1, 'Join_3_Value_1'), (4, 'Join_3_Value_4'); SELECT t1.id, t1.value, t2.id, t2.value, t3.id, t3.value FROM t1 INNER JOIN t2 ON t1.id = t2.id AND t1.value = 'Join_1_Value_0' LEFT JOIN t3 ON t2.id = t3.id ORDER BY ALL SETTINGS query_plan_optimize_join_order_algorithm = 'dpsub', query_plan_enable_optimizations = 0; ``` Observed (both versions, deterministic 5/5): ``` 0 Join_1_Value_0 0 Join_2_Value_0 0 Join_3_Value_0 1 Join_1_Value_1 1 Join_2_Value_1 1 Join_3_Value_1 ``` Expected (and returned at default settings): only the first row — for `id = 1`, `t1.value = 'Join_1_Value_0'` is false, so the `INNER JOIN` must not match. The second row is the `ON` conjunct being ignored. The trigger surface: - only `dpsub` misbehaves; `greedy` and `dpsize` return the correct result with `query_plan_enable_optimizations = 0`; - either co-factor alone is fine: `dpsub` with optimizations on is correct (the #109638 fix), and `query_plan_enable_optimizations = 0` with the default algorithm is correct; - the chain is required — the two-table `t1 JOIN t2 ON t1.id = t2.id AND t1.value = '...'` alone is correct; - a `RIGHT JOIN` first hop misbehaves the same way (matched rows appear where the unmatched-side defaults are expected); - `EXPLAIN` shows the reordering did run despite optimizations being disabled: the plan becomes `t1 ⋈ (t2 ⋈ t3)` and no filter step carries `t1.value = 'Join_1_Value_0'` anywhere. ## Expected behavior The same rows as with default settings. `query_plan_optimize_join_order_algorithm` and `query_plan_enable_optimizations` are plan-shape settings and must never change query results. Workaround: `query_plan_enable_optimizations = 1` (default), or a different join-order algorithm. Related: https://github.com/ClickHouse/ClickHouse/issues/109617 Related: https://github.com/ClickHouse/ClickHouse/pull/109638 Found by an automated optimizer-correctness differential tester (unoptimized-vs-optimized `opt_toggle` oracle on a stateless-suite replay).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/111898",
        "createdAt": "2026-07-25T12:58:39Z",
        "updatedAt": "2026-08-13T10:58:43Z",
        "timestamp": "2026-08-13T10:58:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "bug",
          "comp-joins",
          "comp-query-optimizer",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "fkastrati"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112018",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "LOGICAL_ERROR on re-run INSERT … SELECT — sparse Tuple subcolumn breaks the insert-deduplication retry pipeline",
        "text": "When a Tuple column has an element stored with sparse serialization, re-running an INSERT … SELECT that is caught by block deduplication throws LOGICAL_ERROR (code 49) instead of being a silent no-op. Repro (tested in 26.4.1): ``` CREATE TABLE sparse_tuple_src ( id UInt64, body Tuple(key UInt64, flag Bool) ) ENGINE = MergeTree ORDER BY id; CREATE TABLE sparse_tuple_dst ( id UInt64, body Tuple(key UInt64, flag Bool) ) ENGINE = MergeTree ORDER BY id SETTINGS non_replicated_deduplication_window = 100; -- body.flag is the default (false) for all but the first 1000 of 2M rows, i.e. ~99.95% -- defaults. That is above ratio_of_defaults_for_sparse_serialization (default 0.9375), -- so the tuple ELEMENT gets sparse serialization. INSERT INTO sparse_tuple_src SELECT number, (number, number < 1000) FROM numbers(2000000); -- First insert: succeeds. INSERT INTO sparse_tuple_dst SELECT * FROM sparse_tuple_src SETTINGS insert_deduplication_token = 'repro'; -- Second, identical insert: blocks are duplicates, so the deduplication -- retry pipeline is built -> LOGICAL_ERROR. INSERT INTO sparse_tuple_dst SELECT * FROM sparse_tuple_src SETTINGS insert_deduplication_token = 'repro'; Code: 49. DB::Exception: Block structure mismatch in function connect between SourceFromSingleChunk and ConvertingTransform stream: different columns: body Tuple(key UInt64, flag Bool) Tuple(size = 0, UInt64(size = 0), Sparse(size = 0, UInt8(size = 1), UInt64(size = 0))) body Tuple(key UInt64, flag Bool) Tuple(size = 0, UInt64(size = 0), UInt8(size = 0)). (LOGICAL_ERROR) ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #111191 - Backported to: `26.7.3.19`, `26.6.3.7`, `26.5.7.21` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112018",
        "createdAt": "2026-07-27T06:17:23Z",
        "updatedAt": "2026-08-13T09:54:35Z",
        "timestamp": "2026-08-13T09:54:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "bug"
        ],
        "author": "romanbukarev-clh",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112241",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "A `WHERE` predicate on an `arrayJoin`'d column can be rewritten into `arrayFilter` so the array is filtered before expansion",
        "text": "### Company or project name ClickHouse customer ### Use case Queries that unnest one or more `Array` columns with `arrayJoin` (or `ARRAY JOIN`) and then select a few specific elements in `WHERE` are a very common pattern (tag/label arrays, event attribute arrays, key-value arrays stored as parallel arrays). Today ClickHouse first materializes the full expansion and only then applies the filter, so a query that ultimately returns one row per source row can transiently produce `length(A) * length(B) * length(C)` rows per source row. The work and the memory spent on those rows is unnecessary, and with several `arrayJoin` calls in one query it's multiplicative. Example: ```sql SELECT arrayJoin(A) AS a, arrayJoin(B) AS b, arrayJoin(C) AS c FROM t WHERE a = 'X-A' AND b = 'X-B' AND c = 'X-C'; ``` is equivalent to, but much slower and more memory hungry than, the manual rewrite: ```sql SELECT arrayJoin(arrayFilter(x -> x = 'X-A', A)) AS a, arrayJoin(arrayFilter(x -> x = 'X-B', B)) AS b, arrayJoin(arrayFilter(x -> x = 'X-C', C)) AS c FROM t; ``` ### Describe the solution you'd like Add an optimization that pushes a WHERE conjunct that constrains an arrayJoin'd column into an arrayFilter over the source array, i.e. rewrite `arrayJoin(expr) AS a ... WHERE f(a)` into `arrayJoin(arrayFilter(x -> f(x), expr)) AS a` and drop the conjunct from WHERE when it becomes redundant. Conditions under which the rewrite is applicable: 1. The predicate is a top-level conjunct of WHERE (an AND operand). OR across different array-join columns must not be rewritten. 2. The conjunct depends on exactly one arrayJoin'd column. Other columns it references must be row-constant (not produced by an arrayJoin); those are captured by the lambda, which arrayFilter already supports. A conjunct relating two different arrayJoin'd columns (e.g. a = b) cannot be rewritten. 3. The predicate must be deterministic and stateless (exclude rand, runningAccumulate, neighbor, and anything else whose result depends on evaluation order or position in the stream). 4. Multi-array `ARRAY JOIN a, b` is not supported ### Describe alternatives you've considered Writing `arrayJoin(arrayFilter(...))` by hand. This works and is the current workaround. ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112241",
        "createdAt": "2026-07-28T08:40:27Z",
        "updatedAt": "2026-08-13T11:45:58Z",
        "timestamp": "2026-08-13T11:45:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "feature"
        ],
        "author": "ayakovlev-clickhouse",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112376",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "KILL QUERY requires SELECT on system.processes (and it should not)",
        "text": "### Company or project name ClickHouse Cloud support on behalf of a customer ### Describe what's wrong `KILL QUERY` requires the user to have `SELECT` privilege on `system.processes` table in the Cloud. In specific use cases (e.g. using a BI tool like Metabase), it can lead to a security hole where a user would be forced to open all records if that table to a used needing to kill a query -- not just their own. Creation of a ROW POLICY, however, is not possible (revoked by the Cloud for system.*). The ask is to make it possible to run KILL QUERY (by query_id) without having SELECT access on system.processes. ### Does it reproduce on the most recent release? Yes ### How to reproduce In the Cloud, create a user without SELECT access to system.processes As that user, execute KILL QUERY: ``` me@myLaptop ch_client % clickhouse client --host foo.bar.aws.clickhouse.cloud --secure --user me_test --password '*****' \\ --query \"KILL QUERY WHERE query_id='hello_kitty'\" Received exception from server (version 26.4.1): Code: 497. DB::Exception: Received from foo.bar.aws.clickhouse.cloud:9440. DB::Exception: me_test: Not enough privileges. To execute this query, it's necessary to have the grant SELECT ON system.processes. (ACCESS_DENIED) (query: KILL QUERY WHERE query_id='hello_kitty') ``` ### Expected behavior An existing query should be deleted if it belongs to the user. Otherwise (query doesn't exist, query doesn't belong to the user) a no-op. ### Error message and/or stacktrace see above ### Related issues and pull requests _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112376",
        "createdAt": "2026-07-28T23:30:15Z",
        "updatedAt": "2026-08-13T15:28:08Z",
        "timestamp": "2026-08-13T15:28:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "comp-rbac",
          "potential bug"
        ],
        "author": "romanbukarev-clh",
        "state": "open",
        "assignees": [
          "pufit"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112444",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "[RFC] Query Acceleration: derived structures for transparent query rewrite",
        "text": "## Summary ClickHouse needs a first-class way to maintain **derived, query-specific data structures** next to a source table and use them **transparently** during planning — without making those structures a second system of record, and without forcing applications to query a side table. **Query Acceleration** is that framework: 1. User creates an **accelerator** bound to a source table and a set of indexed columns. 2. Background work **materializes** derived data from active source parts, with bounded resources and full ops visibility. 3. On a normal `SELECT` against the **source**, the optimizer asks registered accelerators whether they can improve the plan. 4. If matched, the plan uses the derived structure to produce **candidates / read hints**, then gathers the final rows from the source. If not matched, the ordinary source plan remains correct. The first concrete family is an experimental **ANN (approximate nearest-neighbor) index** for vector top‑k / bounded range queries. This issue is about the **acceleration principle and framework contract**; family-specific algorithms and knobs belong in per-family docs and follow-up issues. --- ## Motivation ### Problem Many expensive query patterns share the same operational needs: - The **source table** must stay the durable, mutable dataset (insert / merge / mutate as today). - A **derived structure** (index, sketch, layout, partial aggregate, …) would make a known shape much cheaper. - Construction should be **asynchronous**, resource-bounded, and observable — not folded only into foreground inserts or ad-hoc ETL. - Freshness may **lag** source changes for a while; queries must still be **correct** under partial coverage. - Applications should keep writing SQL against the **source**, not against an internal index table. Today’s alternatives each miss part of that story: | Approach | Gap | |---|---| | Full scan / exact evaluation only | Correct but does not scale for selective ranked or specialized filters | | Native secondary indexes tied to part metadata | Limited independence of lifecycle, algorithm choice, and ops surface | | Manual side tables / external systems | Dual-write or ETL; no transparent rewrite; correctness and freshness are the user’s problem | ### Why a framework (not one-off indexes) If each specialized structure re-implements catalog binding, rebuild, plan matching, hybrid fallback, and telemetry, the product fragments. Query Acceleration factors the common contract so a new family only owns: - on-disk / in-memory format and build; - query-shape recognition and cost; - how to search the derived data and map hits back to source rows; - family-specific settings and metrics. Candidate future families (design directions, not commitments): text / sparse retrieval, geospatial, join layouts, aggregation state, domain-specific predicates. --- ## Goals 1. **Source is authoritative.** User DML and ordinary reads target the source. Accelerators are recomputable from the source; they are not independently backed up as a second truth. 2. **Transparent acceleration.** Create once; query the source with natural SQL. No application dual-path to an index table. 3. **Correct under lag.** Incomplete materialization must not omit required rows from the query contract. Uncovered regions use a family-defined **exact / residual** path (or decline the rewrite), not silent omission. 4. **Declarative selection.** Matchers offer or decline with stable reasons. The planner never blindly forces an accelerator. 5. **Isolation of concerns.** Framework: catalog, lifecycle hooks, matching plumbing, coverage snapshot semantics, SYSTEM commands, shared telemetry. Family: format, build, search, cost, algorithm settings. 6. **Operable.** Coverage, backlog, jobs, failures, and per-query outcomes are first-class (system tables + logs). 7. **Fail closed.** Capability mismatch, corrupt lineage/mapping, and execution failures raise errors. They are not silently reclassified as “uncovered” or swapped to another interface. ## Non-goals - Turning an accelerator into a general-purpose user-facing table (`SELECT`/`INSERT`/`ALTER` as primary API). - Guaranteeing that every source query shape is accelerated. - Making approximate families exact by default (approximation, when used, is an explicit family semantic). - Shipping many families in the first GA; the first family validates the path, not the entire roadmap. --- ## Principle: how query acceleration works ### Lifecycle of an accelerator ```text CREATE ACCELERATOR acc ON source (cols) ENGINE = <Family>(...) │ ▼ Background scheduler compare active source parts ↔ recorded lineage build / compact / publish derived parts │ ▼ Query on source optimizer extracts a plan shape matcher: match or decline (reason_code) │ matched│ declined ▼ ▼ Hybrid / hinted plan Ordinary source plan accelerated regions + residual for uncovered │ ▼ Gather payload from source by identity ``` **Users never query the accelerator as a data table.** Lifecycle is `CREATE` / `DROP` / `DETACH` / `ATTACH` / rename, plus `SYSTEM` controls for refresh, start/stop builds, and bounded sync for tests and rollout. ### Four layers | Layer | Responsibility | |---|---| | **Catalog & binding** | Accelerator is catalog-visible, bound to one source and indexed columns; dependency prevents accidental drop when checks are enabled | | **Materialization** | Async jobs reconciling source parts with derived parts; resource pools and per-source limits | | **Planning & execution** | Shape recognition → matcher offer → rewrite that uses derived data only where coverage and capability allow | | **Operations & telemetry** | Live state, job history, query decision log, coverage and fallback outcomes | ### Match kind today: `ReadHint` The first rewrite pattern is **candidate-first / read-hint**: 1. Derived structure returns a bounded set of **source-row identities** (and optional scores). 2. Optional global selection (e.g. top‑k or quota) merges accelerated and residual producers. 3. Pipeline **gathers** full row payloads from the source using those identities (hints), instead of scanning the whole selected range for ranking. Other match kinds (e.g. partial aggregate, join probe) are possible later when a family needs them; they should not be invented without a concrete consumer. ### Fail-closed matching Matchers decline with stable `reason_code`s (shape unsupported, metric/config mismatch, no ready parts, competing native index preferred, …). Optional strict settings can turn “would fall back to full source plan” into an exception for tests and benchmarks, so silent non-acceleration is visible. ### Observability contract Operators should answer four questions without reading code: 1. What accelerators exist, and what is their **coverage / backlog / health**? 2. What **background jobs** are running or failing? 3. Did this query **match**, and if not, **why**? 4. Was the outcome fully accelerated, partial residual, full fallback, or failed? Shared objects (names may gain family-specific companions): - `system.accelerators` - `system.accelerator_jobs` / `accelerator_job_log` - `system.accelerator_query_log` (final outcomes + stage/decision rows) --- ## Core abstractions (implementation anchors) | Concept | Role | |---|---| | `IAccelerator` | Catalog storage base: source binding, family/impl, observability snapshot; exposes matcher + scheduler; rejects direct user R/W by default | | `IAcceleratorMatcher` | Plan-time: plan shape → offer (`ReadHint`) or decline with `reason_code` | | `IAcceleratorScheduler` | Background entry: snapshot + schedule jobs | | Optimizer glue | Extract shape, enumerate matchers, inject hybrid / hinted plan when accepted | | Family package | Build, format, search, cost, residual evaluator, family system tables | Framework code must stay family-agnostic enough that a **second family** does not require forking catalog DDL, SYSTEM commands, or query-log infrastructure. --- ## First implementation (brief) The first family is **`ANNIndex` / `ReplicatedANNIndex`**: approximate nearest-neighbor indexes over a vector column on a MergeTree-family source. - **Purpose:** accelerate source queries of the form ranked distance top‑k and certain bounded distance ranges. - **User model:** create accelerator → query **source** → optional `SYSTEM SYNC ACCELERATOR` for tests/rollout. - **Gate:** experimental setting for create/attach. - **Validates the framework:** async part-level builds, hybrid covered/uncovered execution, match/decline reasons, job and query telemetry, replication of derived parts as a separate lifecycle from the source. Algorithm choice, index parameters, filtered-search modes, and ANN-only ops details are **out of scope for this epic issue**; see the ANN accelerator guide and family follow-ups. --- ## Known framework gaps | Area | Gap | Direction | |---|---|---| | Families | Only one family exists | Land checklist + second family to prove generality | | Match kinds | Only `ReadHint` | Add kinds only with a real consumer | | Multi-accelerator choice | Limited arbitration when several could match | Cost-based selection + deterministic tie-break | | Distributed / parallel replicas | Accelerated path not universally available | Explicit policy per execution mode; residual ordinary plan | | Strict shapes | Some query features force full source plan | Document; extend rewrite only when ownership of predicates is clear | | Coupling risk | First-family types can leak into optimizer entry points | Keep matcher/plan-shape APIs generic; confine family pipelines | | GA readiness | Experimental gates, format stability, runbooks | Soak first family; freeze public contracts | --- ## Roadmap ### Phase A — Framework contract solid with first family as reference 1. Freeze DDL / SYSTEM / shared system-table and log contracts. 2. Document and test hybrid coverage, residual semantics, and fail-closed matching. 3. Golden `EXPLAIN` / outcome codes for match, partial residual, and full fallback. 4. Reduce first-family leakage into shared optimizer types. 5. Performance of shared paths: materialization scheduling, candidate production, payload gather. ### Phase B — Production readiness of the framework + first family 1. Format / upgrade story for derived parts. 2. Narrow experimental gates after soak. 3. Runbooks: coverage lag, degraded state, build storms, fallback rate alerts. 4. Clear distributed and parallel-replicas policy. ### Phase C — Second family (proves the framework) Pick one high-ROI shape (e.g. geospatial, constrained join, or repeated aggregation). Acceptance: lands **without** forking catalog, SYSTEM, or shared query-log infrastructure; only family package + thin matcher/optimizer registration. --- ## One-line pitch > Query Acceleration maintains derived structures beside a MergeTree source table and rewrites matching source queries transparently, with asynchronous builds, hybrid residual evaluation under partial coverage, and first-class operational telemetry — starting with an experimental ANN family for vector search.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112444",
        "createdAt": "2026-07-29T14:26:22Z",
        "updatedAt": "2026-08-13T01:17:56Z",
        "timestamp": "2026-08-13T01:17:56Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [],
        "author": "fastio",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112569",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Request for office hours for new contributors.",
        "text": "### Company or project name N/A ### Question -- Is there any office hours for new contributors such that we can contribute to the project?",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112569",
        "createdAt": "2026-07-30T09:59:20Z",
        "updatedAt": "2026-08-13T11:46:42Z",
        "timestamp": "2026-08-13T11:46:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "question"
        ],
        "author": "harshil15999",
        "state": "closed",
        "assignees": [
          "antonkovalenko"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112586",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Range condition in JOIN ON against a single-row side is not used for index analysis (no part pruning) since the logical join step became default",
        "text": "### Company or project name ClickHouse Inc. Found while comparing two server versions on an internal test cluster, on a reporting workload that derives its date window from a joined single-row subquery. ### Describe the situation When a range predicate on a `MergeTree` primary/partition key column is expressed in the `ON` clause of a `JOIN` against a single-row constant side, the condition is no longer turned into a `KeyCondition`. Index analysis reports `Condition: true` and every part is read, instead of pruning. The logical join step turns the inequality `ON` into a `cross` join plus a `Filter (Post Join Actions)` sitting above the join, and that filter never reaches index analysis. The same predicate written as a plain `WHERE` (or as scalar subqueries in `WHERE`) prunes correctly, and results are identical either way, so this is purely lost part/granule pruning. This is a common shape for \"current week / current month\" reporting queries, where the bounds are computed once in a CTE and joined in: ```sql WITH bounds AS (SELECT toStartOfWeek(today(), 1) AS lo, today() AS hi) SELECT ... FROM big_table AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi ``` On a real workload (a ~365M row table partitioned by month) this was the difference between reading 0 parts and reading all 302 parts, and between **1.4s and 35-68s (~25x)** for a query whose correct answer was an empty result set. ### Which ClickHouse versions are affected? The pruning is lost whenever the logical join step is used, i.e. whenever `query_plan_use_new_logical_join_step = 1`. **The regression is 26.5.** That is where https://github.com/ClickHouse/ClickHouse/pull/104017 made the setting `Obsolete`, hardcoding it to `true`. Up to and including 26.4 it was a normal `Production` setting, so anyone hitting this could simply set it to `0` and get pruning back. From 26.5 on there is **no setting-level mitigation at all** and the only fix is to rewrite the query. For completeness on the earlier history: the declared default went `false` -> `true` in 25.2 (https://github.com/ClickHouse/ClickHouse/pull/74909), so a 25.2..26.4 user who never touched the setting also saw no pruning. But that was recoverable, and deployments that pinned `query_plan_use_new_logical_join_step = 0` (ClickHouse Cloud among them, via its `compatibility` profile) kept correct pruning all the way through 26.4 and only lost it on upgrade past 26.5. Verified with official builds: | Version | Default behavior | With `query_plan_use_new_logical_join_step = 0` | |---|---|---| | 26.3.1.876 | `Parts: 14/14` (no pruning) | `Parts: 1/14` (prunes) | | 26.7.1.2043 | `Parts: 14/14` | setting is `Obsolete`, no effect | | 26.8.1.48 (master, `d38cb514a590e60af07777d8f48c34afdf20aa1f`) | `Parts: 14/14` | setting is `Obsolete`, no effect | `SETTINGS compatibility = '26.4'` does **not** restore the pruning on master either. ### How to reproduce Works in `clickhouse local`, no special settings needed. ```sql CREATE TABLE repro_join_prune (d Date, v UInt32) ENGINE = MergeTree PARTITION BY toYYYYMM(d) ORDER BY d; INSERT INTO repro_join_prune SELECT toDate('2025-01-01') + number, number FROM numbers(400); -- (1) bounds arrive via JOIN ON: no pruning on 25.2+ EXPLAIN indexes = 1 WITH bounds AS (SELECT toDate('2025-06-01') AS lo, toDate('2025-06-10') AS hi) SELECT count() FROM repro_join_prune AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi; -- (2) same predicate as a plain WHERE: prunes correctly on every version EXPLAIN indexes = 1 SELECT count() FROM repro_join_prune AS t WHERE t.d >= toDate('2025-06-01') AND t.d <= toDate('2025-06-10'); -- (3) same predicate as scalar subqueries: also prunes correctly on every version EXPLAIN indexes = 1 SELECT count() FROM repro_join_prune AS t WHERE t.d >= (SELECT toDate('2025-06-01')) AND t.d <= (SELECT toDate('2025-06-10')); ``` All three queries return `10`, so results are correct in every case. Plan for query (1) on master. Note that the range predicate is present as a `Filter (Post Join Actions)` but index analysis still gets `Condition: true`: ``` Aggregating │ Keys: │ Aggregates: count() │ Skip merging: 0 └──Filter (Post Join Actions) │ Filter column: d >= '2025-06-01' AND d <= '2025-06-10' └──Join (JOIN FillRightFirst) │ t[400] ⋈ system.one[1] │ Type: cross | Strictness: all | Algorithm: HashJoin │ Result rows: 400 │ Output: │ Left: __join_result_dummy, d │ Right: hi, lo ├──ReadFromMergeTree (default.repro_join_prune) │ Read type: Default │ Parts: 14 | Granules: 14 │ Output: d │ Indexes: │ Min-Max │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ Partition │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ PrimaryKey │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ Ranges: 14 └──ReadFromSystemOne ``` Plan for query (2) on master, for comparison: ``` Aggregating │ Keys: │ Aggregates: count() │ Skip merging: 0 └──Filter ((WHERE + Change column names to column identifiers)) │ Filter column: d >= '2025-06-01' AND d <= '2025-06-10' └──ReadFromMergeTree (default.repro_join_prune) Read type: Default Parts: 1 | Granules: 1 Output: d Indexes: Min-Max Keys: d Condition: and((d in (-Inf, 20249]), (d in [20240, +Inf))) Parts: 1/14 Granules: 1/14 Partition Keys: toYYYYMM(d) Condition: and((toYYYYMM(d) in (-Inf, 202506]), (toYYYYMM(d) in [202506, +Inf))) Parts: 1/1 Granules: 1/1 PrimaryKey Keys: d Condition: and((d in (-Inf, 20249]), (d in [20240, +Inf))) Parts: 1/1 Granules: 1/1 Search Algorithm: binary search Ranges: 1 ``` To see the \"good\" plan for query (1), run it on 26.4 or earlier with `SETTINGS query_plan_use_new_logical_join_step = 0`. ### Expected performance Query (1) should prune the same way queries (2) and (3) do, since the join is against a single-row constant side and the `ON` conditions are a plain range over the sorting/partition key. It did prune before the logical join step became the default, and it still prunes on 26.4 and earlier with `query_plan_use_new_logical_join_step = 0`. Note that a `CROSS JOIN` with the same predicate moved into `WHERE` does **not** prune on any version tested, including 26.4 with the old join step. That looks like a separate, pre-existing limitation rather than part of this regression, but it may share a root cause. ### Related issues and pull requests Caused by: https://github.com/ClickHouse/ClickHouse/pull/104017 Related: https://github.com/ClickHouse/ClickHouse/pull/74909 ### Additional context The following settings were tried on 26.6+ and none of them restore the pruning: - `query_plan_merge_filter_into_join_condition = 1` - `query_plan_convert_join_to_in = 1` - `allow_general_join_planning = 1` - `query_plan_filter_push_down = 1` - `query_plan_convert_outer_join_to_inner_join = 1` - `query_plan_optimize_join_order_limit = 10` - `use_join_disjunctions_push_down = 1` - `compatibility = '26.4'` The only workaround we found is a query rewrite: duplicate the bounds as a plain `WHERE` on the scanned table. The `JOIN` can stay in place, so the rewrite is semantics-preserving, and it restored the real workload from 35-68s to 1.3s. ```sql WITH bounds AS (SELECT toDate('2025-06-01') AS lo, toDate('2025-06-10') AS hi) SELECT count() FROM repro_join_prune AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi WHERE t.d >= toDate('2025-06-01') AND t.d <= toDate('2025-06-10'); ``` Because the escape-hatch setting is now obsolete, users upgrading from 26.4 (or from any version where they had pinned `query_plan_use_new_logical_join_step = 0`) to 26.5+ hit this as a hard performance regression with no setting-level mitigation. <!-- ch-version-info:start --> ### Version info - Resolved by: #113484 - Backported to: `26.7.4.24`, `26.6.3.33` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112586",
        "createdAt": "2026-07-30T12:41:56Z",
        "updatedAt": "2026-08-13T14:32:19Z",
        "timestamp": "2026-08-13T14:32:19Z",
        "metrics": {
          "reactions": 2,
          "comments": 6
        },
        "labels": [
          "performance"
        ],
        "author": "fm4v",
        "state": "closed",
        "assignees": [
          "vdimir"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112903",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "serialize_query_plan = 1: WITH ROLLUP / WITH CUBE / Join-engine lookup join in a distributed subquery fails with NOT_IMPLEMENTED \"Method serialize is not implemented\" — no fallback",
        "text": "`serialize_query_plan = 1` sends the shard-side fragment as a serialized query plan. Three query-plan steps have no `serialize` implementation, and there is no fallback to the text path — any query whose shipped fragment contains one of them fails with `NOT_IMPLEMENTED` at execution time, while the same query succeeds with `serialize_query_plan = 0`: - `Rollup` (`GROUP BY ... WITH ROLLUP` inside the distributed subquery) - `Cube` (`GROUP BY ... WITH CUBE`) - `JoinStepLogicalLookup` (a lookup join against a `Join`-engine table inside the distributed subquery) Notably `GROUP BY GROUPING SETS ((h), ())` — which subsumes `ROLLUP` semantically — serializes fine, as does `WITH TOTALS`; the holes are specifically the `Rollup`/`Cube` transform steps and the lookup-join step. **How to reproduce** On a server with any single-shard cluster whose replica is the server itself (e.g. `test_shard_localhost` from the standard test configs), version 26.8.1.561: ```sql CREATE TABLE t_r (h UInt32) ENGINE = MergeTree ORDER BY h; INSERT INTO t_r SELECT number % 10 FROM numbers(1000); CREATE TABLE dist_t_r AS t_r ENGINE = Distributed(test_shard_localhost, currentDatabase(), t_r); -- OK: returns 11 SELECT count() FROM (SELECT sum(h) FROM dist_t_r GROUP BY h WITH ROLLUP) SETTINGS serialize_query_plan = 0, prefer_localhost_replica = 0; -- Code: 48. DB::Exception: Method serialize is not implemented for Rollup: While executing Remote. (NOT_IMPLEMENTED) SELECT count() FROM (SELECT sum(h) FROM dist_t_r GROUP BY h WITH ROLLUP) SETTINGS serialize_query_plan = 1, prefer_localhost_replica = 0; -- Code: 48 ... Method serialize is not implemented for Cube ... SELECT count() FROM (SELECT sum(h) FROM dist_t_r GROUP BY h WITH CUBE) SETTINGS serialize_query_plan = 1, prefer_localhost_replica = 0; -- lookup-join variant: CREATE TABLE j_r (h UInt32, name String) ENGINE = Join(ANY, LEFT, h); INSERT INTO j_r SELECT number % 10, concat('n', toString(number % 10)) FROM numbers(10); -- OK with serialize_query_plan = 0; with 1: -- Code: 48 ... Method serialize is not implemented for JoinStepLogicalLookup ... SELECT count() FROM (SELECT t.h, j.name FROM dist_t_r AS t ANY LEFT JOIN j_r AS j USING (h)) SETTINGS serialize_query_plan = 1, prefer_localhost_replica = 0; ``` Deterministic: 20/20 failures for both the `Rollup` and the `JoinStepLogicalLookup` shape; 20/20 success with `serialize_query_plan = 0`. Required settings: `serialize_query_plan = 1` plus a genuinely remote read (`prefer_localhost_replica = 0` here — with the local-replica shortcut the plan is never serialized, which is why a plain local read hides the bug). Everything else is at defaults. **Expected**: either these steps get a `serialize` implementation, or `serialize_query_plan` falls back to sending the fragment as text (the way `make_distributed_plan` refuses gracefully with a code 344 `not serializable for remote execution` *before* execution). Failing a valid query at execution time on a plain production `Bool` setting is neither. Related: https://github.com/ClickHouse/ClickHouse/issues/112167 Related: https://github.com/ClickHouse/ClickHouse/issues/112079",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112903",
        "createdAt": "2026-08-01T13:37:49Z",
        "updatedAt": "2026-08-13T06:17:24Z",
        "timestamp": "2026-08-13T06:17:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "bug",
          "comp-query-optimizer",
          "comp-distributed",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "yakov-olkhovskiy"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112905",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Unexpected results with a minmax index and a NULL-containing `SOME` comparison",
        "text": "### Company or project name _No response_ ### Describe what's wrong `c1 = SOME([1, NULL])` is lowered to `has([1, NULL], c1)`. For `c1 = 0`, that is a definite, non-NULL **false** — `0` is not `1`, and \"not found, but the array also contains NULL\" evaluates to false, not unknown, per the row's own projection. So `NOT (c1 = SOME([1, NULL]))` must be true. With a `minmax` skip index on `c1`, the query instead returns **zero rows**, and so do the direct predicate and the `IS NULL` check — the row vanishes from all three. ```sql CREATE TABLE t (c1 INT) ENGINE = MergeTree ORDER BY tuple(); CREATE INDEX idx ON t(c1) TYPE minmax GRANULARITY 1; INSERT INTO t VALUES (0); SELECT c1, NOT (c1 = SOME([1, NULL])) AS p FROM t; -- c1=0, p=1 (true) SELECT count() FROM t WHERE NOT (c1 = SOME([1, NULL])); -- Expected: 1 -- Actual: 0 ``` *Found by an automated fuzzing & triage agent; the repro is verified but the analysis may be wrong.* ### Does it reproduce on the most recent release? Yes ### How to reproduce Can be reproduced on 26.7.1.1315 ```sql CREATE TABLE t (c1 INT) ENGINE = MergeTree ORDER BY tuple(); CREATE INDEX idx ON t(c1) TYPE minmax GRANULARITY 1; INSERT INTO t VALUES (0); SELECT count() FROM t WHERE (c1 = SOME([1, NULL])); -- 0 (correct) SELECT count() FROM t WHERE NOT (c1 = SOME([1, NULL])); -- 0 (WRONG, expect 1) SELECT count() FROM t WHERE ((c1 = SOME([1, NULL]))) IS NULL; -- 0 (correct) -- 0 + 0 + 0 = 0, but the table has 1 row. -- Disabling data-skipping indexes for this query alone (no schema/data -- change) restores the correct result: SELECT count() FROM t WHERE NOT (c1 = SOME([1, NULL])) SETTINGS use_skip_indexes = 0; -- 1 (correct) -- Removing the NULL array element also restores correct indexed execution, -- isolating NULL inside the membership set as part of the trigger: SELECT count() FROM t WHERE NOT (c1 = SOME([1])); -- 1 (correct) ``` ### Expected behavior As mentioned above. ### Error message and/or stacktrace _No response_ ### Related issues and pull requests [#106948](https://github.com/ClickHouse/ClickHouse/issues/106948) / [#110266](https://github.com/ClickHouse/ClickHouse/issues/110266) — `minmax`/`KeyCondition` mis-negating a float range in the presence of NaN. ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112905",
        "createdAt": "2026-08-01T14:07:36Z",
        "updatedAt": "2026-08-13T09:04:34Z",
        "timestamp": "2026-08-13T09:04:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "bug",
          "comp-skip-index",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "suyZhong",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112908",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Constant projection + ORDER BY ALL + LIMIT n BY + LIMIT WITH TIES: exception 10 \"Not found column 1_UInt8 in block\" at default settings",
        "text": "A valid table-free query combining a constant projection, `ORDER BY ALL`, `LIMIT n BY`, and `LIMIT ... WITH TIES` throws an exception at **default settings**: ```sql SELECT 1 FROM numbers(10) ORDER BY ALL LIMIT 1 BY number LIMIT 4 WITH TIES ``` ``` Code: 10. DB::Exception: Not found column 1_UInt8 in block. There are only columns: __table1.number. (NOT_FOUND_COLUMN_IN_BLOCK) ``` All four ingredients are required — removing any one produces the correct result on 26.8.1.561: - without `WITH TIES`: `SELECT 1 ... LIMIT 1 BY number LIMIT 4` → OK - non-constant projection: `SELECT number ... WITH TIES` → OK - explicit key instead of `ORDER BY ALL`: `SELECT 1 ... ORDER BY number ... WITH TIES` → OK - without `LIMIT 1 BY number` → OK Deterministic (20/20), no settings involved, no tables involved. `WITH TIES` needs the sort-description columns in the stream to compare ties; the constant projection under `ORDER BY ALL` + `LIMIT BY` apparently drops the constant from the block header the tie-comparison expects. Found while fuzzing plan-shape combinations; the same shape fails identically through `Distributed` and parallel-replica reads (the local exception is the root).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112908",
        "createdAt": "2026-08-01T14:44:00Z",
        "updatedAt": "2026-08-13T06:17:28Z",
        "timestamp": "2026-08-13T06:17:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "bug",
          "comp-query-analyzer",
          "comp-query-execution",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "KochetovNicolai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:112909",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "FULL JOIN USING over a Distributed left table: qualified column t1.a returns the coalesced USING value for right-only rows (or exception 8 at pure defaults)",
        "text": "`FULL JOIN ... USING` where the **left** table is read through `Distributed`: selecting the qualified join column `t1.a` is wrong for right-only rows — and at pure defaults the same query throws an exception. **How to reproduce** (26.8.1.561, any single-shard cluster whose replica is the server itself, e.g. `test_shard_localhost` from the standard test configs): ```sql CREATE TABLE t1 (a UInt16, b UInt16) ENGINE = MergeTree ORDER BY tuple(); CREATE TABLE t2 (a Int16, b Nullable(Int64)) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t1 SELECT number + 1, number FROM numbers(10); INSERT INTO t2 SELECT number - 4, number FROM numbers(10); CREATE TABLE dist_t1 AS t1 ENGINE = Distributed(test_shard_localhost, currentDatabase(), t1); CREATE TABLE dist_t2 AS t2 ENGINE = Distributed(test_shard_localhost, currentDatabase(), t2); -- LOCAL, correct: right-only rows show t1.a = 0 (the UInt16 default) SELECT a, t1.a, t2.a FROM t1 FULL JOIN t2 USING (a) ORDER BY (t1.a, t2.a); -- DISTRIBUTED, wrong values: right-only rows show t1.a = -4..-1 -- (the COALESCED USING value leaks into the qualified left column) SELECT a, t1.a, t2.a FROM dist_t1 AS t1 FULL JOIN dist_t2 AS t2 USING (a) ORDER BY (t1.a, t2.a) SETTINGS prefer_localhost_replica = 0; -- DISTRIBUTED, pure defaults (prefer_localhost_replica = 1): exception -- Code: 8. DB::Exception: Cannot find column `a` in source stream, -- there are only columns: [a, __table1.a, t2.a]. (THERE_IS_NO_COLUMN) SELECT a, t1.a, t2.a FROM dist_t1 AS t1 FULL JOIN dist_t2 AS t2 USING (a) ORDER BY (t1.a, t2.a); ``` Deterministic 20/20 for both manifestations. Characterization: - Wrapping ONLY the left table in `Distributed` is sufficient (right side local: same wrong values / exception). - `RIGHT JOIN` is correct; the defect is FULL-specific (left-only default rows vs right-only coalesced rows). - The bare `a` (USING projection) is correct in all variants — only the QUALIFIED `t1.a` is corrupted/lost. - `prefer_localhost_replica` selects which manifestation appears (1, the default → exception 8; 0 → silent wrong values); everything else is at defaults. Related: https://github.com/ClickHouse/ClickHouse/issues/66739",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/112909",
        "createdAt": "2026-08-01T14:44:45Z",
        "updatedAt": "2026-08-13T09:34:32Z",
        "timestamp": "2026-08-13T09:34:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "bug",
          "comp-joins",
          "comp-distributed",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "KochetovNicolai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113182",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "`optimize_inverse_dictionary_lookup` rewrite inside a correlated EXISTS makes decorrelation fail: 48 \"Cannot decorrelate query, because 'DelayedCreatingSets' step is not supported\"",
        "text": "## Describe the problem A valid query with a correlated `EXISTS` subquery fails with exception 48 (`NOT_IMPLEMENTED`) at pure default settings when the subquery's `WHERE` contains a `dictGet(...) >= <const>` comparison. `optimize_inverse_dictionary_lookup` (default `1`) rewrites the `dictGet('d', 'attr', key) >= c` predicate into `key IN __set_...`. When the predicate sits inside a correlated subquery, the injected set adds a `DelayedCreatingSets` step to the subquery plan, and the correlated-subquery decorrelation then refuses the plan: ``` Code: 48. DB::Exception: Cannot decorrelate query, because 'DelayedCreatingSets' step is not supported. (NOT_IMPLEMENTED) ``` Both settings involved are on by default (`optimize_inverse_dictionary_lookup = 1`, `allow_experimental_correlated_subqueries = 1`), so an ordinary query errors out of the box. Disabling the rewrite (`optimize_inverse_dictionary_lookup = 0`) makes the same query run fine and return the correct result. This is the only setting that matters: `correlated_subqueries_use_in_memory_buffer = 0` does not help, and the same failure occurs when the `EXISTS` is used as a scalar (e.g. `o.g >= (EXISTS ...)`), not just in `WHERE`. ## How to reproduce Version: `26.8.1.653` (public master build; also reproduced on `26.8.1.561`). All settings at defaults. ```sql CREATE TABLE i_src (id UInt64, val UInt32) ENGINE=MergeTree ORDER BY id; INSERT INTO i_src SELECT number, number % 97 FROM numbers(500); CREATE DICTIONARY i_dict (id UInt64, val UInt32 DEFAULT 0) PRIMARY KEY id SOURCE(CLICKHOUSE(TABLE 'i_src' DB 'default')) LIFETIME(0) LAYOUT(FLAT()); CREATE TABLE i_t (k UInt32, g Int32, s String) ENGINE=MergeTree ORDER BY k; INSERT INTO i_t SELECT number, number % 7, toString(number % 5) FROM numbers(100); SELECT o.k FROM i_t AS o WHERE EXISTS ( SELECT 1 FROM i_t AS i WHERE i.s = o.s AND dictGet('i_dict', 'val', toUInt64(abs(g)) % 500) >= 76); ``` Observed (deterministic, 20/20 runs): ``` Code: 48. DB::Exception: Cannot decorrelate query, because 'DelayedCreatingSets' step is not supported. (NOT_IMPLEMENTED) ``` Expected: the query executes and returns the rows for which the `EXISTS` holds (here an empty result — confirmed by running with `SETTINGS optimize_inverse_dictionary_lookup = 0`, which succeeds). Settings analysis: - REQUIRED: none beyond defaults. `optimize_inverse_dictionary_lookup = 0` is the only toggle that avoids the error. - Incidental: `correlated_subqueries_use_in_memory_buffer` (both values fail), the exact comparison constant, table sizes, dictionary layout. ## Additional context Mechanism: `InverseDictionaryLookupPass` rewrites the `dictGet` comparison into a membership test against an internally built set (`__set_...`). Building that set inside the correlated subquery introduces a `DelayedCreatingSets` plan step, and the decorrelation of correlated subqueries does not support that step, so planning aborts with 48. Either the pass should be skipped inside correlated subqueries it would break, or the decorrelator should learn to handle `DelayedCreatingSets`. Same \"optimization-pass output escapes into an unsupported context\" family as the `ColumnSet`-reaches-`FINAL`-merge manifestation of the same pass, but a distinct trigger and error code. Related: https://github.com/ClickHouse/ClickHouse/issues/112030 Related: https://github.com/ClickHouse/ClickHouse/issues/99500",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113182",
        "createdAt": "2026-08-03T20:23:51Z",
        "updatedAt": "2026-08-13T06:17:33Z",
        "timestamp": "2026-08-13T06:17:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "bug",
          "comp-query-optimizer",
          "comp-query-analyzer",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "nihalzp"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113184",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Materialized CTE over Distributed: 49 LOGICAL_ERROR \"Reading from materialized CTE before its materialization completed - DelayedPortsProcessor gate is missing\" survives the #108924 fix",
        "text": "## Describe the problem With `enable_materialized_cte = 1`, a twice-referenced materialized CTE reading a `Distributed` table, filtered by `IN (SELECT ... FROM <another materialized CTE>)`, fails with exception 49 (`LOGICAL_ERROR`): ``` Code: 49. DB::Exception: Reading from materialized CTE 'ct' before its materialization completed - DelayedPortsProcessor gate is missing in the query plan: While executing Memory. (LOGICAL_ERROR) ``` This is the fail-fast invariant introduced by the materialized-CTE scheduling fix itself (PR https://github.com/ClickHouse/ClickHouse/pull/108924, merged) — verified on a master build that CONTAINS that merge: the PR's own `04036_materialized_cte_distributed_race` shapes pass, but this neighbouring shape on the same `Distributed` path still trips the retained gate check. The identical shapes over a plain `MergeTree` table pass (the `04227_materialized_cte_reused_with_in_subquery` cases). Release builds throw; debug builds abort on the logical-error check. ## How to reproduce Version: `26.8.1.653` (public master build, includes the #108924 merge). Single server; the single-shard `test_shard_localhost` cluster is enough: ```sql SET enable_materialized_cte = 1; CREATE TABLE t (c Int32) ENGINE = MergeTree ORDER BY c; INSERT INTO t VALUES (1), (2), (3); CREATE TABLE dist_t AS t ENGINE = Distributed(test_shard_localhost, currentDatabase(), t); WITH ct AS MATERIALIZED (SELECT 1 AS c), rs AS MATERIALIZED (SELECT * FROM dist_t WHERE c IN (SELECT c FROM ct)) SELECT count() FROM rs AS a, rs AS b; ``` Deterministic: 20/20 exceptions with `max_threads = 1`, 20/20 with `max_threads = 8`, and the same with `serialize_query_plan` 0 and 1 (80/80 total). Expected: `1` (the count the same query returns over the plain `MergeTree` table). Required ingredients (each verified by removing it): - `enable_materialized_cte = 1`. - The twice-referenced materialized CTE (`rs`) reads a `Distributed` table. Only `rs` needs it — `ct` can be a constant `SELECT 1 AS c`, and moving the `Distributed` read into `ct` while `rs` reads the local table makes the query pass. - `rs` filters by `IN (SELECT ... FROM ct)` where `ct` is another *materialized* CTE. With a plain (non-materialized) CTE, a bare `IN (SELECT 1)`, or a literal `IN (1, 2, 3)` the query passes — this looks like the `forceMaterializeCTE` path for IN-set subqueries, which the fix intentionally kept \"collected and gated exactly as before\". - `rs` referenced twice in the join tree. `rs AS a, rs AS b`, `ANY LEFT JOIN`, and `SELECT * FROM rs UNION ALL SELECT * FROM rs` all fail alike. Incidental: the join kind, `max_threads`, `serialize_query_plan`, the number of shards. ## Additional context A *single* reference to `rs` over `Distributed` fails differently — `Code: 60 UNKNOWN_TABLE: Unknown table expression identifier '_materialized_cte_ct_...'` on the shard-local query — the temp-table shipping defect family (#112642 is the parallel-replicas variant), so the double reference selects this gate-missing channel rather than merely amplifying it. Exception site: `src/Processors/QueryPlan/ReadFromMemoryStorageStep.cpp` (`MemorySource::generate`), reached from the materialization pipeline. Distinct from the other open \"gate is missing\" fixes: #111194 (`WITH TOTALS`, per #110176) and #113043 (extremes + `UNION`) — neither involves the `Distributed` + IN-set path. Found by query fuzzing on master (`tests/optimizer_tester`, branch `optimizer-tester-framework`). Related: https://github.com/ClickHouse/ClickHouse/pull/108924 Related: https://github.com/ClickHouse/ClickHouse/issues/112642 Related: https://github.com/ClickHouse/ClickHouse/issues/111194 Related: https://github.com/ClickHouse/ClickHouse/issues/113043",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113184",
        "createdAt": "2026-08-03T20:24:05Z",
        "updatedAt": "2026-08-13T10:01:33Z",
        "timestamp": "2026-08-13T10:01:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "bug",
          "comp-query-optimizer",
          "comp-distributed",
          "comp-query-execution",
          "common table expressions"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113320",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Text index direct read drops the source column another PREWHERE step needs: NOT_FOUND_COLUMN_IN_BLOCK on a valid query",
        "text": "**Describe what's wrong** With a `text` index on column `c`, a query that filters on `c` in both `PREWHERE` and `WHERE` fails with `Code: 10. NOT_FOUND_COLUMN_IN_BLOCK` on a perfectly valid query, when `c` itself is not in the SELECT list. The direct-read optimization (`query_plan_direct_read_from_text_index`, default `1`) rewrites one of the two text-searchable predicates (the `ILIKE`) to read the index's virtual column (`__text_index_tix_ilike_<hash>`), and the source column `c` is dropped from the block — but the other read step still needs `c` to evaluate its predicate. **How to reproduce** ```sql SET allow_experimental_full_text_index = 1; CREATE TABLE tt (a String, c String, INDEX tix c TYPE text(tokenizer = 'splitByNonAlpha') GRANULARITY 1) ENGINE = MergeTree ORDER BY a; INSERT INTO tt SELECT toString(number % 10), concat('8Gamma', toString(number)) FROM numbers(1000); SELECT a FROM tt PREWHERE c LIKE '8%' WHERE c ILIKE '%Gamma%' ORDER BY a; ``` ``` Code: 10. DB::Exception: Not found column c: in block a String String(size = 0), __text_index_tix_ilike_dd7947548bd47c991b09b7e00eb35dbd UInt8 UInt8(size = 0). (NOT_FOUND_COLUMN_IN_BLOCK) ``` Deterministic: 20/20 runs on current master (`clickhouse local` and server), also reproduces on 26.7.1.1077. **Observations** - `SETTINGS query_plan_direct_read_from_text_index = 0` returns the correct result — the direct-read rewrite is the trigger. - Adding `c` to the SELECT list avoids the exception. - Removing either of the two predicates on `c` avoids the exception: the bug needs one predicate rewritten to the index virtual column while the other still wants the raw column. **Expected behavior** The query returns the matching rows; whether the text index answers one of the predicates is an internal optimization and must not make the plan lose a column another step consumes. Related: https://github.com/ClickHouse/ClickHouse/issues/109329 (the `make_distributed_plan` manifestation of a direct-read/`__text_index_*` column mismatch — this report is single-node, no distributed plan involved) Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113320",
        "timestamp": "2026-08-12T21:54:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "bug",
          "comp-query-optimizer",
          "comp-text-index"
        ],
        "author": "zlareb1",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113337",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Paimon tables with a nullable ARRAY or MAP column cannot be read: \"Nested type ... cannot be inside Nullable type\"",
        "text": "A Paimon table containing a nullable `ARRAY` or `MAP` column cannot be read at all: schema parsing wraps the composite type in `Nullable`, which ClickHouse forbids, so `DESC` and `SELECT` throw an exception before any data is read. Since Paimon columns are nullable by default (e.g. any Spark-created table with an array column), this makes fairly ordinary Paimon tables entirely unreadable — the exception is raised while parsing the table schema in `PaimonMetadata::create`, so it affects the whole table, not just queries touching the offending column. **How to reproduce** Write a Paimon table with a nullable array column (Paimon 1.1.1, Spark 3.5): ```sql -- spark-sql with org.apache.paimon:paimon-spark-3.5:1.1.1, -- spark.sql.catalog.paimon = org.apache.paimon.spark.SparkCatalog, -- spark.sql.catalog.paimon.warehouse = file:/tmp/paimon_wh CREATE TABLE paimon.default.t (f ARRAY<INT>) TBLPROPERTIES ('file.format'='parquet'); INSERT INTO paimon.default.t VALUES (array(1,2)); ``` Read it with ClickHouse (master, commit `7c826d816cd5`): ``` $ clickhouse local --query \"DESC paimonLocal('/tmp/paimon_wh/default.db/t')\" Code: 43. DB::Exception: Nested type Array(Nullable(Int32)) cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT) ``` The same error occurs via `paimonS3` / `paimonAzure` and the `Paimon*` table engines. A nullable `MAP` column fails identically with `Nested type Map(...) cannot be inside Nullable type`. **Root cause** `Paimon::DataType::parse` in `src/Storages/ObjectStorage/DataLakes/Paimon/Types.h` wraps the result in `DataTypeNullable` when the Paimon field is nullable, including for the `ARRAY` and `MAP` branches: ```cpp if (real_type == \"ARRAY\") { ... type.clickhouse_data_type = std::make_shared<DataTypeArray>(nested_type.clickhouse_data_type); if (nullable) type.clickhouse_data_type = std::make_shared<DataTypeNullable>(type.clickhouse_data_type); } ``` `DataTypeNullable`'s constructor rejects composite nested types (`src/DataTypes/DataTypeNullable.cpp`), hence the exception. The `Nullable` wrap should be skipped for `ARRAY` and `MAP` (reading a null array/map as empty), which is how the Iceberg and Delta Lake type mappers handle nullable composites. The existing test fixtures (`tests/queries/0_stateless/data_minio/paimon_all_types`) never hit this because their array/map fields are declared non-nullable. Found during E2E tests: https://github.com/ClickHouse/clickhouse-private/pull/67431; the new type-matrix test works around it with `NOT NULL` composite columns for now. Caused by: https://github.com/ClickHouse/ClickHouse/pull/102343",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113337",
        "createdAt": "2026-08-04T14:18:21Z",
        "updatedAt": "2026-08-13T15:43:09Z",
        "timestamp": "2026-08-13T15:43:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "bug",
          "unfinished code",
          "experimental feature",
          "comp-datalake"
        ],
        "author": "zlareb1",
        "state": "closed",
        "assignees": [
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113360",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Paimon primary-key tables silently return duplicate/stale rows: no merge-on-read, primary keys ignored",
        "text": "Reading a Paimon primary-key table returns duplicate/stale rows silently: the reader collects the raw union of base and delta data files with no merge-on-read, so superseded row versions from upserts are returned alongside current ones. Primary keys are parsed from the table schema (`PaimonSchemaProcessor::getPrimaryKeys`) and the LSM merge metadata is parsed from manifests (`_LEVEL`, `_MIN_SEQUENCE_NUMBER`, `_MAX_SEQUENCE_NUMBER` in `PaimonClient`), but none of it is used by the read path in `PaimonMetadata`. Since primary-key tables are the default choice for upsert/CDC workloads in Paimon, this produces silently wrong results on a very common table type — no exception, no warning. **How to reproduce** Write a primary-key table with an upsert (Paimon 1.1.1, Spark 3.5): ```sql CREATE TABLE paimon.default.pk_t (id INT, val STRING) TBLPROPERTIES ('primary-key'='id', 'bucket'='1', 'file.format'='parquet'); INSERT INTO paimon.default.pk_t VALUES (1, 'old'), (2, 'two'); INSERT INTO paimon.default.pk_t VALUES (1, 'new'); SELECT * FROM paimon.default.pk_t ORDER BY id; -- Spark (correct): 1 new -- 2 two ``` Read it with ClickHouse (master, commit `7c826d816cd5`): ``` $ clickhouse local --query \"SELECT * FROM paimonLocal('/tmp/paimon_pk_wh/default.db/pk_t') ORDER BY id, val\" 1 new 1 old <-- superseded row version returned silently 2 two ``` The same happens via `paimonS3` / `paimonAzure` and the `Paimon*` table engines. **Suggested behavior** Until merge-on-read is implemented, reading a table whose schema declares `primary-key` should throw an exception (fail-close), the same way other unsupported Paimon features already do (`ROW` type, unsupported `scan.mode` values). Silent wrong data is strictly worse than an error. Found while adding Paimon-over-object-storage integration tests (the suites deliberately avoid primary-key tables because of this). Related: https://github.com/ClickHouse/ClickHouse/issues/113337 Caused by: https://github.com/ClickHouse/ClickHouse/pull/102343",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113360",
        "createdAt": "2026-08-04T17:16:17Z",
        "updatedAt": "2026-08-13T13:22:19Z",
        "timestamp": "2026-08-13T13:22:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "bug",
          "unfinished code",
          "unexpected behaviour",
          "experimental feature",
          "comp-datalake"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": [
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113711",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "MATERIALIZED CTE is not materialized when queried through a VIEW",
        "text": "### Company or project name _No response_ ### Describe the unexpected behaviour ClickHouse runs CTE twice if we select from a view based on a query with MATERIALIZED CTE. ### Which ClickHouse versions are affected? Latest (26.7) ### How to reproduce https://fiddle.clickhouse.com/eb4a0c4a-7efa-4554-a3a6-67c964162b15 ```sql -- 1. Create a sample table CREATE TABLE test_cte ( customer_id UInt32, amount Decimal(10, 2) ) ENGINE = MergeTree ORDER BY customer_id; -- 2. Insert sample data INSERT INTO test_cte (customer_id, amount) VALUES (1, 500.00), (1, 750.00), (2, 200.00), (2, 300.00), (3, 1500.00); --3. CREATE VIEW with MATERIALIZED CTE CREATE VIEW view_cte AS WITH summary AS MATERIALIZED ( SELECT customer_id, sum(amount) AS total_amount FROM test_cte GROUP BY customer_id ) SELECT * FROM summary AS s1 INNER JOIN summary AS s2 ON s1.customer_id = s2.customer_id SETTINGS enable_materialized_cte = 1, final = 1; -- 4. Check usage of MATERIALIZED CTE -- raw query use MATERIALIZED ! EXPLAIN indexes=1 WITH summary AS MATERIALIZED ( SELECT customer_id, sum(amount) AS total_amount FROM test_cte GROUP BY customer_id ) SELECT * FROM summary AS s1 INNER JOIN summary AS s2 ON s1.customer_id = s2.customer_id SETTINGS enable_materialized_cte = 1, final= 1; MaterializingCTEs (Materialize CTEs before main query execution) ... └──MaterializingCTE (Materializing CTE: summary) └──Aggregating │ Keys: customer_id │ Aggregates: sum(amount) │ Skip merging: 0 └──ReadFromMergeTree (default.test_cte) ... -- select from the view => NO Materialized CTE EXPLAIN indexes=1 SELECT * FROM view_cte; Join (JOIN FillRightFirst) ... ├──Aggregating │ │ Keys: customer_id │ │ Aggregates: sum(amount) │ │ Skip merging: 0 │ └──ReadFromMergeTree (default.test_cte) ... └──BuildRuntimeFilter (Build runtime join filter on customer_id) │ Filter id: RF1 │ Source table: default.test_cte └──Aggregating │ Keys: customer_id │ Aggregates: sum(amount) │ Skip merging: 0 └──ReadFromMergeTree (default.test_cte) ... ``` ### Expected behavior MATERIALIZED CTE is materialized when queried through a VIEW ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113711",
        "createdAt": "2026-08-06T18:31:10Z",
        "updatedAt": "2026-08-13T11:59:13Z",
        "timestamp": "2026-08-13T11:59:13Z",
        "metrics": {
          "reactions": 2,
          "comments": 2
        },
        "labels": [
          "bug",
          "experimental feature"
        ],
        "author": "SaltTan",
        "state": "open",
        "assignees": [
          "novikd"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113741",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Custom-key parallel replicas over a Merge table with a Distributed child: children are offloaded to a finalized stage (CANNOT_CONVERT_TYPE, logical error in GroupingAggregatedTransform)",
        "text": "🕵 Reading a `Merge` table (or the `merge` table function) with custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_sampling'` or `'custom_key_range'`) fails with an exception when the common processing stage of the children is `WithMergeableState` — for example, when one of the underlying tables is a `Distributed` table. ```sql CREATE TABLE t_mrg_ck_1 (k UInt64) ENGINE = MergeTree ORDER BY k AS SELECT number FROM numbers(100000); CREATE TABLE t_mrg_ck_2 (k UInt64) ENGINE = MergeTree ORDER BY k AS SELECT number FROM numbers(100000); CREATE TABLE t_mrg_ck_3 (k UInt64) ENGINE = Distributed('test_shard_localhost', currentDatabase(), 't_mrg_ck_1'); SET enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_sampling', parallel_replicas_custom_key = 'k'; SELECT count() FROM merge(currentDatabase(), '^t_mrg_ck_'); ``` ``` Code: 70. DB::Exception: Conversion from UInt64 to AggregateFunction(count) is not supported: while converting source column `count()` to destination column `count()`: Child table: default.t_mrg_ck_1. (CANNOT_CONVERT_TYPE) ``` Depending on the aggregate function, it instead trips an assertion in the pipeline — an exception in debug and sanitizer builds (found by the AST fuzzer, STID `3970-479a`): ```sql SELECT 47, quantileExactInclusive(visibleWidth(['1', '2'])) IGNORE NULLS FROM merge(currentDatabase(), '^t_mrg_ck_') GROUP BY ALL LIMIT 973; ``` ``` Logical error: 'Chunk should have AggregatedChunkInfo/ChunkInfoWithAllocatedBytes in GroupingAggregatedTransform.'. ``` **Root cause.** The `Distributed` child reports `WithMergeableState` from its `getQueryProcessingStage`, so `StorageMerge::getQueryProcessingStage` sets the common stage of all children to `WithMergeableState`, and `ReadFromMerge::createPlanForTable` plans each `MergeTree` child through an interpreter with `SelectQueryOptions(WithMergeableState)`. Inside that child interpreter, the custom-key branch of `PlannerJoinTree` (`src/Planner/PlannerJoinTree.cpp`, the `canUseParallelReplicasCustomKey` block) offloads the child query to the replicas at the hard-coded stage `WithMergeableStateAfterAggregationAndLimit`, ignoring the requested `to_stage`. The child plan therefore produces finalized rows (post-aggregation, post-`LIMIT`) where the parent `ReadFromMerge` pipeline expects partial aggregation states: `convertAndFilterSourceStream` throws `CANNOT_CONVERT_TYPE` when the finalized type differs from the state type, and when the types coincide structurally, the chunks without `AggregatedChunkInfo` reach `GroupingAggregatedTransform` and trip the assertion. Only the analyzer path is affected (`enable_analyzer = 0` returns correct results). Reproduced on current master (verified on a binary with no unrelated changes); found by the targeted AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/110972 (which is unrelated: the failure reproduces without `parallel_replicas_allow_merge_tables`, a setting that does not exist on master), report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110972&sha=ac0a584ce316ace31f4dbc16b38e1262a2344751&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29 The custom-key offload should be skipped when the requested `to_stage` is below `WithMergeableStateAfterAggregationAndLimit`: a plan that must stop at a partial stage cannot accept a finalized remote read. Related: https://github.com/ClickHouse/ClickHouse/pull/110972 <!-- ch-version-info:start --> ### Version info - Resolved by: #113742 - Backported to: `26.7.4.29` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113741",
        "createdAt": "2026-08-06T23:14:04Z",
        "updatedAt": "2026-08-13T16:23:00Z",
        "timestamp": "2026-08-13T16:23:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "comp-distributed",
          "potential bug",
          "comp-parallel-replicas"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113763",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Propagating `getArgumentsThatCanBeOnlyNull` through the combinators silently changes the result type and value of every pre-existing `<agg>If<Combinator>(x, NULL)` expression",
        "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ ### Describe what's wrong After this PR, `countIfOrNull(number, NULL)` returns `Nullable(UInt64)` `NULL` where it returned `UInt64` `0`; `uniqIfOrNull(number, NULL)` likewise flips `0` -> `NULL`; `sumIfResample(0, 2, 1)(number, NULL, number % 2)` returns `Array(UInt64)` `[0,0]` where it returned the scalar `Nullable(Nothing)` `NULL`; and `sumIfState(number, NULL)` now produces a real `AggregateFunction(sumIf, UInt64, Nullable(Nothing))` state where it produced the constant `Nullable(Nothing)`. The same shift applies to `count`, `uniq`, `any`, `min`, `max`, `avg`, `groupArray`, `quantile`, `varSamp` and `topK` under `-State`, `-OrNull`, `-OrDefault`, `-Resample`, `-ArgMin` and `-ArgMax`. None of this is mentioned in the PR title, description or changelog entry, all of which speak only about the new `gini` function. **Root cause:** src/AggregateFunctions/Combinators/AggregateFunctionState.h:163 (and the six sibling overrides listed in affected_locations) newly forward `getArgumentsThatCanBeOnlyNull` from the nested function. Combined with the existing consumer at src/AggregateFunctions/Combinators/AggregateFunctionNull.cpp:88, this removes the `AggregateFunctionNothing` fold for every `<agg>If<Combinator>(x, NULL)` expression - a user-visible change to functions completely unrelated to `gini`, shipped under the changelog category `New Feature`. **Why we believe this is a bug:** `AggregateFunctionFactory::get` (src/AggregateFunctions/AggregateFunctionFactory.cpp:106-127) applies the `Null` combinator whenever any argument type is `Nullable`. `AggregateFunctionCombinatorNull::transformAggregateFunction` (src/AggregateFunctions/Combinators/AggregateFunctionNull.cpp:88) asks the nested function for `getArgumentsThatCanBeOnlyNull` and, if an only-null (`Nullable(Nothing)`) argument is NOT in that set, folds the whole aggregate into `AggregateFunctionNothing` (line 116-118). Before this PR only `AggregateFunctionIf` overrode the method, so as soon as a second combinator wrapped the `-If` function the information was lost and the fold happened. The PR makes `-State`, `-SimpleState`, `-OrFill`, `-Distinct`, `-Resample`, `-ArgMin`/`-ArgMax` and `AggregateFunctionNullBase` forward the nested set, so the fold no longer happens and the real combinator chain is built instead. Risk if broken: result types of persisted `AggregateFunction` columns and materialized-view schemas change on upgrade, and aggregate values silently flip from `0` to `NULL`. **Affected locations:** - `src/AggregateFunctions/Combinators/AggregateFunctionState.h:163` — new getArgumentsThatCanBeOnlyNull forwarding to nested_func - `src/AggregateFunctions/Combinators/AggregateFunctionOrFill.h:388` — new forwarding override, drives -OrNull/-OrDefault - `src/AggregateFunctions/Combinators/AggregateFunctionResample.h:247` — new override, nested set plus last_col - `src/AggregateFunctions/Combinators/AggregateFunctionDistinct.h:375` — new forwarding override - `src/AggregateFunctions/Combinators/AggregateFunctionCombinatorsArgMinArgMax.cpp:225` — new override, nested set plus key_col - `src/AggregateFunctions/Combinators/AggregateFunctionNull.h:285` — AggregateFunctionNullBase now forwards to nested_function - `src/AggregateFunctions/Combinators/AggregateFunctionSimpleState.h:108` — new forwarding override - `src/AggregateFunctions/Combinators/AggregateFunctionNull.cpp:88` — consumer: decides whether to fold into AggregateFunctionNothing **Impact:** On upgrade, existing queries and materialized views that stack a second combinator on `-If` with a constant-`NULL` (or `Nullable(Nothing)`) filter change their result type, and `countIfOrNull`/`uniqIfOrNull` change their value from `0` to `NULL`. `<agg>IfResample(...)` changes shape from a scalar to an `Array`, which makes downstream expressions that consumed the scalar fail. `<agg>IfState(x, NULL)` becomes a real `AggregateFunction(...)` state, which is schema-affecting for `CREATE MATERIALIZED VIEW ... AS SELECT sumIfState(...)`. ## Assumptions The bot recorded these claims it could not verify directly from source. A maintainer ✅ confirms; ❌ flags a wrong premise (the bot should rework or close the finding). - [ ] **The new results are semantically better than the old ones (a real state beats a folded constant, and `-Resample` must return an Array), so the intent is a fix rather than a regression.** - *Why unverifiable:* The PR description does not mention the combinator change at all, so the author's intent for non-`gini` functions is not stated anywhere. - *Falsifiable test:* Ask the author/maintainer whether the change in `countIfOrNull(x, NULL)` from `0` to `NULL` and in `sumIfResample(...)(x, NULL, k)` from scalar `NULL` to `[0,0]` is intended; if yes, the changelog category must become `Backward Incompatible Change` and the behavior must be covered by a test. ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `31081d9f050140`). ### How to reproduce ```sql -- Test: result of an -If aggregate with a constant NULL filter, combined with a second combinator. SELECT toTypeName(countIfOrNull(number, NULL)), countIfOrNull(number, NULL) FROM numbers(5) ORDER BY 1; SELECT toTypeName(uniqIfOrNull(number, NULL)), uniqIfOrNull(number, NULL) FROM numbers(5) ORDER BY 1; SELECT toTypeName(sumIfResample(0, 2, 1)(number, NULL, number % 2)), sumIfResample(0, 2, 1)(number, NULL, number % 2) FROM numbers(5) ORDER BY 1; SELECT toTypeName(x) FROM (SELECT sumIfState(number, NULL) AS x FROM numbers(5)) ORDER BY 1; ``` [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/f1146cf3-cf28-4eea-bbd7-9ecef5328573) ### Expected behavior ``` UInt64 0 UInt64 0 Nullable(Nothing) \\N Nullable(Nothing) ``` ### Error message and/or stacktrace ``` Nullable(UInt64) \\N Nullable(UInt64) \\N Array(UInt64) [0,0] AggregateFunction(sumIf, UInt64, Nullable(Nothing)) ``` ### Additional context **Open risks:** - `AggregateFunctionArray`, `AggregateFunctionForEach`, `AggregateFunctionMap` and `AggregateFunctionMerge` also expose `getNestedFunction` but were NOT given the override. I audited all four: each rejects a top-level `Nullable(Nothing)` argument before the `Null` combinator can matter (`Illegal type Nullable(Nothing) of argument for aggregate function with Array/ForEach suffix. Must be array.`, `Aggregate function Map requires map as argument.`, and `-Merge` requires an `AggregateFunction(...)` argument), so the omission is not reachable - not filed as a separate finding. - State binary compatibility is unaffected: `CREATE TABLE (s AggregateFunction(sumIf, UInt64, Nullable(Nothing)))` and `CAST(... AS AggregateFunction(sumIf, UInt64, Nullable(Nothing)))` behave identically on 26.7 and on the PR build (both expect 8 bytes), so only the inferred type of `sumIfState(x, NULL)` moved. **Suggested fix:** Keep the combinator propagation (it is the more consistent behavior), but re-categorize the changelog entry as `Backward Incompatible Change`, spell out the affected expression shapes in the entry, and add the non-`gini` cases to a stateless test so the new behavior is pinned. If the change to `countIfOrNull(x, NULL)` (`0` -> `NULL`) is not intended, restrict the propagation to the combinators `gini` actually needs. **Analysis details:** Confidence HIGH | Severity P2 | Testability: `STATELESS_SQL` Found during automated review of [PR #112280](https://github.com/ClickHouse/ClickHouse/pull/112280). --- _ClickGapAI · Confidence: HIGH · Severity: P2 · Finding: `h_pr112280_001`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #113868 - Merged into: `26.8.1.1317` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113763",
        "createdAt": "2026-08-07T04:57:14Z",
        "updatedAt": "2026-08-13T13:36:04Z",
        "timestamp": "2026-08-13T13:36:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "comp-aggregate-functions"
        ],
        "author": "clickgapai",
        "state": "closed",
        "assignees": [
          "Algunenano"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113891",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "[CI crash] Incorrect type cast during serialization",
        "text": "<details><summary>Stack trace details</summary> The sipHash64(st.trace_full) is 7748280104269371522 The trace is from the master or release branch: True The query for CIDB to compare the trace with the known one: ```sql WITH ( SELECT groupArrayDistinct(cleanStackTrace(trace_full) AS trace) FROM default.stack_traces WHERE sipHash64(trace) IN (7748280104269371522, {ANOTHER_TRACE_HASH}) -- FIXME: replace with the known hash ) AS traces, 1.97 AS alpha, stack_frame_weights AS ( WITH ( SELECT count() FROM default.stack_traces FINAL ) AS total, 2.0 AS beta, 3.7 AS gamma SELECT arrayJoin(cleanStackTrace(trace_full)) AS frame, countDistinct(trace_full) AS count, log(total / count) AS IDF, sigmoid(beta * (IDF - gamma)) AS weight FROM default.stack_traces FINAL GROUP BY frame ), (SELECT groupArray(weight) AS w, groupArray(frame) AS f FROM stack_frame_weights) AS weights, (trace -> arrayMap((_frame, pos) -> (pow(pos, -alpha) * arrayFirst(w, f -> (f = _frame), weights.w, weights.f)), trace, arrayEnumerate(trace))) AS get_trace_weights, (arr -> arrayStringConcat(arr, '\\n')) AS joinArr SELECT arraySimilarity(traces[1], traces[2], get_trace_weights(traces[1]) AS weights1, get_trace_weights(traces[2]) AS weights2) AS similarity, arrayLevenshteinDistanceWeighted(traces[1], traces[2], weights1, weights2), joinArr(traces[1]), joinArr(traces[2]), joinArr(weights1), joinArr(weights2) ``` </details> The following new stack trace from CI Logs `system.crash_log` found: ``` throwBadTypeidCast(std::type_info const&, std::type_info const&) DB::SerializationString::serializeBinaryBulk(DB::IColumn const&, DB::WriteBuffer&, unsigned long, unsigned long) const DB::ISerialization::serializeBinaryBulkWithMultipleStreams(DB::IColumn const&, unsigned long, unsigned long, DB::ISerialization::SerializeBinaryBulkSettings&, std::shared_ptr<DB::ISerialization::SerializeBinaryBulkState>&) const DB::MergeTreeDataPartWriterOnDisk::initColumnsSubstreamsIfNeeded() DB::MergeTreeDataPartWriterWide::write(DB::Block const&, DB::PODArray<unsigned long, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, DB::Block*) DB::MergedBlockOutputStream::write(DB::Block const&) DB::PartMergerWriter::mutateOriginalPartAndPrepareProjections() DB::MutateAllPartColumnsTask::executeStep() DB::MutateTask::execute() DB::MutatePlainMergeTreeTask::executeStep() DB::TaskRuntimeData::executeStep() const DB::MergeTreeBackgroundExecutor<DB::DynamicRuntimeQueue>::routine(std::shared_ptr<DB::TaskRuntimeData>) DB::MergeTreeBackgroundExecutor<DB::DynamicRuntimeQueue>::threadFunction() ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() void* std::__thread_proxy[$ABI]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) ``` <details><summary>Raw stack trace (with frame addresses)</summary> ``` 3. __pthread_kill @ 0x00000000000969fd 4. gsignal @ 0x0000000000042476 5. __lgamma_r_finite@GLIBC_2.15 @ 0x00000000000287f3 6. /ClickHouse/src/Common/Exception.cpp:60:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x00000000145811ee 7. /ClickHouse/src/Common/Exception.cpp:109:13: DB::Exception::handleErrorCode(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x0000000014582230 8. /ClickHouse/src/Common/Exception.cpp:165:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x0000000014582699 9. /ClickHouse/src/Common/Exception.h:201:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d4c5656 10. /ClickHouse/src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d4c5118 11. /ClickHouse/src/Common/Exception.h:219:77: DB::Exception::Exception<String, String>(int, FormatStringHelperImpl<std::type_identity<String>::type, std::type_identity<String>::type>, String&&, String&&) @ 0x000000000d4c3586 12. /ClickHouse/src/Common/typeid_cast.cpp:15:11: throwBadTypeidCast(std::type_info const&, std::type_info const&) @ 0x00000000146e3722 13.0. inlined from /ClickHouse/src/Common/typeid_cast.h:22: T typeid_cast<DB::ColumnString const&, DB::IColumn const>(DB::IColumn const&) 13. /ClickHouse/src/DataTypes/Serializations/SerializationString.cpp:135:5: DB::SerializationString::serializeBinaryBulk(DB::IColumn const&, DB::WriteBuffer&, unsigned long, unsigned long) const @ 0x000000001955a1d1 14. /ClickHouse/src/DataTypes/Serializations/ISerialization.cpp:258:9: DB::ISerialization::serializeBinaryBulkWithMultipleStreams(DB::IColumn const&, unsigned long, unsigned long, DB::ISerialization::SerializeBinaryBulkSettings&, std::shared_ptr<DB::ISerialization::SerializeBinaryBulkState>&) const @ 0x0000000019428852 15. /ClickHouse/src/Storages/MergeTree/MergeTreeDataPartWriterOnDisk.cpp:733:24: DB::MergeTreeDataPartWriterOnDisk::initColumnsSubstreamsIfNeeded() @ 0x000000001e0c4607 16. /ClickHouse/src/Storages/MergeTree/MergeTreeDataPartWriterWide.cpp:328:5: DB::MergeTreeDataPartWriterWide::write(DB::Block const&, DB::PODArray<unsigned long, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, DB::Block*) @ 0x000000001e0c6ab8 17.0. inlined from /ClickHouse/src/Storages/MergeTree/MergedBlockOutputStream.cpp:479: DB::MergedBlockOutputStream::writeImpl(DB::Block const&, DB::PODArray<unsigned long, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, DB::Block*) 17. /ClickHouse/src/Storages/MergeTree/MergedBlockOutputStream.cpp:94:13: DB::MergedBlockOutputStream::write(DB::Block const&) @ 0x000000001e2aa1f9 18. /ClickHouse/src/Storages/MergeTree/MutateTask.cpp:1901:19: DB::PartMergerWriter::mutateOriginalPartAndPrepareProjections() @ 0x000000001e2bd3dc 19.0. inlined from /ClickHouse/src/Storages/MergeTree/MutateTask.cpp:1760: DB::PartMergerWriter::execute() 19. /ClickHouse/src/Storages/MergeTree/MutateTask.cpp:2295:21: DB::MutateAllPartColumnsTask::executeStep() @ 0x000000001e2d9456 20. /ClickHouse/src/Storages/MergeTree/MutateTask.cpp:3248:23: DB::MutateTask::execute() @ 0x000000001e2c1cf7 21. /ClickHouse/src/Storages/MergeTree/MutatePlainMergeTreeTask.cpp:114:34: DB::MutatePlainMergeTreeTask::executeStep() @ 0x000000001e2bb08b 22. /ClickHouse/src/Storages/MergeTree/MergeTreeBackgroundExecutor.h:74:22: DB::TaskRuntimeData::executeStep() const @ 0x000000001df86aa0 23. /ClickHouse/src/Storages/MergeTree/MergeTreeBackgroundExecutor.cpp:383:36: DB::MergeTreeBackgroundExecutor<DB::DynamicRuntimeQueue>::routine(std::shared_ptr<DB::TaskRuntimeData>) @ 0x000000001df8c415 24. /ClickHouse/src/Storages/MergeTree/MergeTreeBackgroundExecutor.cpp:446:13: DB::MergeTreeBackgroundExecutor<DB::DynamicRuntimeQueue>::threadFunction() @ 0x000000001df8e672 25.0. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:502: ? 25.1. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:754: ? 25. /ClickHouse/src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000014774dc9 26.0. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:502: ? 26.1. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:754: ? 26.2. inlined from /ClickHouse/src/Common/ThreadPool.cpp:1293: operator() 26.3. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) 26.4. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) 26.5. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) 26. /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x000000001477de4b 27.0. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:502: ? 27.1. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__functional/function.h:754: ? 27. /ClickHouse/src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x0000000014771f7c 28.0. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) 28.1. inlined from /ClickHouse/contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) 28. /ClickHouse/contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x000000001477b20e 29. start_thread @ 0x0000000000094ac3 30. __clone3 @ 0x00000000001268d0 ``` </details> Possible causes: - Mismatch between expected and actual type during serialization - Incorrect type information in typeid cast - Serialization format not compatible with data type - Invalid data being serialized leading to type mismatch - Serialization settings not properly configured for the data The stack trace appeared in the following checks: - [Stateless tests (amd_debug, distributed plan, s3 storage, parallel)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=971938b5ab0f27eb0a9ea60156c7e2d087cdee96&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_debug%2C%20distributed%20plan%2C%20s3%20storage%2C%20parallel%29&name_1=Stateless%20tests%20%28amd_debug%2C%20distributed%20plan%2C%20s3%20storage%2C%20parallel%29) - [Stateless tests (amd_debug, parallel)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107134&sha=bc64592ec4a3b06b37b2c1d432bd1b68b0c3383b&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20parallel%29&name_1=Stateless%20tests%20%28amd_debug%2C%20parallel%29) - [Stateless tests (amd_msan, WasmEdge, parallel, 2/2)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112973&sha=149a6b975b5c213d4aaa0fba85658ff78249b6e4&name_0=PR&name_1=Stateless%20tests%20%28amd_msan%2C%20WasmEdge%2C%20parallel%2C%202%2F2%29&name_1=Stateless%20tests%20%28amd_msan%2C%20WasmEdge%2C%20parallel%2C%202%2F2%29)",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113891",
        "timestamp": "2026-08-12T23:12:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "crash-ci",
          "comp-mergetree"
        ],
        "author": "robot-clickhouse-ci-2",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:113993",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Minmax index incorrectly prunes fractional DateTime64 values for integer bounds",
        "text": "### Company or project name _No response_ ### Describe what's wrong A minmax data-skipping index can change the result of a strict comparison between DateTime64 and an integer. Without skip-index pruning, `time > 0` correctly matches a value of 0.01 seconds. With a minmax index, the same row is pruned. ### Does it reproduce on the most recent release? Yes ### How to reproduce https://fiddle.clickhouse.com/66f7a197-cd87-4248-8e8c-6c44a4b85e4b ```sql CREATE TABLE t ( time DateTime64(9), INDEX idx_time time TYPE minmax GRANULARITY 1 ) ENGINE = MergeTree ORDER BY tuple() SETTINGS index_granularity = 1; INSERT INTO t VALUES (0.01::Decimal(9, 2)::DateTime64(9)); SELECT count() FROM t WHERE time > 0 SETTINGS use_skip_indexes = 0; -- 1 SELECT count() FROM t WHERE time > 0 SETTINGS force_data_skipping_indices = 'idx_time'; -- 0 EXPLAIN indexes = 1 SELECT * FROM t WHERE time > 0; ``` `EXPLAIN` shows the minmax condition as `time in [1, +Inf)`. ### Expected behavior Both queries should return 1. A data-skipping index must not change the query result. ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context `Range::shrinkToIncludedIfPossible` changes the open UInt64 bound `> 0` into the closed bound `>= 1`. That transformation is valid for an integer value domain, but not for DateTime64(9), which can contain values between 0 and 1. KeyCondition treats native integers and DateTime64 as directly comparable, so the integer bound reaches this normalization without first being converted into the DateTime64 domain.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/113993",
        "createdAt": "2026-08-08T23:59:27Z",
        "updatedAt": "2026-08-13T14:36:45Z",
        "timestamp": "2026-08-13T14:36:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "clickgap-analyzed",
          "culprit-pr-pinned"
        ],
        "author": "EmeraldShift",
        "state": "open",
        "assignees": [
          "amosbird"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114004",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Inconsistent AST formatting: `view((SELECT ...))` in table-function arguments loses subquery parentheses and cannot be parsed back (STID: 1941-1bfa)",
        "text": "🕵 Found by `AST fuzzer (amd_debug, targeted, old_compatibility)` on an unrelated PR (https://github.com/ClickHouse/ClickHouse/pull/91993, which only touches hex encoding): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=91993&sha=eae3a3cbbfb5d81bed061634d53c701167dce04a&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The fuzzer produced a query containing `view((SELECT ...))` as an argument of a comparison function nested inside `file(...)` table-function arguments. Formatting the AST drops the parentheses around the `view` subquery argument (`view((SELECT ...))` → `view(SELECT ...)`), and the formatted text cannot be parsed back, so the format-consistency check fails with a logical error (`abortOnFailedAssertion` in the debug build): ``` Logical error: 'Inconsistent AST formatting: the query: DESCRIBE TABLE file(concat(xor(if(indexHint(and(greaterOrEquals(alias5764._time, toLowCardinality(toNullable(1048575)) + number), lessOrEquals(alias5764._time, view((SELECT toUInt16OrDefault(65536, equals(minus(materialize(NULL, arrayElement(b)), greater(lowCardinalityIndices('\\\\\\\\', toLowCardinality(NULL)), -1)), toInt8OrZero(2147483646)))))))), a <= toInt256OrDefault(-2147483647), materialize(isNullable(NULL))), and(b >= 257, 1.1754943508222875e-38 > b)), currentDatabase(), toFixedString(toNullable('_04513.orc'), assumeNotNull(256)))) SAMPLE 100 / 5 SETTINGS input_format_orc_use_fast_decoder = 0 cannot parse query back from DESCRIBE TABLE file(concat(xor(if(indexHint(and(greaterOrEquals(alias5764._time, toLowCardinality(toNullable(1048575)) + number), lessOrEquals(alias5764._time, view(SELECT toUInt16OrDefault(65536, equals(minus(materialize(NULL, arrayElement(b)), greater(lowCardinalityIndices('\\\\\\\\', toLowCardinality(NULL)), -1)), toInt8OrZero(2147483646))))))), a <= toInt256OrDefault(-2147483647), materialize(isNullable(NULL))), and(b >= 257, 1.1754943508222875e-38 > b)), currentDatabase(), toFixedString(toNullable('_04513.orc'), assumeNotNull(256)))) SAMPLE 100 / 5 SETTINGS input_format_orc_use_fast_decoder = 0'. ``` The two texts differ only in the parentheses around the `view` subquery: the original has `view((SELECT ...))`, the re-formatted text has `view(SELECT ...)`. On a recent master-based **release** build (`clickhouse local`) the asymmetry is observable in the opposite direction — the parenthesized form is the one that does not parse inside table-function arguments: ```sql -- parses fine (special `view` handling in expression context): SELECT lessOrEquals(t, view((SELECT 1))); -- parses fine (unparenthesized subquery): SELECT * FROM file(if(indexHint(lessOrEquals(t, view(SELECT 1))), 'a', 'b')); -- SYNTAX_ERROR (parenthesized subquery inside table-function arguments): SELECT * FROM file(if(indexHint(lessOrEquals(t, view((SELECT 1)))), 'a', 'b')); ``` So whether `view((SELECT ...))` / `view(SELECT ...)` parses depends on context (plain expression vs. table-function argument) and on build/compatibility configuration, while the formatter always emits the unparenthesized form. Either the parser should accept both forms in all contexts where `view` is accepted at all, or the formatter should print the form that is guaranteed to parse back in the surrounding context. Previous distinct manifestations of this dedup bucket (STID: 1941-1bfa) were fixed individually: #109501 (back-quoted numeric type name, fixed by #109720), #106850, #106358, #100131. This `view` parenthesization case is a new one. CIDB shows this check failing on 5 unrelated PRs in the last 30 days (112921, 110886, 113533, 91993, 111061) and never on master — expected, since the AST fuzzer only runs per-PR with random queries.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114004",
        "createdAt": "2026-08-09T02:45:03Z",
        "updatedAt": "2026-08-13T15:47:35Z",
        "timestamp": "2026-08-13T15:47:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "fuzz",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114026",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "`rewrite_in_to_join`: `PREWHERE x IN (subquery)` throws `Unknown function exists` while the `WHERE` spelling works (also hit via `make_distributed_plan`)",
        "text": "**Describe what's wrong** With `rewrite_in_to_join = 1`, any non-constant `IN (subquery)` predicate placed in `PREWHERE` makes the query fail with `Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION)`. The same predicate in `WHERE` works and returns the correct result. `NOT IN`, tuple `IN`, and the same query with an explicit `JOIN` fail the same way; `GLOBAL IN` and scalar subqueries (`PREWHERE s = (SELECT ...)`) are unaffected. This is hit at default settings by users of `make_distributed_plan`: since `allow_experimental_correlated_subqueries` defaults to `1`, `adjustSettingsForMakeDistributedPlan` (`src/Core/SettingsQuirks.cpp`) silently force-enables `rewrite_in_to_join`, so `SELECT ... FROM distributed_table PREWHERE x IN (subquery) SETTINGS make_distributed_plan = 1` fails with the same bogus `Unknown function exists` instead of a declared unsupported-feature error (the `WHERE` spelling of that distributed query gets the honest `Code: 48. Correlated subqueries are not supported with remote tables`). **How to reproduce** Reproduces on current master (26.8.1.1008) and on every version tried back to 26.6, with `clickhouse local` (20/20 deterministic; the `WHERE` control arm is 20/20 correct): ```sql CREATE TABLE t (k UInt64, s String) ENGINE = MergeTree ORDER BY k; INSERT INTO t SELECT number, toString(number % 2) FROM numbers(1000); SELECT count() FROM t PREWHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1; -- Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION) SELECT count() FROM t WHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1; -- 500 (correct) ``` Required: `rewrite_in_to_join = 1` (or `make_distributed_plan = 1`, which force-enables it), the analyzer (`enable_analyzer = 1`, the default; the old analyzer is unaffected), a non-constant left-hand side, and the predicate in `PREWHERE`. **Mechanism** (from the stack trace on 26.8.1.1008) 1. Under `rewrite_in_to_join`, the analyzer rewrites `x IN (subquery)` into a special `exists(...)` `FunctionNode` (`src/Analyzer/Resolve/resolveFunction.cpp`, \"Replace IN (subquery)\"). `exists` is resolved by a dedicated special-function path — it is not registered in `FunctionFactory`. 2. `PREWHERE` resolution then clones the predicate tree and runs `ReplaceColumnsVisitor` to undo JOIN-changed column types (`src/Analyzer/Resolve/QueryAnalyzer.cpp`, \"Expressions in PREWHERE with JOIN should not change their type\"). That visitor calls `rerunFunctionResolve` on every resolved `FunctionNode` it visits. 3. `rerunFunctionResolve` (`src/Analyzer/Utils.cpp`) resolves ordinary functions through `FunctionFactory::instance().get(name, context)`. It already special-cases `grouping` (also resolved outside the factory) but not `exists`, so the lookup throws `UNKNOWN_FUNCTION`: ``` 4. src/Functions/FunctionFactory.cpp:89: DB::FunctionFactory::getImpl(String const&, ...) 5. src/Functions/FunctionFactory.cpp:108: DB::rerunFunctionResolve(DB::FunctionNode*, ...) 6. src/Analyzer/InDepthQueryTreeVisitor.h:62: DB::QueryAnalyzer::resolveQuery(...) ``` **Expected behavior** The query returns the same result as its `WHERE` spelling (or, where the rewritten form is genuinely unsupported, a declared `SUPPORT_IS_DISABLED`/`NOT_IMPLEMENTED` error naming the feature — not `Unknown function exists`). Related: https://github.com/ClickHouse/ClickHouse/issues/109476 Related: https://github.com/ClickHouse/ClickHouse/issues/109326 Related: https://github.com/ClickHouse/ClickHouse/issues/102630 Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114026",
        "createdAt": "2026-08-09T10:21:45Z",
        "updatedAt": "2026-08-13T16:04:02Z",
        "timestamp": "2026-08-13T16:04:02Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug",
          "comp-joins",
          "comp-query-optimizer",
          "comp-query-analyzer"
        ],
        "author": "zlareb1",
        "state": "closed",
        "assignees": [
          "KochetovNicolai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114105",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Unconditional `std::adjacent_find` offsets scan in the `ColumnArray` constructor costs 15-31% on array-heavy queries in release builds",
        "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ ### Describe what's wrong Every query that builds arrays row-by-row gets measurably slower. ClickHouse's own Performance Comparison job reports `tests/performance/array_join.xml` queries #0-#5 as `slower` on all three of this PR's benchmarked commits, by +14.6% to +30.9%, while the same queries move by +0.3% on master over 1190 measurements. **Root cause:** src/Columns/ColumnArray.cpp:75 adds an O(number_of_rows) sequential scan with an early-exit predicate to the two-argument `ColumnArray` constructor, with no `#ifdef DEBUG_OR_SANITIZER_BUILD` / `chassert` guard, so it executes in release builds on the hottest column-construction path in the server. The comment on the line above it (src/Columns/ColumnArray.cpp:74) asserts 'The scan is linear, but it is cheap compared to filling the offsets in the first place' — the measurement contradicts that claim: filling the offsets in `FunctionArray` is a single fused add-and-store loop, so re-reading them with a non-vectorisable early-exit `adjacent_find` roughly doubles the offsets traffic and adds ~25 ms to a 97 ms query. **Why we believe this is a bug:** `FunctionArray::executeImpl` (src/Functions/array/array.cpp:119) -> `ColumnArray::create(out_data_column, out_offsets_column)` -> `ColumnArray::ColumnArray(MutableColumnPtr &&, MutableColumnPtr &&)` (src/Columns/ColumnArray.cpp:44) -> `std::adjacent_find` over the whole offsets array (src/Columns/ColumnArray.cpp:75). The same path is taken by all 265 `ColumnArray::create` call sites that pass an offsets column, once per block per array-producing expression. **Affected locations:** - `src/Columns/ColumnArray.cpp:75` — unconditional `std::adjacent_find` monotonicity scan added to the two-argument constructor - `src/Columns/ColumnArray.cpp:74` — comment claiming the scan is cheap, contradicted by the Performance Comparison measurement - `src/Functions/array/array.cpp:119` — `array` now constructs through the validating two-argument constructor once per block; this is the function used by `tests/performance/array_join.xml` **Impact:** A 15-31% slowdown on `ARRAY JOIN` over row-wise constructed arrays, and a proportional (smaller for wide arrays, larger for single-element arrays) tax on every other array-producing expression in the server, since the scan cost depends on the number of rows and not on the number of elements. Workloads that build one-element arrays per row — the `[x]` idiom, `Nested` writes, `Map` construction via `ColumnMap`, Arrow/Parquet/ORC array decoding — pay the most. ## Assumptions The bot recorded these claims it could not verify directly from source. A maintainer ✅ confirms; ❌ flags a wrong premise (the bot should rework or close the finding). - [ ] **The regression is attributable to the constructor scan rather than to the producer rewrites themselves** - *Why unverifiable:* Only the PR build is available in this environment, so no local A/B against master was possible - *Falsifiable test:* Build the head sha with the `std::adjacent_find` block at src/Columns/ColumnArray.cpp:75-79 deleted and re-run `tests/performance/array_join.xml`; the +15-25% deltas should disappear while the producer rewrites remain in place. Reading the diff supports the attribution: the `array` rewrite performs the same allocations and the same offsets fill as before, so the constructor validation is the only added work on that path. - [ ] **`std::adjacent_find` with an early-exit predicate is not auto-vectorised by the release toolchain** - *Why unverifiable:* No disassembly of the release binary was taken - *Falsifiable test:* objdump the `ColumnArray::ColumnArray(MutableColumnPtr&&, MutableColumnPtr&&)` symbol in a release build and check for SIMD compares; the ~1.25 ns/element implied by the measured delta over the ~20M offsets scanned per query is consistent with a scalar loop. ### Does it reproduce on most recent release? Likely yes — see testability note in additional context. ### How to reproduce ```sql 1. Reproduce the CI measurement: curl -s 'https://play.clickhouse.com/?user=play' --data-binary \"SELECT commit_sha, replaceRegexpOne(test_name,'::(new|old)$','') AS q, maxIf(test_duration_ms, endsWith(test_name,'::old')) AS old_ms, maxIf(test_duration_ms, endsWith(test_name,'::new')) AS new_ms, round((new_ms-old_ms)/old_ms*100,1) AS pct FROM default.checks WHERE pull_request_number=112504 AND check_name LIKE 'Performance%' AND test_name LIKE 'array_join #%' GROUP BY commit_sha, q ORDER BY commit_sha, q FORMAT PrettyCompact\" 2. Run the same query with pull_request_number=0 for the master control and with pull_request_number=109212 for the natural A/B. 3. Local A/B: build the head sha twice, once unmodified and once with src/Columns/ColumnArray.cpp:71-79 removed, and run `tests/performance/array_join.xml` against both. ``` ### Expected behavior ``` No significant difference, i.e. the same profile as the master control. Positive control, same test, same framework, `pull_request_number = 0` over 1190 measurements in the last 10 days: #0 46.2 -> 46.9 ms, #1 110.9 -> 111.9 ms, #2 195.3 -> 195.8 ms, #3 195.3 -> 195.9 ms, #4 235.2 -> 235.6 ms, #5 235.8 -> 236.4 ms, #6 6035.9 -> 6028.9 ms; the `slower` verdict fires on 4-10 of 1190 measurements (0.3-0.8%) per query. ``` ### Error message and/or stacktrace ``` ClickHouse Performance Comparison (arm_release, master_head), from CIDB `default.checks`, PR 112504, old = master reference build, new = PR build: sha bf5adf14 (2026-08-04): #0 40->48 ms (+20.0%), #1 96->123 ms (+28.1%), #2 205->246 ms (+20.0%), #3 210->251 ms (+19.5%), #4 239->278 ms (+16.3%), #5 239->279 ms (+16.7%), #6 5335->5326 ms (-0.2%) sha 32002174 (2026-08-05): #0 39->47 ms (+20.5%), #1 ``` ### Additional context **Open risks:** - The same scan is now on the Arrow/Parquet/ORC/Puffin array decode paths (src/Processors/Formats/Impl/Parquet/Reader.cpp:3187, src/Processors/Formats/Impl/ArrowIPC/RecordBatchDecoder.cpp:664 and :848, src/Processors/Formats/Impl/NativeORCBlockInputFormat.cpp:2704) — no perf test in `tests/performance/` covers array-typed Parquet/Arrow ingestion, so a similar regression there would be invisible to CI. - For untrusted inputs (a crafted Arrow/Parquet file with a non-monotonic offsets buffer) the new check converts an out-of-bounds read into a `LOGICAL_ERROR`, which aborts in debug and sanitizer builds. That is an improvement over the out-of-bounds read, but the error code for a user-supplied file should arguably be `INCORRECT_DATA` rather than `LOGICAL_ERROR`. **Suggested fix:** Keep the invariant but stop paying for it in release builds: wrap the `std::adjacent_find` block at src/Columns/ColumnArray.cpp:71-79 in `#ifdef DEBUG_OR_SANITIZER_BUILD` (or express it as a `chassert`-style debug-only validation, which is how ClickHouse enforces similar structural invariants). The O(1) `data->size() != last_offset` check at src/Columns/ColumnArray.cpp:62-68 can stay unconditional — it is the part that actually closes the reported hole and costs nothing. If a release-build monotonicity check is considered mandatory, it should at least be branchless/vectorisable (compare the whole array against a shifted copy and reduce) rather than `adjacent_find`. **Analysis details:** Confidence HIGH | Severity P2 | Testability: `THEORETICAL` Found during automated review of [PR #112504](https://github.com/ClickHouse/ClickHouse/pull/112504). --- _ClickGapAI · Confidence: HIGH · Severity: P2 · Finding: `h_pr112504_001`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #114469 - Merged into: `26.8.1.1291` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114105",
        "timestamp": "2026-08-12T23:50:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "comp-regular-function",
          "comp-data-types"
        ],
        "author": "clickgapai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114169",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "ClickHouse writes an ORC file whose schema it cannot infer: `Variant(String, Int128)`",
        "text": "### Describe the bug Writing a `Variant` whose branches map to ORC `string` and ORC `binary` produces a valid ORC file, but schema inference on that file then fails, so `SELECT * FROM file('x.orc', ORC)` does not work on a file ClickHouse just wrote. The data itself is fine and is returned correctly if the structure is supplied explicitly — this is an inference/round-trip bug, not data loss. `Variant` → ORC union writing was added in #110085, which included a guard (`e17144f9682`, \"Address review: reject a `Variant` whose branches collapse to the same ORC type\"). The guard is keyed on the **ORC** type, while the reader's equivalent check is keyed on the **ClickHouse** type the branch parses back to. ORC `STRING` and ORC `BINARY` both parse back to `String`, so a union with one of each passes the writer and is rejected by inference. ### How to reproduce ClickHouse 26.8.1.1087 (master). No settings needed. ```bash clickhouse local -q \"SELECT NULL::Variant(String, Int128) FORMAT ORC\" \\ | clickhouse local --input-format ORC -q \"SELECT * FROM table\" ``` or equivalently: ```sql SELECT NULL::Variant(String, Int128) INTO OUTFILE 'q.orc' FORMAT ORC; SELECT * FROM file('q.orc', ORC); ``` ### Error message and/or stacktrace ``` Code: 636. DB::Exception: The table structure cannot be extracted from a ORC format file. Error: Code: 50. DB::Exception: ORC union type 'uniontype<binary,string>' has branches with identical types, which is not supported. (UNKNOWN_TYPE) (version 26.8.1.1087 (official build)). You can specify the structure manually. (CANNOT_EXTRACT_TABLE_STRUCTURE) ``` Supplying the structure works and returns the right values, so the file is sound: ```bash clickhouse local -q \"SELECT multiIf(number%3=0, 'abc'::Variant(String, Int128), number%3=1, (-170141183460469231731687303715884105728::Int128)::Variant(String, Int128), NULL::Variant(String, Int128)) AS c FROM numbers(3) FORMAT ORC\" \\ | clickhouse local --input-format ORC --structure \"c Variant(String, Int128)\" \\ -q \"SELECT c, variantType(c) FROM table\" ``` ``` abc String -170141183460469231731687303715884105728 Int128 \\N None ``` ### Affected combinations Two groups of ClickHouse types are written as ORC `BINARY` and `STRING` respectively, and inference maps **both** back to `String`. Any cross pairing escapes the guard: - → ORC `BINARY`: `Int128`, `UInt128`, `Int256`, `UInt256`, `Decimal256`, `IPv6` - → ORC `STRING`: `String`, `FixedString(N)` (when `output_format_orc_string_as_string = 1`) | Variant | union written | result | |---|---|---| | `Variant(String, Int128)` | `uniontype<binary,string>` | **written, not inferrable** | | `Variant(String, UInt128)` | `uniontype<string,binary>` | **written, not inferrable** | | `Variant(String, Int256)` | `uniontype<binary,string>` | **written, not inferrable** | | `Variant(String, UInt256)` | `uniontype<string,binary>` | **written, not inferrable** | | `Variant(String, Decimal256(0))` | `uniontype<binary,string>` | **written, not inferrable** | | `Variant(FixedString(4), Int128)` | `uniontype<string,binary>` | **written, not inferrable** | | `Variant(String, IPv6)` | `uniontype<binary,string>` | **written, not inferrable** | | `Variant(IPv6, FixedString(8))` | `uniontype<string,binary>` | **written, not inferrable** | | `Variant(IPv6, Int128)`, `Variant(String, FixedString(8))` | — | correctly rejected by the guard | | `Variant(IPv4, Int32)`, `Variant(Int32, UInt32)` | — | correctly rejected by the guard | | `Variant(Int64, String)`, `Variant(IPv4, Date32)`, ... | | fine | Only reachable at the **default** `output_format_orc_string_as_string = 1`. With `= 0` the writer maps `String` to `binary` as well, both branches then share one ORC type, and the existing guard correctly rejects the write. ### Root cause Writer — dedups on the ORC type (`ORCBlockOutputFormat.cpp`, `case TypeIndex::Variant`): ```cpp std::unordered_map<String, String> variant_by_orc_type; for (const auto & nested_type : variant_type.getVariants()) { auto child_type = getORCType(nested_type, ...); auto [it, inserted] = variant_by_orc_type.emplace(child_type->toString(), nested_type->getName()); if (!inserted) throw Exception(ErrorCodes::ILLEGAL_COLUMN, \"... both written as ORC type '{}' ...\"); ``` Inference — dedups on the ClickHouse type (`NativeORCBlockInputFormat.cpp:288`): ```cpp if (!seen_type_names.insert(recursiveRemoveLowCardinality(parsed_type)->getName()).second) throw Exception(ErrorCodes::UNKNOWN_TYPE, \"ORC union type '{}' has branches with identical types, which is not supported\", ...); ``` and the mapping that collapses them (`NativeORCBlockInputFormat.cpp:326`): ```cpp case orc::TypeKind::CHAR: case orc::TypeKind::VARCHAR: case orc::TypeKind::BINARY: case orc::TypeKind::STRING: { DataTypePtr type; if (orc_type->getKind() == orc::TypeKind::CHAR) type = std::make_shared<DataTypeFixedString>(orc_type->getMaximumLength()); else type = std::make_shared<DataTypeString>(); // BINARY and STRING both land here ``` `binary` != `string` as ORC types, so the writer's `emplace` succeeds; both become `String` during inference, so the reader's `insert` fails. The sibling check on the read path (`NativeORCBlockInputFormat.cpp:2441`) does not fire when a structure is given, because there `branch_types` are the caller's hints, which are distinct — which is why the explicit-structure read above works. ### Expected behaviour The writer's guard should use the same key inference will use — the ClickHouse type each branch reads back as — rather than the ORC type. `Variant(String, Int128)` would then be rejected up front with the existing clear `ILLEGAL_COLUMN` message, instead of producing a file that only reads back if you already know its schema. (Reader-side alternatives, e.g. mapping ORC `BINARY` to something other than `String`, would change the type of every existing ORC `binary` column, so the writer-side fix looks preferable.) cc @alexey-milovidov",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114169",
        "timestamp": "2026-08-12T20:35:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "bug"
        ],
        "author": "PedroTadim",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114286",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "date_time_overflow_behavior is ignored when reading out-of-range Date32 values via IcebergS3 table engine (works with icebergS3() table function)",
        "text": "### Company or project name _No response_ ### Describe the unexpected behaviour Reading an Iceberg table that contains Parquet DATE values outside the ClickHouse Date32 range ([-25567, 120530] / 1900-01-01..2299-12-31) throws VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE when querying through the IcebergS3 table engine, even with: SETTINGS date_time_overflow_behavior = 'saturate' The same data, same settings, and same named collection work when reading via the icebergS3() table function. Server default is already date_time_overflow_behavior = ignore, but the table-engine path still throws. ### Which ClickHouse versions are affected? Reproduced on ClickHouse 26.7.2.59 (official build). ### How to reproduce Iceberg table on S3 with at least one Parquet DATE column containing an out-of-range day number (example: -645804, -639653). Mount as table engine: ``` CREATE TABLE iceberg.example ENGINE = IcebergS3( s3_iceberg_bucket, url = 'https://s3.example/lakehouse/db/table', format = 'Parquet' ) SETTINGS allow_dynamic_metadata_for_data_lakes = 1; ``` Fails: ``` SELECT * FROM iceberg.example LIMIT 1 SETTINGS iceberg_timestamp_ms = <snapshot_ts>, date_time_overflow_behavior = 'saturate'; ``` Error (example): Code: 321. DB::Exception: Input value -639653 is out of allowed Date32 range, which is [-25567, 120530]: read stage: ColumnData: column: value_date: (in file/uri .../data.parquet): While executing ParquetV3BlockInputFormat: While executing ReadFromObjectStorage. (VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE) Works: ``` SELECT min(start_date) FROM icebergS3( s3_iceberg_bucket, url = 'https://s3.example/lakehouse/db/table', format = 'Parquet' ) SETTINGS iceberg_timestamp_ms = <snapshot_ts>, date_time_overflow_behavior = 'saturate'; ``` Also observed: input_format_parquet_use_native_reader_v3 = 1 by default; stack mentions ParquetV3BlockInputFormat. Disabling V3 + saturate in dbt query_settings on SELECT from the table engine still failed in our runs. ### Expected behavior Out-of-range dates should be saturated to 1900-01-01 / 2299-12-31 (or ignored, depending on the setting) for both: ``` ENGINE = IcebergS3 icebergS3() table function ``` Behavior should be consistent. ### Error message and/or stacktrace ``` Code: 321. DB::Exception: Input value -639653 is out of allowed Date32 range, which is [-25567, 120530]: read stage: ColumnData: column: value_date: While executing ParquetV3BlockInputFormat: While executing ReadFromObjectStorage. (VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE) version 26.7.2.59 (official build) ``` ### Related issues and pull requests _No response_ ### Additional context Iceberg type date is mapped to ClickHouse Date32. Source data legitimately contains sentinel / dirty dates (e.g. year 0005, 0202, 9999, 5005) that already exist in a MergeTree copy of the same dataset; MergeTree SELECT works, IcebergS3 table engine SELECT throws. Related history: [#51402](https://github.com/ClickHouse/ClickHouse/issues/51402) / [#55696](https://github.com/ClickHouse/ClickHouse/pull/55696) introduced date_time_overflow_behavior for Parquet Date32 overflow; this looks like a remaining gap on the Iceberg table engine read path (or settings not applied there), while the table function path respects the setting. Cluster settings of interest: date_time_overflow_behavior = ignore (default) input_format_parquet_use_native_reader_v3 = 1 Workaround: read via icebergS3(...) + SETTINGS date_time_overflow_behavior = 'saturate' instead of SELECT from ENGINE = IcebergS3.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114286",
        "createdAt": "2026-08-11T08:41:13Z",
        "updatedAt": "2026-08-13T09:38:02Z",
        "timestamp": "2026-08-13T09:38:02Z",
        "metrics": {
          "reactions": 1,
          "comments": 2
        },
        "labels": [
          "unexpected behaviour"
        ],
        "author": "alexsubota",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114404",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Inlined ALIAS body adds constants to the shipped WITH FILL header; parallel replicas throw",
        "text": "_Found via ClickGap automated review — close or comment if this is wrong._ ### Describe what's wrong With parallel replicas enabled, `SELECT k, a_v FROM t ORDER BY k WITH FILL ... INTERPOLATE (...)` where `a_v` is an `ALIAS` column whose body is an expression fails with `Code: 20. Number of columns doesn't match (source: 5 and result: 4). (NUMBER_OF_COLUMNS_DOESNT_MATCH)`. The same query returns rows correctly without parallel replicas, and returns rows correctly under parallel replicas if the alias body is written inline in the `SELECT` list instead of declared as an `ALIAS` column. **Root cause:** `inlineAliasColumns` at src/Interpreters/ClusterProxy/executeQuery.cpp:1048 changes the structure of the shipped query tree, but the initiator's `expected_header` is still computed from the un-inlined tree; when the shipped tree's `WithMergeableState` header gains columns (any plan step that keeps ActionsDAG intermediates, e.g. `Filling`), no reconciliation path can absorb it. <details> <summary>Analysis details (evidence, affected locations, impact)</summary> **Why we believe this is a bug:** `PlannerJoinTree::buildQueryPlanForTableExpression` (src/Planner/PlannerJoinTree.cpp:2048) calls `ClusterProxy::executeQueryWithParallelReplicas`, which since this PR runs `inlineAliasColumns` on the shipped tree (src/Interpreters/ClusterProxy/executeQuery.cpp:1048) and derives the replica `header` from that inlined tree (executeQuery.cpp:1052). Back in `buildQueryPlanForTableExpression`, `expected_header` is re-planned from the ORIGINAL, un-inlined `select_query_info.query_tree` (PlannerJoinTree.cpp:2262-2268). For `ORDER BY ... WITH FILL ... INTERPOLATE` the `WithMergeableState` header is the `Filling` step header, which keeps every ActionsDAG output including the constants of the inlined body — so the inlined tree yields one more column (`2_UInt8` for a `v * 2` body) than the un-inlined tree. `buildShardCollapseFanOut` bails out because it only handles a SMALLER shard header (src/Storages/buildQueryTreeForShard.cpp:1313), and the positional `ActionsDAG::makeConvertingActions` at PlannerJoinTree.cpp:2293 then throws. **Affected locations:** - `src/Interpreters/ClusterProxy/executeQuery.cpp:1048` — new `inlineAliasColumns` call on the shipped tree; `header` at line 1052 is derived from it - `src/Planner/findParallelReplicasQuery.cpp:622` — sibling new `inlineAliasColumns` call; `initial_header` at line 615 is taken from the un-inlined tree and reconciled positionally at line 650 - `src/Planner/PlannerJoinTree.cpp:2293` — positional `makeConvertingActions` that throws NUMBER_OF_COLUMNS_DOESNT_MATCH - `src/Storages/buildQueryTreeForShard.cpp:1313` — `buildShardCollapseFanOut` returns {} when the shard header is not strictly smaller, so a LARGER shard header is unhandled **Impact:** Any query combining an `ALIAS` column whose body is an expression with `ORDER BY ... WITH FILL ... INTERPOLATE` is rejected once `enable_parallel_replicas = 1`. Deterministic, reproduces with both `parallel_replicas_local_plan = 0` and `= 1`. It also masks the correct user error: `INTERPOLATE (k AS k)` (an ORDER BY column as interpolate target) reports `NUMBER_OF_COLUMNS_DOESNT_MATCH` instead of `INVALID_WITH_FILL_EXPRESSION`. </details> ## Assumptions _Unverified assumptions — tick to confirm, comment to refute:_ - [ ] **Before this PR the same query succeeded on the parallel-replicas path** - *Why unverifiable:* no pre-PR binary is available in this environment to run the query against - *Falsifiable test:* Build master (without this PR) and run the repro; expect `0 10 / 2 10 / 4 14 / 6 14 / 8 18`. Pre-PR `executeQueryWithParallelReplicas` shipped the un-inlined tree, so `header` and `expected_header` came from the same query node and matched structurally. ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `48b91073fbabef`). ### How to reproduce ```sql CREATE TABLE t (k UInt32, v Int64, a_v Int64 ALIAS v * 2) ENGINE = MergeTree ORDER BY k; INSERT INTO t VALUES (0,5),(4,7),(8,9); then run SELECT k, a_v FROM t ORDER BY k WITH FILL FROM 0 TO 10 STEP 2 INTERPOLATE (a_v AS a_v) SETTINGS enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, automatic_parallel_replicas_mode = 0, serialize_query_plan = 0; ``` ### Expected behavior ``` 0 10 2 10 4 14 6 14 8 18 0 10 2 10 4 14 6 14 8 18 ``` ### Error message and/or stacktrace ``` 0 10 2 10 4 14 6 14 8 18 Received exception from server (version 26.8.1): Code: 20. DB::Exception: Received from 127.0.0.1:19020. DB::Exception: Number of columns doesn't match (source: 5 and result: 4). (NUMBER_OF_COLUMNS_DOESNT_MATCH) ``` ### Additional context **Open risks:** - The same repro over a 2-shard `Distributed` table fails identically. That path has always inlined, so it is likely broken on master too and is out of scope for this PR — but a fix should cover both, since they share `buildQueryTreeForShard`. - Only `WITH FILL`/`INTERPOLATE` was found to leak ActionsDAG intermediates into the `WithMergeableState` header. Other steps with the same property would fail the same way; not audited exhaustively. **Suggested fix:** Compute the initiator-side `expected_header` from the SAME inlined tree that is shipped (or run `inlineAliasColumns` on the tree used for `expected_header` in `PlannerJoinTree::buildQueryPlanForTableExpression`), instead of relying on the shipped and expected headers happening to agree. Alternatively extend `buildShardCollapseFanOut` to drop shard columns the initiator does not expect, not only to fan out missing ones. Found during automated review of [PR #107700](https://github.com/ClickHouse/ClickHouse/pull/107700). --- _ClickGapAI · Severity: P2 · Finding: `h_pr107700_001`_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114404",
        "createdAt": "2026-08-12T00:43:39Z",
        "updatedAt": "2026-08-13T12:21:58Z",
        "timestamp": "2026-08-13T12:21:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "bug",
          "comp-query-analyzer",
          "comp-parallel-replicas"
        ],
        "author": "clickgapai",
        "state": "open",
        "assignees": [
          "yakov-olkhovskiy"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114406",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Lossy codec (SZ3) on a sorting-key column silently produces mis-sorted parts after a merge (`Sort order of blocks violated` in debug)",
        "text": "🕵 A lossy codec is accepted on a column that is part of the sorting key. Because the loss is applied on every write, the values that come back from a part are not the values the merge sorted, so a merged part can end up out of order. In a debug build the merge aborts with `Sort order of blocks violated`; in a release build nothing is reported and the part - together with its primary key index - is silently mis-sorted. **Reproduction** (release build, `clickhouse local` is enough): ```sql SET allow_experimental_codecs = 1; CREATE TABLE t (i Float64 CODEC(SZ3('ALGO_INTERP_LORENZO', 'REL', 0.01)), f Int64) ENGINE = MergeTree ORDER BY i SETTINGS min_bytes_for_wide_part = 0; INSERT INTO t SELECT number / 80, number FROM numbers(200000); INSERT INTO t SELECT number / 80 + 0.5, number FROM numbers(200000); OPTIMIZE TABLE t FINAL; SELECT countIf(i < prev) AS descending_pairs, min(i - prev) FROM (SELECT i, lagInFrame(i) OVER (ORDER BY _part_offset) AS prev FROM t) WHERE prev > 0; ``` ``` ┌─descending_pairs─┬──────min(minus(i, prev))─┐ │ 1 │ -0.2436889648437699 │ └──────────────────┴──────────────────────────┘ ``` A single `INSERT` produces a sorted part here; the descending pair appears only after the merge. **In CI** this shows up as the AST fuzzer `Logical error: 'Sort order of blocks violated for column number 0, left: Float64_12.415000000000006, right: Float64_12.168837890625007. Chunk 25, rows read 199842.'` in `CheckSortedTransform` during a background merge. The fuzzer reaches it by rewriting a table's sorting-key column to `Float64 CODEC(SZ3('ALGO_INTERP_LORENZO', 'REL', 0.01))`, and the reported neighbours differ by ~2 %, which is the codec's relative error bound rather than any data property. The signature has hit at least 10 CI runs on unrelated pull requests in the last 30 days, including `master`. Suggested direction: reject a lossy codec (`SZ3`, and any future lossy codec) on a column used in the sorting key, primary key or partition key, the same way other structurally unsafe codecs are rejected at DDL time. Silently producing a mis-sorted part is the worst outcome, because index analysis then skips rows that do match. Related: https://github.com/ClickHouse/ClickHouse/issues/111139 Related: https://github.com/ClickHouse/ClickHouse/pull/113575",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114406",
        "createdAt": "2026-08-12T00:56:46Z",
        "updatedAt": "2026-08-13T06:17:45Z",
        "timestamp": "2026-08-13T06:17:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "clickgap-analyzed",
          "culprit-pr-pinned"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114455",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Function `resetSerialID` to reset/remove a `generateSerialID` series",
        "text": "### Company or project name _No response_ ### Use case `generateSerialID(series_identifier)` (introduced in 25.x) creates a named auto-increment counter whose state is persisted as a znode in [Zoo]Keeper under series_keeper_path. There is currently no SQL-level way to reset or delete a series once it has been created — the state is effectively permanent for the lifetime of the Keeper data. This matters in particular for: 1. **Multi-tenant applications:** a common pattern is one series per tenant (e.g. generateSerialID('serial_<tenant_uuid>')) to produce per-tenant sequential IDs. When a tenant is deleted, its series znode remains forever. Over time this accumulates orphaned znodes and consumes the max_autoincrement_series budget with dead series. 2. **ClickHouse Cloud / managed environments**: users have no access to Keeper at all, so they cannot clean up series state out-of-band the way a self-hosted user could with zkCli/keeper-client. Whatever series they create is permanent unless support intervenes manually. 3. **Testing / re-ingestion workflows:** dropping and recreating a table does not reset its associated series, which is surprising — reloading data into a fresh table continues numbering from the old counter, with no way to start from 0 again under the same series name. ### Describe the solution you'd like A function or statement to delete a series and its Keeper state, gated by an appropriate grant, e.g.: ``` sql SELECT resetSerialID('my_series'); -- resets counter to 0 (or deletes the znode) ``` or, arguably more idiomatic as a DDL-like/system operation rather than a function with side effects: ``` sql SYSTEM DROP SERIAL ID 'my_series'; -- or SYSTEM RESET SERIAL ID 'my_series' [TO <value>]; ``` ### Describe alternatives you've considered - Manual znode removal via keeper-client — works for self-hosted, impossible for ClickHouse Cloud users. - Using a new series name after tenant deletion / reload — leaks znodes and consumes max_autoincrement_series indefinitely. ### Additional context Docs: https://clickhouse.com/docs/reference/functions/regular-functions/other-functions#generateSerialID",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114455",
        "createdAt": "2026-08-12T09:37:29Z",
        "updatedAt": "2026-08-13T03:50:56Z",
        "timestamp": "2026-08-13T03:50:56Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "feature",
          "easy task"
        ],
        "author": "Yonatan-Dolan",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114481",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Dropping a nested column group bypasses every ALTER protection",
        "text": "### Company or project name _No response_ ### Describe what's wrong when shared nested offsets are enabled a name in `DROP COLUMN <name>` that is not a column of the table denotes the whole group of columns `<name>.*` and the drop removes all of them The stages of an `ALTER` disagree about this meaning: validation and the mutation stage treat the name as the group, while the metadata update and every protective check (dependent views, unfinished mutations, key columns, dependencies) compare column names exactly and do not see the group. Each mismatch below is one observable consequence ### Does it reproduce on the most recent release? Yes ### How to reproduce ## 1. `DROP COLUMN IF EXISTS <group>` silently destroys the group's data `ALTER` reports OK, `DESCRIBE` still shows `n.a` and `n.b` and the `SELECT` returns `0 0 100`. The columns stay in the table while their values are gone ```sql CREATE TABLE t (`n.a` UInt64, `n.b` UInt64, x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t VALUES (1, 10, 100); ALTER TABLE t DROP COLUMN IF EXISTS n; SELECT * FROM t; -- 0 0 100 ``` Fiddle: https://fiddle.clickhouse.com/7de96186-e49d-45c1-b8af-a56803993cd0 The same split also swallows pending updates: ```sql CREATE TABLE t (`n.a` UInt64, x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t VALUES (1, 1); ALTER TABLE t UPDATE `n.a` = 2 WHERE 1 SETTINGS mutations_sync = 0; ALTER TABLE t DROP COLUMN IF EXISTS n; SELECT * FROM t; -- 0 1 ``` ## 2. A dependent materialized view does not protect the group's columns ```sql CREATE TABLE src (`n.a` UInt64, `n.b` UInt64, x UInt64) ENGINE = MergeTree ORDER BY x; CREATE MATERIALIZED VIEW mv ENGINE = Null AS SELECT `n.a` FROM src; ALTER TABLE src DROP COLUMN n; INSERT INTO src (x) VALUES (1); ``` Drop succeeds, and from that point every insert into `src` fails with `UNKNOWN_IDENTIFIER`, because the view still selects `n.a`. With `IF EXISTS` the outcome is the data destruction from problem 1, with the view still subscribed Fiddle: https://fiddle.clickhouse.com/4dd388c7-af62-41d4-b097-bf4007e25041 ## 3. An unfinished mutation on a group's column does not block dropping the group ```sql CREATE TABLE t (`n.a` UInt64, x UInt64, c UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t VALUES (1, 1, 1); ALTER TABLE t UPDATE c = `n.a` + 1 WHERE 1 SETTINGS mutations_sync = 0; ALTER TABLE t DROP COLUMN n; SELECT command, is_done, latest_fail_reason FROM system.mutations WHERE table = 't'; ``` When the background mutation has not finished yet, the drop passes, the group is removed, and the queued mutation now references a column that no longer exists - it stays in the queue forever until `KILL MUTATION`. Whether the statement is safe depends on a race with the background pool ### Expected behavior _No response_ ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context #114468 #114163",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114481",
        "createdAt": "2026-08-12T12:11:22Z",
        "updatedAt": "2026-08-13T15:20:53Z",
        "timestamp": "2026-08-13T15:20:53Z",
        "metrics": {
          "reactions": 1,
          "comments": 1
        },
        "labels": [
          "potential bug"
        ],
        "author": "m7kss1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114487",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Malformed Iceberg schema-evolution metadata aborts the server via LOGICAL_ERROR instead of a normal exception",
        "text": "🕵️ ## Describe what's wrong Reading an Iceberg table whose metadata describes an invalid schema evolution (per the Iceberg spec) crashes the server instead of raising a normal, catchable exception: ``` Thread ... received signal SIGABRT, Aborted. ... #11 DB::Iceberg::IcebergSchemaProcessor::getSchemaTransformationDag (..., old_id=0, new_id=1) at SchemaProcessor.cpp:729 ``` `getSchemaTransformationDag` correctly detects the spec violation (\"a required column can't be added during schema evolution without a default\") but reports it via `ErrorCodes::LOGICAL_ERROR`, which `Exception::handleErrorCode` treats as an assertion failure and **aborts the process** in debug and sanitizer builds (`Common/Exception.cpp`, comment: \"In debug builds and builds with sanitizers, treat LOGICAL_ERROR as an assertion failure\"). In a plain release build without assertions it would merely surface as an oddly-labeled `Code: 49` exception - still wrong, but not fatal. This is reachable purely through metadata content - no ALTER, no live write path - so any corrupted or adversarially-crafted `.metadata.json` (or a bug in a third-party writer) that describes an invalid evolution takes the server down on the next read, in any debug/sanitizer build (which includes ClickHouse's own CI debug/sanitizer stateless-test configurations). Found via a lake-corruption fuzzer that feeds semantically-mutated Iceberg metadata to `clickhouse local` - this was, by a wide margin, the single most frequent abort signature across two independent 15-minute fuzzing sessions (22/635 and 46/1017 iterations respectively landed on this exact frame), yet distinct from the already-filed #114350 (empty-column-name `chassert`) - different function, different throw site, different trigger condition. ## Root cause `SchemaProcessor.cpp`, inside `IcebergSchemaProcessor::getSchemaTransformationDag`, has (at least) two throw sites that validate externally-supplied schema content but use `ErrorCodes::LOGICAL_ERROR` - a code reserved for \"this should never happen, it's our own bug\", not \"the input data is invalid\": ```cpp // line ~697 - old field was a struct/list/map, new field is a primitive if (old_json->isObject(f_type) && !field->isObject(f_type)) { throw Exception( ErrorCodes::LOGICAL_ERROR, \"Can't cast primitive type to the complex type, field id is {}, old schema id is {}, new schema id is {}\", id, old_id, new_id); } ``` ```cpp // line ~727 - a new column with no counterpart in the old schema is `required` (no default) if (!type->isNullable() && !field->isObject(f_type)) { throw Exception( ErrorCodes::LOGICAL_ERROR, \"Cannot add a column with id {} with required values to the table during schema evolution. \" \"This is forbidden by Iceberg format specification. Old schema id is {}, new schema id is {}\", id, old_id, new_id); } ``` Both are genuine, correct spec-violation detections - the bug is only the error code choice. Notably, the *sibling* function in the same file, `getSchemaTransformationDagByIds` (a few lines down), validates its own external input (an unknown schema-id) with `ErrorCodes::BAD_ARGUMENTS` instead - so these two throw sites are also inconsistent with the file's own established convention for \"the metadata says something invalid.\" ## Does it reproduce on the most recent release? Reproduced on `master`, `clickhouse-master/BUILD/bin/clickhouse` built 2026-08-12 (debug build, assertions enabled). ## How to reproduce No lake infrastructure needed beyond a normal `IcebergLocal` table plus one JSON edit: ```bash BIN=/path/to/clickhouse rm -rf lake \"$BIN\" local -q \" SET allow_experimental_insert_into_iceberg = 1; SET allow_insert_into_iceberg = 1; CREATE TABLE t (id Int64, s String) ENGINE = IcebergLocal('lake/', 'Parquet'); INSERT INTO t SELECT number, toString(number) FROM numbers(5); \" # Add an evolved schema (id 1) with a new *required* field that has no counterpart in schema 0 # (the schema the existing data file was written under), and make it current - a real writer # would never emit this (it violates the Iceberg spec), simulating corrupted/adversarial metadata. python3 - lake <<'PY' import json, copy, sys, glob path = sorted(glob.glob(f\"{sys.argv[1]}/metadata/v*.metadata.json\"))[-1] with open(path) as f: doc = json.load(f) evolved = copy.deepcopy(doc[\"schemas\"][0]) evolved[\"schema-id\"] = 1 evolved[\"fields\"].append({\"id\": 3, \"name\": \"extra\", \"required\": True, \"type\": \"long\"}) doc[\"schemas\"].append(evolved) doc[\"current-schema-id\"] = 1 with open(path, \"w\") as f: json.dump(doc, f) PY \"$BIN\" local -q \"SELECT * FROM icebergLocal('lake/') ORDER BY id\" # Aborted (core dumped) ``` Full self-contained script: `tmp/repro_schema_le/repro.sh`. ## Expected behaviour A normal, catchable exception, matching how the sibling `getSchemaTransformationDagByIds` already handles \"the metadata references an unknown schema-id\" a few lines below. Should not abort in any build type. ## Suggested fix Trivial - swap the error code at both throw sites in `getSchemaTransformationDag` (`SchemaProcessor.cpp:~697` and `~727`) from `ErrorCodes::LOGICAL_ERROR` to `ErrorCodes::ICEBERG_SPECIFICATION_VIOLATION`. That code (743) is *already declared* as `extern const int ICEBERG_SPECIFICATION_VIOLATION;` at the top of this exact file (`SchemaProcessor.cpp:50`) - it's simply never used by these two sites. No new include, no new declaration, no signature change. This is also already the established convention for this exact class of check elsewhere in the same directory: - `MetadataGenerator.cpp:452`: `throw Exception(ErrorCodes::BAD_ARGUMENTS, \"Iceberg spec doesn't allow to add non-nullable columns\")` - near-identical wording, on the write path. - `Utils.cpp:1188`: picks `BAD_ARGUMENTS` vs `ICEBERG_SPECIFICATION_VIOLATION` depending on `format_version` for the same category of spec check. Only remaining work is a regression test (no existing test references either throw message) - same shape as `04846_iceberg_null_current_snapshot_id.sh` from #114465: rewrite metadata to trigger it, assert a normal exception instead of a crash. Not yet compiled/tested. ## Additional context Distinct from #114350 (`ASTIdentifier` `chassert` on an empty field `name`, a different function entirely) despite both being \"malformed Iceberg metadata aborts the server via an internal-invariant-style check.\" Given how much more frequently this one surfaced in fuzzing than the already-filed issue, worth treating as its own report rather than folding in. Also checked PR #114465 (\"Fix reads of a null `current-snapshot-id`\") on the chance it was a fix for this - it isn't: different function (`getHistory`/`expireSnapshots`/`mutate`), different root cause (`Poco::JSON::Object::has` returning true for a JSON-null value), no overlap with `getSchemaTransformationDag`. Searched GitHub (issues and PRs, open and closed, plus a code search for `getSchemaTransformationDag`) and checked the file's recent commit history - no existing issue or fix for this bug as of 2026-08-12.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114487",
        "createdAt": "2026-08-12T13:30:48Z",
        "updatedAt": "2026-08-13T08:40:40Z",
        "timestamp": "2026-08-13T08:40:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "comp-datalake"
        ],
        "author": "PedroTadim",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114498",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Add a `string_bounds` column statistic for String columns",
        "text": "### Company or project name ClickHouse ### Use case Column statistics currently give string predicates almost nothing: - Range predicates (`url < 'https://m'`, `tenant BETWEEN 'a' AND 'b'`) fall through to the hard-coded `default_cond_range_factor = 0.33`, because `tdigest`/`basic`/`minmax` estimation is numeric-only. - `LIKE` / `ILIKE` always get the hard-coded `default_like_factor = 0.1`, whether the pattern matches 30% of rows or zero. - Equality against an impossible constant (e.g. `WHERE code = 'USA'` on a column where every value is 2 bytes long, or a constant outside the column's value range) is still estimated as if it could match. - `StatisticsPartPruner` only supports numeric MinMax, so parts can never be skipped based on statistics for string predicates — even on tables naturally clustered by a string column (tenant/customer IDs), unless the column is in the primary key or a skipping index. - There are no min/max string length or alphabet (all-ASCII) statistics for width estimates or future execution fast paths. All of these gaps are served by _bounds_, not frequencies — a statistic that is tiny, cheap to build, and losslessly mergeable. ### Describe the solution you'd like A new opt-in, string-specific column statistic, `string_bounds(K)` (default `K = 16`), storing per column per part: - truncated **min/max value bounds** — first `K` bytes of the smallest/largest string, each with an explicit exactness state (exact vs. truncated), compared bytewise (`memcmp` order); - exact **min/max string byte length**; - an **all-ASCII flag**; - a non-null row counter for coverage accounting. ```sql CREATE TABLE t (k UInt64, url String STATISTICS(basic, uniq_v2, string_bounds(16))) ENGINE = MergeTree ORDER BY k; ALTER TABLE t ADD STATISTICS url TYPE string_bounds(16); ALTER TABLE t MATERIALIZE STATISTICS url; ``` Initial uses: 1. **Impossibility detection** for `=` / `IN`: constants outside the value bounds or length bounds estimate to zero (before `mcv`/`countmin`/NDV machinery runs). 2. **Range estimation** for `<`, `<=`, `>`, `>=`, `BETWEEN` on strings: deterministic 0/all classification at the bounds, with an optional (setting-gated) interpolated point estimate in between. 3. **`LIKE 'prefix%'` / `startsWith`**: convert extractable fixed prefixes into ranges (reusing the existing `KeyCondition` prefix machinery) instead of the constant `default_like_factor`. 4. **Part pruning**: extend `StatisticsPartPruner` to string columns via the per-part bounds. Properties that make this the cheapest member of the statistics family: ~60 bytes per column per part, one `memcmp`-dominated pass per block to build, and exact, order-independent merging with no error terms — unlike `mcv`/`histogram`, merging never degrades accuracy. Scope for v1: `String`, `Nullable(String)`, and `LowCardinality` wrappers; explicit opt-in only (not in `auto_statistics_types`). `FixedString(N)` (which needs zero-padded comparison normalization) and nullable part pruning are follow-ups. ### Describe alternatives you've considered - **Extending the existing numeric `minmax` statistic to strings** — rejected: it stores exact typed values, whereas string bounds must be truncated with explicit exactness states, and the length/alphabet fields have no home there. - **Relying on `mcv` / `countmin` for strings** — those cover per-value frequencies of heavy hitters; they cannot answer range, prefix, or impossibility questions, and don't support part pruning. - **Putting length bounds / ASCII flag into `basic`** — viable later (its serialization allows additions), but keeping v1 a self-contained opt-in statistic leaves the default write path untouched. - **Storing full (untruncated) min/max values** — rejected: a single pathological long string would balloon a payload loaded during planning for every selected part; truncated bounds with exactness markers are the approach proven by Parquet (`is_min/max_value_exact`) and ORC (`lowerBound`/`upperBound`). ### Additional context DuckDB's per-segment string statistics drew my attention to this area in ClickHouse, which can benefit from zonemap-style min/max metadata applied to strings; similar truncated string min/max statistics also exist in Parquet and ORC.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114498",
        "createdAt": "2026-08-12T14:46:41Z",
        "updatedAt": "2026-08-13T03:17:29Z",
        "timestamp": "2026-08-13T03:17:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "feature",
          "st-need-info"
        ],
        "author": "cv4g",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114499",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "[RFC] Add most-common-values (`mcv`) column statistics for better equality/IN selectivity on skewed columns",
        "text": "### Company or project name ClickHouse ### Use case Cardinality estimation for equality and `IN` predicates on skewed columns is currently weak. When only `uniq`/`uniq_v2` statistics are available, `ColumnStatistics::estimateEqual()` falls back to a uniform assumption of `rows / ndv`, which can be off by orders of magnitude for skewed data: ```sql WHERE country = 'US' -- 'US' may be 45% of rows, not 1/ndv WHERE status = 500 WHERE event_type IN ('purchase', 'signup') ``` Example: a scan of 100M rows with 250 distinct countries yields a uniform estimate of 400K rows for `country = 'US'`. If `US` is actually 45% of the column, the true answer is ~45M rows — a 100x error that propagates into join order, PREWHERE, and other plan choices. Status-, enum-, and country-like `LowCardinality` columns are the primary target, where skew is common and the statistic is cheap to build. ### Describe the solution you'd like Add an opt-in, parameterized per-column statistic `mcv(N)` (most common values), reusing the existing `IStatistics` / `ColumnStatistics` framework: ```sql CREATE TABLE events ( user_id UInt64, country LowCardinality(String) STATISTICS(basic, uniq_v2, mcv(64)), status UInt8 STATISTICS(basic, uniq_v2, mcv(16)) ) ENGINE = MergeTree ORDER BY user_id; ALTER TABLE events ADD STATISTICS country TYPE mcv(64); ALTER TABLE events MATERIALIZE STATISTICS country; ``` Key properties: - **Bounded, mergeable heavy-hitter summary**, not an unbounded exact map and not a plain local top-N list. Canonical representation: a Misra-Gries-family summary (per-value lower-bound counts plus one global error term), with an exact mode for low-cardinality data. Local exact top-N lists are insufficient because they don't merge reliably: a value can be a global heavy hitter without appearing in any single part's top-N. - **Query-scope merging**: queries scan many parts, so serialized summaries must be mergeable at planning time (in `ConditionSelectivityEstimatorBuilder`), with well-defined error accumulation. Physical merges rebuild the statistic from the merge output stream when the row set changes. - **Estimation**: for tracked values, use the MCV count (with explicit lower/upper bounds in sketch mode); for untracked values, subtract the MCV head and spread the residual over the remaining NDV — the classic MCV treatment in relational optimizers (cf. PostgreSQL `most_common_vals`). Add a batch estimate for `IN` lists so shared tail mass isn't double-counted. - **Opt-in and guarded**: require the `N` parameter, keep `mcv` out of default `auto_statistics_types`, and bound build memory and serialized payload size via server-level settings. - **Scope of v1**: equality/`IN` selectivity only, per-shard planning. MCV is complementary to `uniq_v2` (head vs. tail) and to a separately planned histogram statistic (head vs. body/tail distribution). A prerequisite is fixing parameter handling for statistics: today `STATISTICS(tdigest(200))` parses but silently drops the argument, and factory/deserialization paths recreate descriptions from the bare type enum. `mcv(N)` needs parameters to be validated, persisted, compared in `structureEquals()`, and survivable through serialization. ### Describe alternatives you've considered - **Exact local top-N per part**: simple, but merging only top-N lists across parts misses true global heavy hitters; slack candidates and an explicit error term are needed for sound query-scope merging. - **`countmin` (existing, requires `USE_DATASKETCHES`)**: estimates the frequency of a known value but cannot enumerate heavy hitters, so it can't drive skew detection or tail-subtraction estimates; it remains useful as a fallback for untracked point values. - **Unbounded exact `value -> count` map**: unbounded memory during background merges; unacceptable. - **Histograms**: serve range predicates and the distribution body/tail; planned separately and complementary — heavy hitters distort buckets, so an MCV head makes a future histogram better, not redundant. - **Naming (`frequent_items`, `topk`, `heavy_hitters`)**: `mcv` matches the established optimizer-statistics term (PostgreSQL `most_common_vals`) and is algorithm-agnostic. ### Additional context - The Misra-Gries family gives deterministic retention (with `k` counters, any value with frequency `> rows/(k+1)` is retained) and provably mergeable summaries (Agarwal et al., \"Mergeable Summaries\", TODS 2013). - A production-proven implementation (Apache DataSketches frequent-items sketch) is already vendored in `contrib/datasketches-cpp`; the in-tree `SpaceSaving.h` (used by `topK`) is a related candidate engine. The persisted format should be engine-agnostic, and `mcv` should not hard-depend on `USE_DATASKETCHES`. - Follow-ups explicitly out of scope for v1: join cardinality from MCV intersections, skew-aware join/aggregation execution, cross-shard statistics transport, multi-column statistics",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114499",
        "createdAt": "2026-08-12T14:46:45Z",
        "updatedAt": "2026-08-13T09:53:59Z",
        "timestamp": "2026-08-13T09:53:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "feature",
          "st-need-info"
        ],
        "author": "cv4g",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114500",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "[RFC] Add a `histogram(N)` column statistic for range predicate selectivity",
        "text": "### Company or project name ClickHouse ### Use case Selectivity estimation for range predicates (`<`, `<=`, `>`, `>=`, `BETWEEN`, and range decompositions from `PlainRanges`) on columns with non-uniform value distributions. Today these are estimated with `tdigest` (if declared), linear interpolation over `[min, max]` from `basic`/`minmax`, or a magic default factor (`default_cond_range_factor = 0.33`). The interpolation fallback assumes a uniform distribution and can be off by an order of magnitude on common data shapes. Example: a `ts` column spanning three years where 80% of rows are in the last three months — for `WHERE ts >= '2024-01-01' AND ts < '2024-02-01'` interpolation predicts ~1/36 of rows while the true answer is ~27%. This misleads PREWHERE ordering, join-order decisions, and `RelationProfile` row estimates. `tdigest` helps, but as an optimizer statistic it gives point estimates with no error bounds, is comparatively expensive to build inside background merges, and cannot compose with a most-common-values statistic (no way to subtract heavy-hitter mass from a region of the digest). ### Describe the solution you'd like An opt-in, parameterized bucket-frequency histogram statistic: ```sql CREATE TABLE t ( k UInt64, ts DateTime STATISTICS(basic, uniq_v2, histogram(128)) ) ENGINE = MergeTree ORDER BY k; ALTER TABLE t ADD STATISTICS ts TYPE histogram(128); ALTER TABLE t MATERIALIZE STATISTICS ts; ``` Key properties: - **Equi-width buckets on a dyadic grid**: bucket width is a power of two, boundaries anchored at zero. This is what makes the statistic fit ClickHouse's per-part architecture: it builds streaming, block-by-block, in bounded memory (a fixed counter array, no sort, no second pass), and grids built independently on different parts are hierarchically nested, so query-scope merging across thousands of selected parts is exact re-binning — merging loses resolution, never accuracy. - **Exact counts, deterministic bounds**: counts cover all rows of the part (not a sample), so the estimator gets hard `lower(v) <= true_count < upper(v)` bounds with a policy-driven interpolated point estimate in between. - **Framework reuse**: implemented as a new `StatisticsType` in the existing `IStatistics` / `ColumnStatistics` framework; row-preserving merges combine summaries exactly, row-changing merges (TTL, deletes, collapsing) rebuild from the output stream — matching the current rebuild-vs-merge split in `MergeTask`. - **Supported types (v1)**: integers, floats (with dedicated NaN/±Inf counters), `Decimal`, `Date`/`Date32`/`DateTime`/`DateTime64`, `Enum`, `IPv4`, plus `Nullable`/`LowCardinality` wrappers. `String` and 128-bit ID types deferred. - **Estimation use (v1)**: range/comparison selectivity in `estimateLess`/`estimateRange` first; equality/`IN` refinement via bucket density and composition with the proposed `mcv` statistic (subtracting the heavy-hitter head at estimation time) as follow-ups. - **Guardrails**: `N` required and capped by server-level settings; payload size bounded (~one `UInt64` per bucket); not included in `auto_statistics_types` initially. ### Describe alternatives you've considered - **Equi-height (equi-depth) histograms** (PostgreSQL, MySQL, Oracle, StarRocks): more accurate per bucket under skew, but require sorted data or precomputed quantiles to build and cannot be merged exactly across parts — a non-starter for per-part streaming statistics built inside inserts and merges. An approximate equi-depth _view_ can still be derived at query scope from merged fine-grained buckets. - **Keeping `tdigest` as the only range estimator**: retained, not replaced — but it carries no formal error bounds, and its centroids shift under merging; the histogram is expected to be the better range estimator, to be validated by benchmarks. - **A mergeable quantile sketch (KLL)**: provably mergeable, but duplicates `tdigest`'s rank-space role, gives probabilistic rather than deterministic bounds, and offers no stable bucket identity for future per-bucket metadata. - **The `histogram()` aggregate function's adaptive bins** (Ben-Haim/Tom-Tov): streaming and bounded, but data-dependent boundaries never align across parts, so merging is heuristic — fine for visualization, wrong for an optimizer statistic. ### Additional context Suggestion to implement histograms raised by @fkastrati.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114500",
        "createdAt": "2026-08-12T14:46:49Z",
        "updatedAt": "2026-08-13T10:25:15Z",
        "timestamp": "2026-08-13T10:25:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "feature"
        ],
        "author": "cv4g",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114512",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "`PREWHERE <bare column>` with `FINAL` throws THERE_IS_NO_COLUMN when PREWHERE is deferred after FINAL",
        "text": "🕵️ ## Describe what's wrong An explicit `PREWHERE` whose expression is a plain column reference fails when PREWHERE is deferred until after `FINAL`: ``` Code: 8. DB::Exception: Cannot find column `b` in source stream, there are only columns: [other]. (THERE_IS_NO_COLUMN) ``` This is reachable **with stock default settings** — no `SETTINGS` clause required. The deferral is triggered automatically by a row policy on a non-sorting-key column, because `apply_row_policy_after_final` defaults to `1` and PREWHERE must run after the policy: ```sql CREATE TABLE t (k UInt64, b UInt8, other UInt8) ENGINE = ReplacingMergeTree() ORDER BY k; INSERT INTO t SELECT number, number % 2, number % 3 FROM numbers(20); CREATE ROW POLICY pol ON t USING other = 1 TO ALL; SELECT count() FROM t FINAL PREWHERE b; -- Code: 8. THERE_IS_NO_COLUMN ``` The same failure occurs without any row policy by opting into the deferral directly, on every FINAL-capable engine tested (`ReplacingMergeTree`, `CoalescingMergeTree`, `AggregatingMergeTree`, `SummingMergeTree`): ```sql SELECT count() FROM t FINAL PREWHERE b SETTINGS apply_prewhere_after_final = 1; ``` ## What makes it fire The deferred filter step resolves its column against the source stream, but nothing adds the PREWHERE expression's own required columns to the set of columns read from the table. So it only succeeds *by accident*, when the column is already in the SELECT list: | projection | `PREWHERE b` (plain `UInt8`) | `PREWHERE n.null` (subcolumn) | |---|---|---| | `SELECT count()` | **fails** (stream: `[]`) | **fails** (stream: `[]`) | | `SELECT k` | **fails** (stream: `[k]`) | **fails** (stream: `[k]`) | | `SELECT b` | ok (`b` happens to be in the stream) | **fails** | | `SELECT *` | ok (`b` happens to be in the stream) | **fails** (subcolumns are not in the stream) | Two further properties, both consistent with \"the required columns are never requested\": - **Only a bare column reference fails.** Any wrapping makes it work, because the wrapping expression's DAG declares the column as an input: `NOT b`, `b = 1`, `toUInt8(b)` and `b AND k > 5` all succeed where bare `b` does not. - **A subcolumn never works**, for any projection, since subcolumns are not present in the source stream even for `SELECT *`. Controls that behave correctly: the same query without `FINAL`, the same query with `apply_prewhere_after_final = 0`, and the same table without a row policy. ## Root cause `ReadFromMergeTree::deferFiltersAfterFinalIfNeeded` (`src/Processors/QueryPlan/ReadFromMergeTree.cpp:2509`) moves the PREWHERE aside for later: ```cpp if (query_info.prewhere_info && isPrewhereDeferredAfterFinal()) deferred_prewhere_info = query_info.prewhere_info; ``` `isPrewhereDeferredAfterFinal` (`:2500`) returns true for `apply_prewhere_after_final`, or whenever `isRowPolicyDeferredAfterFinal` does — which for a non-sorting-key policy is the default: ```cpp /// PREWHERE must run after the row policy, so deferred row policy defers PREWHERE as well return context->getSettingsRef()[Setting::apply_prewhere_after_final] || isRowPolicyDeferredAfterFinal(); ``` Once deferred, the PREWHERE's required columns still have to be read from the table so the later filter step can evaluate them, and that does not appear to happen for a bare column reference. ## Does it reproduce on the most recent release? Reproduced on `master`, version 26.8.1.1226 (debug build). ## How to reproduce Full matrix script: `tmp/repro_prewhere_subcol/repro.sh`. Minimal case is the 4-statement default-settings snippet above. ## Expected behaviour `SELECT count() FROM t FINAL PREWHERE b` returns the count of matching rows, exactly as `SELECT count() FROM t FINAL PREWHERE NOT b` and `... PREWHERE b = 1` already do. ## Additional context Adjacent but distinct from #109703 (\"Do not move conditions to PREWHERE when PREWHERE is deferred after FINAL\"), which stopped the *optimizer* from moving `WHERE` into a deferred PREWHERE. That fix is why `WHERE b` works here while an explicit, user-written `PREWHERE b` still takes the broken path — the deferral machinery itself was not made safe, only the automatic move into it. The deferral feature comes from #91065 (\"Apply row policies and PREWHERE after FINAL\"). Found by a PREWHERE correctness fuzzer (`tmp/fuzz_prewhere/`) that compares every PREWHERE variant against the same filter with all PREWHERE paths disabled; this surfaced as `apply_prewhere_after_final` disagreeing with the ground truth on `PREWHERE n.null`, and minimising showed the subcolumn was incidental — a plain `UInt8` column fails the same way.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114512",
        "createdAt": "2026-08-12T16:01:41Z",
        "updatedAt": "2026-08-13T02:59:32Z",
        "timestamp": "2026-08-13T02:59:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug"
        ],
        "author": "PedroTadim",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114579",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Lazy FINAL with `optimize_aggregation_in_order` merges one group at a time (25000 `mergeBlocks` calls and 50000 log lines for a 35000-row table)",
        "text": "### Describe the situation When `query_plan_optimize_lazy_final` and `optimize_aggregation_in_order` are both on, the aggregation that the lazy `FINAL` replacement builds is merged **one group at a time**: `Aggregator::mergeBlocks` is called once per distinct key, and each call writes two log lines (a `Trace` \"Merging partially aggregated blocks\" and a `Debug` \"Merged partially aggregated blocks for bucket #-1\"). On a 35000-row table with 25000 distinct keys that is 25000 merges and 50000 log lines for a single query. A plain in-order `GROUP BY` of the same size does 7 merges, so this is specific to the lazy `FINAL` shape rather than to aggregation in order in general. The cost is dominated by the logging, so it scales with how verbose the server is configured to be. On a debug/sanitizer build with the log level CI uses, the query below took **525 s**, of which `LoggerElapsedNanoseconds` attributes **285 s** to the logger; the same query with `optimize_aggregation_in_order = 0` took ~3 s. This is not a correctness problem - the results match - and it is not new; it reproduces on builds well before the report below. ### How to reproduce Any recent `master`. With `clickhouse-local`: ```sql CREATE TABLE lf (k UInt64, version UInt64, is_deleted UInt8, v UInt64) ENGINE = ReplacingMergeTree(version, is_deleted) ORDER BY k; INSERT INTO lf SELECT number, 1, 0, number FROM numbers(20000); INSERT INTO lf SELECT number, 2, if(number % 10 = 0, 1, 0), number * 2 FROM numbers(10000, 15000); SELECT count(), sum(v) FROM lf FINAL WHERE k % 7 != 6 SETTINGS max_threads = 4, max_block_size = 8192, query_plan_optimize_lazy_final = 1, max_rows_for_lazy_final = 10000000, min_filtered_ratio_for_lazy_final = 0, optimize_aggregation_in_order = 1; ``` Run it with `--send_logs_level=trace` and count the merges: ``` optimize_aggregation_in_order = 1 -> 25000 \"Merging partially aggregated blocks\" lines optimize_aggregation_in_order = 0 -> 0 ``` The plan shows the replacement's own aggregation (`GROUP BY k` with `argMax` states) under `LazyReadReplacingFinal`; with aggregation in order it emits one chunk per group into the merge stage. ### Expected performance The merge stage should batch groups the way it does for an ordinary in-order `GROUP BY` (7 merges for 50000 groups), instead of one merge per group. ### Additional context Found while triaging a test timeout in https://github.com/ClickHouse/ClickHouse/pull/111459, where the randomized `optimize_aggregation_in_order = 1` setting made a small lazy `FINAL` test query take ~525 s in every flaky check. It is unrelated to that pull request - it reproduces with the feature under test switched off and on builds that predate it - and the test there now pins the setting, but the underlying pathology is worth fixing. Related: https://github.com/ClickHouse/ClickHouse/issues/113704",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114579",
        "createdAt": "2026-08-13T03:29:56Z",
        "updatedAt": "2026-08-13T17:10:18Z",
        "timestamp": "2026-08-13T17:10:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114581",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "IN (SELECT ...) inside a higher-order-function lambda always evaluates to 0 when the query is a derived table or the lambda is in WHERE",
        "text": "**Describe what's wrong** An `IN (SELECT ...)` predicate inside a higher-order-function lambda (`arrayExists`, `arrayFilter`, `arrayMap`, ...) always evaluates to `0` when the enclosing SELECT is used as a derived table (or when the lambda sits in an outer `WHERE`). The identical expression at the top level returns the correct result. **Does it reproduce on the most recent release?** Reproduces on current master, `26.8.1.1068` and `26.8.1.1194` (official builds). Deterministic: 20/20 runs. **How to reproduce** No tables needed: ```sql SELECT arrayExists(x -> x IN (SELECT 2), [2]); -- 1 (correct) SELECT * FROM (SELECT arrayExists(x -> x IN (SELECT 2), [2])); -- 0 (wrong: the same expression, only wrapped in a derived table) ``` More shapes of the same mechanism: ```sql SELECT arrayFilter(x -> x IN (SELECT 2), [1, 2, 3]); -- [2] (correct) SELECT * FROM (SELECT arrayFilter(x -> x IN (SELECT 2), [1, 2, 3])); -- [] (wrong) SELECT arrayMap(x -> x IN (SELECT '2'), [2, 3]); -- [1,0] (correct) SELECT * FROM (SELECT arrayMap(x -> x IN (SELECT '2'), [2, 3])); -- [0,0] (wrong) -- WHERE context loses rows: SELECT count() FROM (SELECT 1 AS k) WHERE arrayExists(x -> x IN (SELECT 1), [k]); -- 0 (wrong: expected 1) WITH tm1 AS (SELECT arrayExists(x -> x IN (SELECT 2), [2])) SELECT * FROM tm1; -- 0 (wrong) ``` All at default settings; `query_plan_enable_optimizations = 0` does not cure it, so it looks like the set for the lambda-captured `IN` is not built/bound when the expression is resolved inside a subquery scope, rather than a plan-optimization issue. Possibly related observation: with `enable_analyzer = 0` the derived-table form fails outright with an exception `Code: 47` `UNKNOWN_IDENTIFIER`, where the required column is spelled `... in(x, _subquery1) ...` but the available column is `... in(x, _subquery2) ...` — the same set-identity confusion visible in the old analyzer. **Expected behavior** Wrapping a SELECT in a derived table (or moving the expression into `WHERE`) must not change the value of `IN (SELECT ...)` inside a lambda: all the wrapped forms above should return the top-level results (`1`, `[2]`, `[1,0]`, `1`). Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114581",
        "createdAt": "2026-08-13T03:36:46Z",
        "updatedAt": "2026-08-13T15:09:25Z",
        "timestamp": "2026-08-13T15:09:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "potential bug"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114582",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Failed INSERT INTO s3(...) PARTITION BY leaves a durable read-visible prefix; default hive strategy silently duplicates it on retry",
        "text": "### Company or project name ClickHouse QA (durability testing) ### Describe what's wrong A failed `INSERT INTO FUNCTION s3(...) PARTITION BY <key>` (and the equivalent object-storage table engines) is **not atomic and leaves a durable, read-visible prefix of the partitions it had already written**, while the statement reports failure. `PartitionedSink` finalizes one object per partition value **sequentially**, with no rollback of the partitions committed before the one that fails: ```cpp // src/Storages/PartitionedSink.cpp:120-125 void PartitionedSink::onFinish() { for (auto & [_, sink] : partition_id_to_sink) { sink->onFinish(); // each StorageObjectStorageSink::onFinish -> finalizeBuffers -> write_buf->finalize() (the durable PUT / CompleteMultipartUpload) } } ``` Each child `StorageObjectStorageSink::onFinish` (`src/Storages/ObjectStorage/StorageObjectStorageSink.cpp:89-119`) finalizes (PUTs) its own object. If the PUT of partition *k* throws, the exception propagates out of the loop; the partitions finalized **before** *k* are already durable and immediately listable/readable, and nothing deletes them or records that only a prefix landed (`cancelBuffers` only aborts the not-yet-finalized buffers). So the error's implied \"nothing happened\" contract is a lie. The consequence on the **default** partition strategy is silent duplication on a natural retry. `file_like_engine_default_partition_strategy` defaults to `HIVE` (`src/Core/Settings.cpp:7754`), and `HiveStylePartitionStrategy::getPathForWrite` appends a fresh `generateSnowflakeID()` to every filename: ```cpp // src/Storages/IPartitionStrategy.cpp:383 path += std::to_string(generateSnowflakeID()) + \".\" + Poco::toLower(file_format); ``` so every attempt writes new object keys. A client that treats the failed INSERT as \"nothing happened\" and re-submits it therefore **duplicates every partition that survived the first attempt** — there is no dedup for object-storage writes and no exists-check that can catch it (the keys differ each time). A client that does *not* re-submit is left with silently half-applied state. The `WILDCARD` strategy (an explicit `{_partition_id}` placeholder in the path) fails safe instead of duplicating: the retry recomputes the same keys and the exists-check rejects it (`Code: 36 ... already exists, enable s3_truncate_on_insert`), which is fail-closed but still leaves the client permanently half-applied and unable to tell which partitions survived without listing the bucket. ### Does it reproduce on the most recent release? Yes — reproduced deterministically on 26.8.1.1. The code paths (`PartitionedSink::onFinish`, `StorageObjectStorageSink::finalizeBuffers`, `HiveStylePartitionStrategy::getPathForWrite`, and the `HIVE` default) are all present on current `master`. ### How to reproduce A durability rig reproduces both facets 3/3 with controls (MinIO behind a fault proxy that returns a non-retryable `403 AccessDenied` on **one** partition's PUT — the classic least-privilege / transient-denial to a single prefix — and a single clean server). `N = 300` rows over 3 partition values `p = number % 3`. **Facet 1 — a failed INSERT leaves a read-visible prefix (WILDCARD path, deterministic):** 1. `INSERT INTO FUNCTION s3('http://.../base/{_partition_id}/data.parquet', ..., 'Parquet', 'id UInt64, p UInt64') PARTITION BY p SELECT number AS id, number % 3 AS p FROM numbers(300)`, with the proxy denying PUTs to the key of the partition that finalizes **last**. 2. The INSERT fails (`Code: 499 ... S3_ERROR`), yet the two partitions finalized before it are durable and queryable through the object store, while the faulted partition is absent. - **Clean control** (proxy disarmed): the 3-partition INSERT round-trips all 300 rows. - **Isolation control** (same fault, but `INSERT ... PARTITION BY p ... WHERE p = <faulted>` — a single partition, no fan-out): the INSERT fails and **nothing** is durable, proving the surviving prefix in step 2 is the fan-out's doing. **Facet 2 — the default HIVE strategy silently duplicates on retry:** 1. Same INSERT into a non-wildcard path (default `partition_strategy = hive`), denying PUTs to one partition's `p=<v>/` prefix. 2. The INSERT fails; a `SELECT count()` over the whole prefix returns **200** (the two surviving partitions). 3. The client re-submits the identical INSERT (its natural response to a failed op). It succeeds, and `SELECT count()` over the prefix now returns **500** — 300 unique rows plus the **200-row surviving prefix duplicated** (fresh snowflake filenames, no dedup, no exists-check). ### Expected behavior A failed `INSERT ... PARTITION BY` should not silently leave a durable, read-visible subset of its partitions with the statement reporting failure. Either the partitioned write should roll back the partitions it already committed on failure (so the error's \"nothing happened\" contract holds), or the surviving prefix should be recorded/surfaced so the outcome is not silently half-applied — and in particular, on the **default** hive strategy a natural client retry of a failed partitioned INSERT should not silently duplicate the rows of the partitions that survived the first attempt. Honest caveat: object-storage writes through the `s3`/`azureBlobStorage` table functions and engines are not transactional, and at-least-once export is a known property. What this report isolates is the specific, undocumented combination — a failed multi-partition INSERT leaves a **query-visible** durable prefix, and the **default** hive strategy makes an ordinary retry duplicate that prefix with no dedup and no way to be idempotent — which is a data-integrity footgun distinct from the concurrent-writer clobber already filed as #112419. ### Error message and/or stacktrace `Code: 499. DB::Exception: ... (S3_ERROR)` on the faulted INSERT; `Code: 36 ... Object in bucket ... already exists (...enable s3_truncate_on_insert)` on the wildcard retry. ### Additional context Found by the ClickFawkes durability framework under its `failed_op_durable_residue` lens (a client op fails but a durable read-visible prefix survives, un-rolled-back). Reproduced via `--mode objectstorage-partition-prefix-residue`.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114582",
        "createdAt": "2026-08-13T03:38:15Z",
        "updatedAt": "2026-08-13T03:38:15Z",
        "timestamp": "2026-08-13T03:38:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114587",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "arrayIntersect overflow guard tests isInteger on a Nullable type, so it never fires",
        "text": "### Describe what's wrong **`arrayIntersect` on `Nullable` integer arrays of different widths silently matches values that do not survive the narrowing cast to the common element type. `arrayIntersect([toNullable(1)], [toNullable(257)])` returns `[1]`; the non-Nullable form `arrayIntersect([1], [257])` correctly returns `[]`. Pre-existing — not introduced by this PR.** - **Root cause:** arrayIntersect.cpp:398 sets `nested_init_type` to the array's nested type WITHOUT removing `Nullable`, and line 401 then calls `isInteger(nested_init_type)`. `WhichDataType(Nullable(UInt32))` is `TypeIndex::Nullable`, so `isInteger` (and `isDate`/`isDateTime`/`isDateTime64`) return false and the whole overflow-mask block is skipped for every `Nullable` argument. <details> <summary>Analysis details (evidence, affected locations, impact)</summary> **Why we believe this is a bug:** `executeImpl` (arrayIntersect.cpp:463) computes the common subtype with `getMostSubtype`, which NARROWS (`Nullable(UInt16)` + `Nullable(UInt8)` -> `Nullable(UInt8)`), then `castColumns` casts the wider argument down. `prepareArrays` is meant to catch the values that changed under that cast by building `arg.overflow_mask` (arrayIntersect.cpp:395-415), and the element loop skips masked elements (arrayIntersect.cpp:711). For a `Nullable` array the guard at line 401 never lets that happen, so the truncated value is inserted and matched as if it were the original. **Affected locations:** - [`src/Functions/array/arrayIntersect.cpp:401`](https://github.com/ClickHouse/ClickHouse/blob/2485f1496f7dff/src/Functions/array/arrayIntersect.cpp#L401) — overflow-mask guard: isInteger on a still-Nullable nested type - [`src/Functions/array/arrayIntersect.cpp:398`](https://github.com/ClickHouse/ClickHouse/blob/2485f1496f7dff/src/Functions/array/arrayIntersect.cpp#L398) — nested_init_type keeps its Nullable wrapper **Impact:** Wrong results from `arrayIntersect` whenever the arguments are `Nullable` arrays of different integer widths (or Date/DateTime widths) and a value in the wider argument aliases a value in the narrower one modulo the narrower type. Reachable from ordinary SQL over two `Array(Nullable(...))` table columns; `Nullable` array columns are common, and the non-Nullable behaviour that users would infer from `00930_arrayIntersect.sql` is the opposite. </details> ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `2485f1496f7dff`). ### How to reproduce [▶ Run on ClickHouse Fiddle](https://fiddle.clickhouse.com/6be77c3a-66e5-40e6-9596-b3c75b04e31c) <details> <summary>Reproducer</summary> ```sql -- Test: a value that does not survive the cast to the common element type is not in the intersection, -- also when the arrays are Nullable. DROP TABLE IF EXISTS t_04869; SELECT arrayIntersect([1], [257]); SELECT arrayIntersect([toNullable(1)], [toNullable(257)]); SELECT arrayIntersect([-100], [156]); SELECT arrayIntersect([toNullable(-100)], [toNullable(156)]); DROP TABLE IF EXISTS t_04869; CREATE TABLE t_04869 (a Array(Nullable(UInt32)), b Array(Nullable(UInt8))) ENGINE = Memory; INSERT INTO t_04869 VALUES ([1024, 1031], [0, 7]), ([256, 300], [0, 44]); SELECT arraySort(arrayIntersect(a, b)) FROM t_04869 ORDER BY a; DROP TABLE IF EXISTS t_04869; ``` </details> ### Expected behavior ``` [] [] [] [] [] [] ``` ### Error message and/or stacktrace ``` [] [1] [] [156] [0,44] [0,7] ``` <details> <summary>Suggested fix</summary> Strip `Nullable` before the type test, e.g. compute `auto init_nested = removeNullable(nested_init_type);` and gate on that. `callFunctionNotEquals` is already handed the null-stripped nested columns (arrayIntersect.cpp:389-392), so the types passed with them at lines 408-409 must be null-stripped too, and the existing `removeNullable(overflow_mask)` at line 412 then becomes redundant rather than load-bearing. </details> <details> <summary>Additional context</summary> **Open risks:** - `arrayUnion` and `arraySymmetricDifference` use `getLeastSupertype` and therefore never narrow — verified `arrayUnion([toNullable(1)], [toNullable(257)])` = `[1,257]`. No sibling call site to fix. Found during automated review of [PR #113021](https://github.com/ClickHouse/ClickHouse/pull/113021). Severity P1 · Finding `h_pr113021_001` </details> cc @alexey-milovidov (author of #113021) ### Additional context _Generated by ClickGap / Claude._",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114587",
        "createdAt": "2026-08-13T03:44:35Z",
        "updatedAt": "2026-08-13T03:44:35Z",
        "timestamp": "2026-08-13T03:44:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "clickgapai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114588",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Column named like an array subcolumn (a.size0) added via ALTER: old parts silently return the subcolumn value instead of the DEFAULT, and merges materialize the wrong values",
        "text": "**Describe what's wrong** After `ALTER TABLE ... ADD COLUMN` adds a column whose name collides with a generated array subcolumn (e.g. a column named `a.size0` next to an `Array` column `a`), reads of that column from parts created **before** the ALTER silently return the **subcolumn's computed value** (the array size) instead of the added column's DEFAULT. Parts created after the ALTER return the stored value, so one `SELECT` returns a mix of two different meanings for the same identifier — and a subsequent merge (`OPTIMIZE ... FINAL`) **materializes the wrong values permanently** into the new part. **Does it reproduce on the most recent release?** Reproduces on current master, `26.8.1.1068` and `26.8.1.1194` (official builds). Deterministic: 20/20 runs. **How to reproduce** ```sql CREATE TABLE t_sub (a Array(UInt64)) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t_sub VALUES ([0]), ([1, 2]); ALTER TABLE t_sub ADD COLUMN `a.size0` UInt64 DEFAULT 7; INSERT INTO t_sub (a, `a.size0`) VALUES ([3], 7); SELECT a, `a.size0` FROM t_sub; -- [0] 1 <- wrong: the array-size subcolumn, expected the DEFAULT 7 -- [1,2] 2 <- wrong: expected 7 -- [3] 7 <- correct (part written after the ALTER) ``` The implicit default has the same problem: ```sql CREATE TABLE t_sub3 (a Array(UInt64)) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t_sub3 VALUES ([10, 20, 30]); ALTER TABLE t_sub3 ADD COLUMN `a.size0` UInt64; SELECT `a.size0` FROM t_sub3; -- 3 <- wrong: expected 0 (the type default) ``` And the merge bakes the wrong values in, so the corruption survives the mixed-parts state: ```sql OPTIMIZE TABLE t_sub FINAL; SELECT a, `a.size0` FROM t_sub ORDER BY a; -- [0] 1 <- now stored physically -- [1,2] 2 -- [3] 7 ``` A control table created with the column from the start returns the stored values everywhere. **Expected behavior** A declared physical column must shadow the generated subcolumn consistently: reads from parts that predate the ALTER should evaluate the column's DEFAULT (`7`, or `0` for the implicit default), exactly as they do for any other added column name, and merges must materialize those DEFAULT values. Related: https://github.com/ClickHouse/ClickHouse/issues/113420 (the same name-collision family: `ALTER MODIFY COLUMN` with mixed-type parts makes every SELECT throw `AMBIGUOUS_COLUMN_NAME`; this issue is the silent wrong-result face under `ADD COLUMN`) Related: https://github.com/ClickHouse/ClickHouse/issues/87375 Related: https://github.com/ClickHouse/ClickHouse/issues/90219 Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114588",
        "createdAt": "2026-08-13T03:45:57Z",
        "updatedAt": "2026-08-13T03:45:57Z",
        "timestamp": "2026-08-13T03:45:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "potential bug"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114591",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Settings profile with a Map setting (`http_response_headers`) does not survive a serialization round trip",
        "text": "### Company or project name ClickHouse ### Describe the unexpected behaviour An access entity with a Map-valued setting — e.g. a settings profile with `http_response_headers` — is serialized into a form that ClickHouse's own parser cannot read back. The stored entity becomes permanently unloadable. ### How to reproduce ```sql -- A string literal is the only accepted way to write a Map setting in a profile -- (map literals {...}, array literals [...], and map(...) are all rejected here): CREATE SETTINGS PROFILE test_profile SETTINGS http_response_headers = '{''Content-Type'':''application/json'', ''Access-Control-Allow-Origin'':''*''}' CONST; SHOW CREATE SETTINGS PROFILE test_profile; -- CREATE SETTINGS PROFILE `test_profile` SETTINGS http_response_headers = [('Content-Type', 'application/json'), ('Access-Control-Allow-Origin', '*')] CONST -- Feeding the emitted statement back — which is exactly what DiskAccessStorage and -- ReplicatedAccessStorage store and re-parse via deserializeAccessEntity: CREATE SETTINGS PROFILE test_profile2 SETTINGS http_response_headers = [('Content-Type', 'application/json'), ('Access-Control-Allow-Origin', '*')] CONST; -- Code: 62. DB::Exception: Syntax error: failed at position 72 ([): -- Expected one of: literal, NULL, NULL, number, Bool, TRUE, FALSE, string literal, end of query ``` No value of the setting round-trips — even an empty map serializes as `[] CONST`, which is rejected the same way. ### Root cause The write path and the read path disagree: - Write: `SettingsProfileElement` (`src/Access/SettingsProfileElement.cpp`) casts the value to the setting's native type (`Map`), and the AST formats the Field as a literal, producing an array-of-tuples literal `[('k', 'v'), ...]`. - Read: `ParserSettingsProfileElement` (`src/Parsers/Access/ParserSettingsProfileElement.cpp`) parses the value with a scalar-only literal parser. Unlike the query-level `SETTINGS` clause parser, it was never extended to accept collection literals (cf. #75065, where the query-level parser expects `literal or map, ..., OpeningCurlyBrace`). So `serializeAccessEntity` → `deserializeAccessEntity` fails for any user/role/profile whose profile elements contain a collection-valued setting (`http_response_headers`, `additional_table_filters`, ...). > *Generated by [Nerve](https://github.com/ClickHouse/nerve)*",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114591",
        "createdAt": "2026-08-13T04:17:50Z",
        "updatedAt": "2026-08-13T11:14:20Z",
        "timestamp": "2026-08-13T11:14:20Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "potential bug"
        ],
        "author": "pufit",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114592",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Lightweight UPDATE: a mid-commit failure leaves a read-visible partial update and a retry double-applies",
        "text": "### Company or project name ClickHouse QA (internal durability testing) ### Describe what's wrong A multi-partition lightweight `UPDATE` is neither atomic on failure nor idempotent on retry. `UPDATE t SET c = ... WHERE ...` over a table with N partitions creates one patch part per partition and commits them sequentially (`MergeTreeSinkPatch::finishDelayedChunk` for `MergeTree`, `ReplicatedMergeTreeSinkPatch::finishDelayed` for `ReplicatedMergeTree`), each with `deduplicate = false` hard-set. If the commit of the k-th partition's patch part fails (for example a filesystem error while renaming the staging directory into place), the patch parts for partitions `1..k-1` are already Active and read-visible, yet the `UPDATE` returns an error. Two problems follow: 1. The failed `UPDATE` has partially applied and that partial state is immediately visible on `SELECT` — with no retry at all. 2. Because `deduplicate = false` is hard-set for patch parts, a client that retries the same `UPDATE` (the natural response to an error) re-applies the change to the partitions that already have it. There is no idempotency: the retried patch is a fresh part with a fresh block number. The net effect is a silent split double-apply: the partition that received the residue ends at `+2` while the partition that did not ends at `+1`, when the user issued a single `SET x = x + 1`. This is distinct from #112133 (please read the distinction before deduplicating on the title): #112133 concerns the mutation-entry commit — a single Keeper `tryMulti` whose response is lost, misreported as failure, so a retry creates a second mutation entry. That path applies a whole N-partition mutation as one atomic znode (all-or-nothing; no partial application — this is exactly why heavy `ALTER TABLE ... UPDATE` does not exhibit this). The finding here is a different code path (the per-partition patch-part INSERT sinks), a different fault (a clean, unambiguous commit failure — no lost Keeper response needed), and has a facet #112133 does not: the partial application is read-visible after the failed op even with zero retries. ### Does it reproduce on the most recent release? Yes — reproduced on a current `26.8.1.1` master build; the mechanism is present in the current source. It reproduces at both the wide-part layout and the default (compact) part layout. ### How to reproduce ```sql CREATE TABLE t (id UInt64, x UInt64) ENGINE = MergeTree PARTITION BY (id % 2) ORDER BY id SETTINGS enable_block_number_column = 1, enable_block_offset_column = 1; INSERT INTO t SELECT number, 0 FROM numbers(200); -- 100 rows in each of 2 partitions, x = 0 ``` Inject a filesystem I/O error (`EIO`) on the rename that commits the **second** partition's patch part (the second `tmp_insert_patch-*` → `patch-*` rename), then run: ```sql UPDATE t SET x = x + 1 WHERE x < 1000000; -- errors with a filesystem_error ``` Observe that the update failed but one partition already reflects the change: ```sql SELECT id % 2 AS p, sum(x) FROM t GROUP BY p ORDER BY p; -- p=0 -> 100 (patch committed before the fault: durable residue of a failed UPDATE) -- p=1 -> 0 ``` Now retry the same statement (fault cleared, as a client would): ```sql UPDATE t SET x = x + 1 WHERE x < 1000000; -- succeeds SELECT id % 2 AS p, sum(x) FROM t GROUP BY p ORDER BY p; -- p=0 -> 200 (double-applied) -- p=1 -> 100 ``` This was verified with a deterministic control-gated harness that arms the `EIO` on exactly the second patch-part commit rename. Controls: with no fault the single `UPDATE` yields a uniform `+1` on both partitions; a single-partition table under the same fault leaves no residue and the retry yields a uniform `+1` (so the split is attributable to the multi-partition fan-out). ### Expected behavior A failed `UPDATE` should apply to no partition (atomic), and a retry of the same `UPDATE` should be idempotent — the final state should be `x + 1` on every partition regardless of the mid-commit failure and the retry. ### Error message and/or stacktrace ``` Code: 1001. std::exception. Code: 1001, type: std::__1::filesystem::filesystem_error, e.what() = filesystem error: in rename: Input/output error [...patch-...] ``` ### Additional context - Commit loops: `src/Storages/MergeTree/MergeTreeSinkPatch.cpp` (`finishDelayedChunk`, per-partition loop) and `src/Storages/MergeTree/ReplicatedMergeTreeSinkPatch.cpp` (`finishDelayed`, per-partition loop). - `deduplicate = false` is hard-set for patch parts, and a deduplicated patch part is itself treated as a `LOGICAL_ERROR` (\"Patch part {} was deduplicated. It's a bug\"), so idempotent retry is unattainable by design — atomicity is the only possible containment, and it is absent. - Lightweight `UPDATE` is a beta feature; its documentation states the query waits for patch-part creation before returning but says nothing about failure atomicity or retry idempotency. - Reachability: the only non-default requirement is `enable_block_number_column` / `enable_block_offset_column`, which are the documented prerequisites for lightweight updates, not an experimental gate.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114592",
        "createdAt": "2026-08-13T05:02:41Z",
        "updatedAt": "2026-08-13T05:02:41Z",
        "timestamp": "2026-08-13T05:02:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114593",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "RESTORE: a mid-attach I/O error leaves an un-rolled-back durable data prefix; the suggested allow_non_empty_tables retry duplicates it",
        "text": "### Company or project name ClickHouse QA (internal durability testing) ### Describe what's wrong An I/O error partway through `RESTORE` leaves already-attached parts Active and read-visible, with no rollback, while `RESTORE` reports failure — and the server's own error message then steers the operator into silently duplicating that residue. `StorageReplicatedMergeTree::attachRestoredParts` attaches a backup's parts in a bare per-part loop of `sink->writeExistingPart(part, /*deduplicate_part*/ false)` with no try/compensation around it. Each part commits by renaming its `tmp_restore_*` staging directory into the final active name. If the rename of the k-th part fails (for example a transient filesystem error), the parts committed before it stay Active and readable, but `RESTORE` fails and reports the table could not be restored. The failed operation has left a durable, read-visible data prefix. The natural next step makes it worse. Because the table is now non-empty, a retry of the same `RESTORE` fails with `CANNOT_RESTORE_TABLE ... use allow_non_empty_tables=true`, and following that suggestion re-attaches every part with fresh block numbers (dedup off), duplicating the rows that the failed restore had already left behind. The operator is guided by the server's own message into duplicating the residue of a failed restore. ### Does it reproduce on the most recent release? Yes — reproduced on a current `26.8.1.1` master build; `attachRestoredParts` is a bare per-part `writeExistingPart` loop with no rollback in the current source. ### How to reproduce On a `ReplicatedMergeTree` table with several partitions (so RESTORE attaches several parts): ```sql CREATE TABLE rc (id UInt64, v String) ENGINE = ReplicatedMergeTree('/clickhouse/tables/rc','r0') PARTITION BY (id % 3) ORDER BY id; INSERT INTO rc SELECT number, toString(number) FROM numbers(300); -- 300 rows, 3 parts BACKUP TABLE rc TO File('rc'); DROP TABLE rc SYNC; CREATE TABLE rc (id UInt64, v String) ENGINE = ReplicatedMergeTree('/clickhouse/tables/rc','r0') PARTITION BY (id % 3) ORDER BY id; -- empty target ``` Inject a filesystem I/O error (`EIO`) on the rename that commits the **second** restored part (the second `tmp_restore_*` rename), then: ```sql RESTORE TABLE rc FROM File('rc'); -- fails SELECT count() FROM rc; -- 100 (a durable, read-visible prefix of a failed RESTORE) ``` Retry as the server's error message suggests: ```sql RESTORE TABLE rc FROM File('rc') SETTINGS allow_non_empty_tables = 1; -- succeeds SELECT count() FROM rc; -- 400 (300 unique rows + 100 duplicated from the residue) ``` Verified with a deterministic control-gated harness that arms the `EIO` on exactly the second restored part's commit rename. Controls: a clean `RESTORE` into an empty table round-trips exactly 300 rows; a single-part backup under the same fault leaves nothing (no prefix) and the retry restores exactly 300 — so the prefix is attributable to the multi-part attach fan-out. ### Expected behavior A `RESTORE` that fails partway through should leave the table unchanged (no durable, read-visible prefix from a failed restore), so that a retry restores exactly the backup's contents without duplication. At minimum, the residue of a failed `RESTORE` should not be counted as pre-existing data that the operator is then told to `allow_non_empty_tables` over. ### Error message and/or stacktrace The `RESTORE` fails with the injected `filesystem error: in rename: Input/output error`; the retry into the now-non-empty table fails with `CANNOT_RESTORE_TABLE` and the message suggesting `allow_non_empty_tables=true` (`src/Backups/RestorerFromBackup.cpp`, `throwTableIsNotEmpty`). ### Additional context - No-rollback attach loop: `StorageReplicatedMergeTree::attachRestoredParts` — per-part `writeExistingPart(part, deduplicate_part=false)`, no compensation. - Distinct from #104464 (RESTORE leaves an orphan **empty** table after a mid-restore **crash** — opposite residue: no data; this finding is a no-crash I/O failure that leaves a **data** prefix plus retry duplication). #104464 being accepted as a bug is precedent that \"RESTORE is not atomic on failure\" is treated as a defect. - Distinct from #114266 (duplication via the emptiness guard reading local state on a lagging replica, where RESTORE **succeeds**) and #112149 (INSERT-sink cancellation during an ambiguous commit — data loss, not residue). - The `allow_non_empty_tables` duplication is documented for legitimately pre-existing data (`src/Backups/RestoreSettings.h`). The defect reported here is the un-rolled-back durable prefix from a **failed** restore; the documented duplication is the consequence the operator is steered into, not the primary bug.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114593",
        "createdAt": "2026-08-13T05:02:42Z",
        "updatedAt": "2026-08-13T05:02:42Z",
        "timestamp": "2026-08-13T05:02:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114595",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "S3Queue (multi-server, hash-ring): a stale per-server Processed cache silently skips a re-appeared object after its tracked-file TTL expires",
        "text": "### Company or project name ClickHouse QA (internal durability testing) ### Describe what's wrong In a multi-server `S3Queue`/`ObjectStorageQueue` setup with `enable_hash_ring_filtering = 1`, a per-server in-memory \"Processed\" cache silently skips a re-appeared object after that object's tracked-file record has been aged out of Keeper by a different server. The row is never ingested — no error, no failed file. The pieces: - `ObjectStorageQueueIFileMetadata::trySetProcessing()` returns `false` immediately when the in-memory per-server `local_file_statuses` entry for a path is `Processed`, **without reading Keeper** (the Keeper check in `setProcessingImpl` is only reached after this in-memory guard). So a server that has already processed key `K` will not even ask Keeper about `K` again. - `tracked_file_ttl_sec` cleanup removes a file's `/processed` node from Keeper, and it removes the file from `local_file_statuses` **only on the server that runs the cleanup** (`ObjectStorageQueueMetadata`, the TTL cleanup path). The other servers keep their stale in-memory `Processed` entry. - `enable_hash_ring_filtering` routes each key to exactly one server (`filterOutForProcessor` → `chooseServer(hash(path))`). Put together: server X processes `K` (X's cache = `Processed`, Keeper `/processed/K` created). Server Y later runs the TTL cleanup and deletes `/processed/K` from Keeper (and Y's own cache entry, which never had `K`). The object at key `K` is then rewritten with new content (a legitimate same-key overwrite). The hash ring still routes `K` to X. X consults its in-memory cache, sees `Processed`, and skips `K` — never reaching Keeper, where the record no longer exists. The rewritten, producer-acked content is silently dropped. ### Does it reproduce on the most recent release? Yes — reproduced 2/2 on a current `26.8.1.1` master build. ### How to reproduce Two ClickHouse servers sharing one Keeper and one bucket, each with an `S3Queue` table on the same `keeper_path` and a local target table via a materialized view: ```sql CREATE TABLE q (id UInt64, val String) ENGINE = S3Queue('http://minio:9000/bucket/*', 'JSONEachRow') SETTINGS mode = 'unordered', enable_hash_ring_filtering = 1, tracked_file_ttl_sec = 4, processing_threads_num = 1; ``` On server X set the cleanup interval high (e.g. `cleanup_interval_min_ms = cleanup_interval_max_ms = 3600000`); on server Y set it low (e.g. `500`/`1000`) so Y deterministically drives the TTL cleanup. (These are per-server local settings, not part of the shared table metadata.) 1. Upload objects; observe (via each server's own target table) which keys the hash ring routes exclusively to X. Pick such a key `K`. 2. Wait for Y's cleanup to age out `/processed` (confirm the `/processed` children are gone via `system.zookeeper`). 3. Re-upload key `K` with new content (a sentinel row). Confirm the object with the new content is present in the bucket. 4. The sentinel never appears in either server's target table — silent loss. Control (proves it is the stale in-memory cache, not a genuine Keeper state): repeat with a different X-routed key, but restart server X (clearing its in-memory `local_file_statuses`) before the re-upload — the rewritten object is then ingested. ### Expected behavior Once a file's tracked record has been aged out of Keeper, a re-appeared object with that key must be eligible for processing by whichever server owns it. The per-server in-memory `Processed` cache must not short-circuit past a Keeper record that no longer exists — `trySetProcessing` should confirm against Keeper before declining a file whose local state is `Processed`, or the TTL cleanup must invalidate the entry on all servers, not only the cleaning one. ### Error message and/or stacktrace _No error_ — the file is silently not processed; `system.s3queue`/`s3queue_log` show no failure. ### Additional context - In-memory short-circuit: `src/Storages/ObjectStorageQueue/ObjectStorageQueueIFileMetadata.cpp` (`trySetProcessing`, the `Processed` early-return before any Keeper request). - Cache removal only on the cleaning server: `src/Storages/ObjectStorageQueue/ObjectStorageQueueMetadata.cpp` (TTL cleanup `local_file_statuses.remove`). - Hash-ring routing: `filterOutForProcessor` / `chooseServer`. - Distinct from #112341 (single-server Processed-while-discarding on MV detach — no ageout, no cross-server cache asymmetry), #111537 (power-loss at-least-once → at-most-once), #112041 (DatabaseReplicated recovery drops the registry), and #109751 (Azure listing continuation-token truncation). This is specifically the stale per-server `Processed` cache winning over an aged-out Keeper record under hash-ring routing. - Reachability: `enable_hash_ring_filtering` is the multi-server distribution mode and `tracked_file_ttl_sec` is the documented tracked-files bound; the trigger is a same-key overwrite.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114595",
        "createdAt": "2026-08-13T05:21:03Z",
        "updatedAt": "2026-08-13T05:21:03Z",
        "timestamp": "2026-08-13T05:21:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114598",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "PREWHERE Map subcolumn costing performs full-table metadata I/O before pruning",
        "text": "### Company or project name _No response_ ### Describe the situation After upgrading from 26.6.1 to 26.7, queries using Map subcolumns on a MergeTree became very slow on their first execution. This appears related to [#110623](https://github.com/ClickHouse/ClickHouse/pull/110623). PREWHERE costing now aggregates requested subcolumn sizes across every active part before partition/primary-key pruning. This causes sequential full-table-scale metadata work. Local disks incur per-part prefix/file reads; with an S3-backed MergeTree, these become remote requests. One example query touching S3-backed MergeTree: Profile event | First run | Second run -- | -- | -- QueryPlanBuildMicroseconds | 48,590,413 | 69,982 QueryPlanOptimizeMicroseconds | 48,573,662 | 39,756 SharedPartsLockHoldMicroseconds | 48,561,615 | 28,923 S3ReadMicroseconds | 46,118,276 | 124,920 ReadBufferFromS3InitMicroseconds | 46,358,270 | 0 S3ReadRequestsCount | 1,270 | 11 During the first run, query progress remains at zero for most of the time because this occurs during planning. The identical second run is fast because the per-part subcolumn-size metadata is cached. cc @Avogar ### Which ClickHouse versions are affected? 26.7 ### How to reproduce _ ### Expected performance _No response_ ### Related issues and pull requests _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114598",
        "createdAt": "2026-08-13T06:06:47Z",
        "updatedAt": "2026-08-13T06:07:33Z",
        "timestamp": "2026-08-13T06:07:33Z",
        "metrics": {
          "reactions": 1,
          "comments": 0
        },
        "labels": [
          "performance"
        ],
        "author": "starpact",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114603",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Logical error: Unexpected token for lazy mode: A. Multi-block postings must be compressed (STID: 4250-5377)",
        "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Unexpected token for lazy mode: A. Multi-block postings must be compressed (STID: 4250-5377) CI report: [AST fuzzer (amd_debug)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=3b9bc037faeea25fb3dc2841ec3f87fbf3f2d8f8&name_0=MasterCI&name_1=AST%20fuzzer%20%28amd_debug%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBVbmV4cGVjdGVkIHRva2VuIGZvciBsYXp5IG1vZGU6IEEuIE11bHRpLWJsb2NrIHBvc3RpbmdzIG11c3QgYmUgY29tcHJlc3NlZCAoU1RJRDogNDI1MC01Mzc3KScKICAgIC0tIEFORCBjaGVja19uYW1lID0gJ0FTVCBmdXp6ZXIgKGFtZF9kZWJ1ZyknCiAgICBBTkQgdGVzdF9zdGF0dXMgSU4gKCdGQUlMJywgJ0VSUk9SJykKICAgIEFORCAocHVsbF9yZXF1ZXN0X251bWJlciA9IDAgT1IgYmFzZV9yZWYgSU4gKCcnLCAnbWFzdGVyJykpCiAgICBBTkQgdGVzdF9jb250ZXh0X3JhdyBMSUtFICclRXhjZXB0aW9uOiUnCkdST1VQIEJZIGRheQpPUkRFUiBCWSBkYXkgREVTQwo=) Test output: ``` Error: Logical error: 'Unexpected token for lazy mode: zrare. Multi-block postings must be compressed'. --- Failed query: SELECT DISTINCT count() FROM cluster('test_unavailable_shard', currentDatabase(), 'tab_postings_order__fuzz_12') PREWHERE hasAllTokens(s, ['filler', 'zrare']) WHERE hasAllTokens(s, ['zra\\0e', 'filler']) --- Reproduce commands (auto-generated; may require manual adjustment): SELECT DISTINCT count() FROM cluster('test_unavailable_shard', currentDatabase(), 'tab_postings_order__fuzz_12') PREWHERE hasAllTokens(s, ['filler', 'zrare']) WHERE hasAllTokens(s, ['zra\\0e', 'filler']); DROP TABLE IF EXISTS cluster('test_unavailable_shard', currentDatabase(), 'tab_postings_order__fuzz_12'); --- Stack trace: pthread_kill @ 0x00000000000969bd gsignal @ 0x0000000000042476 __lgamma_r_finite@GLIBC_2.15 @ 0x00000000000287f3 src/Common/Exception.cpp:66:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x00000000143ff02e src/Common/Exception.cpp:115:13: DB::Exception::handleErrorCode(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x000000001440006c src/Common/Exception.cpp:171:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x00000000144004ac src/Common/Exception.h:201:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d6533d6 src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d652e9e src/Common/Exception.h:219:77: DB::Exception::Exception<std::basic_string_view<char, std::char_traits<char>>&>(int, FormatStringHelperImpl<std::type_identity<std::basic_string_view<char, std::char_traits<char>>&>::type>, std::basic_string_view<char, std::char_traits<char>>&) @ 0x000000000df7df83 src/Storages/MergeTree/MergeTreeReaderTextIndex.cpp:370:15: DB::MergeTreeReaderTextIndex::makeLazyCursor(std::basic_string_view<char, std::char_traits<char>>, DB::TokenPostingsInfo const&) @ 0x000000001e281f96 src/Storages/MergeTree/MergeTreeReaderTextIndex.cpp:793:30: DB::MergeTreeReaderTextIndex::fillColumnLazy(DB::IColumn&, unsigned long, unsigned long, unsigned long, roaring::Roaring&) @ 0x000000001e285955 src/Storages/MergeTree/MergeTreeReaderTextIndex.cpp:524:17: DB::MergeTreeReaderTextIndex::readRows(unsigned long, bool, unsigned long, unsigned long, std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) @ 0x000000001e28325a src/Storages/MergeTree/MergeTreeRangeReader.cpp:187: DB::MergeTreeRangeReader::DelayedStream::readRows(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&, unsigned long) src/Storages/MergeTree/MergeTreeRangeReader.cpp:259:47: DB::MergeTreeRangeReader::DelayedStream::finalize(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) @ 0x000000001e25f0d7 src/Storages/MergeTree/MergeTreeRangeReader.cpp:375: DB::MergeTreeRangeReader::Stream::finalize(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) src/Storages/MergeTree/MergeTreeRangeReader.cpp:1201:31: DB::MergeTreeRangeReader::startReadingChain(unsigned long, DB::MarkRanges&) @ 0x000000001e268190 src/Storages/MergeTree/MergeTreeReadersChain.cpp:296:36: DB::MergeTreeReadersChain::read(unsigned long, DB::MarkRanges&, std::vector<DB::MarkRanges, std::allocator<DB::MarkRanges>>&, std::function<void (std::vector<DB::ColumnWithTypeAndName, AllocatorWithMemoryTracking<DB::ColumnWithTypeAndName>> const&, std::unordered_set<String, std::hash<String>, std::equal_to<String>, std::allocator<String>> const&, unsigned long, std::optional<bool>&)> const&) @ 0x000000001e2bb9ba src/Storages/MergeTree/MergeTreeReadTask.cpp:442:38: DB::MergeTreeReadTask::read() @ 0x000000001e2b82de src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:270:31: DB::MergeTreeSelectProcessor::readCurrentTask(DB::MergeTreeReadTask&, DB::IMergeTreeSelectAlgorithm&) const @ 0x000000001e2c6e27 src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:477:23: DB::MergeTreeSelectProcessor::read() @ 0x000000001e2c906d src/Storages/MergeTree/MergeTreeSource.cpp:233:41: DB::MergeTreeSource::tryGenerate() @ 0x000000001f16ae23 src/Processors/ISource.cpp:119:26: DB::ISource::work() @ 0x000000001ea558c1 src/Processors/Executors/ExecutionThreadContext.cpp:57: DB::executeJob(DB::ExecutingGraph::Node*, DB::ReadProgressCallback*) src/Processors/Executors/ExecutionThreadContext.cpp:127:28: DB::ExecutionThreadContext::executeTask() @ 0x000000001ea7715e src/Processors/Executors/PipelineExecutor.cpp:387:26: DB::PipelineExecutor::executeStepImpl(unsigned long, DB::WorkloadResources&&, std::atomic<bool>*) @ 0x000000001ea694e8 src/Processors/Executors/PipelineExecutor.cpp:355:5: DB::PipelineExecutor::executeSingleThread(unsigned long, DB::WorkloadResources&&) @ 0x000000001ea69bac src/Processors/Executors/PipelineExecutor.cpp:692: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>::type std::__invoke[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:17: ? @ 0x000000001ea6cae7 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000014613869 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1293: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x000000001461c8d2 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x0000000014610a1e contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x0000000014619c8e start_thread @ 0x0000000000094a83 __GI___clone3 @ 0x0000000000126890 ```",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114603",
        "createdAt": "2026-08-13T08:37:23Z",
        "updatedAt": "2026-08-13T09:48:01Z",
        "timestamp": "2026-08-13T09:48:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "testing",
          "fuzz"
        ],
        "author": "PedroTadim",
        "state": "open",
        "assignees": [
          "CurtizJ"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114605",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Infinite uncancellable loop in function `hop` at analysis time: a window interval whose span wraps to 0 modulo 2^32 dodges the time-overflow guard",
        "text": "🕵 Found by the AST fuzzer in the stress test of a CI run ([Stress test (arm_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112930&sha=1c5bf0b613146bd6666b6d2bd2be55f539f6c0d8&name_0=PR&name_1=Stress%20test%20%28arm_debug%29), `Hung check failed, possible deadlock found`): the fuzzed query hung for 1338 s with `is_cancelled: 1` and was still spinning when the hung check gave up. Reproducer (hangs forever, also in `clickhouse local` on current master): ```sql SELECT hop(toDateTime32('1969-12-31'), toIntervalDay(1), toIntervalDay(2147483648), 'US/Samoa') ``` Because the arguments are constants, the loop runs during constant folding in `QueryAnalyzer::resolveFunction`, i.e. at analysis time, where nothing checks the cancellation - the query cannot be killed (`KILL QUERY` marks it cancelled, but the thread never looks). The stack of the hung thread is `TimeWindowImpl<HOP>::dispatchForColumns` <- `IExecutableFunction::execute` <- `QueryAnalyzer::resolveFunction` <- ... <- `InterpreterExplainQuery::execute` (the fuzzed query was an `EXPLAIN`). The mechanism is in `executeHop` (`src/Functions/FunctionsTimeWindow.cpp`). `AddTime<Day>` is a wrapping `UInt32` computation, `static_cast<UInt32>(t + delta * 86400)`. With `window_num_units = 2147483648 = 2^31`, the subtraction of the window is `2^31 * 86400 = 43200 * 2^32 ≡ 0 (mod 2^32)`, so ```cpp wstart = AddTime<kind>::execute(wend, -window_num_units, time_zone); // wstart == wend: subtracting the window is a no-op modulo 2^32 if (wstart > wend) throw Exception(ErrorCodes::BAD_ARGUMENTS, \"Time overflow in function {}\", name); // does not fire: equal, not greater ``` the time-overflow guard added by https://github.com/ClickHouse/ClickHouse/pull/61523 (for the same class of hang, https://github.com/ClickHouse/ClickHouse/issues/61521) does not fire. The following loop ```cpp do { wend_latest = wend; wend = static_cast<ToType>(AddTime<kind>::execute(wend, -hop_num_units, time_zone)); } while (wend > time_data[i]); ``` then decrements `wend` by one day at a time past zero, where it wraps back to ~2^32. Every value it takes stays congruent to the starting point modulo `gcd(86400, 2^32) = 128`, so it never hits a value `<= time_data[i]` exactly when the start is not congruent to a small value, and the loop never terminates. The same loop shape exists in `executeHopSlice` (`windowID`). The guard needs to catch a wrapped subtraction that lands exactly on (or above) `wend`, and the loop needs to detect the wrap of `wend -= hop` (the new `wend` coming out greater than the old one) instead of relying on the comparison with the time.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114605",
        "createdAt": "2026-08-13T09:03:36Z",
        "updatedAt": "2026-08-13T09:03:36Z",
        "timestamp": "2026-08-13T09:03:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug",
          "fuzz"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114611",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Logical error: Invalid number of columns in chunk pushed to OutputPort. Expected A, found B (STID: 2270-3258)",
        "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Invalid number of columns in chunk pushed to OutputPort. Expected A, found B (STID: 2270-3258) CI report: [AST fuzzer (amd_debug)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=108522&sha=da345caa0c75914b7749446668b09bb6ca9355e5&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBJbnZhbGlkIG51bWJlciBvZiBjb2x1bW5zIGluIGNodW5rIHB1c2hlZCB0byBPdXRwdXRQb3J0LiBFeHBlY3RlZCBBLCBmb3VuZCBCIChTVElEOiAyMjcwLTMyNTgpJwogICAgLS0gQU5EIGNoZWNrX25hbWUgPSAnQVNUIGZ1enplciAoYW1kX2RlYnVnKScKICAgIEFORCB0ZXN0X3N0YXR1cyBJTiAoJ0ZBSUwnLCAnRVJST1InKQogICAgQU5EIChwdWxsX3JlcXVlc3RfbnVtYmVyID0gMCBPUiBiYXNlX3JlZiBJTiAoJ21hc3RlcicpKQogICAgQU5EIHRlc3RfY29udGV4dF9yYXcgTElLRSAnJUV4Y2VwdGlvbjolJwpHUk9VUCBCWSBkYXkKT1JERVIgQlkgZGF5IERFU0MK) Test output: ``` Error: Logical error: 'Invalid number of columns in chunk pushed to OutputPort. Expected 0, found 1 Header: Chunk: String(size = 1) '. --- Failed query: SELECT '-92233\\02036854775808', rank() OVER (RANGE BETWEEN UNBOUNDED PRECEDING AND UNBOUNDED FOLLOWING) FROM remote('127.0.0.{2..3}:9000', view(SELECT DISTINCT '1' AS c0 LIMIT 318)) AS v0 QUALIFY rank() OVER (ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW) > 0 LIMIT 511 --- Stack trace: pthread_kill @ 0x00000000000969bd gsignal @ 0x0000000000042476 __ieee754_lgamma_r @ 0x00000000000287f3 src/Common/Exception.cpp:66:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x00000000143ef46e src/Common/Exception.cpp:115:13: DB::Exception::handleErrorCode(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x00000000143f04ac src/Common/Exception.cpp:171:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x00000000143f08ec src/Common/Exception.h:201:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d64c0d6 src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d64bb9e src/Common/Exception.h:219:77: DB::Exception::Exception<unsigned long, unsigned long, String, String>(int, FormatStringHelperImpl<std::type_identity<unsigned long>::type, std::type_identity<unsigned long>::type, std::type_identity<String>::type, std::type_identity<String>::type>, unsigned long&&, unsigned long&&, String&&, String&&) @ 0x000000001a484b4e src/Processors/Port.h:431: DB::OutputPort::pushData(DB::Port::State::Data) src/Processors/ISource.cpp:49:19: DB::ISource::prepare() @ 0x000000001ea44754 src/Processors/Sources/RemoteSource.cpp:102:30: DB::RemoteSource::prepare() @ 0x000000001eed2ef9 src/Processors/Executors/ExecutingGraph.cpp:384:59: DB::ExecutingGraph::updateNode(DB::ExecutingGraph::Node*, std::queue<DB::ExecutingGraph::Node*, boost::container::devector<DB::ExecutingGraph::Node*, AllocatorWithMemoryTracking<DB::ExecutingGraph::Node*>, void>>&, std::queue<DB::ExecutingGraph::Node*, boost::container::devector<DB::ExecutingGraph::Node*, AllocatorWithMemoryTracking<DB::ExecutingGraph::Node*>, void>>&) @ 0x000000001ea5eb4d src/Processors/Executors/PipelineExecutor.cpp:407:38: DB::PipelineExecutor::executeStepImpl(unsigned long, DB::WorkloadResources&&, std::atomic<bool>*) @ 0x000000001ea5888c src/Processors/Executors/PipelineExecutor.cpp:355:5: DB::PipelineExecutor::executeSingleThread(unsigned long, DB::WorkloadResources&&) @ 0x000000001ea58eac src/Processors/Executors/PipelineExecutor.cpp:692: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>::type std::__invoke[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:17: ? @ 0x000000001ea5bde7 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000014603329 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1293: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x000000001460c392 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x00000000146004de contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x000000001460974e start_thread @ 0x0000000000094a83 __clone3 @ 0x0000000000126890 ```",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114611",
        "createdAt": "2026-08-13T09:20:40Z",
        "updatedAt": "2026-08-13T09:48:04Z",
        "timestamp": "2026-08-13T09:48:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "testing",
          "fuzz"
        ],
        "author": "kssenii",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114612",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Heap-use-after-free: Parquet v3 prefetcher reads and writes through a `ReadBuffer` freed by `IInputFormat::onFinish`",
        "text": "🕵️ ## Describe what's wrong `ParquetV3BlockInputFormat` does not override `resetReadBuffer()`, so `IInputFormat::onFinish()` frees the format's owned `ReadBuffer` while the Parquet `Prefetcher`'s IO tasks are still reading through it on the prefetch thread pool. The tasks then dereference freed memory. ASan on a plain `clickhouse local` run of the reproducer below (master, 26.8.1.1310): ``` ==ERROR: AddressSanitizer: heap-use-after-free on address 0x7112bd001800 WRITE of size 1048576 at 0x7112bd001800 thread T9 (ParquetPrefetch) #3 in DB::ReadBuffer::next() src/IO/ReadBuffer.cpp:113:15 #5 in DB::ReadBuffer::read(char*, unsigned long) src/IO/ReadBuffer.h:169:37 #6 in DB::Parquet::Prefetcher::readSync(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:113:29 #7 in DB::Parquet::Prefetcher::runTask(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:540:13 #8 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:435 #15 in DB::ThreadPoolCallbackRunnerFast::threadFunction() src/Common/threadPoolCallbackRunner.cpp:225:13 ``` It is a **1 MiB write into freed heap**, so this is memory corruption, not only a bad read. The same defect was hit independently by the `La Casa Del Dolor (arm_asan_ubsan)` CI job, where the report is a read and carries the full free/allocation stacks: ``` ERROR: AddressSanitizer: heap-use-after-free on address 0xfcdd2aec39c0 READ of size 8 at 0xfcdd2aec39c0 thread T423 (ThreadPool) #0 in DB::Parquet::Prefetcher::readSync(...) Prefetcher.cpp:110:21 <- reader->setReadUntilEnd() #1 in DB::Parquet::Prefetcher::runTask(...) Prefetcher.cpp:523:13 #2 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() Prefetcher.cpp:424:17 #9 in DB::ThreadPoolCallbackRunnerFast::threadFunction() threadPoolCallbackRunner.cpp:225:13 0xfcdd2aec39c0 is located 0 bytes inside of 272-byte region freed by thread T468 (ThreadPool) here: #1 in std::default_delete<DB::ReadBuffer>::operator()(DB::ReadBuffer*) #7 in std::vector<std::unique_ptr<DB::ReadBuffer>>::~vector() #8 in DB::ISource::work() src/Processors/ISource.cpp:140:13 ... #18 in DB::IPolygonDictionary::loadData() src/Dictionaries/PolygonDictionary.cpp:337:8 previously allocated by thread T468 (ThreadPool) here: #2 in DB::FileDictionarySource::loadAll() src/Dictionaries/FileDictionarySource.cpp:60:19 ``` `ISource.cpp:140` is the `onFinish()` call inside the `catch (...)` handler, and the freed 272-byte region is the `ReadBufferFromFile` that `FileDictionarySource::loadAll` handed to the format via `addBuffer`. ## How to reproduce Reproduces on master (26.8.1.1310) with the official `build_amd_asan_ubsan` binary, **8 out of 10 runs, with no ThreadFuzzer and no unusual settings**. Under ThreadFuzzer it hit on the first run. Run Fiddle: https://fiddle.clickhouse.com/06f26a34-d286-40a1-85f5-7d3e952d47b8 On a release build this prints only the load error and exits 53: `Code: 53. CAST AS Array can only be performed between same-dimensional Array ... While executing ParquetV3BlockInputFormat. (TYPE_MISMATCH)` — the 1 MiB write into freed memory is silent. On an ASan build the process dies with the report above instead (5/5 runs of exactly the commands above; 8/10 in an earlier variant, so it is a race, but a very wide one). The polygon dictionary is only a convenient way to make the pipeline throw mid-read while owning its `ReadBuffer` — the defect is in the format, not in the dictionary. ## Root cause `Prefetcher` protects only *its own* lifetime. `scheduleTask` captures the shutdown handle and each task takes `std::shared_lock(*_shutdown, std::try_to_lock)`; `~Prefetcher` calls `shutdown->shutdown()` (`ShutdownHelper`, `src/Common/threadPoolCallbackRunner.h:507`), which blocks until every in-flight task has released. As its own comment says, \"`this` is safe to access as long as `shutdown_lock` is held\". But `Prefetcher::reader` points at a `ReadBuffer` owned by `IInputFormat::owned_buffers`, an unrelated lifetime that the handshake does not cover: - `IInputFormat::onFinish()` -> `resetReadBuffer()` -> `resetOwnedBuffers()` -> `owned_buffers.clear()` - `ParquetV3BlockInputFormat` overrides `resetParser()` and `onCancel()`, but **not `resetReadBuffer()`**. `resetParser()` happens to be safe only because it destroys the reader *before* delegating to the base. `resetReadBuffer()` has no such ordering, so the buffer dies while the `ReadManager` -> `Reader` -> `Prefetcher` chain is still alive with tasks running. Three routes reach it: `onFinish()` from the `catch (...)` in `ISource::work()` (the one above), `onFinish()` on the normal completion path (`ISource.cpp:133`) with speculative prefetches still outstanding, and `IInputFormat::setReadBuffer` when a format is reused across files. ## Suggested fix Mirror what `resetParser()` already does, so `~Prefetcher` drains the IO tasks before the base frees the buffer: ```cpp void ParquetV3BlockInputFormat::resetReadBuffer() { /// ~Prefetcher waits for in-flight IO tasks, which read through the buffer that /// IInputFormat::resetReadBuffer() is about to free. { std::lock_guard lock(reader_mutex); reader.reset(); } IInputFormat::resetReadBuffer(); } ``` ## Additional context Distinct from #109678 / #112573. That was a data race on `AsynchronousBoundedReadBuffer::prefetch_future` consumed by concurrent `readBigAt` (the `RandomRead` branch, `Prefetcher.cpp:102`). This one is a lifetime bug in the `SeekAndRead` branch, which already holds `read_mutex` — locking cannot help once the object is freed — and the freed buffer here is a plain `ReadBufferFromFile`, so `AsynchronousBoundedReadBuffer` is not involved at all. The binary used above already contains #112573 (merged into 26.8.1.1240). Found by this Dolor run: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=94148&sha=a2ef409981fc77a168eaf8cc68d7ddf67c9fadec&name_0=PR&name_1=La+Casa+Del+Dolor+%28arm_asan_ubsan%29",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114612",
        "createdAt": "2026-08-13T09:21:28Z",
        "updatedAt": "2026-08-13T16:56:05Z",
        "timestamp": "2026-08-13T16:56:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "bug",
          "blocker",
          "comp-parquet-reader-v3"
        ],
        "author": "PedroTadim",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114616",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Row policy over a file-backed table breaks `DEFAULT` columns missing from the data file (`UNKNOWN_IDENTIFIER`; silently wrong results on 26.7)",
        "text": "🕵 Found while working on https://github.com/ClickHouse/ClickHouse/pull/114262 (lazy materialization for local Parquet files); the behavior is identical with that PR's optimization on or off, and reproduces on current master without it. **Describe what's wrong** When a table over a data file (`File(Parquet)`, and by code inspection the same applies to the object-storage path) has a `DEFAULT` column that is missing from the file, and a row policy references either that column or interacts with reading it, the query fails with `UNKNOWN_IDENTIFIER` — the source prunes the `DEFAULT` expression's input columns before `AddingDefaultsTransform` can compute the column. On the official 26.7 build, the first case below returned **silently wrong results** (0 rows instead of 900): the policy was evaluated against type defaults instead of real row values. Current master turned that into a fail-safe exception, which is safer, but the queries are legitimate and should work. **How to reproduce** ```bash clickhouse local ``` ```sql INSERT INTO FUNCTION file('rp.parquet', Parquet) SELECT number AS k, number % 10 AS a, concat('val_', toString(number)) AS s FROM numbers(1000) SETTINGS engine_file_truncate_on_insert = 1; -- Case 1: the row policy references a DEFAULT column missing from the file. CREATE TABLE t_rp_j (k UInt64, a UInt64, s String, j JSON DEFAULT toJSONString(map('user', map('name', concat('u', toString(a)))))) ENGINE = File(Parquet, 'rp.parquet'); CREATE ROW POLICY pol_j ON t_rp_j USING j.user.name != 'u0' TO ALL; SELECT k, s FROM t_rp_j ORDER BY k LIMIT 3; -- Code: 47. DB::Exception: Unknown expression or function identifier `a` in scope -- _CAST(toJSONString(map('user', map('name', concat('u', toString(a))))), 'JSON') AS j ... While executing File. -- Official 26.7 instead returns 0 rows (silently wrong: expected 900 rows admitted by the policy). -- Case 2: the policy is on a plain file column, and the query selects a DEFAULT -- column depending on another file column, together with a PREWHERE. CREATE TABLE t_rp_d (k UInt64, a UInt64, s String, d UInt64 DEFAULT a * 2) ENGINE = File(Parquet, 'rp.parquet'); CREATE ROW POLICY pol_d ON t_rp_d USING a != 0 TO ALL; SELECT k, d FROM t_rp_d PREWHERE s != 'val_2' ORDER BY k LIMIT 3; -- Code: 47. DB::Exception: Unknown expression or function identifier `a` in scope _CAST(a * 2, 'UInt64') AS d ... -- (the same statement without PREWHERE works and returns 1→2, 2→4, 3→6) ``` **Expected behavior** Both queries return the same result as with the row policy applied above the read (e.g. as `MergeTree` does): case 1 → 900 rows admitted by the policy computed from real `a` values; case 2 → `d` computed from `a` after prewhere filtering. **Where it comes from** `updateFormatPrewhereInfo` / `SourceStepWithFilter::applyPrewhereActions` shrink the format header to the filter DAG's outputs, dropping input columns that a `DEFAULT` expression of another (missing) column needs; `AddingDefaultsTransform` inside the source then cannot compute the default. For case 1 the policy's own input is a missing defaulted column, so its dependency `a` is never part of the read set to begin with. **Additional context** Related: https://github.com/ClickHouse/ClickHouse/pull/114262 Related: https://github.com/ClickHouse/ClickHouse/pull/110970",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114616",
        "createdAt": "2026-08-13T09:56:28Z",
        "updatedAt": "2026-08-13T09:56:28Z",
        "timestamp": "2026-08-13T09:56:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "potential bug"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114630",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "KeyCondition's pointInPolygon primary-key analysis skips is_valid validation (unlike pointInPolygon/spatial_bbox/GeoParquet pruning)",
        "text": "### Describe the unexpected behaviour `KeyCondition`'s primary-key range analysis for `pointInPolygon` (`analyze_point_in_polygon` in `src/Storages/MergeTree/KeyCondition.cpp:3583-3663`) builds the query polygon directly from the literal argument, calls `boost::geometry::correct` and `boost::geometry::envelope`, but never calls `boost::geometry::is_valid`. This is inconsistent with the other two places in the codebase that parse the same kind of literal: - `FunctionPointInPolygon::parseConstPolygon`/`parseConstMultiPolygon` (`src/Functions/pointInPolygon.cpp:842-861,894-913`), which validate the assembled polygon with `bg::is_valid` and throw `BAD_ARGUMENTS` when `validate_polygons` is enabled (the default). - The `spatial_bbox` `MergeTree` skip index and GeoParquet row-group pruning (`src/Common/GeoBbox.h`, added in #104437), which validate every constant geometry argument and fail closed (decline to prune) rather than derive a bbox from an invalid one. `analyze_point_in_polygon` also only recognizes the plain 2-argument `pointInPolygon(point, ring)` form (see the `Case1 no holes in polygon` comment at line 3715); a polygon-with-holes literal (3+ arguments) isn't analyzed for primary-key pruning at all. That part is safe (it just declines to prune, so `KeyCondition` falls back to scanning), it's the missing validity check on the single-ring case that's the actual gap. ### Practical impact Under the default `validate_polygons = 1`, this gap is effectively unreachable through SQL: ClickHouse evaluates `WHERE`-clause constant expressions once on a zero-row block before any index analysis or pruning runs, so an invalid constant polygon literal passed to `pointInPolygon` always raises `BAD_ARGUMENTS` immediately — `KeyCondition`'s analysis, which runs later during part/granule selection, never gets a chance to act on the unvalidated bbox. However, if a user explicitly sets `validate_polygons = 0` — a documented, supported way to bypass `pointInPolygon`'s own geometry validation — the dry-run exception no longer fires, and `KeyCondition`'s PK-range analysis still unconditionally computes a bbox/envelope from the same, now genuinely unvalidated, polygon and may use it to prune primary-key ranges. `boost::geometry`'s query algorithms (`within`, `intersects`, etc.) have undefined results for invalid geometries, so pruning decisions derived from such a polygon are not guaranteed to be sound. ### How to reproduce Not reproducible with a concrete wrong-result example yet — this is a code-review finding, not an observed bug. The scenario would require: a `MergeTree` table with a `pointInPolygon`-friendly primary key, `SETTINGS validate_polygons = 0`, and a self-intersecting/otherwise invalid constant polygon literal in the `WHERE` clause, compared against a full scan of the same query. ### Expected behavior `analyze_point_in_polygon` should either validate the polygon the same way `FunctionPointInPolygon` and `Common/GeoBbox.h` do (and fail closed / decline to prune when invalid, honoring `validate_polygons` the same way the function itself does), or the inconsistency should be a deliberate, documented decision. ### Additional context Found while auditing ClickHouse's geospatial pruning code for duplicated/drifted logic during #104437 (which unified the equivalent bbox-extraction-and-validation logic between the `spatial_bbox` skip index and GeoParquet row-group pruning in `src/Common/GeoBbox.h`). Filing this as a lower-priority follow-up for future reference rather than addressing it in that PR, since it's a different subsystem (primary-key range analysis) and not currently reachable under default settings. Related: https://github.com/ClickHouse/ClickHouse/pull/104437",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114630",
        "createdAt": "2026-08-13T12:47:16Z",
        "updatedAt": "2026-08-13T15:28:59Z",
        "timestamp": "2026-08-13T15:28:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "bug",
          "comp-geo"
        ],
        "author": "bacek",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114637",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Trivial GROUP BY LIMIT optimization: support aggregate functions in the projection (2x on ClickBench Q17)",
        "text": "**Use case.** ClickBench Q17: ```sql SELECT UserID, SearchPhrase, COUNT(*) FROM hits GROUP BY UserID, SearchPhrase LIMIT 10; ``` The trivial `GROUP BY ... LIMIT` optimization (`optimize_trivial_group_by_limit_query`, default `true`) rewrites such queries to `max_rows_to_group_by = n + offset` with `group_by_overflow_mode = 'any'`, so the aggregation stops accepting new keys once enough distinct keys are produced. Today it is restricted to projections **without aggregate functions**, so it never fires for Q17 (or for any ClickBench query). The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 (commit `adda6e0ce599`, `Trivial group by limit`) applied the same rewrite **with** aggregate functions in the projection. Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64): | build | Q17 hot | |---|---| | original bench-opt (25.9.1.1) | 0.092 s | | master `405e218ff` | 0.201 s | i.e. **2.1x**. The official ClickBench numbers on `c7a.metal-48xl` show the same shape: 0.048 s vs 0.116 s. **Why it is restricted.** With `group_by_overflow_mode = 'any'`, rows for already-kept keys continue to be aggregated, so the aggregate values of the returned keys stay correct in the single-threaded case. But with parallel aggregation, each thread keeps its own first `n + offset` keys: a key kept by thread A and rejected by thread B loses B's rows, and the merged result returns that key with an undercounted aggregate. Returning an *unspecified subset* of keys is fine for `LIMIT` without `ORDER BY`; returning *wrong aggregate values* for them is not. The restriction to aggregate-free projections sidesteps this, at the cost of never firing on realistic queries. **Possible directions.** - Share the cutoff key set across threads (e.g. a concurrent filter of \"kept keys\"): once the global set reaches `n + offset`, all threads keep aggregating rows whose key is in the set and drop the rest — aggregate values for kept keys stay exact. - Alternatively, keep per-thread sets but make the final merge drop keys that were not kept by *every* participating thread (kept-by-all keys have exact values; needs `n + offset` sized so enough survive). - Or fall back to a single aggregation thread when the rewrite fires and the estimated cardinality is small — for `LIMIT 10`-style queries the single-threaded early-exit can still beat the full parallel scan. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114637",
        "createdAt": "2026-08-13T12:58:40Z",
        "updatedAt": "2026-08-13T12:58:40Z",
        "timestamp": "2026-08-13T12:58:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114638",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Aggregation: precompute hashes and prefetch for preallocated hash-table variants (up to 1.5x on multi-key GROUP BY)",
        "text": "**Use case.** Multi-key `GROUP BY` over large blocks, e.g. ClickBench Q16/Q18: ```sql SELECT UserID, SearchPhrase, COUNT(*) FROM hits GROUP BY UserID, SearchPhrase ORDER BY COUNT(*) DESC LIMIT 10; SELECT UserID, extract(minute FROM EventTime) AS m, SearchPhrase, COUNT(*) FROM hits GROUP BY UserID, m, SearchPhrase ORDER BY COUNT(*) DESC LIMIT 10; ``` The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 carried `Precompute hash and prefetch for prealloc variants` (commit `7663500fdebf` in the `bench-opt` branch): for aggregation methods that go through the preallocated/two-level hash table path, compute the key hashes for the whole block up front and software-prefetch the target buckets before the insert/lookup pass, hiding DRAM latency on the hash-table probes. This part of the PR was never upstreamed (the single-`String`-key improvements landed separately as the packed string keys, #93271). Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64), original bench-opt build (25.9.1.1) vs master `405e218ff`: | query | original | master | ratio | |---|---|---|---| | Q16 | 0.209 s | 0.326 s | 1.53x | | Q18 | 0.387 s | 0.487 s | 1.25x | | Q30 | 0.059 s | 0.078 s | 1.28x | | Q35 | 0.060 s | 0.078 s | 1.26x | | Q31 | 0.103 s | 0.122 s | 1.17x | | Q13/Q14/Q15 | — | — | ~1.1x | The official ClickBench numbers on `c7a.metal-48xl` agree in shape (e.g. Q18: 0.186 s vs 0.264 s). These queries all use multi-key aggregation (`keys128`/`serialized`-family methods), where master currently issues dependent random accesses per row. Precomputed hashes + prefetching is the remaining unmerged aggregation win from that PR. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772 Related: https://github.com/ClickHouse/ClickHouse/pull/93271",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114638",
        "createdAt": "2026-08-13T12:58:50Z",
        "updatedAt": "2026-08-13T12:58:50Z",
        "timestamp": "2026-08-13T12:58:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114639",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Push dynamic TopN thresholds into MergeTree reads for ORDER BY ... LIMIT (2-3x on ClickBench Q24/Q26)",
        "text": "**Use case.** `ORDER BY ... LIMIT n` over a sorted-by-something-else table with a narrow projection, e.g. ClickBench Q24/Q26: ```sql SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10; SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime, SearchPhrase LIMIT 10; ``` The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 implemented a **dynamic TopN threshold pushdown** (`Push TopN threshold to MergeTreeSource`, commit `2e2f594308f4`, plus `Better query condition cache: make TopN dynamic filters deterministic and reusable`, gated by `max_limit_to_push_down_topn_predicate = 100`): while the partial-sort transform maintains the current top-`n`, the running n-th-best value of the `ORDER BY` key is pushed down into the `MergeTree` read as a dynamic threshold, so granules whose key range cannot beat the current top-`n` are skipped instead of read, decompressed, and sorted. This mechanism was dropped during the upstreaming of that PR (the scaffolding was removed from the branch; the surviving `RewriteOrderByLimitPass` is a different, row-offset-based approach — it is off by default and measures performance-neutral on ClickBench when enabled). Master's lazy materialization (`query_plan_optimize_lazy_materialization`) covers the wide-`SELECT *` case (Q23), but does not prune reads for narrow projections: every granule passing the `WHERE` is still fully processed. Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64), original bench-opt build (25.9.1.1) vs master `405e218ff`: Q24 0.009 s vs 0.013 s, Q26 0.008 s vs 0.013 s (1.2–1.3x with times this small). The gap is much larger on the official `c7a.metal-48xl` numbers: Q24 0.014 s vs 0.046 s, Q26 0.014 s vs 0.044 s (**2–3x**). A related consideration from the original design: the dynamic filter interacts with the query condition cache, so the thresholds need to be deterministic/reusable (or excluded from the cache key) — the original branch had a follow-up commit specifically making the TopN dynamic filters deterministic for that reason. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114639",
        "createdAt": "2026-08-13T12:58:53Z",
        "updatedAt": "2026-08-13T16:33:01Z",
        "timestamp": "2026-08-13T16:33:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114640",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "uniqExact regression between 26.3 and 26.7 at high core counts (up to 3.3x on ClickBench Q10/Q11 on c7a.metal-48xl)",
        "text": "Between the official ClickBench mainline runs of 2026-03-27 (≈26.3) and 2026-08-11 (latest stable, 26.7.x) on `c7a.metal-48xl` (192 vCPU), the `COUNT(DISTINCT ...)` queries regressed while overall hot performance stayed flat (hot geomean 0.97): | query | hot 2026-03-27 | hot 2026-08-11 | ratio | |---|---|---|---| | Q10 `SELECT MobilePhoneModel, COUNT(DISTINCT UserID) AS u FROM hits WHERE MobilePhoneModel <> '' GROUP BY MobilePhoneModel ORDER BY u DESC LIMIT 10` | 0.114 s | 0.196 s | **1.7x** | | Q11 `SELECT MobilePhone, MobilePhoneModel, COUNT(DISTINCT UserID) AS u FROM hits WHERE MobilePhoneModel <> '' GROUP BY MobilePhone, MobilePhoneModel ORDER BY u DESC LIMIT 10` | 0.078 s | 0.257 s | **3.3x** | | Q37 (possibly related) | 0.023 s | 0.042 s | 1.8x | `COUNT(DISTINCT ...)` maps to `uniqExact` by default (`count_distinct_implementation`), and the evidence points at `uniqExact` state merging at high thread counts specifically: - The versions benchmark (`c7a.4xlarge`, 16 vCPU) shows **no** regression on these queries across 26.3 → 26.7 → master — but it substitutes `uniq(UserID)` for `COUNT(DISTINCT UserID)` (oldest-compatible query set), so it does not exercise `uniqExact` at all. - On a 96-core aarch64 box with the same 100M-row dataset, 26.7.1.1315 and current master both run Q10/Q11 in ~0.07 s hot — no regression at 96 threads. - So the degradation manifests only on the 192-vCPU machine, i.e. it scales badly with thread count rather than being a general codegen regression. Q10/Q11 have few groups (mobile phone models) with large per-group `uniqExact` states over ~100M rows — the merge of many per-thread large states is the hot path. A bisect of the 26.3 → 26.7 window on a high-core-count machine is needed; candidates are changes to parallel aggregation state merging or `uniqExact`-specific merge paths. Found while investigating the ClickBench dashboard delta for PR #81944 — this regression is unrelated to that PR but inflates the hot side of its dashboard comparison. Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114640",
        "createdAt": "2026-08-13T12:58:55Z",
        "updatedAt": "2026-08-13T12:58:55Z",
        "timestamp": "2026-08-13T12:58:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114641",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Text index: support arbitrary `LIKE` patterns with `tokenizer = 'array'`",
        "text": "### Company or project name ClickHouse Inc. ### Use case A text index with `tokenizer = 'array'` stores the whole column value as a single token, so its dictionary is the set of distinct values of the column, and a `LIKE` filter over that dictionary is exactly the query predicate. Today only patterns of the shape `%needle%`, with an alphanumeric needle, are evaluated using the index. Any other pattern reads the whole column, although the index already contains everything needed to answer it. ```sql CREATE TABLE t ( name String, INDEX idx name TYPE text(tokenizer = 'array') ) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t VALUES ('alpha-service-prod'), ('beta-service-prod'), ('gamma-svc-4999-dev'); ``` | Query | Uses the index today | Why not | |---|---|---| | `SELECT * FROM t WHERE name LIKE '%4999%'` | yes | | | `SELECT * FROM t WHERE name LIKE '%svc-4999%'` | no | punctuation in the needle | | `SELECT * FROM t WHERE name LIKE 'alpha%'` | no | anchored at the start | | `SELECT * FROM t WHERE name LIKE '%prod'` | no | anchored at the end | | `SELECT * FROM t WHERE name LIKE '%alpha%prod%'` | no | more than one needle | | `SELECT * FROM t WHERE name LIKE 'alpha_service%'` | no | `_` wildcard | ### Describe the solution you'd like For `tokenizer = 'array'`, evaluate arbitrary `LIKE` and `ILIKE` patterns using the text index: anchors, punctuation, `_` wildcards, several `%`-separated needles, escaped metacharacters. The existing cost guards should keep working — the minimum required literal length and the limit on how much of the index may be read, falling back to reading the column when the limit is exceeded. ### Describe alternatives you've considered An additional `ngrams(3)` index on the same column answers these patterns, but doubles index storage and write amplification for data the `array` dictionary already describes exactly. ### Additional context Actually, any predicate over the column with an `array` tokenizer can be supported, but it requires non-trivial development. The `LIKE` improvement comes almost for free, so let's start with this. Related: https://github.com/ClickHouse/ClickHouse/pull/98149 Related: https://github.com/ClickHouse/ClickHouse/issues/97723",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114641",
        "createdAt": "2026-08-13T13:03:00Z",
        "updatedAt": "2026-08-13T14:11:21Z",
        "timestamp": "2026-08-13T14:11:21Z",
        "metrics": {
          "reactions": 1,
          "comments": 0
        },
        "labels": [
          "feature",
          "comp-text-index"
        ],
        "author": "CurtizJ",
        "state": "open",
        "assignees": [
          "ahmadov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114649",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Segmentation fault in `UniqExactSet::merge()` with `countDistinct(String)` and GROUP BY",
        "text": "### Company or project name _No response_ ### Describe what's wrong A SELECT query crashes `clickhouse-server` with SIGSEGV while merging aggregate states for: ```sql countDistinct(note) = 1 ```` ClickHouse version: ```text 26.3.17.110 (official build) git hash: 59141459d999fabdcf3d1dd88cdd5ed6f10136ee ``` The relevant part of the query can be simplified to: ```sql SELECT count() FROM ( SELECT countIf(event_id, note = '') AS cnt, min(event_time) AS first_event, max(event_time) AS last_event FROM events WHERE event_time >= ... AND event_time <= ... GROUP BY key1, key2, key3, key4, key5, (topic, partition) HAVING cnt > 1 AND countDistinct(note) = 1 ) ``` The server crashes with: ```text Received signal Segmentation fault (11) Address: NULL pointer. Access: read. ``` Relevant stack frames: ```text TwoLevelHashTable<...>::TwoLevelHashTable(...) DB::UniqExactSet<...>::merge(...) DB::AggregateFunctionUniqExactData<String, true>::mergeBatch(...) DB::Aggregator::mergeStreamsImplCase(...) DB::Aggregator::mergeBlocks(...) DB::MergingAggregatedBucketTransform::transform(...) ``` So the crash happens while merging `uniqExact` states, apparently around single-level / two-level hash table merging or conversion. ### Workaround / additional evidence In this particular query: ```sql countIf(event_id, note = '') > 1 AND countDistinct(note) = 1 ``` can be replaced with the equivalent condition: ```sql countIf(event_id, note = '') > 1 AND countIf(note != '') = 0 ``` After replacing `countDistinct(note) = 1` with: ```sql countIf(note != '') = 0 ``` the same query completes successfully and no longer crashes the server. This strongly suggests that the crash is specifically related to the `uniqExact` / `countDistinct` aggregation path rather than the rest of the query. ### Does it reproduce on the most recent release? Yes ### How to reproduce See the description section. We didn't found easy way to reproduce this bug. ### Expected behavior The query should either complete successfully or fail with a ClickHouse exception. A SELECT query should not crash the server process. ### Error message and/or stacktrace [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.565819 [ 136891 ] <Fatal> BaseDaemon: Address: NULL pointer. Access: read. Sent by the kernel. [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.565835 [ 136891 ] <Fatal> BaseDaemon: Stack trace: 0x0000000019771551 0x0000000019773138 0x0000000019769925 0x000000001c7844ca 0x000000001c67f1c5 0x000000001c704c60 0x000000001c70f634 0x000000001f0dca4b 0x000000001ed504cb 0x000000001ed7157d 0x000000001ed62584 0x000000001ed66983 0x0000000017cc8816 0x0000000017ccfa1d 0x0000000017cc597d 0x0000000017ccd15a 0x00007f2901cc11f5 0x00007f2901d418ec [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566004 [ 136891 ] <Fatal> BaseDaemon: 2. TwoLevelHashTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>, HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>>, 8ul>::TwoLevelHashTable<HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>>>(HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>> const&) @ 0x0000000019771551 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566061 [ 136891 ] <Fatal> BaseDaemon: 3. DB::UniqExactSet<HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>>, TwoLevelHashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>>>::merge(DB::UniqExactSet<HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>>, TwoLevelHashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>>> const&, ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>*, std::atomic<bool>*) @ 0x0000000019773138 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566118 [ 136891 ] <Fatal> BaseDaemon: 4. DB::IAggregateFunctionHelper<DB::AggregateFunctionUniq<String, DB::AggregateFunctionUniqExactData<String, true>>>::mergeBatch(unsigned long, unsigned long, char**, unsigned long, char* const*, ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>&, std::atomic<bool>&, DB::Arena*) const @ 0x0000000019769925 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566196 [ 136891 ] <Fatal> BaseDaemon: 5. void DB::Aggregator::mergeStreamsImplCase<DB::ColumnsHashing::HashMethodSerialized<PairNoInit<std::basic_string_view<char, std::char_traits<char>>, char*>, char*, true, true>, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>>(DB::Arena*, DB::ColumnsHashing::HashMethodSerialized<PairNoInit<std::basic_string_view<char, std::char_traits<char>>, char*>, char*, true, true>&, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>&, bool, char*, unsigned long, unsigned long, std::vector<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, std::allocator<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*>> const&, std::atomic<bool>&, DB::Arena*) const @ 0x000000001c7844ca [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566288 [ 136891 ] <Fatal> BaseDaemon: 6. void DB::Aggregator::mergeStreamsImpl<DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>>(DB::Arena*, DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>&, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>&, char*, DB::ColumnsHashing::LastElementCacheStats&, bool, unsigned long, unsigned long, std::vector<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, std::allocator<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*>> const&, std::vector<DB::IColumn const*, std::allocator<DB::IColumn const*>> const&, std::atomic<bool>&, DB::Arena*) const @ 0x000000001c67f1c5 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566373 [ 136891 ] <Fatal> BaseDaemon: 7. void DB::Aggregator::mergeStreamsImpl<DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>>(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>> const&, unsigned long, DB::Arena*, DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>&, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>&, char*, DB::ColumnsHashing::LastElementCacheStats&, bool, std::atomic<bool>&, DB::Arena*) const @ 0x000000001c704c60 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566410 [ 136891 ] <Fatal> BaseDaemon: 8. DB::Aggregator::mergeBlocks(std::list<DB::Aggregator::AggregatedChunk, std::allocator<DB::Aggregator::AggregatedChunk>>&, bool, std::atomic<bool>&) @ 0x000000001c70f634 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566480 [ 136891 ] <Fatal> BaseDaemon: 9. DB::MergingAggregatedBucketTransform::transform(DB::Chunk&) @ 0x000000001f0dca4b [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566528 [ 136891 ] <Fatal> BaseDaemon: 10. DB::ISimpleTransform::work() @ 0x000000001ed504cb [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566562 [ 136891 ] <Fatal> BaseDaemon: 11. DB::ExecutionThreadContext::executeTask() @ 0x000000001ed7157d [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566596 [ 136891 ] <Fatal> BaseDaemon: 12. DB::PipelineExecutor::executeStepImpl(unsigned long, DB::IAcquiredSlot*, std::atomic<bool>*) @ 0x000000001ed62584 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566669 [ 136891 ] <Fatal> BaseDaemon: 13. void std::__function::__policy_func<void ()>::__call_func[abi:fe210105]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0>(std::__function::__policy_storage const*) @ 0x000000001ed66983 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566712 [ 136891 ] <Fatal> BaseDaemon: 14. ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000017cc8816 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566774 [ 136891 ] <Fatal> BaseDaemon: 15. void std::__function::__policy_func<void ()>::__call_func[abi:fe210105]<ThreadFromGlobalPoolImpl<false, true>::ThreadFromGlobalPoolImpl<void (ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::*)(), ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool*>(void (ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool*&&)::'lambda'()>(std::__function::__policy_storage const*) @ 0x0000000017ccfa1d [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566801 [ 136891 ] <Fatal> BaseDaemon: 16. ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x0000000017cc597d [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566842 [ 136891 ] <Fatal> BaseDaemon: 17. void* std::__thread_proxy[abi:fe210105]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x0000000017ccd15a [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566914 [ 136891 ] <Fatal> BaseDaemon: 18. ? @ 0x00000000000891f5 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566934 [ 136891 ] <Fatal> BaseDaemon: 19. ? @ 0x00000000001098ec [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:56.281387 [ 136891 ] <Fatal> BaseDaemon: Integrity check of the executable successfully passed (checksum: 4079FE6D02AC0B0EFCF8B7F04A4EBE95) [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:56.281479 [ 136891 ] <Fatal> BaseDaemon: ClickHouse version 26.3.17.110 is old and should be upgraded to the latest version. [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:56.281657 [ 136891 ] <Fatal> BaseDaemon: Changed settings: max_query_size = 800000, connect_timeout_with_failover_ms = 1000, connect_timeout_with_failover_secure_ms = 2000, use_uncompressed_cache = false, distributed_foreground_insert = true, load_balancing = 'random', enable_positional_arguments = false, log_queries_cut_to_length = 250000, distributed_product_mode = 'global', insert_quorum = 0, select_sequential_consistency = 0, max_http_get_redirects = 1, any_join_distinct_right_table_keys = true, distributed_ddl_task_timeout = 30, max_bytes_before_external_group_by = 15000000000, max_bytes_before_external_sort = 15000000000, max_execution_time = 7200., max_ast_depth = 2000, max_ast_elements = 100000, max_expanded_ast_elements = 1000000, max_memory_usage = 15000000000, formatdatetime_parsedatetime_m_is_month_name = false, allow_drop_detached = true, deduplicate_blocks_in_dependent_materialized_views = false, async_insert = false, allow_experimental_analyzer = false, allow_experimental_window_functions = true, background_schedule_pool_size = 100, output_format_json_quote_64bit_integers = true, output_format_pretty_max_value_width = 250000 Error on processing query: Code: 32. DB::Exception: Attempt to read after eof: while receiving packet from 127.0.0.1:9017, local address: 127.0.0.1:14871. (ATTEMPT_TO_READ_AFTER_EOF) (version 26.3.17.110 (official build)) ### Related issues and pull requests This looks related to the previous `uniqExact` parallel merge issues: * #108912 * #108928 * #109389 However, the crashing build already contains the fixes from those changes, so this may be another issue in the same code path. ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114649",
        "createdAt": "2026-08-13T14:34:08Z",
        "updatedAt": "2026-08-13T15:25:32Z",
        "timestamp": "2026-08-13T15:25:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "bug",
          "crash"
        ],
        "author": "vasyaabr",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:114657",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "The remote leg of a Distributed query re-serializes a lenient toDecimal64 constant into a strict CAST that throws: local and distributed results diverge",
        "text": "**Describe what's wrong** A query whose `WHERE` clause contains a constant expression producing an out-of-precision `Decimal` value — e.g. `toDecimal64(1000000000000000000, 0)`, a 19-digit value in `Decimal(18, 0)` — executes fine against a local table, but the same query through a `Distributed` table throws `Code: 69. DB::Exception: Too many digits (19 > 18) in decimal value` from the remote leg. The distributed layer re-serializes the folded constant as `_CAST('1000000000000000000', 'Decimal(18, 0)')`, and that string-literal form validates digits strictly while the original integer-argument `toDecimal64` call does not (the long-standing leniency described in #14124). So the rewritten query the initiator ships to the shards is not executable, although the user's original query is — the same statement returns a result locally and an exception through `Distributed`. **Does it reproduce on the most recent release?** Reproduced on master `26.8.1.1307` (official build), 20/20 through the `Distributed` table (exception) and 20/20 on the local table (correct result). **How to reproduce** ```sql CREATE TABLE t (a Decimal(18, 0)) ENGINE = MergeTree ORDER BY a; INSERT INTO t SELECT toDecimal64(number * 100000000000, 0) FROM numbers(100); CREATE TABLE dist_t AS t ENGINE = Distributed(test_shard_localhost, currentDatabase(), 't'); -- local: returns 0 SELECT count() FROM t WHERE intDiv(a, toDecimal64(1000000000000000000, 0)) NOT IN (0, 1); -- distributed: Code: 69. DB::Exception: Too many digits (19 > 18) in decimal value: -- In scope SELECT count() FROM ... WHERE intDiv(a, _CAST('1000000000000000000', 'Decimal(18, 0)')) NOT IN (0, 1) SELECT count() FROM dist_t WHERE intDiv(a, toDecimal64(1000000000000000000, 0)) NOT IN (0, 1) SETTINGS prefer_localhost_replica = 0; ``` `prefer_localhost_replica = 0` is needed only because the repro cluster is a localhost loopback; on a cluster with genuinely remote shards the remote leg always takes the serialization path and the exception fires at default settings. Simpler predicates hit the same thing: `WHERE a != toDecimal64(1000000000000000000, 0)` and `WHERE a < toDecimal64(1000000000000000000, 0)` both error 69 through `Distributed` and succeed locally. The constant itself evaluates leniently on the initiator: ```sql SELECT toDecimal64(1000000000000000000, 0); -- returns 1000000000000000000 SELECT CAST('1000000000000000000', 'Decimal(18, 0)'); -- Code: 69, Too many digits (19 > 18) ``` **Expected behavior** The same query should either succeed on both paths or fail on both. Whichever way the `toDecimal64` integer-form leniency is resolved, the SQL the distributed layer generates for the remote legs should be executable whenever the original query is — e.g. by serializing the folded constant in a form that round-trips (keeping the original function call, or a value form that does not re-validate into a stricter type). Related: https://github.com/ClickHouse/ClickHouse/issues/14124 Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/114657",
        "createdAt": "2026-08-13T15:31:16Z",
        "updatedAt": "2026-08-13T15:31:16Z",
        "timestamp": "2026-08-13T15:31:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "potential bug"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:58242",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Non-constant right hand side of IN",
        "text": "``` milovidov-desktop :) SELECT number FROM numbers(10) WHERE number % 2 IN (number % 3, number % 5) Received exception: Code: 47. DB::Exception: Missing columns: 'number' while processing query: 'number % 3', required columns: 'number' 'number': While processing (number % 2) IN (number % 3, number % 5). (UNKNOWN_IDENTIFIER) ``` It can be done in the following way: - support non-constant case for IN [array] as full scan (`has`); - interpreting non-constant IN (set) in the same way as IN [array].",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/58242",
        "createdAt": "2023-12-27T11:35:09Z",
        "updatedAt": "2026-08-13T16:15:22Z",
        "timestamp": "2026-08-13T16:15:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "feature",
          "comp-sql-syntax",
          "warmup task"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "novikd"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:68428",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "[Umbrella] JSON type improvements",
        "text": "This issue contains tasks for further improvements of new [JSON](https://github.com/ClickHouse/ClickHouse/pull/66444) data type. Continuation of https://github.com/ClickHouse/ClickHouse/issues/54864. Important tasks: - [x] Reduce memory usage during insertion into JSON column and during merges by better buffer size managing - https://github.com/ClickHouse/ClickHouse/pull/69272 - [x] Support ALTER to new `JSON` type from `String`/`Map`/`Tuple`/ - https://github.com/ClickHouse/ClickHouse/pull/70442 - [x] Support ALTER to new `JSON` type from `Object` - https://github.com/ClickHouse/ClickHouse/pull/71784 - [x] Support CAST from `Map`/`Tuple`/ `Object` types to new `JSON` - https://github.com/ClickHouse/ClickHouse/pull/71320 - [x] Support CAST between `JSON` types with different parameters (skip paths, limits, type hints) - https://github.com/ClickHouse/ClickHouse/pull/72303 - [x] Try to implement default implementation for functions for `Dynamic` type (so any function will be executed on underlying types dynamically and throw an exception on first incompatible type) - https://github.com/ClickHouse/ClickHouse/pull/69691 - [x] Add aggregate functions `distinctDynamicTypes`/`distinctJSONPaths`/`distinctJSONPathsAndTypes` for better introspection of the content in the JSON column - https://github.com/ClickHouse/ClickHouse/pull/68463 - [x] Allow to insert/select JSON from binary strings in RowBinary format (https://github.com/ClickHouse/ClickHouse/issues/69443) - https://github.com/ClickHouse/ClickHouse/pull/70288 - [x] Allow to insert/select JSON from single String column in Native format (https://github.com/ClickHouse/ClickHouse/issues/70281) - https://github.com/ClickHouse/ClickHouse/pull/70312 - [x] Support JSON subcolumns in ORDER BY and data-skipping indices expression during table creation - https://github.com/ClickHouse/ClickHouse/pull/72644 - [x] Support `Dynamic` type in `ifNull`/`coalesce` functions - https://github.com/ClickHouse/ClickHouse/pull/72772 - [x] Support `Dynamic` type in functions `toFloat64` and similar - https://github.com/ClickHouse/ClickHouse/pull/72989 - [x] Support equal comparison between JSON values - https://github.com/ClickHouse/ClickHouse/pull/72991 - [x] Try to improve subcolumns formatting - https://github.com/ClickHouse/ClickHouse/pull/73085 - [x] Try to support `Nullable(JSON)` - https://github.com/ClickHouse/ClickHouse/pull/73556 - [x] Support subcolumns in MaterializedView queries - https://github.com/ClickHouse/ClickHouse/pull/74030 - [x] Support referring JSON subcolumns in default and materialized exressions - https://github.com/ClickHouse/ClickHouse/pull/74403. - [x] Improve performance of reading the whole JSON column from wide parts - https://github.com/ClickHouse/ClickHouse/pull/74827 - [x] Improve performance of subcolumns reading from compact parts - https://github.com/ClickHouse/ClickHouse/pull/77940 - [x] Add new serialization for Dynamic and JSON without SharedVariant/SharedData for better integrations - https://github.com/ClickHouse/ClickHouse/pull/80499 - [x] Support ALTER UPDATE for `JSON`/`Dynamic` types - https://github.com/ClickHouse/ClickHouse/pull/82419 - [x] Use `Array(Dynamic)` instead of unnamed Tuple for arrays of values with different types - https://github.com/ClickHouse/ClickHouse/issues/74937 - [x] Improve deserialization of JSON subcolumns from shared data - https://github.com/ClickHouse/ClickHouse/pull/83777 - [x] Remove old `Object` data type - https://github.com/ClickHouse/ClickHouse/pull/85718 - [x] Support JSON in `tupleElement` function - https://github.com/ClickHouse/ClickHouse/pull/91327 - [x] Add setting `type_json_skip_invalid_typed_paths ` to use default value during bad parsing of a path with a hint type - https://github.com/ClickHouse/ClickHouse/pull/89886 - [x] Optimize `distinctJSONPaths` aggregate function so it reads only metadata files with the list of paths - https://github.com/ClickHouse/ClickHouse/pull/92196 - [x] Optimize insertion into JSON column - https://github.com/ClickHouse/ClickHouse/pull/93614 and https://github.com/ClickHouse/ClickHouse/pull/94247 - [x] Add syntax `json.$a.b` for union of `json.a.b` and `json.^a.b` or function for such use case - https://github.com/ClickHouse/ClickHouse/pull/98788 - [x] Support `JSON` type in `JSONExtract*` functions - https://github.com/ClickHouse/ClickHouse/pull/96711 - [x] Support creating indexes on `JSONAllPaths` function similar to `mapKeys` for `Map` type - https://github.com/ClickHouse/ClickHouse/pull/98886 Tasks with no priority: - [ ] Improve alters of JSON/Dynamic columns in Wide part (avoid whole part rewrite) - [ ] Replace JSON column in the subquery to requested subcolumns https://github.com/ClickHouse/ClickHouse/issues/75538 - [ ] Add setting `max_subcolumns_to_read`. - [ ] Add support for EPHEMERAL subcolumns in JSON paths types declaration. - [ ] Add support for TTL for JSON subcolumns. - [ ] Add support for CODEC for JSON subcolumns. - [ ] Support ALTER UPDATE for individual subcolumns - [ ] Add function and aggregate function to merge JSON objects - [ ] Match json and jsonb type in PostgreSQL to new JSON type under a setting - [ ] Support JSON type in MongoDB engine (https://github.com/ClickHouse/ClickHouse/issues/71970) - [ ] Add function `JSONAllValues` that returns the list of values (casted to String maybe) stored in JSON column in sorted by path order - [ ] Add new JSON type to schema inference - [ ] Consider adding new syntax to read subcolumns by regex ([comment](https://github.com/ClickHouse/ClickHouse/issues/68428#issuecomment-2478594230)) - [ ] Consider implementing function for querying paths dynamically ([comment](https://github.com/ClickHouse/ClickHouse/issues/68428#issuecomment-2499106228)) - [ ] Add information about subcolumn sizes in system table. - [ ] Support writing `json.array[N].key` for `Array(JSON)` subcolumns ([comment](https://github.com/ClickHouse/ClickHouse/issues/68428#issuecomment-2587018860)) - [ ] Support declaring the type for all dynamic subcolumns ([comment](https://github.com/ClickHouse/ClickHouse/issues/68428#issuecomment-2609822747)) - [ ] Support `Dynamic` type as first argument of function `has` - [ ] Add function like `jsonuntuple(json, array_of_paths)` to select paths as separate columns - [ ] Support `arrayElement` function for JSON to read subcolumns like `json['key']` and use optimization functions to subcolumns here - [ ] Add new syntax in JSON type definition to define type for path regexp ([comment](https://github.com/ClickHouse/ClickHouse/issues/68428#issuecomment-2727234716)) - [ ] Support JSON in BSONEachRow format - [ ] Add functions `JSONRemoveKeys` and `JSONUpdateKeys` - [ ] Support subcolumns in ReplacingMergeTree arguments for `version` and `is_deleted` columns - [ ] Consider adding function that converts JSON to `Map(String, Dynamic)` Feel free to add any ideas for improvements/feature requests in the comments in this issue",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/68428",
        "timestamp": "2026-08-12T22:38:16Z",
        "metrics": {
          "reactions": 67,
          "comments": 107
        },
        "labels": [
          "feature"
        ],
        "author": "Avogar",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:72380",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "`RENAME DATABASE` query doesn't work for materialized views",
        "text": "**Company or project name** Prefer not to specify **Describe what's wrong** `RENAME DATABASE` query doesn't seem to be working correctly with materialized view statements. After renaming a database with `RENAME DATABASE old_name TO new_name` query, the materialized views would still reference tables with the `old_name` as could be seen with `SHOW CREATE new_name.materialized_view` statement in the repro. That leads to a bunch of issues with database permissions and `SELECT` queries. And while `INSERT` queries into the main table seem to be working, they are not triggering any relative materialized views. Repro: https://fiddle.clickhouse.com/797f730d-d037-4b64-8d17-ce47911fccdd **Does it reproduce on the most recent release?** Yes **Enable crash reporting** Doesn't crash **How to reproduce** I first got this on `24.8.1.10452` in ClickHouse Cloud, but I'm also able to reproduce it on current latest. ``` CREATE DATABASE test; CREATE TABLE test.sample ( id UUID DEFAULT generateUUIDv4(), data TEXT DEFAULT '' ) ENGINE = MergeTree PRIMARY KEY (id) ORDER BY (id); -- Create materialized view without explicitly specifying target table CREATE MATERIALIZED VIEW test.inline_mat_view ( uuid UUID, data TEXT ) ENGINE MergeTree ORDER BY (uuid) AS SELECT id as uuid, data FROM test.sample; -- Create table for materialized view explicitly CREATE TABLE test.explicit_table ( uuid UUID, data TEXT ) ENGINE MergeTree ORDER BY (uuid); -- Create materialized view pointing to explicitly created table CREATE MATERIALIZED VIEW test.explicit_mat_view TO test.explicit_table AS SELECT id as uuid, data FROM test.sample; -- Output original CREATE statements for both materialized views SHOW CREATE test.inline_mat_view; SHOW CREATE test.explicit_mat_view; RENAME DATABASE test TO dev; -- Inserting data into main table doesn't trigger any error, but materialized views are not updated INSERT INTO dev.sample (data) VALUES ('test1'), ('test2'), ('test3'), ('test4'), ('test5'); -- CREATE statements for materialized views after renaming the database still points to the old database name SHOW CREATE dev.inline_mat_view; SHOW CREATE dev.explicit_mat_view; -- Exception due to missing 'test' database SELECT * FROM dev.inline_mat_view; SELECT * FROM dev.explicit_mat_view; ``` **Expected behavior** When renaming the database I would expect materialized view to be updated accordingly and to point to the correct table in the renamed database.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/72380",
        "createdAt": "2024-11-25T10:53:38Z",
        "updatedAt": "2026-08-13T14:39:02Z",
        "timestamp": "2026-08-13T14:39:02Z",
        "metrics": {
          "reactions": 1,
          "comments": 5
        },
        "labels": [
          "potential bug"
        ],
        "author": "biased-badger",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:80759",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "TCP server threads metrics (i.e. TCPThreads) is broken for <protocols> (over <tcp_port>/...)",
        "text": "In case of <protocols> is used over <tcp_port> (e.t.c.), the name of the server is different and the code in AsynchronouseMetrics.cpp fails to match them. Suggestion - add proper ServerType::Type enum into the ProtocolServerAdapter, and use it over matching by \"port name\".",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/80759",
        "createdAt": "2025-05-23T21:31:37Z",
        "updatedAt": "2026-08-13T06:15:14Z",
        "timestamp": "2026-08-13T06:15:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "easy task",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "azat",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:80787",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "MaterializedPostgreSQL DB Engine: Support for TLS",
        "text": "### Company or project name _No response_ ### Describe the unexpected behaviour As of now, it is impossible to connect to a PostgreSQL with enforced SSL connections, with or without verifying the certificates ### How to reproduce Setup a PostgreSQL with required SSL (e.g. with zalando postres operator) Try to create a MaterializedPostgreSQL Database Connection will fail - because clickhouse won't use SSL. ### Expected behavior I propose adding pgsslmode into the Connection creation and adding settings to set the certificates when pgsslmode=verify or pgsslmode=verify-full. In the code there's already a connection string builder, which should be capable of adding these settings. ### Error message and/or stacktrace _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/80787",
        "createdAt": "2025-05-25T09:07:21Z",
        "updatedAt": "2026-08-13T00:06:52Z",
        "timestamp": "2026-08-13T00:06:52Z",
        "metrics": {
          "reactions": 1,
          "comments": 0
        },
        "labels": [
          "help wanted",
          "unfinished code"
        ],
        "author": "martin31821",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:85293",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Dictionary credentials rotation",
        "text": "### Company or project name Cabify ### Use case We are using Clickhouse Cloud. We are using dictionaries to populate data from other ClickHouse tables. The source looks sth like this: ``` SOURCE(CLICKHOUSE( USER '${CH_DICT_USER}' PASSWORD '${CH_DICT_PASSWORD}' DATABASE '${CH_DB}' TABLE '${CH_TABLE}' ) ``` The issue we are encountering is that we cannot rotate the credentials; doing so would make the dictionary stop working, as the credentials used during its creation would become invalid. We use short-lived dynamic credentials (managed by [Vault](https://github.com/hashicorp/vault)) as a general approach, so not being able to rotate provided dictionary credentials seems not ideal for us. Requiring users with the `default_role` to provide static credentials for creating a dictionary appears to be quite restrictive and may contradict certain security best practices implemented by different organizations. We are aware that no password is required if the `DEFAULT` user is used for creating the dictionary, sth like: ``` SOURCE(CLICKHOUSE( TABLE '${CH_TABLE}' )) ``` That approach could work, as rotating `DEFAULT` user credentials won't affect the already created dictionaries. However, again this would force to use `DEFAULT` user for creating a dictionary, which goes against PoLP (Principle of Least Privilege). We have followed the following approach for now: > Use a user with `default_role` privileges, static credentials and HOST LOCAL. Although it could met some of our security requirements (as that user can NOT access from outside the instance), the solution is not ideal. ### Goal Allow the provision of user credentials for creating the dictionary, ensuring that the dictionary remains functional even if the credentials are rotated. Additionally, I’m unclear on why the user needs to have the `default_role` just to create a dictionary, as this seems to grant excessive permissions. ### Describe the solution you'd like I expect that any user with the `default_role` will have similar behaviour as the `DEFAULT` user when creating a dictionary. Once dictionaries are created by the `DEFAULT` user, the user’s credentials can be changed, and the dictionaries continue working. So ideally, we provide a user with HOST LOCAL with credentials that can be rotated without affected related dictionaries. For example, instead of hardcoding user credentials at the dictionary level, verify the authentication and authorization of the user during the dictionary creation process. If the user is valid, link their username to the dictionary. Corresponding validations should also be performed during dictionary reloads, similar to the current approach with the `DEFAULT` user. ### Describe alternatives you've considered _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/85293",
        "createdAt": "2025-08-08T12:23:39Z",
        "updatedAt": "2026-08-13T06:48:59Z",
        "timestamp": "2026-08-13T06:48:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "feature"
        ],
        "author": "x-martinez",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:88957",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Unexpected syntax error when using `paste join` with `IcebergS3Cluster`",
        "text": "### Company or project name _No response_ ### Describe the unexpected behaviour ```sql SELECT * FROM icebergS3Cluster('replicated_cluster', 'http://minio:9000/warehouse/data2', 'admin', 'password') PASTE JOIN ( SELECT * FROM iceberg('http://minio:9000/warehouse/data3', 'admin', 'password') ) AS t1 ORDER BY tuple(*) ASC FORMAT Values ``` ``` Elapsed: 0.016 sec. Received exception from server (version 25.8.10): Code: 62. DB::Exception: Received from localhost:9000. DB::Exception: Received from clickhouse1:9000. DB::Exception: You must not specify ANY or ALL for PASTE JOIN.. (SYNTAX_ERROR) ``` ### Which ClickHouse versions are affected? 25.8 ### How to reproduce Run query ```sql SELECT * FROM icebergS3Cluster('replicated_cluster', 'http://minio:9000/warehouse/data2', 'admin', 'password') PASTE JOIN ( SELECT * FROM iceberg('http://minio:9000/warehouse/data3', 'admin', 'password') ) AS t1 ORDER BY tuple(*) ASC FORMAT Values ``` If using this syntax, it will work ```sql SELECT * FROM ( SELECT * FROM icebergS3Cluster('replicated_cluster', 'http://minio:9000/warehouse/data2', 'admin', 'password') ) AS t1 PASTE JOIN ( SELECT * FROM iceberg('http://minio:9000/warehouse/data3', 'admin', 'password') ) AS t2 ORDER BY tuple(*) ASC FORMAT Values ``` But I saw in tests, that both types of query are supported. ### Expected behavior Correct result of Paste Join. ### Error message and/or stacktrace _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/88957",
        "createdAt": "2025-10-24T13:17:11Z",
        "updatedAt": "2026-08-13T06:15:25Z",
        "timestamp": "2026-08-13T06:15:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "unexpected behaviour",
          "comp-datalake",
          "clickgap-analyzed",
          "culprit-pr-not-found"
        ],
        "author": "alsugiliazova",
        "state": "open",
        "assignees": [
          "thevar1able"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:92350",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "JWT Auth for native and HTTP protocols",
        "text": "### Company or project name Arenadata ### Use case Support of JWT auth in ClickHouse (non Clould version) in native and HTTP protocols. ### Describe the solution you'd like We found an attempt to implement it in [PR](https://github.com/ClickHouse/ClickHouse/pull/68634) with `Closed` status. After analysis it seems more like a draft PR: * payload validation - `<claims>{\"resource_access\":{\"account\": {\"roles\": [\"view-profile\"]}}} </claims>` it seems does not work (in debugger I see that parser does not load claims at all). * `<claims>` is not covered by integration tests. * `CREATE USER jwt_user6 IDENTIFIED WITH jwt CLAIMS '{\"resource_access\":{\"account\": {\"roles\": [\"view-profile\"]}}}'` does not work and not covered by integration tests. Overall: for now we don't need claims. Probably `<claims>` in `users.xml` and in `CREATE USER` should be removed. * probably `<validator>` with name of validator should be added into `<jwt>` section of user in users.xml (the same as `server` in LDAP or external HTTP auth). Or even allow to define just one JTW validator in `config.xml` instead of list of validators (the same as one `<kerberos>` section for Kerberos auth). * `#if USE_JWT_CPP && USE_SSL` is not used. * `curl -H \"X-ClickHouse-JWT-Token: Bearer <token>\" \"http://localhost:8123/?query=SELECT currentUser()\" ` it seems like we should remove Bearer because we use `X-ClickHouse-JWT-Toke`. Or remove (don't use) `X-ClickHouse-JWT-Toke` itself, because I can be used in standard way: `curl -H \"Authorization: Bearer <token>\" \"http://localhost:8123/?query=SELECT currentUser()\"`. Also there is another [PR](https://github.com/ClickHouse/ClickHouse/pull/88857) which is subset of the PR above and with lack of activity in last 2 months (abandoned?) ### Describe alternatives you've considered _No response_ ### Additional context _No response_",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/92350",
        "timestamp": "2026-08-12T20:41:51Z",
        "metrics": {
          "reactions": 3,
          "comments": 9
        },
        "labels": [
          "feature"
        ],
        "author": "rvasin",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:95963",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "Idempotency key for DDL operations",
        "text": "### Company or project name _No response_ ### Use case Retries after network connection errors are really complex in distributed systems. There are many operations that are quite risky to be retried such as `EXCHANGE TABLE` and we do not have easy way to check that previous operation succeeded or not for example, when disconnect was result of completely disappearing node and we have query log local to this node. ### Describe the solution you'd like - Add optional query setting `idempotency_key` that can be passed from a client with a query similar to `insert_deduplication_token`. - When it is passed operations that change state(or go through keeper) will commit the `idempotency_key` to a special node in keeper. - idempotency_key work the same way as insert_deduplication_token - it will be cleaned up with the similar window logic. - Client libraries can set this key automatically and retry on network error preserving `idempotency_key`. It will make most retries completely safe. - Server will check if key exists before running query and also include it during operation commit to keeper. This way you can also cancel one of 2 parallel EXCHANGE TABELS so only one will succeed. ### Describe alternatives you've considered Other option is to pass some key similar to query_if from a client and on error check that this query was not running before. Saving this information somehow to keeper with operations that change metadata in keeper will allow to later check if we had one already finished. But we do not have guarantee that there is none inflight with unreachable node, that is still connected to keeper. So it will add more complexity if we can not store a tombstone saying this operation should not succeed after you checked. ### Additional context Language clients suffer a lot of disconnects as sometimes there is a race for connections to be not closed for up to 5 minutes: https://github.com/ClickHouse/ClickHouse/issues/92046 They can't easily detect if this case was there i.e. no bytes were sent from client and just ordinary disconnect when network is bad.",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/95963",
        "createdAt": "2026-02-04T16:27:50Z",
        "updatedAt": "2026-08-13T04:48:48Z",
        "timestamp": "2026-08-13T04:48:48Z",
        "metrics": {
          "reactions": 2,
          "comments": 1
        },
        "labels": [
          "feature"
        ],
        "author": "qoega",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:issue:96651",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "issue",
        "title": "NATS JetStream pull consumer does not recover after failover (or change server)",
        "text": "### Company or project name _No response_ ### Describe the unexpected behaviour When using ClickHouse ENGINE = NATS with JetStream (pull mode, durable consumer), the consumer stops fetching messages after a NATS cluster failover. The TCP connection is successfully re-established (PING/PONG continues), but the JetStream pull loop does not resume. Messages remain unprocessed until the table is manually restarted using: ``` DETACH TABLE <table>; ATTACH TABLE <table>; ``` ### Which ClickHouse versions are affected? ClickHouse server version 26.1.2 ### How to reproduce 1. Run ClickHouse with NATS JetStream consumer. 2. Ensure messages are being consumed normally. 3. Stop the NATS node currently serving as JetStream connect by Clickhouse. 4. Wait for Clickhouse reconnect to next node 5. TCP reconnect happens (PING/PONG visible) 6. Clickhouse send message connect like: ``` CONNECT {\"verbose\":false,\"pedantic\":false,\"user\":\"clickhouse\",\"pass\":\"mysuperPass\",\"tls_required\":false,\"name\":\"\",\"lang\":\"C\",\"version\":\"3.9.2\",\"protocol\":1,\"echo\":true,\"headers\":true,\"no_responders\":true} 7. Then Clickhouse send: ``` SUB _INBOX.PXZCFR8UL2QF2MTNAXU2OL.* 1 ``` 8. After that Clickhouse send only ping, without fetch message 9. After detatch/attach table: ``` DETACH TABLE nats_consumer; ATTACH TABLE nats_consumer; ``` clickhouse send: ``` PUB $JS.API.CONSUMER.INFO.test.clickhouse _INBOX.PXZCFR8UL2QF2MTNAXU2OL.9 0 SUB _INBOX.PXZCFR8UL2QF2MTNAXU47V.* 11 PUB $JS.API.CONSUMER.MSG.NEXT.test.clickhouse _INBOX.PXZCFR8UL2QF2MTNAXU47V.1 13 {\"batch\":128} ``` and consumption resumes immediately. ### Expected behavior _No response_ ### Error message and/or stacktrace I don't see any errors in the logs ### Additional context DDL: ``` --connect to nats CREATE TABLE IF NOT EXISTS nats_connect ( uuid String, send UInt8, metainfo String ) ENGINE = NATS SETTINGS nats_server_list = 'node1:4222,node2:4222,node3:4222,node4:4222', nats_subjects = 'test.>', nats_consumer_name = 'clickhouse', nats_format = 'JSONEachRow', nats_num_consumers = 1, nats_max_rows_per_message = 1, nats_username = 'clickhouse', nats_password = 'mysuperPass', nats_stream = 'test', nats_handle_error_mode = 'stream'; --main table to insert data from nats CREATE TABLE IF NOT EXISTS nats_data' ( uuid String, send UInt8, metainfo String, ts DateTime DEFAULT now() ) ENGINE = MergeTree() ORDER BY ts; --consumer CREATE MATERIALIZED VIEW nats_consumer TO nats_data AS SELECT uuid, send, metainfo FROM nats_connect; ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #112828 - Merged into: `26.8.1.1343` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/issues/96651",
        "createdAt": "2026-02-11T10:11:19Z",
        "updatedAt": "2026-08-13T15:51:47Z",
        "timestamp": "2026-08-13T15:51:47Z",
        "metrics": {
          "reactions": 7,
          "comments": 2
        },
        "labels": [
          "unexpected behaviour"
        ],
        "author": "echohes",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100160",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Feature: Paimon minmax index",
        "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Supports Paimon min-max index pushdown by setting `use_paimon_minmax_index_pruning=1` ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100160",
        "createdAt": "2026-03-20T08:23:32Z",
        "updatedAt": "2026-08-13T02:19:03Z",
        "timestamp": "2026-08-13T02:19:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "submodule changed",
          "can be tested",
          "pr-experimental"
        ],
        "author": "JiaQiTang98",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100185",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add assertions in IColumn::mutate to verify deep unique ownership of sub-columns",
        "text": "Investigation for https://github.com/ClickHouse/ClickHouse/issues/99920 This is a draft PR with diagnostic assertions to help find a COW violation bug where columns stored in `HashJoin` have their `allocatedBytes()` change because internal sub-columns are modified by someone who treats immutable columns as mutable. ## What was found The bug manifests as `data->allocated_size != debug_allocated_size` exception in debug builds. Investigation narrowed it down to: 1. **Removing the single-chunk optimization** in `Squashing.cpp:70-72` fixes the bug — this optimization passes chunks through without `IColumn::mutate`, so columns remain shared with Memory engine blocks 2. The **modification happens between** `addBlockToJoin` calls — during pipeline processing of the next chunk 3. The mutation is a **reallocation** (reserve/grow), not a data write — `protect()` (mprotect) doesn't catch it 4. The assertions in this PR (`IColumn::mutate` deep ownership check) **do NOT fire** — meaning the mutation happens outside the `IColumn::mutate` path entirely. Someone treats an immutable column as mutable without going through `mutate`. ## Reproducer The test `04049_fix_hashjoin_shared_column_allocated_size` reproduces the exception in debug builds. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100185",
        "createdAt": "2026-03-20T11:00:24Z",
        "updatedAt": "2026-08-13T06:34:50Z",
        "timestamp": "2026-08-13T06:34:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 21
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100371",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `borrow_from_cache` object storage and `memory` metadata types",
        "text": "Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache using ephemeral `FileSegment`s. Each stored object is backed by a cache segment held alive via `FileSegmentsHolder`; when released, the cache reclaims the space. Add a new metadata type `memory` that keeps all file-to-blob mappings and directory structure entirely in memory with no persistence. On server restart all data is lost, which is acceptable for the intended use case of temporary tables. Configuration example: ``` disk(type=object_storage, object_storage_type='borrow_from_cache', metadata_type='memory', cache_name='some_cache') ``` ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A new object storage type and metadata type suitable for temporary tables. Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache and holds it from eviction. Add a new metadata storage type `memory` that keeps the mapping in memory. For example, this could allow creating temporary tables with custom engines in the Cloud without using S3.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100371",
        "createdAt": "2026-03-22T15:18:44Z",
        "updatedAt": "2026-08-13T17:29:44Z",
        "timestamp": "2026-08-13T17:29:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 32
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100374",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add symbols and lines to error tables",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add fields last_error_symbols, last_error_lines to system.errors and system.error_log table. Resolves https://github.com/ClickHouse/ClickHouse/issues/74561 ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100374",
        "timestamp": "2026-08-12T21:34:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-improvement",
          "can be tested",
          "pr-autogenerated-docs"
        ],
        "author": "npakeer",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100377",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Disable read-in-order when primary key selectivity is poor",
        "text": "When the WHERE clause cannot use the primary key effectively (e.g., `LIKE '%...'`), the `optimize_read_in_order` optimization kills parallelism: each part is read by a single stream instead of many. This can make `ORDER BY` queries 4x slower than the same query without `ORDER BY`. Example on a 2.5B row table with `ORDER BY path` and `WHERE path LIKE '%stderr.log'`: - Without ORDER BY: ~9.5s - With ORDER BY (read-in-order): ~40s - With `optimize_read_in_order = 0`: ~9.6s The fix adds a runtime check in `ReadFromMergeTree::spreadMarkRanges`: if the primary key index selected more than a configurable fraction of all granules and there is no LIMIT, we fall back to parallel reading with per-stream `PartialSortingTransform` + `MergeSortingTransform`, so the output is still sorted as the `SortingStep` above expects. New setting `read_in_order_max_primary_key_ratio` controls the threshold. The default is 1.0, which preserves the previous behavior (never disable read-in-order based on primary key selectivity): the heuristic is opt-in until the default threshold is tuned against the performance report. Set it below 1.0 (e.g. 0.5) to enable the guard. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add a new setting `read_in_order_max_primary_key_ratio` that can disable the read-in-order optimization when the primary key selectivity is poor (more than the given fraction of granules selected), falling back to parallel reading with sorting. This avoids severe parallelism loss for queries like `SELECT ... WHERE path LIKE '%...' ORDER BY path`. The default 1.0 preserves the previous behavior; set a lower value (e.g. 0.5) to enable the heuristic. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) New setting `read_in_order_max_primary_key_ratio` (Float, default 1.0): maximum ratio of selected to total primary key granules for `optimize_read_in_order` to stay enabled. When the ratio exceeds this value, read-in-order is disabled in favor of parallel reading with per-stream sorting. The default 1.0 preserves the old behavior.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100377",
        "createdAt": "2026-03-22T18:02:39Z",
        "updatedAt": "2026-08-13T12:55:23Z",
        "timestamp": "2026-08-13T12:55:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 40
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100391",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallelize read-in-order from a single part with PrefetchingConcatProcessor",
        "text": "When a single MergeTree part is split into multiple streams for read-in-order, use the new `PrefetchingConcatProcessor` instead of relying on the downstream `MergingSortedTransform`. The key difference: - `MergingSortedTransform` marks all inputs as needed but only consumes from one stream at a time (for non-overlapping ranges), and each port can only buffer 1 block — so after initial prefetch, reading becomes sequential. - `PrefetchingConcatProcessor` marks all inputs as needed AND pulls data from non-current inputs into internal buffers, keeping upstream sources busy. It outputs data in strict input order (concatenation), which is correct because ranges from a single part are non-overlapping and pre-sorted. This avoids both the merge comparison overhead and the sort overhead, while enabling true parallel IO and filtering across range groups. Benchmark on 100M rows (4 threads, single part, poor PK selectivity): - Single-thread read-in-order: 1.57s - Parallel + sort (no read-in-order): 1.27s - PrefetchingConcat: 0.61s (2.6x faster) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve the performance of queries that are reading in the order of the primary key. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes read-in-order execution and introduces a new buffering processor in the query pipeline; incorrect gating or backpressure/buffering behavior could impact result ordering or memory usage in some MergeTree reads. > > **Overview** > **Adds `PrefetchingConcatProcessor`**: a new concat-like processor that keeps *all* inputs marked needed and buffers per-input chunks (bounded by `max_buffered_chunks`) to prefetch non-current streams while still outputting streams strictly in order. > > **Wires it into MergeTree read-in-order**: `ReadFromMergeTree` now conditionally replaces `MergingSortedTransform` with `PrefetchingConcatProcessor` when reading in ascending order from a *single part* split into multiple streams and only when there is filtering work (`PREWHERE` or row-level filter); it also adds `prefer_multiple_streams`/`setPreferMultipleStreams()` so aggregation-/distinct-in-order paths can opt out to avoid collapsing parallel streams. > > **Tests/bench coverage updated**: adds a new stateless test asserting `PrefetchingConcat` appears only for single-part reads and preserves sort order, adjusts an existing buffering test to force multi-part behavior, and adds a performance scenario guarding against regressions for filter-less `ORDER BY` reads. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit c0ba40d2f630976681024d2ac941261c4abdbe57. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100391",
        "createdAt": "2026-03-22T20:21:02Z",
        "updatedAt": "2026-08-13T11:12:54Z",
        "timestamp": "2026-08-13T11:12:54Z",
        "metrics": {
          "reactions": 1,
          "comments": 41
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:100394",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallel read in order with multiple parts",
        "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve the performance of queries that are reading in the order of the primary key. Based on #100391.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/100394",
        "timestamp": "2026-08-12T23:22:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 70
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101039",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add leader election for non-replicated MergeTree on shared storage",
        "text": "Add leader election for non-replicated MergeTree tables on shared object storage (currently `S3`; `Azure` is implemented but not yet enabled, pending test coverage), enabling active/standby failover without external coordination (no Keeper). Uses conditional writes (`If-Match` / `If-None-Match`) on object storage to maintain a lease file with JSON content `{\"version\":1,\"leader_id\":\"...\",\"timestamp\":...}`. The leader renews its lease periodically; followers monitor and claim leadership when the lease expires. New MergeTree settings: - `leader_election` (Bool, default false) — enable leader election - `leader_election_heartbeat_interval` (Seconds, default 10) — lease renewal interval - `leader_election_session_timeout` (Seconds, default 30) — lease expiry threshold; must be at least 3x the heartbeat interval When not the leader, inserts, merges, mutations, and DDL are blocked; background data processing is skipped. Participating nodes should keep their clocks synchronized (NTP) to within `leader_election_session_timeout`; the conditional-write protocol always prevents split-brain, but excessive clock skew can cause leadership churn. Closes https://github.com/ClickHouse/ClickHouse/issues/91613 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `leader_election` setting for non-replicated MergeTree tables on shared object storage, enabling active/standby failover using conditional writes without external coordination. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) Three new MergeTree-level settings are added: - `leader_election` (Bool) — When enabled on a non-replicated MergeTree table stored on an `S3` object storage disk (`Azure` is implemented but rejected at table creation until it has test coverage), multiple ClickHouse instances sharing the same data path will elect a single leader. Only the leader performs writes, merges, and mutations. Followers act as read-only replicas and automatically claim leadership when the current leader's lease expires. - `leader_election_heartbeat_interval` (Seconds, default 10) — How often the leader renews its lease and followers check for an expired lease. - `leader_election_session_timeout` (Seconds, default 30) — How long a lease remains valid without renewal. Must be at least 3x `leader_election_heartbeat_interval`. Example: ```sql CREATE TABLE shared_table (x UInt64) ENGINE = MergeTree ORDER BY x SETTINGS leader_election = true, leader_election_heartbeat_interval = 10, leader_election_session_timeout = 30; ``` <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > High risk because it introduces new leader-election coordination and modifies core `MergeTree` write/DDL/drop and catalog cleanup paths; bugs could cause write unavailability or accidental deletion/incorrect visibility on shared storage. > > **Overview** > Adds a **Beta** `leader_election` mode for non-replicated `MergeTree` tables on shared S3/Azure, using a conditional-write lease file to ensure only one instance performs inserts/merges/mutations while followers remain read-only and can take over on failure. > > Introduces new MergeTree settings (`leader_election`, `leader_election_heartbeat_interval`, `leader_election_session_timeout`) with validation, new `MergeTreeLeaderElection*` metrics/events, and follower part-refresh + leader takeover sync to load new parts and advance block counters before enabling writes. > > Hardens destructive operations for shared-storage tables: blocks most `ALTER`/partition-mutation operations and `RENAME` under leader election, gates background processing on leadership, and updates drop/cleanup logic (`dropSkipsDataDirectoryCleanup`, `DatabaseCatalog`/`StorageTableProxy` fail-closed behavior) to avoid recursive deletion or hangs when tables are shared or cannot be materialized. Adds documentation and integration/stateless tests covering failover, metrics, validation, and rejection cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 78c1252915ba512bebf94a147932158e1bacecf0. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101039",
        "createdAt": "2026-03-28T21:56:21Z",
        "updatedAt": "2026-08-13T16:57:36Z",
        "timestamp": "2026-08-13T16:57:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 87
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101273",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add geo aggregate functions #80186",
        "text": "### Functions for geospatial aggregation #80186 Add three SQL aggregate functions for polygon set operations and convex hull computation: - `groupPolygonUnion` — union of polygonal geometries in a group - `groupPolygonIntersection` — intersection of polygonal geometries in a group - `groupConvexHull` — convex hull of grouped point, linear, and polygonal geometries The functions support typed Geo inputs and `Geometry` variant columns, `State` / `Merge` combinators, binary aggregate-state serialization with writer/reader invariant checks and corruption guards, and well-defined empty-geometry semantics. Closes: https://github.com/ClickHouse/ClickHouse/issues/80186 ### Changelog category: - New Feature ### Changelog entry Added geospatial aggregate functions `groupPolygonUnion`, `groupPolygonIntersection`, and `groupConvexHull`. ### Documentation entry for user-facing changes - [x] Documentation is written in code",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101273",
        "createdAt": "2026-03-30T21:25:08Z",
        "updatedAt": "2026-08-13T15:50:16Z",
        "timestamp": "2026-08-13T15:50:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 36
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "zhemalb",
        "state": "open",
        "assignees": [
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101512",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `date_time_overflow_behavior` being ignored for integer and float casts to DateTime64/Time64",
        "text": "Fixes `date_time_overflow_behavior` being silently ignored when casting integer and `Float32`/`Float64` values to `DateTime64` and `Time64`. Previously, overflowing values could skip the `VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE` exception in `throw` mode and skip clamping to the target boundary in `saturate` mode. The fix routes native and wide integer sources (`Int8`–`Int256`, `UInt8`–`UInt256`) as well as `Float32`/`Float64` through overflow-aware, scale-aware transforms, and scales `Float32` inputs in the `Float64` domain so the result is no longer distorted by source-float precision. `Decimal*` and `BFloat16` sources are out of scope here — their scale-aware overflow handling belongs in `DataTypesDecimal.cpp`, not in `FunctionsConversion.h` — and are left for a follow-up. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `date_time_overflow_behavior` for integer and `Float32`/`Float64` casts to `DateTime64` and `Time64`: overflowing values now throw `VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE` in `throw` mode and clamp to the target boundary in `saturate` mode; `Float32` inputs are scaled in `Float64` precision, correcting previously imprecise results. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101512",
        "createdAt": "2026-04-01T14:42:05Z",
        "updatedAt": "2026-08-13T01:35:01Z",
        "timestamp": "2026-08-13T01:35:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 25
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yariks5s",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101514",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support `SIMILAR TO` pattern matching predicate",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `SIMILAR TO` pattern matching predicate ### Documentation entry for user-facing changes - [ x ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> - Motivation: SQL standard has SIMILAR TO predicate that combines SQL LIKE wildcards (%, _) with regular expression metacharacters `|, *, +, ?, [...], {m,n}, (...)`. Note that excluded are `^, $, .`. It is supported by PostgreSQL and others. See the [PostgreSQL documentation on pattern matching](https://www.postgresql.org/docs/current/functions-matching.html#FUNCTIONS-SIMILARTO-REGEXP) for reference. I also added: negation and substring pattern shortcut. Syntax highlighting is also hooked in to highlight the metacharacters. - Parameters: None. - Example use: ```sql SELECT similarTo('ClickHouse', 'Cl_ck[hH]ouse'); -- Returns: 1 SELECT 'ClickHouse' NOT SIMILAR TO 'Cl%ck[0-9]ouse'; -- Returns: 1 ``` - Corner cases - Character class inside bracket expression. - Bracket expressions don't nest. So `[[:a]]` is invalid, but up to RE2 to process. Addresses https://github.com/ClickHouse/ClickHouse/issues/99608. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds new SQL operator/function surface area and a new pattern-to-regexp translator, which can affect query parsing and string filtering semantics; risk is moderate due to edge cases in escaping/bracket handling and regex compilation paths. > > **Overview** > Adds SQL-standard `SIMILAR TO` pattern matching (and `NOT SIMILAR TO`) implemented via new `similarToPatternToRegexp()` conversion and new `similarTo`/`notSimilarTo` functions wired into `MatchImpl`/`Regexps::createRegexp`. > > Extends the parser to recognize and pretty-print `SIMILAR TO` / `NOT SIMILAR TO`, updates CNF NOT-pushing inversions, and adds syntax-highlighting support for SIMILAR TO metacharacters. > > Updates call sites to the new `Regexps::createRegexp<like, similar_to, ...>` template signature and expands query fuzzer swap sets plus a new stateless test suite covering SIMILAR TO behavior and corner cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit e17d9fd7d3d079c5528a281d7a9af6781ecff4dd. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101514",
        "timestamp": "2026-08-12T21:17:33Z",
        "metrics": {
          "reactions": 1,
          "comments": 17
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "zheguang",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101783",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support specifying or auto-assigning Parquet field_ids for output columns",
        "text": "### What this PR does Two related knobs for writing Parquet files with column `field_id` metadata — needed for Apache Iceberg compatibility, which identifies columns by `field_id` rather than by name. Both knobs can be used independently or together. #### 1. Explicit per-column overrides New setting `output_format_parquet_column_field_ids : Map(String, Int32)`. ```sql SET output_format_parquet_column_field_ids = '{\"col_a\": 10, \"col_b\": 20, \"col_c\": 30}'; INSERT INTO FUNCTION file('data.parquet') SELECT 1::UInt32 AS col_a, 'x'::String AS col_b, 42::Int64 AS col_c; ``` Each output column gets the id specified in the map. The map must cover every output column (unless auto-assign below is enabled to fill the gaps), and ids must be unique; violations raise `BAD_ARGUMENTS` with specific messages. #### 2. Auto-assign (Iceberg writer convention) New setting `output_format_parquet_auto_assign_field_ids : Bool`, default `false`. When enabled, every output column is assigned a unique sequential `field_id` starting at `1`, matching the convention used by Apache Iceberg writers. ```sql SET output_format_parquet_auto_assign_field_ids = 1; INSERT INTO FUNCTION file('iceberg.parquet') SELECT 1 AS a, 'x' AS b, 42 AS c; -- a -> 1, b -> 2, c -> 3 ``` #### 3. Mixed mode The two settings compose: override-map entries win for the columns they mention; auto-assign fills the remaining columns with the smallest unused positive ids. ```sql SET output_format_parquet_auto_assign_field_ids = 1, output_format_parquet_column_field_ids = '{\"b\": 1}'; INSERT INTO FUNCTION file('mixed.parquet') SELECT 1 AS a, 2 AS b, 3 AS c; -- b -> 1 (override), a -> 2, c -> 3 (auto-assign skipping 1) ``` ### Implementation - `src/Core/FormatFactorySettings.h` — declare both settings. - `src/Formats/FormatSettings.h` — carry a `std::vector<std::pair<String, Int32>>` of overrides and the `auto_assign_field_ids` bool; no string parsing on the hot path. - `src/Formats/FormatFactory.cpp` — convert the `Map` setting into the pair vector, with structural validation (tuple shape, string keys, integer values in `Int32` range). - `src/Processors/Formats/Impl/ParquetBlockOutputFormat.cpp` — `buildColumnFieldIds()` resolves overrides against the actual output header, auto-assigns unused ids if the flag is on, and rejects unknown columns, duplicate ids, and non-covering overrides. ### Test `tests/queries/0_stateless/04321_parquet_column_field_ids.sh` writes Parquet files with `clickhouse-local`, reads the `field_id` metadata back with `pyarrow`, and covers: - Explicit per-column overrides. - Auto-assign only. - Mixed override + auto-assign. - Writing with neither setting (no `field_id`s, unchanged behavior). - Nested types: auto-assign and dotted-path overrides for `Array.element`, `Tuple` subfields, and `Map.key`/`Map.value`. - Geo columns written as WKB under GeoParquet (the `Point`/`Array`/`Tuple` shape collapses to a single top-level field). - Error cases: unknown column, non-covering map (top-level and nested), duplicate id, non-integer value, negative id, and a dotted top-level name colliding with a nested path (in both auto-assign and override modes). Closes: https://github.com/ClickHouse/ClickHouse/issues/58753 ### Changelog category (leave one): - New Feature ### Changelog entry: Added `output_format_parquet_column_field_ids` (`Map(String, Int32)`) to set explicit Parquet `field_id` overrides per column and `output_format_parquet_auto_assign_field_ids` (`Bool`) to auto-assign sequential `field_id`s to every output column, matching the Apache Iceberg writer convention. The two settings can be combined. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches the Parquet write path and schema generation to emit `field_id` metadata, which could affect interoperability and schema expectations; defaults are preserved unless new settings are enabled. > > **Overview** > Adds two new Parquet output settings to control column `field_id` metadata for Iceberg compatibility: `output_format_parquet_column_field_ids` (explicit per-column overrides) and `output_format_parquet_auto_assign_field_ids` (sequential auto-assignment starting at 1). > > `FormatFactory` now parses/validates the `Map` setting into typed overrides up-front, while `ParquetBlockOutputFormat` resolves the final `field_id` mapping (including error checks for unknown columns, duplicates, negative ids, and incomplete coverage when auto-assign is off) and explicitly rejects using these settings when a datalake writer provides its own column-id mapping. > > Documentation and a new stateless test are added to verify correct `field_id` emission via `pyarrow` and expected failure modes. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit ce5f371493174b5b0ee58307dfc05f5df32af0b2. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101783",
        "createdAt": "2026-04-04T18:49:42Z",
        "updatedAt": "2026-08-13T04:43:18Z",
        "timestamp": "2026-08-13T04:43:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 43
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested"
        ],
        "author": "Onyx2406",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101791",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "In case of trivial views, push whole outer query to shards.",
        "text": "### Changelog category: - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): In case of trivial views over distributed table push whole outer query to shards. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Description / Proposed Solution When a VIEW is defined over a Distributed table, ClickHouse traditionally executes it on the shards without enclosing outer query. This means filters and expressions declared in the outer query are evaluated on the coordinator after pulling raw data from shards. For views whose body is a plain SELECT (column references, *, or arbitrary expressions — but no aggregation, grouping, ordering, joins, window functions, or scalar subqueries) over a single Distributed table, we can do better: inline the view body as a subquery and hand the whole thing to StorageDistributed. Each shard then receives the full outer query with the view body inlined, evaluates it against its local table, and only ships the result back. A view qualifies as \"trivial\" if its inner query: - Has a single SELECT (no UNION) - Selects only column references, *, or expressions — but no window functions (require the full dataset) and no scalar subqueries in the SELECT list - Has no WITH, PREWHERE, GROUP BY, HAVING, QUALIFY, ORDER BY, LIMIT, LIMIT BY, DISTINCT, or ARRAY JOIN - Has no subqueries in the WHERE clause - Reads from exactly one table with no joins, no table functions, no FINAL, no SAMPLE - Is not a parameterized view and does not use SQL SECURITY DEFINER **The optimization can be disbaled by setting (enabled by default):** ``` SET optimize_trivial_view_pushdown_to_distributed = 0; ``` ### Example: Env setup: ``` create table x engine = MergeTree ORDER BY tuple() AS SELECT intDiv(number,100000) as a, number as b FROM numbers(1000000000); SET prefer_localhost_replica = 0; CREATE TABLE x_dist AS x ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), x); CREATE VIEW v_computed AS SELECT a + 1 AS x, b AS y FROM x_dist WHERE a != 0; ``` Performance: ``` :) SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: 6baedc55-c8e7-4b2b-9946-b0d828abaf25 ┌─plus(a, 1)─┬──────────sum(b)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └────────────┴─────────────────┘ 1 row in set. Elapsed: 17.199 sec. Processed 2.00 billion rows, 32.00 GB (116.29 million rows/s., 1.86 GB/s.) Peak memory usage: 38.58 MiB. :) SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 7c4f2854-3de2-4d40-b960-efcc44a7b26d ┌─────x─┬──────────sum(y)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └───────┴─────────────────┘ 1 row in set. Elapsed: 16.497 sec. Processed 2.00 billion rows, 32.00 GB (121.24 million rows/s., 1.94 GB/s.) Peak memory usage: 38.82 MiB. ``` Plan: ``` :) explain SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; EXPLAIN SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: e2d32d5f-7024-4e6a-b686-45c7c0d5f2ef ┌─explain──────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression (Before GROUP BY) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression (Before GROUP BY) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └──────────────────────────────────────────────────────────────────────────────────┘ :) explain SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; EXPLAIN SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 88213305-46a5-493e-862c-a80c675c9452 ┌─explain───────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes query planning/execution for a subset of views over `Distributed` tables and touches access checks/row policy enforcement and SQL SECURITY semantics, which can affect correctness and security-sensitive behavior. > > **Overview** > Adds a new default-on setting `optimize_trivial_view_pushdown_to_distributed` to inline *trivial* views over `Distributed` tables and push the full outer query down to shards, reducing coordinator-side filtering/processing and network transfer. > > Implements planner rewrites to swap the view table expression with an analyzed subquery, merge `FINAL`/`SAMPLE` modifiers, and preserve semantics by suppressing pushdown when the outer query contains non-deterministic functions, while also explicitly handling SQL SECURITY modes, row-policy injection/logging, and column-pruned privilege checks. > > Extends integration/stateless tests to cover modifier propagation, non-determinism suppression, row-policy enforcement, SQL SECURITY behavior, and interactions with `max_rows_to_read_leaf`. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 583e1e6c0e8e25081391d7a07af086c6f9888c6f. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101791",
        "createdAt": "2026-04-04T19:32:01Z",
        "updatedAt": "2026-08-13T17:53:08Z",
        "timestamp": "2026-08-13T17:53:08Z",
        "metrics": {
          "reactions": 3,
          "comments": 16
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "simonmichal",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:101841",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add groupBloomFilter aggregate function and bloomFilterContains scalar function",
        "text": "Adds the `groupBloomFilter` aggregate function and the `bloomFilterContains` scalar function for memory-efficient probabilistic set-membership testing. This can be used to detect new values, perform approximate deduplication checks, and compare large datasets without materializing exact sets. Related: https://github.com/ClickHouse/ClickHouse/issues/11700 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `groupBloomFilter` aggregate function for building Bloom filter states and the `bloomFilterContains` scalar function for testing whether values are probably present. The functions support configurable expected element counts, false-positive rates, and seeds, and work with the `-State` and `-MergeState` combinators and `AggregatingMergeTree`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) Documentation is added in `docs/en/sql-reference/aggregate-functions/reference/groupbloomfilter.md` and `docs/en/sql-reference/functions/bloom-filter-functions.md`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/101841",
        "createdAt": "2026-04-06T07:19:46Z",
        "updatedAt": "2026-08-13T13:39:51Z",
        "timestamp": "2026-08-13T13:39:51Z",
        "metrics": {
          "reactions": 13,
          "comments": 5
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested"
        ],
        "author": "otselnik",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:102033",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix LOGICAL_ERROR crash in IcebergMetadata::iterate when datalake_table_state is missing",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix `LOGICAL_ERROR` exceptions when reading `Iceberg` or `DeltaLake` data lake tables through paths that can reach the read pipeline without a pinned `datalake_table_state`, such as concurrent `Iceberg` metadata updates or `merge` reads over `DeltaLake` tables. ### What this PR does **Problem:** `IcebergMetadata::iterate` and `isDataSortedBySortingKey` throw a `LOGICAL_ERROR` (`Can't extract iceberg table state from storage snapshot`) when `datalake_table_state` is not set in the storage metadata snapshot. The same family of failures hits `DeltaLake` (`No version found in table state snapshot`). This raises an exception (and, in stress tests, a server abort under `abort_on_logical_error`) — tracked as STID `2606-4a47`, a chronic failure with master and PR hits across many sanitizer/build variants over the last 30 days. `PR #101577` fixed one such path (materialized views accessing `Iceberg` target tables), but the failure continued on master because other code paths can reach the read pipeline without `datalake_table_state`: - Concurrent metadata updates create a TOCTOU race between `updateExternalDynamicMetadataIfExists` setting the state (via `setInMemoryMetadata`) and the next `getInMemoryMetadataPtr` reading it. - Cluster table functions and other paths that bypass the analyzer/interpreter reach `read`/`iterate` without having had the state pinned. - `merge` reads over `DeltaLake` tables reach the read step with a snapshot that lacks the pinned state. **Fix:** Pin a single, internally consistent data lake snapshot at the source instead of papering over the missing state deep inside the callees. This follows the reviewer's guidance to *fix all the places to create the table state snapshot instead of ad-hoc fallbacks in the callees*: - `StorageObjectStorage::read` pins one coherent `datalake_table_state` **before** `prepareReadingFromFormat`, so the requested columns, the field-id mapping used to list/read files, and the read-in-order sorting key all come from the same data lake snapshot. For cluster table functions it first calls `lazyInitializeIfNeeded`, otherwise `update`, so `getTableStateSnapshot` does not hit an uninitialized-metadata assertion. When `shouldReloadSchemaForConsistency` is set, columns and sorting key are rebuilt from the same state via `buildStorageMetadataFromState` so they cannot diverge. - `StorageObjectStorageSource::createFileIterator` repopulates `datalake_table_state` from `getTableStateSnapshot` before `iterate`, as a safety net for the remaining callers. - `ReadFromObjectStorageStep::getDataOrder` reads the sorting key from the pinned `storage_snapshot->metadata` rather than re-deriving it. The `LOGICAL_ERROR` assertions inside `iterate` and `isDataSortedBySortingKey` are preserved untouched as safety nets — they should now be unreachable on the fixed paths. A test-only failpoint `datalake_simulate_missing_table_state` strips the pinned state from the snapshot right before the read step, so the regression test reproduces the missing-state condition deterministically. ### Tests - `04305_iceberg_missing_table_state.sh` — deterministic reproducer using the `datalake_simulate_missing_table_state` failpoint; a non-trivial read (`sum`) and an `ORDER BY` query exercise both the `iterate` and `isDataSortedBySortingKey` sites. - `04306_delta_lake_merge_missing_table_state.sh` — `merge` read over a `DeltaLake` table, the `DeltaLake` variant of the same bug. - `04141_iceberg_concurrent_no_logical_error.sh` — concurrent readers and a writer producing fresh snapshots, exercising the TOCTOU race in the regular planner path. Closes: https://github.com/ClickHouse/ClickHouse/issues/102037 Closes: https://github.com/ClickHouse/ClickHouse/issues/107334 Related: https://github.com/ClickHouse/ClickHouse/issues/93278 <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.1.448` (included in `26.7` and later) - Backported to: `26.5.7.46` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/102033",
        "createdAt": "2026-04-08T09:17:23Z",
        "updatedAt": "2026-08-13T01:36:03Z",
        "timestamp": "2026-08-13T01:36:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 36
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "can be tested",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "SmitaRKulkarni"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:102192",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix \"Not-ready Set\" exception when buildOrderedSetInplace fails",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/107924 `FutureSetFromSubquery::buildOrderedSetInplace` (the speculative set build run during primary key analysis for `IN` subqueries) used to consume the subquery `source` plan up front. If the in-place build then failed silently (e.g. due to subquery timeout with `timeout_overflow_mode = 'break'`, where the executor stops without throwing and without setting `is_created`), the set remained permanently unbuilt and `FunctionIn` threw \"Not-ready Set is passed as the second argument\". The core fix — running the in-place pipeline against a clone of the source plan so that `DelayedCreatingSetsStep::makePlansForSets` can still build the set after a silent failure — has meanwhile been merged in #107924, which falls back to the original destructive build when the source plan contains a step that does not implement `IQueryPlanStep::clone`. After merging master, this PR carries the remaining hardening on top of #107924: - `clone` implementations for eleven more query-plan steps: `JoinStep`, `FillingStep`, `ExtractColumnsStep`, `FractionalLimitStep`, `FractionalOffsetStep`, `MergingAggregatedStep`, `NegativeLimitStep`, `NegativeLimitByStep`, `NegativeOffsetStep`, `StreamInQueryResultCacheStep`, `ReadFromQueryResultCacheStep`. `IN` subqueries whose plans contain these steps (e.g. `ORDER BY ... WITH FILL`, joins planned as `JoinStep`, subqueries going through the query result cache) previously took the destructive fallback, where a silent in-place failure still reproduces the \"Not-ready Set\" exception; with clone support they take the non-destructive path and recover. - `JoinStep::clone` covers only the join algorithms whose `IJoin` implementation supports cloning: `HashJoin`, `ConcurrentHashJoin`, `ConstantJoin` and `FullSortingMergeJoin`. `JoinSwitcher` (`join_algorithm = 'auto'`), `SpillingHashJoin`, `MergeJoin` / `PartialMergeJoin` and `GraceHashJoin` keep the pre-existing destructive fallback: `IJoin::isCloneSupported` also gates `optimizeJoin`, `convertOuterJoinToInnerJoin` and `QueryPipelineBuilder`, so widening it would silently enable join swapping and outer-to-inner conversion for never-validated stateful spilling algorithms. This limitation is documented at the throw site, and `04649_not_ready_set_with_non_clonable_join_algorithms` pins that those algorithms still produce correct results (the `NOT_IMPLEMENTED` stays internal). - `FilledJoinStep` (`StorageJoin` and dictionary joins) and `JoinStepLogicalLookup` (direct key-value joins) intentionally remain non-clonable as well: they wrap live storage-backed join state with no safe copy semantics (`IJoin::clone` constructs an empty join rather than copying the filled state). The constraint is documented at both classes. - Copying a `QueryResultCacheWriter` (which is what `StreamInQueryResultCacheStep::clone` does) now carries over both the original's `query_start_time` and its `skip_insert` decision. On the successful in-place build path the copy is the only writer that is ever finalized, so otherwise `query_cache_min_query_duration` would be measured from the clone point and an already cache-resident key would be buffered into a throwaway buffer instead of being skipped. Pinned by `04652_query_result_cache_subquery_clone_min_query_duration` and `04653_query_result_cache_subquery_clone_skip_insert`. - `QueryPlan::clone` now preserves `max_threads` and `concurrency_control`, so a pipeline built from a cloned plan runs under the same resource contract as the original instead of defaulting to `max_threads == 0` with concurrency control disabled. - The Planner no longer plants query result cache steps into a *logical* plan. A logical plan is not executed where it is built: it is serialized and shipped to another node (parallel replicas with `serialize_query_plan = 1`, see `createRemotePlanForParallelReplicas`), while `StreamInQueryResultCacheStep` and `ReadFromQueryResultCacheStep` hold node-local state (a `QueryResultCacheWriter`, or the cached chunks) that has no serialized representation. Before this, `query_cache_for_subqueries = 1` together with parallel replicas and `serialize_query_plan = 1` failed the whole query with `Method serialize is not implemented for StreamInQueryResultCache`. The cache is still populated and read by the plan the initiator executes itself. - Regression tests using the `prepared_sets_build_ordered_set_inplace_fail` failpoint: `04095_global_not_in_parallel_replicas` (the original parallel-replicas reproducer), `04492_not_ready_set_with_fill_subquery` (a `WITH FILL` subquery source plan), `04550_not_ready_set_with_join_subquery`, `04648_not_ready_set_with_query_result_cache_subquery` (both the cache write and the cache read path) and `04649_not_ready_set_with_non_clonable_join_algorithms`. The original failure was observed in the [Stress test (amd_tsan)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=d0432097aed783bd35054dce2edcefe0c4e5122c&name_0=MasterCI&name_1=Stress%20test%20%28amd_tsan%29) on master, where `considerEnablingParallelReplicas` triggers `selectRangesToRead`, which calls `buildOrderedSetInplace` for primary key analysis. Note on the changelog category: the \"Not-ready Set\" exception is a `LOGICAL_ERROR`, so reproducing it on the unfixed build aborts the server on the sanitizer builds that the per-arch Bugfix validation jobs run on, and the validation harness deliberately treats a mid-run server death as inconclusive rather than as a reproduced bug. Automated bugfix validation is therefore structurally impossible for this fix, and the user-visible bug fix for the common plan shapes already shipped with #107924; the remaining hardening here is classified as an Improvement. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Extend the non-destructive speculative in-place set build for `IN` subqueries (introduced in [#107924](https://github.com/ClickHouse/ClickHouse/pull/107924)) to more query-plan shapes: `clone` is implemented for eleven more query-plan steps, including clone-supported joins via `JoinStep`, `ORDER BY ... WITH FILL` via `FillingStep`, and the query result cache steps. Such subqueries now preserve the original source plan and can recover from a silent in-place build failure instead of throwing \"Not-ready Set is passed as the second argument\". A cloned query plan now preserves `max_threads` and `concurrency_control` of the original plan. Also fixed `Method serialize is not implemented for StreamInQueryResultCache` when `query_cache_for_subqueries = 1` is used together with parallel replicas and `serialize_query_plan = 1`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/102192",
        "createdAt": "2026-04-09T08:31:34Z",
        "updatedAt": "2026-08-13T12:18:14Z",
        "timestamp": "2026-08-13T12:18:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 53
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:102440",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "#91631 / Extension of CREATE TABLE IF NOT EXISTS ... AS SELECT",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added CREATE TABLE ... AS INSERT and CREATE TABLE IF NOT EXISTS ... AND INSERT syntax to populate a table immediately from SELECT, VALUES, or FORMAT data, with AND INSERT inserting even when the table already exists. This feature was requested on issue #91631 ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/102440",
        "timestamp": "2026-08-12T21:05:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "ivancadena98",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:102499",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Added Parquet Shredded VARIANT Support to ParquetReaderv3",
        "text": "This PR introduces the parquet Shredded VARIANT standard to CLickHouse. duckdb has this already and the main benefit is that it can read less bytes from disk. This is of course critical for a main parquet use case, which is remote reads from S3. Benchmark results for JSONBench reading from file() on NVME on the final binary are below. | query | method | wall s | user CPU s | OSReadBytes | |---|---|---:|---:|---:| | q1_collection_counts | parquet_variant_shredded | 0.330 | 0.08 | 4.1 MB | | | parquet_variant_unshredded | 0.620 | 0.69 | 133.9 MB | | | parquet_string | 0.750 | 1.25 | 133.6 MB | | | parquet_json | 4.830 | 4.57 | 131.8 MB | | q2_collection_users | parquet_variant_shredded | 0.420 | 0.22 | 17.4 MB | | | parquet_variant_unshredded | 0.660 | 0.83 | 133.7 MB | | | parquet_string | 0.830 | 2.08 | 133.3 MB | | | parquet_json | 4.910 | 5.26 | 131.2 MB | | q3_hourly_events | parquet_variant_shredded | 0.370 | 0.13 | 9.4 MB | | | parquet_variant_unshredded | 0.640 | 0.78 | 133.7 MB | | | parquet_string | 0.810 | 1.93 | 133.3 MB | | | parquet_json | 4.730 | 4.64 | 131.2 MB | | q4_first_posts | parquet_variant_shredded | 0.400 | 0.13 | 19.2 MB | | | parquet_variant_unshredded | 0.670 | 1.13 | 133.3 MB | | | parquet_string | 0.800 | 1.93 | 133.0 MB | | | parquet_json | 4.910 | 5.23 | 130.9 MB | | q5_activity_span | parquet_variant_shredded | 0.400 | 0.15 | 19.7 MB | | | parquet_variant_unshredded | 0.670 | 1.06 | 133.6 MB | | | parquet_string | 0.780 | 1.92 | 133.4 MB | | | parquet_json | 4.840 | 5.00 | 131.2 MB | ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add Parquet shredded VARIANT support, including read/write paths and subcolumn-aware read optimizations for semi-structured data. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/102499",
        "createdAt": "2026-04-12T14:56:07Z",
        "updatedAt": "2026-08-13T00:35:07Z",
        "timestamp": "2026-08-13T00:35:07Z",
        "metrics": {
          "reactions": 3,
          "comments": 31
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested"
        ],
        "author": "rorylshanks",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "scanhex12"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:103158",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix TSAN data race in executeASTFuzzerQueries clearing caller's transaction",
        "text": "The server-side AST fuzzer (`executeASTFuzzerQueries`, active only when `ast_fuzzer_runs > 0`) used to clear the transaction on the caller's query and session `Context` with no lock held: ```cpp context->getQueryContext()->getSessionContext()->setCurrentTransaction(NO_TRANSACTION_PTR); context->setCurrentTransaction(NO_TRANSACTION_PTR); ``` `Context::setCurrentTransaction` writes `merge_tree_transaction` unsynchronized, while other threads read the same field under the shared `Context::mutex` via `Context::createCopy`. Under `RESTORE ASYNC` with `ast_fuzzer_runs > 0`, `RestorerFromBackup` background workers keep copying the context while the fuzzer finish callback overwrites the transaction pointer, producing the TSAN race (STID 2604-385d read side, 3336-2c6d write side, seen on `Stress test`/`BuzzHouse (experimental, serverfuzz, *_tsan)`). ## Fix 1. Move the transaction reset onto the fuzz session-context copy, so the caller context is never mutated. 2. Skip fuzzed `BACKUP`/`RESTORE` queries. An async backup/restore returns from `executeQuery` immediately while `BackupsWorker` keeps the (fuzz) context alive for background workers that read it via `Context::createCopy`; mutating that escaped copy afterwards reintroduces the same race. The fuzzer has negligible value on backup/restore (no real backup target), so skipping them is safe. Tests: `04104_ast_fuzzer_preserves_caller_transaction` (caller transaction preserved across `ast_fuzzer_runs > 0`) and `04305_ast_fuzzer_skips_backup_restore` (server stays alive and data intact with fuzzed backup/restore). The whole change lives in fuzzer-only code gated on the `ast_fuzzer_runs` CI/stress-test setting, never a production user path, so this is categorized as a CI fix. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/103158",
        "createdAt": "2026-04-20T13:48:10Z",
        "updatedAt": "2026-08-13T02:17:00Z",
        "timestamp": "2026-08-13T02:17:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 20
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "Algunenano"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:103182",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Lightweight Updates v2",
        "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Lightweight `UPDATE` patch parts now use a new v2 on-disk format sorted by (`sorting_key..., _block_number, _block_offset`) and applied with a new merging algorithm. Peak memory is bounded by the largest equal-sort-key run instead of the full patch, and updates that cross merge boundaries no longer fall back to in-memory Join apply. Old-format patch parts remain readable. During a rolling upgrade from a version before 26.8, keep `patch_parts_version = 'v1'` or use the `compatibility` setting until all replicas are upgraded. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/103182",
        "createdAt": "2026-04-20T17:45:11Z",
        "updatedAt": "2026-08-13T16:29:39Z",
        "timestamp": "2026-08-13T16:29:39Z",
        "metrics": {
          "reactions": 2,
          "comments": 8
        },
        "labels": [
          "pr-backward-incompatible"
        ],
        "author": "CurtizJ",
        "state": "open",
        "assignees": [
          "alesapin"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:103483",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Improve sanitizer robustness in parser and UTF-8 subsequence paths",
        "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> This follows up on sanitizer failures discovered while running CI for https://github.com/ClickHouse/ClickHouse/pull/99740. The branch keeps the original `getURLScheme` empty-input guard from that work and adds two more hardening fixes that became visible after the first fix: - guard `Lexer::nextToken` `max_query_size` check against null input pointers (UBSan path); - return early in `HasSubsequenceImpl::hasSubsequenceUTF8` for empty haystack before first UTF-8 decode (MSan path). It also adds focused regression coverage: - parser gtest for null lexer input with `max_query_size`; - stateless SQL test for empty-haystack UTF-8 subsequence behavior. Original PR where issues were discovered: https://github.com/ClickHouse/ClickHouse/pull/99740 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved sanitizer robustness in parser and string-function edge cases: `protocol` now handles empty input safely in `getURLScheme`, `Lexer::nextToken` no longer performs null-pointer arithmetic in `max_query_size` checks, and UTF-8 subsequence evaluation returns early on empty haystacks before decoding. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-b9b84d83-fd58-456b-97e2-c3386af36d84\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-b9b84d83-fd58-456b-97e2-c3386af36d84\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/103483",
        "createdAt": "2026-04-23T23:58:04Z",
        "updatedAt": "2026-08-13T13:09:23Z",
        "timestamp": "2026-08-13T13:09:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [],
        "author": "yakov-olkhovskiy",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:103706",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "ReaderExecutor: pipeline-based read orchestration",
        "text": "Builds on top of #103234 (ReadPipeline). Gated by `SET use_reader_executor = 1` (experimental, default off). This commit replaces the matryoshka of `ReadBuffer` wrappers with a `ReaderExecutor` that owns offset mapping, cache decisions, prefetch and decryption in one place. Caches plug in via a uniform `ICacheProvider` / `ICacheHandle` API. ## What this brings **Zero-copy page-cache hits.** Reads served from `PageCache` reach the caller as a `shared_ptr` into the mmap'd cell — no `memcpy` between cache and working buffer (compare to `CachedInMemoryReadBufferFromFile`'s per-block copy). **Connection reuse across remote read calls.** Sequential reads against the same S3/Azure/HTTP object reuse an open buffer instead of issuing a new HTTP request per window. `SourceBufferLimit` caps live connections globally (`max_remote_read_connections`, default 1000); over-limit reads fall back to stateless open-read-close. Visible via `system.remote_read_connections`. **Coordinated cache + prefetch decisions.** The matryoshka design had each layer making independent choices (prefetch ahead in `AsynchronousBoundedReadBuffer`, decide what to cache in `CachedOnDiskReadBufferFromFile`, gather across objects in `ReadBufferFromRemoteFSGather`). The executor sees the whole window: - prefetches only what isn't already cache-hit - retains source over-read only when paired with a live connection (drops it otherwise) - knows which bytes are user-requested vs cache-fill, so it doesn't account cache-fill bytes against the user's read budget **Per-object cache identity, not per-pipeline.** `DiskCacheProvider` derives `FileCacheKey` / `FileCacheOriginInfo` per `StoredObject` (etag-keyed caching for `StorageObjectStorageSource`, segment-key-type classification for `Data` vs `System` queues). The old single-key-per-executor approach silently shared cache identity across all objects in a gather-mode read. **Foundation for new cache backends.** New caches (vector cache, distributed cache, etc.) implement one interface (`ICacheProvider::lookup`) and plug into the chain without touching read paths. ## Observability - `system.remote_read_connections` — open connections (path, query id, position, elapsed) - `system.reader_executor_log` — per-executor stats (cache hit/miss/populate bytes, prefetch hits/cancellations, source request count, decrypt time, …) - `ProfileEvents`: `ReaderExecutor*` counters and microsecond timers; `LiveSourceBuffer*` counters - `HistogramMetrics`: cache get/populate/source-read/prefetch-wait latencies ## Server settings - `reader_executor_prefetch_pool_size` (default 8) — shared prefetch thread pool size - `reader_executor_prefetch_queue_size` (default 0 = `pool_size * 10`) — pool queue depth - `max_remote_read_connections` (default 1000) — live-buffer slot count, reload-able ## Internals - `Rope` / `RopeNode` / `OwnedRopeBuffer` — refcounted buffer chains with mixed provenance (owned, page-cache-pinned), built-in cursor (`peek`/`advance`/`tryRewind`), sorted-on-insert nodes, disjoint-interval coverage tracking. - `ICacheProvider` / `ICacheHandle` — uniform cache API (lookup → status/get/put). Two implementations: `PageCacheProvider` (file-level, zero-copy), `DiskCacheProvider` (per-object, FileCache-backed). - `ISourceReader` — stateless range-read interface. `LocalSourceReader`, `ObjectStorageSourceReader`, `BufferSourceReader` (adapter for backup / `BufferCreator`). - `OffsetMap` — replaces `ReadBufferFromRemoteFSGather`'s gather logic with a logical-to-(object, object-offset) lookup; supports single-object unknown-size (S3 `HEAD` without `Content-Length` → streams to EOF). - `ReaderExecutor` — owns: position, offset map, cache chain, prefetch handle, live buffer, over-read tail, source-buffer slot. - `PipelineReadBuffer` — thin `ReadBufferFromFileBase` exposing the executor through `BufferBase::set/next/seek` (so legacy callers see no change). - `PrefetchThreadPool` — shared bounded pool, returns `nullptr` on overflow (sync fallback), task cancellation is race-free. ## Testing ~3,700 lines of tests (≈46% of the PR), 130+ gtest cases: - `gtest_rope` — 31 cases (append / peek / advance / tryRewind / slice / copyTo / coverage queries / shift). - `gtest_reader_executor` — 41+ cases (cache-chain combinations, per-object lookup, prefetch races, EOF release, slot leak fixes, unknown-size streaming, …). - `gtest_filecache` — 22 cases including over-read / bypass-mode / partial coverage scenarios. - `gtest_read_pipeline` — known/unknown size selection through the executor. - Functional: `04262_reader_executor_observability` plus opt-outs (`SET use_reader_executor = 0`) on tests that intentionally exercise the legacy path. ## Load tesing the simpliest cache API\" ``` ┌─────────┬────────────┬──────────┬───────┬────────────────────────┐ │ regime │ legacy QPS │ exec QPS │ QPS Δ │ per-query-type latency │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ cold │ 0.252 │ 0.320 │ +27% │ −12% (faster) ✅ │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ warm │ 0.407 │ 0.252 │ −38%¹ │ +27% (slower) ❌ │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ partial │ 0.247 │ 0.251 │ +2% │ +37% (slower) ❌ │ └─────────┴────────────┴──────────┴───────┴────────────────────────┘ ``` changed cache API (stream aware): ``` Full clean verdict — 9e9abe9d (26.6.1.1) ┌─────────────┬─────────────────────────┬───────────────────────────────────────────────────┐ │ Regime │ Executor vs legacy │ Note │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ cold │ −15% (win) │ 0 slot fail; coalesced reads │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ warm │ −8% (win) │ 0 slot fail; cache reads healthy │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ⚠️ partial │ +7% │ S3-source over-read / churn on uncached half │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ❌ populate │ +17% (worst) │ 50% connection reuse on cache-fill path — churn │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ stress │ 27 ok / 1 OOM vs 4 / 11 │ controlled degradation = net win (bounded, alive) │ └─────────────┴─────────────────────────┴───────────────────────────────────────────────────┘ stress details: ┌───────────────────────┬───────────┬──────────┐ │ │ legacy │ executor │ ├───────────────────────┼───────────┼──────────┤ │ queries finished (ok) │ 4 │ 27 │ ├───────────────────────┼───────────┼──────────┤ │ failed │ 11 │ 1 │ ├───────────────────────┼───────────┼──────────┤ │ OOM │ 11 │ 1 │ ├───────────────────────┼───────────┼──────────┤ │ connection reuse │ 86.8% │ 98.8% │ ├───────────────────────┼───────────┼──────────┤ │ connection resets │ 233,618 │ 38,581 │ ├───────────────────────┼───────────┼──────────┤ │ peak TCP recv-buffer │ 4,152 MiB │ 813 MiB │ ├───────────────────────┼───────────┼──────────┤ │ TCP sockets │ 17,333 │ 1,548 │ ├───────────────────────┼───────────┼──────────┤ │ cgroup mem │ 24.4 GiB │ 24.6 GiB │ └───────────────────────┴───────────┴──────────┘ sha -- 494fcfe8825c ┌─────────┬──────────────┬──────────────┬──────────────┐ │ regime │ baseline QPS │ executor QPS │ exec/base │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ cold │ 0.207 │ 0.292 │ 1.41× (+41%) │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ warm │ 0.330 │ 0.307 │ 0.93× (−7%) │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ partial │ 0.375 │ 0.420 │ 1.12× (+12%) │ └─────────┴──────────────┴──────────────┴──────────────┘ ``` That is good outcome. I need to optimize the work with caches. Need to stream data from available cache segments without any additional costs. sha -- dd4ae00fdaac (26.7.1.1) ``` ┌──────────┬──────────────┬──────────────┬──────────────┐ │ regime │ baseline QPS │ executor QPS │ exec/base │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ cold │ 0.230 │ 0.328 │ 1.43× (+43%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ warm │ 0.244 │ 0.302 │ 1.24× (+24%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ partial │ 0.332 │ 0.325 │ 0.98× (−2%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ populate │ 0.348 │ 0.384 │ 1.10× (+10%) │ └──────────┴──────────────┴──────────────┴──────────────┘ ``` First build with every regime at parity or better. The warm coordination-CPU tax and the populate cache-write regression are both gone; populate over-read down to 14%. Run under production-default networking (`disk_connections_rcvbuf=204800`, `max_remote_read_connections=1000`), with `reader_executor_use_long_connections=1` set explicitly (its default flipped to 0 on this build). Long-connection hit rate 99.5–100% on all regimes. sha -- fcccec1c7e6c (26.7.1.1) — equal-workload methodology ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 959 s │ 705 s │ 1.36× (+36%) │ │ warm │ 665 s │ 450 s │ 1.48× (+48%) │ │ partial │ 592 s │ 540 s │ 1.10× (+10%) │ │ populate │ 552 s │ 516 s │ 1.07× (+7%) │ │ evict │ 601 s │ 565 s │ 1.06× (+6%) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Methodology fix: both arms now run the identical query multiset (`--iterations`, sequential — no `--randomize`), metric = total wall time. The earlier windowed-QPS numbers systematically penalized the faster arm (a free benchmark worker draws a new random query, so the faster arm attracts more heavy queries) — warm was reported −9…−14% but is actually **+48%**; `evict` = new regime with the FileCache as a transit buffer (continuous eviction, no cross-query hits). The executor is faster in all five cache regimes; the 128-thread stress arm now completes without stuck queries (previously required `KILL QUERY`), with 4× fewer sockets than legacy. No correctness errors or crashes across the campaign since `c089b7e5`. Known remaining costs (per-query analysis): (1) cold/partial S3 over-read 2.2–2.8× from prefetch speculation; (2) on mixed cache regimes small queries regress 2-3× — `cache_get` returns zero bytes for data that is present (`FileSegmentWait` on DOWNLOADING segments, then reads from source anyway) plus block-granular request amplification on narrow columns (~90 MiB requested for a 5 MiB column), paid at S3 first-byte latency with ~90% of prefetches cancelled. sha -- 018c0257567b (26.7.1.1) — plan-look-ahead window build ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 810 s │ 619 s │ 1.31× (+31%) │ │ warm │ 477 s │ 450 s │ 1.06× (+6%) │ │ partial │ 785 s │ 518 s │ 1.52× (+52%) │ │ populate │ 532 s │ 478 s │ 1.11× (+11%) │ │ evict │ 579 s │ 581 s │ 1.00× (par) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Same equal-workload methodology (identical 186-query multiset per arm). Since the workload is fixed, executor wall times are directly comparable across builds: vs `fcccec1c` the executor improved on cold (705→619 s, −12%), populate (516→478 s, −7%) and partial (540→518 s, −4%) — the plan-window changes helped. Baseline wall times swing between rounds (legacy path + master merge + day variance), so cross-build conclusions should use the executor-vs-executor columns, not the ratios. Also corrected with the fixed workload: the true equal-work S3 over-read is ~2.0× on cold and ~2.3× on partial (the earlier 2.8–2.9× figures were inflated by the windowed methodology — the faster arm simply ran more queries per window). Stress arm self-completed again (4th consecutive build). No `Code: 33`, no crashes. sha -- 2b43930ebbcf (26.8.1.1) — read-path log cut + settings-history dedup ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 496 s │ 315 s │ 1.57× (+57%) │ │ warm │ 223 s │ 226 s │ 0.99× (par) │ │ partial │ 326 s │ 254 s │ 1.28× (+28%) │ │ populate │ 242 s │ 240 s │ 1.01× (par) │ │ evict │ 281 s │ 304 s │ 0.92× (−8%) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Two rounds this cycle. The morning round (head `a6465d03`) exposed a hot-read-path logging tax: `PipelineReadBuffer` logged one `Trace` message per window advance (6.3M messages per warm arm) plus ~1M per-buffer `Debug` \"Created\" messages, costing 795 s of thread time per arm (baseline: 6.7 s) — roughly 14% of the CPU budget on cache-served regimes, which showed as warm/populate 0.95×. The same-day fix (`da34b431`) is verified by this round: logger time dropped 795 → 170 s, allocation volume dropped 5.85 → 1.94 TB ≈ baseline's 1.84 TB (most of the long-standing \"executor allocates ~4× more\" observation was log-message formatting), and executor `UserTime` is now below baseline on warm. Populate flipped to parity-plus; warm is at 0.99× with the residual logger time still 34× baseline — a few more sites may be worth demoting. S3 over-read: populate 1.01× and evict 1.03× (the executor no longer over-fetches where the cache absorbs writes), cold 1.20× while making 7× fewer GETs at ~11 MiB each (this is what buys the 1.57× cold wall), partial 1.38× — the remaining cost item. The evict 0.92× turned out to be regime noise: a clean repeat pair on the settled cluster came back 323/279 s = 1.16× in the executor's favor, matching the morning round's 1.17× (evict has a history of single-pair swings — 0.78× → 1.18× between two runs of one build earlier). Verdict: executor ≥ baseline in all five regimes. Environment notes for cross-entry comparison: starting with this cycle the staging instance's default profile `compatibility` was bumped `25.12` → `26.8`, which changes effective defaults for both arms — all walls in this entry are ~2× faster than in previous entries for that reason; compare ratios, not walls, across entries. Relatedly, the settings-history dedup in this head means Cloud `compatibility` no longer resurrects the pre-release 8 MiB `reader_executor_plan_look_ahead_max_window` introduction default on private builds. Zero failed queries, no `Code: 33`, no crashes across all 20 arms of both rounds. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental `ReaderExecutor` for pipeline-based read orchestration with unified cache API (`PageCacheProvider`, `DiskCacheProvider`), connection reuse for remote reads, shared prefetch pool, and `system.remote_read_connections` / `system.reader_executor_log` observability tables. Enable with `SET use_reader_executor = 1`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/103706",
        "createdAt": "2026-04-29T07:27:30Z",
        "updatedAt": "2026-08-13T17:03:28Z",
        "timestamp": "2026-08-13T17:03:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "CheSema",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104217",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix SQLite WHERE predicate pushdown for strings with special characters",
        "text": "`StorageSQLite::read` used `LiteralEscapingStyle::Regular`, which escapes single quotes as `\\'`. SQLite does not recognise backslash escapes; its only valid string escape is `''`. A pushed-down predicate like `WHERE col = 'it\\'s'` causes SQLite to parse `'it\\'` as a closed string and `s'` as a stray token — a SQL syntax error or injection vector. Switching to `LiteralEscapingStyle::PostgreSQL` would fix single quotes but still emit `\\n`, `\\r`, `\\t` as backslash sequences (which `writeAnyEscapedString` applies unconditionally). SQLite does not interpret those, so predicates on control-character strings would silently return no rows. This PR adds a dedicated `LiteralEscapingStyle::SQLite` backed by `writeQuotedStringSQLite`: only `'` → `''`; all other bytes (including `\\`, newline, tab) are embedded literally. NUL bytes cannot be embedded — SQLite's tokenizer loop in `sqlite3GetToken` terminates on `c==0` even inside a string literal, returning `TK_ILLEGAL` — so a predicate whose string literal (possibly nested in an `IN` tuple, array or map) contains a NUL byte is not pushed down at all: ClickHouse evaluates it locally, and with `external_table_strict_query = 1` the query is rejected instead of silently returning wrong rows. This is a follow-up to PR #74144 which fixed the DDL/PRAGMA and INSERT paths for SQLite but left the SELECT pushdown path using the wrong escaping style. During review the same class of bug was fixed on the PostgreSQL pushdown path as well: strings nested inside `Array` / `Tuple` / `Map` literals (e.g. the elements of a pushed-down `IN` list) now stay in the selected dialect all the way down instead of falling back to the regular ClickHouse escaping, and PostgreSQL string literals that contain backslashes or control characters are emitted as escape string constants (`E'...'`), so the server reads back exactly the original bytes regardless of `standard_conforming_strings` (a real tab used to be sent as the two characters `\\t`). Predicates whose string literals contain a NUL byte are not pushed down to PostgreSQL either, since a PostgreSQL string value cannot contain NUL. The same row-value restriction is applied on the normal `WHERE` pushdown path: a multi-column tuple is written as the row value `(a, b)`, which SQLite and MySQL accept only next to a comparison or `IN`, so a predicate such as `WHERE (id, val) IS NOT NULL` is no longer pushed down to them (ClickHouse evaluates it, and with `external_table_strict_query = 1` the query is rejected) instead of being sent as SQL the external database cannot parse (SQLite reports `row value misused`). For PostgreSQL, whose row constructors are ordinary value expressions, it is still pushed down. A tuple used as the whole condition is ClickHouse's list-of-predicates form and keeps being pushed down as a conjunction, `WHERE (\"a\" > 0) AND (\"column\" > 10)`, for every dialect - no external database accepts a row value as a condition. The user-provided `(SELECT ...)` table argument of `sqlite` / `postgresql` / `mysql`, which is re-serialized from the parsed AST and sent to the external database as is, no longer leaks ClickHouse-only syntax into that SQL: `Array` / `Map` literals and tuples with fewer than two elements (which could only be written back as `tuple(...)`) now throw `BAD_ARGUMENTS` instead of producing SQL the external database cannot parse, an explicit `tuple(a, b)` call is re-serialized as the parenthesized row value `(a, b)` - for SQLite and MySQL only in positions where those databases accept a row value (an operand of a comparison or `IN`); in any other position, such as the SELECT list, both the `tuple(...)` call and the equivalent tuple literal throw `BAD_ARGUMENTS`, because the parenthesized form is a syntax error there (SQLite reports `row value misused`). PostgreSQL row constructors are ordinary value expressions, valid in any expression position (`SELECT (a, b)`, `WHERE (a, b) IS NOT NULL`), so for PostgreSQL such tuples are sent through as row values everywhere instead of being rejected - everywhere except a boolean position, since no database accepts a record as a condition. A tuple in a boolean position - the `WHERE` / `HAVING` of the passed query, or an operand of `AND` / `OR` / `NOT` - is ClickHouse's list-of-predicates form, and is lowered to a conjunction for every dialect: `(SELECT ... WHERE (a > 0, b > 10))` reaches the external database as `WHERE (a > 0) AND (b > 10)`, the same rewrite the normal pushdown path applies; `PREWHERE`, which is ClickHouse-only syntax no external database can parse, is lowered into `WHERE` on that path as well (merging with an existing `WHERE` via `AND`), and the lowered filter gets the same boolean-position normalization. the equivalent tuple literal of constants is not a list of predicates the external database could evaluate and throws `BAD_ARGUMENTS` there instead. `array` / `map` calls on that path are rejected for all three databases. The internal `_CAST(literal, 'Type')` wrapper that the analyzer's `ConstantNode::toAST` puts around a tuple literal used as a plain expression operand (e.g. `WHERE (id, val) = (2, 'y')`) when it re-serializes the subquery argument from the query tree is unwrapped back to the literal, instead of leaking the ClickHouse-internal `_CAST` function into the SQL sent to the external database. A single-row multi-column `IN` set keeps its outer parentheses for both carriers - the fast-path literal `(a, b) IN ((1, 'x'))` and the explicit call `(a, b) IN (tuple(1, 'x'))` - so it reaches the external database as `IN ((1, 'x'))` instead of collapsing to the scalar list `IN (1, 'x')`. This normalization applies to MySQL as well: it shares the same re-serialization path, and although its `Regular` literal escaping style is correct for MySQL string literals (MySQL interprets backslash escapes like ClickHouse), the `tuple(...)` / `array(...)` / `map(...)` forms and `Array` / `Map` / single-element-tuple literals are not MySQL syntax either. The JDBC/ODBC (`StorageXDBC`) pushdown path is intentionally out of scope: the bridge protocol only reports the identifier quoting style, not the literal escaping dialect of the remote database, so plumbing a dialect-aware escaping style through it needs a bridge protocol extension. That path keeps the historical `Regular` escaping, and the limitation is now documented at the call site. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect SQL literal escaping in `StorageSQLite` and `sqlite()` table function when pushing `WHERE` predicates to SQLite: single quotes and control characters (`\\n`, `\\r`, `\\t`, `\\`) were escaped with backslashes, which SQLite does not interpret, causing syntax errors or wrong query results. Also fixed the escaping of string literals pushed down to PostgreSQL: strings nested inside `IN` lists kept ClickHouse escaping, and control characters were sent as backslash sequences that PostgreSQL reads back as different bytes. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104217",
        "createdAt": "2026-05-06T11:55:23Z",
        "updatedAt": "2026-08-13T15:33:55Z",
        "timestamp": "2026-08-13T15:33:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 18
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104350",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix correlated subquery + GROUP BY ROLLUP under group_by_use_nulls",
        "text": "A correlated subquery references an outer column that is a GROUP BY key of an outer query under `group_by_use_nulls = 1` with `WITH ROLLUP`/`CUBE`/`GROUPING SETS`. The actual values fed into the inner subquery are post-rollup `Nullable`s, but the analyzer used to leave the inner column reference at its non-Nullable type. The planner then built CAST wrappers and aggregate functions with the original signature, and at header inference (after decorrelation replaced `PLACEHOLDER` nodes with `INPUT` nodes of the actual `Nullable` types) the precomputed wrappers raised `Logical error: 'Bad cast from type DB::ColumnNullable to DB::ColumnVector<...>'`. The minimal reproducer from issue #91119 hits `createUInt8ToBoolWrapper`: ```sql SELECT (SELECT c0) FROM (SELECT 1::Bool) t0(c0) GROUP BY c0 WITH ROLLUP SETTINGS group_by_use_nulls = 1; ``` The same surface bug class also appears for plain numeric types (modulo, plus), aggregate functions over correlated outer columns (e.g. `(SELECT anyLastOrDefault(number))`), correlated `HAVING`/`WHERE` filter steps, and `Nullable` source columns. The root cause is in `QueryAnalyzer::resolveExpressionNode`: the `nullable_group_by_keys` walk stops at the first enclosing `QUERY` scope, so the outer scope where the column was defined is never consulted for an inner correlated reference. This change makes the walk continue past the inner `QUERY` scope when the `node` is a column whose source table expression is registered in a deeper outer scope (i.e. the column is correlated). The aggregate-function exclusion is now evaluated against each scope's `expressions_in_resolve_process_stack`, so an outer scope's `nullable_group_by_keys` applies even when the inner subquery is currently inside an inner aggregate (the inner aggregate operates on the outer's already-`Nullable` values). For correlated matches the existing node is mutated in place instead of being replaced with a clone of the GROUP BY key. The same `shared_ptr` is referenced by the enclosing `QueryNode::correlated_columns` list (registered via `addCorrelatedColumn`); cloning would leave the planner's `correlated_columns_set` pointing to the original non-`Nullable` node, and `PlannerActionsVisitor::visitColumn` (which uses type-sensitive equality) would then fail to identify the projected column as correlated and skip the `PLACEHOLDER` step that the decorrelation pass replaces. Closes #91119. Closes #106377. Closes https://github.com/ClickHouse/ClickHouse/issues/109509 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `Bad cast` exception in correlated subqueries that reference an outer column promoted to `Nullable` by `group_by_use_nulls = 1` with `GROUP BY` `WITH ROLLUP`/`CUBE`/`GROUPING SETS`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104350",
        "createdAt": "2026-05-08T08:49:47Z",
        "updatedAt": "2026-08-13T06:41:12Z",
        "timestamp": "2026-08-13T06:41:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 61
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "novikd"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104431",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallelize reads from a single Parquet file in StorageFile, again",
        "text": "Reverts ClickHouse/ClickHouse#104359",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104431",
        "createdAt": "2026-05-08T22:04:53Z",
        "updatedAt": "2026-08-13T03:55:52Z",
        "timestamp": "2026-08-13T03:55:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 68
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104435",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Geoparquet rowgroup pruning",
        "text": "Part of making ClickHouse fastest spatial analytical engine on Earth https://github.com/bacek/chgeos/blob/main/BENCHMARK.md ;) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Adds GeoParquet spatial pruning at row group and page levels to the Parquet reader. Row group pruning skips entire row groups whose bounding box doesn't overlap the query geometry. Page-level pruning uses the `covering.bbox` column index to skip irrelevant pages within row groups. Also generalizes spatial predicate pushdown through `IFunctionBase::isSpatialPredicate()` and adds a `GeoFilter` that evaluates spatial predicates during Parquet row reading. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes Parquet read-time pruning logic by adding spatial predicate extraction and bbox-based skipping at row-group and page granularity, which can impact query correctness if predicate/bbox handling is wrong. Also adjusts Iceberg manifest min/max serialization/pruning behavior and adds new settings/events that need validation across varied Parquet/GeoParquet metadata. > > **Overview** > Adds **GeoParquet spatial filter pushdown** to the Parquet reader, gated by new `input_format_parquet_spatial_filter_push_down`, to skip row groups (and optionally pages) when WHERE contains *conjunctive-only* spatial predicates whose constant-geometry bbox is disjoint from bbox statistics (via `covering.bbox` columns or `geospatial_statistics.bbox`). This introduces a new `Parquet::GeoFilter` extractor/bbox utilities, injects covering bbox columns for stats-only evaluation, wires spatial `KeyCondition`s into row-group/page pruning, and tracks pruned pages via new `ProfileEvents::ParquetPrunedPages`. > > Generalizes spatial predicate identification by adding `IFunction*::isSpatialPredicate()` (propagated through adaptors), marking built-in spatial predicates and enabling WebAssembly UDFs to opt in via `is_spatial_predicate` setting; also improves GeoParquet metadata parsing to capture `covering.bbox` column paths and exposes this flag in `FunctionNode` dumps. > > Extends Iceberg min/max handling to support Float32/Float64 stats correctly and to *filter out only non-serializable columns* instead of disabling bounds entirely, and adds Iceberg manifest pruning based on bbox column bounds for spatial predicates. New integration/stateless tests cover spatial row-group/page pruning, OR-safety, and Float32 stats round-tripping. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 135fb227c7e9a0b9b671b483ae561d0c83e31365. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104435",
        "createdAt": "2026-05-08T22:45:37Z",
        "updatedAt": "2026-08-13T11:24:21Z",
        "timestamp": "2026-08-13T11:24:21Z",
        "metrics": {
          "reactions": 1,
          "comments": 11
        },
        "labels": [
          "pr-performance",
          "can be tested",
          "pr-synced-to-cloud",
          "comp-datalake"
        ],
        "author": "bacek",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104437",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add spatial_bbox skip index for MergeTree geometry columns",
        "text": "Part of making ClickHouse fastest spatial analytical engine on Earth https://github.com/bacek/chgeos/blob/main/BENCHMARK.md ;) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Adds `spatial_bbox` skip index for MergeTree geometry columns. The index stores a bounding box per granule and skips granules whose geometry cannot intersect the query geometry, reducing work for spatial predicates. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104437",
        "createdAt": "2026-05-08T22:48:54Z",
        "updatedAt": "2026-08-13T16:55:01Z",
        "timestamp": "2026-08-13T16:55:01Z",
        "metrics": {
          "reactions": 1,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "bacek",
        "state": "open",
        "assignees": [
          "nihalzp"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104489",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support PromQL set binary operators",
        "text": "Adds native lowering for PromQL set binary operators over ClickHouse `TimeSeries` data. The implementation supports `and`, `or`, and `unless` with PromQL vector matching rules over packed vector grids. Matching is evaluated per timestamp so a series present on one side at one step does not accidentally match every step. Split out from the integration draft #104271; shared base #104484. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added native PromQL-to-SQL lowering for set binary operators `and`, `or`, and `unless` over ClickHouse `TimeSeries` data. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104489",
        "timestamp": "2026-08-12T21:48:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "can be tested",
          "pr-experimental",
          "comp-promql"
        ],
        "author": "BadLiveware",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104591",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add optimize_row_order_if_no_order_by (reopen #103919)",
        "text": "Reopen of https://github.com/ClickHouse/ClickHouse/pull/103919. This closes #103839. Adds a new `MergeTree` setting `optimize_row_order_if_no_order_by` (default `1`) that enables `optimize_row_order` automatically for tables without an explicit `ORDER BY`. **Behavior change and migration path.** With an empty sorting key (`ORDER BY ()` / `ORDER BY tuple()`) no query can rely on the physical row order, so the rows of every inserted block are reordered to improve compressibility. This makes such inserts slower (a 5M-row insert benchmark shows roughly `+180%`..`+230%` on the insert itself) in exchange for a smaller on-disk size and faster filters on low-cardinality columns; the CI `clickbench` and `tpch_adapted` runs report no significant change. Existing tables are affected on upgrade. To keep the old behavior: - per table: `SETTINGS optimize_row_order_if_no_order_by = 0` (or an explicit `optimize_row_order = 0`, which also opts out); - server-wide: set it in the `merge_tree` config section; - by version: it is recorded in the `MergeTree` settings changes history, so `compatibility` set to a version before `26.8` keeps it off. Note that `MergeTree` setting defaults are resolved once, when the server materializes its global `MergeTreeSettings`, so `compatibility` has to come from the default profile (`users.xml`) - a `SET compatibility` in an already-running session does not change them. The `MergeTree` documentation is updated accordingly. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `optimize_row_order_if_no_order_by` `MergeTree` setting. When enabled (default), row order optimization is applied to inserts into tables without an explicit `ORDER BY` clause. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes default insert-time behavior for `MergeTree` tables with `ORDER BY ()`, which can affect CPU cost and on-disk row ordering/compression. Also alters row-order optimization internals to swallow some `NOT_IMPLEMENTED` errors, which could mask type-specific issues if incorrect. > > **Overview** > Introduces a new `MergeTree` setting `optimize_row_order_if_no_order_by` (default **on**) and wires it into `MergeTreeDataWriter` so row-order optimization is automatically applied on insert for tables that *lack a sorting key* (`ORDER BY ()`), while preserving existing behavior for tables with an explicit `ORDER BY`. > > Hardens `RowOrderOptimizer` by catching `NOT_IMPLEMENTED` during cardinality estimation and falling back to an all-distinct upper bound instead of failing inserts, and records the setting in settings-change history. > > Updates many stateless/perf tests to explicitly disable the new default (`SETTINGS optimize_row_order_if_no_order_by = 0`) to keep deterministic baselines, adds new coverage for the setting’s default/override behavior and the regression on unsupported nested types, and extends `tests/performance/scripts/perf.py` to strip this setting from `CREATE TABLE` when running against older servers that don’t recognize it. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 02d68c628d9a3625824d33b91f87bd12d6fd3938. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104591",
        "createdAt": "2026-05-11T14:01:19Z",
        "updatedAt": "2026-08-13T15:33:39Z",
        "timestamp": "2026-08-13T15:33:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 44
        },
        "labels": [
          "pr-improvement",
          "pr-performance",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104691",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Lazy-load column statistics during query planning",
        "text": "When the query planner needs column statistics — for join reordering, for prewhere selectivity estimation, or for part pruning — it currently loads statistics for every column of the table from disk on the first access, even when the query only filters or joins on a handful of columns. This PR reduces statistics-file I/O during query planning on wide tables. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Column statistics are now loaded on demand for only the columns the query planner needs instead of for every column of the table on first access.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104691",
        "createdAt": "2026-05-12T11:13:36Z",
        "updatedAt": "2026-08-13T09:29:44Z",
        "timestamp": "2026-08-13T09:29:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 15
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "zoomxi",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104809",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `NOT_FOUND_COLUMN_IN_BLOCK` in `query_plan_convert_join_to_in` with `arrayJoin` JOIN key",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `NOT_FOUND_COLUMN_IN_BLOCK` exception in `INNER JOIN ... ON arrayJoin(...) = ...` queries when `query_plan_convert_join_to_in` is enabled and the SELECT references the source array column. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- Not applicable -- bug fix, no new feature or API change. --> --- ### Problem With `SET query_plan_convert_join_to_in = 1` (off by default), the optimizer rewrites a hash `INNER JOIN ON arrayJoin(L.col) = R.col` into `[Expression(\"Calculate join left keys\"), Filter(\"IN\"), DelayedCreatingSets, Expression(\"Join output actions\")]`. If the SELECT projection (or any post-JOIN expression) references the source array column `L.col` itself, or `arrayJoin(L.col)` again, the rewritten plan throws ``` Code: 10. DB::Exception: Not found column __table1.tags in block ... (NOT_FOUND_COLUMN_IN_BLOCK) ``` at execution (and at `EXPLAIN actions = 1` planning time). Minimal repro: ```sql CREATE TABLE lt (id UInt64, tags Array(String)) ENGINE = MergeTree ORDER BY id; CREATE TABLE rt (tag_id String) ENGINE = MergeTree ORDER BY tag_id; INSERT INTO lt VALUES (1, ['a','b','c']), (2, ['d','e']); INSERT INTO rt VALUES ('a'), ('d'); -- Throws NOT_FOUND_COLUMN_IN_BLOCK: SELECT lt.id, lt.tags FROM lt INNER JOIN rt ON arrayJoin(lt.tags) = rt.tag_id SETTINGS query_plan_convert_join_to_in = 1; -- Same with arrayJoin in the projection -- also throws: SELECT lt.id, arrayJoin(lt.tags) FROM lt INNER JOIN rt ON arrayJoin(lt.tags) = rt.tag_id SETTINGS query_plan_convert_join_to_in = 1; ``` ### Root cause `tryConvertJoinToIn` builds `left_pre_join_actions = JoinExpressionActions::getSubDAG(<left join keys>)`. The output set is only the JOIN-key expressions — `arrayJoin(__table1.tags)` and `__table1.id` in the repro — **not** the source `__table1.tags` array column. Later, `cloneSubDAGWithHeader(output_header, JoinExpressionActions::getSubDAG(join_output_actions))` clones the post-JOIN expressions onto a new DAG seeded from `output_header`. `cloneSubDAGWithHeader` calls `mergeInplace(..., remove_dangling_inputs=true)`, which remaps `INPUT` nodes by name to the corresponding inputs in `output_header`. When `join_output_actions` references `__table1.tags` (because SELECT did), the source `__table1.tags` INPUT in `second_dag` has no match in `output_header` — `mergeInplace`'s `remove_dangling_inputs` only removes input nodes that collide with `first`'s inputs, not input nodes that have no counterpart at all. The non-INPUT nodes that depend on it (the cloned ARRAY_JOIN node when SELECT had `arrayJoin(lt.tags)`, or the column reference itself when SELECT had `lt.tags`) survive, and the resulting `ExpressionStep` cannot find the column at execution. ### Fix Decline the conversion in `tryConvertJoinToIn` whenever any input of `join_output_actions` is not among the outputs that `left_pre_join_actions` forwards. The check runs immediately after `left_pre_join_actions` / `right_pre_join_actions` are computed and **before** any plan mutation (`makeExpressionNodeOnTopOf` etc.), so the bail-out is safe — the query then runs through the normal JOIN path and produces the correct result. ```cpp { auto join_output_actions_subdag = JoinExpressionActions::getSubDAG(join_output_actions); std::unordered_set<std::string_view> forwarded_columns; for (const auto * out : left_pre_join_actions.getOutputs()) forwarded_columns.insert(out->result_name); for (const auto * input : join_output_actions_subdag.getInputs()) if (!forwarded_columns.contains(input->result_name)) return 0; } ``` The check mirrors the established `appendInputsForUnusedColumns` style (`ActionsDAG.cpp:1453-1461`), which is the idiomatic \"verify every input is present in a sample block\" pattern in this codebase. ### Related fixes for the same root cause This is the third bug in 12 months rooted in `mergeInplace` / `clone` of an ARRAY_JOIN-bearing DAG being merged into a context whose stream header does not provide a required INPUT column. The prior two are: - #96989 (prevention strategy: blacklist `ARRAY_JOIN` in the pushdown eligibility check). - #97239 (bail-out strategy: detect ghost `ARRAY_JOIN` after the merge and abandon the optimization). - #104785 (post-hoc cleanup strategy: snapshot the merged-in `ARRAY_JOIN` pointers and drop them after the merge via a new `ActionsDAG::removeNodes` helper). This PR follows the bail-out pattern of #97239 — `tryConvertJoinToIn` is an optional optimization, so declining the rewrite is safe and the query falls back to the regular JOIN path. A more invasive follow-up could forward the missing left-side columns through `left_pre_join_actions` so the optimization stays on even in these cases; left for a separate PR once we have evidence that real workloads hit the bail-out frequently. The setting `query_plan_convert_join_to_in` is off by default, so the bug is reachable only when explicitly enabled. ### Regression test `tests/queries/0_stateless/03918_convert_join_to_in_arrayjoin_dangling.{sql,reference}` — co-located with #96989's `03918_arrayjoin_function_with_join_and_where`. Four queries: 1. `SELECT lt.id ... INNER JOIN ... ON arrayJoin(lt.tags) = rt.tag_id` — the optimization stays on, returns the correct rows (validates the fix's \"no-regression\" guarantee for the cases it does not bail out on). 2. `SELECT lt.id, lt.tags ...` — before fix: NOT_FOUND_COLUMN_IN_BLOCK; after fix: correct result via the fallback JOIN path. 3. `SELECT lt.id, arrayJoin(lt.tags) ...` — same shape as (2). 4. Control: same as (3) with `query_plan_convert_join_to_in = 0`. Confirms the fallback path returns the same answer the fix produces in (3). Repro-first verified: queries (2) and (3) were observed to fail with NOT_FOUND_COLUMN_IN_BLOCK on un-fixed upstream master (`7bd0fa28146`) before the reference was locked.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104809",
        "createdAt": "2026-05-13T09:09:21Z",
        "updatedAt": "2026-08-13T09:33:28Z",
        "timestamp": "2026-08-13T09:33:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "vdimir"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104948",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Declarative function signatures, continuation of #3775",
        "text": "This is an experimental continuation of [#3775 (2018)](https://github.com/ClickHouse/ClickHouse/pull/3775) by @alexey-milovidov, which proposed a declarative way to describe function signatures so that argument validation and return-type inference can be driven by a small DSL instead of hand-written `getReturnTypeImpl` logic per function. ### What's in this branch A `getSignatureString()` method on `IFunction` and `IFunctionOverloadResolver`. When a function exposes a non-empty signature, the base `getReturnTypeImpl(ColumnsWithTypeAndName)` parses and applies it via a small grammar of type **matchers** (`UInt`, `Number`, `Array(T)`, `MaybeNullable(T)`, `Function((args), R)`, `T : Any` capture, …) and type **functions** (`leastSupertype`, `nativeNumber`, `aggregateFunctionReturnType(AggregateFunction(name, …))`, `subcolumnTypeOf`, `typeFromString`, `DateTime64(scale, tz)`, …). The DSL supports variadic positions (`…`), ellipsis grouping for repeated argument units (`T1, V1, …` repeats the pair), `OR` between alternatives, optional positions `[T]`, const-value capture (`const name String`), lambdas, etc. The grammar / parser / type matchers / type functions live under `src/DataTypes/FunctionSignature.h`, `FunctionSignature.cpp`, `TypeMatchers.cpp`, `TypeFunctions.cpp`. ### Coverage ~191 commits on top of master, each is a small per-family round titled `Function signatures: round N — …`. Current state on this branch: ``` SELECT count(*) AS total, countIf(signature != '') AS with_sig, round(countIf(signature != '') * 100.0 / count(*), 1) AS pct FROM system.functions WHERE NOT is_aggregate AND alias_to = '' AND origin = 'System'; 1416 1395 98.5 ``` (With \\`allow_experimental_nlp_functions = 1\\`; the few remaining unset ones are setting / config-gated functions like \\`aiClassify\\`, \\`region*\\`, \\`synonyms\\`, whose \\`create\\` throws unless the relevant config block / setting is present, so \\`system.functions.tryGet\\` returns null even though the signature is in source. With the right config + setting they all surface — verified in tmp/ test config.) Roughly: - **Authoritative** (DSL drives type-check and return type, the legacy `getReturnTypeImpl` is either gone or bypassed): higher-order array functions (`arrayMap`, `arrayFilter`, `arrayFirst*`, …), `toIntervalX`, comparison, `multiSearch*` / `multiMatch*`, vector L-norms / distances / dot product, `UUIDv7ToDateTime`, `arrayReduce` / `arrayReduceInRanges`, `reverse`, `mapKeys` / `mapValues`, `getSubcolumn`, the `least` / `greatest` resolver, paired-variadic `timeSeriesTagsToGroup` / `timeSeriesStoreTags`, … - **Documentation-only** (signature surfaced via `system.functions` but the legacy `getReturnTypeImpl` still runs because the result type uses promotion / widening / setting-dependent dispatch the current DSL can't express): arithmetic (`plus`, `minus`, `multiply`, `divide`, modulo / intDiv family), array widening (`arraySum` / `arrayCumSum*` / `arrayDifference`), `mapContains*Like`, `dateTrunc`, `toStartOfWeek`, `parseDateTime*`, `transform`, `range`, `toX` conversions, etc. There is a `signature_documentation` opt-in alongside `signature` in the binary-arithmetic, unary-arithmetic, FunctionArrayMapped, and FunctionMapToArrayAdapter families specifically so a function can advertise a signature without the DSL accidentally hijacking the legacy widening logic. ### Status Draft / experimental. - This is an experiment — I'm not asking for it to be merged. There's significant disruption (191 commits touching every function family), and the gain is mostly documentation: today only ~30% of the converted functions are actually DSL-authoritative; the rest are decorative. The arithmetic promotion matrix and the per-Op widening rules in particular would need a richer type-function vocabulary (or new matchers) before they can be expressed declaratively. - The branch has been kept rebased on master throughout and builds cleanly. I've run the stateless test suite on each round. The pre-existing-on-master test failures I hit are listed in commit messages of the rounds where I encountered them (none introduced by this work). - The original PR has been open since 2018 — this branch is intended as a concrete data point on \\\"what fraction of ClickHouse functions can be reasonably described by a declarative signature, and what would the DSL need to grow to cover the rest.\\\" That's the question I'd love feedback on. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104948",
        "createdAt": "2026-05-14T14:35:16Z",
        "updatedAt": "2026-08-13T14:11:52Z",
        "timestamp": "2026-08-13T14:11:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 38
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104965",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add server setting `additional_memory_tracking_per_thread`",
        "text": "Each thread accumulates up to `max_untracked_memory` (4 MiB by default) of allocations before reporting them to the server-wide `MemoryTracker`. With many threads, this unreported memory can sum to a large amount, causing the server's tracked memory usage to under-count actual consumption and leading to OOM. This PR introduces a server-level setting `additional_memory_tracking_per_thread` (default 4 MiB) and speculatively charges this amount to the server-wide `MemoryTracker` around every job executed in our `ThreadPool` workers. The global tracked memory becomes a safe upper bound on actual consumption. The reservation is charged on the server-wide (total) tracker only — deliberately not through the query's tracker chain. Query-level accounting feeds heuristics that compare memory deltas against byte thresholds (conversion of aggregation hash tables to two-level via `group_by_two_level_threshold_bytes`, spill-to-disk decisions), and phantom reservations of `num_threads * 4 MiB` trip those thresholds immediately: an earlier revision of this PR that charged the query tracker showed consistent slowdowns of GROUP BY queries in performance tests for exactly this reason. With the server-wide-only reservation, query-level and user-level accounting (`max_memory_usage`, `memory_usage` in `system.processes`) are unaffected. The speculative reservation uses the throwing path of the memory tracker, so when it would exceed the server memory limit the corresponding job is treated as failed with `MEMORY_LIMIT_EXCEEDED` — the same behavior as if the job itself had exceeded the limit. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new server setting `additional_memory_tracking_per_thread` (default 4 MiB) which speculatively reserves this amount on the server-wide memory tracker around every `ThreadPool` job. It compensates for the up to `max_untracked_memory` of un-reported allocations per thread, making the server's tracked memory a safe upper bound on actual consumption and reducing the risk of OOM with many concurrent threads. Query-level and user-level memory accounting are not affected. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASan - [ ] <!---ci_include_tsan--> All with TSan - [ ] <!---ci_include_msan--> All with MSan - [ ] <!---ci_include_ubsan--> All with UBSan - [ ] <!---ci_include_coverage--> All with Coverage - [ ] <!---ci_include_aarch64--> All with Aarch64 #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASan - [ ] <!---ci_exclude_tsan--> All with TSan - [ ] <!---ci_exclude_msan--> All with MSan - [ ] <!---ci_exclude_ubsan--> All with UBSan - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 #### Extra options: - [ ] <!---ci_set_arm--> Add tests with aarch64 builds - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104965",
        "createdAt": "2026-05-14T16:39:55Z",
        "updatedAt": "2026-08-13T15:50:37Z",
        "timestamp": "2026-08-13T15:50:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 37
        },
        "labels": [
          "pr-improvement",
          "memory"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "azat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:104993",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Handle non-constant RHS for `IN`",
        "text": "Fix `IN` and `NOT IN` expressions with non-constant right-hand side operands that reference columns from the current row. # old analyzer Previously, the old analyzer tried to build a standalone `Set` for expressions such as `number % 2 IN (number % 3, number % 5)`, which made the right-hand side unable to resolve `number` and produced an `UNKNOWN_IDENTIFIER` exception. This change rewrites such expressions to row-wise `has` expressions instead. ```sql --- error was reported for the query below for old analyzer as mentioned in #58242 SET enable_analyzer = 0; SELECT number FROM numbers(10) WHERE number % 2 IN (number % 3, number % 5) ORDER BY number; ``` ### Out of scope: bare source-column RHS under the old analyzer Under the old analyzer (`enable_analyzer = 0`), a bare column of the `FROM` source as the right-hand side, such as `x IN (arr)` where `arr` is a column of the current row, still fails with `UNKNOWN_TABLE`: `MarkTableIdentifiersVisitor` rewrites `x IN ident` into `x IN (SELECT * FROM ident)` before source columns are collected, so the expression never reaches the row-wise rewrite. The new analyzer resolves the same query as a column and succeeds; tests `04234` and `04812` pin this divergence explicitly. Closing it needs either reordering that visitor after source columns are known, or falling back from a table to a column when the table does not exist, plus a compatibility decision for `x IN t` when a column shadows an existing table name - that is tracked in the review discussion and left out of this PR on purpose. # new analyzer The new analyzer already handled the basic non-constant right-hand side case, but some tuple and NULL cases still failed. This change fixes: * tuple-typed right-hand side expressions produced by functions other than tuple * tuple left-hand side membership checks that previously tried to create `Nullable(Tuple(...))` * `NULL` operands in non-constant tuple right-hand side operands, where the old cast target could become `Nullable(Nothing)` Examples: ```sql --- Before this fix, the new analyzer treated the tuple-typed if RHS as a single tuple value and failed with a type error; after this fix, it expands the tuple value one level for scalar IN, so the query returns 1 SELECT number IN (if(number >= 0, tuple(number, number + 1), tuple(0, 0))) FROM numbers(1); ``` ```sql --- tuple in tuple, user exepects some rows to match, however, error like `Cannot create column with type 'Nullable(Tuple(UInt8, UInt8))' because Nullable Tuple type is not allowed` will be reported before this fix SELECT number, (1, 1) IN ((number % 3, number % 2), (2, 2)) FROM numbers(6) ORDER BY number; ``` ```sql --- user expects `NULL` but error like `Conversion from UInt8 to Nothing is not supported` will be reported before this fix SELECT x IN (y, 1) FROM ( SELECT materialize(NULL) AS x, materialize(2) AS y ); ``` Issue: https://github.com/ClickHouse/ClickHouse/issues/58242 ### Changelog category (leave one): - Bug Fix ### Changelog entry: - Fix IN and NOT IN expressions with non-constant right-hand side operands referencing columns from the current row, and align new analyzer tuple right-hand side handling with existing ClickHouse IN semantics. Under the old analyzer, a bare source column as the right-hand side (`x IN (arr)`) still resolves as a table name and stays out of scope. This closes [#58242](https://github.com/ClickHouse/ClickHouse/issues/58242)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/104993",
        "createdAt": "2026-05-15T04:11:24Z",
        "updatedAt": "2026-08-13T17:00:52Z",
        "timestamp": "2026-08-13T17:00:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 27
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "niyue",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105045",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support column matcher expansion for default value expressions and index expressions",
        "text": "This PR closes https://github.com/ClickHouse/ClickHouse/issues/92266 Support column matchers in column `DEFAULT`, `ALIAS`, `MATERIALIZED`, and `EPHEMERAL` expressions, and in data skipping index expressions. This allows expressions such as `*`, `COLUMNS('...')`, `COLUMNS(a, b)`, `EXCEPT`, `APPLY`, and `REPLACE` to be expanded before expression validation and execution. ~~The change also adds `namedTuple` function to make matcher-expanded named tuple expressions ergonomic in tests and user queries.~~ The tests cover these use cases, direct and indirect cyclic default-expression dependency detection, and nested matcher expansion. ### Changelog category: - New Feature ### Changelog entry: - Support column matchers such as `*` and `COLUMNS` in column default value expressions, `DEFAULT`, `ALIAS`, `MATERIALIZED`, and `EPHEMERAL` expressions, and in data skipping index expressions. ### Note ~~About the newly added `namedTuple` function: I read the previous discussions in [1], [2], and [3], and my impression is that the existing `enable_named_columns_in_function_tuple` setting is not very discoverable. A separate function name may make the intention clearer and the feature easier to use, especially in this use case. It also avoids changing the behavior of the existing tuple function, so it should not introduce compatibility issues. Please let me know if this direction is not desirable.~~ [1] https://github.com/ClickHouse/ClickHouse/issues/63524 [2] https://github.com/ClickHouse/ClickHouse/issues/54921 [3] https://github.com/ClickHouse/ClickHouse/pull/54881 <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches core expression/DDL validation paths (defaults, aliases, ALTER, mutations, skip indexes), which can affect query planning and error behavior if matcher expansion or cycle detection is incorrect. > > **Overview** > Enables column matchers (e.g. `*`, `COLUMNS(...)` plus `EXCEPT`/`APPLY`/`REPLACE`) inside column `DEFAULT`/`MATERIALIZED`/`ALIAS`/`EPHEMERAL` expressions by expanding matchers before validation/execution, honoring `asterisk_include_*` settings and rejecting qualified matchers. > > Updates DDL/default validation, alias expansion, read-order optimization, merge/SELECT paths, and mutation/materialize flows to use a shared `cloneAndExpandColumnDefaultExpression` helper and adds cycle detection that accounts for matcher-expanded dependencies. > > Extends skip index parsing/analysis to normalize matcher and alias usage (including cyclic-alias detection) before `TreeRewriter` analysis, and adds docs + comprehensive stateless tests covering matcher expansion, errors, and ALTER/mutation/index scenarios. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 48ab80fd074599c63f967f624435e0a2e7f166ed. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105045",
        "createdAt": "2026-05-15T15:10:34Z",
        "updatedAt": "2026-08-13T12:39:00Z",
        "timestamp": "2026-08-13T12:39:00Z",
        "metrics": {
          "reactions": 3,
          "comments": 49
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested"
        ],
        "author": "niyue",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105249",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Accessing tables as files, query construction and out-of-band modification in HTTP interface",
        "text": "Implements [#46925](https://github.com/ClickHouse/ClickHouse/issues/46925) according to the [updated spec](https://github.com/ClickHouse/ClickHouse/issues/46925#issuecomment-4475417259). Closes: https://github.com/ClickHouse/ClickHouse/issues/46925 Related: https://github.com/ClickHouse/clickhouse-docs/pull/6398 ## What's in the box - **Path → database/table/format/compression.** New settings `http_allow_database_as_path`, `http_allow_table_as_file`, `http_allow_filters_as_path`, `http_allow_filters_as_unrecognized_url_parameters` let the HTTP interface interpret `/database/table.format.compression` (and hive-style `/name=value/` partitions) in the URL path. When a path identifies a table, the request is processed as `SELECT * FROM database.table`. - **Query construction settings.** `select`, `order`, `sort`, `filter`, `page` wrap the base query as `SELECT [select] FROM (...) [WHERE filter] [ORDER BY order]`. `page=N` translates to `offset = limit * (N - 1)` (errors if `limit` is unset or `offset` is also set). Multiple `filter` URL parameters are combined with `AND`. - **Format/compression overrides.** `default_format`, `format`, `input_format`, `output_format` are now first-class settings; the explicit `format` / `output_format` overrides win over the FORMAT clause in the query and the file extension in the path, while `default_format` is only the fallback used when nothing else selects a format (so the path extension wins over it). A generic `compression` setting wraps the response body (independent of HTTP `Content-Encoding`). The URL path file extension is equivalent; specifying both with conflicting values throws. - **Implicit table.** When a path table coexists with a user `query`, the path table is exposed via `implicit_table_at_top_level`, so `?query=SELECT a, b` against `/hits.csv` reads from `hits` with no `FROM`. If the user's query already has `FROM`, the path component only seeds the download filename. - **Content-Disposition.** Binary or compressed responses get `Content-Disposition: attachment; filename=…`, where the filename comes from the URL path component (or `result.<format>.<compression>` as fallback). - **`database` and the format settings are now real settings.** The `database` URL parameter and the `X-ClickHouse-Database` / `X-ClickHouse-Format` headers all flow through the regular settings pipeline (the header still overrides the URL parameter to preserve historical precedence). `X-ClickHouse-Database` is an alias for `database`; `X-ClickHouse-Format` is an alias for `output_format` (not `default_format`), i.e. an explicit override of the response format that also wins over a `FORMAT` clause in the query. It maps to `output_format` rather than to the bidirectional `format` because the header has always described the response only, so it must not reinterpret the body of an `INSERT`. For the same reason, `--format` in `clickhouse-client` maps to `output_format` (it has always been output-only there), while in `clickhouse-local` it keeps its historical bidirectional meaning and maps to `format`. They are marked `changeable_in_readonly` in the shipped default profile so they remain settable on HTTP GET (which forces `readonly = 2`) and for `readonly = 1` users. Because `database` is a real setting, a profile that constrains it (e.g. marks it readonly or restricts its values) is enforced consistently on every way of choosing the current database: `USE`, `SET database = ...`, the HTTP `database` parameter/header, and the connect-time database of the native TCP, MySQL, and PostgreSQL protocols. In `clickhouse-local`, an explicitly configured database (`--database` or a config-file `database` key) is mirrored into the `database` setting so a profile-inherited value cannot override it. `HTTPHandler::processQuery` is restructured so that authentication and the user profile are applied first, then the auth-related parameter (`role`), then the read-only enforcement, and only then the general settings. ## Changelog category - New Feature ## Changelog entry Added a way to access tables and databases via URL paths in the HTTP interface (e.g. \\`/database/table.format.gz?filter=a>0\\`), plus new settings (\\`http_allow_database_as_path\\`, \\`http_allow_table_as_file\\`, \\`http_allow_filters_as_path\\`, \\`http_allow_filters_as_unrecognized_url_parameters\\`, \\`select\\`, \\`order\\`, \\`sort\\`, \\`filter\\`, \\`page\\`, \\`compression\\`, \\`format\\`, \\`input_format\\`, \\`output_format\\`, \\`default_format\\`, \\`database\\`) that compose with one another and with the existing \\`query\\` URL parameter. ## Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) — [ClickHouse/clickhouse-docs#6398](https://github.com/ClickHouse/clickhouse-docs/pull/6398) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > High risk because it significantly changes HTTP request routing/parsing and introduces many new per-query settings that affect query construction, output formatting, compression, and readonly constraints, which could impact security and compatibility. > > **Overview** > **Enables “table-as-file” and query shaping via the HTTP interface.** The server can now interpret URL paths like `/db/table.format[.compression]` (plus optional hive-style `name=value` path filters) to auto-generate `SELECT * FROM ...`, apply `select`/`filter`/`order`/`sort` wrappers, and support paging via new `page` setting (translated into `limit`/`offset`). > > **Promotes HTTP-specific knobs into first-class settings and aligns client behavior.** Adds new settings for `database`, `default_format`, `format`, `input_format`, `output_format`, `compression`, and HTTP path feature gates; updates HTTP handler to route headers/URL params through the normal settings/constraints pipeline (including readonly carve-outs), adds generic response-body compression and `Content-Disposition` for binary/compressed outputs, and updates client/local tooling to mirror config/CLI format/database options into per-query settings. > > **Adjusts core settings/exec behavior for compatibility.** Widens `limit`/`offset` to `Double` (supporting negative/fractional values), updates analyzers/interpreters/cluster proxy to handle/strip initiator-only settings in distributed queries, refines handler factory ordering/prefix mounting, and adds/updates stateless tests and configs for the new HTTP path semantics and hints. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 8efcce89edbde35344024d1f9306be66470f3e1a. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105249",
        "createdAt": "2026-05-18T16:22:57Z",
        "updatedAt": "2026-08-13T06:53:49Z",
        "timestamp": "2026-08-13T06:53:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 83
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105429",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Release pull request for branch 26.5",
        "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches security-sensitive internal/DDL execution gating by replacing `query_kind` checks with a new server-set flag, which could affect ON CLUSTER/replication/backup behavior if mis-propagated. Also tightens Arrow/Native input validation, which may reject previously-accepted malformed inputs and impact ingestion edge cases. > > **Overview** > Introduces a new `Context` flag `is_ddl_or_on_cluster_internal` (server-set and non-spoofable) and switches multiple code paths from `ClientInfo::QueryKind::SECONDARY_QUERY` to this flag for *security-sensitive* decisions, including internal backup/restore gating, Replicated DB DDL handling, `ON CLUSTER`/UUID-macro allowances, and context creation in DDL/replication/system operations. > > Hardens Arrow ingestion by reading geo metadata from the Arrow schema, validating BinaryArray offsets/lengths against buffer bounds (and handling absent/empty buffers), and avoiding `mutable_data()` usage in geo parsing; adds regression tests for corrupted Arrow offsets and geo metadata. > > Improves error classification for Variant deserialization under Native format (reports `INCORRECT_DATA` instead of `LOGICAL_ERROR`), adds an integration test for malformed Native Variant payloads, and updates planner behavior to disable/forbid parallel replicas when `additional_table_filters` is used without `serialize_query_plan` (with new stateless tests). > > Minor: CI style job skips test-number gap checks on release/backport branches, MySQL protocol test fixes dotnet working dir, and release metadata (version/contributors) is updated. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 1a682e2727943363cb077a81015b9242ab666852. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105429",
        "createdAt": "2026-05-20T13:36:34Z",
        "updatedAt": "2026-08-13T17:59:14Z",
        "timestamp": "2026-08-13T17:59:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "release"
        ],
        "author": "robot-clickhouse",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105499",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Drop for detached tables",
        "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental `DROP DETACHED TABLE` support, gated by `allow_experimental_drop_detached_table`, to remove metadata and data for detached tables. Continues #62490",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105499",
        "createdAt": "2026-05-21T09:59:53Z",
        "updatedAt": "2026-08-13T14:42:08Z",
        "timestamp": "2026-08-13T14:42:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "manual approve",
          "can be tested",
          "pr-experimental"
        ],
        "author": "UberDever",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105710",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject divergent same-named parts in parallel replicas coordinator",
        "text": "The parallel replicas coordinator deduplicated parts purely by `MergeTreePartInfo` (name + version), so when two replicas announced a same-named part holding genuinely different data the second announcement was silently merged into the first. The coordinator then dispatched ranges from the first replica's snapshot to the second replica, whose local part was different, producing an exception `Trying to get non existing mark N, while size is M` inside `MergeTreeIndexGranularityConstant::getMarkRows` — or, when the layouts happened to line up, potentially incorrect results. This is reachable with `parallel_replicas_for_non_replicated_merge_tree = 1` against a cluster whose members each have independent local `MergeTree` data: block numbers of a non-replicated `MergeTree` come from a node-local `SimpleIncrement`, so each member's local first part is named `all_1_1_0` but can store different data. The fix makes the coordinator validate the identity of same-named parts across announcements: - `RangesInDataPartDescription` now carries the underlying part's total mark count, a content fingerprint (`checksums.getTotalChecksumUInt128`), and a `part_name_identity` tri-state saying whether a part name identifies the same content on every cluster member. The fields are gated on a new parallel replicas protocol version (`DBMS_PARALLEL_REPLICAS_PROTOCOL_VERSION` bumped to 10). - `part_name_identity` is `ClusterWide` when the engine coordinates block numbers through Keeper (`ReplicatedMergeTree` and descendants) *or* when all of the table's disks keep their metadata in storage shared by every cluster member (`MetadataStorageType::Plain`, `PlainRewritable`, `StaticWeb`, `WebIndex`, `Keeper`) — there every member enumerates literally the same parts. It is `NodeLocal` for a plain `MergeTree` on ordinary local or per-node remote disks. - `InOrderCoordinator` and `DefaultCoordinator` snapshot these fields from the first announcement of each part and compare later announcements against the snapshot. A fingerprint mismatch raises `BAD_ARGUMENTS` naming the diverging part and pointing the user at `ReplicatedMergeTree` or at disabling `parallel_replicas_for_non_replicated_merge_tree`. - When the fingerprint is unavailable on either side (a replica running an older server version, or a part whose checksums are not loaded) and either side reports `NodeLocal` part names, the coordinator fails closed with `BAD_ARGUMENTS` instead of weakening the identity check — such same-named parts cannot be verified, and merging them blindly could return incorrect results. For `ClusterWide` part names the part name implies identical content, so a mark-count fallback keeps mixed-version clusters working during rolling upgrades. Only analyzed-view fields (`rows`, `ranges`) are deliberately NOT compared: per-replica PK and skip-index analysis can legitimately select different mark subsets from the same underlying part, and `ranges` is consumed in place as the coordinator dispatches work. The protocol bump also required pinning one existing consumer of the serializer. The `read_bucket` task parameter of a distributed read plan (`make_distributed_plan`) embeds a `RangesInDataPartsDescription` blob but travels as an opaque query parameter with no handshake to negotiate a version on, so producer and consumer both used the build's own `DBMS_PARALLEL_REPLICAS_PROTOCOL_VERSION` and would disagree about the field layout across a rolling upgrade. The blob is now pinned to `DBMS_PARALLEL_REPLICAS_DISTRIBUTED_READ_BUCKET_VERSION = 8`, the layout in effect when it was introduced. Surfaced by the AST fuzzer (STID `4920-51f2`) on PR #105706. The 5-replica `parallel_replicas` cluster used by `tests/queries/0_stateless/02275_full_sort_join_long.sql.j2` exposes the divergence whenever local `t2` parts differ in size across replicas. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105706&sha=84790f83ec78aee26b08ab6e0bc6712f6e4f1745&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29 Related: https://github.com/ClickHouse/ClickHouse/pull/105706 Coverage is via `gtest_parallel_replicas_coordinator`, which exercises both coordination modes and both announcement orderings: rejection on divergent mark counts and divergent fingerprints, acceptance of identical announcements and of divergent analyzed views over the same part, the fail-closed path for `NodeLocal` part names without a fingerprint, the mark-count fallback for `ClusterWide` part names in mixed-version clusters, and a serializer round trip across protocol versions 8, 9 and 10 that requires the reader to consume the whole buffer. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an exception `Trying to get non existing mark N, while size is M` and potential incorrect parallel-replica range assignment when `parallel_replicas_for_non_replicated_merge_tree = 1` is used on divergent local `MergeTree` data: the parallel replicas coordinator now validates same-named parts by a content fingerprint of the underlying part instead of merging announcements blindly, and fails closed when the fingerprint is unavailable and part names are not guaranteed to identify the same content on every cluster member. ### Documentation entry for user-facing changes - [x] Documentation is unchanged",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105710",
        "createdAt": "2026-05-23T23:10:40Z",
        "updatedAt": "2026-08-13T17:29:05Z",
        "timestamp": "2026-08-13T17:29:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 28
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105714",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add INSERT ... RETURNING (non-atomic, user-supplied SELECT)",
        "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `INSERT ... RETURNING (SELECT ...)` to run a user-supplied SELECT after INSERT and return its result in one round-trip. Closes #21697. **Backward-incompatible change:** `RETURNING` is now a reserved keyword and can no longer be used as an unquoted implicit column alias (e.g. `SELECT 1 returning`). Use `AS returning` or a quoted identifier instead. --- Closes #21697 ## Summary Draft implementation of `INSERT ... RETURNING` for ClickHouse. The INSERT runs first using the existing pipeline; if it succeeds, a parenthesized `RETURNING (SELECT ...)` subquery runs in the same session and its result set is returned to the client in one round-trip. This is intentionally **not** Postgres-style atomic `RETURNING`. The user supplies the trailing SELECT and owns the `WHERE` clause used to identify inserted rows. ## Syntax For `INSERT VALUES` and `INSERT FORMAT`, `RETURNING` appears **before** the data clause (parser constraint: bytes after `VALUES`/`FORMAT` are raw input): ```sql INSERT INTO t (id, name) RETURNING (SELECT * FROM t WHERE id = 123) VALUES (123, 'foo'); INSERT INTO t RETURNING (SELECT count() FROM t WHERE batch_id = 'X') FORMAT JSONEachRow {\"batch_id\":\"X\", ...} ``` For `INSERT SELECT`, `RETURNING` appears **after** the source query: ```sql INSERT INTO t SELECT * FROM src RETURNING (SELECT count() FROM t WHERE batch_id = 'X'); ``` The parenthesized subquery is required (distinct from Postgres column-list `RETURNING id, name`). ## Semantics - INSERT runs first with its existing pipeline untouched. - If INSERT throws, the RETURNING SELECT does not run. - If INSERT succeeds, a select interpreter is built for the trailing subquery using the same Context; its pipeline is returned as the statement result. - The trailing SELECT inherits session, user, and settings. It can reference any table and use the full SELECT grammar. - Client receives **one result set** (the RETURNING SELECT's), not an INSERT ack plus a result set. ## Limitations (documented, not bugs) - **Not atomic.** Concurrent writers can land rows between INSERT completion and SELECT execution. - **No automatic \"the rows I just inserted.\"** User provides the SELECT and owns its `WHERE`. - **Replication lag is visible** on Replicated tables unless consistency settings are used. - **Incompatible with `async_insert=1`.** Returns `NOT_IMPLEMENTED`. - **No special MV interaction.** MVs fire on INSERT; RETURNING SELECT is a normal SELECT. - If INSERT succeeds but RETURNING SELECT fails, inserted data is **not rolled back**. ## Backward-incompatible change `RETURNING` is added to the parser's reserved-keyword list (same list as `FROM`, `WHERE`, `SETTINGS`, etc.) to prevent it from being consumed as an implicit table alias in `INSERT ... SELECT ... FROM tablename RETURNING (...)`. As a side-effect, `RETURNING` can no longer be used as an unquoted implicit alias anywhere in ClickHouse SQL: ```sql -- Before: worked (returning was a plain identifier) SELECT 1 returning SELECT * FROM t returning -- After: must use AS or quoting SELECT 1 AS returning SELECT * FROM t AS returning ``` ## Implementation - **Parser** (`ParserInsertQuery.cpp`): parse `RETURNING` + parenthesized SELECT; placement rules above. - **AST** (`ASTInsertQuery.h`): `returning_select` field; clone/format updated. - **Interpreter** (`InterpreterInsertQuery.cpp`): reject `async_insert`; wrap completed INSERT pipeline with RETURNING via `DelayedSource`. - **Pipeline** (`buildInsertReturningPipeline.cpp`): run INSERT to completion, then build SELECT pipeline; native push inserts swap to RETURNING SELECT after data is received. - **executeQuery** / **TCPHandler** / **LocalConnection** / **ClientBase**: inlined-data wrap, native-protocol push path, query-log accounting, process-list propagation. - **Docs**: section added to `insert-into.md` ## Tests added - `tests/queries/0_stateless/04266_insert_returning_21697.sql` — core `INSERT ... RETURNING` semantics, delayed planning/failure behavior, query-cache rejection paths, transactional rollback checks (`implicit_transaction=1` + explicit transaction), and parser keyword regression coverage (`SELECT 1 returning` rejected; explicit/quoted alias accepted). - `tests/queries/0_stateless/04267_insert_returning_native_format.sh` — native TCP `INSERT FORMAT` + external data + `RETURNING`, including `input(...)` + `FORMAT` placement/round-trip stability. - `tests/queries/0_stateless/04410_insert_returning_trailing_settings.sql` — source-side trailing `SETTINGS` scoping/precedence, nested source settings handling, source-side query-cache settings rejection (`NOT_IMPLEMENTED`), and a plain `INSERT ... SELECT` regression proving nested source subquery settings stay local (non-`RETURNING`). ## Alternatives considered - **`AND SELECT` / `THEN SELECT` keywords** — parser ambiguity or poor discoverability for Postgres migrants. - **Atomic Postgres-style RETURNING** — much larger surface (formats, MVs, distributed, async); deferred as a possible follow-up. - **Multi-statement `;`** — out of scope. ## Open questions for maintainers 1. Is the non-atomic, user-supplied-SELECT model the right v1 framing? 2. Confirm canonical placement rules (before data for VALUES/FORMAT, after source for INSERT SELECT). 3. Restrict trailing SELECT to read-only queries only? (Current implementation: yes, via SELECT parser.) 4. Confirm single-result-set protocol behavior across HTTP/native/JDBC. ## Test plan - [x] Stateless test `04266_insert_returning_21697` - [x] `INSERT VALUES` + `RETURNING` with row filter - [x] `INSERT SELECT` + `RETURNING` with aggregate - [x] `RETURNING` against a table other than INSERT target - [x] INSERT failure prevents `RETURNING` subquery from running - [x] `NOT_IMPLEMENTED` when `async_insert=1` - [x] Native TCP `INSERT FORMAT` + `RETURNING` (`04267_insert_returning_native_format`) - [x] `input(...)` source with `FORMAT` before `RETURNING` (parser/formatter round-trip + execution) - [x] Source-side `SETTINGS` precedence and nested-source scoping (`04410_insert_returning_trailing_settings`) - [x] Plain `INSERT ... SELECT` keeps nested source subquery `SETTINGS` local (no non-`RETURNING` leakage) - [x] Source-side query-cache settings rejection in `INSERT ... RETURNING` source `SETTINGS` - [x] `02310_profile_events_insert` local `clickhouse-local` profile-events check is non-fatal in CI (`InsertedRows` is not emitted consistently there) - [x] Transactional rollback when delayed `RETURNING` planning fails (`implicit_transaction=1` and explicit transaction) - [x] `RETURNING` keyword parser regression (`SELECT 1 returning` rejected; `SELECT 1 AS returning` and quoted alias accepted) - [x] `04266` rollback subtests stabilized by avoiding fixed database-name collisions in flaky/parallel runs (use `t_insert_returning_tx` in the current test DB) - [ ] Replicated-table non-atomic behavior (integration test, future)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105714",
        "createdAt": "2026-05-24T01:08:43Z",
        "updatedAt": "2026-08-13T00:44:27Z",
        "timestamp": "2026-08-13T00:44:27Z",
        "metrics": {
          "reactions": 2,
          "comments": 32
        },
        "labels": [
          "pr-backward-incompatible"
        ],
        "author": "tylerhannan",
        "state": "open",
        "assignees": [
          "nikitamikhaylov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105780",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add SCANN Vector Index Support",
        "text": "Add [Google ScaNN](https://github.com/google-research/google-research/tree/master/scann) as a new approximate nearest neighbor (ANN) backend for the `vector_similarity` skip index in `MergeTree` tables, accessible via `TYPE vector_similarity('scann', ...)` syntax. ScaNN uses an IVF-based index with asymmetric hashing (LUT16) and exact reranking, complementing the existing HNSW backend. The implementation integrates three new contrib dependencies — `scann`, `highway` (portable SIMD), and `eigen` (header-only linear algebra) — without modifying any upstream submodule source files; all platform adaptation is in the corresponding `-cmake/` wrappers. **Supported distance functions:** `L2Distance`, `cosineDistance`, `dotProduct`. **New query-tuning settings:** - `scann_num_leaves_to_search` — number of partitions to search at query time (0 = use the index-configured default, typically `num_vectors^0.25`). - `scann_candidate_pool_size` — asymmetric hashing candidate pool size fed into the exact reranker (0 = automatic: `1000 × topK`). All training artifacts (partitioner proto, asymmetric hashing codebook, hashed dataset, datapoint-to-token mapping) are serialized to disk on index build via `SingleMachineSearcherConfig`, so search results are deterministic across server restarts without retraining. **Example:** ```sql CREATE TABLE tab ( id UInt64, vec Array(Float32), INDEX idx vec TYPE vector_similarity('scann', 'cosineDistance', 128) GRANULARITY 100000000 ) ENGINE = MergeTree ORDER BY id; -- Query-time tuning SELECT id, cosineDistance(vec, [...]) AS dist FROM tab ORDER BY dist LIMIT 10 SETTINGS scann_num_leaves_to_search = 80, scann_candidate_pool_size = 5000; ``` ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added Google ScaNN as a new backend for the vector_similarity index type (syntax: TYPE vector_similarity('scann', 'cosineDistance', N)). Supports L2Distance, cosineDistance, and dotProduct distance functions with deterministic results across restarts. Two new settings — scann_num_leaves_to_search and scann_candidate_pool_size — allow per-query tuning of recall/latency trade-offs. ### Related issue: Closes #103664. The issue mentions ScaNN, but the reply references SPANN — not sure if that was a typo. I went ahead and implemented ScaNN as originally requested, and would welcome any response.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105780",
        "createdAt": "2026-05-25T11:29:39Z",
        "updatedAt": "2026-08-13T06:46:39Z",
        "timestamp": "2026-08-13T06:46:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 27
        },
        "labels": [
          "submodule changed",
          "manual approve",
          "can be tested",
          "pr-experimental"
        ],
        "author": "lvzhipin03",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105823",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Iceberg: skip eager getContent() JSON serialization when metadata log is disabled",
        "text": "`AvroForIcebergDeserializer::getContent(row_index)` was passed as an argument to `insertRowToLogTable` and serialized a manifest entry to JSON on every call, even when `iceberg_metadata_log_level` (default `none`) discarded the result. Hoist the level check above the five call sites so the payload is only built when the log will consume it. CPU profile of a query over ~300k pruned manifest entries (`max_threads=1`): ~59% of CPU was JSON serialization (`writeJSONString`, `SerializationTuple::serializeTextJSON`, `SerializationObjectPool::getOrCreate` and the resulting allocator churn). ``` ┌─samples─┬───pct─┬─function─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ 3278 │ 13.15 │ DB::SerializationObjectPool::getOrCreate(wide::integer<128ul, unsigned int>, std::__1::function<DB::ISerialization* ()>) │ │ 2797 │ 11.22 │ DB::WriteBuffer::write(char) │ │ 2679 │ 10.75 │ syscall │ │ 1283 │ 5.15 │ DB::writeJSONString(char const*, char const*, DB::WriteBuffer&, DB::FormatSettings const&) │ │ 1009 │ 4.05 │ DB::SerializationTuple::serializeTextJSON(DB::IColumn const&, unsigned long, DB::WriteBuffer&, DB::FormatSettings const&) const │ │ 737 │ 2.96 │ je_sallocx │ │ 715 │ 2.87 │ DB::Iceberg::ManifestFileIterator::processRow(unsigned long) │ │ 635 │ 2.55 │ DB::SerializationNamed::getHash(std::__1::shared_ptr<DB::ISerialization const> const&, std::__1::basic_string<char, std::__1::char_traits<char>, std::__1::allocator<char>> const&, DB::ISerialization::Substream::Type) │ │ 632 │ 2.53 │ DB::SerializationObjectPool::getOrCreate(wide::integer<128ul, unsigned int>, std::__1::function<DB::ISerialization* ()>)::$_0::operator()(DB::ISerialization const*) const │ │ 558 │ 2.24 │ CurrentMemoryTracker::allocImpl(long, bool) │ │ 536 │ 2.15 │ DB::Iceberg::(anonymous namespace)::deserializeFieldFromBinaryRepr(std::__1::basic_string<char, std::__1::char_traits<char>, std::__1::allocator<char>>, std::__1::shared_ptr<DB::IDataType const>, bool) │ │ 499 │ 2 │ DB::SerializationTuple::create(std::__1::vector<std::__1::shared_ptr<DB::SerializationNamed const>, std::__1::allocator<std::__1::shared_ptr<DB::SerializationNamed const>>>, bool) │ │ 463 │ 1.86 │ │ │ 444 │ 1.78 │ DB::(anonymous namespace)::writeTraceInfo(DB::TraceType, int, siginfo_t*, void*) │ │ 444 │ 1.78 │ je_malloc │ │ 420 │ 1.68 │ je_nallocx │ │ 383 │ 1.54 │ DB::DataTypeTuple::doGetSerialization(DB::SerializationInfoSettings const&) const │ │ 381 │ 1.53 │ auto DB::Field::dispatch<DB::Field::create(DB::Field const&)::'lambda'(auto&), DB::Field const&>(auto&&, DB::Field const&) │ │ 310 │ 1.24 │ DB::Iceberg::IcebergSchemaProcessor::tryGetFieldCharacteristics(int, int) const │ │ 304 │ 1.22 │ rtree_read │ │ 250 │ 1 │ CurrentMemoryTracker::free(long) │ │ 226 │ 0.91 │ DB::Field::~Field() │ │ 220 │ 0.88 │ auto DB::Field::dispatch<DB::Field::create(DB::Field&&)::'lambda'(auto&), DB::Field&>(auto&&, DB::Field&) │ │ 195 │ 0.78 │ itoa(long, char*) │ │ 188 │ 0.75 │ DB::SerializationTuple::supportsPooling() const │ │ 177 │ 0.71 │ je_sdallocx │ │ 172 │ 0.69 │ DB::ISerialization::pooled(wide::integer<128ul, unsigned int>, std::__1::function<DB::ISerialization* ()>) │ │ 166 │ 0.67 │ DB::SerializationArray::serializeTextJSON(DB::IColumn const&, unsigned long, DB::WriteBuffer&, DB::FormatSettings const&) const │ │ 165 │ 0.66 │ DB::SerializationNamed::~SerializationNamed() │ │ 155 │ 0.62 │ operator new[](unsigned long) │ └─────────┴───────┴──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` related to #104169 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Avoid eager JSON serialization of Iceberg manifest entries when `iceberg_metadata_log_level` is below the call site's threshold.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105823",
        "createdAt": "2026-05-26T06:14:12Z",
        "updatedAt": "2026-08-13T03:41:06Z",
        "timestamp": "2026-08-13T03:41:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-performance",
          "manual approve",
          "can be tested"
        ],
        "author": "starpact",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105848",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use text index for LIKE/ILIKE with ESCAPE",
        "text": "Follow-up to #99774 (qoega's LIKE ESCAPE clause). Before this change, a query like `... WHERE col LIKE pattern ESCAPE 'c'` silently bypassed the primary key (`KeyCondition`), the text index (`TYPE text`), and the bloom-filter text indexes (`ngrambf_v1`, `tokenbf_v1`, `sparse_grams`), falling back to a full scan even on tables with such an index defined or with `col` in the sorting key. Root cause: `LIKE pattern ESCAPE 'c'` is parsed into a 3-argument function call `like(col, pattern, escape_char)`. `MergeTreeIndexConditionText::traverseAtomNode`, `MergeTreeConditionBloomFilterText::extractAtomFromTree`, and `KeyCondition::extractAtomFromTree` only accepted 2-argument forms; any 3-arg `like`/`ilike` was rejected without ever reaching the LIKE handler. The execution layer in `FunctionsStringSearch::executeImpl` already calls `likePatternWithCustomEscapeToLikePattern` to fold the custom escape into standard backslash escapes before matching, so result correctness was never affected, only performance. Apply the same fold at index-condition time: - Text index (`MergeTreeIndexConditionText`): handles 3-arg `like` and `ilike` (the same arities `isSupportedFunction` allows for the index). - Bloom-filter text indexes (`MergeTreeConditionBloomFilterText`, for `ngrambf_v1` / `tokenbf_v1` / `sparse_grams`): handles 3-arg `like` and `notLike` (the like-family forms this index supports; `ilike` is not supported by this index type). - Primary key (`KeyCondition`): handles 3-arg `like` and `notLike` (the like-family forms present in `atom_map`). Case-insensitive `ilike`/`notILike` are not in `atom_map` and continue to fall back to row-level evaluation, same as the existing 2-arg behavior. In all paths the pattern must be a constant String and the escape argument a constant String of length 1; otherwise we bail out. Invalid escape sequences cause analysis to skip the index/key so row-level evaluation throws the same error the user would have seen before. The `nextInStringLike` tokenizer used by both text-index variants drops the backslash for an unknown escape `\\c` while row-level matching keeps it, so an unknown/trailing backslash is declined (the `likePatternHasUnknownBackslashEscape` guard) and falls back to row-level matching to avoid wrongly pruning a matching granule. In the bloom-filter path this guard is placed in the shared `traverseTreeEquals`, so it protects both the folded 3-argument form and the pre-existing 2-argument `like` / `notLike` / `mapContainsKeyLike` / `mapContainsValueLike` paths. Regression tests use `EXPLAIN indexes = 1` to assert that pruning now happens: - `04277_text_index_function_like_escape.sql` covers the `TYPE text` index for `LIKE`/`ILIKE`, both `splitByNonAlpha` and `array` tokenizers, and both the operator form and the functional form `like(a, p, 'c')`. - `04292_105885_primary_key_like_escape.sql` covers the primary key for `LIKE`/`NOT LIKE` (with a remaining wildcard so the perfect-prefix extraction succeeds) and the functional form. - `04351_bloom_filter_text_index_like_escape.sql` covers the `ngrambf_v1` / `tokenbf_v1` / `sparse_grams` text indexes for `LIKE`/`NOT LIKE` ESCAPE and the functional form, including a consumed escape character, `force_data_skipping_indices`, and the unknown/trailing-backslash decline for both the 2-argument and 3-argument forms. The tests FAIL on master without this change and PASS with it. Closes #105885 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `LIKE ... ESCAPE 'c'` and `ILIKE ... ESCAPE 'c'` predicates (added in #99774) now use `TYPE text` skip indexes to prune granules; `LIKE ... ESCAPE 'c'` and `NOT LIKE ... ESCAPE 'c'` additionally use the `ngrambf_v1`, `tokenbf_v1`, and `sparse_grams` skip indexes, and the primary key when the column is in the sorting key, instead of falling back to a full scan. ### Documentation entry for user-facing changes - [x] Documentation is unchanged (the existing text-index docs already describe LIKE-based pruning; this change only fills a coverage gap for the recently added ESCAPE clause).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105848",
        "createdAt": "2026-05-26T11:56:24Z",
        "updatedAt": "2026-08-13T11:49:00Z",
        "timestamp": "2026-08-13T11:49:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 38
        },
        "labels": [
          "pr-improvement",
          "can be tested",
          "hold"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "ahmadov",
          "rschu1ze"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:105987",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Constant filter folding under materialize",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allows to constant-fold filters through materialize wrappers. Closes https://github.com/ClickHouse/ClickHouse/issues/78166#issuecomment-2758343581",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/105987",
        "createdAt": "2026-05-27T17:00:57Z",
        "updatedAt": "2026-08-13T05:53:36Z",
        "timestamp": "2026-08-13T05:53:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yariks5s",
        "state": "open",
        "assignees": [
          "vdimir",
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106006",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix LOGICAL_ERROR in DatabaseReplicatedDDLWorker with max_replication_lag_to_enqueue=0",
        "text": "Reject `max_replication_lag_to_enqueue = 0` at parse time everywhere it can be supplied. `0` was never useful: it asks the post-recovery unsynced check `max_log_ptr + max_replication_lag_to_enqueue <= new_max_log_ptr` to be trivially true (the ZooKeeper counter is monotonically non-decreasing), which mis-flags a caught-up replica as unsynced and then trips `chassert(our_log_ptr < max_log_ptr)` in `DatabaseReplicatedDDLWorker::initAndCheckTask`, aborting the server on debug and sanitizer builds. Switching the setting type from `UInt64` to `NonZeroUInt64` lets the parser reject `0` from every source: `CREATE DATABASE ... SETTINGS`, the `<database_replicated>` server-config block, `ATTACH` replay, and upgrade-time replay of existing database metadata. The post-recovery comparison stays at the original `<=` because the bad value can no longer reach the worker. The smallest valid value is `1`. New regression test `04298_98823_replicated_zero_lag_assertion` checks both the rejection (`A setting's value has to be greater than 0`, BAD_ARGUMENTS) and the smallest valid value (`1`) going through the post-recovery DDL path that previously aborted. Adapted from the analysis in #100043. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reject `max_replication_lag_to_enqueue = 0` on `Replicated` databases. The value made the post-recovery unsynced check trivially true and aborted the server with `LOGICAL_ERROR` in debug and sanitizer builds. `0` is now rejected at parse time with `BAD_ARGUMENTS` from every source (`CREATE DATABASE ... SETTINGS`, `<database_replicated>` server-config block, `ATTACH` replay, and upgrade-time replay of existing metadata). The smallest valid value is `1`. Closes #98823. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106006",
        "createdAt": "2026-05-27T19:40:08Z",
        "updatedAt": "2026-08-13T00:37:31Z",
        "timestamp": "2026-08-13T00:37:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 20
        },
        "labels": [
          "pr-bugfix",
          "manual approve",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alesapin"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106011",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "DeltaLake: Use new create table transaction",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Use new create table transaction resolves https://github.com/ClickHouse/ClickHouse/issues/103155 this needs https://github.com/ClickHouse/ClickHouse/pull/105861",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106011",
        "createdAt": "2026-05-27T21:12:24Z",
        "updatedAt": "2026-08-13T13:00:43Z",
        "timestamp": "2026-08-13T13:00:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-feature"
        ],
        "author": "SmitaRKulkarni",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106019",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parquet: map TIME logical type to Time64 ",
        "text": "The native Parquet reader currently maps both `logical.TIMESTAMP` and `logical.TIME` (and the deprecated converted types `TIMESTAMP_MILLIS`/`TIMESTAMP_MICROS`/`TIME_MILLIS`/`TIME_MICROS`) to `DataTypeDateTime64`. That was fine before ClickHouse 25.6, when there was no time-of-day type, but it now produces the Parquet analogue of #104038: when the target column is `Time64` and `session_timezone` is non-UTC, the `DateTime64 -> Time64` cast (PR #90310) shifts the value by the timezone — silent data corruption for a value that has no date/timezone to begin with. This PR splits the schema-converter branch so that: - `logical.TIME` / `TIME_MILLIS` / `TIME_MICROS` → `DataTypeTime64(scale)` (timezone-unaware). - `logical.TIMESTAMP` → `DataTypeDateTime64(scale, timezone)` (unchanged). It also extends `is_output_type_decimal` to recognize `Time64` so min/max stats stay enabled when the output column is `Time64`, and removes the stale comment that claimed ClickHouse has no time-of-day type. While covering explicit type hints in the test (`SELECT * FROM file(..., 'Parquet', 't Time')`), I hit a second pre-existing gap from 25.6: `Time64 -> Time` was never registered as a cast — it threw `CANNOT_CONVERT_TYPE` because `IsDataTypeNumber<DataTypeTime>` is `false` (it inherits `DataTypeNumberBase<Int32>`, not `DataTypeNumber<Int32>`). Before this PR the path went `DateTime64 -> Time` and returned a wrong, timezone-shifted value; after the Parquet fix it would have thrown. The second commit adds an inline `Time64 -> Time` branch in `ConvertImpl` that divides by the scale multiplier — no timezone offset, because both Time64 and Time are timezone-unaware seconds-of-day. A new stateless test `04277_parquet_time_of_day` covers: - Schema-less inference (Parquet `time32[ms]`/`time64[us]`/`time64[ns]` → `Nullable(Time64(3/6/9))`; `TIMESTAMP` stays `Nullable(DateTime64(6, 'UTC'))`). - Import into `Time64(scale)` columns under `session_timezone = 'Asia/Shanghai'` (values must not shift). - Explicit type hints under non-UTC session: `Time64(3)` / `Time` / `DateTime64(3, 'UTC')` / `Int64`. Review follow-up: making `Time64 -> DateTime` legal also exposed that `ToDateTimeMonotonicity` treated `Time64` sources as monotonic, while the cast wraps negative time-of-day values with the default `date_time_overflow_behavior = 'ignore'` (`-00:00:01` -> `2106-02-07 06:28:15`), i.e. it drops at zero. `toDateTime` now reports non-monotonic for `Time`/`Time64` sources; the regression test `04828_time64_to_datetime_not_monotonic` shows that without the guard the key condition prunes every granule and range/point queries over a `Time64` key silently return empty results. The Parquet docs (`docs/reference/formats/Parquet/Parquet.mdx`) are updated for the new `TIME` -> `Time64` mapping. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix Parquet `time32`/`time64` columns being read as `DateTime64` and then silently shifted by `session_timezone` when inserted into a `Time64` target column. Parquet `TIME` logical types now map directly to `Time64`, and a previously-missing `Time64 -> Time` cast is added. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106019",
        "timestamp": "2026-08-12T22:53:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-autogenerated-docs"
        ],
        "author": "tiandiwonder",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106199",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Push down volume-reducing functions in query plan",
        "text": "Adds a query plan optimization that pushes volume-reducing functions (`length`, `lengthUTF8`, `empty`, `notEmpty`) below the `Sorting` and `Filter` steps, so those steps carry the fixed-size result instead of the original `String` / `FixedString` / `Array` / `Map` argument. The rewrite is done with `ActionsDAG::split`: the function nodes form the first part, which becomes a new step below the child, and the second part becomes the parent step's `ActionsDAG` and reads the results as inputs. It only fires when the wide argument column really stops flowing through the child step — the column is removed from the child's output, and unless the child reads it itself it is not even produced below the child. `tryExecuteFunctionsAfterSorting` and `trySplitFilter` keep the pushed functions where they are, so the optimizations do not move the same nodes in opposite directions. For the shape from the issue below, `SELECT avg(length(s)) FROM test WHERE notEmpty(s)`, the filter still needs `s` to evaluate its condition, but `length(s)` is now computed before the filter, so the wide column is not copied for the surviving rows. Enabled by default; can be turned off with `query_plan_push_down_volume_reducing_functions = 0`. An earlier attempt at the same optimization was https://github.com/ClickHouse/ClickHouse/pull/86139 by @talmawash, credited as a co-author. Closes: https://github.com/ClickHouse/ClickHouse/issues/82378 Related: https://github.com/ClickHouse/ClickHouse/pull/86139 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a query plan optimization that pushes volume-reducing functions (`length`, `lengthUTF8`, `empty`, `notEmpty`) below the `Sorting` and `Filter` steps, replacing the wide `String` / `FixedString` / `Array` / `Map` argument with the fixed-size result, so it is neither buffered by a sort nor copied by a filter. Controlled by the new setting `query_plan_push_down_volume_reducing_functions`, enabled by default. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106199",
        "createdAt": "2026-05-31T13:26:13Z",
        "updatedAt": "2026-08-13T11:21:58Z",
        "timestamp": "2026-08-13T11:21:58Z",
        "metrics": {
          "reactions": 2,
          "comments": 12
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "fastio",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106215",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Enable read_in_order_use_virtual_row by default",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/52624 When reading in order of the primary key over many parts (the `optimize_read_in_order` optimization), the merging pipeline opens one reader per part to feed the `MergingSortedTransform`, and keeps a `CompressedReadBuffer` per column resident for every one of them. With a large number of parts this makes even `ORDER BY pk LIMIT n` consume memory proportional to the number of parts, even though only a few of them can contribute to the result. The original report describes a query needing ~42 GB for this reason (1803 parts × 3 columns × 8 MB compress blocks). The `read_in_order_use_virtual_row` optimization lets `MergingSortedTransform` reprioritize sources using primary key values taken from the sparse index, so the parts that cannot contribute to the answer are never read and their readers are never created. The optimization already existed but was disabled by default; this change enables it. Measured on a table with 300 parts and an 8 MB compress block size, the peak memory of `SELECT * FROM t ORDER BY k LIMIT 10` drops from ~700 MB to a few MB. Related issues: #53101, #53287, #7413 (closed), and #89638 (open) describe the same underlying behavior. Note: the more aggressive `read_in_order_use_virtual_row_per_block` (which also disables `read_in_order_use_buffering` and the preliminary merge) is intentionally left disabled by default, since it trades some parallelism for the extra memory savings. **Update.** The first CI run showed consistent 1.3x–5x slowdowns in read-in-order perf tests (`distinct_in_order`, `optimize_sorting_for_input_stream`, `read_in_order_many_parts`, `monotonous_order_by`, `redundant_functions_in_order_by`) on both architectures. The cause: a source that starts with a virtual row is not read until the merge reaches its key, so the merge requests deferred sources strictly one by one and reading degrades to a single thread (locally, `SELECT DISTINCT` over 100M rows ran at the same speed with `max_threads=1` and `max_threads=16`). This was addressed by letting deferred sources read ahead: without a `LIMIT` a window of `number of threads` deferred sources, and with a `LIMIT` only sources provably needed soon (those whose virtual row is not greater than the key at which the merge must stop within an already-read chunk). **Update 2.** The second perf run and review showed that the \"provably needed\" rule for the `LIMIT` case was wrong in both directions: - it was unbounded: with many overlapping parts, every deferred source whose virtual row preceded the merge stop row was scheduled at once, which could re-create the O(parts) resident-reader memory issue this PR fixes; - it was still too conservative: with a selective filter (`SELECT * FROM hits WHERE UserID = ... ORDER BY pk LIMIT 100`), filtered chunks rarely reach the merge, the proof never arrives, and reading degrades to a single stream again — `order_by_read_in_order` showed 3.3x–4.7x and `lazyMaterialization` 1.9x–2.5x slowdowns on both architectures. Both are fixed by making the read-ahead window unconditional: the merge always keeps the next `number of threads` deferred sources (in the order it will need them) reading ahead, with or without a `LIMIT`. Under a `LIMIT` the read-ahead is speculative — the merge may finish before reaching a prefetched source — but the waste is bounded by O(threads) chunks, because a source past its virtual row only buffers one chunk in its output port until the merge consumes it. Peak reader memory stays bounded by the window instead of the number of parts: the 300-part `LIMIT 10` query above runs with ~7 MB peak memory and reads only the window's first chunks, vs ~1.3 GB and a full read without the optimization. Pending read-ahead is dropped as soon as the merge finishes, so a query that already produced its `LIMIT` does not start new reads. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable `read_in_order_use_virtual_row` by default. When reading in order of the primary key (e.g. `ORDER BY primary_key LIMIT n`) over a table with many parts, only the parts that can actually contribute to the result are read (plus a read-ahead window of at most `max_threads` parts that keeps reading parallel), which significantly reduces peak memory consumption. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106215",
        "createdAt": "2026-05-31T22:32:25Z",
        "updatedAt": "2026-08-13T05:06:22Z",
        "timestamp": "2026-08-13T05:06:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 37
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106231",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Implement CREATE HANDLER: SQL-defined HTTP handlers",
        "text": "Implements [#100000](https://github.com/ClickHouse/ClickHouse/issues/100000): create and manage custom HTTP handlers from SQL, without editing the server configuration file. New SQL statements: - `CREATE HANDLER [IF NOT EXISTS] name [PROTOCOL p] URL [PREFIX|REGEXP] '/x' [METHODS (GET, POST)] [TYPE query] AS <query>` - `ALTER HANDLER name ...` — partial update of any subset of clauses - `DROP [IF EXISTS] HANDLER name` Handlers are matched after configuration-defined handlers, in the lexicographical order of their names. The URL is matched without the `?` query string and `#` fragment; exact/prefix URLs are checked for ambiguity at create/alter time. The query is parsed (not analyzed) at creation, can be parameterized (URL params, form variables, headers, and named regexp capture groups), and an `INSERT` handler reads its data from the HTTP body. Handlers are persisted in a local or Keeper storage, mirroring named collections, configured via the `query_rules_storage` config section, and kept in sync across replicas. Access control adds `CREATE HANDLER`, `ALTER HANDLER` and `DROP HANDLER` grants. Introspection adds the `currentHandler` and `currentRequestURL` functions, the `http_handler_name` and `http_request_url` columns in `system.query_log`, and the `system.handlers` table. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `CREATE HANDLER`, `ALTER HANDLER` and `DROP HANDLER` statements to define custom HTTP handlers from SQL, persisted in a local or Keeper storage. Added the `currentHandler` and `currentRequestURL` functions, the `system.handlers` table, and the `http_handler_name` / `http_request_url` columns in `system.query_log`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106231",
        "createdAt": "2026-06-01T12:09:45Z",
        "updatedAt": "2026-08-13T05:52:33Z",
        "timestamp": "2026-08-13T05:52:33Z",
        "metrics": {
          "reactions": 1,
          "comments": 79
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106673",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add LossyQuantile codec for lossy compression of floating-point numbers",
        "text": "Resolves #106219 Implement a new experimental compression codec `LossyQuantile` that encodes floating-point values as K-bit indices into per-group empirical quantile codebooks. **Syntax:** `CODEC(LossyQuantile(bits [, group_size [, stripe_size]]))` The codec accepts parameters: - `bits` (1-8): number of bits per encoded value, controls compression ratio vs. accuracy tradeoff - `group_size` (default 1048576): number of values per quantile group - `stripe_size` (default 1): interleaving stride for multi-dimensional data (e.g., fixed-size embedding arrays) **Compression:** For each group, sort the values, compute 2^K interpolated quantiles at levels `(q+1)/(2^K+1)`, then encode each value as the index of its nearest centroid. Bit-pack all K-bit indices. **Decompression:** Read centroids from the header, unpack indices, look up centroid values. **Special value handling:** NaN is mapped to the lowest centroid, Inf/-Inf are mapped to extreme centroids. These are excluded from quantile calculation. **Use case:** Brute-force vector search on slow media (S3) where I/O bandwidth dominates compute cost. Achieves 4-32x compression depending on K, with bounded distribution-adaptive reconstruction error. Requires `SET allow_experimental_codecs = 1` (same as ALP). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental `LossyQuantile` codec for lossy compression of floating-point columns using quantile-based scalar quantization. Encodes each value as a K-bit index (1-8 bits) into per-group empirical quantile codebooks, achieving 4-32x compression with bounded reconstruction error. Supports stripe mode for multi-dimensional array data. Made with [Cursor](https://cursor.com)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106673",
        "createdAt": "2026-06-07T11:09:19Z",
        "updatedAt": "2026-08-13T14:17:34Z",
        "timestamp": "2026-08-13T14:17:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "can be tested",
          "pr-experimental"
        ],
        "author": "anandheritage",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:106734",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "More settings to randomize",
        "text": "Extends the stateless-test randomizers in `tests/clickhouse-test` with settings added since they were last swept, and pins the tests that were implicitly relying on the old defaults. `SettingsRandomizer`: * Broadens `query_plan_optimize_join_order_algorithm` to cover `dpsub` and `dphyp` (always with `greedy` kept in the fallback chain, since exhausting the chain throws `EXPERIMENTAL_FEATURE_ERROR`), and randomizes `query_plan_optimize_join_order_max_searched_plans`. * Adds ~35 further query-level settings (`LIMIT BY` / `FINAL` / lazy-`FINAL` / runtime-filter / text-index / statistics / regexp-compilation toggles and their numeric companions). * Adds `use_statistics_for_part_pruning` → `use_statistics` and `use_projection_index_in_read_pools` → `optimize_use_projection_filtering` to `conditional_settings`, so a child setting is not enabled while its parent is off. * Randomizes `function_implementation` over the values the server actually supports, probed at startup: it globally filters every `ImplementationSelector`-backed function and each has a different arch-tag set, so a forced value that some function lacks would raise `NO_SUITABLE_FUNCTION_IMPLEMENTATION`. * A few settings are deliberately left commented out with the reason recorded in place, e.g. `use_constant_folding_in_index_analysis` (unsafe per-part min-max prune, and the trigger for https://github.com/ClickHouse/ClickHouse/issues/109893), `correlated_subqueries_use_in_memory_buffer`, `min_filtered_ratio_for_lazy_final` (https://github.com/ClickHouse/ClickHouse/issues/112332) and `max_streams_for_union_step`. `MergeTreeSettingsRandomizer`: * Randomizes `concurrent_part_removal_threshold_for_remote_disk`. * The implicit min-max index settings (`add_minmax_index_for_block_number_column`, `add_minmax_index_for_block_offset_column`, `part_minmax_index_columns`) are explicitly kept out: they are not transparent to queries. The first two add real `MINMAX` skipping indices named `auto_minmax_index__block_number` / `auto_minmax_index__block_offset`, which surface in `SHOW INDEXES`, `system.data_skipping_indices`, part checksums and the `Skip` section of `EXPLAIN indexes = 1`; `part_minmax_index_columns = with_block_number_offset` gives every table a part-level min-max index even with no partition key, so `EXPLAIN indexes = 1` grows an extra `Min-Max / Condition: true` block everywhere. This is the same reason their siblings `add_minmax_index_for_{numeric,string,temporal}_columns` were never randomized — all five are `isReadonlySetting`, i.e. part of the table schema. The rest of the diff pins the affected settings in the stateless tests whose output depends on them. Related: https://github.com/ClickHouse/ClickHouse/pull/107586 (fixes the `Invalid binary search result in MergeTreeSetIndex` logical error this randomization made much more likely to be hit; merged, so it is pulled in by the branch update) Related: https://github.com/ClickHouse/ClickHouse/issues/109893 Related: https://github.com/ClickHouse/ClickHouse/issues/112332 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Running CI a few times before merging.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/106734",
        "createdAt": "2026-06-08T16:44:21Z",
        "updatedAt": "2026-08-13T15:43:04Z",
        "timestamp": "2026-08-13T15:43:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-ci"
        ],
        "author": "PedroTadim",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107028",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/106364 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a rare `LOGICAL_ERROR` \"Attempt to release query context that does not exist\" and the accompanying server crash when reading `MergeTree` tables through a filesystem cache disk created with `enable_filesystem_query_cache_limit = 1`. ### Description `FileCacheQueryLimit::query_map` (a `std::unordered_map`) was reachable from two different cache locks with no single lock serializing it: - the read `tryGetQueryContext` (`query_map.find`) runs under `CacheStateGuard::Lock` (called from `FileCache::doTryReserve`, once per reservation); - the writes `getOrSetQueryContext` (`query_map.emplace`) and `removeQueryContext` (`query_map.erase`) run under `CachePriorityGuard::WriteLock` (called from `FileCache::getQueryContextHolder` and `~QueryContextHolder`). Since the read takes a different mutex than the writes, a `find()` can run concurrently with an `emplace()`/`erase()` that rehashes the table and invalidates buckets and iterators. That is undefined behaviour on `std::unordered_map` and surfaces as either `Attempt to release query context that does not exist` (the map is corrupted so a present key looks absent in `removeQueryContext`) or a server crash while reading a half-rehashed bucket. Both were seen together in the same CI run on `Stateless tests (amd_debug, sequential)`. The split appeared when `tryGetQueryContext` was moved from the priority write lock to the cheaper cache state lock, leaving the map reachable from two locks at once. Fix: guard all three `query_map` accessors with a dedicated leaf mutex, so the map has a single owning lock regardless of which cache lock the caller holds. The mutex is taken last and never nests another cache lock under it, so it does not change the existing lock order, and the optimization of keeping the read off the priority write lock is preserved. Reproduced with a concurrent reader/writer witness over an S3-backed cache disk with `enable_filesystem_query_cache_limit = 1`: 2073 overlapping accesses to `query_map` without the fix, 0 with it. ### Follow-up: last-holder release leak (#109508) @ Algunenano found a related issue by code analysis (#109508): the last-holder decision in `~QueryContextHolder` still dropped this holder's own reference to the context outside the cache lock, and `removeQueryContext` gated the erase on `use_count() > 2` before that drop. `use_count()` is not a synchronization primitive, so this is a TOCTOU. A query with parallel read streams has several holders for the same `query_id` (each `CachedOnDiskReadBufferFromFile` creates its own), so `use_count > 2`. When two holders release at the same time both observe the shared count, both skip the erase, and after both drop their reference only the map entry remains and is never removed. That orphans `query_map[query_id]` for the lifetime of the cache, so a later query reusing the same `query_id` picks up stale per-query limit state. Fix: drop this holder's reference inside `removeQueryContext` under `query_map_mutex` and erase only once the map entry is the sole owner (`use_count() == 1`). Every reference change to the context is then serialized by the lock that also guards `getOrSetQueryContext`, so the check-and-erase is atomic with respect to concurrent holders. Added `FileCacheTest.QueryLimitConcurrentReleaseNoLeak`, which drives the concurrent-release interleaving deterministically (fails on the previous `use_count() > 2` logic, passes now). Related: https://github.com/ClickHouse/ClickHouse/issues/109508 <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.1.840` (included in `26.7` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107028",
        "createdAt": "2026-06-10T16:02:03Z",
        "updatedAt": "2026-08-13T03:34:12Z",
        "timestamp": "2026-08-13T03:34:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "can be tested",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107091",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not terminate the server on retryable errors while loading outdated parts",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/106736 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Do not terminate the server when a transient retryable error (such as `MEMORY_LIMIT_EXCEEDED` or a network error) occurs while loading outdated or unexpected data parts in the background. Such errors are not a sign of an inconsistent on-disk set of parts, so the loading task is now retried instead of aborting the whole server. The existing fail-fast behaviour is preserved for genuinely inconsistent state. ### Description Closes #106736 `MergeTreeData::loadOutdatedDataParts()` and the sibling `loadUnexpectedDataParts()` wrap their whole body in a `catch (...)` that calls `std::terminate()`, to fail fast on a genuinely inconsistent on-disk set of parts. The same `catch` also fires for transient runtime errors such as `MEMORY_LIMIT_EXCEEDED` (including the stress test memory fault injector) or network errors, which are not on-disk corruption. This takes the whole server down in the experimental `serverfuzz` stress jobs. `loadDataPartWithRetries()` already classifies and retries such exceptions via `isRetryableException()`, but only around `loadDataPart()` itself; an exception thrown elsewhere in the loader (part removal, ZooKeeper cleanup, rename, the runner machinery, or newly memory-tracked allocations) escapes straight to the terminating `catch`. The fix reuses `isRetryableException()` in both `catch` blocks: for a retryable error in the asynchronous background loader, it logs a warning and reschedules the loading task instead of terminating, so the server stays up and retries once the transient condition clears. Non-retryable errors and the synchronous drop-table path keep the existing fail-fast behaviour. The unexpected-parts loop is made idempotent so a rescheduled retry skips already-loaded parts. A `REGULAR` failpoint `mergetree_load_outdated_parts_inject_retryable_exception` and an integration test (`test_load_outdated_parts_retryable`) inject a retryable error into the outdated-parts loader and verify the server survives and finishes loading. The retry must not requeue a part whose retryable error happened after it was already published. `loadDataPartWithRetries()` inserts a normal Outdated part into `data_parts_indexes` before its post-load cleanup (`preparePartForRemoval` writes the removal TID and can throw retryably); requeueing it as-is made the retry reload the same directory as a fresh part, hit the duplicate-part path, and `remove()` could delete the directory while the published part still referenced it. The published part is now rolled back out of the index before requeueing so the retry reloads it cleanly. The unexpected-parts loop has the same shape (its optional `broken-on-start` detach runs after the part is set), so it now uses a separate `finished` marker set only after the detach succeeds. Two more failpoints exercise these post-load paths and the integration test gains a third case asserting nothing is wrongly detached. The broken-part detach in both loaders no longer hides its errors. It used to call `renameToDetached(\"broken-on-start\", /*ignore_error=*/ replicated)`, and `renameToDetached` swallows `ErrnoException` and `fs::filesystem_error` when `ignore_error` is set - exactly the exception types a failing rename raises, and exactly the class `isRetryableException` treats as transient. On a replicated table a transient detach failure therefore never reached the new classification: the worker continued as if the cleanup had succeeded, leaving the broken outdated part on its original path with its ZooKeeper entry already removed, or setting the unexpected part's `finished` marker so the retry skipped it. Both call sites now pass `ignore_error = false` and classify the error themselves; the tolerance `ignore_error` gave replicated tables is preserved, but only for a permanent failure and only after the retryable case has been handled. A failpoint that throws the swallowed exception type and a regression case per loader cover it. #### Provenance Surfaced on master as the `'px != 0' failed` family whose actual fatal is the `loadOutdatedDataParts` terminate. Latest master hit: `Stress test (experimental, serverfuzz, arm_release)`, commit `81a8bcef7e8ba2dae3dfdc762dde87a853f40070`, STID 3782-2f00. Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=81a8bcef7e8ba2dae3dfdc762dde87a853f40070&name_0=MasterCI&name_1=Stress%20test%20%28experimental%2C%20serverfuzz%2C%20arm_release%29",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107091",
        "createdAt": "2026-06-10T21:59:10Z",
        "updatedAt": "2026-08-13T13:02:49Z",
        "timestamp": "2026-08-13T13:02:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 37
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "v26.6-must-backport"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107125",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add query plan cache for complex queries (views, joins, subqueries)",
        "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add an experimental query plan cache supporting complex `SELECT` queries: views (including nested views), joins, `UNION`, `IN` subqueries and multiple tables over supported local table engines. When `allow_experimental_query_plan_cache = 1` and `enable_query_plan_cache = 1`, repeated identical queries skip query analysis and logical planning (the query is still parsed) while still executing against current data. `SYSTEM DROP QUERY PLAN CACHE` clears the cache; events `QueryPlanCacheHits`/`QueryPlanCacheMisses`/`QueryPlanCacheValidationMisses`/`QueryPlanCacheStaleMisses` and metrics `QueryPlanCacheBytes`/`QueryPlanCacheEntries` provide observability. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) --- ### Description An experimental **query plan cache** for `SELECT` queries that supports **complex queries** — views (including nested views), joins, `UNION`, `IN` subqueries and multiple tables over supported local table engines. It complements https://github.com/ClickHouse/ClickHouse/pull/99309, which covers single-table `MergeTree` queries with a universalize/materialize design; this PR takes a different approach to reach query shapes where planning cost is actually concentrated (deep view chains, many joins). Unlike the query result cache, a plan cache hit still executes the query: only query analysis and logical planning are skipped (the query is still parsed, since the cache key is built from the parsed AST), so hits are transactionally consistent and always see current data. #### Design On a miss, the plan is built in the planner's existing `build_logical_plan` mode (the machinery used by distributed plan shipping), with a new `SelectQueryOptions::cacheable_logical_plan` mode that makes the logical plan self-contained: - Leaf reads are storage-agnostic `ReadFromTable` placeholders for any eligible local engine. - **Views are expanded at plan time** (only `StorageView`; a materialized view stays a leaf and is read from its target table like an ordinary table), recursively in logical mode (propagated through `StorageView::readImpl` via `SelectQueryInfo::build_logical_plan`), so the cached plan embeds the analyzed view bodies instead of deferring the expensive expansion to execution. - Key-value direct-join lookup steps (`JoinStepLogicalLookup`), which bind live storages into the plan, are not used; a regular join is planned instead and the physical join algorithm is chosen at materialization. Distributed and parallel-replicas logical plans do not set this mode and are unaffected. - The plan is serialized with the standard query plan serialization and built with `compile_expressions = 0` (JIT-fused function nodes cannot be serialized; the JIT still applies when pipelines are built from the plan). Both hits and misses execute through `QueryPlan::resolveStorages`, which re-binds every leaf to a fresh storage snapshot. #### Cache key and invalidation The key is the normalized AST hash, a hash of plan-affecting settings, the current database, the user and the sorted role set. Each entry stores a **dependency fingerprint over every referenced storage**, discovered from the plan's `ReadFromTable` leaves plus an AST closure through view definitions (which also catches tables referenced only from scalar subqueries): UUID, metadata version (or a schema content hash including the view body), and the row policy hash. Every hit revalidates all dependencies and re-checks `SELECT` access, so `DROP`/`CREATE`, `ALTER`, view redefinition (including **nested** views), row policy changes and permission revocation are handled transparently. #### Eligibility - Non-deterministic functions anywhere — including inside expanded view bodies — make a query uncacheable (`arrayJoin` is exempt: it is multi-valued but pure). - Scalar subqueries are evaluated during analysis and baked into the plan as constants, so they are gated behind the opt-in setting `query_plan_cache_allow_scalar_subqueries`. - Table functions, temporary tables, system tables (except `system.one`), remote/`Merge` storages and views whose SQL security is not `INVOKER` (`DEFINER` and `NONE`) are rejected. - Plans containing steps without serialization support (e.g. window functions) execute normally and are simply not stored. #### Performance, proven on DOOMbench [DOOMbench](https://github.com/cedardb/doombench) is a DOOM-like raycasting engine implemented entirely in SQL; its `screen` query renders a frame through a deep chain of views over joins, unions, array joins and scalar subqueries. Ported to ClickHouse SQL, the workload is **query-analysis-bound**: profiling showed `IQueryTreeNode::getTreeHash` and analyzer resolution dominating, and `SELECT ... LIMIT 0` was as slow as the full query. With the plan cache (local server, single connection): | DOOMbench frame render | median | speedup | | --- | --- | --- | | cache off | 551-577 ms | — | | cache on (hit) | 227-246 ms | **2.4x** | | cache on, `max_threads = 4` | 148-176 ms | **3.3x** | On hits the analyzer cost disappears from profiles (`getTreeHash`: 25.8k samples → 174); the remaining time is genuine execution. Correctness validated on the same workload: frames are byte-identical between cache off/miss/hit, game-state mutations are visible on every cached render (HTAP freshness), nested view redefinition is picked up immediately, and `rand()` inside a view body refuses to cache. #### Notes - This PR includes the `ActionsDAG` input-constant serialization fix https://github.com/ClickHouse/ClickHouse/pull/107124 as its first commit; it will be rebased once that PR is merged. - The miss path executes through the same logical-plan machinery as hits, so hit and miss behavior is identical by construction; storage-specific planning shortcuts (e.g. trivial `count()`) are not applied to cache-enabled queries. Related: https://github.com/ClickHouse/ClickHouse/pull/99309 Related: https://github.com/ClickHouse/ClickHouse/pull/107124",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107125",
        "createdAt": "2026-06-11T01:46:36Z",
        "updatedAt": "2026-08-13T10:14:20Z",
        "timestamp": "2026-08-13T10:14:20Z",
        "metrics": {
          "reactions": 0,
          "comments": 28
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107292",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add setting to omit CSV quotes for date and time types",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/34668 Adds `output_format_csv_quote_date_time_types`, a backward-compatible CSV output setting for users that need date and time values emitted without surrounding double quotes while keeping existing `CSV` behavior by default. The setting applies to `Date`, `Date32`, `DateTime`, `DateTime64`, `Time`, and `Time64` values in `CSV` output. Strings remain quoted according to the existing `CSV` serializer and numbers are unchanged. CSV input parsing is unchanged and already accepts both quoted and unquoted date and time values. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds the setting `output_format_csv_quote_date_time_types`. When disabled, `Date`, `Date32`, `DateTime`, `DateTime64`, `Time`, and `Time64` values are written without surrounding double quotes in `CSV` output, while preserving the existing quoted behavior by default. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107292",
        "createdAt": "2026-06-11T22:23:38Z",
        "updatedAt": "2026-08-13T08:23:12Z",
        "timestamp": "2026-08-13T08:23:12Z",
        "metrics": {
          "reactions": 1,
          "comments": 10
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "cristhiank",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107305",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Make ALTER MODIFY COLUMN on named Tuple metadata-only when adding subfields",
        "text": "`ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` that only adds new subfields is now metadata-only (no mutation). Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1` (default `false`). Subfield additions through `Array`/`Map`/nested `Tuple` wrappers are also handled. `Nullable(Tuple(...))` is blocked (null map incompatibility). Removing/renaming subfields or changing types still triggers a mutation. Key/index/projection guards reject the metadata-only path when `primary.idx` or skip-index bytes would become invalid (whole tuple in key, or subcolumn whose type changes). Subcolumn references with unchanged types (e.g. `ORDER BY t.a` when only `t.c` is added) are allowed. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` is now metadata-only when only adding subfields, matching the speed of top-level `ADD COLUMN`. Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1`. ### Documentation entry for user-facing changes - [x] Documentation is not required (behavioral improvement; semantics unchanged)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107305",
        "createdAt": "2026-06-12T07:42:58Z",
        "updatedAt": "2026-08-13T17:45:47Z",
        "timestamp": "2026-08-13T17:45:47Z",
        "metrics": {
          "reactions": 1,
          "comments": 8
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "amosbird",
        "state": "open",
        "assignees": [
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107436",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Automatic LowCardinality serialization based on uniq statistics",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/85145 --> This PR adds automatic `LowCardinality` serialization for `MergeTree`, analogous to the existing automatic `Sparse` serialization. A `String` or `FixedString` column whose declared type is **not** `LowCardinality` can now be stored on disk in dictionary-encoded form when it has a `uniq` statistic with a low cardinality estimate. The encoding is transparent: the column keeps its declared data type, flows through the query pipeline as a (non-native) `ColumnLowCardinality`, and is materialized to a full column at the boundaries that require it. **Motivation:** save disk space and read I/O for low-cardinality string columns without requiring users to declare the column type as `LowCardinality(...)`. **How it works** - A new per-part serialization kind `LOW_CARDINALITY` (shown as `LowCardinality` in `system.parts_columns.serialization_kind`). - The decision is made at write time from the **existing** `uniq` statistic — there are **no changes to statistics serialization**. In `MergeTreeDataWriter`, a `String`/`FixedString` column whose `uniq` estimate does not exceed the new `MergeTree` setting `max_uniq_number_for_low_cardinality` (default `0` = disabled), and which was not already chosen for `Sparse`, is stored with `SerializationLowCardinality`. The column must therefore declare `STATISTICS(uniq)` and have `materialize_statistics_on_insert` enabled. - A new `is_native` flag on `ColumnLowCardinality` makes `getDataType` report the nested type, so the \"column class matches data type\" invariant holds while such a column flows through the pipeline. It is materialized to a full column in `removeSpecialRepresentations` (aggregation, joins, squashing, blocks), in the merge-input step of `IMergingAlgorithm`, in the sorting and sparse-removing transforms, and in `IFunction`. The `optimize_functions_to_subcolumns` pass skips columns stored as non-native `LowCardinality`, since the dictionary encoding does not expose the data type's regular subcolumns. - The kind is independent of sparse serialization: `SerializationInfo` entries for eligible columns are created in the writer, in merges, in mutations and in the per-table serialization hints even when `ratio_of_defaults_for_sparse_serialization = 1` disables `Sparse` entirely. **Notes** - Opt-in: disabled by default and requires an explicit `uniq` statistic, so existing tables are unaffected. Requires `allow_experimental_statistics`. - Reading a subcolumn of an encoded column (for example `s.size`) is not supported and reports `NOT_IMPLEMENTED`: the subcolumn's streams do not exist in a dictionary-encoded part. Making it read the full column and extract the subcolumn in memory needs a dedicated serialization and is left for a separate change. - Extracts and reimplements the automatic `LowCardinality` part of the draft #85145 against current `master`, without the statistics-serialization refactoring. That refactoring is still unmerged (#85145 and #109128 are open drafts). Validated locally with a server built from this branch: kind selection for `Wide` and `Compact` parts, `String` and `FixedString`, query correctness, functions, `GROUP BY`/`ORDER BY`/`DISTINCT`/`LIMIT BY`/window functions, joins, `IN` subqueries, `INSERT SELECT`, `UNION ALL`, `OPTIMIZE FINAL` (LowCardinality preserved through merges), mutations, `ALTER RENAME`/`MODIFY COLUMN`, `DETACH`/`ATTACH`, `PREWHERE`, input/output formats with and without `allow_special_serialization_kinds_in_output_formats`, and tables with mixed `Default` and `LowCardinality` active parts. The three tests also pass with `ratio_of_defaults_for_sparse_serialization` forced to `0.0` and `1.0`, with `enable_block_offset_column`/`enable_block_number_column`, with compact parts, and over repeated fully randomized runs. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added automatic `LowCardinality` serialization: a `String`/`FixedString` column that has a `uniq` statistic and low cardinality can be stored in dictionary-encoded form on disk while keeping its declared data type. Controlled by the new `MergeTree` setting `max_uniq_number_for_low_cardinality` (disabled by default). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107436",
        "createdAt": "2026-06-14T04:04:43Z",
        "updatedAt": "2026-08-13T04:27:33Z",
        "timestamp": "2026-08-13T04:27:33Z",
        "metrics": {
          "reactions": 1,
          "comments": 6
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107567",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallelize listing of globbed `s3` table function paths",
        "text": "Speeds up listing files for the `s3` table function over globbed paths, in two complementary ways. The listing was a single serial stream of `ListObjectsV2` requests paginated by a continuation token, bound by the per-request latency of S3 (~75 ms measured against `clickhouse-public-datasets`); this dominates queries over buckets with many objects. An earlier attempt (#66504) used speculative keyspace splitting with fixed split points assuming a uniform key distribution and a fixed character set, and was closed as fragile. Controlled by the new setting `s3_list_object_parallelism` (default `1` = the previous serial behavior). **1. Hierarchical layouts** — walk the \"directory\" tree formed by the `/` delimiter with several worker threads in parallel, pruning subtrees that cannot contain a matching key (so for partitioned layouts like `year=*/month=*/` fewer objects are fetched, not merely listed in parallel). Recursive globs (`**`) keep serial listing. **2. Big flat directories** (e.g. files named by hash/UUID — `musicbrainz/mlhdplus-complete/<uuid>.txt.zst`, ~600k files in one prefix) — when a listed prefix is truncated but has no sub-directories, split it by keyspace: tile the remaining range into contiguous `(start_after, end]` sub-ranges listed concurrently, with boundaries derived from the byte alphabet observed in the listed page. Because the sub-ranges are contiguous they tile the interval **with no gaps** — a key whose byte is absent from the sampled alphabet still falls into the range that brackets it — so the split is complete regardless of the key distribution or character set. The split is a single level and is **probe-gated**: before splitting, one cheap `ListObjectsV2` checks whether any key exists beyond the current bucket; if not (keys share a common prefix, e.g. `pageviews-YYYYMMDD-HH`), the directory is paginated serially instead of issuing a fan of empty boundary probes. This is the robust counterpart to #66504 — it never regresses for clustered keys. On S3 Express / directory buckets — which reject `StartAfter` and only allow the `/` delimiter — keyspace splitting is disabled automatically (detected via `isS3ExpressBucket`), so flat ranges are paginated serially while hierarchical layouts are still listed in parallel. Correctness of glob matching is guaranteed by the existing per-file `RE2::FullMatch`, so the directory pruning only needs to be conservative. **Measured against real S3** (`clickhouse-public-datasets`, ~75 ms/list-request) across ~16 datasets — all return identical results (same count, no duplicates, identical path checksum), serial vs parallel: - Big flat, uniform keys — `musicbrainz/mlhdplus-complete` (594415 files): **37s → 3.7s (~10x)**. - Big flat, clustered keys — `wikistat` (97529) / `gharchive` (100091): probe sends them serial; parallel matches serial (no regression). - Hierarchical — `web/store`: ~2.5x (bounded by its small directory count; scales toward the parallelism factor for larger trees). - Small/medium directories (`nyc-taxi`, `ontime`, `hits/native`, `sensors`, `tranco`, `youtube`, `adsblol`, `bluesky`, …): correct, no regression. A unit test (`gtest_object_storage_parallel_listing`) asserts every key is produced exactly once across parallelism 1–64, including adversarial cases (bytes outside the sampled alphabet, shared long prefixes, boundary keys, mixed hierarchical+flat). Stateless tests (`04339_s3_parallel_glob_listing`, `04340_s3_parallel_flat_listing`) verify parallel and serial listing return identical results. Closes: https://github.com/ClickHouse/ClickHouse/issues/65572 Related: https://github.com/ClickHouse/ClickHouse/pull/66504 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a setting `s3_list_object_parallelism` to list the files of the `s3` table function over a globbed path in parallel. Hierarchically partitioned layouts are listed by walking the tree of common prefixes concurrently; a single big flat directory is listed by splitting its keyspace into contiguous sub-ranges. This speeds up queries over buckets with many objects (e.g. ~10x for a directory of 600k files).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107567",
        "createdAt": "2026-06-15T21:43:46Z",
        "updatedAt": "2026-08-13T11:15:29Z",
        "timestamp": "2026-08-13T11:15:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 32
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107637",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Merge filters into join during join reordering",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Added a new setting `query_plan_merge_filters_into_join` that allows merging `Filter` steps into the `JOIN` step during join reordering, so `WHERE` predicates participate in join order optimization",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107637",
        "createdAt": "2026-06-16T15:25:08Z",
        "updatedAt": "2026-08-13T18:01:48Z",
        "timestamp": "2026-08-13T18:01:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "vdimir",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107669",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `input_format_json_max_object_size` setting to limit JSON object size on parsing",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `input_format_json_max_object_size` setting to limit JSON object size on parsing. Closes https://github.com/ClickHouse/ClickHouse/issues/106704",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107669",
        "createdAt": "2026-06-16T20:49:08Z",
        "updatedAt": "2026-08-13T17:55:50Z",
        "timestamp": "2026-08-13T17:55:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "Avogar",
        "state": "open",
        "assignees": [
          "Manerone"
        ],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107865",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add per-user filesystem cache disk usage metrics",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/105020 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added opt-in Prometheus gauges for current filesystem cache usage (`filesystem_cache_size_bytes` and `filesystem_cache_elements`), labeled by `cache_name` and `user_id`. The gauges are exposed through `system.dimensional_metrics` and the Prometheus endpoint and controlled by the `expose_prometheus_cache_usage_metrics_per_user` cache setting, which is disabled by default. --- ### Description Adds two gauges exposed through `system.dimensional_metrics` and the Prometheus endpoint: | Name | Labels | |---|---| | `filesystem_cache_size_bytes` | `cache_name`, `user_id` | | `filesystem_cache_elements` | `cache_name`, `user_id` | **Motivation.** Operators can identify which users currently occupy filesystem cache space, measured both in bytes and in file segments, for each cache. **Mechanism.** Each enabled cache owns a `FileCacheUsageTracker` containing shared per-user atomic counters. Main cache-priority entries retain the corresponding counters and update them together with the existing cache size and element accounting. `ServerAsynchronousMetrics` periodically obtains a per-user snapshot through `getUsageStatPerClient` and updates the dimensional gauges. This keeps `DimensionalMetrics` updates out of cache mutation paths. Composite LRU, SLRU, and split-cache priorities share the same tracker, including during SLRU queue transitions. Counters use shared ownership so inactive users can be reclaimed safely. When a user no longer has cache entries and its counters are zero, the next snapshot removes it from the tracker. Stale dimensional metric label combinations are also removed. **Sampling.** The metrics reflect the most recent asynchronous metrics update. **Cardinality.** In-process cardinality is bounded by users currently retained by each cache, plus labels awaiting the next asynchronous metrics update. Stale `(cache_name, user_id)` label combinations are removed. The feature remains disabled by default because enabling per-user metrics can still create significant Prometheus time-series cardinality. ### Documentation entry: - [x] Documentation is updated.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107865",
        "createdAt": "2026-06-18T14:22:58Z",
        "updatedAt": "2026-08-13T17:48:49Z",
        "timestamp": "2026-08-13T17:48:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "sacheendra",
        "state": "open",
        "assignees": [
          "kssenii"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107943",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix NOT_IMPLEMENTED exception in system.detached_tables for DatabaseDictionary and similar engines",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/104868 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry: Fixed `SELECT * FROM system.detached_tables` throwing `NOT_IMPLEMENTED` exception when databases like `DatabaseDictionary`, `DatabaseOverlay`, `DatabaseFilesystem`, or `DatabaseS3` exist. Such databases are now silently skipped during iteration. ### Problem `SELECT * FROM system.detached_tables` fails with `Cannot get detached tables for DatabaseDictionary. (NOT_IMPLEMENTED)` when any Dictionary or similar database exists on the server. Root cause: `getDetachedTablesIterator` was called without error handling in both `StorageSystemDetachedTables.cpp` and `StorageSystemTables.cpp`. Databases like `DatabaseDictionary`, `DatabaseOverlay`, `DatabaseFilesystem`, `DatabaseHDFS`, and `DatabaseS3` do not override this method and throw `NOT_IMPLEMENTED`. ### Fix Wrapped `getDetachedTablesIterator` calls in `try-catch` blocks in both affected files. Databases that throw `NOT_IMPLEMENTED` are now silently skipped, allowing the query to return results from supported databases. ### Test Added stateless test `04357_system_detached_tables_not_implemented` covering: - `DatabaseDictionary` coexisting with a normal `Atomic` database - Detached table correctly visible from normal database - No exception thrown when unsupported database engines exist",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107943",
        "createdAt": "2026-06-19T09:03:36Z",
        "updatedAt": "2026-08-13T09:34:35Z",
        "timestamp": "2026-08-13T09:34:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 16
        },
        "labels": [
          "pr-bugfix",
          "manual approve",
          "can be tested"
        ],
        "author": "adityaksolves",
        "state": "open",
        "assignees": [
          "tuanpach"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:107960",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cache vended credentials for REST catalogs",
        "text": "Now, new vended credentials are requested on each metadata request. This PR adds an (optional) cache for creds with configurable TTL. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add option to cache vended credentials for REST catalogs; add a setting `vended_credentials_cache_ttl` (seconds). 300 by default. 0 means no caching.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/107960",
        "createdAt": "2026-06-19T12:17:13Z",
        "updatedAt": "2026-08-13T04:40:44Z",
        "timestamp": "2026-08-13T04:40:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-improvement",
          "manual approve",
          "can be tested"
        ],
        "author": "zvonand",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108017",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add explicit DB::NsyncSharedMutex and use it for KeeperLogStore changelog locking",
        "text": "### Description We would like to try [google/nsync](https://github.com/google/nsync) mutex for ClickHouse. We used the existing microbenchmarks and did a comparison: - nsync: https://pastila.nl/?01726aec/f309fa62466e02f4514ab51b0166be86#jxPxN6DPD2sau29I1XHTmw==GCM - current: https://pastila.nl/?000ffb86/366fc7712e4de4898668d9c7ebee6d9b#nM+GdsaPsVmmynDDP8vVIQ==GCM **Readers-only: current SharedMutex wins** | Threads | `Self` current | `Nsync` | Winner | |---:|---:|---:|---| | 32 | 66.7755 ns | 160.620 ns | `Self` | | 64 | 69.9920 ns | 151.107 ns | `Self` | | 128 | 67.0653 ns | 150.980 ns | `Self` | So `NsyncSharedMutex` is worse for pure read-heavy workloads. That confirms we should not replace DB::SharedMutex globally. **Writers-only: NsyncSharedMutex wins strongly** | Threads | `Self` current | `Nsync` | Winner | |---:|---:|---:|---| | 32 | 466.074 ns | 56.4393 ns | `Nsync` | | 64 | 461.444 ns | 56.1529 ns | `Nsync` | | 128 | 460.664 ns | 57.4841 ns | `Nsync` | `Nsync` is about **8x faster** for contended writers. **Mixed read/write: NsyncSharedMutex wins strongly** | Threads | `Self` current | `Nsync` | Winner | |---:|---:|---:|---| | 32 | 268.006 ns | 97.0517 ns | `Nsync` | | 64 | 291.993 ns | 92.9774 ns | `Nsync` | | 128 | 325.001 ns | 94.2072 ns | `Nsync` | `Nsync` is about **3.45x faster** for mixed read/write contention. Since `KeeperLogStore::changelog_lock` has a mixed shared/exclusive access pattern, we selected it as the candidate to try. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Use google/nsync for KeeperLogStore changelog locking",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108017",
        "createdAt": "2026-06-20T05:22:29Z",
        "updatedAt": "2026-08-13T08:42:57Z",
        "timestamp": "2026-08-13T08:42:57Z",
        "metrics": {
          "reactions": 1,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "submodule changed",
          "can be tested"
        ],
        "author": "chhetripradeep",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108078",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Re-land: make Ctrl+C terminate the output of a result set in the client promptly",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/22426 Related: https://github.com/ClickHouse/ClickHouse/pull/106430 Related: https://github.com/ClickHouse/ClickHouse/pull/107845 Related: https://github.com/ClickHouse/ClickHouse/pull/107828 ### Description This re-lands #106430 (\"Make Ctrl+C terminate the output of a result set in the client promptly\"), which was reverted by #107845, and folds in the data-race fix from the (now closed) #107828 so the feature can land safely this time. Why it was reverted: `BuzzHouse (amd_tsan)` reported a ThreadSanitizer data race on master (STID 5296-4b74) on the client's persistent stdout buffer (`std_out`, an `AutoCanceledWriteBuffer<WriteBufferFromFileDescriptor>`). The interruptible-output feature installed a cancellation hook on `std_out` at the start of `processOrdinaryQuery` and cleared it via `SCOPE_EXIT({ std_out->setCancellationHook({}); })` at the end of the same scope. When the result set is written with parallel formatting, `ParallelFormattingOutputFormat` runs a background collector thread that writes through `std_out` and reads its `cancellation_hook` in `WriteBufferFromFileDescriptor::nextImpl`. On the normal end-of-stream path the collector is joined by `output_format->finalize` before the scope guard runs, so the clear is safe; but on paths that leave `processOrdinaryQuery` by an exception without first finalizing the output (a `NetException`, or a `LocalFormatError` rethrown from `receiveResult`), the scope guard cleared the hook while the collector thread was still alive and reading it — a data race on the `std::function`. The fix (from #107828): the hook is cleared in `resetOutput` right after `output_format.reset()`, which destroys the parallel formatter and joins its collector thread, so the mutation is single-threaded with respect to `std_out`'s hook. The hook is also re-pointed for the teardown flushes that follow: on normal completion the interrupt handler is still armed and is honored, so a fresh Ctrl+C keeps interrupting a flush to a slow or stuck stdout; on exception paths the handler is already stopped (its `cancelled` is then unconditionally true), so the genuine cancellation flag is used instead to avoid discarding already-produced output. `resetOutput` runs on every exit path and at the start of the next query, so `std_out` (which outlives a single query) never carries a stale hook into the next query. The pager and `INTO OUTFILE ... AND STDOUT` buffers are per-query and destroyed in `resetOutput` after the collector join, so they were never affected. The change is split into two commits: the first reapplies #106430 verbatim (a revert of the revert #107845); the second folds in the #107828 fix, confined to `src/Client/ClientBase.cpp`. The low-level `WriteBufferFromFileDescriptor` (used server-wide on a hot single-threaded write path) is unchanged by the fix. Revert that motivated this: https://github.com/ClickHouse/ClickHouse/pull/107845 TSAN report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=73e97deb9abbc3258e75d297aabfc3fbe7c3ca5e&name_0=MasterCI&name_1=BuzzHouse%20%28amd_tsan%29&name_1=BuzzHouse%20%28amd_tsan%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Pressing Ctrl+C in `clickhouse-client`/`clickhouse-local` now promptly stops printing a large result set, even when the output is being written to a slow or stuck destination such as a slow terminal. Previously the first Ctrl+C could appear to have no effect. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108078",
        "timestamp": "2026-08-12T22:46:17Z",
        "metrics": {
          "reactions": 0,
          "comments": 30
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108090",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support DEFAULT expressions inside Tuple data types",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/2797 `DEFAULT` expressions are now supported for named elements of `Tuple` data types, e.g.: ```sql CREATE TABLE t (id UInt8, c Tuple(a UInt8, s String DEFAULT 'Hello')) ENGINE = MergeTree ORDER BY id; ``` Such defaults exist only at the syntax level: they are pulled up to the column level (for `CREATE TABLE` as well as `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN`), so the stored schema is normalized without any `DEFAULT` inside the data type. The column above is stored as type `Tuple(a UInt8, s String)` with a column-level `DEFAULT tuple(defaultValueOfTypeName('UInt8'), 'Hello')`. A default expression may reference other columns, but not other elements of the same tuple/nested; a reference that collides with an element name is rejected as ambiguous. Building an actual data type while a default is set throws. `DEFAULT` inside `Nested` (which is `Array(Tuple(...))`) or `Array` is not supported yet and is rejected with a clear `NOT_IMPLEMENTED` error, because a scalar element default cannot be represented as a static array column default. Note this is why issue #2797 (which asks specifically for `DEFAULT` on `Nested` elements) is linked as `Related` rather than `Closes`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `DEFAULT` expressions inside `Tuple` data types, e.g. `Tuple(a UInt8, s String DEFAULT 'Hello')`. The default is normalized away from the type and pulled up to the column level. Works for `CREATE TABLE` and `ALTER TABLE ... ADD`/`MODIFY COLUMN`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108090",
        "createdAt": "2026-06-22T01:17:05Z",
        "updatedAt": "2026-08-13T17:13:35Z",
        "timestamp": "2026-08-13T17:13:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108329",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Bump `libssh` to 0.12.0",
        "text": "Update the `contrib/libssh` submodule to the upstream release `libssh-0.12.0` (from an upstream master snapshot, `libssh-0.11.0-368-g47305a2f`), tracked via the `ClickHouse/libssh-0.12.0` branch of the fork mirror. Both the old and new pins are clean upstream commits, so no ClickHouse patches needed to be carried forward. Build integration changes in `contrib/libssh-cmake`: - Bump the `libssh_VERSION_*` variables to 0.12.0. - Add the new ML-KEM sources. In 0.12.0, `kex.c` references `ssh_client_hybrid_mlkem_remove_callbacks` from an unguarded switch case, so `hybrid_mlkem.c` and `mlkem.c` are now mandatory. ClickHouse bundles OpenSSL 3.5, which provides the EVP ML-KEM API, so the OpenSSL backend `mlkem_crypto.c` is used (matching upstream's `OPENSSL_VERSION >= 3.5.0` logic). - In every per-platform `config.h`: define `GLOBAL_CONF_DIR` (now used unconditionally via string concatenation in `options.c`), and enable `HAVE_OPENSSL_MLKEM` / `HAVE_MLKEM1024` to select the OpenSSL ML-KEM backend and the ML-KEM-1024 variants. ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Update `libssh` to 0.12.0. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.1.24` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108329",
        "createdAt": "2026-06-23T20:48:38Z",
        "updatedAt": "2026-08-13T12:00:33Z",
        "timestamp": "2026-08-13T12:00:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-build",
          "submodule changed",
          "pr-synced-to-cloud"
        ],
        "author": "thevar1able",
        "state": "closed",
        "assignees": [
          "Algunenano"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108336",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Release pull request for branch 26.6",
        "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108336",
        "createdAt": "2026-06-23T22:33:53Z",
        "updatedAt": "2026-08-13T15:55:58Z",
        "timestamp": "2026-08-13T15:55:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "release"
        ],
        "author": "robot-clickhouse",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108371",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Array subscript operator supports array of integers as index.",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/108095 The array subscript operator now accepts an array of integers as the index, so `arr[indexes]` gathers the elements at all of those positions at once. It is equivalent to `arrayMap(i -> arr[i], indexes)`, including the result type, the handling of negative indexes and the value produced for an out-of-range position, but it has its own implementation: a constant source array is not materialized per row, and numeric element types are gathered through the `PODArray` directly. ```sql SELECT [10, 20, 30, 40][[2, 4, 1]]; -- [20,40,10] SELECT [10, 20, 30][[1, 5, -1]]; -- [10,0,30] SELECT arrayElementOrNull([10, 20, 30], [1, 5]); -- [10,NULL] SELECT [10, 20, 30][[1, NULL, 2]]; -- [10,NULL,20] ``` The positions may be nullable, and a `NULL` position produces `NULL`, just as for a scalar index. The main use case is a lookup table: a constant dictionary array indexed by a per-row array of positions. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The array subscript operator supports an array of integers as the index: `arr[indexes]` returns the elements at all of the given positions, equivalently to `arrayMap(i -> arr[i], indexes)`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108371",
        "createdAt": "2026-06-24T12:15:02Z",
        "updatedAt": "2026-08-13T14:13:55Z",
        "timestamp": "2026-08-13T14:13:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "ucasfl",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108468",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `merge_use_batch_sorting_queue` `MergeTree` setting for ordinary merges",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/38022 Related: https://github.com/ClickHouse/ClickHouse/pull/38859 Adds `merge_use_batch_sorting_queue`, a Boolean `MergeTree` setting for ordinary `MergeTree` merges. It has no effect on merges that change rows, such as those performed by `ReplacingMergeTree` or `AggregatingMergeTree`. The setting defaults to `false`, preserving the existing merge behavior. When set to `true`, it uses the batch sorting queue in the ordinary merge sorting stage. The batch sorting queue is already used in several query sorting paths. This change makes it available for background merges as an opt-in table setting, so users can enable it for merge-heavy workloads where `MergingSortedTransform` is a significant part of merge time. The PR also adds targeted stateless coverage for: - `merge_use_batch_sorting_queue = false` vs `true` correctness across horizontal and vertical merges with a broad set of data types; - vertical merges with small granules and larger merge blocks; - equal primary-key values spread across multiple source parts, checked by merged-part order. Benchmark results from local runs with `merge_use_batch_sorting_queue` disabled and enabled: | Dataset | Wall time speedup | Wall time disabled | Wall time enabled | User CPU change | Peak RSS disabled | Peak RSS enabled | `MergeTotalMilliseconds` change | `MergingSortedMilliseconds` change | |---|---:|---:|---:|---:|---:|---:|---:|---:| | ClickBench hits | 23.9% | 10.270s | 7.813s | -28.3% | 1.38 GiB | 1.36 GiB | -27.1% | -69.2% | | OnTime 2019 | 35.7% | 6.855s | 4.410s | -41.4% | 1.79 GiB | 1.83 GiB | -31.5% | -91.4% | | StackOverflow votes | 13.0% | 0.730s | 0.635s | -17.4% | 374 MiB | 372 MiB | -14.4% | -82.2% | | HackerNews | 2.8% | 3.565s | 3.465s | -3.1% | 798 MiB | 803 MiB | -0.9% | -74.8% | | NYC Taxi | 2.4% | 3.505s | 3.420s | -1.9% | 998 MiB | 1000 MiB | -2.2% | -5.2% | | StackOverflow posts | 2.7% | 4.275s | 4.160s | -2.0% | 1.07 GiB | 1.08 GiB | -2.3% | -32.8% | The benefit is workload dependent. It is largest when sorted merging is a significant fraction of the total merge time, and smaller when writing, compression, or other merge work dominates. This PR was developed with AI assistance from Codex and Claude. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `merge_use_batch_sorting_queue` `MergeTree` setting to optionally use the batch sorting queue for ordinary `MergeTree` merges, reducing CPU and wall-clock time for merge workloads where sorted merging is a significant cost.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108468",
        "createdAt": "2026-06-25T12:18:51Z",
        "updatedAt": "2026-08-13T09:00:02Z",
        "timestamp": "2026-08-13T09:00:02Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-performance",
          "manual approve",
          "can be tested"
        ],
        "author": "rorylshanks",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108522",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add system.s3(azure)_queue_metadata",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add system.s3(azure)_queue_metadata. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1316` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108522",
        "createdAt": "2026-06-25T16:42:44Z",
        "updatedAt": "2026-08-13T09:54:17Z",
        "timestamp": "2026-08-13T09:54:17Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-improvement",
          "pr-synced-to-cloud"
        ],
        "author": "kssenii",
        "state": "closed",
        "assignees": [
          "bharatnc"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108642",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix DateTime64 parser consumes Bool column value when small epoch",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes parsing for DateTime64 on small values. Closes https://github.com/ClickHouse/ClickHouse/issues/101487",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108642",
        "timestamp": "2026-08-12T21:10:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yariks5s",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108653",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support `GROUPS` frame mode for window functions",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> The query below applies the same `1 PRECEDING AND 1 FOLLOWING` bounds as a `ROWS`, a `RANGE`, and a `GROUPS` (this PR) frame. The `order` column contains duplicate and non-consecutive values, so the three modes cover different rows: ```sql CREATE TABLE wf_frame_groups (`order` UInt64, value UInt64) ENGINE = Memory; INSERT INTO wf_frame_groups FORMAT Values (10, 1), (10, 2), (20, 3), (30, 4), (30, 5); SELECT order, value, groupArray(value) OVER (ORDER BY order ROWS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS rows_frame, groupArray(value) OVER (ORDER BY order RANGE BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS range_frame, groupArray(value) OVER (ORDER BY order GROUPS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS groups_frame FROM wf_frame_groups ORDER BY order, value; ``` ```response ┌─order─┬─value─┬─rows_frame─┬─range_frame─┬─groups_frame─┐ │ 10 │ 1 │ [1,2] │ [1,2] │ [1,2,3] │ │ 10 │ 2 │ [1,2,3] │ [1,2] │ [1,2,3] │ │ 20 │ 3 │ [2,3,4] │ [3] │ [1,2,3,4,5] │ │ 30 │ 4 │ [3,4,5] │ [4,5] │ [3,4,5] │ │ 30 │ 5 │ [4,5] │ [4,5] │ [3,4,5] │ └───────┴───────┴────────────┴─────────────┴──────────────┘ ``` Each mode interprets the bounds differently: - `ROWS` counts physical rows, so the frame is at most three adjacent rows: the current row plus one on each side. - `RANGE` counts `order` values, so `1 PRECEDING` and `1 FOLLOWING` cover rows whose `order` is within 1 of the current row's. With gaps of 10, no neighbouring row qualifies, so the frame holds only the rows that share the current `order`. - `GROUPS` (added by this PR) counts peer groups, so `1 PRECEDING` and `1 FOLLOWING` always include the adjacent groups in full, whatever the gaps between `order` values. cc: @cwurm ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `GROUPS` frame mode for window functions (SQL:2011), e.g. `any(price) OVER (PARTITION BY symbol ORDER BY ts GROUPS BETWEEN CURRENT ROW AND 1 FOLLOWING)`. In a `GROUPS` frame the boundaries count whole peer groups — sets of rows that are equal on the `ORDER BY` key — so `N PRECEDING`/`N FOLLOWING` mean `N` peer groups before/after the current row's peer group, rather than physical rows (`ROWS`) or `ORDER BY` value distances (`RANGE`). <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1354` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108653",
        "createdAt": "2026-06-26T20:45:31Z",
        "updatedAt": "2026-08-13T17:53:45Z",
        "timestamp": "2026-08-13T17:53:45Z",
        "metrics": {
          "reactions": 2,
          "comments": 3
        },
        "labels": [
          "pr-feature"
        ],
        "author": "nihalzp",
        "state": "closed",
        "assignees": [
          "antaljanosbenjamin"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108721",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Detect when tables behind a query have changed",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/108713 Adds a way to detect when the data behind a query has changed, and uses it to make the query cache consistent and to skip unnecessary refreshes of materialized views. A new virtual method `getModificationHash` on `IStorage` returns an optional `UInt128` that changes whenever the data behind the table changes (similar to an HTTP ETag). It is not a hash of the data and the way it is computed is engine-specific. It returns `NULL` when the engine cannot give a usable value, so callers fail closed. Implemented for engines that can expose a loop-free (no-ABA) value - one that never returns to an earlier value across a change-and-change-back (an `A -> B -> A` transition): - `MergeTree` family: the block-number ranges of the active data parts, their content checksums, the table structure/key metadata, a per-lifetime counter that advances on every active-part-set change (so a drop that restores an identical part set does not reproduce an earlier hash), and a per-lifetime metadata version that advances on every metadata change (so a metadata `ALTER` and back does not either). Changes on insert, merge, mutation, `ALTER`. - `Memory`: identity of the current set of blocks plus the row count and a monotonic version. - `Log`, `TinyLog`, `StripeLog`: total rows and bytes plus structure and a monotonic version. - `Merge` and `Distributed`: combine the hashes of the underlying tables (`Distributed` asks each shard through `system.tables`), folding each table's identity (database, name, UUID) so two different tables with the same hash are distinguished; they fail closed if any underlying table does. - `View` and `MaterializedView` are looked through: a view hashes the tables behind its stored `SELECT` (plus its own UUID, columns, and security metadata), a materialized view hashes its target table. They fail closed for parameterized views and in databases without table UUIDs (`Ordinary`), where incarnations of a re-created view cannot be told apart. `URL` and object storage (`S3`, ...) trust the resource's strong (non-weak) `ETag` when one is exposed, and fail closed (report `NULL`) otherwise - for glob/failover patterns, a weak or absent `ETag`, or an unreachable source. The `ETag` is loop-free for content (the same `ETag` denotes the same content), but unlike the engines above there is no monotonic version to fold, so an `A -> B -> A` rewrite back to byte-identical content within a single query's read window can in principle repeat it. That residual is narrow and both consumers below are opt-in, so we keep these engines in the feature rather than drop them. `File` fails closed: its only change signals are size and modification time, both weak. It is exposed as a lazily-computed `modification_hash` column in `system.tables`, and used for two features: - New setting `query_cache_use_only_when_data_was_not_changed`: when enabled, the combined modification hash of the tables referenced by a query is folded into the query cache key, so a cached result is reused only while none of those tables changed. If consistency cannot be guaranteed (e.g. a query calling a non-deterministic function, whose result can change while every referenced table is unchanged; a table function; a `File` table; a `URL`/object-storage table without a strong `ETag`; an object-storage read pruned by a `_path`/`_file`/Hive-partition filter, where the consumed subset cannot be compared with the full listing; a query inside a transaction, which reads the transaction's snapshot while the hash samples the live table state; or a referenced table with an active row policy for the current user, which changes what the user reads while every referenced table is unchanged - row policies existing only on remote shard servers of a `Distributed` table cannot be seen and remain part of the best-effort window), the query cache is bypassed for that query. - `REFRESH ... IF CHANGED` for refreshable materialized views: a scheduled refresh is skipped when none of the tables the view reads from changed since the last refresh that rebuilt the view (e.g. `REFRESH EVERY 1 MINUTE IF CHANGED`). It always rebuilds when a source table cannot prove it is unchanged. The `clickhouse` binary was built and all new tests pass locally, including a functional check that `Distributed.modification_hash` changes when a remote table changes and that `IF CHANGED` skips refreshes while the source is unchanged. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a `modification_hash` column to `system.tables` that changes whenever the data behind a table changes. Based on it, added a setting `query_cache_use_only_when_data_was_not_changed` to make the query cache consistent (a cached result is reused only while the referenced tables are unchanged) and a `REFRESH ... IF CHANGED` option for refreshable materialized views to skip refreshes when the source data did not change.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108721",
        "createdAt": "2026-06-28T00:01:15Z",
        "updatedAt": "2026-08-13T16:59:40Z",
        "timestamp": "2026-08-13T16:59:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 56
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108786",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Switch the default compression to ZSTD(3) for table data and network",
        "text": "Switch the default compression in ClickHouse from `LZ4` to `ZSTD(3)`, for both on-disk table data and network communication. ### Motivation `LZ4` has been the default for a long time and optimizes for speed, but `ZSTD(3)` gives substantially better compression ratios at a still-modest CPU cost, reducing both storage footprint and network traffic out of the box. ClickHouse Cloud already uses a stronger default; this aligns the self-managed defaults. ### What changes - **On-disk table data.** `CompressionCodecFactory::getDefaultCodec` now returns `ZSTD(3)` instead of `LZ4`. For `MergeTree` column data the built-in default is *size-aware*: when a column specifies no codec and no `<compression>` server-config case matches, `CompressionCodecSelector::choose` uses the faster `LZ4` for parts smaller than 100 MB and `ZSTD(3)` for larger parts, so freshly inserted data starts as `LZ4` and the bigger parts produced by background merges switch to `ZSTD(3)`. The direct `getDefaultCodec` users — the `Log` family, part checksums, the HTTP `compress=1` output, the `StripeLog` data stream, and the persistent `Set`/`Join` files — are not size-aware and switch uniformly from `LZ4` to `ZSTD(3)`. External temporary (spill) files are unaffected: they keep their own `temporary_files_codec` default (`LZ4`). - **Network.** The defaults of `network_compression_method` (`LZ4` → `ZSTD`) and `network_zstd_compression_level` (`1` → `3`) are changed, so the native client/server and Distributed (server/server) protocol uses `ZSTD(3)` by default. The changes are recorded in `SettingsChangesHistory.cpp` so the `compatibility` setting restores the previous behavior for these paths. The distributed-query streaming exchange (`StreamingExchangeSink`) does not read `network_compression_method` and always uses the server default codec; this is safe because each frame is self-describing (the receiver auto-detects the codec) and the exchange is a transient, same-version channel — `StreamingExchangeProtocol` rejects peers on a different protocol version, so a stream is never read back by a node expecting a different codec. - **Upgrade safety.** The append-only `Log`/`TinyLog`/`StripeLog` engines resolve the default codec at write time, so a table written before the upgrade (`LZ4`) and appended to after it (`ZSTD(3)`) ends up with mixed-codec blocks in one file. Their readers now pass `allow_different_codecs = true` (as the `MergeTree` reader already does) so such files still read back correctly. The legacy custom-frame Keeper snapshot format (`compress_snapshots_with_zstd_format = false`) is pinned to `LZ4` so it stays the documented format. Columns, tables, and connections that specify a codec or method explicitly are unaffected — with one exception: streams written through the built-in default codec directly, which ignore per-column codecs and the `<compression>` config. The `StripeLog` data stream, the persistent `Set`/`Join` backup files (written by `SetOrJoinSink` and `StorageJoin::mutate` through a `CompressedWriteBuffer` with no codec), and `MergeTree` auxiliary streams such as part checksums (`checksums.txt`, via `MergeTreeDataPartChecksums::write`) always follow the new default; see the changelog entry below for the per-path rollback. ### Validation Built locally and verified against the new binary: - A `MergeTree` table created with no explicit codec follows the size-aware default: small parts report `LZ4` and parts larger than 100 MB report `ZSTD(3)` as their `default_compression_codec` in `system.parts`. Direct `getDefaultCodec` streams (e.g. `StripeLog`) report `ZSTD(3)` regardless of size. - `system.settings` shows `network_compression_method = ZSTD` and `network_zstd_compression_level = 3`. - End-to-end `SELECT`/`INSERT` over the native protocol with default (now `ZSTD`) network compression work; all of `LZ4`/`lz4hc`/`zstd`/`none` still work and an invalid method still errors with `BAD_ARGUMENTS`. - The `02995_new_settings_history` consistency check passes (both setting changes are recorded). - The on-disk byte-dump tests (`02047_log_family_*_data_file_dumps`) were regenerated for the new default: `Log`/`TinyLog` keep `LZ4`-pinned columns, while `StripeLog`'s shared `data.bin`/`index.mrk` reflect the `ZSTD(3)` default (it uses `getDefaultCodec` and ignores the per-column codec). Size-sensitive stateless and integration tests that assert compressed byte sizes were pinned back to `LZ4` (cache segment sizes, distributed-batch corruption offsets, full-disk thresholds, frozen-part checksums). The CI performance comparison does **not** measure this tradeoff: the harness pins the built-in default codec to `LZ4` (`tests/performance/scripts/config/config.d/compression.xml`) and `network_compression_method` to `LZ4` (`tests/performance/scripts/config/users.d/perf-comparison-tweaks-users.xml`) on *both* the reference and the tested server, deliberately, so the report measures query logic rather than drowning in expected codec regressions. It therefore serves only as a regression filter for the non-codec parts of this change; the storage/network-vs-CPU tradeoff of `LZ4` → `ZSTD(3)` itself is a well-known property of the two codecs and is not measured by this pull request's CI. Documentation updated accordingly, including the `Native` format and protocol specifications and the `<compression>` config examples. ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The default compression method is changed from `LZ4` to `ZSTD(3)`. For client/server and server/server network communication it switches uniformly to `ZSTD(3)`. For on-disk table data, `MergeTree` column data now uses a size-aware built-in default — parts smaller than 100 MB use `LZ4` and larger parts use `ZSTD(3)` — while the direct built-in-default streams (the `StripeLog` data stream, the persistent `Set`/`Join` files, the HTTP `compress=1` framed output, and part checksums) switch uniformly from `LZ4` to `ZSTD(3)`. This improves compression ratios and reduces storage and network usage out of the box, at a modest increase in CPU usage. None of the on-disk or HTTP paths are controlled by `compatibility`; the previous behavior can be restored per path as follows. For client/server and Distributed network compression, use the `compatibility` setting (or set `network_compression_method = 'LZ4'`). For `MergeTree` column data, set a column/table `CODEC(LZ4)` or configure the server `<compression>` default to `LZ4`. For the `Log`/`TinyLog` engines, set a column `CODEC(LZ4)` (their default does not consult the `<compression>` config). The `StripeLog` data stream, the persistent `Set`/`Join` backup files, the HTTP `compress=1` framed output, and `MergeTree` auxiliary streams such as part checksums (`checksums.txt`) use the built-in default codec directly — they ignore per-column codecs and the `<compression>` config — so they always use the new `ZSTD(3)` default and have no runtime rollback; reading stays correct because every compressed frame is self-describing.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108786",
        "createdAt": "2026-06-29T02:28:41Z",
        "updatedAt": "2026-08-13T13:36:12Z",
        "timestamp": "2026-08-13T13:36:12Z",
        "metrics": {
          "reactions": 1,
          "comments": 69
        },
        "labels": [
          "pr-performance",
          "pr-backward-incompatible",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108820",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allow Distributed and Remote tables in Replicated databases",
        "text": "`database_replicated_allow_only_replicated_engine` rejects table engines that keep their own unreplicated data on disk in a `Replicated` database. Before this change, `Distributed` was also rejected because its optional background `INSERT` queue makes `storesDataOnDisk` return true, even though the queue is a transient send buffer and the table's actual data belongs to its destination shards. This change distinguishes table-owned on-disk data from auxiliary delivery state. The restriction applies to non-`ATTACH` `CREATE` queries when `database_replicated_allow_only_replicated_engine = 1`. | Engine or group | Result | Reason | |---|---|---| | `ReplicatedMergeTree` family and `SharedMergeTree` | Allow | The engine provides a replication or shared-storage contract. | | Writable non-replicated `MergeTree` with a local or remote storage policy | Reject | It owns unreplicated on-disk table data; a remote disk alone does not establish replication. | | Static read-only non-replicated `MergeTree` | Allow | It cannot create new table-owned data. | | `Log`, `TinyLog`, `StripeLog`, `Set`, `Join`, `EmbeddedRocksDB`, and database `File` | Reject | They own unreplicated on-disk table data. | | `MaterializedPostgreSQL` | Reject | It owns a local nested table. | | `Distributed`, `Remote`, and `RemoteSecure` | Allow | Their optional local queue is auxiliary delivery state rather than data of the table itself. | | `Memory`, `Buffer`, and `Null` | Allow | They do not own on-disk table data. | | `Merge`, `Alias`, and `View` | Allow | They are metadata-only. | | Views with inner tables | Depends | Each generated inner table is created and checked separately. | | External-storage engines, data lakes, external databases, and external queues | Allow | Their data is managed outside ClickHouse-owned table storage. | | Lazy `StorageTableProxy` | Reject conservatively | The nested storage is unknown without loading it. | `ATTACH` remains outside the existing gate. The `Distributed` background `INSERT` queue is local to the node accepting an insert and is not replicated. Users requiring acknowledgement only after data reaches the destination shards should set `distributed_foreground_insert = 1`. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow `Distributed`, `Remote`, and `RemoteSecure` tables in `Replicated` databases when `database_replicated_allow_only_replicated_engine` is enabled, while continuing to reject writable non-replicated `MergeTree` tables using local or remote storage policies.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108820",
        "createdAt": "2026-06-29T15:29:14Z",
        "updatedAt": "2026-08-13T16:32:56Z",
        "timestamp": "2026-08-13T16:32:56Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "UberDever",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:108862",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use `HashSet` for aggregations without aggregates",
        "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Aggregation queries without aggregate functions now use `HashSet`-based methods instead of `HashMap` (for supported key types). Speedups up to 1.8x times were observed. --- <img width=\"1015\" height=\"430\" alt=\"Screenshot 2026-07-03 at 00 34 25\" src=\"https://github.com/user-attachments/assets/15cc6a20-7568-49a4-94a5-854c7751ae02\" /> Further steps are `distinct` -> `group by` rewrite and key-columns-only (and perhaps semi) joins.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/108862",
        "createdAt": "2026-06-29T22:01:12Z",
        "updatedAt": "2026-08-13T11:42:45Z",
        "timestamp": "2026-08-13T11:42:45Z",
        "metrics": {
          "reactions": 1,
          "comments": 7
        },
        "labels": [
          "pr-performance"
        ],
        "author": "nickitat",
        "state": "open",
        "assignees": [
          "nihalzp"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109004",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Assign merges for all partitions at once for OPTIMIZE FINAL",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/46770 For a non-replicated `MergeTree` table, `OPTIMIZE TABLE ... FINAL` without an explicit partition used to process partitions one by one: it selected and ran the merge for one partition, waited for it to finish, and only then moved on to the next one. On a table with many partitions this serialized all the work and only ever showed a single merge in `system.merges`. Now the per-partition merges are assigned and executed in parallel, so all partitions are merged at once. The degree of parallelism is bounded by the configured background merge concurrency (`background_pool_size` * `background_merges_mutations_concurrency_ratio`). This mirrors how `StorageReplicatedMergeTree::optimize` already assigns a merge per partition and then waits for all of them. Notes: - Tables inside an explicit transaction keep the sequential path, because parallel merges would otherwise share a single transaction object that is not made for concurrent use. - The wait for already-running merges inside `selectPartsToMerge` (the `OPTIMIZE FINAL` path) is now scoped to the current partition: merges in other partitions cannot prevent selecting all the parts of this one, and waiting for them would needlessly serialize the parallel assignment. Verified on a local build: on a table with 8 partitions, `OPTIMIZE TABLE ... FINAL` now runs 8 merges concurrently (observed via `system.merges`) instead of one at a time, and every partition is still correctly merged into a single part (checked for `MergeTree`, `OPTIMIZE ... FINAL DEDUPLICATE`, and `ReplacingMergeTree`, plus the `optimize_skip_merged_partitions` / `optimize_throw_if_noop` no-op paths and two concurrent `OPTIMIZE FINAL` queries). ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `OPTIMIZE TABLE ... FINAL` on a non-replicated `MergeTree` table now assigns and runs the merges for all partitions at once instead of processing them one partition at a time.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109004",
        "createdAt": "2026-06-30T23:57:47Z",
        "updatedAt": "2026-08-13T17:51:12Z",
        "timestamp": "2026-08-13T17:51:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109005",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallel full sorting merge join (parallel_full_sorting_merge)",
        "text": "Resolves https://github.com/ClickHouse/ClickHouse/issues/48165 Adds a new `join_algorithm` value `parallel_full_sorting_merge`. `full_sorting_merge` streams both sides and joins them with a single merge, so it keeps memory bounded but runs the merge on one thread — on many-core machines it is often slower than `parallel_hash` even though it uses far less memory. `parallel_full_sorting_merge` keeps the streaming, low-memory profile of a merge join but shards the join by the hash of the join keys into independent per-shard merge joins that run in parallel (up to `max_threads`). A query-plan optimization (`optimizeParallelFullSortingMergeJoin`) switches each side's pre-join `SortingStep` to scatter the rows by the hash of the join keys into a fixed number of shards (`max_threads`) and sort each shard, then marks the join to run shard-by-shard, reusing the existing sharded pipeline (`joinPipelinesYShapedByShards`). Because the partitioning depends only on the join-key values — and `FullSortingMergeJoin` already requires matching key types — equal keys land in the same shard on both sides, so shards join independently. The result is unordered. Unlike the by-primary-key-ranges sharding (`query_plan_join_shard_by_pk_ranges`), this works on unsorted inputs by scattering each side before sorting. Already-sorted inputs are not scattered by this rewrite; they fall back to a single merge join, while in-order MergeTree reads can still be sharded at the source by primary-key ranges. When a side reads a single stream the scatter still produces the fixed shard count on both sides, so joins between inputs with different parallelism (e.g. different part counts) stay co-partitioned. Benchmark, `20M ⋈ 20M` on `UInt64` keys, `max_threads = 8`: | `join_algorithm` | wall | peak RSS | |---|---|---| | `parallel_hash` | 1.27 s | 3201 MB | | `parallel_full_sorting_merge` | **0.52 s** | **974 MB** | | `full_sorting_merge` (single merge) | 0.76 s | 552 MB | i.e. ~2.4× faster and ~3.3× less memory than `parallel_hash`, and ~1.5× faster than the single-threaded merge join. Correctness is verified against the `hash` algorithm for `INNER`/`LEFT`/`RIGHT`/`FULL` joins, many-to-many keys, and `join_use_nulls` in `04492_parallel_full_sorting_merge_join`. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new `join_algorithm` value `parallel_full_sorting_merge`: for hash-compatible equality joins, it shards a full sorting merge join by the hash of the join keys into independent per-shard merge joins running on all threads. It keeps the low, streaming memory usage of a merge join while parallelizing it (in a benchmark, ~2.4x faster and ~3.3x less memory than `parallel_hash`). `ASOF` joins fall back to a single `full_sorting_merge`; hash-incompatible key types (floating-point, `JSON`, `Object`, `Dynamic`) skip only the hash-scatter rewrite and can still be sharded at the source by primary-key ranges when `query_plan_join_shard_by_pk_ranges` is enabled. The result is not ordered. ### Documentation entry for user-facing changes - [x] Documentation is written (the new value is described in the `join_algorithm` setting). <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.592` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109005",
        "createdAt": "2026-07-01T00:38:58Z",
        "updatedAt": "2026-08-13T13:02:56Z",
        "timestamp": "2026-08-13T13:02:56Z",
        "metrics": {
          "reactions": 2,
          "comments": 37
        },
        "labels": [
          "pr-performance",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "m-selmi"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109130",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Stress test: do not force use_query_cache for non-throw overflow-mode tests",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/107907 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description The stress runner enables `use_query_cache=1` as a global per-runner client option (`ci/jobs/scripts/stress/stress.py`, probability `1/15`). The server refuses that together with a non-throw `*_overflow_mode` and raises `QUERY_CACHE_USED_WITH_NON_THROW_OVERFLOW_MODE` (error 731), because such a mode can truncate the result, which must never be cached (Bug 67476, guarded in `executeQuery.cpp`). Tests that set such a mode then fail. `00107_totals_after_having` sets `group_by_overflow_mode = 'any'`, and three consecutive failures trip `--max-failures-chain`, aborting the runner. 72 stateless tests set a non-throw overflow mode. Reported by @ alexey-milovidov while triaging a `Stress test (amd_msan)` failure on #107907 (report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107907&sha=f7e21726f54a9&name_0=PR). That block now also pins the overflow modes to `'throw'`, but only as a client option, so it does not cover this: a test's own `SET` overrides a client option. Verified against a local server: with `use_query_cache=1` and every mode pinned to `'throw'` on the command line, a query after `SET group_by_overflow_mode = 'any'` still returns 731. Fix: detect tests that set a non-throw overflow mode and turn the forced query cache back off for exactly those, on the command line and on the HTTP carrier used in cloud mode. The trailing value wins in both. The detected settings are the ten the server guards, so `read_overflow_mode_leaf` is covered; comment-only mentions are ignored. Query-level `SETTINGS` still override the override, so the intentional 731 coverage in `02494_query_cache_bugs` is preserved, and the query-cache coverage of every other test is unchanged. Treating 731 as benign in the runner was rejected: it would let a test's queries silently abort and could mask real 731 regressions.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109130",
        "createdAt": "2026-07-02T09:28:03Z",
        "updatedAt": "2026-08-13T13:14:34Z",
        "timestamp": "2026-08-13T13:14:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109225",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix wrong results with parallel_hash JOIN and read-in-order-through-join",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/issues/109216 Related: https://github.com/ClickHouse/ClickHouse/pull/110671 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results (silently dropped or mis-grouped rows) when `join_algorithm = 'parallel_hash'` is combined with a sorted consumer such as `optimize_aggregation_in_order`, `optimize_distinct_in_order` or `LIMIT BY`. With several join slots and a single-level hash map, `ConcurrentHashJoin` scatters the left block across slots, so the read-in-order-through-join optimization must no longer advertise the left sort order in that case. ### Description The `partial_merge` / `prefer_partial_merge` half of this PR has since been fixed on master by #110671, which added `IJoin::preservesLeftBlockOrder()` (defaulting to `true`) plus the `MergeJoin` / `JoinSwitcher` overrides and the `findReadingStep` gate. What remains here is a different carrier, and it is still a live wrong result on master. **`parallel_hash` (`ConcurrentHashJoin`) with several slots and a single-level map.** It inherits the `true` default from master, but `chooseMethod` leaves a key that materializes to one or two bytes (`key8` / `key16`) single-level - wider keys, including string and fixed-string ones, get a two-level variant. For a single-level map `joinBlock` scatters the left block across slots and `ConcurrentHashJoinResult` emits slot 0, then slot 1, and so on, so equal left-key values stop being contiguous while `findReadingStep` still installs the ordered read. Measured on master (`54dee101`, which contains #110671) against a debug build of this branch: ```sql CREATE TABLE t3 (a UInt32, j UInt8) ENGINE = MergeTree ORDER BY (a, j); CREATE TABLE t4 (j UInt8, v UInt64) ENGINE = MergeTree ORDER BY j; INSERT INTO t3 SELECT intDiv(number, 8)::UInt32, (number % 8)::UInt8 FROM numbers(64); INSERT INTO t4 SELECT (number % 8)::UInt8, number FROM numbers(8); SET join_algorithm = 'parallel_hash', max_threads = 8, optimize_aggregation_in_order = 1, max_bytes_before_external_join = 0, max_bytes_ratio_before_external_join = 0; SELECT a, count() FROM t3 LEFT ALL JOIN t4 ON t3.j = t4.j GROUP BY a ORDER BY a; ``` Master returns `1, 1, 1, 1, 1, 1, 1, 57`; the correct answer is 8 per group, which this branch returns. Ground truth was confirmed three independent ways (`optimize_aggregation_in_order = 0`, `join_algorithm = 'hash'`, `query_plan_read_in_order_through_join = 0`). The fix flips the `IJoin::preservesLeftBlockOrder()` default from `true` (fail-open) to `false` (fail-closed) and makes each join that really does stream the left side through once opt in: `HashJoin`, `DirectKeyValueJoin`, `ConstantJoin`, `PasteJoin` unconditionally, and `ConcurrentHashJoin` only when it does not scatter (`slots == 1 || twoLevelMapIsUsed()`). Flipping the default is what makes the contract hold by property rather than by accident. On master `FullSortingMergeJoin` has no override, so it inherits `true` - it is safe today only because `JoinStepLogical` inserts a `Sorting (Sort Left before JOIN)` step that `findReadingStep` does not descend through. That is a property of the current plan shape, not of the join, so any future plan change would silently reintroduce a wrong result. Under the fail-closed default it is safe by property. Precision was verified in both directions, so the stricter default does not cost the optimization anywhere it was previously correct: a two-level `UInt64` key still reads in order, a single-slot (`max_threads = 1`) `parallel_hash` join still reads in order, and `hash` / `direct` are unchanged. `GraceHashJoin` and `SpillingHashJoin` remain excluded through `hasDelayedBlocks()` as before. `topKThroughJoin.cpp` keeps its explicit `FullSortingMergeJoin` type check for its own mode 2 (a pre-JOIN `Sort` on the preserved input); its comment is updated to say the `preservesLeftBlockOrder()` read already covers that join and the type check is now belt-and-braces. Tests: `04498_distinct_in_order_partial_merge_join` fails on current master on exactly the `parallel_hash` block and passes here, so it is a live regression test rather than a restatement of #110671. `04500_read_in_order_through_constant_join` covers the `ConstantJoin` and `DirectKeyValueJoin` opt-ins, and `04500_limit_by_in_order_partial_merge_join` guards the `LIMIT BY` consumer. Each assertion was verified by mutation: with the corresponding override removed the assertion flips. All of them pin the whole read-in-order trio (`optimize_read_in_order`, `query_plan_read_in_order`, `query_plan_read_in_order_through_join`), since the stateless runner randomizes all three and a drawn `0` would make the plan assertions blind. #109216 is downgraded to `Related:` because the shape it reports (`prefer_partial_merge` + `optimize_distinct_in_order`) no longer reproduces on master after #110671; its reproducer now returns the correct 6 rows over repeated runs. This PR covers the sibling `parallel_hash` carrier of the same class, so it should not auto-close that issue.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109225",
        "createdAt": "2026-07-02T20:34:33Z",
        "updatedAt": "2026-08-13T14:45:46Z",
        "timestamp": "2026-08-13T14:45:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 42
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "vdimir"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109252",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Make JSONExtract honour cast_string_to_date_time_mode when parsing DateTime values",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `JSONExtract` now honours `cast_string_to_date_time_mode` when converting string JSON values to `DateTime`/`DateTime64`, consistently with `CAST`. Closes #109126. --- Extracting a string JSON value into `DateTime`/`DateTime64` is a string-to-type cast, but `JSONExtract` keyed its parsing mode off `date_time_input_format` (an input-format parsing setting), while the equivalent `CAST` honours `cast_string_to_date_time_mode`. With `date_time_input_format = 'basic'` and `cast_string_to_date_time_mode = 'best_effort'` (reproduced on current master): ```sql SELECT JSONExtract('{\"date\":\"2020-01-01 00:00:00.123Z\"}', 'date', 'DateTime64(3)'); -- 1970-01-01 00:00:00.000 (silently returns default) SELECT toDateTime64('2020-01-01 00:00:00.123Z', 3); -- 2020-01-01 00:00:00.123",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109252",
        "createdAt": "2026-07-03T03:40:20Z",
        "updatedAt": "2026-08-13T16:34:17Z",
        "timestamp": "2026-08-13T16:34:17Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "Utkal059",
        "state": "open",
        "assignees": [
          "george-larionov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109299",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Respect `date_time_overflow_behavior` for numeric temporal casts",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/101131 Related: https://github.com/ClickHouse/ClickHouse/pull/110459 Related: https://github.com/ClickHouse/ClickHouse/pull/101512 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Makes `date_time_overflow_behavior` take effect for numeric casts to `Date`, `Date32`, `DateTime` and `Time`. The `throw` and `saturate` paths were dead code for these targets, so an out-of-range value was always silently reinterpreted or truncated. ### Description `ConvertImpl` instantiated the numeric-to-temporal transforms with the compile-time constant `default_date_time_overflow_behavior` (`ignore`) instead of the runtime setting, so their `throw` and `saturate` branches were never instantiated. Threading the runtime value in makes them reachable, exposing three latent defects there: the rejected value was formatted with a narrowing `static_cast<Int64>`, undefined for a huge, infinite or `NaN` float (floats are now widened to `double`); `NaN` passed every range comparison into a narrowing cast, so it is guarded explicitly; and the clamp narrowed to `time_t` before `std::min`, so a `UInt64` above `INT64_MAX` wrapped negative and gave `1970` instead of the maximum (now clamped in the source domain first). The wide integer types missed every branch of the `DateTime` dispatch and fell through to `convertNumericGeneral`, which truncates and ignores the setting; they now use the overflow-aware transforms, so `toDateTime32(toInt128(99999999999999999999999999))` saturates instead of returning `1970-02-04`. The numeric `Time` dispatch and its bounds already landed on the base as a714d76b341fa8c from an earlier round here. `convertFieldToType` (the `VALUES`/`IN` coercion path) ignored the setting too. Serving two kinds of caller, in `ignore` mode it keys on the exactness flag: an exact target (`DROP`/`OPTIMIZE PARTITION`, strict `IN`, `KeyCondition`, sharding key) gets the canonical storage value, so an unstorable literal raises `ARGUMENT_OUT_OF_BOUND` instead of addressing a clamped partition; value materialization clamps like `CAST` (`Date`/`Date32` now clamp where they returned `NULL`). `values()` is the exception: built with a default `FormatSettings`, it keeps clamping while `CAST` raises under `throw`. The test pins that divergence. `Date32` was the only temporal target that accepted a non-finite float, saturating `inf`/`NaN` to a boundary date in every mode while the other three raise `CANNOT_CONVERT_TYPE`; it now rejects them too, on both the `CAST` and the materialization side. PR 110459 (merged) is already reconciled in this branch.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109299",
        "createdAt": "2026-07-03T14:13:03Z",
        "updatedAt": "2026-08-13T16:03:53Z",
        "timestamp": "2026-08-13T16:03:53Z",
        "metrics": {
          "reactions": 0,
          "comments": 37
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109367",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Scheduler: reclaimable memory tracking and dynamic spilling",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/109064 --> Scheduler-side support for reclaimable memory tracking and dynamic spilling, built on top of the memory reservations subsystem. The design is described in #109064. Queries can report the portion of an allocation that can be spilled or discarded on request (`IAllocationQueue::setReclaimable`), which is aggregated bottom-up as a new `reclaimable` field on every scheduler node. Each `AllocationLimit` gains a soft limit: when a workload's allocated memory exceeds it and the subtree has reclaimable memory, the scheduler asks a victim to reclaim memory (`ResourceAllocation::spillAllocation`, mirroring `killAllocation`) instead of waiting for the hard `max_memory` limit to force a kill. The request is replied: the query finishes it with `IAllocationQueue::finishSpill` after issuing the decreases for the freed memory (zero reclaimable doubling as a decline), the reply travels to the root as `Update::spilled` on the same propagation path as the state it describes, and at most one spill request is outstanding per subtree until it arrives; a victim that leaves the queue counts as having replied. Victim selection is deterministic and matches the kill order (largest usage, least precedence, largest allocation), descends a single root-to-leaf path via reclaimable-filtered ordered sets, and is fail-close: with nothing reclaimable, or with no soft limit configured, behavior is exactly as before. The soft limit is configured per workload via two new settings: `max_memory_before_spill` (absolute) and `max_memory_to_spill_ratio` (a fraction of the workload's own `max_memory`), smaller wins. The new state is observable in `system.scheduler`: `reclaimable`, `spills`, and the effective `soft_limit`. This is the scheduler side only. `MemoryReservation::spillAllocation` is currently a no-op; the query side that reports reclaimable memory and reacts to spill signals is a separate change. Documentation for the scheduler-side workload settings (`max_memory_before_spill`, `max_memory_to_spill_ratio`) and a `Spilling reclaimable memory` section are included here; the set of operators that can spill will be documented together with the query-side reaction. Invariants for the new machinery are documented on `ISpaceSharedNode` (I1-I8). Added unit tests cover fail-close, largest-first selection, fair descent skipping unreclaimable subtrees, the gate held until the victim's reply, the reply deferred to a pending decrease, re-signalling until under the soft limit, clamping reclaimable on a shrink, a declining victim, a victim removed mid-spill, the soft-limit enable transition via `CREATE OR REPLACE WORKLOAD`, the settings (absolute, ratio, and smaller-wins combination), a concurrency stress, and a parametrized throughput test. Verified under ThreadSanitizer (145 scheduler/workload gtests, 0 data races). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added workload settings `max_memory_before_spill` and `max_memory_to_spill_ratio` for the experimental memory reservation scheduling. They configure a soft memory limit above which a workload's queries are asked to spill reclaimable memory, before the hard `max_memory` limit forces an eviction. This is the scheduler-side foundation: queries do not yet report reclaimable memory or react to spill requests, so the settings have no effect until the query-side change lands. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109367",
        "createdAt": "2026-07-03T21:49:24Z",
        "updatedAt": "2026-08-13T17:44:40Z",
        "timestamp": "2026-08-13T17:44:40Z",
        "metrics": {
          "reactions": 1,
          "comments": 10
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "serxa",
        "state": "open",
        "assignees": [
          "azat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109368",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Require a join subquery alias only when it removes a real ambiguity",
        "text": "`joined_subquery_requires_alias = 1` (the default) rejected every unaliased subquery, table function or union used in a multi-table join, even when the missing alias could not cause any ambiguity. That is stricter than necessary: an alias only serves to qualify a column, so it is only needed when the unaliased table expression exposes a column whose name also occurs in another table expression of the same join. `validateJoinTableExpressionWithoutAlias` now throws `ALIAS_REQUIRED` only on such a name collision (computed with the existing `getColumnsFromTableExpression` helper over the sibling table expressions), and otherwise allows the missing alias. Genuine ambiguities involving non-sibling table expressions are still caught later by the normal `AMBIGUOUS_IDENTIFIER` resolution, exactly as they are for ordinary tables. This lets standard queries such as TPC-DS q14 (whose `cross_items` derived table has no correlation name) run without setting `joined_subquery_requires_alias = 0`: ```sql SELECT i_item_sk FROM item, (SELECT iss.i_brand_id AS brand_id FROM store_sales, item AS iss ...) WHERE i_brand_id = brand_id; -- no shared column name -> no alias needed ``` Notes: - Validation in `resolveJoin`/`resolveCrossJoin` is moved to run after all table expressions of the join are resolved, so sibling columns are known when the collision is checked. - The change is purely permissive: it never turns a previously-succeeding query into an error. When the columns of any side cannot be determined it falls back to the old strict behavior. - Only the analyzer is affected; the deprecated non-analyzer path keeps the stricter behavior. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): With `joined_subquery_requires_alias = 1` (the default), the analyzer now rejects an unaliased subquery or table function in a join only when the missing alias would make one of its columns unreachable, for example because the name collides with another joined table expression or is shadowed by an in-scope alias; otherwise the query is allowed. If the analyzer cannot determine the exposed names up front, it keeps the old strict behavior. Unambiguous queries, such as some standard TPC-DS queries, no longer require adding an alias. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109368",
        "createdAt": "2026-07-03T21:53:19Z",
        "updatedAt": "2026-08-13T16:54:37Z",
        "timestamp": "2026-08-13T16:54:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 26
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "novikd",
          "m-selmi"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109369",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support Variant type in aggregate functions (sum, avg, min, max, ...)",
        "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Aggregate functions that do not handle the `Variant` type natively (`sum`, `avg`, `min`, `max`, `argMin`, `argMax`, `quantile`, `stddevPop`, ...) can now be applied to `Variant` arguments. Previously they failed, e.g. `Illegal type Variant(Decimal(7, 2), Float64) of argument for aggregate function sum`. The aggregate functions that do accept a `Variant` argument natively (`count`, `any`, `groupArray`, `groupConcat`, the `uniq` family, ...) now skip the rows where the `Variant` holds a NULL value, like they skip the NULL values of a `Nullable` argument (only the row skipping matches `Nullable`: the result type and the empty-set result of the function are preserved) -- previously `count` counted those rows, `any` could return NULL from a group that has non-NULL values, `groupArray` stored the NULLs, and the `uniq` family counted NULL as a distinct value. Set `aggregate_functions_skip_variant_nulls = 0` (or `SET compatibility = '26.7'`) to restore the previous behavior. Persisted `AggregateFunction(f, Variant(...))` states (materialized views, `AggregateFunction` columns) keep the same type name, layout and serialization in both modes, so states written before and after the upgrade merge freely but mean different things -- the old ones aggregated the NULL rows, the new ones skip them -- and a merged result over the mix follows neither rule consistently. To keep the pre-`26.8` results over such data, set `aggregate_functions_skip_variant_nulls = 0` before inserting new data; the NULL rows already aggregated into the historical states cannot be removed by re-merging, only by re-inserting the data from the source. ### Description Such functions now aggregate over the least common supertype of the variants, wrapped in `Nullable`: ``` f(variant) == f(CAST(variant AS Nullable(supertype(T1, ..., TN)))) ``` `Nullable` preserves the implicit NULLs of the `Variant`, which the aggregation skips. **How it works.** `AggregateFunctionFactory::get()` resolves the function normally first; only if that fails for a `Variant` argument with an \"unsupported argument type\" error — `ILLEGAL_TYPE_OF_ARGUMENT`, or `NOT_IMPLEMENTED` for the creators that use it instead (`rankCorr`, `mannWhitneyUTest`, `kolmogorovSmirnovTest`, `largestTriangleThreeBuckets`) — does it fall back to a new `AggregateFunctionVariantAdapter`, which casts the `Variant` argument column(s) to `Nullable(supertype)` on the fly and forwards everything else to the nested function (whose state layout it shares). If the supertype route is not possible either, the original error is reported unchanged. Functions that already accept `Variant` (`count`, `any`, `uniq`, `groupArray`, ...) are unaffected. **Supertype.** `getLeastSupertype` is strict and has no common type for e.g. `Decimal`/`Int64` ↔ `Float64` (no lossless conversion). For an aggregate whose result is a floating-point value computed by arithmetic over its input (the `sum`/`avg`/variance/... families), a mix of numeric variants with no lossless common supertype falls back to `Float64` — but only when the user has opted into the lossy numeric supertype with the `allow_lossy_numeric_supertype` setting, and under the same promotion rule that setting applies at type inference for `if`/`multiIf`/`coalesce`/`ifNull`/`array`/`map` (all-numeric variants with at least one floating-point member; an integer-only mix such as `Variant(Int64, UInt64)` is not promoted). With the setting off (the default) the adapter handles only lossless supertypes and the original `ILLEGAL_TYPE_OF_ARGUMENT` error — which points at the setting when enabling it would help — is reported unchanged. Reconstruction of an already declared state type such as `AggregateFunction(sum, Variant(Int64, Float64))` always allows the promotion, so a state validated when it was declared never becomes unreadable because of the setting's current value: this covers the background paths that have no query context (table load / `ATTACH` at startup, merges) and, through `from_declared_state_type`, every query-time path that names the state type explicitly — a column type, a `CAST` target, a decoded binary type, and the nested function of `-Merge` over such a state. The fallback is deliberately **not** applied to exact/order-based aggregates (`min`, `max`, `argMin`, `argMax`, `any`, `quantileExact`, ...) even under the setting: a lossy `Float64` cast would silently return wrong results for them (two distinct integers above 2^53 collapse to the same `Float64`), so they keep reporting the original `ILLEGAL_TYPE_OF_ARGUMENT` when there is no lossless common supertype. Whether a function is float-promoting is declared by `AggregateFunctionProperties::is_float_promoting` (set at registration, consulted through the factory), fail-closed: a function that does not set it is treated as not float-promoting. Clean supertypes are preserved regardless of the function and the setting (`Variant(UInt8, UInt32) -> UInt32`, `Variant(Int32, Float64) -> Float64`, `Variant(Date, DateTime) -> DateTime`), so `min`/`max` still work over such variants. **Combinators.** The adapter is applied as the outermost wrapper (the combinator recursion resolves nested functions without it), so combinators compose in the usual order, e.g. `Adapter(Null(If(sum)))`. `-If`, `-State`/`-Merge` and `GROUP BY` work; state serialization for `sumState` uses the supertype (`AggregateFunction(sum, Nullable(Float64))`), so distributed / two-phase aggregation is unchanged. Only the argument positions the function actually rejects are adapted: `argMin`/`argMax` (and the `*ArgMin`/`*ArgMax` combinators) accept a `Variant` in the returned \"arg\" position and reject it only in the comparable key, so the \"arg\" keeps its original `Variant` type while just the key is cast to `Nullable(supertype)`. **Examples.** ```sql SET allow_lossy_numeric_supertype = 1; SELECT sum(v), toTypeName(sum(v)) FROM values('v Variant(Decimal(7, 2), Float64)', 1.5, 2.5, NULL, 10); -- 14 Nullable(Float64) ``` **Scope.** Only top-level `Variant` arguments whose common supertype can be wrapped in `Nullable` are handled — the adapter uses `Nullable` to carry the implicit NULLs of the `Variant`. A `Variant` whose common supertype is a container type (`Array`/`Tuple`/`Map`) is therefore still rejected, even for an orderable aggregate such as `min`/`max` over `Variant(Array(UInt8), Array(UInt16))` (supporting it would require tracking the `Variant`'s NULLs separately from the value column). A `Variant` nested inside `Tuple`/`Array`, and the `Dynamic` type, are likewise still rejected as before. These are all natural follow-ups. The result is `Nullable` because a `Variant` value can always be NULL. `singleValueOrNull` is deliberately excluded from the adapter (`AggregateFunctionProperties::is_distinctness_sensitive`): its contract keys on how many distinct values there are, and the cast to `Nullable(supertype)` collapses `Variant` values that are distinct because their alternative types differ (`1::UInt8` vs `1::UInt64`, which `uniq` counts as 2), so it would silently return a non-NULL value where the contract requires NULL — including inside the `x = ALL (SELECT ...)` rewrite. It keeps reporting `ILLEGAL_TYPE_OF_ARGUMENT` for a `Variant` argument, unchanged from before. The `-Distinct` combinator makes any combined function distinctness-sensitive in the same way (`sumDistinct` must deduplicate the genuine `Variant` values, not their casts to the supertype), so the combinator propagates the property through `AggregateFunctionFactory::tryGetProperties` (`IAggregateFunctionCombinator::isDistinctnessSensitive`) and every `...Distinct` form over a `Variant` argument keeps its original error too. `groupArrayInsertAt` and `groupArraySorted` do not claim native `Variant` support either: their generic implementations keep the state as `Field`s, which loses the original alternative type of a `Variant` value on ingest and reinfers the first compatible one on output (`1::UInt8` and `1::UInt64` collapse), and `groupArraySorted` would order by `Field` comparison instead of `Variant` order. Like `sum` / `avg`, they go through the adapter over the least common supertype of the variants, and reject a `Variant` (or `Dynamic`) argument at resolution when there is no lossless supertype. **Error codes.** `BAD_ARGUMENTS` is deliberately not part of the \"unsupported argument type\" retry signal, because creators use it for genuine semantic failures — above all invalid parameters, as in `kolmogorovSmirnovTest('bogus')` — and retrying on it would report the unrelated type error of the original, unadapted call instead of the parameter error. The few creators that rejected argument *types* with `BAD_ARGUMENTS` now reject them with `ILLEGAL_TYPE_OF_ARGUMENT`, like every other function, which keeps them adaptable to a `Variant` argument: `analysisOfVariance`, `studentTTest`, `studentTTestOneSample`, `welchTTest`, `meanZTest` and `boundingRatio` report `ILLEGAL_TYPE_OF_ARGUMENT` instead of `BAD_ARGUMENTS` when applied to an argument of an unsupported type. **`count` over a `Variant`.** `count` accepts a `Variant` natively, without the adapter, but `count(expr)` counts the not-NULL values of its argument, and a `Variant` row can hold a NULL value (`isNull` is true for it). A `Variant` is not `Nullable`, so the `Null` combinator -- and with it `AggregateFunctionCountNotNullUnary` -- is never applied to it, and the native path counted every row. `count` over a single `Variant` argument now resolves to `AggregateFunctionCountNotNullVariant`, which counts the rows whose local discriminator is not `NULL_DISCRIMINATOR`; its state is the same single counter and normalizes to `AggregateFunction(count)`, so the state forms stay byte-compatible with the plain `count` state. **The NULL-skipping contract of the natively-supported functions.** The other functions that accept a `Variant` natively are held to the same contract by `AggregateFunctionVariantNull`, a wrapper mirroring the `Null` combinator: the rows where a `Variant` argument is NULL are skipped, like the `Null` combinator skips the NULL values of `Nullable` arguments (only the row skipping is mirrored; the result-type promotion and the all-NULL-group result of the combinator are deliberately not). Without it, `any` returned NULL from a group that has non-NULL values, `groupArray` stored the NULLs its documentation promises to remove, `groupConcat` concatenated them as data, and the uniq family counted NULL as a distinct value. The wrapper is applied at the factory's leaf resolution point (`getImpl`), which every path funnels through -- the top-level native resolution as well as nested functions reconstructed from declared `AggregateFunction(f, Variant(...))` state types -- so the state layouts always match. Unlike the `Null` combinator, the wrapper never changes the result type, the state layout or the serialization of the nested function: a function that accepts a `Variant` natively already accepted it before this change, with exactly the nested result type and state representation, and both are persisted (in `AggregateFunction(f, Variant(...))` columns and in the schemas of the materialized views reading them), so changing either would break an upgrade. An all-NULL group produces the empty nested state and returns the nested function's empty-set result (`groupConcat` over a `Variant` keeps returning `String`, and an empty string for an all-NULL group); for functions whose result is the `Variant` itself (`any`, `argMin`, ...) an all-NULL group still reports NULL, because a `Variant` represents NULL on its own. Window functions keep handling their argument types themselves, so the `RESPECT NULLS` forms and `estimateCompressionRatio` still see the NULL rows, and `count` keeps its dedicated implementation above and declares `AggregateFunctionProperties::skips_variant_nulls`. **Compatibility of the NULL-skipping contract.** The functions above accepted a `Variant` argument before this change, so skipping its NULL rows is a change of their behavior and is gated by a compatibility setting: `aggregate_functions_skip_variant_nulls = 0` (implied by `compatibility` below `26.8`) restores the previous behavior, where those rows were aggregated as ordinary values. The setting only controls which rows are added to a state -- the result type, the state layout and the serialization are identical in both modes, so an `AggregateFunction(f, Variant(...))` state written by either mode stays readable and mergeable by the other. For the same reason the setting cannot change the meaning of a state that has already been written: a state written by an older version keeps the values that went into it, including the ones that came from the NULL rows, and no state-type version can retroactively tell the two apart because their bytes are identical. The consequence for existing data is spelled out in the changelog entry and in the documentation: states written before and after the upgrade merge freely but follow different accumulation rules, so a deployment that must keep the old results over persisted states has to keep writing with `aggregate_functions_skip_variant_nulls = 0` until the historical data is re-inserted. It is deliberately not consulted when there is no query context (a background operation, or a table loaded at startup). The one state that replays raw values into a nested function -- the distinct-key history of the `-Distinct` combinator -- keeps the rule by construction: the nested function of `-Distinct` over a `Variant` argument is resolved with the NULL skipping disabled, so the replay always aggregates the NULL keys a history contains (exactly like the versions that wrote states before the contract existed, so a stored `countDistinctState` over a `Variant` reads back to the same result under either value of the setting), and the NULL rows of newly aggregated data are skipped in front of the history instead, by an outer `AggregateFunctionVariantNull` over the combined function under the same setting. The aggregation of `Variant` arguments, the NULL-skipping contract and this compatibility note are documented in the `Variant` docs (the embedded documentation block in `DataTypeVariant.cpp` and `docs/reference/data-types/variant.mdx`). Added `tests/queries/0_stateless/04504_variant_aggregate_functions.sql`, `tests/queries/0_stateless/04644_variant_aggregate_state_setting_independence.sql`, `tests/queries/0_stateless/04652_variant_count_skips_nulls.sql` `tests/queries/0_stateless/04657_variant_aggregate_functions_skip_nulls.sql` `tests/queries/0_stateless/04692_variant_aggregate_nulls_compatibility_setting.sql` and `tests/queries/0_stateless/04817_variant_distinct_states_setting_independence.sql`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109369",
        "createdAt": "2026-07-03T21:55:20Z",
        "updatedAt": "2026-08-13T04:04:51Z",
        "timestamp": "2026-08-13T04:04:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 24
        },
        "labels": [
          "pr-backward-incompatible",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109433",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reserve memory for merges up front",
        "text": "Add proactive memory reservation for background merges, enhancing the existing `merges_mutations_memory_usage_soft_limit` mechanism. `background_memory_tracker` measures memory that background tasks have *already* allocated, so `canEnqueueBackgroundTask` is purely reactive: when many merges are scheduled at the same time — for example right after a mutation produces many parts — their IO buffers are not allocated yet, all of them pass the gate, and only then grow their memory usage and collide. Every merge now estimates the memory of its input and output IO buffers (`CompactionStatistics::estimateNeededMemoryForMerge`) as the number of input column streams (over all source parts) times the read IO buffer size, plus the number of output column streams (of the result part) times the write IO buffer size. Object storage (S3) write buffers are large and double-buffered, so they are accounted separately. Since IO buffers only ever hold data that flows through them, the estimate is capped by the data volume of the merge (beyond the eagerly allocated per-stream compressor and file buffers, which honor adaptive write buffers): without this cap, a merge of tiny parts in a many-column table on object storage would reserve gigabytes it can never touch, and concurrent merges would saturate the soft limit and starve each other. The merge reserves this amount at start (`MergeMemoryReservation`) and releases it when it finishes. A merge that will certainly run the vertical algorithm is priced by the streams that are alive at once (the merging columns, one gathering column at a time, and up to `max_merge_delayed_streams_for_parallel_write` delayed streams) rather than all output streams concurrently, and a stream whose data volume is not derivable from the source parts (a rebuilt projection, a `DEFAULT`-filled column of variable size) is priced at the buffers its writer allocates before any data flows through it (its compressor block and file buffer) plus a projected-volume bound, rather than at any multipart upload size - a multipart writer starts from the buffer size its caller passes and grows it only with the data written into it, leaving further growth to the reactive tracker — over-reservation starves all merges, while under-reservation merely degrades to the reactive behavior of the current code for that stream. `canEnqueueBackgroundTask` now consults both the actual usage and the reservation: - a non-replicated background merge reserves at selection and is not scheduled when the reservation would exceed the limit (it retries later); once `CurrentlyMergingPartsTagger` chooses the actual destination disk, the reservation is corrected if the pre-selection guess made from the source parts was wrong, and a merge that waited in the background queue re-prices its reservation at task start against the destination disk's live multipart upload settings (a config reload could have raised them), as the replicated path does. Whether a background merge runs as a `ReplacingMergeTree` cleanup merge is decided at selection and carried to the scheduler, so a row-reducing cleanup merge is priced as one; - a replicated merge, already committed to run locally, reserves unconditionally at execution so the reservation still throttles selection of further merges; - a user-initiated merge (`OPTIMIZE`) also reserves unconditionally, so it is throttled by the gate for other merges but can never be silently skipped by it. A single merge whose estimate exceeds the whole limit is always allowed to proceed alone, so progress is never blocked. Adds the `MergesMutationsMemoryReservation` metric for observability, a `gtest` for the reservation accounting, and stateless tests: a smoke test and a regression test that `OPTIMIZE TABLE ... FINAL` merges everything down to a single part under a pathologically small soft limit instead of silently doing nothing. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Background merges now reserve the memory of their input/output IO buffers up front against `merges_mutations_memory_usage_soft_limit`, so that scheduling many merges at once (for example right after a mutation) can no longer oversubscribe memory as they all start. A user-initiated `OPTIMIZE` is never silently skipped by this mechanism.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109433",
        "createdAt": "2026-07-05T07:51:55Z",
        "updatedAt": "2026-08-13T09:51:04Z",
        "timestamp": "2026-08-13T09:51:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 69
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109450",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix split-dependent hash of JSON/Object columns",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109428 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `JSON`/`Object`/`Dynamic` columns hashing logically equal values differently depending on their physical layout: how `JSON`/`Object` paths were split between dynamic subcolumns and shared data, or whether a `Dynamic` value sat in a typed or the shared variant. Hash-based joins on such a key silently missed matches, and a `grace_hash` spill could raise a `LOGICAL_ERROR` (\"Invalid state transition\"). A `JSON` key needed no non-default setting; `Dynamic` also required `allow_dynamic_type_in_join_keys`. ### Description **The implementation changed since the earlier approval here, so please treat it as needing a fresh look.** The approved version hashed canonical `serializeValueIntoArena` bytes and cost up to +568% CPU; this one produces none. Both measured below. `computeHashInto` is the per-row weak hash behind the in-memory scatter paths (sharded aggregation, `grace_hash` bucketing, window `PARTITION BY`, join scatter). For `JSON`/`Object` and `Dynamic` it hashed the physical layout: `ColumnDynamic` forwarded to `ColumnVariant`, which hashes a shared-variant value by its serialized blob but a typed one by the column's representation; `ColumnObject` chained sub-columns in section order, so a path's contribution moved with it. Both splits follow insertion history and can change across a temp-file round-trip, while `compareAt` already treats the representations as equal. `JSON` hits this at default settings, since `hasDynamicType` misses its inner `Dynamic`. `ColumnDynamic::computeHashInto` forwards to `ColumnVariant` as master does, then overwrites only shared-discriminator rows with the leaf hash the value would have when typed. `ColumnObject::computeHashInto` folds `(path, value)` over the sorted union of the dynamic paths and `shared_data`, one cursor per row. `updateHashFast` and `updateHashWithValueRange` stay layout-dependent. `04505_json_object_hash_split_invariance` checks that equal values with different layouts match under the three hash joins, collapse under `GROUP BY`/`DISTINCT`/`uniqExact`, and that a spill flushes real payload rather than only constructing buckets: six assertions fail on master, all pass here. Three `ComputeHashInto` gtests cover the per-row hash and scratch state SQL cannot see. `ShuffleSendStep` shuffles rows to remote workers with this hash and marks no basis in its serialized step, so mixed versions could disagree on a shuffle key. Capability gate, or unsupported anyway? <details><summary>Provenance and performance</summary> Found by the AST fuzzer on #109428 as `Invalid state transition, expected WRITING_BLOCKS, got JOINING_BLOCKS` in `GraceHashJoin::FileBucket`, STID 3913-4579 ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109428&sha=15a6faf088a1b097b3ed56c69ec1fb4d5d602d17&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)): a key re-hashing into an earlier bucket mid-spill violates that state machine. Release-shaped build, three arms from one revision (master with these files reverted; the approved implementation; this one), `UserTimeMicroseconds` min over 11 alternating reps; all arms agreed on all 12 workloads first. | workload | vs master | vs appr. | |---|---|---| | JSON `grace_hash`, no shared | +2.8% | -61.2% | | JSON `grace_hash`, mixed | +8.2% | -33.8% | | JSON `grace_hash`, all shared | +11.7% | +6.5% | | window JSON, no shared | +12.8% | -88.3% | | window JSON, mixed | **-10.6%** | -1.7% | | window JSON, 4 shared | +51.5% | -46.8% | | `Dynamic` `grace_hash` | +44.3% | -34.3% | | `Dynamic`, 100% shared | +70.3% | -1.2% | Non-JSON controls are flat and `ArenaAlloc*` matches master. Making a byte-producing basis cheap failed three times: even hoisting the buffer so the block allocated almost nothing still cost +1253%. The residual tracks the shared representation only: each shared value pays one decode, one leaf hash, one fresh scratch column. Of the worst arm's +70.3%, about +30.4% is that scratch reset (+24.9%, +18.5% on the next two). It uses `cloneEmpty()` rather than `popBack` because `popBack` is a row operation: `ColumnLowCardinality::popBack` drops indexes but not the dictionary, which `computeHashInto` re-hashes in full per row, and `ColumnVariant` nests the shape. Conditioning it on a measured property of the scratch column did not survive review, so it stays unconditional. The cliff is bounded per `computeHashInto` call, not by the table; a 5k-40k ladder is linear. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109450",
        "createdAt": "2026-07-05T18:43:19Z",
        "updatedAt": "2026-08-13T13:56:00Z",
        "timestamp": "2026-08-13T13:56:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "harikrishnan94"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109453",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add ALTER TABLE ... RECOMPRESS COLUMN",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109432 Introduce a new `ALTER TABLE ... RECOMPRESS COLUMN col` statement that re-compresses the existing data of a column with the column's current compression codec. Changing a column's codec with `MODIFY COLUMN col CODEC(...)` is metadata-only: the new codec applies to newly written data, while data already stored in existing parts keeps its old codec until the parts happen to be merged. `RECOMPRESS COLUMN` rewrites the data of `col` in existing parts so that it is compressed with the codec currently set in the table metadata. Because a compression codec does not change the serialized representation of a column, for `Wide` parts the recompression is done **without deserializing the values**: each compressed block of every substream `.bin` is decompressed and re-compressed one-to-one with the new codec. This keeps the decompressed content and granule boundaries byte-identical, so the marks file only needs its compressed offsets remapped (the decompressed offsets, per-granule row counts, the primary index and skip indexes are preserved and hardlinked). The decompressed content is unchanged, so the `uncompressed_size`/`uncompressed_hash` checksums are carried over from the source part and only the on-disk `file_size`/`file_hash` are recomputed. `Compact` parts (and dynamic-subcolumn types) cannot recompress a single column in isolation, so they fall back to a normal whole-part re-serialization that writes every column with its current codec. Implemented as a mutation. A new `ALTER RECOMPRESS COLUMN` grant is added. The issue also asks to check `RECOMPRESS` TTL mutations, which currently go through a full deserialize/re-serialize merge; that is left for a follow-up. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `ALTER TABLE ... RECOMPRESS COLUMN col`, which re-compresses a column's existing data with its current codec. For wide parts the data is recompressed without deserializing the column values.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109453",
        "createdAt": "2026-07-05T21:19:23Z",
        "updatedAt": "2026-08-13T07:01:14Z",
        "timestamp": "2026-08-13T07:01:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 24
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109454",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Materialize column statistics on INSERT for small tables by default",
        "text": "Cost-based join reordering relies on column statistics (in particular the number of distinct values of join keys) to avoid pathological plans. Column statistics are auto-declared by default (`auto_statistics_types = 'minmax, uniq'`), but they are only materialized during merges (`materialize_statistics_on_merge`) — so a freshly bulk-loaded table has no usable statistics at query time until it happens to be merged. This bites the TPC-H benchmark. At scale factor 40 on a 32 GB machine, Q5 and Q8 time out (>100s, spilling to disk): - Without the NDV of the low-cardinality `nationkey` (25 distinct values), the optimizer cannot see that joining `customer` and `supplier` on `nationkey` (via the transitive `c_nationkey = s_nationkey`) produces a ~19 billion row many-to-many intermediate, and it materializes that as the hash-join build side. - The dimension tables that carry the decisive join keys are small (SF40: `customer` 509 MiB, `supplier` 32 MiB, `region`/`nation` tiny), so materializing their statistics on insert is cheap — but the fact tables (`lineitem` 11 GiB, `orders` 2.5 GiB) are large and should not pay per-insert statistics cost. This PR enables `materialize_statistics_on_insert` by default, bounded by a new setting `materialize_statistics_on_insert_max_table_size` (default 25 GiB): tables whose current size is at or below the threshold build statistics at insert time, larger tables skip it and materialize during merges as before. The size check is an `O(1)` atomic load of the table's current active size; `0` disables the limit. Measured on TPC-H SF40 with the server memory capped at 28.8 GiB (0.9 × 32 GiB), out of the box (no manual `MATERIALIZE STATISTICS`): | Query | Before | After | | --- | --- | --- | | Q5 | >200s, 8.6 GB spilled to disk | 0.5s, 636 MiB, no spill | | Q8 | 143s, 4.35 GB spilled to disk | 0.6s, 1.37 GiB, no spill | The per-insert cost is gated correctly: inserting into a table already above the threshold builds no statistics (`MergeTreeDataWriterStatisticsCalculationMicroseconds = 0`), while an uncapped insert of the same block spends the usual time. ## Validation on other benchmarks To check the effect more broadly, I compared having statistics available at load time (this change) against not having them (`use_statistics=0`, the pre-change planning state), on the ClickBench `/versions` query sets, on the same freshly-loaded data under the 32 GiB cap: - **TPC-DS (SF40, 103 queries):** 4 queries go from OOM/timeout to succeeding (q24, q51, q71, q93); many large speedups (q74 98×, q66 31×, q7 22×, q83 18×, …); ~35% faster in aggregate; no genuine regressions. - **JOB / IMDB (113 queries):** net ~12% faster; wins up to ~8× (q1, q91, q92). The dataset is small enough that nothing fails either way; a couple of queries regress mildly (q55, q26). - **Coffee Shop (fact 359M rows + 2 dimensions, 17 queries):** wins up to ~50× (q8 42×, q11 and q16 go from timeout to 1.4s / 63s, q13 15×); one query regresses ~2.3× (q17). Overall, making dimension statistics available out of the box is strongly net-positive — many multi-× speedups and several queries rescued from OOM/timeout. A small number of queries regress because the cost-based join reorder occasionally makes a worse choice with statistics than without; that is a pre-existing optimizer-quality concern that enabling statistics by default surfaces more often, not a regression in the materialization itself. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Column statistics are now materialized on `INSERT` by default when the table's current active size plus the written block size is at most the new `materialize_statistics_on_insert_max_table_size` setting (default 25 GiB); the check is per written block, so a first bulk load into an empty table may still materialize statistics for each written block. This gives the cost-based join optimizer accurate estimates for freshly-loaded dimension tables and avoids pathological join orders (for example, TPC-H Q5 and Q8 no longer time out at scale factor 40), while large established fact tables keep materializing statistics during merges. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1307` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109454",
        "createdAt": "2026-07-05T21:20:22Z",
        "updatedAt": "2026-08-13T04:46:37Z",
        "timestamp": "2026-08-13T04:46:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 44
        },
        "labels": [
          "pr-performance",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "rschu1ze"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109455",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Native Google Cloud Storage integration (google-cloud-cpp)",
        "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental native Google Cloud Storage integration built on the Google Cloud C++ SDK (`google-cloud-cpp`, the GCS JSON API), as an alternative to the existing S3-compatibility path. Enable it with the new `use_native_gcs` setting for the `gcs` table function and the `GCS` table engine, or use it as a MergeTree storage disk via `object_storage_type: gcs`. ### Documentation entry for user-facing changes Documented in `docs/reference/functions/table-functions/gcs.mdx` (new \"Native GCS integration\" section). --- ## Description ClickHouse currently talks to GCS **only through its S3-compatible XML API** (the AWS SDK against a `storage.googleapis.com` endpoint, plus a cluster of GCS quirk-toggles in the S3 client). This PR adds a first-class native backend on top of the already-vendored `google-cloud-cpp` storage client. **Opt-in, non-breaking.** Everything is gated behind the new experimental setting `use_native_gcs` (default `false`). With it off, `gcs()` / `ENGINE = GCS` behave exactly as before (S3-compatibility, HMAC keys). With it on, they route through the native client. ### What's added - `ObjectStorageType::GCS` + `GCSObjectStorage : IObjectStorage` (`src/Disks/DiskObjectStorage/ObjectStorages/GCS/`), with `ReadBufferFromGCS`/`WriteBufferFromGCS` over the SDK's `ObjectReadStream`/`ObjectWriteStream` (ranged seeks, resumable uploads), native listing, `RewriteObject`-based copy, `DeleteObject`, `GetObjectMetadata`. - Registered as the `gcs` object-storage type + a `gcs` disk-type alias, so MergeTree data can live natively on GCS (`type: object_storage, object_storage_type: gcs`). - `StorageGCSConfiguration` (reuses `StorageS3Configuration`'s argument parsing) and a `TableFunctionGCS` / `ENGINE = GCS` selector that picks native vs S3-compat by the `use_native_gcs` setting. - Auth mirrors the compat surface on the native side: Application Default Credentials, service-account JSON, the `google_adc_*` OAuth refresh-token flow (reusing `IO/GCPOAuth`), anonymous / `NOSIGN`, and a REST endpoint override for the GCS emulator. - `USE_GOOGLE_CLOUD` is now exposed in `system.build_options`. ### Testing - Every changed/new translation unit compiles against the real headers and the vendored SDK (`USE_GOOGLE_CLOUD=1`, `USE_AWS_S3=1`). - Unit test for endpoint parsing (`gtest_gcs_endpoint`). - Stateless test for the setting wiring (`04502_use_native_gcs_setting`). - Integration test `test_native_gcs` runs against a `fake-gcs-server` emulator (a new `with_gcs` helper in `cluster.py`): `gcs()` INSERT/SELECT + glob with `use_native_gcs=1`, and a MergeTree-on-GCS-disk round trip. Note: `fake-gcs-server` speaks the GCS API (unlike minio, which is S3-only). The module is skipped on builds without the SDK. ### Known follow-ups - `GCSObjectStorage::iterate()` currently materializes the full listing (a lazy paginating iterator like S3's is a future optimization). - The `google_adc_*` refresh-token access token is minted eagerly (no auto-refresh yet for long-lived disks). - `readSmallObjectAndGetObjectMetadata` uses the base default (relevant only for a future native Iceberg-over-GCS path, which is not wired here). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109455",
        "createdAt": "2026-07-05T22:03:13Z",
        "updatedAt": "2026-08-13T02:51:58Z",
        "timestamp": "2026-08-13T02:51:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "submodule changed",
          "pr-experimental",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109472",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Log tolerated connection failures to remote MySQL/PostgreSQL databases as warnings instead of errors",
        "text": "Fix CI upgrade test failure: Example failing report — Upgrade check (amd_release) on the unrelated PR https://github.com/ClickHouse/ClickHouse/pull/108255: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=108255&sha=28ca539bd4bf66de37e388238de45d840bc30c52&name_0=PR&name_1=Upgrade%20check%20%28amd_release%29 (`Error message in clickhouse-server.log (see upgrade_error_messages.txt)` with the `getTablesIterator` line above; the report shows green now because the check passed on a later rerun — which is exactly the flaky-victim pattern). **Problem.** A `DatabasePostgreSQL` or `DatabaseMySQL` pointed at an unreachable server writes the connection failure to the server log at `<Error>` level on code paths that explicitly tolerate and retry the failure. A temporarily unavailable remote server is a normal operational state on these paths, yet it floods the log with error-level messages. Concretely: - every `system.tables` / `system.columns` scan that touches the database logs an error (`DatabasePostgreSQL::getTablesIterator`: `Code: 614. DB::Exception: ... Connection to ... failed`); - the background cleaner logs one more error every rescheduling cycle; - `ATTACH DATABASE ... ENGINE = MySQL(...)` logs three error lines for a single tolerated probe failure. This also trips the Upgrade check (which fails on any unexpected error message in `clickhouse-server.log`) on unrelated PRs: the stateless test `04210_show_remote_databases_in_system_tables` (added in #104416) leaves `PostgreSQL`/`MySQL` databases pointed at the unroutable `192.0.2.1` during the check's restart window, and since #109082 made remote databases visible to `system.tables` scans by default, the check failed on 137 distinct PRs in 14 days, and 0 times on master. **Root cause.** Seven log sites report a caught-and-tolerated (or caught-and-propagated) connection failure at error level: the catches in `DatabasePostgreSQL::getTablesIterator` (kept non-throwing so `system.tables` scans do not fail) and `DatabasePostgreSQL::removeOutdatedTables` (the cleaner reschedules and continues) use `tryLogCurrentException` with the default `LogsLevel::error`; the same for the `DatabaseMySQL` constructor catch on `ATTACH` and `DatabaseMySQL::getCreateTableQueryImpl` with `throw_on_error = false`; and the connection-pool layers (`mysqlxx::Pool::allocConnection`, `mysqlxx::PoolWithFailover::get`, `postgres::PoolWithFailover::get`) log at error level before propagating the exception to the caller — who is the one deciding how severe the failure actually is. **Fix.** Downgrade those seven sites to warning. Exception propagation is unchanged everywhere: a non-tolerated failure (e.g. `CREATE DATABASE` with an unreachable host) still reaches the client as an error through the normal query-error path. The new test `04506_remote_database_unreachable_no_error_log` attaches `PostgreSQL` and `MySQL` databases pointing at an unreachable host, exercises the tolerated paths (attach probe, `system.tables` scan, background cleaner), and asserts through `system.text_log` that each path produced a warning (proving the failure path fired) and no error-level lines at all for the corresponding queries. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109472",
        "createdAt": "2026-07-06T09:15:05Z",
        "updatedAt": "2026-08-13T10:09:29Z",
        "timestamp": "2026-08-13T10:09:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-ci"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109531",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix text index direct read with patch parts (lightweight updates)",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/106460 (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/99543 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/106460 Related: https://github.com/ClickHouse/ClickHouse/pull/99543 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results (and an `UNKNOWN_IDENTIFIER` error for `MATERIALIZED` indexed columns) when a query reads from a `text` index together with a column updated by a lightweight update (patch parts). Direct reading from the text index is now disabled for parts that have patch parts. ### Description Direct reading from a `text` index (`query_plan_direct_read_from_text_index`, enabled by default) produces the search result as a virtual column via a dedicated index-read step prepended to the reader chain. That step reads no physical data, so it cannot anchor patch application, which aligns patches to the block by `_part_offset`. When a queried part has patch parts (from a lightweight `UPDATE`) and a patch-applied column is read together with the direct-index virtual column, the direct-read path and the patch-application path do not line up. The result was either dropped rows (wrong results) or, when the indexed column was `MATERIALIZED`, an `UNKNOWN_IDENTIFIER` error while evaluating the virtual column's default expression. This happens even when the patched column is unrelated to the indexed column, so the existing per-index `canUseIndex` check was not sufficient. The fix disables direct text-index reading for the whole query when any queried part has patch parts, falling back to regular index reading (identical to `query_plan_direct_read_from_text_index = 0`, which always produced correct results). Bisected to #99543. Reproducer from the issue (`json_title String MATERIALIZED data_as_json.title::String` with a `text` index) plus a plain-column variant are added as `03100_lwu_51_text_index_patched_column`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109531",
        "createdAt": "2026-07-06T15:52:24Z",
        "updatedAt": "2026-08-13T13:36:08Z",
        "timestamp": "2026-08-13T13:36:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 20
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "v26.5-must-backport"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "CurtizJ"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109594",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support `arrayExists` predicates for text-like indexes",
        "text": "Previously, `text`, `tokenbf_v1`, and `ngrambf_v1` indexes on `Array(String)` columns could not be used for predicates such as `arrayExists(x -> x LIKE '%needle%', arr)`. These predicates had to read every granule even though the index already stores tokens from the array elements. This PR lets index analysis use `arrayExists` lambdas where the lambda tests the element against a constant, for example `arrayExists(x -> x LIKE '%needle%', arr)` or `arrayExists(x -> x IN ('a', 'b'), arr)`. The index can then check the tokens for `arr` before reading rows. This is safe because any matching element must have added the required tokens to the index granule. The original `arrayExists` expression is still evaluated for rows that pass the index filter. The supported functions are listed explicitly for each index type. They include positive string predicates such as `equals`, `LIKE`, `ILIKE`, `startsWith`, `endsWith`, `match`, `multiSearchAny`, `hasToken`, and `IN`; the `text` index also supports `hasAnyTokens`, `hasAllTokens`, `hasPhrase`, `multiSearchAnyUTF8`, and `multiMatchAny`. Negative predicates such as `notEquals`, `NOT LIKE`, and `NOT IN` are not supported because they are not valid filters for empty arrays. For `text` indexes, functions whose result depends on tokenization, such as `hasToken`, are supported only when the index uses `splitByNonAlpha` without preprocessors or postprocessors. Other `arrayExists` lambdas are unchanged: they simply do not use these indexes. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `text`, `tokenbf_v1`, and `ngrambf_v1` indexes can now prune granules for predicates such as `arrayExists(x -> x LIKE '%needle%', arr)` on `Array(String)` columns.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109594",
        "createdAt": "2026-07-07T02:19:49Z",
        "updatedAt": "2026-08-13T11:45:00Z",
        "timestamp": "2026-08-13T11:45:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "EmeraldShift",
        "state": "open",
        "assignees": [
          "ahmadov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109602",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix the streaming-insert block wait not expiring under the query profiler",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109592 Related: https://jira.mariadb.org/browse/CONC-834 **Problem.** The file-descriptor poll timeout can silently never expire on a query thread. `ReadBufferFromFileDescriptor::poll` restarts the interrupted `poll` with the full original timeout after `EINTR`, resetting the deadline on every signal. Query threads receive periodic sampling-profiler timer signals (`query_profiler_real_time_period_ns`, default 1 s; the handler's `SA_RESTART` does not apply — per `signal(7)`, `poll` is never auto-restarted), so whenever the signal period does not exceed the timeout, the wait becomes unbounded while the fd stays silent. The user-visible path is the streaming-insert block wait: `IRowInputFormat::read` calls `poll` with the remaining `input_format_max_block_wait_ms` budget over a `StorageFile` fd/file or stdin buffer. With the profiler active, a stalled input source delays the partial-block flush indefinitely instead of flushing when the wait limit is reached. (All socket paths use `ReadBufferFromPocoSocket*::poll`, which is already deadline-aware, and are not affected.) **Root cause.** Same defect class as the MySQL `connect_timeout` fix in #109592 (mariadb-connector-c, upstream [CONC-834](https://jira.mariadb.org/browse/CONC-834)): an `EINTR` retry loop that passes the original timeout instead of the remainder. This was the last such site in `src/` — a sweep of all timed waits (`poll`/`epoll_wait`/`select`/`nanosleep`/`sigtimedwait`/io_uring/timerfd) found every other one deadline-aware (Poco sockets, `Epoll`, `KeeperTCPHandler`, `ShellCommandSource`, `base/sleep`). **Fix.** Re-poll with the remaining time computed **in microseconds from a single monotonic anchor**, reporting a timeout once the budget is exhausted. Per-retry whole-millisecond accounting (as in some existing call sites) would truncate a sub-millisecond retry interval to zero and make no progress under a sub-millisecond signal cadence, so the remainder is derived from the untouched anchor instead. The unit test (`gtest_fd_read_buffer_poll_under_signals.cpp`) waits on a pipe with no writer while a per-thread kernel timer (`timer_create` + `SIGEV_THREAD_ID`, the query profiler's own mechanism, so delivery deterministically targets the polling thread) interrupts the poll at two cadences: 10 ms (the classic deadline reset) and 0.5 ms (pins the microsecond-precision accounting). A watchdog thread disarms the timer after 3 s so a regressed build fails the elapsed assertion at ~3.2 s instead of hanging; the fixed poll returns in ~200 ms. Verified in both directions: the naive loop fails both cadences, a millisecond-accounting variant fails exactly the sub-millisecond one. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the block-wait timeout of streaming inserts (`input_format_max_block_wait_ms`) potentially never expiring while the query profiler is active: the file-descriptor poll restarted with the full timeout after every profiler signal, so a stalled input source could delay the partial-block flush indefinitely instead of flushing when the wait limit is reached.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109602",
        "createdAt": "2026-07-07T07:53:58Z",
        "updatedAt": "2026-08-13T14:15:36Z",
        "timestamp": "2026-08-13T14:15:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "Algunenano"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109636",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reduce peak memory during JSON advanced shared data merges",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reduce peak memory during merges of JSON columns with many unique paths in advanced shared data by materializing columns per-bucket instead of all at once, and by introducing a new `advanced_chunked` shared data serialization that splits row ranges into smaller chunks so only one flattened chunk needs to be in memory at a time. The chunk size is controlled by the new `object_shared_data_target_chunk_rows` MergeTree setting. Closes https://github.com/ClickHouse/ClickHouse/issues/108992",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109636",
        "timestamp": "2026-08-12T20:51:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "Avogar",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109710",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "[DRAFT] Remove CatBoost integration",
        "text": "One fine day ... Requires ~https://github.com/ClickHouse/ClickHouse/pull/80363~ https://github.com/ClickHouse/ClickHouse/pull/109999 first. Removes the CatBoost integration: the `catboostEvaluate` function, the `system.models` table, the `SYSTEM RELOAD MODEL(S)` queries and the matching `SYSTEM RELOAD MODEL` privilege, together with their documentation, tests and the bundled model artifacts. The `CANNOT_LOAD_CATBOOST_MODEL` and `CANNOT_APPLY_CATBOOST_MODEL` error codes are intentionally kept so their numbers are not reused. Refs: - https://github.com/ClickHouse/ClickHouse/issues/70771 - https://github.com/ClickHouse/ClickHouse/issues/45052 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Removed the CatBoost integration. The `catboostEvaluate` function, the `system.models` table and the `SYSTEM RELOAD MODEL(S)` queries are no longer available. Queries, views and grants referring to them have to be migrated before the upgrade: in particular, revoke `SYSTEM RELOAD MODEL` from all users and roles beforehand, because an access entity whose grants still mention the removed privilege cannot be parsed on startup. Also let any in-flight `SYSTEM RELOAD MODEL(S) ON CLUSTER` entries drain from the distributed DDL queue before the upgrade: upgraded hosts can no longer parse such entries and will finish them with an error status instead of executing them. The same applies to persisted table metadata: a table whose column `DEFAULT`/`MATERIALIZED`/`ALIAS` expressions, secondary indices, constraints, projections, partition/order/sample keys, or TTL expressions still reference `catboostEvaluate` cannot be loaded after the upgrade (`UNKNOWN_FUNCTION`), so drop or rewrite such definitions beforehand. The revoke also has to cover every other carrier of serialized grants: a `GRANT SYSTEM RELOAD MODEL` line in the `grants` section of `users.xml` must be removed as well, because it can no longer be parsed and the server then refuses to load the users configuration at startup; on clusters with replicated (Keeper-backed) access storage, revoke the privilege before upgrading any replica, because with the default `access_control_improvements.throw_on_invalid_replicated_access_entities = 0` an upgraded replica that reads a user or role still carrying `SYSTEM RELOAD MODEL` from ZooKeeper cannot parse it and silently drops that entity from its in-memory copy (set `throw_on_invalid_replicated_access_entities = 1` to fail loudly instead); and access backups (`BACKUP ... ACCESS`) taken while the grant was still present cannot be restored with `RESTORE ... ACCESS` on the new version, so re-create such backups after the revoke.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109710",
        "createdAt": "2026-07-07T20:39:00Z",
        "updatedAt": "2026-08-13T00:56:28Z",
        "timestamp": "2026-08-13T00:56:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 17
        },
        "labels": [
          "pr-backward-incompatible",
          "hold",
          "pr-autogenerated-docs"
        ],
        "author": "rschu1ze",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109881",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Read Keeper changelogs in parallel at startup",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Speed up Keeper startup by reading multiple changelog files concurrently instead of serially, controlled by new settings `log_startup_read_max_streams` and `log_startup_read_buffer_size`. ### Documentation entry: - [ ] Documentation is written",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109881",
        "createdAt": "2026-07-09T10:41:36Z",
        "updatedAt": "2026-08-13T17:41:41Z",
        "timestamp": "2026-08-13T17:41:41Z",
        "metrics": {
          "reactions": 2,
          "comments": 3
        },
        "labels": [
          "pr-improvement",
          "submodule changed",
          "comp-keeper"
        ],
        "author": "antonio2368",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109891",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reintroduce borrowed threadgroup async uaf fix",
        "text": "Reintroduce #108988 Related: https://github.com/ClickHouse/ClickHouse/pull/107030 Related: https://github.com/ClickHouse/ClickHouse/pull/108577 CI: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105890&sha=3dc0e76362eb18e27f1fffcd3f61ca9f13725fe8&name_0=PR&name_1=Stateless%20tests%20%28amd_tsan%2C%20s3%20storage%2C%20sequential%2C%201%2F2%29 Fixes a use-after-free risk in asynchronous work scheduled while a borrowed `ThreadGroup` is current. Borrowed `ThreadGroup` objects used by materialized view and async insert flush paths point their `performance_counters` and `memory_tracker` to the parent query group, so they are valid only while that parent group is alive. Async callbacks could capture such a borrowed group and later attach it on a pool thread after the parent query group had finished. Instead of keeping the parent `ThreadGroup` alive with a `shared_ptr`, this change keeps borrowed accounting scoped. Borrowed groups are marked explicitly, async callback capture drops borrowed groups, and thread pool callback runners capture the normalized group at task enqueue time rather than when a potentially long-lived runner object is created. This preserves synchronous borrowed accounting, but async work started from a borrowed scope runs under normal thread/global accounting instead of writing into, or prolonging the lifetime of, an already finished query group. Full ASAN reports https://gist.github.com/filimonov/1ec59047c65e3a5367c5c83f6021cc27 Compared to #108988 - added one commit with code comments + fix of the test failure https://github.com/ClickHouse/ClickHouse/issues/109841 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a memory safety issue where asynchronous work scheduled from materialized view processing could keep using query-level accounting after the query had finished.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109891",
        "createdAt": "2026-07-09T12:34:20Z",
        "updatedAt": "2026-08-13T17:56:28Z",
        "timestamp": "2026-08-13T17:56:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 47
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "comp-query-execution"
        ],
        "author": "filimonov",
        "state": "open",
        "assignees": [
          "azat",
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109896",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix use_constant_folding_in_index_analysis issues",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109893 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix three issues with `use_constant_folding_in_index_analysis`: a logical error `Invalid partition key size` that could abort a `SELECT` using a normal projection when the table sets `part_minmax_index_columns = 'with_block_number_offset'`; wrong results (silently dropped rows) for a filter on a modulo partition key such as `PARTITION BY id % 200`; and a crash (data race) when a `text` index with a `sparseGrams` tokenizer is queried with a `LIKE` predicate over many partitions with `max_threads > 1`. ### Description Three issues in the `use_constant_folding_in_index_analysis` path. **1. Logical error `Invalid partition key size`** (fuzzer STID 2677-496b, [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=106734&sha=6220f1e884aa1a725214970d897d217a75ba8c6a&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20parallel%29)). In `selectPartsToRead`, `minmax_idx_condition` comes from a normal projection's metadata, whose partition key is empty, but the loop passed the *parent* part's partition (size 1) to per-partition specialization, so `MergeTreePartition::getID` aborted against the size-0 key. Fix: projection parts are not specialized; they get the unsubstituted condition. **2. Wrong results for modulo partition keys** ([repro](https://fiddle.clickhouse.com/3a7a3f59-b256-420f-908f-09459e46aa8d)). The stored value uses the backward-compatible `moduloLegacy` rewrite (8-bit: `moduloLegacy(-199, 200) = 57`) while the filter evaluates modern `modulo` (16-bit: `-199`). Matching the modern predicate against the modern key substituted the stored legacy value, turning `id % 200 < 0` into `57 < 0` and pruning parts that do hold matching rows (498 rows became 356). Fix: match against the legacy-adjusted key, so a modern `modulo` node no longer matches. Non-modulo keys fold unchanged. **3. Crash with a stateful `sparseGrams` tokenizer** ([repro](https://fiddle.clickhouse.com/c8e3cc24-f366-4413-835e-4bb7afc0d0d8)). With folding on, the skip-index condition is rebuilt per partition inside a thread pool, and every condition got the index's single tokenizer as a raw pointer. `sparseGrams` advances a mutable iterator, so concurrent builds corrupted its state (SIGSEGV in `SparseGramsTokenizer::nextInStringLike`). Fix: add `ITokenizer::isStateful` and give each condition a private clone; stateless tokenizers stay shared. Aggregators clone too, for symmetry with `MergeTreeIndexAggregatorText`, but that is defence in depth rather than a fixed race: each part-writer owns its aggregator, so writes never share one tokenizer concurrently. Tests in `04510_projection_partition_minmax_key_size`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109896",
        "createdAt": "2026-07-09T14:01:22Z",
        "updatedAt": "2026-08-13T01:48:45Z",
        "timestamp": "2026-08-13T01:48:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 28
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "Michicosun"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109925",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Warm statistics estimator on commit to avoid first-SELECT cold load",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/109454 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Warm the table statistics estimator right after an INSERT commits so the first `SELECT` does not synchronously cold-load and rebuild column statistics from disk in the query planner. ### Description Helper PR against `stats-on-insert-size-threshold`, folding in the root-fix direction discussed on #109454. Opened at @egor-click's request. Root fix (instead of gating the planner around the cold load): - Retain the full per-column `ColumnsStatistics` that INSERT already builds on the in-memory part (`IMergeTreeDataPart::statistics_cache` + `setStatistics`/`tryGetCachedStatistics`/`hasCachedStatistics`). Disk-loaded parts leave it empty and keep using the on-disk files. - Warm `MergeTreeData::cached_estimator` in `Transaction::commit` once committed parts become Active, so the first `SELECT` skips the synchronous planner cold-load. - The estimator builder now copies statistics into its own aggregate (`cloneEmpty()` + `merge`) before merging, so it never mutates the part-owned `ColumnStatistics` that `loadStatistics()` returns. This also protects `refreshStatistics` and the what-if estimator, which both route through `addStatistics`. Hardening applied (item 1 of the 3 reported on #109454): - Commit-time warming is gated on committed parts actually carrying retained in-memory statistics (`hasCachedStatistics`), not on `use_statistics_cache` alone. Without this gate, tables whose statistics are NOT materialized on insert (large tables above `materialize_statistics_on_insert_max_table_size`, or the setting off) would loop all active parts and do the full on-disk cold-load synchronously on every INSERT. Measurements (debug, this branch's base): - Smoke (MergeTree, tdigest+uniq+minmax+countmin on 2 cols): server-side planner cold-load marker = 0 across SELECT-after-INSERT / repeat / SELECT-after-2nd-INSERT; commit-warm fired 3x; no crash / LOGICAL_ERROR / sanitizer. - Gate verified both directions: implicit-stats table warms on commit under the INSERT query id; `auto_statistics_types=''` + `materialize_statistics_on_insert=0` skips warming under the INSERT id. - A/B on the exact `join_convert_outer_to_inner` query, same binary/data: fix = 0 planner cold-loads on first SELECT-after-INSERT; baseline (`use_statistics_cache=0`) = 4 cold-loads (2 parts x 2 tables). Remaining follow-ups (reported on #109454, not in this PR): incremental warm instead of O(parts) rebuild per commit; bound `statistics_cache` retention by a table-size budget or drop after first warm; and the `disk_connections_hard_limit` interaction the bot flagged for remote-storage inserts.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109925",
        "createdAt": "2026-07-09T16:30:04Z",
        "updatedAt": "2026-08-13T05:28:57Z",
        "timestamp": "2026-08-13T05:28:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "can be tested"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:109946",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix propagation of settings in `accurateCastOrDefault`",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix propagation of settings in `accurateCastOrDefault`. Closes https://github.com/ClickHouse/ClickHouse/issues/109943",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/109946",
        "createdAt": "2026-07-09T23:21:47Z",
        "updatedAt": "2026-08-13T17:52:55Z",
        "timestamp": "2026-08-13T17:52:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport"
        ],
        "author": "Avogar",
        "state": "open",
        "assignees": [
          "antonio2368"
        ],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110029",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Read-through filesystem cache for the experimental ReaderExecutor",
        "text": "First of two PRs adding a read-through cache to the experimental `use_reader_executor` read path (default off), split out of the `ReaderExecutor` series (#103706), after decryption (#109702). This PR adds the cache-provider interface and the filesystem cache tier; the page cache and the richer, coordinated driver follow in a second PR. Adds `ICacheProvider` and the FileCache-backed `DiskCacheProvider`, consulted and populated per read window. The interface and provider are adopted from #103706's redesigned head: a whole-range `resolve(object, offset, range)` returns the window's residency as an ordered list of `Resolution`s in one cache transaction — each hit carrying a `CacheReader`, each populating miss carrying its open `CacheWriter` (a read-only/bypass tier returns writer-less misses). The executor drives it with the simplest possible per-window loop: `resolve` the window start, serve a cache hit straight from the tier's buffer (zero-copy), or on a miss claim the covering cell(s), fetch from the source, populate, and serve one block. The miss read goes through the long connection, so a cold sequential scan streams from one held connection; the window is returned as a `ChainedBuffers` (block-chunked) and decrypted per node on an encrypted disk. Concurrent readers of the same cold cell elect a single downloader via a claim taken before the fetch, so only one populates each cell; a cell another reader is already downloading is fetched through from the source (its populate lands zero bytes) rather than waited on. Coordinated waiting arrives with the page cache PR. Scope of this PR: - The filesystem cache serves known-size sources only — an FS cache is never attached to an unknown-size object, so the cache path is never entered for one. The executor's handling of unknown-size sources is otherwise unchanged from `master`. - No page cache, no cross-window plan, no prefetch, and none of #103706's plan machinery (`ResidencyIterator`, `CoverageMap`, `MemoryPressureMonitor`). - `ChainedBuffers` is functionally unchanged from `master` (only a couple of over-long comments trimmed). - Everything is gated behind `use_reader_executor`; the executor also falls back for the distributed cache and async prefetch, which it does not implement. New settings: `reader_executor_window_size` (serve window, 4 MiB) and `reader_executor_block_size` (buffer chunk, 1 MiB), each at least 4 KiB (rejected at settings load otherwise). Tests: `04511_reader_executor_disk_cache` (an `s3_cache` MergeTree) asserts the executor engages and consults the filesystem cache; `04604_reader_executor_min_size` asserts the sub-4-KiB window/block rejection; IO gtests cover the executor, the provider, and the offset map. Related: https://github.com/ClickHouse/ClickHouse/pull/103706 Related: https://github.com/ClickHouse/ClickHouse/pull/109702 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a read-through filesystem cache to the experimental `ReaderExecutor` read path (`use_reader_executor`, disabled by default). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110029",
        "createdAt": "2026-07-10T16:50:28Z",
        "updatedAt": "2026-08-13T17:36:45Z",
        "timestamp": "2026-08-13T17:36:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "CheSema",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110072",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix bitmap subset functions for small-set and promoted signed bitmaps",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109974 Related: https://github.com/ClickHouse/ClickHouse/issues/106208 `subBitmap` and `bitmapSubsetOffsetLimit` apply offset/limit in ascending value order. The small-set bitmap path iterated keys in insertion order instead, producing wrong subsets when values were not inserted sorted (e.g. `bitmapBuild([5, 4, 1, 2, 3])`). The small paths of `rb_range` and `rb_limit` also compared element values as `UInt32`, so values above `2^32` in `UInt64` / `Int64` bitmaps were truncated and failed to match their own thresholds — for example `bitmapSubsetLimit(bitmapBuild([4294967297]::Array(UInt64)), 4294967297, 1)` returned an empty bitmap. Beyond that, this aligns `bitmapSubsetInRange` / `bitmapSubsetLimit` / `bitmapMin` / `bitmapMax` / `bitmapContains` / `bitmapTransform` so that small and promoted bitmaps compare elements in the same unsigned element-type domain. That is the domain signed element types were introduced with in https://github.com/ClickHouse/ClickHouse/pull/20171: `01702_bitmap_native_integers` has asserted `bitmapMin` = `251` and `bitmapMax` = `255` for `Int8` `[-1, -2, -3, -4, -5]` since 2021, while `rb_range` and `rb_limit` were left comparing sign-extended `UInt32` values and were annotated at the time as \"currently only support UInt32\". Master therefore contradicted itself: `bitmapSubsetInRange(bm, bitmapMin(bm), bitmapMin(bm) + 1)` returned nothing for `Int8`, `Int16` and `Int64` bitmaps, even though `bitmapMin` reported an element that is present. `BSINumericIndexedVector` looks indexes up in the same bitmaps, so it now uses that domain as well. Without it, `groupNumericIndexedVector` returned different results for `Int8` / `Int16` index columns than for wider index types once a bit-slice bitmap was promoted past the small set, and `numericIndexedVectorGetValue` never found a negative index at all. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect results from `subBitmap`, `bitmapSubsetInRange` and `bitmapSubsetLimit` for bitmaps still held in the small representation, including `UInt64` and `Int64` element values above `2^32`, which were truncated to 32 bits. Comparisons in bitmap functions over signed element types now consistently use the unsigned value of the element type, so in an `Int8` bitmap the element `-1` is compared as `255` instead of as the sign-extended `4294967295`; this also fixes `bitmapMin`, `bitmapMax`, `bitmapContains` and `bitmapTransform` on bitmaps that have grown past the small representation. Queries that passed sign-extended thresholds have to be adjusted: over an `Int8` bitmap, `bitmapSubsetInRange(bm, 4294967168, 4294967296)` becomes `bitmapSubsetInRange(bm, 128, 256)`. Also fixed `groupNumericIndexedVector` returning different results for `Int8` and `Int16` index columns than for wider index types, and `numericIndexedVectorGetValue` returning `0` for negative indexes. Corrected the bitmap function documentation, including the subset functions that were described as using 1-based indexing and the signed `bitmapBuild` / `bitmapToArray` support. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1309` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110072",
        "createdAt": "2026-07-11T01:18:01Z",
        "updatedAt": "2026-08-13T09:33:43Z",
        "timestamp": "2026-08-13T09:33:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "RamiDarwiche",
        "state": "closed",
        "assignees": [
          "yakov-olkhovskiy"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110073",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allow creating a Distributed table over a table function",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/106189 Allow creating a `Distributed` table over a table function, in the same way the `remote`/`cluster` table functions already accept an arbitrary table expression. Previously `CREATE TABLE ... ENGINE = Distributed(...)` only accepted `Distributed(cluster, database, table[, sharding_key[, policy_name]])`, where `database` and `table` had to be string literals or identifiers. `StorageDistributed` itself already supported being backed by a table function (`remote_table_function_ptr`), but that capability was only reachable through the `remote`/`cluster`/`clusterAllReplicas` table functions. This change lets the engine take a table function directly: ```sql CREATE TABLE t ENGINE = Distributed(cluster, table_function()[, sharding_key[, policy_name]]); -- e.g. CREATE TABLE distributed_numbers ENGINE = Distributed(test_cluster_two_shards, numbers(100)); ``` mirroring the `cluster('cluster_name', table_function())` signature. The second engine argument is treated as a table function only when it is a call to a registered table function; any other expression is still interpreted as a database name, so the existing `Distributed(cluster, database, table, ...)` form is unaffected. The structure is inferred from the table function when the columns are omitted. `INSERT` into a table-function-backed `Distributed` table is rejected with `NOT_IMPLEMENTED`, since there is no concrete remote table to route rows to. As a side effect, `INSERT INTO FUNCTION cluster('cluster', table_function())` now fails with this clear error instead of silently building a broken `INSERT INTO <empty>` query for the shards. ### Relationship to #106189 This is complementary to #106189 (`Remote`/`RemoteSecure` storage engines): that PR adds address-based engines over an ad-hoc cluster, while this one lets the configured-cluster `Distributed` engine take a table function. The table-function `INSERT` guard in `StorageDistributed::write` is aligned with #106189 (identical exception message, guard placed only in `write`). The `read` path here additionally handles the case where the remote database/table names are empty (a table-function-backed `Distributed` stores empty names rather than the `system.one` placeholder that #106189's engines inherit from the `remote` table function); this also makes the shared `read` path robust for #106189's engines. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow creating a `Distributed` table over a table function, e.g. `CREATE TABLE t ENGINE = Distributed(cluster, numbers(100))`, in the same way as the `cluster`/`remote` table functions accept a table function. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110073",
        "timestamp": "2026-08-12T20:50:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 47
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110084",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add UUID2 data type with correct sorting",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/110066 Introduces `UUID2`, a variant of the `UUID` data type with correct (lexicographic) sorting. ## Motivation For historical reasons, the `UUID` data type sorts by the *second half* of the value. This is unexpected and, in particular, hurts the performance of primary indexes built on `UUIDv7` columns, whose most significant bits are a timestamp: with sorting by the second half, primary-key analysis cannot prune granules by the timestamp. ## What this does `UUID2` stores the 128-bit value as a plain big-endian integer of the 16 canonical bytes, so that natural integer comparison of the underlying value matches the textual (lexicographic) order and the canonical byte order used by most other systems. It reuses the `UUID` column and `Field` representation (like `DateTime` reuses `UInt32`), so sorting is correct with no extra comparison code, and its binary/interchange serialization is the canonical big-endian byte order. - `UUID1` is an alias of the current `UUID` type. - A new setting `uuid_type_version` (default `1`) controls whether the bare name `UUID` resolves to `UUID` (`1`) or `UUID2` (`2`) at `CREATE`/`ALTER` time. The resolved concrete type is materialized into the stored table definition (including nested types such as `Array(UUID)`), so reads never depend on the session setting and existing tables are never rewritten. The default will be flipped to `2` in a later, separate change. - Conversions to/from `String`, `UInt128`, `FixedString(16)` and `UUID`, plus `toUUID2` / `toUUID2OrZero` / `toUUID2OrNull`. - Parity across functions (`hex`/`bin`, `reinterpretAs*`, `UUIDv7ToDateTime`, `UUIDToNum`, `empty`/`notEmpty`, `min`/`max`, hashing, `uniq`), formats (`RowBinary`, `Native`, `JSON`, `CSV`, `TSV`, `Arrow`, `Parquet`, `Avro`, `BSON`, `MsgPack`, `Protobuf`, `CapnProto`, `JSONExtract`) and storage (`generateRandom`, `bloom_filter` skip index). The `UUID` type is unchanged (verified: still sorts by second half, identical conversions, default `uuid_type_version = 1`). ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new data type `UUID2`, a variant of `UUID` that sorts by its textual (lexicographic) representation instead of by the second half of the value. The setting `uuid_type_version` (default `1`) selects whether the type name `UUID` resolves to `UUID` or `UUID2`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110084",
        "createdAt": "2026-07-11T11:29:15Z",
        "updatedAt": "2026-08-13T13:19:35Z",
        "timestamp": "2026-08-13T13:19:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 37
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110102",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallelize aggregation-in-order via key-hash reshuffle (aggregation_in_order_shuffle)",
        "text": "## Description `optimize_aggregation_in_order` is not enabled by default because it is often slower than the default hash aggregation. The reason is structural: for a multi-stream read the whole aggregation is funneled through a single `FinishAggregatingInOrderTransform`, so it runs on ~2–6 cores regardless of `max_threads` and is 8–16× slower than the default parallel hash path for high-cardinality `GROUP BY` (it does *less* total CPU work, but cannot parallelize it). This PR adds an experimental setting **`aggregation_in_order_shuffle`** (default `0`) that removes that funnel by repartitioning, reusing the shuffle primitive introduced for the sharded aggregator (`BufferedShardByHashTransform`): ``` read N sorted streams -> scatter each stream by hash(GROUP BY keys) into num_shards (BufferedShardByHashTransform) -> per shard: MergingSortedTransform(N -> 1) -- groups are disjoint by key -> per shard: streaming AggregatingInOrderTransform + FinalizeAggregatedTransform ``` Each shard aggregates a disjoint set of keys, so there is no cross-shard coordination and no single-threaded merge, while each shard still streams out completed key groups (bounded, O(1) memory). It is only used when the `GROUP BY` order is not relied upon downstream (`!memoryBoundMergingWillBeUsed()`, no bucket-order requirement, no `LIMIT` push-down), and the output is therefore not ordered by the keys. ### Making the shuffle deadlock-free and bounded-memory The M per-shard merges share the N scatters, so a naive scatter deadlocks: a slow/exhausted lane blocks the shared scatter from feeding the others, and a per-shard *sorted* merge is a selective consumer. Two fixes in `BufferedShardByHashTransform` (unbounded mode, used only by this path): - **Demand-driven scheduling**: push to every ready lane first, and pull a new input chunk only to feed an output that is ready *and* starving (empty queue). The shared scatter never stalls on a slow lane (its data is just buffered), so there is no cross-lane cycle, and read-ahead — hence memory — is bounded by how far the fastest consumer runs ahead of the slowest. - **Per-output drain on EOF**: when the input is exhausted, finish each output whose queue is already empty, so a sorted merge gets EOF on exhausted inputs instead of waiting forever (this hung deterministically for non-overlapping parts). Verified: 0 deadlocks over 120+ runs across all cardinalities, overlapping and non-overlapping parts, and `max_threads` 8..128; results are byte-identical to the default for order-independent aggregates. ### Results (200M rows, 50M keys, 96 threads; median of 5) | GROUP BY | default | in-order (funnel) | **shuffle** | |---|---|---|---| | `sum`, 50M groups | 870 ms / 23 GB / 42 cores | 5779 ms / 0.8 GB / 4 cores | **1341 ms / 0.83 GB / 20 cores** | | `uniqExact`, 50M groups | 12040 ms / 56 GB | 5744 ms / 1.5 GB | **1049 ms / 2.0 GB / 22 cores** | For high-cardinality `GROUP BY` the shuffle is **4–11× faster than the current aggregation-in-order at the same O(1) memory**, and for `uniqExact` it is faster than the default while using ~28× less memory. ### Known limitation (why it stays off by default) It scatters raw rows, so it regresses for low/medium cardinality (where in-order is already fast and the shuffle gives no benefit, and long key-runs make the scatter buffer more). Gating it to high cardinality (e.g. from the primary-key granule estimate) is a natural follow-up. `EXPLAIN PIPELINE` is also fixed to render the scatter/merge stage of the in-order pipeline (it was previously omitted). ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added an experimental setting `aggregation_in_order_shuffle` that parallelizes `optimize_aggregation_in_order` by repartitioning the sorted input by the hash of the `GROUP BY` keys into independent shards, removing the single-threaded merge bottleneck while keeping the bounded memory of aggregation-in-order. For high-cardinality `GROUP BY` it is several times faster than the ordinary aggregation-in-order. Disabled by default. ### Documentation entry for user-facing changes - [x] Documentation is written (the new setting is documented in `src/Core/Settings.cpp`).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110102",
        "createdAt": "2026-07-11T16:31:59Z",
        "updatedAt": "2026-08-13T07:25:50Z",
        "timestamp": "2026-08-13T07:25:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110104",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check for cancellation in AggregatingInOrderTransform",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107941 --> Related: https://github.com/ClickHouse/ClickHouse/issues/107941 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `optimize_aggregation_in_order` ignoring query cancellation. `AggregatingInOrderTransform` now checks for cancellation while aggregating a chunk, so a query stopped by `KILL QUERY` or by `max_execution_time` (in the default `timeout_overflow_mode = 'throw'`) stops promptly instead of running the whole chunk to completion. ### Description `AggregatingInOrderTransform::consume()` splits one input chunk into runs of equal keys in a loop. Query time and cancellation limits are only checked between pipeline steps (between `work()` calls), so a chunk with many distinct keys makes a single `consume()` call run for a long time (`O(distinct_keys)` iterations, each an `upper_bound` over the remaining rows) with no cancellation checkpoint. As a result a cancelled query (`KILL QUERY`, `max_execution_time`) using `optimize_aggregation_in_order` kept aggregating until the whole chunk was done; the connection thread then blocked in `PullingAsyncPipelineExecutor::cancel() -> ThreadFromGlobalPool::join()` waiting for that loop. The server-side AST fuzzer repeatedly hit this as `Hung check failed, possible deadlock found` (Stress test, all sanitizers), with `system.processes` showing `is_cancelled = 1` and `elapsed` far past the 90s hung-check window while the worker thread sat in `AggregatingInOrderTransform::consume -> Aggregator::executeImpl`. The loop now checks `isCancelled()` once per key interval (cheap) and returns early; the partial aggregation state is discarded because the pipeline is being torn down. This mirrors the existing per-loop cancellation checks in `WindowTransform` and `FillingTransform`. Scope: this covers cancellation that sets `is_cancelled` on the pipeline, i.e. `KILL QUERY` and `max_execution_time` in the default `timeout_overflow_mode = 'throw'`, which is what the reproduced hung check hit (`system.processes` showed `is_cancelled = 1`). The non-default `break` mode is a soft limit that returns a partial result and never sets `is_cancelled`; honoring it mid-chunk (as `FillingTransform` does via `process_list_element->checkTimeLimit()`) is a separate partial-result change, out of scope here. The regular `AggregatingTransform` behaves the same way. Regression test `04512_aggregation_in_order_cancellation` forces one long `consume()` over 40M distinct-key rows in a single chunk, `KILL QUERY ... SYNC` once every row is read: with the fix the KILL returns in a fraction of a second, without it it blocks for the several seconds the loop needs to finish. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1344` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110104",
        "createdAt": "2026-07-11T17:17:10Z",
        "updatedAt": "2026-08-13T16:41:03Z",
        "timestamp": "2026-08-13T16:41:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "yakov-olkhovskiy"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110105",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add clickhouse-proxy application mode",
        "text": "Implements a new application mode in the `clickhouse` binary, named `proxy`. The proxy accepts connections over end-user ClickHouse protocols, finds the upstream backend based on configurable rules (hostname from TLS SNI or HTTP header, user name, database name, and — for HTTP — query type), and forwards the traffic to it. It is built on the `silk` fiber framework so that many concurrent connections are handled with low RAM usage. Start it with `clickhouse proxy --config-file proxy_config.xml`; a fully commented example config is in `programs/proxy/proxy_config.xml`, and the feature is documented at `docs/en/operations/clickhouse-proxy.md`. **What it does** - Protocol frontends for HTTP(S), native TCP, MySQL, PostgreSQL, transparent TLS-by-SNI, and opaque TCP streams. It handles only end-user protocols (not Keeper or inter-server replication). The user name and database are parsed from the first packets where the protocol allows it (HTTP headers/params/Basic auth, native `Hello`, PostgreSQL `StartupMessage`, HTTP query type). MySQL is server-speaks-first with in-band TLS, so it is forwarded transparently and routed by peer address or the default pool. - TLS routing options: terminate and re-encrypt (the proxy and each backend hold their own certificates), terminate only (unwrap: TLS to clients, plaintext to backends), and transparent (route by SNI without decrypting). Optional ACME certificate provisioning, as in `clickhouse-server`. - Multi-criteria routing rules matching host / user / database / query type / protocol, by exact value or regular expression; regexp captures can be substituted into a backend address template (e.g. route users `ch-<tenant>` to per-tenant backends). - Pools with pluggable load balancing (`random`, `round_robin`, `least_connections`, `lowest_latency`, `least_resources`) behind a common interface; a pool may list several backends for load balancing. - Session stickiness by `session_id` (from the HTTP URL) or by peer address, using a consistent (rendezvous) hash of the backends. - Backend health monitoring (connect latency, consecutive failures) and optional CPU/memory polling with per-backend credentials, feeding the `least_resources` strategy; a JSON status endpoint, `/ping`, and static pages are served by the proxy itself. - An abstract routing table exposing hooks (unknown route, no backends available, first time a user or database is seen) that run a shell command — for example to provision a backend on demand and wait for it to become available. Validated end to end against a live server: HTTP `/ping`, static pages, the status endpoint, GET/POST forwarding, the no-backend error paths, native protocol round-trips, and clean shutdown. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new `proxy` mode to the `clickhouse` binary: a lightweight, protocol-aware proxy that routes end-user connections (HTTP, native, MySQL, PostgreSQL, TLS-by-SNI, and raw TCP) to backend servers based on a configurable routing table, with TLS termination/passthrough, load balancing, health checks, and session stickiness. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110105",
        "createdAt": "2026-07-11T17:34:42Z",
        "updatedAt": "2026-08-13T17:46:26Z",
        "timestamp": "2026-08-13T17:46:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-feature",
          "submodule changed"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110127",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Framing formats: multiplex data, totals, extremes, progress, logs, and profile events in the HTTP response stream",
        "text": "A framing format multiplexes different response parts of the query in a single stream: chunks of data, totals and extremes, progress packets, profile events (metrics), and server logs — everything that the native protocol supports. This allows rich data exchange in the HTTP protocol. Framing formats are independent of output formats: they encapsulate bytes produced by any output format, by separating and potentially encoding these chunks of bytes. The concatenation of the payloads of all `data`, `totals`, and `extremes` packets is exactly what the output format would have produced without framing. Auxiliary packets (progress, logs, profile events, exceptions) are represented as JSON. The framing format is selected by the new query setting `framing_output_format` (BETA tier). It currently applies to the HTTP protocol and is ignored for other interfaces. The implemented framing formats: - `None` — transparently routes everything applicable (data, totals, extremes, progress) to the output format, and ignores everything that is not applicable (metrics, logs), so everything works as it is by default. - `EventStream` — frames packets as HTTP server-sent events (`text/event-stream`). It integrates with the HTTP protocol and throws an exception when not applicable. - `JSONEachPacketBase64` — every packet is a JSON object on a separate line; the formatted data is base64-encoded (suitable for binary output formats). - `JSONEachPacketString` — every packet is a JSON object on a separate line; the formatted data is put into a string. Example: ``` $ curl \"http://localhost:8123/?framing_output_format=EventStream\" -d \"SELECT number FROM numbers(3) FORMAT JSONEachRow\" event: progress data: {\"read_rows\":\"3\",\"read_bytes\":\"24\",\"total_rows_to_read\":\"3\",\"elapsed_ns\":\"684907\"} event: data data: {\"number\":0} data: {\"number\":1} data: {\"number\":2} event: profile_events data: {\"host_name\":\"localhost\",\"current_time\":\"2026-07-11 22:38:21\",\"thread_id\":\"0\",\"type\":\"increment\",\"name\":\"SelectedRows\",\"value\":\"3\"} ... ``` Implementation: a framing format works as a multiplexor. The output format writes into the framing format's payload buffer, and `IOutputFormat` notifies the framing format on packet boundaries (under its writing mutex), which wraps everything accumulated since the previous boundary into a packet of the corresponding kind. Progress is routed through the existing throttled concurrent-progress path. Server logs (with the `send_logs_level` setting) and profile events reuse `InternalTextLogsQueue` and `ProfileEvents::getProfileEvents` — the same mechanisms as the native protocol — newly attached for HTTP queries. Exceptions are always written as the last packet of the stream (regardless of `http_write_exception_in_output_format`), so the client can always parse the response as a stream of packets. Parallel formatting is not used when framing is enabled, because the framing format needs to know the packet boundaries. Processing of multiple queries at once is out of scope of the first implementation, but the design allows it: every packet can be extended with the information about the query index along multiple queries. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add framing formats, selected by the new setting `framing_output_format`: they multiplex different response parts of the query in a single HTTP response stream — chunks of data, totals and extremes, progress packets, profile events, server logs, and exceptions. Implemented framing formats: `None` (default, everything works as before), `EventStream` (HTTP server-sent events), `JSONEachPacketBase64`, and `JSONEachPacketString` (a JSON object per packet with base64-encoded or string data). - [x] Documentation entry for user-facing changes: `docs/en/interfaces/framing-formats.md`",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110127",
        "createdAt": "2026-07-11T22:45:02Z",
        "updatedAt": "2026-08-13T03:10:18Z",
        "timestamp": "2026-08-13T03:10:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110130",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Automatically choose between the plain and the secure port in clickhouse-client",
        "text": "When neither `port` nor `secure`/`no-secure` is specified, `clickhouse-client` now probes both the default port 9000 and the secure port 9440 concurrently and uses the protocol of the port that answers first. A server that answers on one port only is connected to without waiting out the connect timeout of the other one, and a server that answers on both is connected to over either of them, since both work (TLS wins a tie, which costs no waiting). This makes the following work out of the box: ``` clickhouse-client --host play.clickhouse.com --user play ``` On `play.clickhouse.com` (and many similarly firewalled servers) the plain port is silently dropped rather than refused, so probing the ports sequentially would stall for the whole connect timeout before TLS could even be attempted — hence the concurrent probe. The addresses of a port, in contrast, are attempted one at a time, the next one only after 250 milliseconds without an answer - the \"Connection Attempt Delay\" of RFC 8305 (Happy Eyeballs) - so a hostname that resolves to several reachable backends is not connected to on all of them at once. The connection the probe establishes to the port it chooses is then handed over to the client instead of being discarded. Together, these keep the automatic choice from leaving sessions that never send anything on the server, which it logs as `Client has not sent any data.` and counts against `max_connections`. When the secure port is the one that answered, the client connects with TLS and the interactive banner shows `Connecting to play.clickhouse.com:9440 (secure) ...`. Because the protocol here is chosen rather than requested, a port that turns out to be unusable is not an error: the client falls back to the other one. It matters the most for the secure port, whose common failure is a self-signed or otherwise untrusted certificate that every client not passing `--accept-invalid-certificate` rejects: the plain port is what the client would have connected to if there were no automatic choice at all, so nothing is taken away from the user. Pass `--secure` to require TLS. The fallback works in the other direction too: when the plain port is the one that answered but the connection to it then fails at the protocol level (e.g. a TLS-terminating proxy in front of the plain port), the secure port is tried. Explicit `--port`, `--secure`, `--no-secure` (on the command line, in the configuration file, or in connection credentials), ClickHouse Cloud hostnames (which already default to TLS), and builds without TLS support bypass the detection entirely. Covered by unit tests for the prober and a new integration test `test_client_auto_secure_port` that reproduces the firewalled-server scenario with `iptables` REJECT/DROP rules: it asserts that the DROP case (the `play.clickhouse.com` one) connects quickly instead of waiting out the connect timeout, that an untrusted certificate falls back to the plain port, and that the probe leaves no extra connection behind on the server. The functional tests now say which protocol the run uses (`--no-secure` for a plain run, symmetrically to the `--secure` a secure run already passes), because the test server listens on both ports. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When neither `port` nor `secure` is specified, `clickhouse-client` tries both the default port 9000 and the secure port 9440 concurrently and uses the one that answers first, so `clickhouse-client --host play.clickhouse.com --user play` connects over TLS without `--secure`, even though the plain port of that server is silently dropped rather than refused. If the port that answered turns out to be unusable — a secure port with an untrusted certificate, for example — the client falls back to the other one, because the protocol was not requested explicitly. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110130",
        "createdAt": "2026-07-12T01:17:48Z",
        "updatedAt": "2026-08-13T07:28:40Z",
        "timestamp": "2026-08-13T07:28:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110144",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support per-authentication-method GRANTS clause in CREATE USER and ALTER USER",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109117 Implements per-authentication-method grant limits, the first part of the linked issue: ```sql ALTER USER vasya ADD IDENTIFIED WITH password BY 'WYmdFyas8PftrHbHQQo8' VALID UNTIL '2026-12-31' GRANTS (SELECT ON db.table) ``` When a user logs in with such a method, the access rights of the session are the intersection of the user's access rights (including granted roles) with the listed elements. The clause never adds rights: a listed privilege that is not granted to the user stays unavailable. This provides a way to create tokens for applications: an additional credential with an expiration date and a limited set of grants, which is tied to the user — it is displayed in `query_log` and `processlist` as the user, stops working if the user is deleted, and is narrowed when the user loses grants. Details: - The clause is parsed after the per-method `VALID UNTIL`, works with `CREATE USER`, `ALTER USER [ADD] IDENTIFIED`, and `NOT IDENTIFIED`, is shown by `SHOW CREATE USER`, and persists through the SQL serialization of access entities (and therefore backups). Elements without a database name are bound to the current database when the query is interpreted. - The intersection erases all grant options (the clause cannot express them), so such sessions cannot `GRANT` anything. Role administration is denied entirely (fail-close), including per-role admin option. `EXECUTE AS`, `ALTER USER`, `CREATE USER` and similar escapes require the corresponding rights to be listed explicitly and granted to the user. - Reattaching to a named session (`session_id`) with a different credential re-applies the limit of the credential used by the new connection, so a limited credential cannot pick up the full rights of a session created by an unrestricted one. - The limit is captured at login: `ALTER USER` affects new sessions, not established ones (same as `VALID UNTIL`). - A new `auth_grants` column in `system.users` exposes the limit of each authentication method. - `ContextData`'s copy constructor now preserves `external_roles` and the new field: previously a recalculation of access rights on a copied context silently dropped external roles; for the new field that would mean silently widening the rights. Known limitations (consistent with `VALID UNTIL` and external roles): the limit is not propagated to other nodes of a cluster in distributed queries or `ON CLUSTER` DDL (the query is checked on the initiator), and the clause is not available in `users.xml`. The `CREATE TOKEN` syntactic sugar and the separate grant for self-service `ADD IDENTIFIED` mentioned in the issue are left for a follow-up. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Authentication methods in `CREATE USER` and `ALTER USER ... ADD IDENTIFIED` support a `GRANTS (SELECT ON db.table, ...)` clause which limits the access rights of sessions authenticated with that method to the intersection with the listed grants. This allows using additional credentials as tokens for applications: `ALTER USER vasya ADD IDENTIFIED WITH password BY '...' VALID UNTIL '2026-12-31' GRANTS (SELECT ON db.table)`. Closes [#109117](https://github.com/ClickHouse/ClickHouse/issues/109117). ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110144",
        "createdAt": "2026-07-12T05:16:00Z",
        "updatedAt": "2026-08-13T15:43:33Z",
        "timestamp": "2026-08-13T15:43:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110171",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add VALID FOR clause for users and credentials",
        "text": "We support `VALID UNTIL` for users and credentials. This adds `VALID FOR <interval>` as a convenience shorthand. Instead of an absolute date and time, `VALID FOR` accepts an interval, and the expiration deadline is computed as the current time plus that interval at the moment the query is executed. The result is stored in the `VALID UNTIL` form, so `SHOW CREATE USER` always displays the resolved absolute deadline. It can be used everywhere `VALID UNTIL` can - at the user level and per authentication method, in both `CREATE USER` and `ALTER USER`. Examples: ```sql CREATE USER u1 VALID FOR INTERVAL 1 DAY; CREATE USER u2 IDENTIFIED WITH plaintext_password BY 'x' VALID FOR INTERVAL 3 MONTH; ALTER USER u1 VALID FOR INTERVAL 1 DAY + INTERVAL 12 HOUR; ``` Implementation notes: the current time is injected as a literal (rather than using `now`, which is non-deterministic and would not fold to a constant expression), and the deadline is computed with `toDateTime64` so that large intervals saturate at the `DateTime64` upper bound instead of overflowing the year-2106 boundary of `DateTime`. Note on the system table schema: to represent deadlines beyond the year 2106 exactly, the `valid_until` column of `system.users` changes from `Array(DateTime)` to `Array(DateTime64(0))`. This is a backward-incompatible change of a documented system table for tooling that introspects the column type; the values themselves keep second precision. ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added the `VALID FOR <interval>` clause to `CREATE USER` and `ALTER USER` as a shorthand for `VALID UNTIL`. The expiration deadline is computed as the current time plus the given interval at query execution time and stored in the `VALID UNTIL` form. The `valid_until` column of the `system.users` table now has the type `Array(DateTime64(0))` instead of `Array(DateTime)`, so that deadlines beyond the year 2106 are represented exactly; tooling that reads this column should handle the new type. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110171",
        "createdAt": "2026-07-12T18:12:03Z",
        "updatedAt": "2026-08-13T11:20:58Z",
        "timestamp": "2026-08-13T11:20:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-backward-incompatible",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "antaljanosbenjamin"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110180",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check access rights in EXPLAIN QUERY TREE and EXPLAIN SYNTAX",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/78938 `EXPLAIN QUERY TREE` and `EXPLAIN SYNTAX` (in the analyzer) resolve the query and dump table metadata such as column names and types, but unlike `EXPLAIN PLAN` they do not build a query plan. The `SELECT` access check that the planner performs in `prepareBuildQueryPlanForTableExpression` was therefore skipped, so a user with no privileges could read the column names and data types (including `Enum` element lists) of tables they are not allowed to access, while `SELECT`, `EXPLAIN PLAN` and `EXPLAIN PIPELINE` are correctly rejected. This adds a `SELECT` access check for every table referenced anywhere in the query tree (including tables inside subqueries in expressions such as `WHERE x IN (SELECT ... FROM t)`). The check mirrors the planner: it validates access to the columns that are actually read, with the trivial-count fallback (access is granted if at least one column is accessible) for queries that read no specific column, e.g. `SELECT count() FROM t`. As a result, a user with a column-level grant sees the same behavior as for a plain `SELECT` (`EXPLAIN QUERY TREE SELECT granted_col FROM t` is allowed, `... other_col ...` is denied). `buildQueryTree` only builds the tree; table identifiers are bound to storages and columns are resolved by the query analysis pass. The check runs on the resolved tree: on `query_tree` directly when the passes already resolved it, or on a throwaway resolved copy otherwise (e.g. `EXPLAIN QUERY TREE run_passes = 0`, which intentionally dumps the unresolved tree, so the tree that gets dumped is unchanged). ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `EXPLAIN QUERY TREE` and `EXPLAIN SYNTAX` not checking access rights on the referenced tables, which allowed a user without privileges to read table column names and data types. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110180",
        "createdAt": "2026-07-12T19:53:08Z",
        "updatedAt": "2026-08-13T01:18:57Z",
        "timestamp": "2026-08-13T01:18:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-must-backport",
          "pr-critical-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110183",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix use-of-uninitialized-value in WITH FILL suffix over a merge",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107074 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-of-uninitialized-value in `ORDER BY ... WITH FILL ... STALENESS` when the fill suffix generates no rows (for example when the staleness window leaves nothing to fill) and the result is read through a merge of sorted streams (such as a two-shard `Distributed` table). ### Description Found by the AST fuzzer under MSan on top of #107074. `FillingTransform`, when all input chunks are processed, may run the suffix path. If the fill constraints are satisfied but no fill rows are produced (e.g. `WITH FILL ... STALENESS` with an exhausted staleness window), `generateSuffixIfNeeded` returns `true` while the result columns are freshly `cloneEmpty()`'d and carry no data. Previously the transform still emitted a 0-row chunk built from those empty columns. A downstream `MergingSortedTransform` (as used when reading from a two-shard `Distributed` table) then built a sort cursor over that empty chunk and compared row 0, reading past the end of the empty column. Fix: do not emit the suffix chunk when it has no rows. Minimal reproducer (needs a two-shard merge on the initiator): ```sql CREATE TABLE m (key Int) ENGINE = Memory; INSERT INTO m VALUES (100); CREATE TABLE d2 AS m ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), m); SELECT _shard_num FROM d2 ORDER BY _shard_num ASC WITH FILL TO 46 STALENESS 1; ``` CI finding: `AST fuzzer (amd_msan)`, report https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107074&sha=3ec54c88d4eb535a5d644fc1ab91af31d717f9a5&name_0=PR&name_1=AST%20fuzzer%20%28amd_msan%29 MSan use-of-uninitialized-value in `ColumnVector<UInt32>::doCompareAt` (`MergingSortedAlgorithm::consume`), origin `FillingTransform::initColumns` (`cloneEmpty`) via the suffix path. Verified on a local `amd_msan` build: reproduces before the fix (server aborts), clean after; the new regression test returns `1\\n2`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110183",
        "createdAt": "2026-07-12T21:10:52Z",
        "updatedAt": "2026-08-13T17:26:43Z",
        "timestamp": "2026-08-13T17:26:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110199",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Web UI: stop the progress spinner when a running query is cancelled",
        "text": "Pressing **Stop** while a query is running keeps the progress spinner (the hourglass in the Web UI toolbar) spinning indefinitely. `cancel()` invalidates the in-flight request (bumps `request_num`) and aborts the fetch, so the awaiting `postSingle`/`postImpl` path bails out early without calling `finish()`/`clear()` — nothing hides the spinner. This fix stops it explicitly in `cancel()`, via a new `stop()` method that hides the hourglass without showing the success check mark (the query was cancelled, not completed). ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the progress spinner in the Web UI (Play) continuing to spin after a running query is cancelled with the Stop button. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110199",
        "timestamp": "2026-08-12T21:35:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110210",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Web UI prototype for framing formats",
        "text": "Prototype changes to the Web UI (`programs/server/play.html`) to test and showcase the framing formats feature. This is a draft for experimentation and review of the client-side experience, not intended to merge as-is. It builds on the framing-formats server work in #110127 (this branch is based on that PR, so the diff includes those commits until it merges into `master`). What the prototype does: - Streams every query with the `EventStream` framing format over HTTP (`framing_output_format=EventStream`, `send_logs_level=trace`), decoding data / progress / log / profile_events / exception packets as they arrive - so progress and logs work regardless of the output format (including `FORMAT Pretty`, binary formats via base64, and images). - Shows realtime CPU, memory, and disk usage like `clickhouse-client`, aggregated per host (total and max/host), switching to peak RAM when the query finishes. The meter state is owned by each tab: the CPU counters in `profile_events` packets are per-packet increments, so a query running in a background tab keeps accumulating them, and reopening the tab continues the meter from its live values (rather than restarting near zero). - Merges elapsed time, realtime metrics, and rows/bytes read stats into a single progress area with the progress-bar gradient rendered behind all of them; the text is tinted by a mask clipped from the same gradient so it stays readable over the fill. - Displays server logs in real time, colored like the client, with a Logs button available even on exception; the log view stays fast for 100k+ lines while keeping native browser search, bounding retention to the most recent 100k lines (older lines are dropped and counted in a marker line). The full response text kept for the tab/history snapshot is also capped as it is collected (at ~100 KB, the size above which the snapshot is dropped), so a single large framed data/log burst is not retained in full only to be discarded later. A failed run whose reply outgrew that cap is persisted as a compact snapshot that keeps the exception carrier in the form its framing kind replays (the terminal `event: exception` block, the `{\"packet\":\"exception\",...}` line, or the plain `{\"exception\":...}` line, captured at stream-read time since the capped text stops growing before the terminal exception arrives), so reopening or reloading a failed tab still shows the error reason. - Base64-decodes every `data`/`totals`/`extremes` payload (the `EventStream` wire encodes each block as a single base64 `data:` field and the `Content-Type` carries `payload=base64`), deciding the rendering by the output format: the default format is reassembled into the table, an image format (or bytes carrying an image signature) is rendered as an inline image once the stream completes, and any other format is decoded and shown incrementally as raw text. - Uses a framing-compatible default format for queries without an explicit `FORMAT` clause: the framed request asks for `JSONCompactStringsEachRowWithNamesAndTypes` (the framing rejects the in-band-progress `JSONStringsEachRowWithProgress`) and the client reassembles the compact rows into the table renderer's shapes; on the server side (#110127) the compact family emits totals and extremes under framing, so `WITH TOTALS` and extremes-based column coloring keep working on the default path. When such a format is shown as raw text instead (an explicit `FORMAT JSONCompactEachRow`), only the `data` packets are concatenated, so the rendered text is exactly the plain output of that format rather than one carrying the totals and extremes rows the format itself drops. - Propagates a framed `exception` packet as a query failure even when the HTTP status was already 200 (so `Run all` stops at the failed statement) - and, when a query chooses its own `JSONEachPacket*` framing that fails before the 200 OK header (coming back as a non-200 `application/x-ndjson` packet stream ending with a `{\"packet\":\"exception\",...}` line), shows those packets verbatim and records `framing_kind = 'ndjson_packets'` for replay rather than rendering the whole stream as one opaque error string; retries framing-incompatible explicit formats (e.g. `FORMAT JSONEachRowWithProgress`, `FORMAT Template`) once without framing for read-only queries (read-only-ness is resolved with a CTE-aware lexer walk - `WITH y AS (SELECT 1) INSERT INTO t SELECT * FROM y` is a write - shared with `Run all`'s grouping, so such a statement is also a barrier there and never runs in parallel with the reads that follow it; the port with regression coverage lives in `src/Parsers/tests/gtest_play_query_is_read_only.cpp`); a retried `JSON*EachRowWithProgress` format that itself reports a failure in-band - a trailing `{\"exception\":...}` object while the HTTP status stays 200 (`http_write_exception_in_output_format`) - is detected as a failure too, keyed off the output format rather than only a user-chosen `JSONEachPacket*` framing, so `Run all` stops after such a retried query fails; and does not add its own framing to a query that sets its own `framing_output_format` to a real framing choice (the response is then dispatched by content type, so the requested packets are shown verbatim); a query that sets `framing_output_format = 'None'` is refused client-side instead, since this page's rendering depends on framing (values the page cannot know upfront - `= DEFAULT`, a reset to the session/server default, and query-parameter placeholders like `= {fmt:String}` - are classified conservatively the same way and refused, rather than sent with a request shape the response might not match); likewise a standalone `SET framing_output_format = ...` is refused, because it would change the setting for the whole session (with a `session_id`) while the page keeps adding its own framing per request - a query-level `SETTINGS framing_output_format = ...` clause is the supported way to choose framing for one query. A query that carries its own `framing_output_format` is also refused for download (the setting would override the download's chosen `default_format`, so the file would be the framing packet stream), and if such a query fails, its history snapshot - the raw `JSONEachPacket*` packet stream - is replayed as raw text on tab-switch/reload rather than as one opaque error string. - Pins `framing_output_format=None` on every request that expects an unframed response (the plain/chart request, the compatibility retry, the download, and the panel/server-status/completion queries), rather than only omitting the setting - otherwise a framing carried by the connection URL or by the HTTP session behind it would frame those responses too. A query-level `SETTINGS framing_output_format = ...` clause is applied after the URL parameters, so a query that intentionally chooses a framing still overrides the pin. - Records the framing kind (`event_stream` / `ndjson_packets` / none) with each result snapshot and keys the history/tab replay off it, instead of guessing from the payload's first bytes - so a raw result whose text happens to start with `event:` (e.g. `SELECT 'event: data' FORMAT RawBLOB`) or `{\"packet\":` is not reparsed as a framing stream after a tab switch or reload. A snapshot recorded as `ndjson_packets` is replayed as raw text regardless of whether the run succeeded and of its underlying output format, matching the live path - so a successful user-framed `JSONEachPacket*` result whose format has its own restore path (a table or a `JSONCompactColumns` chart) is not reparsed as that format's JSON. The snapshot also records whether the framed stream was truncated, so replay keeps the live fail-closed behavior for images: a cut-off framed `FORMAT PNG` response that showed only its error live reopens from history / Back / Forward as that error too, never as a partially decoded picture (a failed but complete stream - a terminal `exception` packet after the payload - still renders its collected image, as live). - Detects both the query's `FORMAT` clause and its `framing_output_format` setting with the WASM lexer rather than a raw text match, so a mention inside a string literal or a comment - e.g. `SELECT 'FORMAT JSONCompactColumns'` - does not make the page silently opt out of its own framing. The detection is positional, not keyword-adjacent: a settings context is recognized by its `name = value` list grammar (a column merely named `settings` does not open one), and a `FORMAT` clause candidate must follow a token that ends an expression (so in `WITH 1 AS format SELECT format JSONCompactColumns SETTINGS max_threads = 1` both `format` words are identifiers, not a clause). The download reuses the same `FORMAT`-clause detection (now returning the clause span) to strip only a real trailing `FORMAT` clause from the download query, leaving text or ordinary SQL like `SELECT 'FORMAT TSV' AS s` untouched. Both detectors also accept a quoted spelling of the name - the server parses setting names and `FORMAT` names with identifier parsers, so a backquoted `framing_output_format` or format name is real - comparing by the unquoted name. Regression coverage for both lives in `src/Parsers/tests/gtest_play_detect_explicit_format.cpp` and `gtest_play_detect_framing_setting.cpp` (ports of the token walking onto the real `DB::Lexer`). - Dispatches on the response output format case-insensitively. Format names are case-insensitive in ClickHouse (`FormatFactory` looks them up by their lowercased name), while `X-ClickHouse-Format` echoes the identifier exactly as the `FORMAT` clause spelled it, so `FORMAT jsoncompactcolumns` used to lose the chart renderer, a lowercased default format lost the table renderer, and the late in-band exception probes did not recognize `FORMAT xml` / `FORMAT json` / `FORMAT jsoneachrowwithprogress` - a query failing after its `200 OK` header was then reported as a success and `Run all` continued past it. Every dispatch now compares a lowercased copy of the format name. - On the server side (#110127), framed pulling `SELECT` queries now also end with the documented final `progress` packet carrying `result_rows` / `result_bytes` / `memory_usage`, matching the native protocol and the no-result path. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Prototype Web UI changes to test framing formats (draft). Related: https://github.com/ClickHouse/ClickHouse/pull/110127",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110210",
        "createdAt": "2026-07-13T04:17:42Z",
        "updatedAt": "2026-08-13T03:34:09Z",
        "timestamp": "2026-08-13T03:34:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 23
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110230",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix S3Queue shutdown hanging on a streaming pipeline stuck inside a blocking call",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/103126 ## Problem `S3Queue`/`AzureQueue` shutdown blocks for as long as an in-flight streaming pipeline stays stuck inside a blocking call — a stalled object storage read, a mutex convoy on the shared file iterator, or an executor thread parked in `epoll_wait`. Consequences observed in production (a pipeline wedged for two hours): - A synchronous `DROP TABLE` on the queue table hangs, and everything serialized behind it hangs too. - Plain server shutdown / `DETACH TABLE` hangs the same way. Independently, the periodic metadata cleanup deletes `processing/` znodes and bucket locks older than `persistent_processing_node_ttl_seconds` (default 3600) by mtime alone, even while the owning streaming execution is still running in the same process. The execution's commit then fails with `Coordination::Exception: Transaction failed ... (No node)`, and the freed bucket lock can be acquired by another server concurrently (Ordered-mode correctness violation). ## Root cause `StorageObjectStorageQueue::shutdown` sets `shutdown_called` and calls `task->deactivate()`, which blocks until the in-flight `streamToViews` returns. The streaming pipeline runs `CompletedPipelineExecutor::execute` with no cancel callback, so all shutdown handling is cooperative at chunk boundaries inside `ObjectStorageQueueSource::generateImpl` (what #103126 merged). A pipeline parked inside a blocking call never reaches a chunk boundary, so shutdown blocks unboundedly. The cleanup TTL (`ObjectStorageQueueMetadata::cleanupPersistentProcessingNodes`) is meant to reap nodes orphaned by dead servers, but it has no notion of ownership: a node whose owner is alive in this very process is deleted just the same once its mtime passes the TTL. ## Background: why #103126's first draft dropped executor-level cancellation, and why it is safe now The first draft of #103126 used exactly this approach (`setCancelCallback` in `streamToViews`), and the review rejected it — not as wrong in principle, but because two correctness gaps had no answer at the time: 1. **Silent success after cancel → data loss** ([review comment](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3193288074): \"reading is safe, but AFAICS writing the data is not\"). `executor->cancel()` can land after the source finished reading but before the sink finalized; `CompletedPipelineExecutor::execute` then returns without an exception, and `commit(insert_succeeded=true)` marks files `Processed` whose rows were never written. 2. **Forced failed-commit → duplicate inserts.** The draft's counter-measures (\"skip commit when `shutdown_called`\", then a `cancel_was_triggered` flag) were each found racy ([1](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3195586146), [2](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3196025042)): a cancel landing between files forces an already-fully-`Processed` batch into the failed-commit path, so those files are reset and re-read — duplicating rows when deduplication is off. Two more factors sealed it: Kafka's identical fix (#100388, `setCancelCallback` + skip offset commits on cancel) had been reverted 18 days earlier (#101646) after making `test_kafka_commit_on_block_write` flaky, and the [proposed alternative](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3193311270) — a chunk-boundary throw inside `generateImpl`, where per-file state is precisely known — was much smaller to review. That is what merged. It solves \"shutdown waits for the in-flight file to reach EOF\" (minutes), but not \"the pipeline never reaches a chunk boundary at all\" (hours, this PR's production case). Both gaps are closed here, on top of the machinery the merged #103126 itself introduced: - Gap 1: when the cancel callback has fired, `streamToViews` **always** throws instead of committing success — the silent-success path no longer exists. - Gap 2: the cancel gate requires `table_is_being_dropped || is_deduplication_v2`, so a forced retry is either moot (drop — no retry) or absorbed by deduplication; the worst case is an extra read of the file, which the #103126 review itself deemed acceptable ([comment](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3213509835)). The failure direction also flips: the draft's races erred toward *not committing what was written* (loss), this PR's over-triggering errs toward *reset-for-retry under dedup* (at most a re-read). - The Kafka precedent does not transfer: Kafka's offset semantics have no deduplication, so cancellation must choose between loss and duplication; `S3Queue` since 26.2 has `deduplication_v2`, plus the `Cancelled`-not-`Failed` file state machine that the merged #103126 built — the very foundation that makes executor-level cancellation safe now. ## Solution 1. `streamToViews` installs `CompletedPipelineExecutor::setCancelCallback` (1 s poll), gated exactly like the chunk-boundary abort: `shutdown_called && (table_is_being_dropped || is_deduplication_v2)`. Without deduplication, a mid-file abort would duplicate already-inserted rows on retry, so dedup-off plain shutdown keeps processing the in-flight file to EOF, as before. `executor->cancel()` cancels the processors and wakes the polling queue, and the in-flight file lands in `Cancelled` state (reset for retry, not `Failed`) through the existing machinery. A cancelled pipeline that finishes without an exception is routed through the failed-commit path instead of being committed as successful, because the sink may not have finalized. **Scope of the cancellation.** `executor->cancel()` wakes the executor's own waits directly: `PollingQueue::finish` writes the self-pipe the driver thread sleeps on — which is exactly the shape of the two-hour production hang (the driver parked in `epoll_wait` with no worker running or blocked, frame-proven from `trace_log`). A worker inside a blocking storage call is not preempted; that call stays bounded by the client's per-socket-operation timeouts, and cancellation takes effect at the first check after it returns — what this PR removes is the previously unbounded continue-to-EOF/next-files work after that point. A blocking call that never returns (e.g. a dead connection outliving every socket timeout) needs an end-to-end per-request deadline; that is deliberately out of scope here and tracked as a follow-up. 2. Local executions register their node paths in a new `ObjectStorageQueueLocalActiveNodes` registry *before* creating them in Keeper (`trySetProcessing`, `prepareSetProcessingRequests`, bucket acquisition), backing off when registration is refused; the cleanup wraps its revalidate-and-remove window in a try-only removal lock that fails while the path is registered. Registration and the removal lock exclude each other under one mutex, so the cleanup can never delete a node owned by a live local execution, in any interleaving — the register-before-create ordering matters because a node deleted and recreated between the cleanup's revalidation read and its remove restarts at Keeper version 0, so no captured version can guard the removal; a recreated node is additionally recognized by its fresh mtime on the revalidation read. The registry is shared between sibling tables on the same keeper path, because `ObjectStorageQueueMetadataFactory` shares the whole `ObjectStorageQueueMetadata` object per `{zookeeper_name, zookeeper_path}`. Another server's cleanup can still delete a >TTL node of a live remote execution — inherent to the mtime-TTL design and unchanged here. Regression tests: a new `object_storage_queue_park_in_generate` failpoint parks a source mid-file until pipeline cancellation or failpoint disable; integration tests cover `DROP TABLE` with a stuck pipeline, both clauses of the cancel gate (`DETACH` with dedup on/off), live-node survival across TTL cleanup runs, and recreation inside the cleanup's collection-to-removal window (via a pauseable failpoint); a gtest pins the registry fencing protocol. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `S3Queue`/`AzureQueue` shutdown and `DROP TABLE` hanging for as long as a streaming pipeline stayed stuck inside a blocking call. Also fix the periodic metadata cleanup deleting the processing nodes and bucket locks of executions still running on the same server, which made their commits fail once `persistent_processing_node_ttl_seconds` elapsed. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110230",
        "createdAt": "2026-07-13T09:09:59Z",
        "updatedAt": "2026-08-13T16:11:22Z",
        "timestamp": "2026-08-13T16:11:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110283",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Compose join-order statistics over parts surviving partition/PK pruning",
        "text": "Closes: [https://github.com/ClickHouse/ClickHouse/issues/110281](<https://github.com/ClickHouse/ClickHouse/issues/110281>) ### Changelog category (leave one): * Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](<https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md>) of the changes that goes into CHANGELOG.md): Join-order cardinality estimation now composes column statistics over the parts surviving partition/PK pruning instead of all active parts. Previously a query pruned to a small partition was planned against table-wide statistics (observed 2500x row overestimation flipping the hash-join build side), and disabling statistics paradoxically produced a better plan. Also fixed: lazy FINAL was silently disabled for MergeTree relations under a JOIN because the join-order optimizer's memoized index-analysis result was mistaken for an applied projection. ### Documentation entry for user-facing changes Not needed: no user-visible interface changes; the fix makes existing settings (`use_statistics`, `use_statistics_cache`, `query_plan_optimize_lazy_final`) behave as documented. --- **Symptom.** With `use_statistics = 1` (default), `estimateReadRowsCount` in `optimizeJoin.cpp` estimated relation sizes over **all** active parts, ignoring partition/PK pruning. On the reproducer from ClickHouse/ClickHouse#110281 (fact: 5M rows in `p=1` + 1k rows in `p=2`, dim: 100k, `WHERE p = 2`): fact side estimated at 2,500,500 rows (= 5,001,000 × 1/NDV(p), a 2,500× error; NDV error 250,000×), identical under `use_statistics_cache = 0/1`. With statistics *disabled* the index-based fallback calls `selectRangesToRead()` and estimates 1,000 rows exactly — i.e. enabling statistics made the plan \\~100× worse by the optimizer's own cost model. **Root cause.** At `optimizeJoin` time `analyzed_result_ptr` is not populated yet, so `ReadFromMergeTree::getParts()` falls back to `prepared_parts` (all parts) when building the `ConditionSelectivityEstimator`. The table-wide `cached_estimator` (populated by the background `refreshStatistics()` task over all active parts) matched that wrong scope, which masked the cache/query-scope divergence. **Fix.** 1. `estimateReadRowsCount` obtains the partition/PK analysis result up front (`getAnalyzedResult()` or `selectRangesToRead()` — exactly what the non-statistics fallback branch already did) and passes it to a new `getConditionSelectivityEstimator(required_columns, analyzed_result)` overload, so statistics are folded over `parts_with_ranges`. 2. `MergeTreeData::getConditionSelectivityEstimator` returns the table-wide cached estimator only when the requested part set matches the set the cache was built over (allocation-free `isStale(RangesInDataParts)` overload; the shared_ptr is copied under `stats_mutex`, the comparison runs outside — the estimator is immutable once published). Without this, fixing (1) would make the previously-masked cache scope divergence real. 3. `optimizeLazyFinal`: the \"projection was applied\" guard checked for a mere non-null analysis result. `selectRangesToRead()` memoizes its result on the reading step, so join-order estimation (the no-statistics fallback before this PR, the statistics path as well after it) made lazy FINAL silently bail for any MergeTree relation under a JOIN. The guard now checks `readFromProjection()`. 4. `optimizeLazyFinal` also re-ran `selectRangesToRead()` unconditionally after the guard, repeating the full part/PK/skip-index analysis that join-order estimation had already memoized (observed: two `SelectExecutor` \"Key condition\" passes over the same parts in one stats-enabled `ReplacingMergeTree FINAL JOIN` query). It now reuses the memoized result — analysis passes drop 2 → 1; PK conditions are pushed in the first optimization pass, before both consumers, so the repeated analysis was provably identical. 5. Early index-analysis exits that prove the read empty now preserve the exact-zero invariant; join ordering short-circuits to zero rather than degrading to unknown when no estimator exists (`WHERE 0`: `f ⋈ d` before, `f[0] ⋈ d[0]` after). 6. With throwing `max_rows_to_read` or `max_rows_to_read_leaf`, join estimation uses a non-memoized range analysis without row-limit checks. The final read repeats analysis after `optimizeReadInOrder`; it remains limited unless InOrder is selected. **Scope note.** Statistics are composed over surviving *whole parts*; mark-range granularity inside a surviving part is not used (follow-up material, see the issue). Filtering by skip indexes deferred to the scan by `use_skip_indexes_on_data_read = 1` is likewise not reflected in join-order estimates: join order is fixed before those results exist. The fix relies only on partition/PK analysis remaining pre-scan. **Verification** (macOS arm64 debug build @ `5a9528b4db5`): * Reproducer from the issue: fact estimate 2,500,500 → **1,000** rows under both cache settings; the pruned query's plan cost becomes bit-identical (996.86) to the physically-pruned counterfactual table; build side flips to the small side. * Unpruned query with a warm cache still hits the cache (0 `Loading statistics` events) — no regression from the part-set check. * `04516_join_order_estimation_pruned_parts` (fails before the fix: `f[50500]` on 26.7.1.448; passes after: `f[1000]`), with an NDV oracle probe (`f[100]` = 1000 × 1/NDV(id) — derivable only from pruned column statistics, the index fallback yields `f[1000]`, so a silently degraded statistics path fails the test) and a PK-only pruning case (no partitioning; a part fully excluded by the primary key). * `04517_lazy_final_join_with_statistics` (new): `InputSelector` present for single-table FINAL and for FINAL under a JOIN with statistics on and off, plus result correctness. * `04518_statistics_cache_pruned_scope` (new): warms the table-wide cache via the real background refresh (deterministic retry on `LoadedStatisticsMicroseconds = 0`, the same oracle as `03707_statistics_cache`), then asserts: pruned query bypasses the cache and gets pruned-scope estimates; the unpruned cache hit is preserved afterwards. * Locally green: lazy_final suite (`03990`, `03991`; `03988`/`04092`/`04093` skipped as `no-debug`), `03707_statistics_cache`, `03279_join_choose_build_table_{,auto_}statistics`, `03788_statistics_part_pruning*`, `02864_statistics_*`, explain-pretty/join-order/estimate suites. **Cost/risk notes.** * The statistics path normally runs index analysis at planning time and **memoizes it** on the reading step (`selectRangesToRead()` writes `analyzed_result_ptr`), so execution reuses the analysis instead of re-running it. This matches what the no-statistics fallback already did before this PR; the new part is that the statistics path joins that behavior. Throwing read limits are the deliberate exception: estimator analysis stays local until `optimizeReadInOrder` determines whether the final read is exempt. * Pruned queries bypass the table-wide estimator cache and fold per-part statistics per query (`Loading statistics`). A part-set-keyed or per-part-decoded statistics cache is follow-up work (discussed in ClickHouse/ClickHouse#110281). * `tests/performance/join_planning_pruned_statistics.xml` (new): planning-only `EXPLAIN` benchmarks over 100- and 1000-part fact tables — per-part fold (cold cache, selective/unselective), a *genuinely warm* table-wide cache (`refresh_statistics_interval = 1` + warm-up; verified: zero statistics loads on the unpruned hit) exercising the O(parts) part-set comparison and the pruned bypass, and a planning-only lazy FINAL join over a 200-part ReplacingMergeTree guarding the duplicate-analysis regression. An end-to-end execution query is kept as a separate benchmark. CI perf compares merge-base vs head on a Linux release build. * If the pre-analysis pass ordering was a deliberate trade-off, happy to hear maintainer context — the issue discusses this explicitly.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110283",
        "createdAt": "2026-07-13T15:11:51Z",
        "updatedAt": "2026-08-13T00:19:28Z",
        "timestamp": "2026-08-13T00:19:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud",
          "pr-must-backport-synced",
          "v26.4-must-backport"
        ],
        "author": "skuznetsov-clickhouse",
        "state": "closed",
        "assignees": [
          "fkastrati"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110321",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support INSERT ... VALUES in the polyglot SQL dialect",
        "text": "Enable `INSERT ... VALUES` with inline data in the polyglot SQL dialect (`dialect = 'polyglot'`). Previously, running e.g. `INSERT INTO t VALUES (1), (2), (3)` with `dialect = 'polyglot'` failed with `Multi-statement queries are not supported in polyglot dialect mode`. The underlying problem is that transpiling inside the parser cannot deliver inline data to the executor: the transpiled buffer is transient, and the executor overwrites `ASTInsertQuery::tail` with the external input stream, so the inline-data pointers (`data`/`end`) must reference a live query buffer. This transpiles the query up front instead of inside the parser: - The server (`executeQuery`) transpiles a foreign-dialect query to ClickHouse SQL before parsing, keeps the transpiled text alive on the query context, and parses it with the standard parser. Inline INSERT data then points into a live buffer and is processed by the normal machinery. `SET` queries are still parsed as-is so `dialect`/`polyglot_dialect` can always be changed back. - The client (`clickhouse-client`/`clickhouse-local`) parses a non-ClickHouse-dialect query into an AST — which, for a foreign dialect, means transpiling it locally only to drive client-side handling (statement classification, output format, INSERT detection) — but then sends the *original* query text verbatim, without splitting off inline data. The server performs the authoritative transpilation whose result is actually executed, so inline INSERT data lives in a server-owned buffer and survives parsing. The client-side transpilation is throwaway; note this means the transpiler must also be available on the client (a client built without `USE_POLYGLOT` fails locally with `SUPPORT_IS_DISABLED`), and the client and server transpilers are assumed to agree — acceptable for this experimental dialect. Every parse-time setting the query was parsed under (`dialect`, `allow_experimental_polyglot_dialect`, `polyglot_dialect`, `allow_settings_after_format_in_insert`, `implicit_select`, and the parse limits `max_query_size`, `max_parser_depth`, `max_parser_backtracks`) is pinned in the per-query settings sent along with the verbatim text, so the query's own `SETTINGS` clause cannot change how the server reparses that same text (it still applies to the query's execution, and a `SET` still takes effect for subsequent queries). All changes are gated on the dialect, so ordinary ClickHouse INSERTs are unaffected. Validated over the HTTP interface, the native client, and `clickhouse-local` (multi-row and single-row `VALUES`, `INSERT ... SELECT`, and PostgreSQL literal transpilation such as `true`/`false`); `SET` passthrough and multi-statement rejection are preserved. External insert data combined with a foreign-dialect `INSERT` is rejected with `NOT_IMPLEMENTED` instead of being silently dropped, on both surfaces: the client rejects piped stdin and `INFILE` (it sends the query verbatim and cannot forward a data tail), and the server rejects a non-empty HTTP request body appended to a streaming `INSERT` (`POST /?query=INSERT ... &dialect=polyglot` with a body). A foreign-dialect `INSERT` is transpiled as a whole, so the body would go through neither the transpiler nor the `max_query_size` guard, mixing two parsing rules in one `INSERT`. An empty body still works, which is the normal way to run a polyglot `INSERT` over HTTP. Limitations (scoped, experimental): because a foreign-dialect query is transpiled as a whole (the transpiler rewrites the inline data too and cannot know where the SQL header ends without parsing the dialect), the inline `INSERT ... VALUES` data counts towards `max_query_size` — unlike a native ClickHouse `INSERT`, whose inline data is streamed and is not bounded by `max_query_size`. An oversized payload fails-close with a dedicated, actionable error rather than silently changing the `INSERT` size contract; increase `max_query_size` to submit larger inline payloads. Only `INSERT ... VALUES` inline data is transpilable by the bundled dialects. `INSERT ... FORMAT ...` is not: `FORMAT` is a ClickHouse-only extension, so a foreign-dialect parser rejects the query at the inline data that follows (empirically, `postgresql`/`mysql`/`sqlite`/`duckdb`/`snowflake`/`bigquery` all fail at the first data row after `FORMAT`; a hypothetical identity transpiler even drops the raw `FORMAT` payload rather than re-emitting it). A foreign-dialect `INSERT ... FORMAT` therefore fails cleanly with a syntax error and inserts nothing — like `EXPLAIN INSERT ... VALUES`, which is also not transpilable by the bundled dialects (rejected at the `VALUES` token). The server-owned transpiled buffer that carries the inline data is itself format-agnostic and would handle `FORMAT` data if a transpiler ever produced such a query; the parser also defensively clears the inline-data pointers of an `EXPLAIN`-wrapped `INSERT` — the same way the client unwraps it — so both forms are safe if a future transpiler supports them. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `INSERT ... VALUES` with inline data when using the experimental `polyglot` SQL dialect. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110321",
        "createdAt": "2026-07-13T21:39:09Z",
        "updatedAt": "2026-08-13T14:00:19Z",
        "timestamp": "2026-08-13T14:00:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110344",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix wrong primary-key pruning for toStartOfDay and relative-number functions on out-of-range DateTime64",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/90461 Related: https://github.com/ClickHouse/ClickHouse/pull/108018 --> Related: #90461 Related: #108018 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong `count()` results and dropped rows when a `DateTime64` primary-key column is filtered through `toStartOfDay`, `toRelativeSecondNum`, `toRelativeMinuteNum`, `toRelativeHourNum`, `toRelativeDayNum`, `toRelativeWeekNum`, `toMonthNumSinceEpoch` or `toYearNumSinceEpoch` with values outside the `UInt32`-seconds range (before 1970 or beyond 2106). These functions claim to be always monotonic to the primary index but their standard-precision results wrapped for out-of-range `DateTime64`, breaking primary-key pruning. ### Description The `DateTime64` sibling of #108018 (which fixed the same class for `Date32`). `toStartOfDay` and the relative-number transforms have `FactorTransform = ZeroTransform`, so `IFunctionDateOrDateTime::getMonotonicityForRange` reports them as always monotonic. Their standard-precision `DateTime64` code paths narrowed the result to `UInt32`/`UInt16` without saturating, so for arguments outside that range the value wrapped and the function stopped being monotonic. This makes primary-key range analysis produce exact ranges that extend before the selected mark range, which: - in release builds: silently drops granules holding matching rows, returning a wrong `count()`; - in debug/sanitizer builds: trips `chassert(exact_ranges[i].begin >= range.begin)` in the trivial-count projection optimization (`optimizeUseAggregateProjection.cpp`). Found by the AST fuzzer (amd_msan) on a query that mutated a `Date32` key column to `DateTime64(5)`: `SELECT count() FROM t WHERE toStartOfDay(d) >= toDateTime('2000-01-01 00:00:00','UTC') SETTINGS force_primary_key = 1`. Report: https://s3.amazonaws.com/clickhouse-test-reports/PRs/110310/e242401ecdb1e933c8646546bc2f905d4ed106cd/ast_fuzzer_amd_msan/fatal.log The fix saturates the `DateTime64` `execute` overloads to `[0, result-type max]`, matching the `Date32` fix in #108018, keeping each function monotonic over the whole `DateTime64` domain. Adds `04538_datetime64_zerotransform_monotonicity_pruning`; updates the references of `01768_extended_range`, `04408_datediff_datetime64_overflow` and `02403_enable_extended_results_for_datetime_functions`, which asserted the previous wrapped standard-precision results.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110344",
        "createdAt": "2026-07-14T07:56:02Z",
        "updatedAt": "2026-08-13T16:35:41Z",
        "timestamp": "2026-08-13T16:35:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110391",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Metrics for tracking insertions via materialized view",
        "text": "### Changelog category: - New Feature ### Changelog entry: Added metrics `DirectInsertedRows`, `DirectInsertedBytes`, `MaterializedViewInsertedRows`, and `MaterializedViewInsertedBytes`. ... ### Documentation entry for user-facing changes: I think there are none user-facing changes. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? Sometimes people need to track table's changes via matirialized view. - Example use: A query or command. I used this to check is materialized view metrics work: ```sql CREATE DATABASE IF NOT EXISTS test; DROP TABLE IF EXISTS test.random_events; CREATE TABLE test.random_events ( id UUID, user_id UInt64, event_type LowCardinality(String), amount Float64, is_success UInt8, event_time DateTime, payload String ) ENGINE = MergeTree ORDER BY (event_time, user_id); INSERT INTO test.random_events SELECT generateUUIDv4(), rand64() % 1000000, arrayElement(['click', 'view', 'purchase', 'login', 'logout'], 1 + rand() % 5), round(randCanonical() * 1000, 2), rand() % 2, now() - toIntervalSecond(rand() % 86400), concat('random_payload_', toString(rand64())) FROM numbers(1000); ``` ```sql CREATE TABLE test.random_events_hourly_stats ( event_type LowCardinality(String), event_hour DateTime, total_amount Float64, success_count UInt64, fail_count UInt64, event_count UInt64 ) ENGINE = SummingMergeTree ORDER BY (event_type, event_hour); CREATE MATERIALIZED VIEW test.mv_hourly_stats TO test.random_events_hourly_stats AS SELECT event_type, toStartOfHour(event_time) AS event_hour, sum(amount) AS total_amount, countIf(is_success = 1) AS success_count, countIf(is_success = 0) AS fail_count, count() AS event_count FROM test.random_events GROUP BY event_type, event_hour; -- Backfill hourly stats INSERT INTO test.random_events_hourly_stats SELECT event_type, toStartOfHour(event_time) AS event_hour, sum(amount) AS total_amount, countIf(is_success = 1) AS success_count, countIf(is_success = 0) AS fail_count, count() AS event_count FROM test.random_events GROUP BY event_type, event_hour; -- Testing results SELECT * FROM test.random_events_hourly_stats ORDER BY event_hour DESC, event_type LIMIT 20; ``` Kinda optional stage: ```sql INSERT INTO test.random_events SELECT generateUUIDv4(), rand64() % 1000000, arrayElement(['click', 'view', 'purchase', 'login', 'logout'], 1 + rand() % 5), round(randCanonical() * 1000, 2), rand() % 2, now() - toIntervalSecond(rand() % 86400), concat('random_payload_', toString(rand64())) FROM numbers(1000000); ``` Checking metrics: ```sql SELECT event, value FROM system.events WHERE (event ILIKE '%InsertedRows') OR (event ILIKE '%InsertedBytes') ORDER BY event; ```",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110391",
        "timestamp": "2026-08-12T20:18:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "UberDever",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110429",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix \"Cannot write to finalized buffer\" in MergeTreeDeduplicationLog::rotate",
        "text": "Fixes a server abort with the logical error `Cannot write to finalized buffer` that was hit by the stress test. `MergeTreeDeduplicationLog::rotate` finalized the current log writer and only afterwards created the writer for the new log file. If creating the new writer threw — a transient I/O error, or, in the CI failure, a memory-tracker fault injection hitting the `WriteBufferFromS3` allocation after `finalize` had already succeeded — `current_writer` was left pointing at the already finalized buffer while `stopped` was still `false`. The next write to the deduplication log then wrote to that finalized buffer and aborted. In the CI failure the write came from the background `MergeTreeCleanupThread`: ``` MergeTreeCleanupThread::iterate -> MergeTreeData::clearEmptyParts -> StorageMergeTree::dropPartNoWaitNoThrow -> MergeTreeDeduplicationLog::dropPart -> writeRecord -> WriteBuffer::write -> Logical error: 'Cannot write to finalized buffer' ``` The failure was reproduced on a disk that does not support writing with append (`s3_plain_rewritable`), where the deduplication log is rotated on every operation, so a failing `writeFile` during rotation is easy to hit. The fix makes `rotate` exception-safe: the writer for the new log file is opened first, before any state is changed. If it throws, nothing has changed and `current_writer` still points to the previous, live writer, so the log stays usable and the operation can be retried; only once the new writer is ready is the previous one finalized and swapped out. This preserves the invariant that `current_writer` is never a finalized buffer. Additionally (a review finding): making `rotate` exception-safe was not enough to honor the retry contract. `addPart` had already published the block IDs into the in-memory deduplication map (and written the `ADD` records) before the rotation ran, while `MergeTreeSink` commits the part only after `addPart` returns. An insert aborted by a rotation failure therefore left its block IDs published, and a client retry of the same insert was wrongly deduplicated against a part that never became active — silently dropping the data. Publication in `addPart` is now all-or-nothing: on any failure, the already published block IDs are removed from the in-memory map, and compensating `DROP` records are written to the still-live writer (best effort), so replaying the log on server startup does not re-publish them either. Additionally (a second review finding): the compensating `DROP` writes above assumed `current_writer` was still a usable, live writer. That only holds when the failure came from `rotate` itself, inside `rotateAndDropIfNeeded`. If instead one of the `ADD` records failed to write directly, `WriteBuffer::next` cancels the buffer on any exception, so `current_writer` was left canceled and refused any further writes — the compensating `DROP` records were silently lost in that case, reopening the same wrongly-deduplicated-retry window. The rollback now rotates to a fresh writer first whenever `current_writer` is canceled; `rotate` already tolerates an already-canceled `current_writer`, so this is also safe (if redundant) when it is still the live writer from before a failed `rotate` call. Additionally (a third review finding, caught by the CI's own gtests): the previous fix for the canceled-writer case introduced a new abort. `rotate`'s rollback called `current_writer->finalize()` unconditionally on the previous writer, but `WriteBuffer::finalize` disallows calling it on an already-canceled buffer — it throws a `LOGICAL_ERROR`, which aborts the process immediately in debug and sanitizer builds (before the surrounding `try`/`catch` even runs). `rotate` and `shutdown` now skip `finalize` when `current_writer` is already canceled, since a canceled buffer has nothing left to flush anyway. Additionally (a fourth review finding): the all-or-nothing rollback in `addPart` still narrowed the deduplication window on a failure. `LimitedOrderedHashMap::insert` evicts the oldest entry once the map is at capacity, and the rollback only erased the block IDs the failing call itself had published — it did not restore entries evicted by those `insert` calls. With a small deduplication window, a failed insert could therefore evict an unrelated, already-active part's block ID from memory, and a retry of that unrelated block ID would then be wrongly accepted instead of deduplicated, until the log was replayed from disk. `addPart` now defers all `deduplication_map.insert` calls until the durable writes and the rotation have both succeeded, so the in-memory map is never mutated on a path that might still need to roll back. Additionally (a fifth review finding): `rotate` logged and suppressed failures to `finalize`/`sync` the previous log file's writer. Since `rotateAndDropIfNeeded` is the success boundary of `addPart`, an insert whose `ADD` records had just been written to that file was still treated as durably recorded: `addPart` published the block IDs and `MergeTreeSink` committed the part, even though the only on-disk `ADD` records may never have reached durable storage. After a restart the deduplication log forgot the committed insert, so a client retry of the same block was wrongly accepted and duplicated the data. `rotate` now rethrows such a failure after switching over to the new writer — so the log itself stays usable — and `addPart` treats the insert as failed: its rollback writes compensating `DROP` records into the freshly opened log file and the insert is aborted instead of committing a part the deduplication log may have already forgotten. Additionally (a sixth review finding): even after the rollback stopped re-publishing rolled-back block IDs, it still narrowed the deduplication window across a server restart. The rollback wrote compensating `DROP` records, but on startup `loadSingleLog` replays the log in order, so the failed insert's `ADD` record is applied — evicting the oldest committed block ID from the bounded in-memory map — before the `DROP` erases the rolled-back one. With `deduplication_window = 1`, a committed `block1` followed by a failed insert of `block2` replayed as `ADD block1`, `ADD block2`, `DROP block2` and ended with an empty map, so after a restart a retry of `block1` was wrongly accepted and its data duplicated (larger windows dropped the oldest committed block the same way). Rolled-back inserts are now written with a distinct `CANCEL` record instead of a `DROP`, and replay cancels each `CANCEL` against its matching preceding `ADD` and skips both before applying the remaining records — so a rolled-back insert consumes no deduplication-window slot on replay and the reloaded state matches the live in-memory state exactly. Additionally (a seventh review finding): even with rolled-back inserts written as `CANCEL` records, log retention still over-counted them. `dropOutdatedLogs` decides which older log files are redundant by summing each file's raw record count from the newest backwards until the deduplication window is covered, but a cancelled `(ADD, CANCEL)` pair contributes nothing to the reconstructed map. A failed multi-block insert could therefore make retention treat those transient records as consumed deduplication-window slots and drop an older log that still held live, committed block IDs. With `deduplication_window = 2`, committed `block1` / `block2` in the first log and a failed `addPart({\"block3\", \"block4\", \"block5\", \"block6\"}, ...)` whose rollback wrote four `CANCEL` records into the second log, the first restart rebuilt the map correctly but, seeing the inflated raw counts, rotated and dropped the first log; a second restart then replayed only the `CANCEL`-only log and forgot `block1` / `block2` — wrongly accepting, and duplicating, a retry of those committed inserts. Retention accounting is now based on the records that survive cancel-pair elimination: `applyRecords` recomputes each log's `entries_count` from only the non-cancelled records on replay, and `addPart`'s rollback undoes the retention count of the rolled-back `ADD` records (and does not count the `CANCEL` records), so the live and replayed accounting stay consistent and a rolled-back insert never shrinks the retained history. Additionally (an eighth review finding): the same durability boundary left `dropPart` inconsistent. It wrote a `DROP` record, erased the block ID from the in-memory map, and rotated, for each covered block ID one at a time; now that `rotate` can rethrow a failure to `finalize`/`sync` the previous log file, a multi-block drop could throw partway through — after some of the dropped part's block IDs had been erased but before the rest — leaving the in-memory map in a partial state that its caller `StorageMergeTree::dropPartNoWaitNoThrow` never repairs (it has already taken the part out of the active set and does not retry the drop). The remaining block IDs then stayed published against a part that no longer exists, and any non-durable `DROP` records resurrected the erased ones after a restart. `dropPart` is now transactional like `addPart`: it collects every covered block ID, writes all the `DROP` records, rotates, and only then erases the block IDs from the map, so a failed drop leaves every covered block ID published (the deduplicating, safe direction) instead of a half-applied mixture. Additionally (a ninth review finding): the reworked `dropPart` was still not all-or-nothing on disk. `writeRecord` flushes every record, so when the write of one `DROP` record failed partway through a multi-block drop, the records written before it were already durable while no block ID had been erased from the in-memory map. After a restart, replaying that durable prefix erased only part of the failed drop: a block ID whose `DROP` record reached the disk stopped deduplicating while its siblings still did — a half-applied drop that matches neither the live map (which kept every covered block ID) nor a completed drop, and that the caller `StorageMergeTree::dropPartNoWaitNoThrow` never repairs. Worse, the failed write left `current_writer` canceled, and a canceled buffer silently discards all further writes, so the `ADD` records of later, successfully committed inserts would never reach the disk either and those inserts would be forgotten after a restart, wrongly accepting and duplicating their retries. `dropPart` now mirrors `addPart`'s rollback: on any failure it writes a compensating `CANCEL` record for each `DROP` record that was written (rotating to a fresh writer first when the failed write canceled the current one) and undoes their retention count, and replay cancels a `CANCEL` against the most recent preceding un-cancelled `ADD` or `DROP` of the same block ID — so a failed drop keeps every covered block ID published both live and across a restart, and the log stays usable and durable afterwards. Additionally (a tenth review finding): the persisted `CANCEL` record had no downgrade contract. A server from before this change replays every log record as either an erase (`DROP`) or an insert (anything else), so after a downgrade with rollback records already on disk, the `CANCEL` written for a rolled-back insert replayed as an insert: the never-committed block ID stayed published, and a client retry of the failed insert was wrongly deduplicated — silently dropping its data. The rollback records are now encoded so that every server version, old or new, replays them with the correct net effect, with no need for a format version. The rollback of a failed insert is written as a plain `DROP` record carrying a reserved part name (`cancel`, which can never collide with a real part name — and the part name of a `DROP` record is never parsed by any server version): an older server replays it as the erase that unpublishes the never-committed block ID, so a retry of the failed insert is accepted, while a server with this change recognizes the marker and cancels the `(ADD, DROP)` pair out of the replay entirely, preserving the deduplication-window and retention guarantees above. The `CANCEL` operation remains only as the rollback of a failed drop, where it carries the real, parseable part name: an older server replays it as the insert that restores the still-published block ID — exactly the rollback's net effect, with only the entry's position in the eviction order diverging. A downgraded server is therefore never worse off on these logs than it would have been running on logs it produced itself. Additionally (an eleventh review finding): two remaining bookkeeping steps could still throw at a point where the `Cannot write to finalized buffer` state or a broken retry contract would return, this time on a memory-allocation (`std::bad_alloc`) rather than an I/O failure. First, `rotate` registered the new log file in `existing_logs` (a `std::map`, whose `emplace` allocates a node) only after finalizing the previous writer; if that allocation threw, `current_writer` was left pointing at the finalized old writer — the exact abort this pull request eliminates — and no rollback path could detect it, because the buffer is finalized, not canceled. `rotate` now performs that registration before finalizing the old writer, so the only remaining throwing bookkeeping step runs while the old writer is still live and usable; the switch-over to the new writer (an integer store and a `unique_ptr` move) is non-throwing. Second, `addPart` deferred `deduplication_map.insert` until after the `ADD` records were durable, but that insert is itself a throwing, evicting step: `LimitedOrderedHashMap::insert` allocates and, when the map is full, evicts the oldest entry before inserting. An allocation failure there — after the records were already durable — propagated an exception without writing any compensating rollback records, so a client retry could be deduplicated against a part that never committed (and, in the window-full case, an unrelated committed block ID was evicted from the live map). Publication is now split so that nothing which can throw runs after durability: the block IDs are inserted up front with a new `LimitedOrderedHashMap::insertWithoutEviction` — strongly exception-safe, and crucially never evicting, so a failure before the durable writes is rolled back with a plain non-allocating `erase` that never drops an unrelated block ID — and the deduplication window is enforced afterwards with `LimitedOrderedHashMap::trimToMaxSize`, which only pops the oldest entries and so cannot throw at a point where the insert could no longer be rolled back. `insert` and `setMaxSize` are expressed in terms of these two primitives, so their behavior is unchanged. Additionally (a twelfth review finding): two more failure modes, on the load and accounting paths. First, `loadSingleLog` appended each record and its originating log number to two parallel vectors as two separate steps; if the second append threw (for example `std::bad_alloc` while loading a large deduplication log) after the first had succeeded, the vectors were left with different lengths, and because `load` tolerates and still replays whatever was read, `applyRecords` — which indexes the two in parallel — then read past the end of the shorter one, turning an allocation failure at startup into undefined behavior. The two appends are now atomic: if appending the log number throws, the just-pushed record is removed again (`pop_back` on a non-empty vector never throws), so the two vectors always stay in lockstep. Second, the retention fix above had repurposed each log file's single record count to mean *records surviving cancel-pair elimination*, but that same field is also the log-growth threshold that drives rotation and compaction. A rolled-back operation's records net to zero surviving records, so repeated transient failures could append arbitrarily many raw rollback pairs while the count stayed at zero — the newest logs never reaching the rotation threshold, growing without bound, and forcing `load` to materialize a number of records proportional to the number of failures rather than to the deduplication window. Each log file now carries two counts: the raw `entries_count` (every physical record, which drives rotation, so a rollback-heavy log still rotates and stays bounded) and `effective_entries_count` (only the records that survive cancel-pair elimination, which drives retention, as in the seventh finding above). `addPart` and `dropPart` count every physical record — including their compensating rollback records — towards the raw count and never decrement it, undoing only the effective count of the rolled-back records; `applyRecords` recomputes both from the replayed stream. Additionally (a thirteenth review finding, following up on the twelfth): splitting the raw and effective counts bounded the growth of a single log file but not the *number* of files. `dropOutdatedLogs` cannot reclaim the `(ADD, rollback)` and `(DROP, CANCEL)` record pairs a rolled-back operation leaves behind — the rollback record sits in a newer file while the record it cancels sits in an older file that is still retained for other, live block IDs, and retention only ever drops an oldest prefix — so under repeated transient write or fsync failures those cancelled-out pairs, and the log files holding them, would still accumulate without bound, and every restart would replay a number of records proportional to the number of failures. A new `compact` step rewrites the whole live deduplication state — which the in-memory map already holds exactly — into a single fresh log file and drops every older file, once the raw record count across all files exceeds the effective (surviving-record) coverage by more than a couple of rotation intervals (which only happens once rolled-back operations have piled up, since the two counts are equal in normal operation). Written in the map's insertion order, the snapshot replays to the identical state, so discarding the accumulated history is safe. It runs at the end of a successful `addPart`/`dropPart` and after `load`, so both the retained files and the load-time replay stay bounded by the deduplication window regardless of how many failures preceded them. `compact` is best effort and never throws — it finalizes the snapshot before removing any old file, and on any failure leaves the existing files and writer untouched. Additionally (a fourteenth review finding): the compaction above was skipped entirely on disks that do not support writing with append (for example `s3_plain_rewritable`, the disk on which the original abort reproduced), on the assumption that its every-operation rotation already keeps the log small. It does not. `rotateAndDropIfNeeded` rotates on every operation there, so each rolled-back insert or drop still leaves a newer log file whose effective (surviving-record) count is zero — its rollback record cancels an `ADD`/`DROP` in an older file that is still retained for other, live block IDs — and `dropOutdatedLogs`, which can only drop an oldest prefix, cannot reclaim any of them. The retained files, and the records `load` replays on every restart, therefore still grew with the number of failures, unbounded, in exactly the regime this pull request targets. Compaction now runs on such disks too: because the finalized snapshot file cannot be reopened for appending there, `compact` writes the live snapshot to a fresh durable file and then starts the next operation in another fresh, empty file, instead of reopening the snapshot — reducing the retained history to the snapshot on any disk. Additionally (a fifteenth review finding, two more accounting issues). First, the all-or-nothing rollback discounted the rolled-back `ADD` (or `DROP`) records from the log file's `effective_entries_count` — the count `dropOutdatedLogs` uses for retention — only once, after the whole rollback loop, as `effective_entries_count -= written`. But `writeRecord` flushes per record, so a compensating write can throw partway through the loop after some records are already durable, and the post-loop decrement was then skipped entirely, leaving the count as if none of the rolled-back records had been cancelled. A replay of the partially written stream, however, cancels out exactly the records whose compensating record reached disk, so the inflated live count could make `dropOutdatedLogs` drop an older log that still held committed block IDs, and a restart then forgot those committed inserts and wrongly accepted — and duplicated — their retries. The discount now happens once per successfully written compensating record, right after it is durable, so a mid-loop failure discounts only the records that reached disk, matching what a replay reconstructs. Second, on a disk without append support (such as `s3_plain_rewritable`) the compaction bounded the number of retained files against operations and failures but not against restarts alone: every rotation, including the one in `load`, starts a fresh file, and `dropOutdatedLogs` can never reclaim a zero-record file that sits after the file holding the live state (a committed file or a compaction snapshot), because it only drops an oldest prefix. So each restart with no new operations left one more empty file behind — `snapshot`, `empty1`, `empty2`, … — and the retained files, and the records `load` had to replay, grew as O(number of restarts). `load` now removes the trailing zero-record log files before its first rotation (only without append support — with append support the last file is reopened and reused instead), so a restart is idempotent: the file holding the live state plus exactly one fresh writer file. Additionally (a sixteenth review finding): the compaction's failure cleanup could leave a stale snapshot behind that corrupts the next replay. `compact` writes its snapshot at a log number one past `current_log_number`; if the snapshot is made durable but the compaction then fails before switching over to it (for example the reopen of the snapshot for appending throws), the cleanup tried to remove the orphan snapshot and treated a failure to remove it as harmless. It is not: the durable snapshot then survives at a *higher* log number than the older files the server keeps appending to, so a later successful insert commits newer block IDs into an older file, and on the next restart `load` replays the stale snapshot last — after those newer records — resurrecting evicted block IDs and forgetting committed ones (a retry can then be wrongly deduplicated, or a committed insert forgotten, after the restart). The cleanup now goes through `neutralizeOrphanLog`, which removes the orphan file and, if the removal fails too, overwrites it with an empty file — an empty log replays as a no-op regardless of its log number, so it can no longer corrupt the reconstructed state, preserving a consistent numbering boundary instead of leaving an invisible higher-numbered log. Additionally (a seventeenth review finding, two replay/cleanup issues). First, `applyRecords` cancelled out the record pairs left by a rolled-back operation by matching a rollback record to the most recent preceding un-cancelled record of the same block ID, regardless of whether it was an `ADD` or a `DROP`. That pairing is too loose on the failed `dropPart` + failed `sync` path, where the rollback `CANCEL` can survive on disk while the `DROP` it was meant to undo never reached durable storage: matching by block ID alone then cancelled the committed `ADD` of that block instead, forgetting a still-published block ID after a restart and wrongly accepting a duplicate retry. Each rollback record is now paired only with a preceding record of the exact kind it undoes — a `CANCEL` with a real `DROP`, a cancelled-add `DROP` with an `ADD` — so a rollback record whose target was lost leaves the unrelated committed record untouched. Second, when `compact` could not remove some of the old, superseded log files, it treated the lingering file as harmless because the snapshot replays after it. That holds for the set of block IDs but not for their FIFO order: a lingering pre-snapshot file replays to a stale intermediate order, and the snapshot's `ADD` records on top do not refresh the position of an already-present key, so the next insert after a restart could evict a different committed block than the live process would. `compact` now neutralizes an un-removable old file (retrying the removal and, failing that, emptying it) so the snapshot alone determines the reloaded state and its eviction order. Additionally (an eighteenth review finding, refining the seventeenth): pairing a rollback record only by record kind and block ID is still too loose, because a block ID can be reused across part generations — committed as one part, dropped, and committed again as another. On the failed-drop + failed-fsync path, where the `DROP` a `CANCEL` undoes never reached durable storage while the `CANCEL` did, replaying `ADD partA`, `DROP partA`, `ADD partB`, `CANCEL partB` let the `CANCEL` consume the older generation's committed `DROP`: the surviving stream became `ADD partA`, `ADD partB`, and the reconstructed map kept the stale first generation (a repeated insert of a present block ID is a no-op). Dropping the current generation after the restart then no longer covered the block ID, so a legitimate reinsert after that drop was wrongly deduplicated. A `CANCEL` carries the part name of the `DROP` it rolls back, so it is now paired only with a preceding `DROP` of the same block ID *and* the same part name — the exact record generation it was written to undo — and a stray `CANCEL` whose target was lost pairs with nothing. The cancelled-add marker still pairs with an `ADD` by block ID alone (its part name field holds the reserved marker), which is sufficient: a second `ADD` of the same block ID can only be written once the live map no longer holds it, so any older surviving `ADD` is followed by a surviving `DROP` that erases it on replay regardless of the pairing. Additionally (a nineteenth review finding, two recovery-path issues). First, a canceled log writer was only healed while rolling back the operation that canceled it: if that rollback's rotation failed too (the disk still down), the writer stayed canceled, and the first retry after the disk recovered still failed with `Cannot write to canceled buffer` before reaching any recovery code — only that failed retry's own rollback reopened a writer, so the first retry always failed in the double-fault case. Second, when a file left behind by a failed compaction could neither be removed nor overwritten with an empty file, the failure was logged and forgotten, and the server carried on appending newer committed records to an older, lower-numbered file while a stale higher-numbered snapshot survived on disk — which a later restart would replay last, forgetting the newer committed block IDs. Both `addPart` and `dropPart` now start with a `prepareToWrite` step that runs before anything is written: it heals a canceled writer by rotating to a fresh one up front (so the first retry after recovery succeeds), and it retries neutralizing any such pending file, failing closed — the operation throws, retryably, with nothing written — for as long as one remains on disk, because failing an insert loudly is recoverable while silently deduplicating wrongly after a restart is not. Additionally (a twentieth review finding): that fail-closed barrier was process-local, so it did not survive a restart. `prepareToWrite` refused to write while a stale file was pending neutralization, but a crash — or a clean shutdown — before a retry neutralized it lost that knowledge, and the next `load` replayed the stale files as ordinary history, which can resurrect evicted block IDs or rebuild a diverged eviction order and silently deduplicate wrongly later. The barrier is now persisted as an on-disk marker, written durably before a compaction starts and cleared (the file removed or, failing that, overwritten empty) only once the on-disk history is provably consistent again — the compaction finished with a clean cleanup, its failure was fully rolled back, or a later retry neutralized every pending file. If `load` finds the marker still active, the previous run died inside that window, so it discards the whole on-disk history and starts afresh instead of replaying it: losing at most one deduplication window of best-effort insert deduplication (a client retry may be accepted again — a visible duplicate) is strictly safer than replaying history that is known to be possibly inconsistent, and unlike a precise per-file recovery it does not depend on reconstructing which of the compaction's steps had completed when the process died. If a file or the marker can neither be removed nor emptied even then, `load` throws, failing closed. The marker needs no format version: it is the deduplication log file with number 0 — a number no real log can ever get — holding a single rollback record that every server version replays as a no-op, so a downgraded server simply treats it as an ordinary, harmless log file. Additionally (a twenty-first review finding): the restart barrier of that marker could still be bypassed when the server restarted with deduplication disabled. A load with `non_replicated_deduplication_window = 0` deliberately leaves the marker alone — nothing is replayed or written while deduplication is disabled — but re-enabling deduplication with `ALTER TABLE ... MODIFY SETTING` then just reopened the newest fenced-off log file for appending, without consulting the marker. New committed records landed on top of exactly the stale history the marker fences off, and the next restart — acting on the still-active marker — discarded them together with the stale ones, wrongly accepting (and duplicating) a retry of an insert committed after the re-enable. Re-enabling deduplication now acts on the marker the same way a load with deduplication enabled would have: it discards the suspect history and clears the marker before anything new is written, so the records committed from then on survive the next restart. Additionally (a twenty-second review finding): changing the deduplication window with `ALTER TABLE ... MODIFY SETTING` is the one path that can rotate the log or reopen a writer without an insert or drop in front, and it bypassed the fail-closed recovery barrier those operations run first. After a failed compaction whose cleanup also failed completely, the stale files awaiting neutralization are known only to the running process (compaction had already unregistered them), so re-enabling deduplication in the same process discarded only the registered files: it cleared the on-disk marker while the stale files kept their content, and its subsequent rotation truncated the oldest of them in place; a restart before the next insert or drop then replayed the surviving stale files as ordinary history with the oldest records missing, forgetting committed block IDs and wrongly accepting — duplicating — their retries. Even while deduplication stayed enabled, the setter's rotation could reclaim the one stale file pending neutralization while the on-disk marker stayed active with nothing left to clear it, so the next restart discarded the records committed after the `ALTER` and wrongly accepted their retries. The setter now runs the same recovery barrier first whenever the new window is non-zero: the stale files are neutralized precisely — without discarding the consistent history — and the marker is cleared only once none remains, failing the `ALTER` closed (retryably) while the disk does not allow it; the discard on the re-enable transition therefore fires only when the marker was left by a previous process, whose precise in-process knowledge is gone. Additionally (a twenty-third review finding): the all-or-nothing insert contract stopped at the deduplication log's own boundary. `MergeTreeSink::commitPart` publishes the block IDs via `addPart` and only then makes the part active (`renameTempPartAndAdd` followed by `transaction.commit`); if one of those later steps threw, the part never became active but its block IDs stayed durably published, so a client retry of the same insert was silently deduplicated against a part that does not exist - both in the same process and after a restart. The sink now unpublishes them on that path with a best-effort `dropPart` (itself all-or-nothing per the earlier findings) before rethrowing; if even the drop fails - for example on the same broken disk that failed the commit - the block IDs stay published, which is no worse than before, and the original error is not masked. A new stateless test injects a failure between the publication and the part commit through the new `merge_tree_sink_fail_part_commit_after_dedup` failpoint and verifies that a retry of the failed insert is inserted rather than deduplicated (the test fails without the rollback). Gtests inject a `writeFile` failure during rotation, a `next()` (flush) failure on the currently open writer, and separately an fsync failure on the previous writer's `sync` while rotating away from it, and verify in all cases that the log remains usable afterwards, that a retry of the failed insert is accepted (and only then deduplicates as usual), and that unrelated, already-committed block IDs are not evicted from memory by the failed insert — both within the same process and, including the eviction check, after reloading the log from disk on a restart. A further gtest fails a four-block insert with an injected `sync` failure and restarts twice, verifying that the log file holding the committed block IDs is not dropped by the retention pass, so the committed inserts still deduplicate after the second restart. A final gtest drops a range covering two block IDs while injecting an fsync failure into the rotation that the first `DROP` triggers, and verifies that the failed drop erases neither block ID, so a retry of either is still deduplicated. Another gtest fails the write of the second of two `DROP` records mid-drop and verifies — both live and after reloading the log on a restart — that neither covered block ID is forgotten, and that a record written after the failed drop is durable (the rollback rotated away from the canceled writer, which would otherwise silently discard it). A final gtest replays the on-disk logs left behind by a rolled-back insert and a rolled-back drop using the exact pre-change replay logic (`DROP` = erase, anything else = insert, part names parsed) and verifies that an older server would not consider the rolled-back insert's block ID published (no silently dropped retry after a downgrade) and would keep both block IDs of a rolled-back drop published. A unit test exercises the new `LimitedOrderedHashMap` primitives directly (an `insertWithoutEviction` keeps every entry, including the oldest, present and looked up correctly while the map temporarily exceeds its limit; a following `trimToMaxSize` evicts the oldest in FIFO order), and a further gtest verifies that a successful insert still evicts the oldest block ID to enforce the deduplication window, so splitting publication into a non-evicting insert and a later trim does not change the observable windowing. A final gtest fails an insert on an fsync error and then commits a single insert into the rollback-only log file, verifying that the file rotates once its raw size (its rollback record plus the new record) reaches the threshold — which it would not if rotation were driven by the surviving-record count, letting a rollback-heavy log grow without bound. A final gtest accumulates many rolled-back inserts on an fsync-failing disk and verifies that a restart compacts them into a single log file while preserving the live deduplication state, rather than retaining every accumulated file. A final gtest repeats that accumulation on a simulated disk without append support and verifies that a restart likewise compacts the accumulated files back down to the snapshot (rather than retaining one file per failure), confirming the bound holds in the every-operation-rotation regime too. A further gtest interrupts a failed insert's rollback partway — injecting an fsync failure into the rotation and then a flush failure after only some of the compensating records are written — and verifies, after a restart, that the older log holding an unrelated committed block ID is not dropped by retention, so that block still deduplicates. A final gtest restarts several times on a disk without append support with no operations in between and verifies that the retained log file count stays bounded (the live-state file plus one fresh writer file) rather than growing by one empty file per restart. A final gtest makes a compaction snapshot durable and then fails both the reopen of the snapshot for appending and its removal during cleanup, verifying that the orphan snapshot is left as an empty file (so it cannot outrank the live state on the next replay) and that a healthy restart still reconstructs the live deduplication state exactly. A further gtest constructs the on-disk state of a failed drop whose `DROP` record was lost to an fsync failure while its `CANCEL` survived, and verifies that the covered block ID still deduplicates after a restart rather than being cancelled out by the stray `CANCEL`. A related gtest repeats that construction with a block ID reused across two part generations and verifies that, after a restart, the stray `CANCEL` does not cancel the older generation's committed `DROP`: the current generation stays in the map, dropping it clears the block ID, and a legitimate reinsert after the drop is accepted. A final gtest accumulates rollback garbage and then restarts on a disk that can rewrite files but cannot unlink them, verifying that compaction empties every pre-snapshot log file (so the snapshot alone determines the reloaded eviction order) and that the live state survives. A further gtest injects a double fault — a record write fails and the rollback's rotation fails as well — and verifies that the first retry after the disk recovers already succeeds (the canceled writer is healed up front) and that the retried insert survives a restart. A final gtest makes a compaction snapshot durable, fails its completion, and defeats the entire cleanup (the orphan can neither be removed nor emptied), verifying that inserts then fail closed while the stale orphan is on disk, that the committed state still deduplicates, that the first insert after the disk recovers neutralizes the orphan and succeeds, and that the state stays exact across a healthy restart. Two further gtests restart the process while a failed compaction's stale files are still on disk — once with the orphan snapshot left behind before the switch-over, once (on a disk that can create but not destroy files) with unremovable pre-snapshot files left behind after it — and verify that the new process discards the history instead of silently replaying the stale files, and that deduplication works normally from the fresh history on. Another gtest restarts with deduplication disabled while the marker is still active, re-enables it through the window-size setter, and verifies that the marker is cleared and the stale history discarded before anything is written, and that an insert committed after the re-enable still deduplicates across a further restart instead of being thrown away with the stale files. A further gtest replays the same-process trace — a failed compaction whose cleanup fails completely, the disk healing, and the window set to 0 and back to a non-zero value in the same process — and verifies that re-enabling neutralizes the pending stale files precisely (only the snapshot remains on disk), and that the committed block still deduplicates after a restart before any further write. (The on-disk record definitions were moved into their own header, `MergeTreeDeduplicationLogRecord.h`, so that the regression tests can also be compiled against the merge-base sources by the `Bugfix validation (unit tests)` CI job; the tests that exercise machinery introduced by this fix are compiled only when that header is present.) CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=a63a425c48d3123d28bef2ad60511c9fc0583468&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a logical error `Cannot write to finalized buffer` (an abort in debug and sanitizer builds) in the non-replicated `MergeTree` deduplication log when log rotation fails, for example after a transient I/O error on the deduplication log's disk. Failed deduplication-log writes, flushes, syncs, and compactions now roll back safely, so retries of failed inserts or part drops are not wrongly deduplicated against parts that never committed, committed block IDs are not forgotten after a restart, and rollback-heavy histories are compacted instead of growing without bound.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110429",
        "createdAt": "2026-07-14T18:08:47Z",
        "updatedAt": "2026-08-13T06:36:43Z",
        "timestamp": "2026-08-13T06:36:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110464",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `aiRedact` function for PII detection and redaction",
        "text": "`aiRedact` detects and redacts PII in text using an LLM provider. For example: SELECT aiRedact('Contact John Doe at john@doe.org', ['name', 'email']) -- 'Contact [REDACTED] at [REDACTED]' Pass an empty `categories` array to fall back to a default set of common categories (name, email, phone number, address, credit card, IP address). The redaction token defaults to `[REDACTED]` and can be changed with `replacement`. `aiRedact` is best-effort: detection and redaction are performed by an LLM, so the output is not reliable and may still contain PII depending on the model, prompt, and input. It must not be treated as a sufficient anonymization mechanism on its own, always review the output before relying on it. On error it behaves like the other AI functions: it throws by default, or returns an empty string when `ai_function_throw_on_error = 0`. Closes: https://github.com/ClickHouse/ClickHouse/issues/110362 Changelog category (leave one): - New Feature Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): New function `aiRedact` that detects and redacts personally identifiable information (PII) in text using an LLM provider. Specify the categories to redact (e.g. `['email', 'name']`) or pass an empty array to use a default set of PII categories. Matched values are replaced with a token (`[REDACTED]` by default). Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110464",
        "createdAt": "2026-07-14T23:07:06Z",
        "updatedAt": "2026-08-13T17:17:41Z",
        "timestamp": "2026-08-13T17:17:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 17
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested"
        ],
        "author": "davidmenggx",
        "state": "open",
        "assignees": [
          "george-larionov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110477",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "EXPLAIN SYNTAX: return the pretty-printed query as a single multi-line record",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/80410 Related: https://github.com/ClickHouse/ClickHouse/pull/107925 --> Closes: #80410 Related: #107925 (closed, superseded by this PR) ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `EXPLAIN SYNTAX` now returns the pretty-printed (multi-line) query as a single result record instead of one record per line. The `oneline` option (`EXPLAIN SYNTAX oneline = 1`) still collapses the output to a single physical line. Queries that consumed the previous per-line output should treat the result as one row whose value contains embedded newlines. ### Description `EXPLAIN SYNTAX <query>` formatted the reformatted query and split it into one result row per physical line. This PR emits the whole formatted, copy-pasteable query as a single record with newlines preserved (issue #80410). Only the `AnalyzedSyntax` code path is changed; `PLAN`, `PIPELINE`, `AST` and the `oneline` option are unchanged. The previous attempt (#107925) was closed because it flipped the `oneline` default to `1`, collapsing the query to a single physical line rather than keeping the multi-line pretty form. This PR keeps `oneline = false` by default and only changes the result shape from N rows to one multi-line record. Reference files for tests consuming EXPLAIN SYNTAX output (including `.oldanalyzer.reference` variants) were regenerated. A regression test (`04545_explain_syntax_single_record`) asserts the single-record multi-line default and the `oneline = 1` collapse.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110477",
        "createdAt": "2026-07-15T00:34:56Z",
        "updatedAt": "2026-08-13T17:11:09Z",
        "timestamp": "2026-08-13T17:11:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-backward-incompatible"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110479",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "EXPLAIN SYNTAX: return a single record (default on)",
        "text": "<!-- Linked issues and pull requests. --> Closes: https://github.com/ClickHouse/ClickHouse/issues/80410 Related: https://github.com/ClickHouse/ClickHouse/pull/107925 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `EXPLAIN SYNTAX` now returns the reformatted query as a single `String` record (with embedded newlines) instead of one record per line, so its output is a recoverable single row that is directly usable (for example, `SELECT count() FROM (EXPLAIN SYNTAX ...)` returns `1`). This is controlled by the new `single_record` option, which defaults to `1`; set `single_record = 0` to restore the historical one-record-per-line output. Other `EXPLAIN` kinds (`PLAN`/`PIPELINE`/`AST`) keep their per-line tree output. ### Description Per the discussion on #80410, `EXPLAIN SYNTAX` is a reformatted, copy-pasteable query, so returning the whole query as a single record is far more usable than splitting it across N rows. This is done in two commits: 1. Add a `single_record` option to `EXPLAIN SYNTAX`, off by default (no behavior change). The single-record code path already existed (used by `EXPLAIN PLAN` JSON output); this wires it to the new option for the `SYNTAX` kind and renames the internal `single_line` flag to `single_record`. 2. Flip the default to `single_record = 1`. This is backward-incompatible: existing `EXPLAIN SYNTAX` consumers now see one row instead of N. The `oneline` option still controls physical-line rendering (default `0`, i.e. embedded newlines). The affected stateless `.reference` files (including `.oldanalyzer.reference` variants) were regenerated for the new default. The `optimize_syntax_fuse_functions` docstring example was switched from `FORMAT TSV` to `FORMAT TSVRaw` so its multi-line output stays literal now that the single record carries the embedded newlines. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110479",
        "createdAt": "2026-07-15T00:40:02Z",
        "updatedAt": "2026-08-13T13:09:34Z",
        "timestamp": "2026-08-13T13:09:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-backward-incompatible"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "Fgrtue"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110493",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix MaterializedPostgreSQL database/table with ON CLUSTER",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/58726 A `MaterializedPostgreSQL` database or table created with `ON CLUSTER` gets the *same* ClickHouse UUID on every replica, because the UUID is generated once on the initiator and shipped in the DDL. With `materialized_postgresql_use_unique_replication_consumer_identifier = 1` the replication slot name was derived from that UUID and the publication name from the PostgreSQL database/table name, so every replica computed the same names and fought over one PostgreSQL slot and publication: a single replica replicated, the others failed in the background with `pqxx::unique_violation` or could not hold the slot, leaving divergent data. The fix hashes the persistent per-server `ServerUUID` together with the object UUID (via `makeUUIDv4FromHash`, which keeps the result within PostgreSQL's identifier length limit) and derives both the slot and the publication name from it, so each replica gets its own pair and replicates independently. Both UUIDs are persistent, so `ATTACH` keeps reusing the same slot. The default behaviour (setting disabled) is unchanged, and since the change is in the shared `PostgreSQLReplicationHandler` it covers both the database and the table engine. A user-managed `materialized_postgresql_replication_slot` has one fixed name that every replica would share, so combining it with this setting is contradictory and is now rejected for a freshly supplied engine definition — a `CREATE`, or an `ATTACH` that carries a full definition, so it cannot be bypassed through `ATTACH ... ON CLUSTER`. A replay of an already stored definition (startup, `RESTORE`, short-syntax `ATTACH`) still accepts it, so existing deployments keep starting. ## Upgrade of existing deployments Salting the names is a rename of the generated identity, so `adoptLegacyReplicationIdentityIfNeeded` — the mechanism that already handles the schema-aware rename — now also adopts the pre-salt slot and publication when the salted ones do not exist. Replication resumes from the same slot position: no re-snapshot into the already-populated nested tables, nothing orphaned. The pre-salt slot name embeds the object's own ClickHouse UUID, so its existence proves ownership. The publication name is schema-blind and does not, so it is adopted only when it publishes *exactly* the tables this engine replicates: the configured `materialized_postgresql_tables_list`, or otherwise the tables the engine actually replicated in the previous run (their nested tables exist on disk) — not the live PostgreSQL schema, which may have grown since `CREATE DATABASE` with tables that are never replicated without an explicit `ATTACH TABLE`. Adoption is skipped for a database that never materialized a single nested table: there is nothing to preserve and its empty on-disk set cannot prove any publication's ownership, so it synchronizes under the current identity and the legacy objects are named in the log for the operator. ## Fail-closed attach paths `pgoutput` resolves publication membership from the historic catalog snapshot at each change's LSN, so streaming through a publication that was created or narrowed after those changes were written silently drops them — unrecoverably, once `confirmed_flush_lsn` advances. Attach therefore no longer papers over a broken PostgreSQL-side state; it fails with a retryable error, so startup keeps retrying and an operator can repair the conflict or rebuild the replica. This holds for every identity, not only for the pre-salt rename: - **Publication gone behind a surviving slot** — previously recreated silently, skipping every change committed while it was missing. - **Slot gone behind a surviving publication**, as after `pg_upgrade` (it keeps publications but not replication slots) or an operator dropping the slot — previously an in-place re-snapshot, whose rows are materialized with `_sign = 1` and `_version = 1` and therefore neither delete rows that disappeared from PostgreSQL nor override rows already at a higher version, silently leaving the replica stale. A clean rebuild recovers every row. A user-managed slot is excluded, since a missing one is already reported as a configuration error. - **Both gone** while the replica already holds data from a previous run. - **Publication drift under the current identity** — `ALTER PUBLICATION ... DROP TABLE`, an altered `publish` set (`CREATE PUBLICATION` defaults to `insert, update, delete, truncate`), or, on PostgreSQL 15 and newer, a row filter or a column list on a published table. The published set is compared by exact `(schema, table)` pair, because in the single-schema modes both the publication table list and the WAL consumer key tables by bare name, so a publication rewritten from `foo.a` to `bar.a` would otherwise resume and replay `bar.a`'s changes into the ClickHouse table for `foo.a`. Extra published tables stay tolerated — the identity already owns the publication name — unless a foreign-schema table's bare name collides with a replicated one. A database that has not materialized a single nested table is exempt throughout: there is nothing a snapshot could make stale, so it bootstraps through whatever survived. If an interrupted first synchronization left both the publication and the slot behind, the leftover slot is dropped so the initial synchronization can run instead of resuming and then failing forever with `UNKNOWN_TABLE` on a nested table that was never created (a user-managed slot, which this engine may not recreate, is refused with an explicit message). Relatedly, the set of tables to materialize on attach now comes from the existing publication — the engine's persisted table set, which the code always claimed to assume correct — instead of the live schema. Previously a PostgreSQL table created after `CREATE DATABASE` made every attach retry fail with `UNKNOWN_TABLE` on its missing nested table, so a whole-schema database could never resume replication after a restart; this reproduces on master with the setting disabled. Published-but-not-materialized tables are reported in a warning that names the recovery (`ATTACH TABLE`), because that state is indistinguishable from a table that was in the original definition but never completed its first snapshot; persisting the configured table set is deferred to the separate persistent-replication-state change, together with the `rebuild required` marker discussed in the review. ## Tests The new `test_postgresql_replica_database_engine_on_cluster` suite covers the fix — two-node `ON CLUSTER` databases and tables, each replica getting its own slot and publication and synchronizing independently, and both removed by `DROP ... ON CLUSTER` — the adoption of a reconstructed pre-salt identity (including a whole-schema database whose PostgreSQL schema has grown, and foreign publications that must not be adopted), every fail-closed case above, the bootstrap paths for never-synchronized runs and for a fresh full-definition `ATTACH TABLE`, and the rejection of a user-managed slot on `CREATE` and on a full-definition `ATTACH ... ON CLUSTER` of either engine. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `MaterializedPostgreSQL` databases and tables created with `ON CLUSTER`. When `materialized_postgresql_use_unique_replication_consumer_identifier` is enabled, the PostgreSQL replication slot and publication names are now unique per server, so replicas no longer collide on a shared replication slot and publication, which previously left all but one replica unable to replicate. Also fixed a whole-schema `MaterializedPostgreSQL` database failing to resume replication after a server restart when a new table had been created in PostgreSQL, and made the attach fail closed instead of silently losing changes when the publication or the replication slot is missing. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110493",
        "createdAt": "2026-07-15T02:39:06Z",
        "updatedAt": "2026-08-13T00:15:34Z",
        "timestamp": "2026-08-13T00:15:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110552",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add support for S3Queue mode='exclusive'",
        "text": "This patch adds a third mode to the S3Queue engine: `exclusive` mode turns off all file tracking and synchronization in ZooKeeper. S3 file processing will be tracked only locally via in-memory structures of this ClickHouse server process. This mode is meant to support high-throughput, high-volume ingestion scenarios. In our particular case, we have a minio instance on each of our ClickHouse nodes, with S3Queue pointing to 127.0.0.1. Exclusive mode allows ingesting many terabytes of data per day while avoid file framing boundary and buffering issues. (S3 multipart uploads solve the buffering issue for ingestion clients.) We have been running our ClickHouse cluster with these patches for many years with no issues. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds support for mode='exclusive' in S3Queue engine, for high-throughput and self-hosted scenarios.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110552",
        "createdAt": "2026-07-15T13:21:15Z",
        "updatedAt": "2026-08-13T07:52:49Z",
        "timestamp": "2026-08-13T07:52:49Z",
        "metrics": {
          "reactions": 1,
          "comments": 17
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "ivan-tkatchev",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110594",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add aiFilter for natural-language boolean filtering via LLMs.",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/110352 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `aiFilter` function that evaluates a natural-language condition against text with an LLM and returns `UInt8` for use in `WHERE`, `PREWHERE`, and `JOIN ... ON`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110594",
        "createdAt": "2026-07-15T16:49:31Z",
        "updatedAt": "2026-08-13T00:19:39Z",
        "timestamp": "2026-08-13T00:19:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-feature",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "ylw510",
        "state": "closed",
        "assignees": [
          "george-larionov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110613",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add experimental PCO compression codec (linking the pcodec Rust crate)",
        "text": "Adds an experimental `PCO` compression codec that links the [pcodec](https://github.com/pcodec/pcodec) (`pco`) Rust crate — the reference implementation — rather than reimplementing it. `pco` is a lossless codec specialized for sequences of fixed-width numbers; on numeric columns with smooth, multimodal, or high-entropy distributions (measurements, timings, counters, identifiers) it often beats `Gorilla`/`FPC`/`ZSTD`/`ALP` on ratio. This is the Rust-library counterpart to the native C++ port explored in https://github.com/ClickHouse/ClickHouse/pull/106222. Instead of maintaining a C++ reimplementation, it links the upstream crate. ### Runtime CPU dispatch `pco` has no explicit SIMD: it relies on the compiler autovectorizing its per-batch loops, which only happens when the crate is compiled with the relevant instruction sets enabled (its `build.rs` warns to build with `-C target-feature=+avx2,+bmi1,+bmi2`). That does not work for ClickHouse, which compiles all Rust once at a fixed baseline micro-architecture level for a binary that must run on a range of CPUs. So `pco` is taken from a ClickHouse fork patched to add runtime CPU-feature dispatch of its hot loops — the same mechanism ClickHouse's C++ uses via `TargetSpecific.h`. The vectorizable loop bodies (decode `read_offsets`; encode `write_short_uints`/`write_uints`/`set_offsets`) are compiled at the crate baseline and again for `x86-64-v3` (AVX2 + BMI1/2 + FMA) and `x86-64-v4` (AVX-512), and the widest variant the running CPU supports is selected once and cached. On aarch64 the baseline already includes NEON, so the baseline body is used directly. Verified at baseline `-C target-cpu=x86-64`: the `read_offsets` v3 trampoline emits AVX2 (`ymm`) and v4 emits AVX-512 (`zmm`). The patch is merged into `ClickHouse/pcodec` (a fork of pcodec/pcodec in the ClickHouse organization) via its own PR: https://github.com/ClickHouse/pcodec/pull/1. The crate is added as the `contrib/pcodec` submodule and linked through a thin FFI wrapper crate (`rust/workspace/pco`). Its one not-yet-vendored dependency, `rand_xoshiro`, was added to `contrib/rust_vendor` in https://github.com/ClickHouse/rust_vendor/pull/72. ### Codec Supported types are all fixed-width numerics of 1/2/4/8 bytes via their underlying integer/float representation: `Int8`..`Int64`, `UInt8`..`UInt64`, `Float32`/`Float64`, and the types backed by them (`Date`, `DateTime`, `Decimal32`/`Decimal64`, `IPv4`, `Enum`, ...). The on-disk block stores a 2-byte header (element width + partial-tail byte count) followed by a raw partial-value tail (as in `Gorilla`/`FPC`) and the payload. The payload is a standalone `.pco` stream, wire-compatible with the reference pcodec implementation; when compression would not shrink a block the raw bytes are stored instead (a per-block \"stored\" flag), so the output never expands by more than the 2-byte header and `getMaxCompressedDataSize` is tight. The codec is gated behind `allow_experimental_codecs`. Because it needs the column type, it can only be specified per column: it is rejected in `TTL ... RECOMPRESS` and in the untyped compression settings that resolve codecs without a type. `CompressionCodecMultiple` propagates the experimental / column-type-requiring properties so a chain such as `CODEC(Delta, PCO)` is still gated, and a codec-only `ALTER TABLE ... MODIFY COLUMN x CODEC(PCO)` validates against the existing column type. Tests: `04512_pco_codec` round-trips every supported numeric and backed type (per-element verification, edge cases, codec chaining, compression-ratio check) and `04513_pco_codec_gating` covers the experimental gate, the `TTL RECOMPRESS` rejection, and the codec-only `ALTER`. The FFI wrapper has its own Rust unit tests (round-trip of all types, the no-expansion fallback, and fail-closed handling of malformed/mismatched streams). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new experimental compression codec `PCO`, which links the [pcodec](https://github.com/pcodec/pcodec) library (patched for runtime CPU dispatch), specialized for fixed-width numeric columns. It is wire-format compatible with pcodec `.pco` streams and is enabled with `allow_experimental_codecs`. ### Documentation entry for user-facing changes: - [x] Documentation is written (the `PCO` codec is documented in `docs/en/sql-reference/statements/create/table.md`, and the compression-frame method byte in `docs/en/interfaces/specs/NativeFormat.md`).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110613",
        "createdAt": "2026-07-15T19:32:58Z",
        "updatedAt": "2026-08-13T15:57:23Z",
        "timestamp": "2026-08-13T15:57:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 15
        },
        "labels": [
          "submodule changed",
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110615",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support TLS/SSL for PostgreSQL connections",
        "text": "Adds TLS/SSL support to all PostgreSQL integrations. Until now, ClickHouse could not establish an encrypted connection to a PostgreSQL server that enforces SSL, nor verify the server certificate: `postgres::formatConnectionString` only emitted `dbname`/`host`/`port`/`user`/`password`/`connect_timeout`, so `libpq`'s `sslmode`/`sslrootcert`/`sslcert`/`sslkey` could never be set. The bundled `libpq` is already compiled with SSL support (`USE_SSL` is defined through `pg_config_manual.h`), so this change is purely about exposing the options — no build change is required. The design follows the model agreed in the review discussion (https://github.com/ClickHouse/ClickHouse/pull/110615#issuecomment-5116481882) and already implemented for MySQL in https://github.com/ClickHouse/ClickHouse/pull/112070: - `sslmode` (`disable`, `allow`, `prefer`, `require`, `verify-ca` or `verify-full`) can be specified anywhere: as a named collection key or as a trailing `key = value` argument after the positional arguments. When unset, the `libpq` default of `prefer` applies. - `sslrootcert` (CA certificate), `sslcert` (client certificate) and `sslkey` (client private key) are **paths to server-local files**. They are only accepted from a named collection defined in the server configuration file (or from a dictionary defined there) and cannot be overridden in a query: the server opens the files with its own privileges, so a path taken from SQL would let anyone who can define a PostgreSQL source probe the local filesystem and authenticate with a client certificate they are not allowed to read themselves. - `sslrootcert_pem`, `sslcert_pem` and `sslkey_pem` accept the **literal contents** of the corresponding file (copy-pasteable, also usable in ClickHouse Cloud where there is no server filesystem to reference). They are accepted from anywhere — a query, a named collection created with SQL, an override of a configuration-defined collection — and are masked in logs and `SHOW` queries like a password. `libpq` can only load credentials from files, so the contents are materialized into a private temporary file (mode 0600, the permission `libpq` requires of a key file) whose lifetime is tied to the connection pool that uses it. These parameters work for the `PostgreSQL` table engine, the `postgresql` table function, the `PostgreSQL` and `MaterializedPostgreSQL` database engines, the `MaterializedPostgreSQL` table engine, and `PostgreSQL` dictionaries. The dictionary source previously accepted an `sslmode` key but silently ignored it; it is now honored. The integration test `test_postgresql_ssl` enables TLS on a PostgreSQL server at runtime, requires SSL via `pg_hba.conf`, and checks positive and falsifying cases on every surface (wrong CA rejected, client certificate required, contents overriding a configured path, restart survival, masking). A stateless test pins the `[HIDDEN]` masking and the rejection of paths from SQL on every surface. Closes: https://github.com/ClickHouse/ClickHouse/issues/80787 Related: https://github.com/ClickHouse/ClickHouse/pull/112070 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support TLS/SSL connections to PostgreSQL for the `PostgreSQL` table engine, the `postgresql` table function, the `PostgreSQL` and `MaterializedPostgreSQL` database engines, and `PostgreSQL` dictionaries: `sslmode` plus the certificates and the key, given either as literal contents (`sslrootcert_pem`, `sslcert_pem`, `sslkey_pem`; masked like passwords) or as paths (`sslrootcert`, `sslcert`, `sslkey`; accepted only from a named collection defined in the server configuration file).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110615",
        "createdAt": "2026-07-15T19:38:24Z",
        "updatedAt": "2026-08-13T01:17:38Z",
        "timestamp": "2026-08-13T01:17:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 54
        },
        "labels": [
          "pr-feature",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "kssenii"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110626",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Explain data/expected structure mismatches in INSERT parse errors",
        "text": "Improves the error message shown when parsing the data of an `INSERT` fails. Previously, inserting data whose structure does not match the destination produced a confusing low-level parse error (for example `Cannot parse input: expected '\\t' before ...`) with no hint about the real cause. In the linked issue, the destination's inferred schema had integer columns while the inserted `TSV` data had string columns, and the reported error pointed at a tab that was actually present. Now, on a parse failure, ClickHouse infers the structure of the data being inserted (only when the format has a data schema reader, and solely for diagnostics) and, if it does not correspond to the expected structure, appends an explanation listing both the inferred and the expected structure. For example: ``` Code: 27. DB::Exception: Cannot parse input: expected '\\t' before: 'page_view... ... The structure of the data being inserted does not match the structure expected by the query, which is likely the cause of the parsing error. Inferred structure of the input data (in format `TSV`): c1 Nullable(Int64) c2 Nullable(String) c3 Nullable(String) Expected structure: c1 Int64 c2 Int64 c3 Int64 ``` The check is wired through a lazy provider on `IInputFormat` that runs only on a genuine parse error, so there is no cost on the happy path. The comparison ignores the artificial `Nullable` wrapper that schema inference adds by default, so inserting valid data into non-nullable columns is not falsely flagged. It covers the synchronous local/server path, client-side parsing, and the asynchronous insert queue (the default path now that `async_insert` is enabled by default). Closes: https://github.com/ClickHouse/ClickHouse/issues/110622 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When parsing of the data being inserted by an `INSERT` fails, the error message now explains a likely structure mismatch by comparing the structure inferred from the data with the structure expected by the query.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110626",
        "createdAt": "2026-07-15T23:41:14Z",
        "updatedAt": "2026-08-13T14:52:22Z",
        "timestamp": "2026-08-13T14:52:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 25
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110653",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add STREAM BOUNDED modifier",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM BOUNDED modifier, which read only the first snapshot of a streaming query, then finish instead of subscribing for updates. cc @alesapin @Michicosun <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1314` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110653",
        "createdAt": "2026-07-16T08:35:24Z",
        "updatedAt": "2026-08-13T10:01:10Z",
        "timestamp": "2026-08-13T10:01:10Z",
        "metrics": {
          "reactions": 1,
          "comments": 6
        },
        "labels": [
          "pr-improvement",
          "pr-synced-to-cloud"
        ],
        "author": "SmitaRKulkarni",
        "state": "closed",
        "assignees": [
          "Michicosun"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110695",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Combine I/O cost with selectivity in PREWHERE condition ordering",
        "text": "When `use_statistics=1` (the default since `auto_statistics_types` was introduced), the PREWHERE optimizer sorted conditions by `estimated_row_count` alone, with `columns_size` only as a tiebreaker. This caused expensive conditions (e.g. Map column, ~500KB) to be placed before cheap ones (e.g. scalar column, ~1KB) whenever the expensive condition appeared more selective — ignoring the I/O cost difference. Apply the classic conjunctive filter ordering rule: sort by `cost / (1 - selectivity)`, i.e. the I/O cost per rejected row. This is computed as `columns_size / max(1, total_rows - estimated_row_count)` and replaces the separate `estimated_row_count, columns_size` pair in the condition comparison tuple. When statistics are unavailable (`estimated_row_count=0`, `total_rows=0`), the formula degrades to `columns_size`, preserving the existing behavior. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Combine I/O cost with selectivity in PREWHERE condition ordering. It fixes performance regression in PREWHERE execution in some cases introduced after https://github.com/ClickHouse/ClickHouse/pull/101275. Part of https://github.com/ClickHouse/ClickHouse/issues/110462 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1177` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110695",
        "createdAt": "2026-07-16T13:16:19Z",
        "updatedAt": "2026-08-13T09:11:02Z",
        "timestamp": "2026-08-13T09:11:02Z",
        "metrics": {
          "reactions": 1,
          "comments": 6
        },
        "labels": [
          "pr-performance",
          "pr-synced-to-cloud"
        ],
        "author": "Avogar",
        "state": "closed",
        "assignees": [
          "hanfei1991"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110781",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add experimental support for reading Iceberg v3 deletion vectors",
        "text": "Resolves: https://github.com/ClickHouse/ClickHouse/issues/107502 **Changelog category (leave one):** - New Feature **Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):** This PR adds experimental read support for Iceberg v3 deletion vectors stored in Puffin files. Supported: - Read Iceberg v3 deletion vector metadata from manifest entries: - `referenced_data_file` - `content_offset` - `content_size_in_bytes` - Read Puffin deletion vector blobs from object storage. - Validate DV blob length, magic bytes, and CRC. - Decode Roaring64 deletion vectors. - Apply deletion vectors during reads. - Support mixed reads with: - v2 parquet position delete files - v3 Puffin deletion vectors - Support table functions and table engines through the common Iceberg read path, including local/S3/Azure object storage. - Add an experimental setting: - `allow_experimental_iceberg_deletion_vectors` Not supported: - Writing Iceberg v3 deletion vectors. - Updating existing deletion vectors. - Compaction/rewrite of deletion vectors. - Container-level lazy decoding of Roaring bitmaps. - Hard memory limit on decoded Roaring bitmap memory. - Full Iceberg v3 feature support beyond deletion vector reads. Notes: - The feature is disabled by default and must be enabled with `allow_experimental_iceberg_deletion_vectors = 1`. - When disabled, existing behavior is preserved. - The implementation decodes the DV blob into a Roaring bitmap and applies it with a streaming iterator during data reads. **Documentation entry for user-facing changes** - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110781",
        "createdAt": "2026-07-17T05:46:20Z",
        "updatedAt": "2026-08-13T05:41:34Z",
        "timestamp": "2026-08-13T05:41:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "linjiayu1025-collab",
        "state": "open",
        "assignees": [
          "asya-ch"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110784",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix memory tracker leak when a parent tracker throws MEMORY_LIMIT_EXCEEDED",
        "text": "`MemoryTracker::allocImpl` increments `amount` optimistically at each level and then recurses into the parent tracker. When some ancestor (typically the total server tracker) threw `MEMORY_LIMIT_EXCEEDED`, only the throwing tracker reverted its own counter. All descendant trackers that had already incremented (thread, query, user) kept the rejected amount forever, together with the already-applied `CurrentMetrics` updates. Under sustained memory pressure this drifts `MemoryTracking` and the query/user accounting upward and can produce spurious `MEMORY_LIMIT_EXCEEDED` errors for queries that use little memory. Now every level undoes its own increment when the allocation fails anywhere in the chain, and the side effects (peak update, memory profiler trace, metric update) are committed only after the whole parent chain has accepted the allocation. As a consequence, rejected allocations no longer emit `TraceType::Memory` profiler events, no longer advance the profiler step, and no longer raise `peak_memory_usage`. Concurrent speculative charges may still affect these values. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix an accounting leak in memory tracking: when an allocation was rejected by a parent memory tracker (for example, the server-wide memory limit), the thread, query, and user level trackers kept the rejected amount. The error accumulated over time, inflating `MemoryTracking` and per-query/per-user memory usage, and could lead to spurious `MEMORY_LIMIT_EXCEEDED` errors.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110784",
        "createdAt": "2026-07-17T06:57:18Z",
        "updatedAt": "2026-08-13T06:11:12Z",
        "timestamp": "2026-08-13T06:11:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "seva-potapov",
        "state": "open",
        "assignees": [
          "azat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110833",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Compare stored table definition expressions by AST instead of formatted text",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/92340 Related: https://github.com/ClickHouse/ClickHouse/pull/110840 Related: https://github.com/ClickHouse/ClickHouse/pull/108590 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed cross-version metadata compatibility for tables whose keys and other definition expressions were written with redundant parentheses (`PARTITION BY (a)`, `ORDER BY (b, c)`, `DEFAULT (a + 1)`, `TTL (d + INTERVAL 10 YEAR)`, `INDEX ix (b * c)`, `CONSTRAINT c CHECK (a > 0)`, `PROJECTION p (SELECT (b) ...)`) or with an explicit single-element `tuple(a)`. #92340 started preserving these parentheses in stored metadata, so `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` between two such tables failed with `Tables have different ...` even on a single server version, and a replica reading metadata written by another version failed with `METADATA_MISMATCH`. Stored table definition expressions are now compared by their ASTs (`getTreeHash`), not by their formatted text. ### Description #92340 started to preserve redundant parentheses in the stored definition ASTs. Older versions stored the canonical form without them, so the metadata serialized by the two versions differs, and every comparison (`ReplicatedMergeTree` replica join, `ATTACH`/`REPLACE`/`MOVE PARTITION FROM`) rejected definitions that are actually equal. Two sub-cases with different affected ranges: - Redundant parentheses (`PARTITION BY (a)` vs `PARTITION BY a`). This is the #92340 regression: measured on released binaries, 25.8, 26.3 and 26.4 accept `ATTACH PARTITION FROM` between the two forms while 26.5, 26.6 and 26.7 reject it, matching the branches that contain `b38a892dfbfab6`. It does **not** require an upgrade or a mixed-version cluster: both tables created by the same 26.7 binary already fail, because the comparison is between two in-memory ASTs of which only one carries the `parenthesized` flag. - An explicit single-element `tuple(a)` vs `a`. This one fails on every version tried, 25.8 included, so it is a long-standing bug rather than a #92340 regression. It is fixed here as well because `extractKeyExpressionList` unwraps the optional `tuple(...)`. Following https://github.com/ClickHouse/ClickHouse/pull/110833#issuecomment-5040213065, this replaces the earlier text-canonicalization approach entirely: - Expressions are compared by their ASTs using `getTreeHash`, not as text or reformatted text. Comparing formatted text is wrong because the text depends on the formatting logic, which can change for unrelated (e.g. aesthetic) reasons, while the stored form may have been written by any past server version. - The comparison method is reusable: `sameAST` in `src/Parsers/IAST.h` (aliases significant, `nullptr`-tolerant overload for optional expressions). - It is applied at every place where serialized, stored expressions are compared to check whether the table was altered or changed unexpectedly: - `ReplicatedMergeTreeTableMetadata::checkImmutableFieldsEquals` / `checkEquals` / `checkAndFindDiff` (replica join, `ALTER` application). The stored strings are parsed purely syntactically (without resolving against any column set), so the fields of an `ALTER` log entry that adds a column and changes a key in one go compare safely; the `columns`/`context` parameters are gone. Keys additionally go through `extractKeyExpressionList`, so `a`, `(a)` and `tuple(a)` are the same single-column key while `a DESC` stays different. - `MergeTreeData::checkStructureAndGetMergeTreeData` (the `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` structure gate): sorting/partition/primary keys and the secondary-index/projection definition sets. - `StorageReplicatedMergeTree::checkTableStructureAttempt`: the columns comparison (`ColumnsDescription`/`ColumnDescription`/`ColumnDefault` equality now compares the default/codec/TTL expressions as ASTs). The column-by-column comparison also has to ignore in-memory-only state that is never serialized (implicit statistics and the auxiliary `data_type` of `ColumnStatisticsDescription`, hence the new `hasSameExplicitStatistics`), otherwise every replicated `CREATE` would report `INCOMPATIBLE_COLUMNS` against the columns it had just written to ZooKeeper. - `StorageReplicatedMergeTree::alter`: detection of which metadata fields an `ALTER` actually changed. The changed fields are also written back to Keeper through the same backward-compatible serializers the `ReplicatedMergeTreeTableMetadata` constructor uses (`formatDefinition` / `formatDefinitionList`), so an `ALTER` never publishes a parenthesized definition that an older replica would reject. - `AlterCommand::isTTLAlter` (whether restating a TTL schedules a `MATERIALIZE TTL` mutation) and the `MODIFY ORDER BY` no-op detection in `AlterCommands::prepare`. - `ProjectionDescription::operator==`. - For `getTreeHash` to be a faithful identity of a definition, AST nodes that keep semantic state outside of `children` now hash that state (`updateTreeHashImpl` overrides): `ASTTTLElement` (mode, destination, `GROUP BY` keys/assignments, recompression codec), `ASTIndexDeclaration` (name, granularity), `ASTConstraintDeclaration` (name, `CHECK`/`ASSUME`), `ASTProjectionDeclaration` (name), `ASTProjectionSelectQuery` (which clause each child belongs to, previously `SELECT a GROUP BY b` and `SELECT a ORDER BY b` hashed equally), `ASTWindowDefinition` (frame type and boundary kinds). - Since a member that is not a child and is not hashed silently makes two different definitions compare equal, `getTreeHash` documents the requirement, and the classes whose member set is the whole point of the override (`ASTTTLElement`, `ASTIndexDeclaration`, `ASTConstraintDeclaration`, `ASTProjectionDeclaration`, `ASTWindowDefinition`, `ASTSetQuery`, `ASTWithElement`, `ASTWindowListElement`) carry a `sizeof` `static_assert`, so adding a member there fails to compile until it is considered. `ASTSelectQuery`, `ASTProjectionSelectQuery` and `ASTOrderByElement` instead iterate the clause/child-role enumerators, so a newly added one is hashed with no code change. The remaining overrides are not asserted: `ASTColumnsApplyTransformer` (which reaches the `parameters` and `lambda` subtrees) and `ASTSelectWithUnionQuery` hash more than one member, `ASTWithAlias` and `ASTSelectIntersectExceptQuery` one each. Measured: the `static_assert` fires for a new `String`, `UInt64` or pointer member; a lone `bool` can still fit in tail padding, which the comment covers. The negative tests found five ways the comparison was too permissive, each of which master rejects. They are fixed here and each is covered by a test that fails if the fix is reverted: - `ColumnDescription::operator==` used `IDataType::equals`, which ignores the `SimpleAggregateFunction` wrapper (as `MergeTreeData::sortingKeyChanged` already documents), so a replica declaring a plain `UInt64` joined a table whose column is `SimpleAggregateFunction(sum, UInt64)`. It compares `getName` now. - `hasSameExplicitStatistics` kept only the `StatisticsType`, so `STATISTICS(tdigest(1))` and `STATISTICS(tdigest(2))` compared equal although the parameters are a part of the stored definition and survive `SHOW CREATE`. The declaration ASTs are compared as well. - `parameters` and `lambda` of `ASTColumnsApplyTransformer` are not children and were hashed with `updateTreeHashImpl`, which stops at the node itself, so a projection using `APPLY quantile(0.5)` and one using `APPLY quantile(0.9)` hashed equally. `stripArtificialParens` could not reach them either. Both now descend. - The alias was hashed without a length prefix, so `fooIdentifier_bar` and `bar AS Identifier_foo` produced the same byte stream and two projections with different output columns compared equal. - `ASTWindowDefinition` keeps its frame type and boundary kinds outside `children` and did not hash them, so a projection aggregating over `ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW` compared equal to one over `ROWS BETWEEN CURRENT ROW AND CURRENT ROW`. The frame offsets are children and were already covered. One latent bug became reachable and is fixed too: `ASTTTLElement::clone` left `recompression_codec` shared with the source, and `formatDefinition` clones the metadata AST specifically in order to canonicalize it, so `stripArtificialParens` would have written into the live metadata snapshot of a `RECOMPRESS` TTL. No parser or formatter changes and no on-disk/`SHOW CREATE` changes: what the user wrote is preserved; only comparisons became insensitive to formatting. Genuinely different definitions still differ (`a` vs `b`, `a` vs `a DESC`, different index granularity, `CHECK` vs `ASSUME`, different TTL destination). The comparison of whole `CREATE` queries during `RESTORE` (`compareRestoredTableDef`) still compares text: it is a whole-query comparison rather than an expression comparison, and its failure direction is safe (refuses the restore). Tested with a real previous-version server (`clickhouse/clickhouse-server:26.4`): `tests/integration/test_backward_compatibility/test_parenthesized_keys.py` exercises a new replica joining an old table and vice versa, an old replica applying parenthesized `ALTER` log entries written by the new version plus restarts of both replicas, and upgrade + `ATTACH PARTITION FROM`. Unit test `src/Storages/MergeTree/tests/gtest_replicated_metadata_compare.cpp` locks the comparison semantics including the negative cases. Stateless tests: `03471_replace_partition_tuple_key_normalization`, `04604_parenthesized_partition_key_attach_from`, `04612_parenthesized_index_projection_attach_from`, `04622_modify_ttl_parenthesized_no_mutation`, `04646_projection_ttl_parenthesized_zk_metadata`, `04648_alter_parenthesized_zk_metadata`, `04650_replica_column_definition_mismatch_rejected`, `04693_projection_apply_function_name_case` (the function name of a projection's `COLUMNS(...) APPLY` transformer is canonicalized for the comparison only, so `APPLY SUM` and `APPLY sum` compare equal while the stored definition keeps the as-written spelling that older replicas compare byte-for-byte).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110833",
        "createdAt": "2026-07-17T10:35:00Z",
        "updatedAt": "2026-08-13T09:01:25Z",
        "timestamp": "2026-08-13T09:01:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 71
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "v26.6-must-backport"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alesapin",
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110838",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add introspection TCP port",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ClickHouse server now has an introspection port. This is a native protocol TCP listener that starts before the server begins attaching tables and stops only after the tables' detach completes. During these windows, an operator can connect to it with `clickhouse client` and run queries such as `SHOW PROCESSLIST`, `SELECT * FROM system.stack_trace`, or `SYSTEM INSTRUMENT ADD 'QueryMetricLog::startQuery' SLEEP ENTRY 0.5`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110838",
        "createdAt": "2026-07-17T11:07:23Z",
        "updatedAt": "2026-08-13T16:45:51Z",
        "timestamp": "2026-08-13T16:45:51Z",
        "metrics": {
          "reactions": 1,
          "comments": 16
        },
        "labels": [
          "pr-feature"
        ],
        "author": "mstetsyuk",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "evillique"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110883",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Disable TopK dynamic filtering when a sorting projection makes the read in-order",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/110862 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a performance regression where `use_top_k_dynamic_filtering` installed a redundant `__topKFilter` prewhere for `ORDER BY ... LIMIT` queries served in-order by a sorting projection, causing the sort column to be read twice. ### Description Closes: #110862 `optimizeTopK` disables `use_top_k_dynamic_filtering` when the `ORDER BY` column is a prefix of the read's sorted order: there the dynamic prewhere filter is counterproductive, because once the running threshold stabilizes it rejects all remaining rows in sorted order, defeating the early pipeline cancellation that `LIMIT` relies on and forcing a full scan. That guard only checked the base table's sorting key. When a sorting projection whose `ORDER BY` differs from the base table is selected, the read is `ReadType: InOrder` with respect to the projection's sort key, but the guard never matched, so the redundant `__topKFilter` prewhere was installed on top of the projection read, re-reading the sort column that in-order reading already provided. `tryOptimizeTopK` runs in the first plan pass, before projection selection and read-in-order (both second pass), so the guard is predictive. It now also disables dynamic filtering when a normal sorting projection whose `ORDER BY` starts with the sort column and which stores every read column is available and projection optimization is enabled — the exact condition under which such a projection is later chosen to serve the read in-order. Results were correct in all cases; this is a performance-only fix. Verified with `EXPLAIN` that the redundant `__topKFilter` is no longer installed while the projection is still selected and the read stays `InOrder`, and that dynamic filtering is still applied when no projection can serve the order.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110883",
        "createdAt": "2026-07-17T15:32:49Z",
        "updatedAt": "2026-08-13T12:38:11Z",
        "timestamp": "2026-08-13T12:38:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "v26.4-must-backport"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110886",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "MaterializedPostgreSQL: coordinated Replicated/Shared nested tables for HA",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/47655 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): `MaterializedPostgreSQL` can now create its nested tables as `ReplicatedReplacingMergeTree`/`SharedReplacingMergeTree` for high availability, using a new Keeper-based single-active-worker coordination of the replication slot. Controlled by the new settings `materialized_postgresql_table_engine`, `materialized_postgresql_keeper_path` and `materialized_postgresql_replica_name`. ### Documentation entry for user-facing changes The three new settings and the coordinated-failover mechanism are documented in the embedded `Documentation` blocks of `DatabaseMaterializedPostgreSQL.cpp` and `StorageMaterializedPostgreSQL.cpp`, from which the published engine pages are generated. --- ## Problem The nested tables that `MaterializedPostgreSQL` auto-creates were hardcoded to plain `ReplacingMergeTree` ([`StorageMaterializedPostgreSQL.cpp`](https://github.com/ClickHouse/ClickHouse/blob/master/src/Storages/PostgreSQL/StorageMaterializedPostgreSQL.cpp)), so the engine could not be made highly available (issue #47655). Simply allowing `ReplicatedReplacingMergeTree` is unsafe on its own: a PostgreSQL logical replication slot permits only **one** active consuming session, so two ClickHouse replicas consuming the same slot would race on `pg_replication_slot_advance` and silently drop WAL. ## Solution Allow the nested engine to be `ReplicatedReplacingMergeTree` / `SharedReplacingMergeTree`, **gated** behind a new Keeper-based coordination that elects exactly one active worker — the same pattern used by `S3Queue`, the Keeper-coordinated `Kafka` engine and refreshable materialized views: - The active worker holds an ephemeral `/leader` node under `materialized_postgresql_keeper_path` and is the only replica that consumes the slot. - Standby replicas create the nested tables as replicas of the same replicated tree and receive data (including the initial snapshot) through ClickHouse replication, without touching the slot. - When the active worker's Keeper session ends, a standby wins `/leader` and resumes consuming from the slot's `confirmed_flush_lsn` — no reload, no duplication. PostgreSQL's own single-active-session rule on the slot is the ultimate backstop against double-advance during a handover. A graceful stop of the active worker (`DETACH`, a non-last `DROP`, server shutdown) does not wait for that: it releases `/leader` with a **confirmed** removal - the node lives under the server's shared Keeper session, which outlives the database, and nothing re-enters the election on that replica after shutdown, so an unconfirmed (lost-response) removal is resolved on the spot with owner- and version-checked re-checks instead of leaving a stale node that would keep every peer on standby for as long as that session lives. - The slot and the publication are **shared state**, and the whole lifecycle honors that: - A durable `snapshot_completed` marker in Keeper records that the initial snapshot loaded every table. A new leader may resume from `confirmed_flush_lsn` only when the marker exists; otherwise the previous worker died mid-snapshot, and the new one **clears the nested tables**, drops the slot and redoes the snapshot from scratch, so pre-slot rows cannot be lost. Clearing first is required for correctness: a row the dead worker had copied and that PostgreSQL then `DELETE`d has no counterpart in the new snapshot, so without a clear the stale copy would survive (a `ReplacingMergeTree` collapses duplicate keys by `_version` but never turns a now-absent row into a tombstone). The marker is fenced on the live leadership session: it is written through the Keeper session that backs `/leader` (in a multi-request with a check on `/leader`), the snapshot load aborts as soon as that session is lost, and a consumer is never started over a dead leadership session - so a worker deposed mid-snapshot can never mask its successor's replacement snapshot with a stale marker. The redo of the snapshot is fenced the same way: a worker whose leadership session is no longer alive aborts before truncating the nested tables (re-checked per table) and before dropping or recreating the shared slot, so a deposed worker cannot wipe the tables its successor has already reloaded or discard the slot the successor just created. A startup attempt that fails before a consumer got running (most importantly, a coordinated single-table engine whose one snapshot load failed) aborts instead of starting a consumer with nothing to apply - which would advance the shared slot's `confirmed_flush_lsn` while applying no rows - and releases the leadership so a healthy peer can take over. That release never leaves a stale leadership claim behind: a `remove` of `/leader` that could not be confirmed does not prove the node survived, so the claim is dropped in either case, and a leader node this replica created under the current Keeper session but no longer tracks is recognized (by its stored replica name and its owning session) and removed on its next election attempt - a replica never keeps acting as the active worker, or touches the shared slot and snapshot state, without provably holding `/leader`. - A second coordinated `CREATE DATABASE` adopts an existing publication instead of dropping it from under the active worker. - Every replica registers itself under `<keeper_path>/replicas`; dropping the database on one replica keeps the shared slot/publication for the others, and only dropping the last replica removes them from PostgreSQL (fail-close when Keeper is unavailable). The last-replica decision is fenced on the shared `replicas` node so concurrent `DROP DATABASE` on different replicas cannot both act as last, it removes the replica's registration only atomically with winning that fence - a replica that is not the last one keeps its registration until its local nested tables are actually gone, so even a server killed mid-drop stays visible to every later last-replica check - and it runs even for a `DROP DATABASE` issued immediately after a restart, before the background startup task has rebuilt the replication handler. - Recreating a coordinated database right after dropping it is safe: the nested tables are dropped without the delayed-drop window (so their shared Keeper subtrees do not outlive the `DROP DATABASE`), the snapshot insert is never silently deduplicated against block hashes surviving from a previous incarnation of the shared table, a publication leaked by an incompletely dropped setup (no surviving coordination state in Keeper) is dropped by a fresh `CREATE` instead of silently adopted with its stale table set, and a refused (failed) drop never leaves the replica silently dead: every refusable Keeper step of the pre-data teardown runs before the replica's consumer is stopped, and if the drop still fails after that point - the last replica's removal of the shared coordination nodes, or the deletion of this replica's own local nested tables (e.g. Keeper disappearing while a nested replicated table removes its own Keeper metadata) - replication is rebuilt in the background - for the database engine through a new generic `IDatabase::onDropDatabaseFailed` hook that discards the stopped handler and re-runs the startup task, for the single-table engine by re-arming the handler's retrying startup path - so the replica rejoins the setup once Keeper is reachable again. - All replicas of one coordinated setup must agree on the naming-affecting settings (`materialized_postgresql_table_engine`, `materialized_postgresql_schema`, `materialized_postgresql_schema_list`, `materialized_postgresql_tables_list_with_schema`) and must replicate the same PostgreSQL source — the same source database and, for the single-table engine, the same source table (so a single-table engine and a database engine can never share one keeper path): these determine how the ClickHouse names of the shared nested tables and the names of the shared slot and publication are derived. The first replica publishes a canonical fingerprint of them at `<keeper_path>/naming`; a disagreeing replica is rejected — synchronously at `CREATE` time when the setup already exists in Keeper, and fail-close at startup before it registers itself — instead of adopting the same publication yet building a disjoint replicated tree that never receives the other replicas' data. - The shared **table set** is fenced the same way, *before* any nested table is built: the first replica publishes its derived set at `<keeper_path>/table_set`, and a replica whose derived set differs is refused fail-close (the shared publication - from which joining replicas later derive their set - is only created by the elected active worker, so without this fence two fresh replicas starting concurrently could silently build diverging nested tables on one keeper path). The refusal converges by itself once the publication exists, because joining replicas then derive their set from the publication, which was created from the fenced set. - The last-replica teardown is generation-safe: winning the last-replica fence atomically (in one Keeper multi-request) creates an ownership token at `<keeper_path>/teardown`, which is removed only after the shared PostgreSQL slot/publication have actually been dropped. While the token exists, a fresh coordinated `CREATE` on the same keeper path is rejected (synchronously by validation and fail-close at startup), so the pending by-name drops can never delete a new setup's freshly created slot/publication. A replica recovering its own refused drop reclaims its own token, and a retried drop resumes its earlier teardown instead of leaking the slot. - Per-table and database-wide destructive changes are refused (`NOT_IMPLEMENTED`) in coordinated mode: `ATTACH TABLE` / `DETACH TABLE PERMANENTLY` / `DROP TABLE` / `RENAME TABLE` / `EXCHANGE TABLES`, and a database-wide `TRUNCATE DATABASE` / `TRUNCATE ALL TABLES FROM`. Each would only change the local replica (or wipe its copy of the shared replicated data) while the shared publication, tables list, slot/snapshot marker and peer replicas keep the old state, silently diverging the replicas. Recreate the database with an updated `materialized_postgresql_tables_list` instead. New settings (both the database engine and the single-table engine): | Setting | Default | Purpose | |---|---|---| | `materialized_postgresql_table_engine` | `ReplacingMergeTree` | `ReplacingMergeTree` / `ReplicatedReplacingMergeTree` / `SharedReplacingMergeTree` | | `materialized_postgresql_keeper_path` | (empty) | opt-in gate that enables coordination; supports the `{shard}` macro; a per-replica/per-server macro (`{replica}`/`{server_uuid}`, including reached through a config macro) is rejected at `CREATE` time, and so is `{uuid}` unless the DDL carries the UUID (`ON CLUSTER`, a table inside a `Replicated` database, or an explicit `UUID '...'` clause) so that it is provably identical on every replica, and so is a misspelled/unsupported macro (in this path or in `materialized_postgresql_replica_name`) — both settings are macro-expanded during validation exactly as the handler expands them later | | `materialized_postgresql_replica_name` | `{replica}` | replica identity for coordination and the nested replicated engine; must resolve to a distinct value on every replica, which is enforced: the `/replicas/<name>` registration node stores the owning replica's identity, and a name already registered by another replica is rejected (synchronously at `CREATE` time when the registration is already visible); it must also resolve to a single Keeper node name (empty or containing `/` is rejected, since a nested path under `/replicas` would break the last-replica bookkeeping); together with `materialized_postgresql_keeper_path` it forms the coordination identity of the replica, which is treated as immutable once the setup exists: a configuration-only change of a macro these settings expand through is refused at startup, while a `DROP` tears down the identity persisted in the nested tables; a name change made in the one window where the registration already exists while no nested table does yet is recovered from the registration itself, which stores an owner identity no macro feeds into, so the stale registration is removed (on startup and on drop) instead of keeping `<keeper_path>/replicas` non-empty forever and stopping every future drop from becoming the last-replica drop | The replicated/shared engines require `materialized_postgresql_keeper_path` and vice versa (coordination with a plain `ReplacingMergeTree` would leave the standbys without data), and coordination is mutually exclusive with `materialized_postgresql_use_unique_replication_consumer_identifier` (which gives each replica its own slot). Coordination also requires Keeper/ZooKeeper to be configured on the server: a coordinated `CREATE DATABASE` on a server with no Keeper is rejected synchronously at `CREATE` time rather than being accepted and left retrying in the background. All of this validation also applies to a user `ATTACH DATABASE` / `ATTACH TABLE` that spells out the full definition - it is fresh user input, exactly like a `CREATE`; only replaying an already-persisted definition (server startup, and the short `ATTACH` syntax, which re-reads the stored definition) is exempt. Behaviour is unchanged when coordination is not configured. ## Testing - Added integration test `test_postgresql_replica_database_engine/test_coordination.py`: convergence + single leader, leader failover with no data loss/duplication, rejoin, takeover before snapshot completion redoes the snapshot without losing pre-slot rows, a mid-snapshot takeover after a row is deleted in PostgreSQL drops the stale copy (`test_takeover_after_partial_snapshot_drops_stale_deleted_rows`), a second `CREATE` adopts the publication, `ATTACH`/`DETACH`/`DROP`/`RENAME`/`EXCHANGE`/`TRUNCATE` rejection, shared slot/publication kept until the last replica is dropped, a `DROP DATABASE` immediately after restart still unregisters the replica (`test_drop_immediately_after_restart_unregisters_replica`), concurrent `DROP DATABASE` on both replicas tears down the shared state exactly once (`test_concurrent_drop_on_both_replicas_removes_shared_state`), a per-server keeper path macro is rejected (`test_keeper_path_rejects_per_server_macro`), a plain coordinated `CREATE` with `{uuid}` in the keeper path is rejected while an explicit `UUID '...'` clause makes it acceptable (`test_keeper_path_rejects_uuid_macro_for_a_plain_create`), a leaked publication is replaced instead of adopted (`test_leaked_publication_is_not_adopted_by_fresh_coordinated_create`), a refused drop in the restart window keeps the startup alive (`test_refused_drop_in_restart_window_does_not_disable_startup`), a drop refused (via a failpoint) after the replication handler was already stopped recovers without a server restart for both the database and the single-table engine (`test_refused_drop_after_handler_shutdown_recovers_database`, `test_refused_drop_after_handler_shutdown_recovers_single_table_engine`), a drop refused because the local nested-table deletion itself fails likewise recovers for both engines (`test_refused_drop_when_nested_table_drop_fails_recovers_database`, `test_refused_drop_when_nested_table_drop_fails_recovers_single_table_engine`), coordinated `CREATE` rejected without Keeper configured (`test_coordination_requires_keeper_configured`), a joining replica with different naming-affecting settings is rejected at `CREATE` while identical settings converge (`test_join_with_different_naming_settings_is_rejected`), a coordinated single-table engine cannot join a database engine's keeper path because the fenced identity includes the PostgreSQL source (`test_single_table_engine_cannot_join_database_engine_keeper_path`), a bad macro in the keeper path or the replica name fails the `CREATE` synchronously (`test_bad_macro_in_coordination_settings_is_rejected_at_create`), a replica name that is not a single Keeper path component (empty, or containing `/`) is rejected while a plain name on the same keeper path is accepted (`test_replica_name_must_be_a_single_keeper_component`), a duplicate `materialized_postgresql_replica_name` is rejected without disturbing the registered replica (`test_duplicate_replica_name_is_rejected`), a joining replica whose derived table set differs from the fenced one is refused before building any nested table (`test_join_with_different_table_set_is_rejected`), a fresh `CREATE` on a keeper path whose teardown is still pending is rejected and succeeds once the teardown token is released (`test_create_is_rejected_while_teardown_token_is_held`), the three coordination settings rejected by `ALTER DATABASE ... MODIFY SETTING` with an actionable CREATE-time-only message (`test_coordination_settings_cannot_be_altered`), public DDL on a plain database in the startup window (before the background task has built the replication handler) does not touch a null handler and a `DROP DATABASE` in that window still removes the PostgreSQL publication and replication slot (`test_plain_database_ddl_and_drop_in_startup_window`), a plain `DROP DATABASE` quiesces the retrying background startup task so a retry waking mid-drop cannot recreate the publication/slot while the drop is in flight (`test_plain_drop_database_quiesces_retrying_startup_task`), a refused plain drop re-arms the startup task instead of leaving the database mounted but dead (`test_plain_refused_drop_rearms_startup_task`), an `ALTER` of a mutable setting on a coordinated standby is accepted before its consumer exists and survives a failover (`test_alter_mutable_setting_on_standby_survives_failover`), the same `ALTER` is accepted on a former leader demoted back to standby and applied on its next takeover (`test_alter_mutable_setting_on_demoted_leader`), a registration racing the very start of a last-replica teardown is refused by the atomic teardown-token fence (`test_registration_is_fenced_against_concurrent_teardown_token`), a joiner that adopted the shared publication's table set over its own mismatching `materialized_postgresql_tables_list` recreates a missing publication from the adopted set rather than its stale local list (`test_adopted_table_set_survives_publication_recreation`), that adopted set also survives a restart of the joiner, so a replica restarted while the publication is missing recreates it from the table set fenced in Keeper instead of its stale local list (`test_adopted_table_set_survives_restart_and_publication_recreation`), a coordination identity changed by a configuration-only macro change is refused at startup - with nothing registered under the new identity - and the setup resumes once the configuration is restored (`test_coordination_identity_must_stay_stable_across_restart`), a `DROP` after such a change (with the macro's value changed, or the macro removed entirely) tears down the original coordination identity persisted in the nested tables for both engines (`test_drop_after_coordination_identity_change_tears_down_original_identity`, `test_single_table_drop_after_coordination_identity_change_tears_down_original_identity`), a worker whose Keeper session expires while it is loading the initial snapshot aborts without publishing the `snapshot_completed` marker or starting a consumer while its successor redoes the snapshot (`test_lost_leadership_during_snapshot_does_not_publish_stale_marker`), a worker deposed right after entering the redo-the-snapshot recovery branch aborts at the leadership fence without truncating the tables its successor reloaded or dropping the successor's slot (`test_deposed_worker_aborts_redo_snapshot_before_touching_shared_state`), a coordinated single-table engine recreates an externally dropped publication idempotently - with the correctly quoted table name - and resumes replicating (`test_single_table_publication_recreated_after_external_drop`), a coordinated worker whose snapshot load keeps failing aborts each attempt before a consumer exists and releases the leadership, so a healthy peer completes the full snapshot and no WAL is discarded, for both the single-table and the database engine (`test_failed_single_table_snapshot_releases_leadership`, `test_failed_database_snapshot_releases_leadership`), a coordinated `DETACH TABLE ... PERMANENTLY` / `DROP TABLE` issued in the attach/restart window (before the background startup has published the table wrappers) is refused as a true no-op that leaves the nested table consuming (`test_coordinated_detach_in_startup_window_is_a_no_op_rejection`), a read in the recovery window of a refused `DROP DATABASE` still hides PostgreSQL-deleted row versions instead of falling back to the raw nested tables (`test_refused_drop_recovery_window_keeps_wrapped_reads`), a stale registration left behind by a replica whose coordination name changed before it owned any nested table is purged so the last-replica teardown still removes the shared slot and publication (`test_stale_registration_of_a_renamed_replica_is_purged`), a server hard-killed in the middle of a non-last `DROP DATABASE` stays registered in Keeper (the last-replica decision is one atomic operation), so a peer's drop keeps the shared state around the killed replica's surviving data and the restarted replica resumes replicating before a retried drop tears everything down (`test_hard_stop_during_non_last_teardown_keeps_replica_registered`), a graceful stop of the active worker whose fenced `/leader` removal fails (via a failpoint) still frees `/leader` through the confirmed-release re-check, so the peer takes over promptly while the stopped replica's server and Keeper session keep running (`test_graceful_stop_releases_leader_even_when_removal_fails`), a user `ATTACH DATABASE` / `ATTACH TABLE` with a full definition goes through the coordination validator like a `CREATE` (`test_full_attach_database_definition_is_validated`, `test_full_attach_table_definition_is_validated`), and the other negative validation cases. - Added `test_plain_single_table_engine_refused_drop_recovers` to `test_postgresql_replica_database_engine/test_3.py`: the plain (non-coordinated) single-table engine now drops its local nested table before the authoritative PostgreSQL teardown, so a refused (thrown) nested-table drop keeps the replication slot/publication and re-arms the handler to resume from the existing slot, instead of leaving the table mounted but dead with the PostgreSQL objects already removed. - Added `test_rename_and_exchange_table_are_rejected`, `test_database_wide_truncate_is_rejected` and `test_drop_of_individual_table_is_rejected` to `test_postgresql_replica_database_engine/test_1.py`: `RENAME TABLE` / `EXCHANGE TABLES`, a database-wide `TRUNCATE`, and a `DROP` / `TRUNCATE` of an individual table are now also rejected for a plain (non-coordinated) `MaterializedPostgreSQL` database, which previously fell through to the generic `Atomic` DDL and silently diverged the local tables from the replication state (a dropped table stayed in `materialized_postgresql_tables_list` and in the publication, so the consumer marked it skipped while the slot kept advancing); `DETACH TABLE ... PERMANENTLY` remains the supported removal path. - Added `test_count_does_not_include_deleted_rows` to `test_postgresql_replica_database_engine/test_1.py`: `SELECT count()` reads the cheapest column of the nested table - the one-byte `_sign` column itself - and the wrapped read used to skip its `_sign = 1` filter whenever the sign column was among the requested columns, so a count included the durable tombstones a PostgreSQL `DELETE` leaves in the nested `ReplacingMergeTree` table and disagreed with `SELECT *`. The filter is now applied unconditionally (an explicit read of `_sign` is filtered the same way). - Added `test_failed_attach_rolls_back_setting_and_table` to `test_postgresql_replica_database_engine/test_2.py`: a failed `ATTACH TABLE` now rolls back the already persisted `materialized_postgresql_tables_list` extension and the published table wrapper together with the nested table, so the database does not keep claiming a table is attached that never joined the publication; a retry starts from a clean state. - Added `test_failed_detach_rolls_back_and_table_keeps_replicating` to `test_postgresql_replica_database_engine/test_1.py`: a `DETACH TABLE ... PERMANENTLY` whose local nested-table drop throws is now rolled back completely - the table is re-added to replication through a fresh snapshot and the persisted `materialized_postgresql_tables_list` is restored - so the table keeps replicating and the `DETACH` can simply be retried, instead of stranding a live nested table outside the logical database. - Added `test_detach_database_and_reattach` to `test_postgresql_replica_database_engine/test_1.py`: `DETACH DATABASE` used to fail half-way (the generic detach path had already stopped replication when the per-table walk hit the unconditional `DETACH TABLE not allowed` guard), leaving the database mounted but no longer replicating. The internal walk is now let through, so `DETACH DATABASE` unmounts the database cleanly and `ATTACH DATABASE` resumes replication, catching up on changes made while it was detached. - Added `test_detach_permanently_of_last_table_is_rejected` to `test_postgresql_replica_database_engine/test_1.py`: detaching the last replicated table would persist an empty `materialized_postgresql_tables_list`, and an empty list does not mean \"replicate no tables\" - the table set would be re-derived from the current PostgreSQL schema on the next startup, so the detach would not stick. It is now refused up front, leaving the table replicating. - Verified end-to-end against a locally built binary with two real `clickhouse-server` nodes + `clickhouse-keeper` + PostgreSQL (`wal_level=logical`): coordinated `CREATE DATABASE` builds the `ReplicatedReplacingMergeTree` nested tables with correctly macro-expanded per-table Keeper paths; snapshot + ongoing INSERT/UPDATE are consumed by the single leader; a standby serves the same data via ClickHouse replication; killing the leader triggers takeover (~8s) that resumes consumption with matching row/key counts (no loss, no duplication); and a restarted node rejoins as a standby without stealing leadership. Note: `SharedReplacingMergeTree` is only available in ClickHouse Cloud; open-source runs are covered with `ReplicatedReplacingMergeTree`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110886",
        "createdAt": "2026-07-17T15:45:03Z",
        "updatedAt": "2026-08-13T12:48:29Z",
        "timestamp": "2026-08-13T12:48:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 46
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110892",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Explain analyze join stats",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/110668 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add information of internal state of joins to `EXPLAIN ANALYZE`",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110892",
        "createdAt": "2026-07-17T16:06:34Z",
        "updatedAt": "2026-08-13T17:13:11Z",
        "timestamp": "2026-08-13T17:13:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "Fgrtue",
        "state": "open",
        "assignees": [
          "vdimir"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110943",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix ILLEGAL_TYPE_OF_ARGUMENT on nullable MySQL spatial columns",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/110933 Related: https://github.com/ClickHouse/ClickHouse/pull/108944 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `Nested type LineString cannot be inside Nullable type (ILLEGAL_TYPE_OF_ARGUMENT)` when reading a nullable MySQL spatial column (`LINESTRING`, `POLYGON`, `MULTILINESTRING`, `MULTIPOLYGON`, `MULTIPOINT`, or the generic `GEOMETRY`) through the `mysql` table function or the `MySQL` table/database engines. Such columns now fall back to `Nullable(String)` holding the value as MySQL returns it (a 4-byte SRID prefix followed by the WKB payload). `POINT` keeps mapping to `Nullable(Point)`. ### Description `convertMySQLDataType` maps the MySQL spatial types to the `Array`/`Variant`-based ClickHouse geometric types (`LineString`, `Polygon`, `MultiLineString`, `MultiPolygon`, `MultiPoint`, `Geometry`) and then unconditionally wrapped the result in `Nullable`. Those types return `canBeInsideNullable() == false`, so a nullable MySQL spatial column threw at schema-inference time and broke `DESCRIBE`/`CREATE`/`SELECT`. Nullable is MySQL's default for spatial columns and the geometry mapping is on by default, so this broke out of the box (regression from #108944). The fix guards the wrap with `canBeInsideNullable()`; a nullable column whose mapped type cannot be inside `Nullable` falls back to `Nullable(String)`, matching what the query-result-set overload already does for `MYSQL_TYPE_GEOMETRY`. The guard is a property test rather than a list of type names, so a future mapping that cannot be inside `Nullable` is covered too. `Point` deliberately does not take that fallback: it is a `Tuple`, so it can be inside `Nullable` and keeps mapping to `Nullable(Point)`. Added an integration test (`test_storage_mysql/test.py::test_mysql_nullable_geometry`) covering nullable spatial columns through both the `mysql` table function and the `MySQL` table engine. It asserts the inferred `Nullable(String)` schema, the exact bytes of a fallback value, that a MySQL NULL reads back as NULL (not a silently-defaulted empty geometry), and that a nullable `POINT` still reads back as `Nullable(Point)` so an over-broad fix would be caught. The `mysql_datatypes_support_level` setting text and the MySQL engine docs are updated to match, including the previously missing `MULTIPOINT` row in the type-mapping table. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1302` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110943",
        "createdAt": "2026-07-18T15:36:57Z",
        "updatedAt": "2026-08-13T02:12:39Z",
        "timestamp": "2026-08-13T02:12:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 24
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110958",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix toTime key-expression type mismatch under use_legacy_to_time",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107951 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server abort (`Bad cast from type ColumnVector<UInt32> to ColumnVector<Int32>`, a `LOGICAL_ERROR`) when inserting into a `MergeTree` table whose `PRIMARY KEY` or `ORDER BY` uses `toTime(...)` from a session whose `use_legacy_to_time` value differs from the one under which the key metadata was built. The `toTime` legacy/new resolution now follows the context that builds the expression, so a persisted key expression's type no longer depends on the writing session. ### Description `toTime()` resolves to two different functions depending on the `use_legacy_to_time` setting: the new `toTime` returns `Time` (Int32-backed), the legacy variant returns `DateTime` (UInt32-backed). The selection in `FunctionFactory::tryGetImpl` read the thread-local query context and ignored the `context` argument the caller passed. A `MergeTree` table with `PRIMARY KEY (toTime(c1))` / `ORDER BY toTime(c1)` persists only the expression AST. Its primary-index on-disk type comes from `metadata_snapshot->getPrimaryKey().data_types`, derived by rebuilding the key expression with the storage's global context (server-default `use_legacy_to_time`). The part-writer serializes the index with that persisted type. When an `INSERT` runs in a session with a different `use_legacy_to_time`, the write-path key expression re-resolved `toTime` to the other variant, producing a column whose physical type mismatched the persisted serialization, so `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow` hit `assert_cast<ColumnVector<Int32>>(ColumnVector<UInt32>)` and aborted the server (a handled exception in release builds, an abort under debug/sanitizers). Fix: resolve the `toTime` legacy swap from the caller-provided `context` (falling back to the thread-local query context only when no context is supplied). Stored key/sorting expressions are rebuilt with the storage global context, so they now resolve `toTime` consistently with the type persisted in the table metadata, while normal query resolution still honors the session setting. Found by the BuzzHouse fuzzer. - CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109351&sha=b0cca3209ae52b188fa092a31eb1e754db40608d&name_0=PR&name_1=BuzzHouse%20%28amd_msan%29 - Check: `BuzzHouse (amd_msan)`; assertion `Bad cast from type DB::ColumnVector<unsigned int> to DB::ColumnVector<int>` at `SerializationNumber<int>::serializeBinary` <- `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow`. Reproducer: ```sql CREATE TABLE t (c0 Int32, c1 DateTime64 MATERIALIZED nowInBlock64()) ENGINE = MergeTree() PRIMARY KEY (toTime(c1)); INSERT INTO t (c0) SETTINGS use_legacy_to_time = 1 SELECT number FROM numbers(1000); ``` ### DDL behaviour change carried by the fix Previously, `CREATE TABLE` in a session whose `use_legacy_to_time` differed from the server-wide default stamped the stored key type with the session's resolution (e.g. `DateTime` when the session set `use_legacy_to_time = 1` on a server defaulting to `0`). With this fix, the stored key type always resolves under the server-wide default, so the session setting at `CREATE` time no longer affects the persisted key type. This is observable via `DESCRIBE mergeTreeIndex(...)` and is pinned by the test. Upgrade note for that narrow window (table created while the session setting differed from the server global, on a pre-fix binary): after the upgrade the same table resolves its key as `Time`, so parts written before and after store different raw key values for the same timestamp (e.g. `90000` vs `3600` for `01:00:00`); reads, inserts and merges succeed, but a key-range predicate may miss rows from old parts. Where session and global agreed (the overwhelmingly common case), nothing changes.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110958",
        "createdAt": "2026-07-18T23:24:10Z",
        "updatedAt": "2026-08-13T16:39:25Z",
        "timestamp": "2026-08-13T16:39:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110968",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do automatic partition pruning for mutations when it's possible",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ReplicatedMergeTree` mutations now automatically prune the affected partitions based on the `WHERE` part of the query when possible. Also, mutation queries now accept multiple values in the `IN PARTITION` clause. Closes: https://github.com/ClickHouse/ClickHouse/issues/98739 Related: https://github.com/ClickHouse/ClickHouse/pull/99933 This supersedes https://github.com/ClickHouse/ClickHouse/pull/99933 by @alesapin, whose commits are preserved in this branch. The original PR became `CONFLICTING` and could not be updated because maintainer edits are disabled on the fork; the branch had also fallen ~4.5 months behind `master` (in particular, the `MutationCommand` refactor to a text-backed lazy AST required reworking how commands carry `partition`/`partitions`/`predicate`). On top of the original feature, this PR: - Merges current `master` and adapts the code to the `MutationCommand` refactor (commands are re-parsed from `ast_text` via `command.ast()`), the two-argument `getInMemoryMetadataPtr` and the three-argument `ActionsDAGWithInversionPushDown`. - Moves the `optimize_mutations_with_partition_pruning` entry to the current settings-history bucket (review thread on `SettingsChangesHistory.cpp`). - Makes the pruning analysis accept every predicate the mutation itself accepts, instead of hiding failures: the analysis follows the same analyzer selection as the mutation execution - the analysis context matches how the commands will actually be interpreted - for an `ALTER` mutation it is derived from the background context exactly like the background mutation workers derive theirs, so session-only analyzer settings cannot make the submit-time analysis and the asynchronous execution diverge (analyzer-only predicate shapes such as qualified column names work, and a predicate the background execution would reject fails fast at submit time; test `04759_mutation_pruning_background_analyzer_mode`), while a lightweight `UPDATE`, which interprets its commands in the foreground, is analyzed with the submitting context, where the session settings and the current database apply (its predicate is not qualified with the default database, unlike `ALTER` commands; covered by the existing test `03100_lwu_43_subquery_from_rmt`), the column list includes the virtual columns (e.g. `_part`, `_partition_id`) and the `ALIAS` / `EPHEMERAL` columns, and the re-parsed predicate's set operations (`UNION` / `INTERSECT` / `EXCEPT`) are normalized exactly as the mutation execution path does. There is no fallback: an analysis error propagates and fails the mutation rather than silently mutating every partition (review). Without the above, `ALTER TABLE t DELETE WHERE _part = '...'` and similar queries would fail since the setting is enabled by default. Covered by the new test `04612_mutation_pruning_exotic_predicates`. - Requires the `block_numbers` version check only for predicate-pruned commands (review thread on `StorageReplicatedMergeTree.cpp`): for explicit `IN PARTITION` the target set is exact and does not depend on the observed partition list, so the `ZBADVERSION` retry loop is not needed. - In the `ZNONODE` recovery of `EphemeralLocksInPartitions`, reports `ZBADVERSION` to the caller after creating missing partition znodes instead of silently refreshing the version, so a concurrently created partition cannot be missed. - `StorageMergeTree::mutate` validates only explicit `IN PARTITION` ids instead of running the full pruning analysis and discarding the result (on the non-replicated path, parts of unaffected partitions are skipped per part by `canSkipMutationCommandForPart`). - Documents the multi-partition `IN PARTITION` syntax and the automatic pruning behavior. - Recomputes the pruned partition set on every `ZBADVERSION` retry, so a new matching partition created by a concurrent insert on the initiating replica cannot escape the mutation (AI review blocker). Covered by the new failpoint-based test `04613_mutation_pruning_new_partition_race`. - Widens the pruned set with ZooKeeper-only partitions by comparing against the partition set the pruning analysis itself iterated (returned by the pruner from its own parts snapshot), not against a separately read local partition list: a partition the pruner analyzed and ruled out is not re-added, while a partition it could not have seen is still widened in, so a same-replica insert racing with the analysis can neither escape the mutation nor drag an unaffected partition back into it (AI review). Covered by the failpoint-based tests `04614_mutation_pruning_local_partition_race` and `04820_mutation_pruning_analyzed_partition_not_widened`. - Scopes the setting description, the documentation and the changelog entry to the `ReplicatedMergeTree` family: on plain `MergeTree`, predicate-based pruning is not wired in, and an explicit `IN PARTITION` clause should be used (AI review). - Rejects the new `alter->partitions` carrier in `StorageSystemWasmModules` (exactly like the single-partition form) and preserves it in `AlterConversions::createLightweightDeleteCommand`, so the multi-partition clause cannot be silently ignored (AI review). One design note kept as in the original: in the `ZNONODE` recovery path, missing partition znodes are created one per multi-op (create + parent version bump), matching the established pattern of the insert path in `allocateBlockNumber`; batching them into a single parent bump was considered but rejected because a partial `ZNODEEXISTS` would fail the whole multi-op and require a more complex retry. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110968",
        "createdAt": "2026-07-19T04:35:37Z",
        "updatedAt": "2026-08-13T07:14:36Z",
        "timestamp": "2026-08-13T07:14:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 36
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110972",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support parallel replicas for Merge tables and the merge() table function",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/67770 Related: https://github.com/ClickHouse/ClickHouse/pull/95128 Queries over `Merge` tables and the `merge` table function can now be executed with parallel replicas, gated behind a new setting `parallel_replicas_allow_merge_tables` (default off). Previously such queries always ran on a single replica ([#67770](https://github.com/ClickHouse/ClickHouse/issues/67770)): on the CI Logs cluster, `SELECT count() FROM text_log WHERE message LIKE '%test%'` reads at 100 GB/sec using 10 replicas, while the same query through `merge()` reads at 10 GB/sec on one replica. The implementation follows the same approach as parallel replicas over views (`parallel_replicas_allow_view_over_mergetree`) rather than the per-table-coordinator UNION expansion attempted in the earlier PR [#95128](https://github.com/ClickHouse/ClickHouse/pull/95128): the whole first stage of the query (including aggregation) is offloaded to the replicas, and reading from every underlying `MergeTree` table is coordinated by a single reading coordinator, where each underlying table forms its own data stream (the same `stream_id` multiplexing that already serves `UNION ALL` views over `MergeTree` and projection splits). - On secondary replicas, the child reading steps created by `ReadFromMerge` pick up the coordination callbacks from the query context naturally, one announcement per underlying table. - On the initiator, `createLocalPlanForParallelReplicas` finds the `ReadFromMerge` step and switches it into a mode where each child reading step is created as a local parallel replicas reading step wired to the initiator's coordinator (`ReadFromMerge::enableParallelReplicasLocalPlan`). - A storage created by a table function does not exist on remote replicas under its generated id (`_table_function` database), so for `merge(...)` the remote reading step is created without a \"main table\"; otherwise the tables-status check would mark every remote replica as unusable. Eligibility (checked on the initiator and re-checked on the replicas): every underlying table must be a `MergeTree` table, and non-replicated underlying tables additionally require `parallel_replicas_for_non_replicated_merge_tree`. A `Merge` table with a non-`MergeTree` child, no children at all, or `FINAL` falls back to regular single-replica execution, because any child that cannot be coordinated at the level of parts and mark ranges would be read in full by every replica and duplicate its rows in the result. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support parallel replicas for `Merge` tables and the `merge` table function (enabled by the setting `parallel_replicas_allow_merge_tables`): reading from every underlying `MergeTree` table, and the whole first stage of the query, is now distributed across the replicas of the cluster. Closes [#67770](https://github.com/ClickHouse/ClickHouse/issues/67770). ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110972",
        "createdAt": "2026-07-19T06:54:53Z",
        "updatedAt": "2026-08-13T00:00:44Z",
        "timestamp": "2026-08-13T00:00:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 32
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:110997",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix data race on DataTypeAggregateFunction version during Native serialization",
        "text": "Related: found by the `arm_tsan` and `azure, amd_tsan` Stress tests (STID 3977-4818, ThreadSanitizer data race). No existing issue. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results reading `AggregateFunction` states after one client requested them at an older protocol revision. The serialization version chosen for that one response was written onto the data type object shared by the whole table, so it stayed there: every later query read the states at that version, and for `sumMap` over `Decimal32` that returns wrong values, while the column also lost the version in `system.columns` and on the wire. The same in-place write was a data race between concurrent queries serializing such a column in the `Native` format. ### Description A single `DataTypeAggregateFunction` instance is shared across query result blocks: it lives once in the table's column description and is aliased by shallow column copies. `NativeWriter`/`NativeReader` called `setVersionToAggregateFunctions`, which walked to the leaf type and wrote its `mutable version` field in place. Two concurrent `Native` serializations of the same aggregate-function-typed column then raced on that field. Reports: * https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109496&sha=62eeb400eafe86cedff9b5a4e9a36aa722633cda&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29 * https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b59441bd06c2fcb6b103a30874528cc398afc723&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 Both racing stacks are `setVersionToAggregateFunctions` -> `DataTypeAggregateFunction` version setter, via `NativeWriter::write` -> `TCPHandler::processOrdinaryQuery`/`sendData`. Fix: instead of mutating the shared type object, replace the versioned leaf with a copy that carries the version via the constructor (the same way the binary-encoding decode path builds versioned types). The in-place `setVersion`/`updateVersionFromRevision` mutators and the `mutable` qualifier are removed so the object is immutable after construction. This also removes a latent issue that was worse than the race itself. `NativeWriter` passes `if_empty = false` for a client older than `DBMS_MIN_REVISION_WITH_AGGREGATE_FUNCTIONS_VERSIONING`, which unconditionally forced version 0 onto the shared type. Every later query then kept that 0 (`if_empty = true` sees a version already set), so it advertised a type name without a version while serializing version-0 states, and the receiving client - deriving the version from the server revision - deserialized them as version 1. #### Preserving custom type names Because the leaf is now replaced rather than mutated, the rebuilt tree is what the caller ends up with, so the rebuild must not lose anything. Rebuilding the wrappers through `transformTypesRecursively` recreates `Array`/`Tuple`/`Map` via `make_shared` and drops custom type names: | type | expected | with a naive rebuild | |---|---|---| | `Nested(x AggregateFunction(sumMap, ...))` | preserved | `Array(Tuple(...))` | | `SimpleAggregateFunction(anyLast, AggregateFunction(sumMap, ...))` | preserved | `AggregateFunction(...)` | Both are user-visible: the type is sent to the client over `Native`, and on `ATTACH` it becomes the column type in the table metadata. Losing the `SimpleAggregateFunction` name is worse than cosmetic - `AggregatingSortedAlgorithm` and `SummingSortedAlgorithm` recognise such a column by `dynamic_cast` on that very name object, so the column would silently start merging as a plain aggregate function state. So `setVersionToAggregateFunctions` walks the type itself, over exactly the wrappers `transformTypesRecursively` descended into, and returns the original pointer when no leaf changes. `Nullable` is among them: a state cannot be directly inside `Nullable`, but a `Tuple` can, and `Nullable(Tuple(AggregateFunction(...)))` is reachable with `enable_nullable_tuple_type`. A custom name can also sit on the wrapper rather than on the leaf, as in `SimpleAggregateFunction(anyLast, Array(AggregateFunction(...)))`, so a rebuilt wrapper carries the customization of the original too. `DataTypeCustomNamePtr` becomes a `shared_ptr` so a copy of a type can carry the very same custom name object, via the new `IDataType::cloneCustomization`. `Nested` is rebuilt with its custom name kept in sync with the new element types, directly rather than through `createNested`: the latter derives the type from the printed name, and version 0 is deliberately not printed, so a name round trip would turn a leaf explicitly pinned to version 0 back into an unversioned one using the latest version. `callOnNestedSimpleTypes` had no other caller and is removed, so `transformTypesRecursively` (shared with schema inference) is left untouched. ### Testing * `gtest_aggregate_function_version_race` covers the shared-object mutation, nested types, both custom-name cases above, and stress-tests concurrent version assignment over one shared type object. * `04612_aggregate_function_version_custom_type_names` round-trips both types through `Native` (which assigns the version in the writer and again in the reader) and through `DETACH`/`ATTACH`. * `04613_aggregate_function_version_not_sticky` is the one that fails on `master` HEAD. The two tests above assert output that is byte-identical to `master` by design, so neither can. It asks for a `Native` response at a revision below the one that introduced versioning, then checks the column again: on `master` the version 0 forced for that one response stays on the shared type, so a later plain `SELECT finalizeAggregation(s)` reads `([1,2],[10.5,20.25])` back as `([1],[10.5])` and `system.columns` loses the version. The revision is pinned explicitly and the type name the response carries is asserted, so the test cannot pass without that path having run - checked against both ways of it not running, a request that fails and a version assignment that does nothing. * The serialized version bytes are unchanged; the `Native` wire type names are byte-identical to `master` for both types. * 472 related stateless tests (`simple_aggregate`, `nested`, `native`, `geo`, `point`, `polygon`, `aggregate_function`) were run against this build and against a `master` build on the same server config: the failure sets are identical, i.e. no test fails only with this change.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/110997",
        "createdAt": "2026-07-19T14:41:39Z",
        "updatedAt": "2026-08-13T16:08:38Z",
        "timestamp": "2026-08-13T16:08:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 24
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111152",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "docs(h3): fix invalid unidirectional-edge examples (#102845)",
        "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ## What The `h3GetOriginIndexFromUnidirectionalEdge` / `h3GetDestinationIndexFromUnidirectionalEdge` examples used directed edge `1248204388774707197`, which fails `h3UnidirectionalEdgeIsValid` and raises `INCORRECT_DATA` instead of returning the documented indexes. ## Fix - Switch the worked examples to the valid edge `1248204388774707199` (already used by the `h3UnidirectionalEdgeIsValid` example on the same page). - Correct origin → `599686042433355775` and destination → `599686043507097599`. - Keep `docs/en/.../h3.md` aligned with the `REGISTER_FUNCTION` examples in the two C++ sources (those strings feed generated docs). ## Why Repro from #102845 — readers should be able to copy-paste the docs without an exception. ## Notes - AI-assisted; human-reviewed. - Docs-only + FunctionDocumentation string updates; no runtime logic change. - Closes #102845 ### Check - [x] Valid edge per H3 (`is_valid_directed_edge`) - [x] Origin/destination pair matches library on that edge <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1324` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111152",
        "createdAt": "2026-07-20T23:19:41Z",
        "updatedAt": "2026-08-13T13:33:44Z",
        "timestamp": "2026-08-13T13:33:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 18
        },
        "labels": [
          "pr-documentation",
          "manual approve",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "Bartok9",
        "state": "closed",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111219",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add basic implementation of  `DROP PARTITION` for Iceberg",
        "text": "This the first which introduces support for ALTER DROP PARTITION on iceberg tables - supports transformations in partition expression e.g - removes only manifests which were present on the moment of execution of a query - does not support catalogs - does not support schema evolution - does not support partition evolution - does not support \"mixed manifests\", when one manifest file has data files from different partitions (AI always mentions this case, but it is only spec-related case, all majors engines do not produces such manifests, anyway we detect and reject such tables) Related: https://github.com/ClickHouse/ClickHouse/pull/105198 Related: https://github.com/ClickHouse/ClickHouse/pull/109288 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `ALTER TABLE ... DROP PARTITION` support for Iceberg tables. It supports only simple tables, no catalogs, without schema and partition evolution.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111219",
        "createdAt": "2026-07-21T11:55:43Z",
        "updatedAt": "2026-08-13T09:36:16Z",
        "timestamp": "2026-08-13T09:36:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-feature",
          "hold"
        ],
        "author": "Diskein",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111287",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix double free when finalizing -State aggregates under looping combinators",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server crash (double free) that could happen when finalizing an aggregate function with the `-State` combinator nested under a looping combinator (`-Resample`, `-ForEach`, `-Map`), for example `groupArrayStateResample`, if a memory limit was reached during finalization. ### Description Reported on https://github.com/ClickHouse/ClickHouse/pull/110975 (unrelated to that PR). Found by the Stress test (amd_debug): a segfault in `Aggregator::prepareChunkAndFillWithoutKey`, reached from `ConvertingAggregatedToChunksTransform::initialize`. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110975&sha=838d0b61235b06939c8acd923ebd396f51cf5b10&name_0=PR&name_1=Stress%20test%20%28amd_debug%29 Root cause: the `-State` combinator transfers its result by aliasing the raw aggregate state pointer into a `ColumnAggregateFunction` (`AggregateFunctionState::insertResultInto` -> `getData().push_back(place)`); ownership passes to the column. `Aggregator::insertAggregatesIntoColumns` relies on this transfer being atomic per place: on an exception it destroys the whole place exactly once. A looping combinator nested over `-State` aliases many sub-states one at a time; the `push_back` into the column's pointer array can reallocate and, being memory-tracked, throw `MEMORY_LIMIT_EXCEEDED` mid-loop. The already-transferred sub-states are then freed once by the aggregator's full `destroy()` and again by `~ColumnAggregateFunction`, i.e. a double free. Reproducer (crashes without the fix, returns a memory-limit error with it): ```sql SELECT arrayMap(x -> finalizeAggregation(x), state) FROM (SELECT groupArrayStateResample(0, 1048576, 1)(number, number % 20) AS state FROM numbers(100000)) SETTINGS max_memory_usage = 150000000, max_rows_to_read = 0; ``` Fix: reserve the destination columns before the transfer loop so the aliasing `push_back`s cannot reallocate (and therefore cannot throw) once a transfer has started. `ColumnAggregateFunction` used the no-op `IColumn::reserve`, so a real `reserve()`/`capacity()` over its state-pointer array is added. For `-Map`, the (possibly variable-width) key inserts are moved into their own loop before the value transfer, keeping the throwing work out of the aliasing loop. Reserving happens before any aliasing, so a throw there is harmless. The transfer loop is now non-throwing at the point of aliasing, restoring the atomic-per-place contract; results are unchanged. The fix covers all three looping transfer combinators (`-Resample`, `-ForEach`, `-Map`), which share the aliasing path; non-looping combinators delegate a single call and are already atomic. The added stateless test reproduces the crash deterministically via `-Resample` (empty buckets keep memory low until the finalization transfer, so a memory limit reliably lands the throw mid-transfer). `-ForEach` and `-Map` build their sub-states eagerly during aggregation, so they are not deterministically reproducible under a memory limit, but are fixed as the same class via the shared transfer path.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111287",
        "createdAt": "2026-07-21T20:34:03Z",
        "updatedAt": "2026-08-13T17:59:29Z",
        "timestamp": "2026-08-13T17:59:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "nihalzp"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111332",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix RIGHT JOIN with parallel_replicas_min_number_of_rows_per_replica",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111206 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NOT_FOUND_COLUMN_IN_BLOCK` / `THERE_IS_NO_COLUMN` errors (and, on some releases, silently wrong results) for a `RIGHT JOIN` when `parallel_replicas_min_number_of_rows_per_replica` is set and a left-table column is projected. ### Description Closes: #111206 With `parallel_replicas_min_number_of_rows_per_replica` > 0, a `RIGHT JOIN` that projects a left-table column failed: ```sql SELECT r.ver, (l.a + 2) FROM tl AS l RIGHT JOIN tr AS r USING (k); -- Code: 10. Column `a` not found in table default.tr (NOT_FOUND_COLUMN_IN_BLOCK) ``` It also surfaced as `Code: 8 THERE_IS_NO_COLUMN` when aggregating a right-table column, through `RIGHT ANTI JOIN`, and as a silently wrong result on some releases. Leaving the setting at 0, and `LEFT` / `INNER` joins, were unaffected. Root cause: the initiator runs index analysis on the leftmost leaf to estimate the replica count and hands that scan to `createLocalPlanForParallelReplicas`. For a `RIGHT JOIN` the local plan parallelizes the right table (`findReadingSteps` descends the right child), so the leftmost leaf's parts and column list were applied to the right table's read, requesting the left table's columns from the right table. Fix: reuse the pre-analyzed result only when the parallelized scan was not reached through a `RIGHT JOIN` right-branch descent; otherwise let that scan analyze itself, which is what already happens when no analysis is passed. This is also correct for self-joins, where the two sides share the same storage but are distinct occurrences. Note this is unrelated to `automatic_parallel_replicas_mode`, which is a separate feature: a non-zero mode forces `enable_parallel_replicas` to 0, so the two never apply at once. An earlier revision of this PR mislabelled the bug as \"automatic parallel replicas\"; thanks to @nickitat for catching it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111332",
        "createdAt": "2026-07-22T08:41:13Z",
        "updatedAt": "2026-08-13T13:33:36Z",
        "timestamp": "2026-08-13T13:33:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "nickitat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111394",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fsync backup files and directories when writing a backup to local disk",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111320 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `BACKUP ... TO File(...)` / `Disk(...)` now fsyncs the backup data files, the `.backup` manifest and the containing directories before reporting `BACKUP_CREATED`, so an acknowledged backup to local storage survives power loss. Controlled by the new backup setting `fsync_backup_files` (default `true`). Object-storage destinations (`S3`/`Azure`) are unaffected. ### Description Fixes #111320. `BACKUP ... TO File()/Disk()` returned `BACKUP_CREATED` without issuing any `fsync`/`fdatasync` at the destination: not the data files, not the `.backup` manifest, and not the destination directories (there was no `fsync` anywhere in `src/Backups/`). On power loss after the acknowledgement the backup could be lost entirely or left torn, even though `BACKUP_CREATED` is exactly what an operator relies on before dropping the source data. Object-storage destinations were already durable (a completed upload is persisted server-side); only local `File()`/`Disk()` were affected. Report URL: https://github.com/ClickHouse/ClickHouse/issues/111320 (reproduced 3/3 with a `dm-flakey` power-loss simulation). Fix, gated on the new backup setting `fsync_backup_files` (default `true`), following the durability audit family (#68958 -> #111346, #111269 -> #111335): - Two writer hooks with a no-op default on `IBackupWriter`, overridden only by the local `File`/`Disk` writers (`S3`/`Azure`/`Memory`/`Null` inherit the no-op): `syncFileToDisk(file_name)` (fdatasync a written file, covering both the buffered and the native `fs::copy`/`IDisk::copyFile` paths) and `syncDirectoriesToDisk()` (fdatasync every directory the backup created, deepest-first, plus the backup root's parent, via `LocalDirectorySyncGuard` / `IDisk::getDirectorySyncGuard`). - Each data file is synced right after it is written in `BackupImpl::writeFile` (safe under the concurrent write path: each call fsyncs its own file). - In `BackupImpl::finalizeWriting` the `.backup` manifest (or, for archives, the archive file) is synced last, after all data files, so a persisted manifest never precedes its payload. Directory syncing runs for every writer, including the internal writers of `BACKUP ON CLUSTER` which write their own data files. Verified locally with ProfileEvents: `fsync_backup_files=1` issues `FileSync`/`DirectorySync` for the whole backup (data files + manifest + every nested directory); `fsync_backup_files=0` issues none (matching the previous behavior); the backup still restores correctly. Regression test `tests/queries/0_stateless/04412_backup_to_file_fsync.sh` asserts, via the `FileSync`/`DirectorySync` ProfileEvents of the `BACKUP` query in `system.query_log`, that the fsyncs are issued when `fsync_backup_files=1` and are absent when `fsync_backup_files=0`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111394",
        "createdAt": "2026-07-22T13:30:01Z",
        "updatedAt": "2026-08-13T17:48:43Z",
        "timestamp": "2026-08-13T17:48:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-bugfix",
          "manual approve",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "jkartseva"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111427",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Refactor columns deserialization to avoid all assumeMutable calls, simplify the substreams cache, and remove rows_offset",
        "text": "PR #109105 replaced `assumeMutable` with `IColumn::mutate` in the some pars of the deserialization code. When a column is shared via the substreams cache, `IColumn::mutate` clones the whole accumulated column, giving O(rows × granules) cost and a 3–4× deserialization slowdown. Rework the deserialize path so it never clones: - `deserializeBinaryBulkWithMultipleStreams` and its helpers take `IColumn &` instead of `ColumnPtr &`. - The substreams cache never shares a `ColumnPtr`; consumers copy the current range out via `insertRangeFrom`. - `assumeMutable`/`IColumn::mutate`/`const_cast` are removed from the deserialize paths; mutable children come from non-const accessors. - Cache-shared members are dropped from the `Deserialize*State` structs. - The MergeTree reader stack reads into `MutableColumns`. Additionally, remove `rows_offset` from the deserialization API. `rows_offset` (the number of leading rows to skip while deserializing a range) is provably always `0` for every caller, so it is dropped from `deserializeBinaryBulkWithMultipleStreams`, `deserializeBinaryBulk` and all their helpers. This deletes the read-side skip machinery: - the skip loops / `istr.ignore(size * rows_offset)` seeks; - the `Map` bucketed reorder-with-skipped-rows path (which also removes a latent substreams-cache over-count on bucketed `Map`); - the `Object` shared-data `granules_offsets` / `last_incomplete_granule_offset` cursors (the load-bearing `StructureGranule` continuous-read state is kept); - the `SerializationArrayOffsets` helper (`rows_offset`-only); - assorted dead offset arithmetic. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Refactor columns deserialization to avoid all assumeMutable calls, simplify the substreams cache, and remove the always-zero rows_offset parameter.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111427",
        "createdAt": "2026-07-22T16:39:29Z",
        "updatedAt": "2026-08-13T13:30:46Z",
        "timestamp": "2026-08-13T13:30:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "Avogar",
        "state": "open",
        "assignees": [
          "KochetovNicolai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111451",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Release pull request for branch 26.7",
        "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111451",
        "createdAt": "2026-07-22T17:40:24Z",
        "updatedAt": "2026-08-13T16:12:25Z",
        "timestamp": "2026-08-13T16:12:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "release"
        ],
        "author": "robot-clickhouse",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111457",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Compare read-in-order virtual row on its covered sort-key prefix",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/106740 Closes: https://github.com/ClickHouse/ClickHouse/issues/106630 Related: https://github.com/ClickHouse/ClickHouse/pull/110725 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a wrong result (mis-ordered merge) for read-in-order queries with a virtual row when `distinct-in-order` or `LIMIT BY` widens the read to a longer sort-key prefix than `ORDER BY` set it up for, and when a key column fixed by the filter is skipped by `ORDER BY` (e.g. `WHERE b = 1 ORDER BY a, c` on key `(a, b, c)`). The virtual row announced a wrong merge boundary: in release builds the merge could be silently mis-ordered, in debug builds the boundary assertion fired, and a `Nullable` key column after the skipped one threw the `Virtual row has different type` exception. The virtual row optimization now stays enabled in these cases. ### Description Alternative to https://github.com/ClickHouse/ClickHouse/pull/110725: instead of dropping the virtual row conversion when the in-order read prefix changes (and disabling it for skipped key columns), keep the optimization enabled and compare the virtual row only on the sort-key prefix it validly covers. **Root cause.** The read-in-order virtual row announced a wrong merge boundary in two ways: 1. *Widened prefix.* `optimizeReadInOrder` builds the virtual row conversion for the prefix `ORDER BY` needs (e.g. `CounterID`). A later optimization (`optimizeDistinctInOrder`, `optimizeLimitByInOrder`) re-requests the read with a longer prefix (`CounterID, EventDate`), but the `pk_block` width was derived from the conversion's input count, so the extra sort column was default-filled with `0` in `setVirtualRow`. In reverse order `0` understates the real values, so a real row exceeded the announced boundary: `Virtual row boundary violated in MergingSortedAlgorithm ... the virtual row announced UInt64_0 but the source then produced UInt64_1` in debug builds, a silently mis-ordered merge in release builds. 2. *Skipped fixed key.* For key `(a, b, c)` and `WHERE b = 1 ORDER BY a, c`, the fixed key `b` is skipped without an `ORDER BY` counterpart, but the conversion DAG indexed key columns densely, mapping `c` onto key column `b` (visible in `EXPLAIN actions=1`: input `b` aliased to `__table1.c`). The wrong value tripped the boundary check; a wrong type (`Nullable` key) threw the `Virtual row has different type` logical error even in release builds. Moreover, index values of the columns after the skipped key are semantically unusable: the index describes pre-filter data, so the entry `(5, 0, 9)` does not bound the filtered row `(5, 1, 3)` projected to `(a, c)`. **Fix.** - The virtual row conversion outputs only the sort-description prefix it can announce exactly: index values while the key prefix is contiguous, plus constants for fixed columns from `ORDER BY`. A skipped fixed key column ends the index-backed part: the index entry at a mark boundary may hold a filtered-out value for it, so its later components bound nothing in the filtered stream. A column fixed by the filter that stays in `ORDER BY` keeps disabling the virtual row, as before this fix. - The merge compares a virtual row only on the covered prefix and places it first on a covered-prefix tie (equivalent to treating the uncovered columns as minus infinity in the merge order, without materializing any values). The covered prefix is derived from the pk block column names in `MergingSortedAlgorithm` and carried per cursor in `SortCursorImpl::sort_prefix_limit`, honored by the generic `SortCursor::greaterAt`. A truncated virtual row can only occur with a multi-column sort description (its coverage is at least the first column), which always uses the generic cursor, so the single-column specialized queues are unaffected; the JIT comparator is bypassed when a truncated cursor participates. - `ReadFromMergeTree::readInOrder` reads index values for the whole used sorting-key prefix instead of only the conversion inputs. The in-order merges inside the read step sort by the full prefix, and index values are exact bounds for it even after filtering (a filter only removes rows), so these merges always see fully covered virtual rows. - `ReadFromMergeTree::requestReadingInOrder` drops the conversion only when a re-request makes it unsound: a prefix narrower than the one it was built for (the conversion could lose its inputs), or one not fully backed by the primary index. - `setVirtualRowConversions` builds the conversion with `project_inputs` so a raw index column cannot shadow a same-named conversion output when the merge looks sort columns up by name (matters with the old analyzer). This handles all key types uniformly — e.g. a descending `String` column after a skipped key keeps the optimization even though the type has no greatest value to pad with. Verified on release and debug builds (the boundary assertion is compiled only into debug builds): the previously aborting repros now return correct results, and `EXPLAIN` keeps `Virtual row conversions` for the widened and skipped-key reads. The test is based on the one from https://github.com/ClickHouse/ClickHouse/pull/110725, extended with checks that the optimization stays enabled, a descending skipped-key case, a `String` case in both directions, and a fixed key kept in `ORDER BY` (still disabled, as before).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111457",
        "createdAt": "2026-07-22T18:29:12Z",
        "updatedAt": "2026-08-13T17:54:10Z",
        "timestamp": "2026-08-13T17:54:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "vdimir",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111459",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Adaptive Aggregator",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> <img width=\"2938\" height=\"1140\" alt=\"image\" src=\"https://github.com/user-attachments/assets/fb27b9b0-958d-42db-88aa-7a1c0dbe3b35\" /> <img width=\"2928\" height=\"1098\" alt=\"image\" src=\"https://github.com/user-attachments/assets/f115fecc-fa83-44fd-ad37-119c8d831b32\" /> <img width=\"2982\" height=\"898\" alt=\"image\" src=\"https://github.com/user-attachments/assets/9f38d0d8-3019-41ce-8e4d-ecfcc5ac0d67\" /> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New adaptive algorithm for parallel `GROUP BY` (controlled via setting `enable_adaptive_aggregator`, enabled by default): each thread aggregates into its own hash table until it holds `adaptive_aggregator_freeze_threshold` keys and then freezes it, so frequent keys keep updating the small cache-resident tables with no coordination, while rare keys are routed by their hash into per-bucket backlogs and aggregated exactly once, inside the bucket-parallel merge.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111459",
        "createdAt": "2026-07-22T18:45:09Z",
        "updatedAt": "2026-08-13T08:25:02Z",
        "timestamp": "2026-08-13T08:25:02Z",
        "metrics": {
          "reactions": 4,
          "comments": 11
        },
        "labels": [
          "pr-performance"
        ],
        "author": "nihalzp",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111464",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix ORDER BY not applied globally when reading Distributed through Merge",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/111211 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results for a global `ORDER BY` (and `DISTINCT`) when reading a `Distributed` table through a `Merge` engine: the per-shard sorted streams were concatenated instead of merge-sorted, so `LIMIT` could return rows local to one shard instead of the global top rows, and `DISTINCT` could return a doubled result multiset. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/111211 `ReadFromMerge::initializePipeline` narrows the united child pipeline with `narrowPipe`, which concatenates streams via `ConcatProcessor` and does not preserve per-stream order. When a child is read at a partial stage that emits already-sorted streams (e.g. a `Distributed` child produces one sorted stream per shard), the outer plan adds a merge-only sorting step that assumes each input stream is individually sorted. With `max_threads = 1` the two sorted shard streams were concatenated into one, the merge-only sort then had a single unsorted input and became a no-op, and `LIMIT` returned one shard's local top rows. The `should_not_narrow` guard already suppresses narrowing for the two other order-sensitive cases (read-in-order and memory-efficient distributed aggregation). This adds the third: a global `ORDER BY` (no aggregation, or any after-aggregation stage) over a child read above `FetchColumns`. The condition mirrors when the planner/interpreter choose a merge-only sort (`Planner.cpp` `isFromAggregationState` / `InterpreterSelectQuery.cpp` `from_aggregation_stage`). Reproducer (before: `297, 294, 291`; after: `1297, 1294, 1291`), see #111211. The same `narrowPipe` concatenation also corrupts the result multiset under `DISTINCT` (with `optimize_distinct_in_order`, the second shard's sorted run survives adjacent-only dedup, doubling the rows). That query shape is already covered by the guard added here, so no extra code change was needed; the regression test now also asserts the `DISTINCT` / `DISTINCT ON` row counts (reported by @ zlareb1 on #111211).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111464",
        "createdAt": "2026-07-22T18:56:16Z",
        "updatedAt": "2026-08-13T13:29:24Z",
        "timestamp": "2026-08-13T13:29:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 15
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111494",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Text index: add trivial count optimization",
        "text": "Currently, the text index direct read optimization deserialize the sparse index, dictionary block and postings when there is token that exists in the index. Once the postings is read from disk, it fills the newly created boolean virtual column with postings data. With this optimization, we aim to reduce reading postings from disk and creating a virtual column. Instead we can answer queries using the token metadata from the dictionary block for specific query patterns as follows:. 1. `SELECT count() FROM table WHERE hasToken(column, 'foo');` 2. `SELECT count() FROM table WHERE hasAnyTokens(column, ['foo', 'bar']);` 3. `SELECT count() FROM table WHERE hasAllTokens(column, ['foo', 'bar']);` For the 1. case, we can avoid reading postings at all and use the cardinality metadata stored in the dictionary block to answer the query. For 2. and 3. cases, we would still read the postings but can avoid creating a virtual column. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Returns `COUNT()` queries directly from the text index cardinality metadata.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111494",
        "createdAt": "2026-07-22T22:33:56Z",
        "updatedAt": "2026-08-13T16:31:42Z",
        "timestamp": "2026-08-13T16:31:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-performance"
        ],
        "author": "ahmadov",
        "state": "open",
        "assignees": [
          "Ergus",
          "CurtizJ",
          "rschu1ze"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111597",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Recover an intact part from an empty columns.txt instead of losing it",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/111373 Related: https://github.com/ClickHouse/ClickHouse/pull/111414 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed data loss where a MergeTree part with an empty (zero-byte) `columns.txt` was detached as broken on load, discarding all of its rows. An empty `columns.txt` is now treated like a missing one: for wide parts the column list is rebuilt from the part's own metadata, matching the existing behavior for an absent file. ### Description `IMergeTreeDataPart::writeMetadata` rewrites `columns.txt` in place (no atomic rename, no fsync). An interrupted rewrite (crash, power loss, `ENOSPC` between truncate and write) can leave a zero-byte `columns.txt` in an already-committed part directory. On load, an absent `columns.txt` already self-heals: for a wide part the else-branch rebuilds the column list from the table metadata and rewrites the file. But an empty `columns.txt` was read anyway, and `NamesAndTypesList::readText` starts with `assertString(\"columns format version: 1\\n\", buf)`, which throws `CANNOT_PARSE_INPUT_ASSERTION_FAILED` on the empty stream. The otherwise-intact part is then detached as broken and every row is lost. An empty file is strictly less broken than a missing one, so bricking the part for it is illogical. This routes an empty `columns.txt` through the same rebuild path as a missing one. Emptiness is decided via `IMergeTreeDataPart::readFile`, which forces `pread`, so a zero-byte file reports `eof` cleanly instead of faulting under a randomized mmap read method. Compact and patch parts still require `columns.txt` (they cannot rebuild it), so their behavior is unchanged. The rebuild reconstructs the persistent virtual columns the part physically carries (`_row_exists`, `_block_number`, `_block_offset`), not only `getAllPhysical()`. Dropping `_row_exists` would silently discard a lightweight-delete mask (deleted rows would reappear); dropping any of them would fail `columns_substreams.txt` validation and detach the part. Column presence during the rebuild is taken from `columns_substreams.txt`, which records exactly the columns physically written, in order, independent of the on-disk serialization. A column's default serialization can enumerate different streams than were stored (a bucketed Map writes `m.buckets_info`, `m.0.keys`, ... instead of `m.keys`, ...), so probing the default serialization would misjudge such a column absent and drop it, turning recovery back into silent data loss and tripping `columns_substreams.txt` validation. Parts predating that file fall back to enumerating each column's own non-ephemeral streams (present only when every such stream exists), matching `MergeTreeDataPartWide::hasColumnFiles`. These gaps also affected the pre-existing missing-`columns.txt` path. Test `04545_empty_columns_txt_not_fatal` covers wide-part shapes: plain, `_block_number`/`_block_offset`, a lightweight-delete `_row_exists` mask, a `Tuple` and a `Map` (including bucketed serialization), a shared-offset Nested sibling added by `ALTER`, and recovery via the legacy stream-enumeration path when `columns_substreams.txt` is also absent. In each, truncating `columns.txt` to zero and reloading keeps the correct rows/values and persists a complete rebuilt file; a missing `columns.txt` still self-heals. Fails on `master`, passes with the fix.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111597",
        "createdAt": "2026-07-23T12:22:17Z",
        "updatedAt": "2026-08-13T07:48:09Z",
        "timestamp": "2026-08-13T07:48:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111720",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Prepare changelog for 26.8",
        "text": "Automated daily preparation of `CHANGELOG.md` for the upcoming 26.8 release. Every day the `NightlyChangelog` CI job appends the raw changelog entries for the pull requests newly merged into `master` (generated with `utils/changelog/changelog.py`) as one commit, and edits them following `.claude/skills/edit-changelog/SKILL.md` as a separate commit, so both the raw and the edited state stay reviewable. The point up to which entries were generated is recorded as a `Changelog-generated-up-to:` trailer in the generate commits. This pull request stays a draft until the release. The release manager finalizes it manually: fills in the release date and the presentation/video links (the `FIXME` placeholders), reviews the entries, and marks it ready. ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111720",
        "createdAt": "2026-07-24T04:09:28Z",
        "updatedAt": "2026-08-13T04:18:24Z",
        "timestamp": "2026-08-13T04:18:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "do not test",
          "pr-not-for-changelog"
        ],
        "author": "clickhouse-gh[bot]",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111770",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Disable uniq, uniq_v2 for high cardinality column types",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111291 This PR adds optional all-distinct materialization for automatic cardinality statistics. It is intended for column types where building a full cardinality sketch can be expensive and the data is commonly close to unique. New user-visible MergeTree settings: - `auto_statistics_assume_floats_distinct` - `auto_statistics_assume_long_strings_distinct` - `auto_statistics_long_string_distinct_min_length` - `auto_statistics_long_string_distinct_probe_rows` The settings are disabled by default. When enabled, eligible automatic `uniq` / `uniq_v2` statistics can be represented by an `assumed_all_distinct` implementation that estimates cardinality from non-NULL row counts instead of scanning all values into a distinct sketch. ### Implementation notes The implementation keeps the user-facing policy separate from the low-level statistics mechanics: - `StatisticsAssumedAllDistinct` is a lightweight `IStatistics` implementation that stores only cardinality. - `StatisticsUniqStringProbe` owns the stateful probing logic for long String / FixedString columns. - `UniqAssumedAllDistinctPolicy` centralizes the build/merge decisions, replacement logic, and compatibility rules. - `ColumnStatistics` remains responsible for orchestration only: it asks the policy for optional build/merge decisions and applies them. - The merge path now uses an optional `AssumedAllDistinctMergeDecision`, matching the existing `decideBuild` style and avoiding a disabled “plan” state. ### Tests Added coverage for: - Float columns with the assumption enabled/disabled. - Long and short String columns. - Nullable values. - Explicit `assumed_all_distinct` materialization. - Materialization on insert, materialization via mutation, and merge behavior. ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added MergeTree settings to materialize automatic `uniq`/`uniq_v2` statistics for Float and long String columns with an `assumed_all_distinct` cardinality model.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111770",
        "createdAt": "2026-07-24T10:54:30Z",
        "updatedAt": "2026-08-13T00:31:40Z",
        "timestamp": "2026-08-13T00:31:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "cv4g",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111794",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add unordered stream modifier",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM UNORDERED modifier: skip the per-snapshot commit-order sort depends on https://github.com/ClickHouse/ClickHouse/pull/110653 (not for functional reason, only test) cc @alesapin @Michicosun",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111794",
        "createdAt": "2026-07-24T13:31:48Z",
        "updatedAt": "2026-08-13T17:56:27Z",
        "timestamp": "2026-08-13T17:56:27Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-improvement",
          "pr-synced-to-cloud"
        ],
        "author": "SmitaRKulkarni",
        "state": "closed",
        "assignees": [
          "Michicosun"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111830",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject truncated/incomplete AI text function responses",
        "text": "The AI text functions (`aiGenerate`, `aiClassify`, `aiExtract`, `aiTranslate`) could silently return a truncated answer. When a provider stops early — most commonly by hitting the `max_tokens` limit — it reports this in the response, but the shared base class `FunctionBaseAI` never inspected the signal and returned whatever partial text came back. For example, `aiGenerate('Write three sentences about the ocean.', map(..., 'max_tokens', '5'))` returned the fragment `The ocean covers more than` with no error. This change normalizes each provider's native stop reason (OpenAI `finish_reason`, Anthropic `stop_reason`) into a canonical `FinishReason` enum, so the base class can make a single completeness decision without knowing the provider dialect: - `Truncated` (token/context limit) → throws `AI_PROVIDER_RESPONSE_TRUNCATED`. - `ContentFilter` (OpenAI `content_filter`, Anthropic `refusal`) and `ToolCall` → throws `AI_PROVIDER_RESPONSE_INCOMPLETE`. - `Complete` (natural end, or a caller stop sequence such as Anthropic `stop_sequence`) and `Unknown` (unrecognized reason) are accepted, so benign non-`stop` reasons are not misclassified as truncation. The rejection is thrown inside the existing per-row `try`, so it is non-retriable (retrying would hit the same limit) and honors `ai_function_throw_on_error`: with `1` the exception propagates; with `0` the row becomes the column default. `aiEmbed` is unaffected (embeddings have no finish reason). Integration tests in `test_ai_functions` cover truncation (throw + graceful), content-filter, an accepted unknown reason, and the Anthropic `stop_sequence` (must not throw) and `max_tokens` (must throw) cases. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): AI text functions (`aiGenerate`, `aiClassify`, `aiExtract`, `aiTranslate`) now reject truncated or otherwise incomplete provider responses (e.g. when the model hits the `max_tokens` limit) instead of silently returning partial output. Behavior follows the `ai_function_throw_on_error` setting.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111830",
        "createdAt": "2026-07-24T17:24:29Z",
        "updatedAt": "2026-08-13T17:37:56Z",
        "timestamp": "2026-08-13T17:37:56Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "george-larionov",
        "state": "open",
        "assignees": [
          "rschu1ze"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111852",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "fix(Silk): honor O_NONBLOCK in the fiber TLS BIO",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... --- This is a latent-interaction bug that appears only when **three** components are combined — none is wrong on its own: 1. **Silk fiber sockets** (#107680): the fiber-aware OpenSSL BIO `silkBioRead`/`silkBioWrite` submits io_uring I/O and parks the caller for the socket's send/receive timeout. That is correct for a blocking read — but the BIO never consults the fd's `O_NONBLOCK` flag, unlike OpenSSL's default socket BIO, which returns `EAGAIN` immediately when the flag is set. 2. **TLS** (`USE_SSL`): the affected path is the TLS BIO. The plain-socket staleness check uses a raw `recv(MSG_PEEK | MSG_DONTWAIT)` and never touches this code, so the problem is TLS-only. 3. **The connection pool's `SSL_peek` staleness probe** (#110402): before reusing a pooled connection, `DB::getSocketState` flips the fd non-blocking via `ScopedNonBlocking` (a raw `fcntl(F_SETFL, O_NONBLOCK)`, behind Poco's back) and calls `SSL_peek`, expecting an immediate `EAGAIN` — its comment reads *\"The socket is non-blocking, so this never blocks.\"* #110402 introduced this probe, replacing the previous `poll`-based keep-alive disconnect check. Only with all three present does that non-blocking `SSL_peek` route through the silk BIO, which ignores `O_NONBLOCK` and blocks for the receive timeout left on the socket by the previous request. Each component is individually correct; the fix lands on the silk BIO because it is the one whose behavior diverges from OpenSSL's socket-BIO contract (honoring `O_NONBLOCK`), while the TLS layer and the `#110402` probe are behaving as intended. On `master` today the silk BIO has no wired production consumer — it is infrastructure — so this three-way combination is not yet reachable in a shipped server; it was reproduced with downstream work that routes object-storage-disk connections through silk fiber sockets. **Impact.** Every borrow of a pooled TLS connection to an object-storage disk pays a timeout it should not. A server loading tables from an HTTPS object-store disk at startup does many such borrows and stalls — a deterministic ~13.5 s in a local reproduction, and an unbounded boot hang (never reaching \"Ready for connections\", no error logged) with production timeouts or a zero/unset receive timeout, where the wait becomes a deadline-less `future.wait()`. **Root-cause evidence** (local TLS-MinIO reproduction): a server-side request trace showed each request arriving only *after* its wait expired; the ~13.5 s decomposed exactly into the adaptive per-method receive timeouts paid in sequence (GET 500 ms + PUT 3000 ms + DELETE 10000 ms, `ConnectionTimeouts.cpp`); `ss` showed frozen `bytes_sent` through each stall; and a live backtrace was parked at `silkBioRead` ← `SSL_peek` ← `getSslSocketState` ← `isStale` ← `getConnection`. A control with silk sockets disabled does the same step in <15 ms. The plain-HTTP staleness probe uses a raw `recv(MSG_PEEK|MSG_DONTWAIT)` and is unaffected — the bug is TLS-specific. **Fix.** When the fd is non-blocking, `silkBioRead`/`silkBioWrite` do a direct `recv`/`send` with `MSG_DONTWAIT` and set the BIO retry flags (immediate `EAGAIN` → `SSL_ERROR_WANT_READ`), matching OpenSSL's default BIO. The fiber/io_uring path is unchanged for blocking sockets. The non-blocking state is read fresh from the fd on each call via `fcntl(F_GETFL)`, because `Poco::Net::SocketImpl::getBlocking()` is a cached flag the raw-`fcntl` probe never updates (and silk sockets reject `setBlocking(false)` outright). Adds a regression test (`SilkFiberSecureSocketTest.NonBlockingPeekDoesNotBlockOnIdleConnection`) that drives the real `getSocketState` path against an idle TLS connection with a 5 s receive timeout and asserts it returns in under 500 ms; without the fix it blocks the full timeout. Not for changelog: the silk fiber BIO is infrastructure with no in-tree production consumer yet, so no released user is affected. Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111852",
        "createdAt": "2026-07-24T20:53:25Z",
        "updatedAt": "2026-08-13T17:09:37Z",
        "timestamp": "2026-08-13T17:09:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "CheSema",
        "state": "open",
        "assignees": [
          "mstetsyuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111867",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use Gaussian centroids for truncated QBit codes",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109405 Related: https://github.com/ClickHouse/ClickHouse/pull/110911 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve reduced-precision `L2DistanceTransposedQuantized`, `cosineDistanceTransposedQuantized`, and `dotProductTransposedQuantized` by reconstructing `p < 8` `QBit(Int8)` codes with Gaussian conditional-mean prefix centroids. Full precision (`p = 8`) remains bit-exact; reduced-precision approximate distances may change, and recall/latency gains are workload-dependent. ### Motivation The existing reduced-precision path selects one middle fine Lloyd-Max reconstruction level for every truncated prefix. That value is not the conditional mean of the complete Gaussian interval represented by the prefix. This change uses the conditional mean `(phi(lo) - phi(hi)) / (Phi(hi) - Phi(lo))` which minimizes scalar MSE for that prefix under the standard-normal source model of the existing Lloyd-Max codec. The 127 positive values are evaluated from the existing `Float32` boundaries at high precision, rounded once to `Float32`, stored as hexadecimal literals, and mirrored exactly for negative prefixes. This keeps the hot distance loop as one LUT lookup and avoids platform-dependent libm work during LUT initialization. ### Validation - A new stateless regression fails on the exact official `26.7.1.1315` binary for all 20 reduced-precision centroid/sign checks and passes its `p = 8` control. The candidate passes all 20 checks and the control. - An independent all-raw probe covers every raw byte and every `p = 1..8`: expected level counts, finiteness, symmetry, prefix-block invariance, index-order monotonicity, independent Gaussian means (maximum 0 ULP), and bit-exact legacy `p = 8` reconstruction. - An engine probe covers all raw bytes and precisions, non-strided and `QBit(Int8, 16, 8)`, `used_dims` 8/16, dot/L2/cosine, and `optimize_qbit_distance_function_reads` 0/1. Partial-read modes match bitwise; bounded SimSIMD tolerances are used for L2/cosine. - Debug `programs/clickhouse` build passes. Updated `04504_transposed_distance_quantized` and new `04628_qbit_lloyd_max_prefix_centroid` match their references with empty stderr in clean `clickhouse local` paths. On 103,000 source-disjoint Nomic embedding vectors (768 dimensions, 200 queries, four randomized-Hadamard seeds), scalar coordinate MSE decreases at every changed precision. The transformed-space retrieval proxy is deliberately reported separately because it is not a production ClickHouse latency benchmark: | `p` | scalar MSE delta | recall@10 delta (pp) | hit@1 delta (pp) | |---:|---:|---:|---:| | 1 | -3.83% | 0.000 | 0.000 | | 2 | -3.49% | -0.250 | -0.375 | | 3 | -0.89% | +0.538 | +0.750 | | 4 | -15.17% | +1.738 | +1.625 | | 5 | -17.70% | +1.300 | +1.875 | | 6 | -24.08% | +0.800 | +0.125 | | 7 | -44.32% | +0.588 | +0.250 | | 8 | unchanged | 0.000 | 0.000 | There is no universal retrieval improvement claim: `p = 2` regresses slightly in this proxy. There is also no latency claim until a matched Release-build p50/p95 benchmark is available. The intentional compatibility boundary is numerical output at `p < 8`; function signatures, storage, and `p = 8` results are unchanged.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111867",
        "createdAt": "2026-07-25T01:52:18Z",
        "updatedAt": "2026-08-13T16:50:45Z",
        "timestamp": "2026-08-13T16:50:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-improvement",
          "can be tested",
          "v26.7-must-backport"
        ],
        "author": "skuznetsov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111895",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add exclude_data_from_backup MergeTree setting",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111827 ### Changelog category (leave one): - New Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added exclude_from_backup and exclude_data_from_backup MergeTree table settings so BACKUP can skip a table entirely or skip only its data while still restoring its DDL. ### Description Implements #111827: a table-level setting so that `BACKUP` can skip a table's data while still including its DDL, so the table is restorable (empty) later. Useful for tables whose data can be regenerated from a source table (e.g. materialized-view targets), to reduce backup size. - New `Bool` MergeTree setting `exclude_data_from_backup` (default `false`). - Hooked into `BackupEntriesCollector::shouldBackupTableData()`: when the setting is enabled on a `MergeTreeData`-derived table, data collection is skipped for that table; DDL collection is unaffected (existing code path already handles DDL/data independently). - Added `tests/integration/test_exclude_data_from_backup/test.py` covering both the default (`false`, data backed up) and enabled (`true`, data skipped) cases; both pass. **Not yet tested** (feedback welcome): behavior on `ReplicatedMergeTree` specifically, and `BACKUP DATABASE`/`BACKUP ... ALL TABLES` paths (the hook is in the shared per-table code path used by all backup forms, so it should behave the same, but I haven't added explicit coverage for these yet). The setting name/interface is intentionally not MergeTree-specific in wording, per @alexey-milovidov's suggestion on the issue, so it could be adopted by other engines later.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111895",
        "createdAt": "2026-07-25T12:25:30Z",
        "updatedAt": "2026-08-13T11:27:24Z",
        "timestamp": "2026-08-13T11:27:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "adityaksolves",
        "state": "open",
        "assignees": [
          "jkartseva"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111923",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix signed integer overflow in the DateLUTImpl sunday-first week helpers",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/107366 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `toStartOfWeek` and `toLastDayOfWeek` with a Sunday-first week mode returning a wrapped day number for a `Date32` value outside the representable calendar, for example after `toDate32('1970-01-01') + INTERVAL 2147483647 DAY`. The week boundary is now computed at the calendar boundary, consistently with the Monday-first week modes. The wrapping was also a signed integer overflow (undefined behavior). ### Description UBSan reported this in `Stress test (arm_asan_ubsan, s3)` on an unrelated PR, from an AST-fuzzer query: ``` src/Common/DateLUTImpl.h:1316:11: runtime error: signed integer overflow: 2147483647 + 6 cannot be represented in type 'int' ``` Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107366&sha=2f1747ab6b4ba58638193a2da4a7c3f0fc8415e1&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 The two `week_mode` overloads `toFirstDayNumOfWeek(v, week_mode)` and `toLastDayNumOfWeek(v, week_mode)` are the only week helpers in `DateLUTImpl.h` that lack the out-of-LUT-range escape branch all their siblings have. `toDayOfWeek(v)` internally takes the escape and returns the weekday of the clamped day, but the arithmetic that follows runs on the raw unclamped day number, so `v += 6` is undefined behaviour for an `ExtendedDayNum` at `INT32_MAX`. A `Date32` reaches such a value through wrapping `addDays`. This PR fixes two bugs with that one root cause: the reported `INT32_MAX` overflow in `toLastDayNumOfWeek`, and its `INT32_MIN` mirror in `toFirstDayNumOfWeek` (`DateLUTImpl.h:1302`, `-2147483647 - 6`). I added the escape branch to both overloads, mirroring the monday-first siblings, so the day number is clamped into the representable calendar before any arithmetic. `day_of_week % 7` maps Sunday (7) to 0 because these overloads start the week on Sunday. As with the monday-first overloads, the resulting week boundary may lie a few days past the calendar boundary; I kept that regime so the two paths agree for the same input. The in-range code is unchanged: `isOutOfLUTRange` passing bounds the day number to `[-25567, 120273]`, where the arithmetic cannot overflow, so every representable date keeps its current result. The bug is observable without a sanitizer, which is why the changelog category is `Bug Fix` rather than the `CI Fix or Improvement` used for the sibling overflows in this file. Before the fix `toLastDayOfWeek(toDate32('1970-01-01') + INTERVAL 2147483647 DAY, 0)` returned the wrapped day number `-2147483648`, and the first and last day of the same week were `-4294967290` days apart instead of 6. I found no open issue for this. The same UBSan signature fired in 5 unrelated pull requests over the last 45 days and never on master.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111923",
        "createdAt": "2026-07-25T23:38:45Z",
        "updatedAt": "2026-08-13T12:54:32Z",
        "timestamp": "2026-08-13T12:54:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111932",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support the NetCDF format",
        "text": "Adds support for the [NetCDF](https://www.unidata.ucar.edu/software/netcdf/) format, a self-describing binary format for multidimensional arrays that is the standard way climate, weather, oceanographic and other scientific data is distributed. There is a lot of public data in it (ERA5, CMIP, NOAA, Copernicus), and until now the only way to query it with ClickHouse was to convert it first. Reading is supported for the three \"classic\" versions of the format: CDF-1, CDF-2 (64-bit offset) and CDF-5 (64-bit data), and writing produces CDF-2 or CDF-5, whichever the data needs. The files are parsed directly, so there is no new dependency. A NetCDF-4 file, which is an HDF5 file with a different data model on top, is recognized and reported with a message that says how to convert it. **Data model.** Every variable of a file becomes a column, and the rows enumerate the Cartesian product of all the dimensions that the variables use; a variable that does not use some of them is repeated along them. The `to_dataframe` method of `xarray` produces a table with the same columns and the same set of rows, though possibly in a different order: `to_dataframe` puts the dimensions in alphabetical order by default, while ClickHouse keeps the order of the dimensions of the variables. So a file with the dimensions `time`, `lat`, `lon` and the variables `time(time)`, `lat(lat)`, `lon(lon)`, `temperature(time, lat, lon)` reads as a table with four columns and `time * lat * lon` rows. The classic format has no string type, so a `char` variable is read as a String whose length is the last dimension of the variable, when that dimension serves only as the length of the strings; a dimension that anything else in the file uses as a real axis stays in the row space, and such a `char` variable is read as one character per row. Only the variables that a query needs are read, the number of rows comes from the header (so `count()` does not read any data), and an input that cannot be seeked is read into memory instead. **Writing.** Every column becomes a variable over a single dimension named `row`, so a file written by ClickHouse is read back with the same column names and the same rows; the types come back as the closest types of the classic format (a `FixedString` as a `String`, an `Enum` or a `LowCardinality` column as the type it wraps, dates and times as plain numbers, a `Nullable` column as its base type unless read with `input_format_netcdf_fill_value_as_null`). The version of the format is chosen automatically: CDF-5 when a column needs a type that only CDF-5 has or takes more than 4 GiB, CDF-2 otherwise. A `Nullable` column is written with the `_FillValue` attribute, and a column with dates or times gets the `units` attribute of the CF conventions, so `xarray` decodes it back into timestamps. **Settings.** `input_format_netcdf_fill_value_as_null` reads the values equal to the `_FillValue` (or `missing_value`) attribute of a variable as NULL, which is how the CF conventions mark missing data such as sea surface temperature over land. `input_format_netcdf_add_dimension_columns` adds a column with the index along every dimension that has no coordinate variable of the same name, for the files that have none. **Testing.** The test files in `tests/queries/0_stateless/data_netcdf` were written by the netCDF C library, not by this code. Beyond the two functional tests, during development the reader was compared value by value against an independent reference implementation for all three format versions, the output was checked by opening it with the netCDF C library for every supported type family, and 321 truncated and bit-flipped files were fed to the reader — all of them either parsed or failed with a clean error, with no crashes. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `NetCDF` format for both reading and writing. It is a self-describing binary format for multidimensional arrays, widely used for climate, weather and other scientific data. Reading supports all three classic versions of the format (CDF-1, CDF-2 and CDF-5), and writing produces CDF-2 or CDF-5, whichever the data needs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111932",
        "createdAt": "2026-07-26T04:47:02Z",
        "updatedAt": "2026-08-13T05:54:18Z",
        "timestamp": "2026-08-13T05:54:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 25
        },
        "labels": [
          "pr-feature",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111946",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a data race on publication of per-user ProfileEvents counters",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/105056 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a data race on the `ProfileEvents::Counters` parent chain. A freshly constructed per-user counters object was published into a chain that other threads traverse lock-free using a relaxed store, so a thread that observed the pointer was not guaranteed to see the pointee fully constructed. The publication is now release-ordered and the loads that dereference it are acquire-ordered. ### Description `ProcessList::insert` constructs a new `ProcessListForUser` for a user that has no entry yet. Its embedded `user_performance_counters` is initialized with ordinary non-atomic writes, and its address is then published into the shared `ThreadGroup`'s counters chain by `setUserCounters`, which stored `parent` with `std::memory_order_relaxed`. A relaxed store pairs with nothing, so the constructor's writes were not ordered before a consumer's reads. Any thread walking the same chain (`Counters::increment` / `incrementNoTrace` / `incrementSignalSafe`) could therefore dereference a pointer to an object it was not guaranteed to see initialized. `setParent` had the same relaxed publication, and it is used on every thread-group attach and by `attachProfileCountersScope`, which publishes a scope-local `Counters`. The fix is local to `Counters`, which owns both the chain and its traversal: the two publication stores become `memory_order_release`, and every load that dereferences what it loads becomes `memory_order_acquire`. The traversal loads were previously implicit `std::atomic` conversions, that is `seq_cst`, so on the increment path this is a small relaxation rather than a strengthening. On x86-64 all of these compile to a plain `mov`; on ARM the loads go from `ldar` to `ldapr`. The publication stores are cold: once per thread-group attach, and once per `ProcessList` insertion. This is a publication-ordering defect, not a use-after-free and not a rehash invalidation. `user_to_queries` entries are never erased (`ProcessListEntry::~ProcessListEntry` documents this, and `getUserInfo` relies on it), and `UserToQueries` is node-based so element addresses are stable. The write side is the initial construction. The report shape only became possible after #105056, which introduced the `cpus` field and `fetchAdd` that the read side touches. Reproduced locally on a ThreadSanitizer build, where the reported stacks are: ``` Read of size 8 by main thread (mutexes: write M0): ProfileEvents::Counters::fetchAdd(...) src/Common/ProfileEvents.cpp ProfileEvents::Counters::incrementSignalSafe(...) src/Common/ProfileEvents.cpp:1980 DB::(anonymous namespace)::writeTraceInfo(...) src/Common/QueryProfiler.cpp:132 DB::QueryProfilerReal::signalHandler(...) src/Common/QueryProfiler.cpp:595 ... std::condition_variable::wait interrupted in DB::ExternalLoader::LoadingDispatcher::loadImpl(...) DB::registerStorageDictionary(...) src/Storages/StorageDictionary.cpp:385 DB::InterpreterCreateQuery::execute() DB::(anonymous namespace)::loadStartupScripts(...) programs/server/Server.cpp:1134 DB::Server::main(...) programs/server/Server.cpp:3541 Previous write of size 8 by thread T190 (mutexes: write M1): ProfileEvents::Counters::Counters(VariableContext, ProfileEvents::Counters*) src/Common/ProfileEvents.cpp:1678 DB::ProcessListForUser::ProcessListForUser(...) src/Interpreters/ProcessList.h:335 ... operator new of the __hash_node<..., DB::ProcessListForUser> for ProcessList::insert ``` The mutexes on the two sides are disjoint, so nothing synchronized the publication with the traversal. The regression test added to `tests/integration/test_startup_scripts/` reproduces this without any private configuration: a startup script creates a non-lazy dictionary whose `CLICKHOUSE` source authenticates a user other than `default`. `registerStorageDictionary` then blocks the main thread in `ExternalLoader::LoadingDispatcher::loadImpl` while the loader thread, sharing the main thread's `ThreadGroup`, performs the first `ProcessList::insert` for that user. With the profiler sampling the startup thread every 1 ms, the signal handler walks the chain while the new counters are being published. Verified in both directions on a ThreadSanitizer build: without the change the test fails on every attempt with the stacks above, with the change it passes 12 consecutive runs (144 server restarts) with no report.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111946",
        "createdAt": "2026-07-26T09:39:06Z",
        "updatedAt": "2026-08-13T12:45:50Z",
        "timestamp": "2026-08-13T12:45:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111973",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Let read-in-order propagate through SpillingHashJoin",
        "text": "`SpillingHashJoin::hasDelayedBlocks` was hardcoded to `true`, even in the `IN_MEMORY_JOIN` state where nothing is ever delayed. That is the flag gating the read-in-order-through-join and top-k-through-join optimizations, so wrapping a hash join for auto-spilling silently disabled both — and `max_bytes_ratio_before_external_join` defaults to `0.5`, which wraps every hash join. The in-tree comment in `topKThroughJoin.cpp` already describes this as the steady state. `IJoin` documents that \"SpillingHashJoin overrides `keepLeftPipelineInOrder` to forbid switching to GraceHashJoin at runtime\", but no such override existed, so the escape hatch the comment describes was never implemented. This implements it. `keepLeftPipelineInOrder` now pins the join to its in-memory algorithm, and `hasDelayedBlocks` reports `false` from that point on. Pinning is required for **correctness**, not just speed: once the plan drops a sort because the join preserves the left order, a later switch to `GraceHashJoin` would scatter rows by hash and silently return them in the wrong order. The optimizer asks before it commits (`findReadingStep` checks feasibility, and `keepLeftPipelineInOrder` is only called later, if reading in order actually turns out to be possible). So a new `IJoin::canKeepLeftPipelineInOrder` carries the question \"would you preserve the order if I asked?\", defaulting to `!hasDelayedBlocks()` so every other join keeps its current answer. `SpillingHashJoin` answers yes while still reporting delayed blocks, and only stops reporting them once actually pinned — if the optimization turns out not to apply, nothing is pinned and the delayed-block transforms are still built. **Trade-off:** a pinned join can no longer spill, so it holds the whole right side in memory and the memory tracker enforces the limit, exactly as it would with no auto-spill threshold configured. Since dropping the sort and then spilling would be a wrong-results bug, the only alternative is the conservative status quo of never propagating read-in-order through these joins. Both are now reachable: the new setting `query_plan_read_in_order_through_spilling_join` (default `1`) turns the optimization off again, and it is registered in `SettingsChangesHistory` with `previous_value = false`, so `compatibility` set to a version before 26.8 restores the old behavior. `ConcurrentHashJoin` does not spill on its own (its threshold argument only bounds preallocation), so `switchToGraceHashJoin` is the single place the invariant has to hold. The same gate applies to the first-pass `topKThroughJoin`: an `ORDER BY left_key LIMIT n` over a spill-capable `LEFT JOIN` now steps aside for the second-pass read-in-order plan instead of materializing a pushed-down `Sort` and `Limit`, and goes back to pushing them down when the setting is `0`. The handoff keeps the deferral's pre-existing conditions — most notably, `query_plan_join_swap_table` must be explicitly `false`, because under the default `auto` a later optimization may swap the join sides and invalidate the read-in-order plan; with `auto`, such queries keep the pushed-down `Sort` and `Limit`. Making this the default path exposed a pre-existing problem in read-in-order through a `JOIN`, which #110283 pinned down with `04516_join_order_estimation_pruned_parts` a few hours before this branch entered the merge queue. `max_rows_to_read` with `read_overflow_mode = 'throw'` is not checked against the rows a query reads, but against the rows the reading steps announce up front: `ReadProgressCallback::onProgress` substitutes `progress.total_rows_to_read` for `progress.read_rows` whenever the estimate is the larger of the two, and `ReadFromMergeTree` announces `min(part rows, InputOrderInfo::limit)`. `buildSortingDAG` dropped the limit at every `JoinStep`, so a plan reading in order through a `JOIN` always announced whole parts, and `ORDER BY left_key LIMIT 1` over a `LEFT JOIN` that reads 20 rows announced 1010 and failed `max_rows_to_read = 100`. Dropping the limit is right for a join that can filter the left stream, but a `LEFT ALL`/`LEFT ANY` join emits at least one row for every left row, so `n` output rows need at most `n` left rows - duplication only makes fewer left rows necessary. The limit now survives those joins and is dropped everywhere else (`INNER`, `SEMI`, `ANTI`, ...). `InputOrderInfo::limit` never truncates a read - it feeds the announcement, the read-pool task size, the `take_full_part` heuristic and `use_buffering` - so this cannot change results. The behavior was already reachable on `master` with `max_bytes_ratio_before_external_join = 0`, which makes `topKThroughJoin` defer to the second pass exactly as it now does by default. Related: https://github.com/ClickHouse/ClickHouse/pull/111248 Related: https://github.com/ClickHouse/ClickHouse/pull/111972 Related: https://github.com/ClickHouse/ClickHouse/pull/110283 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reading in order through a JOIN now also works when a hash join has an automatic spill-to-disk threshold configured (the default since 26.5). When this optimization applies, the join stays in memory so it preserves the left-side order; set `query_plan_read_in_order_through_spilling_join = 0` to restore the previous behavior, where such joins are not used for read-in-order and remain free to spill. As part of this, an `ORDER BY ... LIMIT` over a `LEFT JOIN` read in order no longer reports the whole table as the number of rows it intends to read, so it no longer trips `max_rows_to_read` with `read_overflow_mode = 'throw'` for a read that stops early.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111973",
        "createdAt": "2026-07-26T19:07:03Z",
        "updatedAt": "2026-08-13T18:00:29Z",
        "timestamp": "2026-08-13T18:00:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111985",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Measure the compressed size of aggregate states in `estimateSizeOfCompressedState`",
        "text": "`Aggregator::estimateSizeOfCompressedState` estimates how many bytes the aggregate states would take when a replica sends them to the initiator, which is how automatic parallel replicas decides whether distributing a query pays off. It builds a `CompressedWriteBuffer` over a `NullWriteBuffer`, but serializes the sampled states into the `NullWriteBuffer` directly and then reads that buffer's counter: ```cpp NullWriteBuffer wb; CompressedWriteBuffer wbuf(wb); // never written to ... aggregate_functions[j]->serialize(place + offsets_of_aggregate_states[j], wb); ... wbuf.finalize(); res += it ? static_cast<size_t>(table.size() * wb.count() / ((it + period - 1) / period)) : 0; ``` So nothing is ever compressed and, despite the function's name and its own `We only interested in the size of compressed state` comment, it returns the plain serialized size. `recordAggregationStateSizes` then stored that single number as all three of `bytes`, `sample_bytes` and `compressed_bytes`, which pins the compression ratio of the aggregation-state statistics to 1. `recordAggregationStateColumnSizes`, documented as `Mirrors the logic of Aggregator::estimateSizeOfCompressedState` for in-order aggregation, feeds the same `AggregationState` counters via `estimateCompressedColumnSize` and does measure the compressed size - so two producers of one statistic disagreed on what it means. This change serializes the sample through the compressing buffer and returns the sample's uncompressed and compressed sizes alongside the extrapolated total, letting the existing compression-ratio machinery in `RuntimeDataflowStatisticsCacheUpdater` apply the states' real ratio. Taking the ratio from the sample rather than compressing the extrapolated size also keeps the per-compressed-block framing overhead out of the extrapolation, where multiplying it by `table.size() / num_samples` would have inflated it. **Effect on the estimate.** Only queries whose output is dominated by compressible aggregate states move. That is exactly the shape that showed up as a CI failure of `03634_autopr_output_bytes_estimation` on master: `query_28` (`MIN(Referer)` states over a two-level hash table with many groups) estimated 59335657 bytes against a recorded 23722663, a ratio of 2.4996 that left the test's 2.5x bound no margin. Reading the estimates out of the failing job's server log, the ten other queries in that test match their recorded values within 1.00..1.41x, because their output is dominated by aggregation keys or by output columns, both of which were already measured compressed. So the recorded values expect the compressed size and this defect is why that one entry looked wrong. **Two small-sample corrections found while validating this in CI.** The compressed format writes a checksum and a block header in front of every block, so a sample of a few bytes comes out of `CompressedWriteBuffer` larger than it went in and the derived ratio drops below one, *inflating* the estimate. With an early conversion to a two-level hash table (`group_by_two_level_threshold=1`, which the test randomization sets) every one of the 256 buckets holds a handful of states and every per-bucket sample is dominated by that framing, which is what made `04034_autopr_dataflow_cache_reuse_between_different_queries` fail on the first run of this pull request: the two cache-reusing queries lost parallel replicas. When the states are really sent, the framing is amortized over `min_compress_block_size` of data, so such a sample is now reported as incompressible rather than as expanding, keeping the ratio at 1 - exactly what the caller assumed before the ratio was measured at all. The new test `04653_autopr_state_size_estimate_small_buckets` pins `group_by_two_level_threshold` to 1 and fails without that. Reviewing the same statistic end to end also turned up the mirror-image defect on the in-order-aggregation side, reported in review: `recordAggregationStateColumnSizes` totalled `IColumn::byteSize`, which for a `ColumnAggregateFunction` whose states live in a foreign arena (as `AggregatingInOrderTransform` hands them over) counts one pointer per row, so arena-backed states - `min(String)`, `groupArray`, `uniqExact` - were under-counted. That column is now sized from its serialized states too, sampling at most as many of them as the hash-table producer samples per bucket. **Recorded values.** The values in `03634_autopr_output_bytes_estimation` are empirical and need `test.hits`, so they are re-measured from this PR's own CI run rather than guessed - the test reports every query that lands outside the 2.5x band, and I will update whatever it reports. `query_1` and `query_20` are single `count()` states whose estimate becomes the ~29-byte compressed-block framing instead of a 2-3 byte varint; the test's `NOT (res.2 < 100 AND res.3 < 100)` guard already covers those. Verified so far: both changed translation units pass a full `-fsyntax-only` typecheck against master. The estimate itself is validated by CI, since reproducing it needs the stateful dataset. Related: https://github.com/ClickHouse/ClickHouse/pull/111981 That PR is the interim, test-only fix for the master CI failure linked below: it re-records `query_28` as the ~58 MB the estimator currently reports. **The two changes touch the same line and are mutually exclusive.** If #111981 merges first, this PR must set `query_28` back to the post-fix measurement (expected to land near the original 23722663); if this PR merges first, #111981 should be closed as unnecessary. https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=5baed0f5a10c333cd220b9646d6ef4647425079d&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20distributed%20plan%2C%20parallel%29 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the estimation of the size of aggregate states used by automatic parallel replicas: it measured the serialized size of the states instead of their compressed size, overestimating how much data a replica would send for queries that aggregate into large, compressible states.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111985",
        "createdAt": "2026-07-26T20:39:07Z",
        "updatedAt": "2026-08-13T09:45:18Z",
        "timestamp": "2026-08-13T09:45:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 25
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:111992",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a mutated part losing files to the temporary directory cleanup",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/107150 Related: https://github.com/ClickHouse/ClickHouse/pull/96376 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a race between a `ReplicatedMergeTree` mutation and the removal of old temporary directories that could publish a mutated part with some of its files missing, making reads of that part and all its later mutations fail. ### Description `MutateFromLogEntryTask::finalize` released `mutate_task` - and with it the RAII guard that registers `tmp_mut_<part>` in `MergeTreeData::TemporaryParts` - before calling `MergeTreeData::Transaction::renameParts`, which is what actually renames the directory on disk. Between those two points the temporary directory of an already precommitted part was not protected from `clearOldTemporaryDirectories`. When the cleanup thread hit that window it started removing the directory while `renameParts` was moving it to the persistent name, so the part became active with part of its files already deleted while `checksums.txt` still listed all of them. From the failing job's `node1` log, the three events are 850 µs apart: ``` 19:34:46.048315 Renaming temporary part tmp_mut_97_0_0_0_7 to 97_0_0_0_7 19:34:46.049102 Removing temporary directory .../tmp_mut_97_0_0_0_7/ <- cleanup thread 19:34:46.049168 Renaming part to 97_0_0_0_7 <- the actual rename ``` Reads of such a part fail, because `MergeTreeMarksLoader` derives the marks path from the part's own checksums and therefore expects the file to exist: ``` Code: 1001. DB::Exception: std::filesystem::filesystem_error: filesystem error: in file_size: No such file or directory [\".../97_0_0_0_7/foo2.cmrk2\"] ``` and every following mutation of the part fails permanently, because it hardlinks the source part's files by the names listed in its checksums: ``` Code: 424. DB::ErrnoException: Cannot link .../97_0_0_0_7/foo2.bin to .../tmp_mut_97_0_0_0_8/num2.bin ... (CANNOT_LINK) ``` This is the flaky `test_rename_column/test.py::test_rename_distributed_parallel_insert_and_select`. That test sets `temporary_directories_lifetime = 1`, which makes the window easy to hit; it has failed this way 12 times in the last two months, including twice on `master`. #107150 reported the same symptom from the reader's side and improved the error message without fixing the race. Now the part is renamed while `mutate_task` is still alive, the way `MergeTreeDataMergerMutator::renameMergedTemporaryPart` already does for merges (\"Explicitly rename part while still holding the lock for tmp folder to avoid cleanup\"). `mutate_task` is still reset before `checkPartChecksumsAndCommit`, so the fallback fetch on a checksum mismatch (#96376) still runs with the guards released. Every other caller of `renameTempPartAndReplace` with `rename_in_transaction=true` renames immediately afterwards; `MutateFromLogEntryTask` was the only one that did not. ### Testing `04512_mutation_temp_dir_cleanup_race` pauses a mutation right before the rename with a new pauseable failpoint and checks that the cleanup thread reports the directory as in use instead of removing it. It fails on the unfixed code (`0 1` - the directory is removed, and the mutation then dies with `Cannot set modification time to file`) and passes with the fix (`1 0`), verified locally on a debug build in both directions. Reported by: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=423ec1ba3ac4288ce5b934235be0c6700e0f7c14&name_0=MasterCI&name_1=Integration%20tests%20%28amd_msan%2C%201%2F8%29 ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/111992",
        "createdAt": "2026-07-26T23:31:56Z",
        "updatedAt": "2026-08-13T08:14:44Z",
        "timestamp": "2026-08-13T08:14:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112152",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keep the function name of a stack frame attributed to a libc++ `__functional` header",
        "text": "Caused by: https://github.com/ClickHouse/ClickHouse/pull/57201 Related: https://github.com/ClickHouse/ClickHouse/pull/100419 `collapseDemangledNames` replaced a stack frame's function name with `?` whenever the frame's source location was a file in a directory ending in `functional`. The intent was to hide `std::function` plumbing frames, whose demangled names spell out the whole captured lambda type and say nothing the neighbouring frames do not already say. But the file of a frame is the source line the *instruction* maps to, which is not necessarily where the enclosing function is defined. An ordinary function can have individual instructions attributed to a libc++ `__functional` header - an inlined `std::function` operation, or compiler-generated code reported with line 0 - and then its name was dropped too, which loses the only useful part of the frame. This PR requires the symbol to name the `std::function` plumbing as well: the type-erasing wrappers (`std::__function::__func`, `__value_func`, `__policy_func`, ...) and the type erasure of `std::function` itself - its `operator()`, its copy, move and callable-taking constructors and assignment operators, and its destructor. The ordinary members of `std::function` (`swap`, `target_type`, `operator bool`, the `nullptr` reset `operator=(std::nullptr_t)`, the empty-constructing `function()` / `function(std::nullptr_t)`, ...) do work of their own, so they keep their names too. Those are still displayed as `?`, exactly as before; every other frame keeps its name - not only an ordinary `DB` function, but also a meaningful libc++ symbol that happens to live in a `__functional` header, such as `std::hash<String>::operator()` from `__functional/hash.h`, or the generic invocation helpers `std::invoke` / `std::__invoke` / `std::mem_fn`, which are not `std::function`-specific and whose frames can name the callable they dispatch to. This is much more likely in a build with ThinLTO enabled, where `std::function` calls are inlined across translation units, so it went unnoticed for a long time: `amd_cfi` is the only integration-test build with ThinLTO on, and its `test_crash_log` failure is what surfaced it. In that build, the frame that actually terminated the server was displayed as ``` 8. contrib/llvm-project/libcxx/include/__functional/function.h:0:7: ? @ 0x1c01f41f ``` both in the fatal log and in `trace_full` in `system.crash_log`, where `0x1c01f41f` is inside `DB::executeQuery` (the symbol resolves correctly - only the display suppressed it). Official release builds also use ThinLTO, so the same frames were being anonymised for users. `collapseDemangledNames` becomes a static member of `StackTrace` so that the heuristic is covered by a unit test in every build, rather than only by the weekly ThinLTO job. Fixes `test_crash_log/test.py::test_crash_log_extra_fields[terminate_with_exception-trace_full]` and `[terminate_with_std_exception-trace_full]`, seen in https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b05161aa67d75ad84151a392f3e87156efd42842&name_0=WeeklyCFI&name_1=Integration%20tests%20%28amd_cfi%2C%202%2F4%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a stack frame being displayed as `?` instead of its function name in fatal log messages and in the `trace_full` column of `system.crash_log`. It affected frames of ordinary functions that have code attributed to a libc++ `__functional` header, which is common in release builds. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1334` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112152",
        "createdAt": "2026-07-27T17:43:33Z",
        "updatedAt": "2026-08-13T14:31:40Z",
        "timestamp": "2026-08-13T14:31:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 16
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112250",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Materialize column statistics on INSERT by default",
        "text": "This PR contains https://github.com/ClickHouse/ClickHouse/pull/109454 minus `materialize_statistics_on_insert_max_table_size` (which I'm happy to introduce in a second step). Made a separate PR to speed up the integration of the feature (the new behavior has high demand and the original PR is stuck since three weeks). If the original PR gets merged first, we can close this one. <!-- Related: https://github.com/ClickHouse/ClickHouse/pull/109454 --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Column statistics are now materialized on `INSERT` by default. This improves estimations for the cost-based join optimizer.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112250",
        "createdAt": "2026-07-28T09:46:11Z",
        "updatedAt": "2026-08-13T09:57:06Z",
        "timestamp": "2026-08-13T09:57:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "rschu1ze",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112304",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Supporting lazily replicated arrays on arrayElement and arraySlice (reduces memory usage and improves performance)",
        "text": "arrayElement and arraySlice now consume lazily replicated arrays (ColumnReplicated, produced by lazy ARRAY JOIN and JOIN under enable_lazy_columns_replication) without materializing them. Follow up of : https://github.com/ClickHouse/ClickHouse/pull/111581 and https://github.com/ClickHouse/ClickHouse/pull/111749 Closes: https://github.com/ClickHouse/ClickHouse/issues/54967 ### Implementation: - A new ReplicatedSource<Base> adapter in GatherUtils wraps any array source (numeric, generic, or nullable) over the compact nested column and remaps each logical row to its nested row through the replication indexes. - arrayElement gets a dedicated path: a per-row gather from the nested data for non-constant indexes (supporting negative indexes, out-of-range defaults, Nullable elements, and arrayElementOrNull), and for constant indexes it executes on the nested rows only and re-wraps the result lazily with the same indexes. - Per-row index arguments (arraySlice offset/length, arrayElement index) are materialized if they arrive replicated - Unsupported shapes (Map, LowCardinality elements) Measured on the test workload (1000-element String arrays, 50× replication by ARRAY JOIN, 100 rows): peak memory drops from 1.12 GB to 29 MB (38×). ### First Query ``` sql WITH materialize(range(1000)) AS large_array SELECT count() FROM system.numbers WHERE NOT ignore( arrayMap(idx -> arraySlice(large_array, idx, 5), arrayEnumerate(large_array))) SETTINGS max_rows_to_read = 262144, read_overflow_mode = 'break', max_memory_usage = 20000000000, enable_lazy_columns_replication = 1 ``` ### Second Query ``` sql WITH materialize(range(1000)) AS large_array SELECT count() FROM system.numbers WHERE NOT ignore( arrayMap(idx -> arraySlice(large_array, idx, 5), arrayEnumerate(large_array))) SETTINGS max_rows_to_read = 16384, read_overflow_mode = 'break', max_block_size = 512, max_threads = 1, max_memory_usage = 20000000000, enable_lazy_columns_replication = {1|0} ``` | Configuration | Rows | Time | Peak RSS | Result | |---|---|---|---|---| | Lazy ON, default block size, 262k rows | 327,045 | ~4.5 s (~72M lambda calls/s) | 3.15 GB | ✅ completes | | Lazy OFF, same settings | — | fails in 0.5 s | — | ❌ `MEMORY_LIMIT_EXCEEDED`: tries to allocate 121.83 GiB for one 65536-row block | | Lazy ON, `max_block_size=512`, 1 thread, 16384 rows | 16,384 | 0.28 s | 272 MB | ✅ | | Lazy OFF, same | 16,384 | 7.26 s | 2.28 GB | ✅ | Tests: a stateless test covering numeric/String/Nullable/Tuple elements, dynamic and constant offsets, negative and Nullable indexes, empty arrays, ARRAY JOIN and JOIN producers, UInt16 replication indexes, plus a performance test comparing both settings. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): arrayElement and arraySlice now read lazily replicated arrays (produced by ARRAY JOIN and JOIN when enable_lazy_columns_replication is enabled) directly, without materializing them. This greatly reduces memory usage and improves performance of queries that index or slice a large array.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112304",
        "createdAt": "2026-07-28T15:36:42Z",
        "updatedAt": "2026-08-13T11:11:30Z",
        "timestamp": "2026-08-13T11:11:30Z",
        "metrics": {
          "reactions": 2,
          "comments": 4
        },
        "labels": [
          "pr-performance"
        ],
        "author": "diegomestre2",
        "state": "open",
        "assignees": [
          "antaljanosbenjamin"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112309",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add hierarchicalKMeans and assignCentroid",
        "text": "Adds functions for computing cluster centroids and assigning new vectors to clusters. Ref : https://github.com/ClickHouse/ClickHouse/issues/112578 ### Changelog category - Experimental Feature ### Changelog entry - Added` hierarchicalKMeans()` and `assignCentroid()` functions.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112309",
        "createdAt": "2026-07-28T16:11:44Z",
        "updatedAt": "2026-08-13T15:14:28Z",
        "timestamp": "2026-08-13T15:14:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "shankar-iyer",
        "state": "open",
        "assignees": [
          "rschu1ze"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112313",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keep an S3Queue file retryable after losing the race for its `processing` node",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `S3Queue`/`AzureQueue` skipping a file forever after losing the race for its `processing` node in Keeper: if the other processor released the file without committing it, the file was not retried by this table until the server restart. ### Description Found during the review of https://github.com/ClickHouse/ClickHouse/pull/108977 (the regression test there has to accept losing one file for exactly this reason). When `ObjectStorageQueueIFileMetadata::trySetProcessing` fails because the `processing` node in Keeper already exists, that node belongs to another processor - another server, or another table on the same server. `afterSetProcessing` then updated the local `FileStatus` to `Processing`, and since the non-processable checks in `trySetProcessing` and `prepareSetProcessingRequests` treat `Processing` as terminal, every later attempt short-circuited on that cached state without ever rechecking Keeper. Unlike `Processed` and `Failed`, `Processing` is not backed by a persistent Keeper node: the foreign processor can release the file without committing it - it can die, or fail the file and reset the node. Stale `processing` node cleanup does not evict `local_file_statuses` either; only the processed/failed node cleanup (`ObjectStorageQueueMetadata.cpp`) does. So the file stayed skipped by this table until its file status was evicted from the cache or the server was restarted. The cached `Processing` state is now remembered as a timestamped observation of the foreign `processing` node (`FileStatus::onProcessingByAnotherProcessor`): it is still reported in `system.s3queue`, and it is respected - the file is skipped without touching Keeper - while the observation is fresh (the new table setting `foreign_processing_node_cache_ttl_seconds`, 5 minutes by default; zero means to always check Keeper). After that, the next attempt probes Keeper again and refreshes the observation if the node is still there, so a file released without a commit is retried within the timeout, while a file being processed by another server costs at most one Keeper probe per timeout instead of one per listing pass. As soon as the foreign processor commits the file, the next probe fails on the `processed` (or `failed`) node instead and the state becomes terminal again. `afterSetProcessing` also keeps the cached state untouched when it is already `Processing` and not marked as foreign: in that case the node belongs to a concurrent local processor sharing this `FileStatus` (tables on the same server using the same Keeper path, insert threads), and its owner updates the state on commit. The cached observation participates in the listing pre-filter (`FileIterator::filterProcessableFiles`) as well: a file with a fresh foreign-`Processing` observation is dropped from the batch before the `processed`/`failed` multi-read is built, so the fresh observation avoids Keeper requests on the listing path too. The terminal states which that pre-filter does discover in Keeper are written back into the cached `FileStatus` (unless the state is owned by an active local processor, whose owner updates it on commit), so `system.s3queue_metadata_cache` follows Keeper instead of keeping a stale `Processing` after another processor has committed the file. The write-back replaces the whole cached record, not only the status: the per-attempt data of an abandoned local attempt is cleared, and a file failed by another processor carries the exception and retries from the `failed` node. The write-back is skipped when the cached state already equals the discovered terminal state: such a record describes a finished local attempt (its `rows_processed` and timings must survive relistings). A cached `Failed`, however, also describes retriable local attempts (`retries < loading_retries`), so it is kept only when its retry count matches the `failed` node payload: when another processor exhausts the retries after a retriable local failure, the cached record follows the terminal node instead of keeping the stale local exception. The set-processing probe (`trySetProcessing`/`prepareSetProcessingRequests`) follows the same contract: when it discovers a `processed`/`failed` node (which can appear after the pre-filter ran), it returns the metadata of that node and `afterSetProcessing` refreshes the whole cached record with the same guards, instead of flipping only the status. A skipped file is not simply dropped from the current listing pass: the file iterator keeps it in a recheck list (also filled when `trySetProcessing`/`prepareSetProcessingRequests` observe a foreign `processing` node), and every batch boundary of the pass takes the files whose observation has expired and runs them through the regular filtering. Within a long listing pass the TTL is therefore honored with batch granularity; files whose observation is still fresh when the listing is exhausted are dropped with the iterator, because the observation timestamps live in the shared file status cache and the next pass re-queues them with the original deadlines. The TTL bounds the retry latency even on an otherwise idle queue, where the polling backoff after zero-row cycles can far exceed it: the streaming task schedules its next cycle no later than the earliest pending recheck deadline. In `Ordered` mode a foreign-held file also blocks the later files of its ordering domain (the scope of one `processed` pointer: a bucket, and a partition within it when partitioning is used) for the current listing pass. Without this, committing a later file advances the `processed` pointer past the held file, and the next listing pass drops it as already processed - losing it forever if the foreign processor never commits it. The file iterator records foreign-held files per ordering domain and drops the later files of the domain both at the listing pre-filter and before handing a file out (a file returned for retry releases its `processing` node); they are re-listed by the next pass. A held file stops blocking its domain as soon as this server wins its `processing` node or a terminal state for it is discovered in Keeper. With several processing threads sharing one ordering domain (`buckets = 1`), the block alone is not enough: it is recorded only when the set-processing attempt of the held file fails, and a later file handed out to another thread before that could still be committed first. The file iterator therefore registers a file whose set-processing outcome is not yet known at hand-out time (following the hand-out order), and a later file of the domain waits until the outcomes of the smaller files are known before starting its own attempt - so the set-processing attempts of one domain serialize for the duration of one Keeper round trip, and a foreign-held discovery blocks the later files before any of them starts processing. `foreign_processing_node_cache_ttl_seconds` is a per-table setting: `ObjectStorageQueueMetadataFactory` shares a single `ObjectStorageQueueMetadata` between all tables with the same `keeper_path`, so the value is not kept there - it travels from `StorageObjectStorageQueue` through `FileIterator` to `ObjectStorageQueueMetadata::getFileMetadata`, and each table uses (and reports in `system.s3_queue_settings`) the value from its own DDL. The setting can be changed on a live table with `ALTER TABLE ... MODIFY SETTING`: the storage keeps the value in an atomic member which the file iterators read through a reference, so the new value (for example, zero, to get a stuck file retried immediately) applies to the running streaming task without recreating the table. Tests: `tests/integration/test_storage_s3_queue/test_foreign_processing_node.py` emulates the foreign processor with a real `processing` node in Keeper, checks that the file is not committed while that node exists, removes it, and asserts that the file is then processed - without the fix the count stays one short. A second test in the same file creates two tables sharing one `keeper_path` with different values of the setting and checks both the reported value and the retry window. A third test emulates the foreign processor committing one file and failing another, and asserts that `system.s3queue_metadata_cache` reports `Processed` (respectively `Failed`, with the exception of the processor which failed the file) instead of a stale `Processing`, and that neither file is ingested by this table. A fourth test checks that `ALTER TABLE ... MODIFY SETTING` shortens the retry window of an already-running table. A fifth test repeats the retry scenario in `ordered` mode, with the foreign processor holding the lexicographically greatest file, asserting that the max processed path does not swallow the skipped file. A sixth test holds the only file of the queue with a two-minute polling backoff configured and asserts that it is retried within the TTL after its holder releases it, not after the backoff. A seventh test parks the `ordered` set-processing attempt at a failpoint between the initial state read and the multi request, fails the file from a fake foreign processor inside that window, and asserts that the cached record carries the exception of the `failed` node instead of an empty one. An eighth test feeds the table a file it cannot parse (a retriable local failure), writes the terminal `failed` node from a fake foreign processor, and asserts that the cached record follows its payload. A ninth test holds a file in the middle of the ordering domain in `ordered` mode and asserts that the later files wait for it: only the files before it are processed until the foreign `processing` node disappears, then the held file and the files after it follow. A tenth test runs two processing threads over one ordering domain, parks the set-processing attempt of the smallest file at a failpoint (the file is held by a fake foreign processor), and asserts that the other thread does not ingest the later files while the outcome of the smallest file is unknown, and that no file is lost after the holder releases it. `src/Storages/ObjectStorageQueue/tests/gtest_file_status_foreign_processing.cpp` covers the shared `FileStatus` state machine, including the same-server contention case and the reset of the per-attempt data (exception, processed rows, processing end time) of a previous local attempt when the file becomes observed as processed by another processor, and the whole-record refresh on a terminal node discovered by the set-processing probe. The members added by the fix are referenced from `if constexpr (requires ...)` branches, so the gtest also compiles at the merge base (where it fails at runtime), which is what the `Bugfix validation (unit tests)` job checks. Related: https://github.com/ClickHouse/ClickHouse/pull/108977",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112313",
        "createdAt": "2026-07-28T16:32:50Z",
        "updatedAt": "2026-08-13T12:05:35Z",
        "timestamp": "2026-08-13T12:05:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112327",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix use-after-free on a sparse join key in a direct dictionary join",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109225 --> Related: https://github.com/ClickHouse/ClickHouse/pull/109225 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-after-free when a `JOIN` with `join_algorithm = 'direct'` onto a dictionary was given a join key that is stored with sparse serialization. `getColumnVectorData` returned a reference to a temporary column, so the dictionary lookup read freed memory: release builds could return wrong results and debug or sanitizer builds aborted. ### Description `getColumnVectorData` (`src/Dictionaries/DictionaryHelpers.h`) materializes its key column into a function-local `ColumnPtr` and then returns a `PaddedPODArray` reference **into that local**. It copied the data into the caller's `backup_storage` only when the input was `Const`. For a dense column that was still safe, because every conversion is a no-op returning `getPtr()` and the caller's own `ColumnPtr` keeps the buffer alive. It is not safe for a column that has to be materialized: `ColumnSparse::convertToFullColumnIfSparse` and `ColumnReplicated::convertToFullColumnIfReplicated` each allocate a **new** column that nothing else owns, so the returned reference dangles as soon as the function returns. The fix takes the copy whenever a conversion actually produced a different column (`full_column.get() != column.get()`). This is pointer identity rather than a type test, so it covers `Const`, `Sparse` and `ColumnReplicated` with one predicate, and it is fail-closed for any representation added later. The previous `Const` check is strictly subsumed: `ColumnConst::convertToFullColumn` returns either the inner column or a `replicate` result, never the `ColumnConst` itself, so `Const` behaviour is unchanged. The dense path is unaffected and adds no copy, which I verified by instrumenting both live call sites: a dense key reports zero copies, a sparse key reports one at each site. A sparse key is the case that is a use-after-free today, and it already paid for a full materialization inside `removeSpecialRepresentations`, so the extra `memcpy` of that same buffer is negligible. Reaching the bug requires a path that hands a non-materialized key to the dictionary. `dictGet`, `dictHas` and the hierarchy functions cannot: `IFunction::useDefaultImplementationForSparseColumns()` and `...ForReplicatedColumns()` both default to true and no dictionary function overrides them, so a dense, caller-owned column arrives. `IDictionary::getByKeys` (the direct join) is the reaching path, because its own `removeSpecialRepresentations` call sits inside a Nullable-only branch and a non-Nullable sparse key passes through untouched. That is also why `04627_direct_join_dictionary_nullable_key`, which does exercise a sparse key, never caught this: its key is Nullable, so it gets materialized. All 14 `getColumnVectorData` call sites are fixed by this single change. Two of them are reachable today, both in `FlatDictionary` and both on the same `getByKeys` call: `hasKeys` (the site in the reports below) and `getColumn`, reached through `getColumns`. The remaining 12 are hierarchy-only and reachable solely through the pre-converting function path. `Hashed` and `HashedArray` never reach the helper on the `getByKeys` path at all, because `DictionaryKeysExtractor` holds its converted column by value and therefore owns it. I also swept every other `convertToFullColumnIf*` / `recursiveRemove*` / `removeSpecialRepresentations` call site under `src/` for the same \"derived data escapes the owning local\" shape and found no second instance, so no sibling fix is needed. The bug dates to 2021 (`b5b624f3d7e9bf`, which introduced the conversion here) and was widened in 2025 by `2b6cb36d1dc936`, which added `ColumnReplicated` as a second carrier. The same code is present on 26.7, 26.6, 26.5, 26.4 and 26.3. Found while triaging CI on #109225 and reproduced on unmodified master. It is latent in CI only because no existing test combined a sparse-serialized left key with a direct join over a `FLAT()` dictionary; CI randomizes `ratio_of_defaults_for_sparse_serialization`, so any test that does hit this combination fails roughly 40% of the time. Reports on `31b4a2961ef4c5183f7d15dda7f755541a77a98c`: - `AddressSanitizer: heap-use-after-free`, allocated by `ColumnSparse::convertToFullColumnIfSparse`, freed at the end of `getColumnVectorData`, read by `FlatDictionary::hasKeys`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20flaky%20check%29 - `Logical error: '(n >= (static_cast<ssize_t>(pad_left_) ? -1 : 0)) && (n <= static_cast<ssize_t>(this->size()))'` from the same read: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20flaky%20check%29 The new test `04652_direct_join_dictionary_sparse_key` covers both live call sites, the aggregation shape from the report, and a key carried through `ARRAY JOIN` over a sparse base column, which is a third shape where the key has to be materialized. It pins its results against a dense table and against `join_algorithm = 'hash'` instead of hand-written constants, and asserts both that the key really is sparse and that `DirectKeyValueJoin` is still chosen, so it cannot pass vacuously. On master it aborts; with the fix it passes 50/50 with and without randomized settings. Reverting only the new predicate makes it abort again. A follow-up cleanup worth doing separately: the helper carries a `/// TODO: Remove` and would be better returning the owning `ColumnPtr` alongside the data, which removes the need for `backup_storage` entirely. That touches all 14 call sites and four dictionary classes, so it does not belong here.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112327",
        "createdAt": "2026-07-28T17:38:47Z",
        "updatedAt": "2026-08-13T16:39:38Z",
        "timestamp": "2026-08-13T16:39:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alexbakharew"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112358",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Run the merge queue's stateless flaky check in PR CI too, at comparable concurrency",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110431 Related: https://github.com/ClickHouse/ClickHouse/pull/110308 `Stateless tests (amd_binary, flaky check)` bounced a pull request from the merge queue with \"Test runs too long (> 180s)\" for `04538_final_read_in_order_limit_no_layers`, while all four flaky checks in that pull request's own CI were green on the very same test: | flaky check job | instance (vCPU) | workers | iterations | min | median | max | | |---|---|---|---|---|---|---|---| | `amd_asan_ubsan` | m7i.4xlarge (16) | **8** | 50 | 28.9 | 35.3 | 72.5 | OK | | `amd_tsan` | m7i.8xlarge (32) | 18 | 50 | 18.7 | 24.0 | 54.5 | OK | | `amd_msan` | m7i.8xlarge (32) | 11 | 50 | 12.4 | 15.9 | 49.5 | OK | | `amd_debug` | m7i.4xlarge (16) | **8** | 50 | 28.7 | 37.4 | 150.3 | OK | | `amd_binary` (merge queue) | m7i.4xlarge (16) | **18** | 50 | 33.9 | 53.4 | **184.4** | FAIL | Two reasons, both addressed here. **The release build is not slower than the sanitizer builds - it ran at 2.25x their concurrency.** The `amd_binary` job takes the \"plain binary job runs fast; allow higher concurrency\" branch, which sets `nproc` to `cpu_count * 1.2`, so the merge queue ran the test with 18 workers on a 16-vCPU runner while `amd_asan_ubsan` ran it with 8 on the identical runner type. That branch is meant for full-suite runs, where every worker picks a different test and most tests are light. In flaky-check mode every worker runs the *same* changed test, so `--jobs 18` multiplies one heavy test (`max_threads = 4`, 500k rows, `FINAL`) by 18, and the check then fails that test on wall-clock time. Exclude flaky and targeted checks from the oversubscription, so per-iteration times, and with them the `TEST_MAX_RUN_TIME_IN_SECONDS` verdict, are comparable across the flaky-check jobs. This does not weaken the gate for genuinely slow tests: the regular suite allows a test 600s, so a test that exceeds 180s only under 18-way self-contention was not failing anything else. **PR CI had no flaky check on a non-sanitizer build at all**, so that configuration first ran in the merge queue - after the merge had already started. Add `amd_binary, flaky check` to the pull request workflow, on the same build and runner as the queue's drift guard. The PR-side run is the stricter of the two (50 reruns and a 45-minute budget against the queue's 20 and 20 minutes), so anything the queue's flaky check can find is now reported in the pull request first. It self-skips on pull requests with no new or changed stateless tests, and `FUNCTIONAL_TEST_FLAKY_CHECK_JOBS` already listed the job name for the `ci-functional-test-flaky` label. Per review, the PR lane is not a separate job config: `ci/workflows/pull_request.py` spreads the very same `JobConfigs.stateless_tests_flaky_mq_jobs` that `ci/workflows/merge_queue.py` uses, so the two lanes cannot drift apart. That does not merge their praktika cache records, and it must not: `ci/praktika/cache.py` keys a record by job name and digest with no workflow in the path, so if the digests coincided the pull request's green run would satisfy the merge queue's lookup and the queue would skip the check as reused from cache - silently disabling the drift guard, whose whole point is to exercise the merge group state the pull-request run never saw. They stay separate because `Digest.calc_job_digest` hashes the job config *after* per-workflow mangling, and the `PR` workflow mangles it differently: `runs_on_label_prefix=\"pr-\"` turns `amd-medium` into `pr-amd-medium`, and the two workflows give the job different `run_after` lists. Neither field is in the digest's `drop_fields`, so the config half of the digest is `29d4` for `PR` against `fe9f` for `MergeQueueCI`. `ci/tests/test_flaky_check_pr_parity.py` pins all of it: that every merge-queue flaky check has a PR-side counterpart with the same build, runner and command; that flaky and targeted checks no longer oversubscribe while the full-suite binary jobs still do; and that the two lanes' cache digests stay distinct. CI report for the merge-queue failure above: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-110431-f734d6a982d0a985751167970f18622947b3b188&sha=b1e55fcc8253c7a995d7a79f121c8f33b26df714&name_0=MergeQueueCI&name_1=Stateless%20tests%20%28amd_binary%2C%20flaky%20check%29 ### Changelog category (leave one): - CI Fix or Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not for changelog. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1285` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112358",
        "timestamp": "2026-08-12T23:12:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112386",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keeper: don't refuse an empty-data `Set` at the memory soft limit",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/109460 **Background.** `ZooKeeper::initSession` writes empty data to `<sessions_path>/<zookeeper_name>/<server-uuid>` purely for the side effect of bumping that znode's version, and records the new version. Critical writes then attach `Check(session_path, recorded_version)` via `addCheckSessionOp`, so once another instance of the same server establishes a session and bumps the version, the older instance's writes are atomically rejected with `ZBADVERSION`. It is a fencing token against a stale process still committing data — the payload is irrelevant, which is why it is empty. **Problem.** When Keeper crosses `max_memory_usage_soft_limit`, clients can no longer re-establish their Keeper session — it fails with `Coordination error: Out of Memory` on a path under the sessions path — so tables that depend on session re-establishment stay read-only for as long as the memory event lasts, rather than as long as the write pressure lasts. In one observed incident this single write accounted for 96.7% of everything Keeper refused (1,330,721 of 1,376,368 requests over roughly two hours), and only 6 refused `Set`s were on any other path — so the limit was almost exclusively blocking recovery rather than holding back data volume. **Root cause.** The soft-limit gate refuses whatever `checkIfRequestIncreaseMem` reports as memory-increasing, and that function decides by op type: it returns true for every `Set`. But a `Set` cannot allocate a znode — the node must already exist, otherwise the request fails with `ZNONODE` and stores nothing — so a `Set` with empty data cannot increase the amount of data Keeper stores. `ZooKeeper::initSession` registers a session by writing empty data to `<sessions_path>/<zookeeper_name>/<server-uuid>`, so that write is refused and the client cannot recover from the condition that caused the refusal. The `Multi` branch has the same defect by a different route: it sums `set_req.bytesSize()`, which includes the path, the version and the xid, so an empty `Set` inside a `Multi` counts as growth proportional to its path length. **Solution.** Classify `Set` by its data, so empty data is not memory-increasing, and make the `Multi` branch sum `set_req.data.size()`. The classifier was duplicated verbatim in both dispatchers, so it moves to `KeeperCommon` where the copies cannot drift — which also makes it directly testable. `Create`, `Remove`, `SetACL` and `Auth` classification are unchanged, and writes that genuinely allocate are still refused: a `Multi` that creates an ephemeral node remains rejected, so this does not by itself return a replicated table to a writable state. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Keeper no longer refuses a `Set` request with empty data when it is over `max_memory_usage_soft_limit`. Such a request cannot increase the amount of stored data, and refusing it prevented clients from re-establishing their session, which could keep tables read-only for the whole duration of a Keeper memory event.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112386",
        "createdAt": "2026-07-29T03:13:12Z",
        "updatedAt": "2026-08-13T13:38:36Z",
        "timestamp": "2026-08-13T13:38:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "al13n321"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112469",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add a regression test for the declared type of `identity` for `Variant`",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/96844 Related: https://github.com/ClickHouse/ClickHouse/pull/110694 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/96844 Related: https://github.com/ClickHouse/ClickHouse/pull/110694 Test-only PR. It pins the declared result type of `identity` for a `Variant` argument. `identity` returns its argument column verbatim, but was still routed through the `Variant` function adaptor, which rebuilds the declared type per alternative and strips nested `LowCardinality`. The declared type then no longer matched the passed-through column, and serializing such a block over the native protocol hit a bad cast. The AST fuzzer on #96844 reported a cast to `ColumnTuple`, because its `Variant` also had a `Point` alternative. The one-line fix (`FunctionIdentityBase` opting out of the adaptor) landed on master through my #110694, which needed the same override to keep the `Geometry` custom type name, so this branch no longer changes any source file. The test is a single query asserting `toTypeName` and `variantType` of `identity(CAST('x', 'Variant(LowCardinality(String), UInt64)'))`. I verified both directions in the harness under randomized settings: it passes on master and fails on a build with `useDefaultImplementationForVariant()` removed from `FunctionIdentityBase`. `__scalarSubqueryResult` is the same `FunctionIdentityBase` object (only the JIT flag differs), so it needs no separate case. Category is CI, since the user-visible entry for this behaviour belongs to #110694. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1274` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112469",
        "timestamp": "2026-08-12T20:14:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112479",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "RabbitMQ related fix",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Security related bugfix. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1209` (included in `26.8` and later) - Backported to: `26.7.4.26`, `26.6.3.31` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112479",
        "createdAt": "2026-07-29T18:56:18Z",
        "updatedAt": "2026-08-13T12:22:33Z",
        "timestamp": "2026-08-13T12:22:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "submodule changed",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "kssenii",
        "state": "closed",
        "assignees": [
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112484",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not deserialize a skip index whose on-disk type is stale",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/112213 Related: https://github.com/ClickHouse/ClickHouse/pull/106988 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes reading a stale secondary (skip) index after an `ALTER TABLE ... MODIFY COLUMN` type change whose mutation never ran, for example because `KILL MUTATION` removed it: the granules on disk were written with the old type but decoded with the new one, raising `LOGICAL_ERROR`, requesting multi-exabyte allocations, or silently returning a wrong result. Also fixes a wrong result from an index over an expression whose meaning changes while every stored byte stays identical, so no mutation is created at all: a `MODIFY COLUMN` altering only a `DateTime` timezone, or only a custom type name such as `UInt8` to `Bool`. Such an index is now skipped for the affected part. Closes #112213. ### Description `MODIFY COLUMN` updates table metadata at once and schedules a mutation to rewrite the parts. Until it runs, a part's granules hold bytes written with the OLD type while index analysis decodes them with the NEW one. The read gate `canUseIndex` infers that mismatch from a *pending mutation entry* and fails open when there is none. Six measured ways reach the deserializer anyway, including `KILL MUTATION` having removed the entry (the report) and conversions for which no mutation is ever created, being metadata-only for the column but not for the index. The top-k minmax read had no gate at all. **The fix** asks the durable question instead: do the part's bytes match the type about to decode them? It lands in `IMergeTreeIndex::getDeserializedFormat`, which already receives the part, so one predicate covers every read path. Physical discovery splits off into a virtual `getPhysicalFormat`, leaving `getDeserializedFormat` **non-virtual** so no override bypasses it. Each required column's part-side type, taken from the part's own list rather than the type-erasing interned cache, is compared against the metadata type: only representation-preserving differences pass, plus a `getName()` same-meaning check on the equals-equal path. The walk recurses pairwise through `Array`, `Nullable` and `LowCardinality`; adding or dropping a wrapper is refused. A non-trivial expression index is refused on any difference. Over-firing costs pruning, not correctness. Two `MergeTask` text-index sites share the predicate, so a stale text index is rebuilt during a merge instead of hardlinked forward. **Out of scope:** index *identity* staleness (a changed expression, a name reused after a killed `DROP INDEX`), which no type check detects. <details><summary>Measured symptoms on master, and validation</summary> | target type | observed on master | |---|---| | `Nullable(UInt64)` | `LOGICAL_ERROR: Sizes of nested column and null map ... not equal after deserialization` | | `Nullable(UInt64)` release / `UInt64` / absent-column part | `Code: 241`, 4 / 2 / 4 EiB | | `Array(UInt64)` | `Code: 33`, \"read just 38 of 18005230136\" | | `Int8` to `Enum8` | wrong: prunes a granule the unindexed read rejects (`Code: 691`) | | `DateTime('UTC')` to `DateTime('Asia/Tokyo')`, `INDEX toHour(dt)` | wrong: 0 vs 3 | | `UInt8` to `Bool`, `INDEX toString(v)` | wrong: 0 vs 32 | | `Tuple(x UInt8)` to `Tuple(x Bool)`, `INDEX toString(p.x)` | wrong: 0 vs 32 | `04165_skip_index_stale_type_after_alter.sql`: 29 cases covering the above, whose 21 control fixtures carry 25 `explain ILIKE` granule assertions pinning the over-fire direction (a simple single-column index across the timezone and `Bool` ALTERs, an unchanged subcolumn index, an unchanged `JSON(a DateTime)` column). On the base commit the test **kills the server** with the reported stack (`SerializationNullable.cpp:178` <- `MergeTreeIndexGranuleSet::deserializeBinary` <- `MergeTreeIndexReader::read` <- `filterMarksUsingIndex`); with the fix it passes 50/50. 36 mutations each confirm one line is load-bearing. A 462-test A/B sweep against pure HEAD leaves the failure set unchanged (27 in both arms, all needing infra this sandbox lacks). Perf over 250 parts and 4 indexes is inside noise. No setting or format changes. #110050 overlaps these files and its helper asks the *physical* question, so it should forward to `getPhysicalFormat` on rebase. My #109616 has since merged and independently reached the same physical-vs-usability split, so it needs no forwarding. Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=108096&sha=25f51f1fddd05350d9c5aa1231aa0e7ee6fc676f&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29 </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112484",
        "createdAt": "2026-07-29T20:03:26Z",
        "updatedAt": "2026-08-13T07:48:24Z",
        "timestamp": "2026-08-13T07:48:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "blocker"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112498",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix segfault reading a Parquet file with an inconsistent bloom filter size",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed a crash when reading a Parquet file with inconsistent bloom filter metadata. Such files could also silently return fewer rows than they should. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ## Problem Reading a Parquet file whose bloom filter metadata is inconsistent segfaults the server. Seen in production: 25 crashes over 12 days on one instance, all with the same stack, while an hourly job read an Iceberg lake on S3. ``` (version 26.2.1.525 (official build), architecture: aarch64) Received signal 11 Signal description: Segmentation fault Address: 0xfffab740826f. Access: <not available>. Address not mapped to object. 2.0. inlined from base/base/../base/unaligned.h:12: unsigned int unalignedLoad<unsigned int>(void const*) 2. src/Processors/Formats/Impl/Parquet/Reader.cpp:906: DB::Parquet::Reader::BloomFilterLookup::findAnyHash(...) 3. src/Storages/MergeTree/KeyCondition.cpp:780: DB::mayExistOnBloomFilter(...) 4. src/Storages/MergeTree/KeyCondition.cpp:4127: DB::KeyCondition::checkInHyperrectangle(...) 5. src/Processors/Formats/Impl/Parquet/Reader.cpp:935: DB::Parquet::Reader::applyBloomAndDictionaryFilters(RowGroup&) 6. src/Processors/Formats/Impl/Parquet/ReadManager.cpp:117: DB::Parquet::ReadManager::finishRowGroupStage(...) ``` Consequences: - The server process dies, so every query on that instance fails, not just the Parquet one. The affected instance was single-replica, so each crash was a full outage until the pod restarted. - The crashing query retries and crashes again, once per retry. - When the out-of-range bloom filter block happens to land inside the read buffer instead of unmapped memory, there is no crash: the row group is pruned on unrelated bytes and rows go missing silently. `SELECT count() FROM file(...) WHERE s = '123456'` returned `0` for a value that is present. - Reproducible on `master`, and on every version since 25.11, when the v3 Parquet reader became the default (`input_format_parquet_bloom_filter_push_down` has defaulted to `1` since 25.5). ## Root cause `Reader::processBloomFilterHeader` takes the bloom filter bitset size from the file (`BloomFilterHeader.numBytes`, validated only for sign and 32-byte alignment) and derives the byte range of each 32-byte bloom filter block from it, at `bloom_filter_offset + header_size + block_idx * 32`. It never checks that `header_size + numBytes` fits inside the byte range it registered for the bloom filter — `ColumnMetaData.bloom_filter_length` when the file declares it, otherwise the \"next known offset\" upper bound computed in `initializePrefetches`. `Prefetcher::splitRange` was the only guard, and it missed the case in two independent ways: - `if (start < range.start || length > range.end - start)` **underflows**: when `start > range.end`, `range.end - start` wraps around, so the comparison is false and a subrange past the end of the range passes. - The check runs only while the parent range is still in state `HasRange`. The bloom filter header range (registered with `likely_to_be_used = true`) and the bloom filter data range start at the same file offset, and the data range is normally smaller than `min_bytes_for_seek`, so starting the header prefetch coalesces the data range into the same read task and flips it to `HasTask`. `splitRange`'s tail path then computes `req->task_offset = subranges[i].first - task->offset` with no validation at all. This is the path a real S3 read takes. `Prefetcher::getRangeData` guarded the resulting span with `chassert` only, which is compiled out in release builds, so it returned a `std::span` pointing outside `task->buf`, and `findAnyHash`'s `unalignedLoad<UInt32>` read unmapped memory. ## Fix Each of the three layers now fails closed: 1. `Reader::processBloomFilterHeader` rejects a header whose `header_size + numBytes` exceeds the bloom filter extent the file declared, with `INCORRECT_DATA` naming `input_format_parquet_bloom_filter_push_down=0` as the escape hatch — the same shape as the two bloom filter validation errors already there. The extent is remembered in the new `ColumnChunk::bloom_filter_data_bytes`, set at both `registerRange` call sites. 2. `Prefetcher::splitRange` makes the `HasRange` check underflow-safe and applies the equivalent check against the read task's byte range on the coalesced `HasTask` path, before touching refcount or `RequestState`s so throwing stays clean. 3. `Prefetcher::getRangeData` turns the buffer-bounds `chassert` into a real check against `task->length` (the invariant that holds for both the `buf` and the zero-copy `cached_region` paths), so a bookkeeping mistake anywhere surfaces as an error instead of an out-of-bounds read. Regression test: `04654_parquet_bloom_filter_bitset_out_of_bounds` reads a 1649-byte fixture whose `s` column `BloomFilterHeader` claims a 1 GiB bitset while its column metadata declares 272 bytes of bloom filter data; the read must report `INCORRECT_DATA`, and the same file still reads correctly with push-down off. The fixture is small on purpose, so its bloom filter stays under the default seek threshold and the test drives the coalesced path that production takes. Each of the three layers was also removed on its own and the test re-run, confirming none of them is dead code: without layer 1 the request is rejected by layer 2 (`Subrange out of bounds: [460964200, 460964232) not in read task [395, 1352)`), without layers 1 and 2 by layer 3, and with all three removed the reader aborts on the buffer-bounds assertion in a Debug build. Note that a file whose `bloom_filter_length` excludes the serialized header (`parquet.thrift` specifies that it includes it) is now rejected rather than read. Those reads were already either crashing or silently over-pruning; `input_format_parquet_bloom_filter_push_down=0` reads such files without their bloom filters. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1303` (included in `26.8` and later) - Backported to: `26.7.4.25`, `26.6.3.35` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112498",
        "createdAt": "2026-07-30T00:31:49Z",
        "updatedAt": "2026-08-13T16:22:14Z",
        "timestamp": "2026-08-13T16:22:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "can be tested",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "tiandiwonder",
        "state": "closed",
        "assignees": [
          "Algunenano"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112573",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix async bounded read buffer readbigat race",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix AsynchronousBoundedReadBuffer's readBigAt data race. Closes https://github.com/ClickHouse/ClickHouse/issues/109678. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1240` (included in `26.8` and later) - Backported to: `26.7.4.27` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112573",
        "createdAt": "2026-07-30T10:54:24Z",
        "updatedAt": "2026-08-13T14:31:46Z",
        "timestamp": "2026-08-13T14:31:46Z",
        "metrics": {
          "reactions": 1,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced",
          "v26.4-must-backport"
        ],
        "author": "kssenii",
        "state": "closed",
        "assignees": [
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112601",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Iterate ColumnObject subcolumns in sorted path order",
        "text": "Make `ColumnObject::forEachSubcolumn`, `forEachMutableSubcolumn` and their recursive variants iterate over the already-maintained `sorted_typed_paths`/`sorted_dynamic_paths` lists instead of the raw `typed_paths`/`dynamic_paths` `unordered_map`s. Motivation: the iteration order of an `unordered_map` is not guaranteed to be preserved by its copy constructor, and `IColumn::mutate` clones the column (copy-constructing those maps) when it is shared. `IColumn::convertToFullIfWrapped` collects the unwrapped subcolumns while iterating the source column and then reassigns them positionally while iterating the mutated (possibly cloned) column, so both passes must visit subcolumns in the same order. If the cloned maps iterated in a different order, converted subcolumns would be assigned to the wrong paths. Iterating the sorted path lists makes the visiting order deterministic and stable across cloning, removing the reliance on unspecified `unordered_map` copy-order behavior. This is currently latent — the bundled `libc++` happens to preserve copy order — so it is a safety change with no user-visible effect. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1131` (included in `26.8` and later) - Backported to: `26.7.4.17`, `26.6.3.32`, `26.5.7.40`, `26.3.18.22`, `25.8.30.14` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112601",
        "createdAt": "2026-07-30T13:55:18Z",
        "updatedAt": "2026-08-13T12:22:37Z",
        "timestamp": "2026-08-13T12:22:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-must-backport",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "Avogar",
        "state": "closed",
        "assignees": [
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112605",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Re-apply LIMIT BY on the initiator for custom-key parallel replicas",
        "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/111555 Follow-up to https://github.com/ClickHouse/ClickHouse/pull/111919, which fixed the `WITH FILL` half of #111555. This fixes the remaining `LIMIT BY` half. ### Problem Under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`), `LIMIT n BY` was applied per replica and the initiator never re-applied it, so it returned up to `n * replicas` rows per group: ```sql CREATE TABLE wf_min (g UInt16, k UInt32) ENGINE = MergeTree ORDER BY k; INSERT INTO wf_min SELECT number % 3, number FROM numbers(30); SELECT g FROM wf_min ORDER BY g LIMIT 2 BY g SETTINGS enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_range', parallel_replicas_custom_key = 'k', parallel_replicas_custom_key_range_upper = 30; -- returned 18 rows (6 per group); correct is 6 (0,0,1,1,2,2) ``` ### Root cause Custom-key parallel replicas splits rows across replicas by an arbitrary key and forces the initiator's input `from_stage` to `WithMergeableStateAfterAggregation(AndLimit)` (`PlannerJoinTree`), which tells the planner \"shards already finalized aggregation-stage processing\". The initiator therefore skips `LIMIT BY` (`Planner.cpp`, guarded by `!isFromAggregationState()`). That is correct for genuine sharding-key-aligned pushdown (each group lives on one shard), but the custom key does not align with the `LIMIT BY` key, so the per-replica `LIMIT BY` is not final. This is the same class of \"no initiator-side finalization over custom-key replica streams\" as the `WITH FILL` half fixed in #111919. ### Fix Re-apply `LIMIT BY` on the finalizing initiator (`isFinalizingStage()`) for custom-key parallel replicas, in addition to the normal `!isFromAggregationState()` case. On a replica that still emits a mergeable state, `LIMIT BY` runs as a preliminary that keeps `offset + length` rows and defers `OFFSET` to the initiator, so `LIMIT n OFFSET m BY` stays correct too. Custom-key parallel replicas is detected via a new `JoinTreeQueryPlan::is_parallel_replicas_custom_key` flag set only where the custom-key path is actually built — the replica custom-key filter, and the `Distributed` / MergeTree initiator dispatch in `PlannerJoinTree` — rather than the ambient `canUseParallelReplicasCustomKey()` setting (which is true whenever the profile enables a `custom_key_*` mode, even when the query used genuine sharding-key pushdown and never built a custom-key plan). This keeps `optimize_distributed_group_by_sharding_key` pushdown (e.g. `01244_optimize_distributed_group_by_sharding_key`) unaffected. The change is analyzer-only; the deprecated legacy interpreter's custom-key path is left unchanged (its custom-key handling is inconsistent across modes and a blanket re-application there regressed `custom_key_sampling` + `OFFSET`). ### Verification Built ClickHouse from an earlier revision of this branch and checked against a running 3-replica custom-key cluster: - `LIMIT 2 BY g` returns 6 rows (was 18) for both `custom_key_range` and `custom_key_sampling`. - `LIMIT 1 OFFSET 1 BY g` returns 3 rows (`0,1,2`), matching the non-distributed result — `OFFSET` applied once. - No regression: normal distributed `LIMIT 2 BY g` over `remote(...)` still returns the correct 6 rows. (The later review fix — switching from the ambient setting to the plan-tied flag — was validated by review/CI; the sandbox VM was recycled so it was not re-run locally. CI runs both `01244_optimize_distributed_group_by_sharding_key` and the new `04657_parallel_replicas_custom_key_limit_by`.) ### Test Added `tests/queries/0_stateless/04657_parallel_replicas_custom_key_limit_by.sql` (tagged `no-old-analyzer`), covering both custom-key modes and the `OFFSET` case. Confirmed it returns 18 (buggy) on unpatched and 6 on patched. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `LIMIT n BY` returning up to `n * replicas` rows per group under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`). `LIMIT BY` is now re-applied on the initiator over the merged replica streams. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112605",
        "createdAt": "2026-07-30T14:04:23Z",
        "updatedAt": "2026-08-13T17:39:00Z",
        "timestamp": "2026-08-13T17:39:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yakov-olkhovskiy",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112648",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix Iceberg Avro writer emitting optional complex fields as required",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/111775 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes an Iceberg table whose write format is `Avro` serializing a field declared `\"required\": false` whose type is a `list`, `map` or `struct` as required, with no `[\"null\", T]` union, so that the data file's own schema disagreed with the table metadata and other Iceberg readers saw a required field. Such a field is now written as the `[\"null\", T]` union the Iceberg spec uses for an optional one. ### Description `AvroSerializer::createSchemaWithSerializeFn` emits a `[\"null\", T]` union only for `Nullable` and `Variant`, so Avro nullability came purely from the type carrying a `Nullable` wrapper. The Iceberg read layer never puts that wrapper on a container: `IcebergSchemaProcessor::getFieldType` calls `makeNullable` only on the scalar branch, while the complex branch returns a bare `Array`/`Map`/`Tuple` regardless of the field's own `required` bit. An optional container's optionality is therefore unrecoverable from the type. This is the Avro half of #111775, which fixed the same defect for ORC and added the per-path metadata this consumes. The fix adds `createFieldSchemaWithSerializeFn`, used at the four sites that own a field (top-level column, tuple field, array element, map value). It consults that metadata and wraps the built schema in `[\"null\", T]`, unless the schema is already a union or null, so the union is added exactly once per path. `getIcebergType` publishes `required: true` for ClickHouse-authored containers, so the carrier is normally an externally authored schema written by ClickHouse with `'Avro'`. The one exception is `Nullable(Tuple)`, the only container ClickHouse DDL can declare optional. Field ids are unchanged and plain `FORMAT Avro` output is byte-identical. At the default `input_format_null_as_default = 1` such a file reads back unchanged. With the setting off the read throws `Cannot insert Avro Union(Null, T) into non-nullable type T`, because `getFieldType` derives a bare container for an optional complex field, the standing reader limitation Spark-written files already hit. For a `Nullable(Tuple)` column it is a change: master wrote a bare record, readable at either value, so with `enable_nullable_tuple_type = 1` (default off) and the setting persisted off at `CREATE`, a read that worked now errors. Optional element, value and field positions are unaffected, their targets being `Nullable`. The reader derivation is the root cause and is tracked separately. <details> <summary>Validation: per-field nullability in the written file's schema, master vs fix</summary> Fixture: hand-written `v1.metadata.json` with optional and required containers, `CREATE TABLE IF NOT EXISTS ... ENGINE = IcebergLocal(dir, 'Avro')`, one INSERT. Values read out of the written file's `avro.schema` header entry. | field (Iceberg) | master | with fix | |---|---|---| | `req_int` required int | required int | required int | | `opt_int` optional int | optional int | optional int | | **`opt_list` optional list** | **required array** | **optional array** | | **`opt_map` optional map** | **required map** | **optional map** | | **`opt_struct` optional struct** | **required record** | **optional record** | | `req_list` / `req_map` / `req_struct` required | required | required | | `list_opt_elem.element`, `map_opt_val.value`, `opt_struct.sy` optional | optional | optional | | **`list_opt_struct_elem.element` optional struct** | **required record** | **optional record** | | **`struct_opt_list_field.inner` optional list** | **required array** | **optional array** | | **`map_opt_struct_val.value` optional struct** | **required record** | **optional record** | Each of the four call sites is pinned separately: reverting one at a time to `createSchemaWithSerializeFn` moves a disjoint set of reference lines, the top-level site moving `opt_list`/`opt_map`/`opt_struct`, and the array-element, tuple-field and map-value sites moving one nested line each. All 13 Iceberg field ids are byte-identical between the arms, and the test also pins them at depth: not descending through the union drops four nested-id lines. 50/50 green, and 20/20 green under `compatibility='20.1'`, the regime whose pre-21.1 `input_format_null_as_default` default the read-back pin defends. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.487` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112648",
        "createdAt": "2026-07-30T18:47:15Z",
        "updatedAt": "2026-08-13T16:51:19Z",
        "timestamp": "2026-08-13T16:51:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112650",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject WITH FILL bounds that do not fit the ORDER BY column type",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/109454 Related: https://github.com/ClickHouse/ClickHouse/issues/109216 --> `WITH FILL FROM`/`TO` values are converted to a type wide enough for the arithmetic - `Int64` for every integer column type - while the generated values are written into a column of the column's own type, which truncates whatever does not fit its range. A truncated value wraps around, so the filled stream stops being monotonic while the query plan keeps claiming that it is still sorted by the fill columns. `DISTINCT` in order relies on that claim and reads the stream as a sequence of sorted runs, so it deduplicates within wrong ranges: ```sql SELECT count() FROM (SELECT DISTINCT x, s FROM (SELECT toUInt8(5) AS x, 'Hello' AS s ORDER BY x ASC WITH FILL FROM 1 TO 1025)); ``` returns `1024` instead of `257` in a release build, because `WITH FILL FROM 1 TO 1025` over a `UInt8` column generates `1..255, 0, 1..255, 0, ...`. In a debug or sanitizer build the same query aborts in `DistinctSortedStreamTransform` with `Equal values are not contiguous within the range assumed to be sorted`, which is what the AST fuzzer hit: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109454&sha=b0030d87c1f4e4b31d2766c17286e42492a2da94&name_0=PR&name_1=AST%20fuzzer%20%28amd_msan%29 https://github.com/ClickHouse/ClickHouse/pull/109454 The fuzzer query is unrelated to that pull request; the abort reproduces on `master`: ```sql SELECT DISTINCT x, isZeroOrNull(materialize(true)), s FROM ( SELECT 5 AS x, 'Hello' AS s ORDER BY x ASC NULLS LAST WITH FILL FROM 1 TO 10 INTERPOLATE (`s` AS concat(s, 'A')) LIMIT 1048576 UNION ALL SELECT 5 AS x, 'Hello' AS s ORDER BY x ASC NULLS LAST WITH FILL FROM 1 TO 1025 INTERPOLATE (`s` AS concatAssumeInjective(s, 'A')) LIMIT 1048576 ) ORDER BY s ASC; ``` ### Fix Reject `FROM`/`TO` values that cannot be represented in the type of the `ORDER BY` column with `INVALID_WITH_FILL_EXPRESSION`, next to the existing check that rejects negative bounds for an unsigned column type. This matches how the equivalent `Date`/`DateTime` bound mismatch is already rejected (https://github.com/ClickHouse/ClickHouse/issues/30421). `TO` is an exclusive bound, so it may still be outside of the range as long as the values that filling actually generates fit: `WITH FILL FROM 0 TO 256` over a `UInt8` column keeps generating `0..255`, and so does `WITH FILL FROM 0 TO 257 STEP 3`, which stops at `255`. The last generated value is known up front only when `FROM` and a plain numeric `STEP` are given. Without `FROM` the sequence is anchored at a data value, so which values are generated is known only at execution time: `WITH FILL TO 257 STEP 3` over a `UInt8` column stops at `254` when the data ends at `11` but reaches `256` when it ends at `13`. Still, whenever filling generates anything at all, its last value lands within one step before `TO`, so such a bound is rejected only when even the value one whole step away from `TO` (clamped towards zero, so that anchors close enough to `TO` to generate nothing keep being accepted) does not fit the column type - that is, only when filling provably wraps for every possible anchor. An `INTERVAL` step performs its calendar arithmetic in the column's own native type, so unlike a plain numeric step it wraps around within the column domain and can never reach a `TO` outside of it: filling would generate wrapped-around values forever (it does on current master, e.g. `SELECT toDate(0) AS d ORDER BY d ASC WITH FILL FROM toDate(0) TO 70000 STEP INTERVAL 100 YEAR` never terminates). Such a `TO` is therefore rejected regardless of `FROM` - but only when it is out of range in the fill direction. A `TO` out of range against the fill direction (e.g. `ORDER BY d DESC WITH FILL TO 70000 STEP INTERVAL -1 YEAR` over a `Date` column) is a guaranteed no-op instead: every possible anchor is already past it, filling never takes a single step, and the bound is accepted - the same clamp towards zero as in the numeric case. `STALENESS`, which is allowed only without `FROM`, replaces `TO` as the effective bound whenever it comes first, so the sequence can stop arbitrarily far below `TO` and such a `TO` is not checked either. Float and Decimal bounds are not checked: they saturate instead of wrapping, and a `Float64` bound is generally inexact in a `Float32` column, so requiring exact representability there would reject ordinary queries. On top of the storage range, the calendar arithmetic of an `INTERVAL` step clamps at the boundaries of the representable calendar - the `[0000-01-01, 9999-12-31]` window of `DateLUTImpl`, taken in the local civil calendar of the column's time zone, so for `DateTime64` the boundary in raw ticks shifts by the UTC offset (the last reachable second of a `DateTime64(0, 'Etc/GMT-14')` column is `253402250399`, not the UTC `253402300799`) - and for `Date32` and `DateTime64` that window is strictly narrower than the storage type. A `TO` beyond the calendar boundary in the fill direction fits the storage type but can never be reached: the filling keeps generating the clamped boundary value forever (it does on current master, e.g. `SELECT toDate32('9999-12-31') AS d ORDER BY d ASC WITH FILL TO 3000000 STEP INTERVAL 1 YEAR` never terminates). Such a `TO` is rejected against the calendar limits, scale-aware for `DateTime64`. For `Date` the calendar clamp coincides with the `UInt16` storage boundary, and `DateTime` wraps within `UInt32`, where any in-range `TO` stays reachable from some anchor, so the storage check covers those two. Beyond reachability, for `Date32` and `DateTime64` the values between the calendar boundary and the boundary of the storage type are invalid in themselves: no conversion produces them (they all clamp at the calendar boundary), yet a `FROM` bound in that gap is materialized into the column as is and serialized as the clamped boundary date - a spurious duplicate of the genuine boundary value next to it (e.g. `SELECT d FROM (SELECT toDate32('2000-01-01') AS d ORDER BY d ASC WITH FILL FROM -719529 TO -719528 STEP INTERVAL 1 YEAR)` on current master returns a `0000-01-01` row holding the day number `-719529`, which is not `0000-01-01`). The representability check therefore tests bounds of these two types against the calendar window (in the local civil calendar of the column's time zone for `DateTime64`), not just the storage range: an out-of-calendar `FROM` is rejected for any kind of step - like a `FROM` out of the storage range already is - and a numeric-step `TO` is rejected when the last value generated under it provably lands in the gap, under the same rules as the storage range. The last-generated-value computation covers the `Decimal64`-carried `DateTime64` bounds as well as the `Int64`-carried types: a numeric step over `DateTime64` advances raw ticks of the column's scale, and ticks beyond the calendar do not wrap but are equally invalid (they all serialize as the clamped boundary date). Finally, the calendar clamp makes an `INTERVAL` step able to stagnate: adding the interval to a value whose result would leave the calendar returns the value unchanged, so the sequence can stop advancing strictly below a perfectly representable `TO` and never terminate (e.g. `WITH FILL FROM toDateTime64('9999-06-01 00:00:00', 0, 'UTC') TO toDateTime64('9999-12-31 00:00:00', 0, 'UTC') STEP INTERVAL 1 YEAR` hangs on current master). With an explicit `FROM` the whole sequence is known up front - `FillingRow::next` advances it by one application of the step function at a time - so it is walked at construction time with the same step function, under a bounded budget (65536 steps), and rejected when it provably stagnates before reaching `TO`. The walk mirrors the runtime step exactly, so it can never misjudge a terminating sequence; a fill whose stagnation lies beyond the budget (a fine-grained step over a huge span) is accepted as before. Note that a query that previously returned wrapped-around values now gets an error instead. There is no in-tree test with such bounds. ### Not fixed here `WITH FILL` has other routes to the same wraparound, all pre-existing and all producing garbage values rather than tripping the sortedness assertion in the shapes I could build. They are data-dependent, so they need a different fix: ```sql -- the INTERVAL step function itself wraps in the column type SELECT groupArray(d) FROM (SELECT toDate('2149-06-01') AS d ORDER BY d ASC WITH FILL FROM toDate('2149-06-01') TO toDate('2149-06-06') STEP INTERVAL 1 YEAR); -- ['2149-06-01','1970-12-26','1971-12-26',...] -- the STALENESS border is computed from a data value and leaves the range SELECT groupArray(x) FROM (SELECT toUInt8(250) AS x ORDER BY x ASC WITH FILL STALENESS 20); -- [250,251,252,253,254,255,0,1,...,13] -- an out-of-range TO without FROM is rejected only when it wraps for every possible anchor; -- when only some anchors wrap, the wrapping ones still do so at execution time SELECT groupArray(x) FROM (SELECT toUInt8(13) AS x ORDER BY x ASC WITH FILL TO 257 STEP 3); -- [13,16,...,253,0] -- the INTERVAL step function returns its input unchanged when the result would leave the representable -- calendar, so an anchor near the boundary stagnates below an in-range TO and the filling never terminates; -- without FROM, which anchors stagnate depends on the data and the step (with an explicit FROM this shape -- is now rejected up front, unless the stagnation lies beyond the bounded walk budget of 65536 steps) SELECT * FROM (SELECT toDateTime64('9999-06-01 00:00:00', 0, 'UTC') AS t ORDER BY t ASC WITH FILL TO toDateTime64('9999-12-31 00:00:00', 0, 'UTC') STEP INTERVAL 1 YEAR); -- hangs: 9999-06-01 + 1 year would be out of range, so the step keeps returning 9999-06-01 -- the same shape over Date32, reachable since https://github.com/ClickHouse/ClickHouse/pull/111534 extended -- the parsed range of Date32 to the whole calendar (before that, such an anchor clamped to 2299-12-31) SELECT * FROM (SELECT toDate32('9999-06-01') AS d ORDER BY d ASC WITH FILL TO 2932896 STEP INTERVAL 1 YEAR); -- hangs the same way; with `FROM toDate32('9999-06-01')` added it is now rejected up front -- the DateTime INTERVAL step arithmetic wraps within UInt32, so a TO near the top of the storage range can be -- unreachable from a given anchor while staying reachable from others; the wrapped sequence cycles instead of -- stagnating, which the construction-time walk does not detect even with an explicit FROM SELECT * FROM (SELECT toDateTime('2106-01-01 00:00:00', 'UTC') AS t ORDER BY t ASC WITH FILL TO 4294967295 STEP INTERVAL 100 YEAR); -- hangs: 2106 + 100 years wraps to 2069 and cycles below TO forever ``` ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong `DISTINCT` results (and a logical error `Equal values are not contiguous within the range assumed to be sorted` in debug builds) for `ORDER BY ... WITH FILL FROM/TO` when a bound does not fit the type of the `ORDER BY` column: the generated values were silently truncated and wrapped around, so the filled stream was no longer sorted. An out-of-range `FROM` is now rejected with `INVALID_WITH_FILL_EXPRESSION`, and an out-of-range `TO` is rejected whenever the values that filling generates provably wrap the column type for every possible starting point - including always under an `INTERVAL` step for a `TO` out of range in the fill direction, whose calendar arithmetic stays within the column domain and can never reach such a `TO`, making the filling non-terminating (a `TO` out of range against the fill direction is a guaranteed no-op and stays accepted). For `Date32` and `DateTime64`, the bounds are additionally checked against the representable calendar (`[0000-01-01, 9999-12-31]` in the column's time zone), which is narrower than the storage type: values in between are invalid - everything else clamps at the calendar boundary, and filling materialized them as is, serialized as a spurious duplicate of the boundary date - so an out-of-calendar `FROM` is rejected for any kind of step, an `INTERVAL`-step `TO` is rejected when it is beyond the calendar in the fill direction (the clamping calendar arithmetic can never reach it), and a numeric-step `TO` over `Date32` or `DateTime64` is rejected when the last generated value provably lands out of the calendar. An `INTERVAL`-step fill with an explicit `FROM` is additionally rejected when the sequence provably stagnates before reaching `TO` (the calendar clamp makes the step return its input unchanged near the boundary, so the filling would never terminate). Data-dependent wraparound or stagnation (e.g. via `STALENESS`, numeric-step starting points that wrap only at execution time, or calendar-boundary anchors without `FROM` that make the `INTERVAL` step stagnate below an in-range `TO`) is not covered by this check.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112650",
        "createdAt": "2026-07-30T19:13:54Z",
        "updatedAt": "2026-08-13T11:58:37Z",
        "timestamp": "2026-08-13T11:58:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112667",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Silk integration",
        "text": "Splits the silk runtime integration out of https://github.com/ClickHouse/ClickHouse/pull/111275, so that it can be reviewed on its own. This adds the plumbing that lets ClickHouse run work on [silk](https://github.com/ClickHouse/silk) fibers, without yet putting any subsystem on them. - `Silk::initializeFiberScheduler` / `Silk::destroyFiberScheduler`, called by the server when the `enable_silk_runtime` server setting is enabled. The fiber stack size is configurable through the `silk.fiber_stack_size` configuration key (320 KiB by default, which leaves enough room for OpenSSL handshakes). - `FiberLocal` - fiber-local storage. A fiber can migrate between operating-system threads, so it must not observe another fiber's `thread_local` state; the values of the registered slots are swapped in and out on every fiber switch instead. `current_thread` (`ThreadStatus`), the OpenTelemetry tracing context, and the memory-tracker and exception blockers are moved to it. - The silk thread-local-storage sanitizer: an LLVM pass in `utils/silk-thread-local-storage-sanitizer` that instruments every `thread_local` access and aborts when a fiber touches raw thread-local storage. Without it, a variable that was not migrated to `FiberLocal` produces silent corruption rather than a diagnostic. It is enabled in the debug and ASan CI builds. - `Silk::ConnectionPool` and `Silk::streamSocketFactory` - a `Connection` pool and a socket factory that suspend the calling fiber instead of blocking the operating-system thread. `PoolBase` and `ConnectionPool` are templated on the lock and the condition variable to make that possible, and `ConnectionPool` stays an alias of the `std::mutex` instantiation, so the existing call sites are unchanged. - Memory that the runtime maps outside the C++ heap - fiber stacks and `io_uring` rings - is charged to `total_memory_tracker` through silk's mmap accounting hooks. - The low-level silk runtime counters are exported to `system.asynchronous_metrics` under a `Silk` prefix. - `Common/Fiber.h` and `Common/FiberStack.h` are renamed to `Common/StackfulCoroutine.h` and `Common/CoroutineStack.h`. They implement the boost-context coroutines used by `AsyncTaskExecutor`, which are unrelated to silk fibers, and having two different things called \"fiber\" in the same codebase is confusing. Related: https://github.com/ClickHouse/ClickHouse/pull/111275 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental support for the [silk](https://github.com/ClickHouse/silk) fiber runtime, enabled with the `enable_silk_runtime` server setting. When it is enabled, the server initializes the silk fiber scheduler at startup, so that subsystems supporting it can run their jobs on fibers instead of occupying an operating-system thread while waiting for I/O.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112667",
        "createdAt": "2026-07-30T21:32:29Z",
        "updatedAt": "2026-08-13T17:55:03Z",
        "timestamp": "2026-08-13T17:55:03Z",
        "metrics": {
          "reactions": 1,
          "comments": 1
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "mstetsyuk",
        "state": "open",
        "assignees": [
          "CheSema"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112679",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallelize delete manifests reads",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description] Queries on Iceberg tables with many delete files now start faster. ClickHouse reads and decodes the delete manifest files concurrently instead of one at a time, so their storage reads overlap. The new setting `iceberg_delete_manifest_decode_concurrency` (default `4`) controls how many are decoded at the same time.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112679",
        "createdAt": "2026-07-30T23:03:10Z",
        "updatedAt": "2026-08-13T16:18:26Z",
        "timestamp": "2026-08-13T16:18:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-performance",
          "submodule changed"
        ],
        "author": "asya-ch",
        "state": "open",
        "assignees": [
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112688",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Push subcolumn reads into subqueries",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/75538 Related: https://github.com/ClickHouse/ClickHouse/issues/92455 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Read only the requested subcolumns of columns exported by subqueries, CTEs and views instead of the whole columns. For example, `SELECT data.a FROM (SELECT * FROM table)` with a `JSON` column `data` now reads only the subcolumn `data.a` from the table. Controlled by the new setting `optimize_push_subcolumns_into_subqueries` (enabled by default). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) --- When a subcolumn of a column exported by a subquery is requested, it is resolved into the `getSubcolumn` function over the whole column: `SELECT data.a FROM (SELECT * FROM test)` reads the whole `data` column from the table and extracts the subcolumn afterwards. The new query tree pass `PushSubcolumnsIntoSubqueries` adds the subcolumn to the subquery projection and replaces the `getSubcolumn` function with a reference to it. If the whole column is not used anywhere else, it is then removed from the subquery projection by the subsequent `RemoveUnusedProjectionColumns` pass, so only the subcolumn is read from the table. The pushdown works through several levels of subqueries, CTEs and views, and also applies to `Tuple`, `Nullable`, `Array` and other types with subcolumns. `UNION ALL` subqueries are supported: the subcolumn is added to every branch at the same position, all-or-nothing, and only when the branch types match exactly (the subcolumn of the least supertype is not guaranteed to be the supertype of the branch subcolumns). The `DISTINCT`, `INTERSECT` and `EXCEPT` modes deduplicate or match rows over all projection columns, so such subqueries are not rewritten (nor are recursive CTEs). The pushdown is skipped when it could change the query result: - when the subquery uses `DISTINCT`, `GROUP BY`, aggregate functions, or `ORDER BY ... WITH FILL` (an added projection column would change the result or would not be valid); - when the subquery is on a side of a `JOIN` that can be filled with default values for non-matched rows (`getSubcolumn` of a default value is not always equal to the default value of the subcolumn type, e.g. for the `null` subcolumn of `Nullable` columns); a shared subquery node (an ordinary CTE referenced several times) found in any such position is ineligible for all of its occurrences; - in the outer query with aggregation, an occurrence of `getSubcolumn` is only replaced when it is evaluated before the aggregation step (`WHERE`, `JOIN ON`, arguments of aggregate functions) or when the whole expression is an aggregation key; - when the column types diverge, e.g. under `join_use_nulls` or `group_by_use_nulls`; - when the whole column is also read in the outer query, including uses by correlated subqueries and uses under a different exported name of the same physical column (`SELECT tup AS x, tup FROM t`, or a trivial `ALIAS` storage column next to its base column): the subquery would then read both the whole column and the subcolumn from the table, while extracting the subcolumn from the already read column is cheaper. For a `UNION ALL` target the alias equivalence of exported names is detected in every branch (the pushdown is applied to all branches or to none), so a single branch exporting the same physical column under another name that stays alive blocks the rewrite. The decision counts only sibling subcolumns that are validated as actually pushable into the target (in a side-effect-free dry run): when two subcolumns of the same exported column are requested and only one of them can be pushed (e.g. the other is shadowed by a same-named storage column inside the subquery), nothing is pushed, since the unpushable sibling keeps the whole column alive. - for references of a reused `MATERIALIZED` CTE (`enable_materialized_cte`): the temporary table serves all references of the CTE, so pruning the parent column there would require proving that no reference needs the whole column. Single-use materialized CTEs are inlined by the analyzer and take the ordinary subquery path (subcolumn reads over materialized CTEs are currently broken on master independently of this change, see https://github.com/ClickHouse/ClickHouse/issues/113623). Subqueries spelled with an alias list, e.g. `SELECT x.a FROM (SELECT json FROM t) AS s(x)`, are supported. Trivial `ALIAS` columns (an `ALIAS` whose body is just another column of the same table, possibly chained) exported by a subquery are followed down to the underlying storage column. A subquery can also export a derived subcolumn (e.g. `SELECT json.a AS x FROM ...` over a deeper subquery keeps a `getSubcolumn(json, 'a')` projection expression); reading a subcolumn of such an export composes the paths (`a` + `b` -> `a.b`), so the pushdown continues through derived exports down to the base table. Conversely, when no name of such an alias-equivalent class of exports remains referenced in the outer queries, the never-referenced sibling exports are dead together with the replaced ones (all of them are removed by `RemoveUnusedProjectionColumns`), so they do not block the pushdown through the deeper levels of subqueries; the alias-equivalent classes are taken from every branch of a `UNION ALL` target.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112688",
        "createdAt": "2026-07-30T23:44:26Z",
        "updatedAt": "2026-08-13T05:56:00Z",
        "timestamp": "2026-08-13T05:56:00Z",
        "metrics": {
          "reactions": 1,
          "comments": 15
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112705",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Improve canceling queries in nested expression functions in FilterTransform",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/103705 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve canceling queries by `KILL QUERY` and Ctrl+C in `clickhouse-client` while a long-running function is evaluated in a filter expression: `FilterTransform` (`WHERE`, including the `FilterSortedStreamByRange` path) and `TotalsHavingTransform` (`HAVING` with `WITH TOTALS`) now forward cancellation into nested expression functions. This supersedes https://github.com/ClickHouse/ClickHouse/pull/103705 by @rvasin (Roman Vasin), whose commits are preserved here — all credit for the feature goes to them. The original branch is in a fork with maintainer edits disabled, so the merge conflicts with `master` and the remaining review feedback could not be pushed there. On top of the original PR, this PR: - Merges `master` and resolves the conflicts (`FailPoint.cpp` failpoint list, `FilterTransform.cpp` includes). - Addresses the two unresolved review threads (the AI review \"Request changes\" verdict): - `FilterSortedStreamByRange` holds a private `FilterTransform` and calls its `transform` directly, but the pipeline cancels only the outer processor. Now `FilterSortedStreamByRange::onCancel` forwards the cancellation into the inner transform, so `cancelExecution` reaches functions of range-filter predicates from `PartsSplitter`. - `TotalsHavingTransform` evaluated the `HAVING` expression without a cancellation callback, so a long-running function in `HAVING` remained uninterruptible. It now mirrors `FilterTransform`: `onCancel` cancels function execution, `expression->execute` gets a `check_cancelled` callback, and the transform returns before the totals/filter bookkeeping when cancelled. - Adds a test `04658_kill_query_having_totals_pause` with a new `totals_having_transform_pause` failpoint, mirroring `04612_kill_query_filter_pause`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112705",
        "createdAt": "2026-07-31T05:28:48Z",
        "updatedAt": "2026-08-13T13:40:43Z",
        "timestamp": "2026-08-13T13:40:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 32
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112717",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Rebuild skip indices, projections, TTL and MATERIALIZED columns left stale by ALTER MODIFY / UPDATE / CLEAR COLUMN",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/112302 Related: https://github.com/ClickHouse/ClickHouse/pull/85985 --> Related: https://github.com/ClickHouse/ClickHouse/pull/112302 Related: https://github.com/ClickHouse/ClickHouse/pull/85985 A mutation decides what to rebuild by comparing the mutated column against the columns that an index / projection / TTL expression requires. When the expression reads a **subcolumn** (`t.a`, `json.a`), those required columns carry the subcolumn name while the mutation tracks the whole stored column (`t`), so the comparison never matched. On **wide parts** the dependent files (`skp_idx_*`, projection parts, TTL info) were then hardlinked into the mutated part, leaving them describing data that no longer exists. Compact parts are unaffected, because a mutation rewrites the whole part. Every such site now resolves an expression's required columns to their top-level columns through one shared helper, `getRequiredColumnsWithSubcolumnsReplaced`, which rewrites subcolumns to `getSubcolumn` — the same rewrite used for `MATERIALIZED` expressions in #85985 — and re-analyzes the expression. The rewrite is applied where the value is recomputed as well, so it is derived from the *updated* parent column instead of a subcolumn pre-extracted from the source part. What is fixed: - **Skip index on a subcolumn** — rebuilt on `ALTER MODIFY COLUMN`, `ALTER UPDATE`, `ALTER CLEAR COLUMN` and patch updates of the parent. The recompute in the partial-rewrite path (`MutateSomePartColumns`) now extracts the subcolumn from the updated parent instead of failing with `NOT_FOUND_COLUMN_IN_BLOCK` or reading stale data. - **Projection** — rebuilt when the altered column feeds its sort key (a subcolumn sort key was missed) or its `WHERE` clause. Both decide what the projection stores: the sort key is persisted as raw bytes in the projection's own primary index, and the `WHERE` fixes the stored rows and aggregate states. A column used only in the projection's `SELECT` needs no rebuild, because that payload is converted on read. - **`alter_column_secondary_index_mode`** — the `throw`/`compatibility` guard in `checkAlterIsPossible` now also fires for an explicit index defined on a subcolumn of the altered column, instead of accepting the ALTER and then failing (or silently dropping the index) during the mutation. - **TTL on a subcolumn** — a `TTL t.a` is recalculated when the parent column is mutated, so rows and columns whose TTL moved into the past actually expire. `SHOW CREATE TABLE` still shows the original expression. - **`MATERIALIZED` column** — recalculated when `ALTER MODIFY COLUMN` changes the type of a column it reads (whole column or subcolumn), including chains of `MATERIALIZED` columns, and everything depending on them is rebuilt too. Previously only `ALTER UPDATE` recalculated them, so a value-changing conversion left them holding values computed from data the mutation had just rewritten. A `MATERIALIZED` column that is part of a key cannot be recalculated in an existing part — its sort order and partition id are fixed when the part is written, which is also why `MATERIALIZE COLUMN` refuses key columns — so such an ALTER is now rejected unless the conversion preserves values. Reproductions, all silent wrong results before this PR: ```sql -- skip index on a subcolumn CREATE TABLE t (id UInt32, t Tuple(a Int32, b String), INDEX idx t.a TYPE minmax GRANULARITY 1) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 4, min_bytes_for_wide_part = 0, min_rows_for_wide_part = 0; INSERT INTO t SELECT number, (number, 'x') FROM numbers(16); ALTER TABLE t MODIFY COLUMN t Tuple(a Float32, b String); SELECT count() FROM t WHERE t.a >= 1; -- 0 before, 15 now SELECT count() FROM t WHERE t.a >= 1 SETTINGS use_skip_indexes=0; -- 15, ground truth -- projection filtered on a column whose type changes CREATE TABLE t2 (id UInt64, x Int64, PROJECTION p (SELECT sum(id) WHERE x < 0)) ENGINE = MergeTree ORDER BY id SETTINGS min_bytes_for_wide_part = 0; INSERT INTO t2 SELECT number, toInt64(3000000000) + number FROM numbers(100); ALTER TABLE t2 MODIFY COLUMN x Int32; -- every value wraps negative, so WHERE x < 0 now matches all rows SELECT sum(id) FROM t2 WHERE x < 0; -- 0 before, 4950 now -- MATERIALIZED column computed from a column whose type changes CREATE TABLE t3 (x Int64, m Int64 MATERIALIZED x) ENGINE = MergeTree ORDER BY tuple() SETTINGS min_bytes_for_wide_part = 0; INSERT INTO t3 VALUES (5000000000); ALTER TABLE t3 MODIFY COLUMN x Int32; SELECT x, m FROM t3; -- 705032704, 5000000000 before; 705032704, 705032704 now ``` Tests: `04617_stale_subcolumn_skip_index_after_mutation` and `04840_recalculate_materialized_column_on_source_type_change`. Known gaps left for follow-up pull requests: a subcolumn in a projection `WHERE` is accepted at `CREATE` but fails at `INSERT` (`Not found column t.a in block`), and the aggregation assignments of a GROUP BY TTL are not covered by the rewrite. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong query results caused by a mutation leaving data derived from the mutated column stale on wide parts. A skip index or a TTL expression defined on a subcolumn (for example a `Tuple`, `Nested`, `Map` or `JSON` element) is now rebuilt or recalculated when `ALTER MODIFY COLUMN`, `ALTER UPDATE` or `ALTER CLEAR COLUMN` changes the parent column; a projection is rebuilt when the altered column feeds its sort key or its `WHERE` clause; and a `MATERIALIZED` column is recalculated when `ALTER MODIFY COLUMN` changes the type of a column it is computed from. An `ALTER MODIFY COLUMN` that would change the values of a `MATERIALIZED` column used in the sorting or partition key is now rejected, because such a column cannot be recalculated in an existing part.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112717",
        "createdAt": "2026-07-31T09:40:47Z",
        "updatedAt": "2026-08-13T17:20:11Z",
        "timestamp": "2026-08-13T17:20:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "Avogar",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112758",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Sync private settings",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112758",
        "createdAt": "2026-07-31T14:55:06Z",
        "updatedAt": "2026-08-13T16:37:02Z",
        "timestamp": "2026-08-13T16:37:02Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "scanhex12",
        "state": "open",
        "assignees": [
          "alesapin"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112788",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allow COMMENT after all the other column modifiers",
        "text": "The parser accepted the modifiers of a column declaration in one fixed order only, so `COMMENT` had to be written before `CODEC`, `STATISTICS`, `TTL` and per-column `SETTINGS`: ```sql CREATE TABLE t (x UInt64 CODEC(ZSTD) COMMENT 'text') ENGINE = Memory; -- Code: 62. DB::Exception: Syntax error: failed at position 38 (COMMENT). -- Expected one of: STATISTICS, TTL, PRIMARY KEY, SETTINGS, ... ``` while the same declaration with the two modifiers swapped was accepted. There is no ambiguity here, only an accident of how the parser was written, and it is annoying: nothing in the syntax hints that the comment has to go first, and the same is true for `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN`. Now the modifiers that follow the type - `COMMENT`, `CODEC`, `STATISTICS`, `TTL`, `COLLATE`, `PRIMARY KEY` and `SETTINGS` - are parsed in a loop, so they can be written in any order, and each of them at most once. In `CREATE TABLE`, everything that was accepted before is still accepted, and `ASTColumnDeclaration::formatImpl` keeps printing the modifiers in the canonical order, so `SHOW CREATE TABLE` output does not change. The `ALTER` surface is made consistent instead of merely permissive. `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN` applies the declared modifiers - `COMMENT`, `CODEC`, `STATISTICS`, `TTL` and per-column `SETTINGS` - and these can now be written in any order. Per-column `SETTINGS` in `ADD COLUMN` and a declared `STATISTICS` in `ADD COLUMN` / `MODIFY COLUMN` used to be silently dropped and are now applied (and validated), like in `CREATE`: `ADD COLUMN` sets the declared statistics on the new column, and `MODIFY COLUMN` replaces the explicit statistics of the column. The column-declaration `STATISTICS` in these `ALTER` commands honors the `allow_statistics` setting and requires the same `ALTER ADD STATISTICS` / `ALTER MODIFY STATISTICS` access rights, like the dedicated `ADD/MODIFY/DROP STATISTICS` commands, and is excluded from the comment-only fast path in the distributed DDL routing. It is also gated on engine support (new `IStorage::supportsStatistics`, true for the `MergeTree` family): storages that reject the dedicated `ADD/DROP/MODIFY STATISTICS` commands, such as `Memory` or `Distributed`, reject the column-declaration spelling with the same `NOT_IMPLEMENTED` error instead of accepting it through the generic column alter. `StorageAlias` forwards `supportsStatistics` to its target table, and `StorageProxy` forwards it (together with `supportsTTL`) to the nested table, so support does not depend on whether the table is addressed directly, through an `Alias`, or through the lazy-loading proxy of a database with `lazy_load_tables = 1`. `COLLATE` and `PRIMARY KEY` in `ADD COLUMN` / `MODIFY COLUMN` now throw an exception instead of being silently ignored: some spellings of them (with a type present) used to parse successfully and do nothing, which was a bug, not a feature - per-column `PRIMARY KEY` cannot be altered at all. This also fixes a formatting round-trip: `formatImpl` prints `COLLATE` after `COMMENT` and after `TTL`, while the parser used to accept `COLLATE` only immediately after the type or after `NULL`/`NOT NULL`, so the formatted result of `x String COLLATE utf8_bin COMMENT 'text'` did not parse back: ```sql SELECT formatQuery(formatQuery('CREATE TABLE t (a String COLLATE utf8_bin COMMENT \\'a comment\\') ENGINE = Memory')); -- before: Code: 62. DB::Exception: Syntax error: failed at position 48 (COLLATE). ``` ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The modifiers of a column declaration - `COMMENT`, `CODEC`, `STATISTICS`, `TTL`, `COLLATE`, `PRIMARY KEY` and per-column `SETTINGS` - can now be written in any order in `CREATE TABLE`, and each of them at most once. Previously only one fixed order was accepted, and, for example, `x UInt64 CODEC(ZSTD) COMMENT 'text'` was a syntax error. In `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN`, the supported modifiers - `COMMENT`, `CODEC`, `STATISTICS`, `TTL` and per-column `SETTINGS` - can also be written in any order; per-column `SETTINGS` in `ADD COLUMN` and a declared `STATISTICS` in `ADD COLUMN` / `MODIFY COLUMN` are now applied instead of being silently dropped, and `COLLATE` and `PRIMARY KEY` in these `ALTER` commands now throw an exception instead of being silently ignored.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112788",
        "createdAt": "2026-07-31T18:32:53Z",
        "updatedAt": "2026-08-13T03:25:41Z",
        "timestamp": "2026-08-13T03:25:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 21
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112805",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not drop a named collection that a detached table still uses",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/96181 Related: https://github.com/ClickHouse/ClickHouse/issues/77366 Related: https://github.com/ClickHouse/ClickHouse/pull/110529 A table detached with a plain `DETACH TABLE` keeps its metadata file, so the server attaches it again on the next start. It is gone from `DatabaseCatalog` though, so `isTableExist` returns false for it, and the `check_named_collection_dependencies` check (added in #96181) treated its dependency as a stale leftover of a failed `CREATE TABLE`: it removed the dependency and let `DROP NAMED COLLECTION` succeed. The `ATTACH` replayed at startup then threw `NAMED_COLLECTION_DOESNT_EXIST`, which aborts loading the metadata, and the server did not start at all. This is how the `Stress test (arm_tsan)` job fails on master with `Cannot start clickhouse-server`: the AST fuzzer makes `04320_url_engine_dispatch_partition_and_format` leave its `URL(named_collection)` table detached (the test's `ATTACH TABLE` never reaches the server), and the test then drops the named collection. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=8aad759007771032aa94d7f4eee3e18103dd63c3&name_0=MasterCI&name_1=Stress%20test%20%28arm_tsan%29 ``` Application: Caught exception while loading metadata: Code: 722. DB::Exception: Waited job failed: Code: 695. DB::Exception: Load job 'load table test_3.04320_..._n' failed: Code: 669. DB::Exception: There is no named collection `04320_..._nc`: Cannot attach table `test_3`.`04320_..._n` from metadata file store/e5b/.../04320_..._n.sql from query ATTACH TABLE ... ENGINE = URL(`04320_..._nc`, format = 'JSON'). (NAMED_COLLECTION_DOESNT_EXIST) ``` The same signature accounts for 6 of the 8 `Cannot start clickhouse-server` failures with a missing named collection in the last 60 days (per `play.clickhouse.com`); the other two come from `03822_named_collection_drop_dependency_check`, which drops the collection with `check_named_collection_dependencies = 0` on purpose, and are the case #110529 handles by tolerating the missing collection at startup. ### Changes Per review feedback, the implementation is a simple in-memory bookkeeping in `NamedCollectionFactory` (an earlier revision inspected the metadata the detached table would be attached from, which required probing database disks and sweeping metadata directories after renames): - `DETACH TABLE` (and `DETACH DATABASE`, for every table inside) moves the dependencies of the table into a list of (collection, database, table) entries. - `DROP NAMED COLLECTION` is refused with `NAMED_COLLECTION_IS_USED` while an entry for the collection exists. - `ATTACH` does not remove the entry by itself: the dependencies are registered while the engine arguments are resolved, and the attach can still fail after that (an unknown format name, a failure creating the storage), leaving the table detached — the entry must keep protecting it. Instead, the entry is removed by the events that prove the metadata under that name is gone or harmless: `DROP TABLE`, `DETACH TABLE ... PERMANENTLY`, and `RENAME` of the (necessarily re-attached) table; `DROP DATABASE` removes the entries of the database's detached tables, and `RENAME DATABASE` re-keys them. The `DROP NAMED COLLECTION` check itself removes nothing: the table's existence in `DatabaseCatalog` is racy against in-flight attaches and detaches (the table can exist while nothing in the drop query has validated its live dependency), so the drop path is read-only and every recorded entry refuses the drop. - `DETACH TABLE ... PERMANENTLY` does not record an entry: a permanently detached table is not loaded at startup, so dropping a collection it references cannot break the server start. A later explicit `ATTACH` of such a table fails cleanly with `NAMED_COLLECTION_DOESNT_EXIST` and is recoverable by recreating the collection. - The list lives in memory only, which is consistent across a restart: a plainly detached table is attached again at the next start, where regular dependency tracking picks it up, and a permanently detached one records no entry at all. The list is deliberately imprecise in one direction: a stale entry may keep refusing the drop for a while after the detached table itself is gone, or after the table was attached back (until the table is dropped or renamed). In exchange, the `DROP NAMED COLLECTION` path performs no disk access at all. - A `RENAME TABLE` that moves a table between an `Ordinary` and an `Atomic` database changes the identity the dependency is keyed by: the move into `Atomic` assigns a fresh UUID to the table and the move out of it drops the UUID, while a dependency is keyed by the UUID for tables of `Atomic` databases and by the name for tables of `Ordinary` ones. The rename interpreter only knows the names, so the entry used to keep the identity the table had before the move and nothing found it afterwards: the detach recorded no entry and, for the `Atomic -> Ordinary` direction, the drop check even classified the entry of the still attached table as a leftover of a failed `CREATE` and dropped the collection from under it. `DatabaseOnDisk::renameTable` now re-keys the entries of the moved table to its new `StorageID`, where both identities are known. `EXCHANGE` is unaffected: it is only supported between two `Atomic` databases, where the UUIDs do not change. - A table of a database with `lazy_load_tables = 1` is attached as a `StorageTableProxy` and its real storage is built only on the first access, so the engine arguments are not resolved at load time and the dependency on the named collection they name stayed unregistered. `DROP NAMED COLLECTION` was then allowed while such a table still referenced the collection - breaking it at the first access with `NAMED_COLLECTION_DOESNT_EXIST` - and a `DETACH` of it had no dependency to move to the list of the detached ones, so the protection above silently disappeared in that mode. `DatabaseOrdinary::loadTableLazy` now registers the dependency straight from the metadata, via the new `tryGetUsedNamedCollectionName` helper. An identifier first argument counts as a collection reference only for the engines that resolve their arguments through named collections (a new `StorageFactory::StorageFeatures::supports_named_collections` flag) - for other engines an identifier means something else, e.g. a cluster name for `Distributed` - and for them the signal is time-stable: whether the collection currently exists is deliberately not checked, so the dependency of a collection that is missing at load time (say, after a drop with `check_named_collection_dependencies = 0`) protects it when it is recreated later. The exception is `Remote`/`RemoteSecure`, where the same identifier is also a valid positional argument - a cluster name - when the named-collection lookup does not resolve (marked by the new `StorageFeatures::named_collection_argument_is_ambiguous`; every other flagged engine treats an unknown collection as an error, not as a fallback to a positional form). Syntax alone cannot prove that such a table uses a collection, so the helper replicates the decision the engine's own argument parsing would make at the same moment - which is exactly what a non-lazy load of the same metadata does: the named-collection branch is taken only when a collection with that name exists, and only `key = value` overrides may follow the collection name, which a positional argument list never looks like. `MongoDB` and `MaterializedPostgreSQL` were the only flagged engines whose eager argument resolution did not pass the dependent table to `tryGetNamedCollectionWithOverrides` (they registered the dependency at the lazy load but not when the storage is built); they now pass it, so both load modes register the same dependency. `addDependency` ignores an exact duplicate, because the same dependency is registered again when the proxy is materialized. Tables created with `CREATE TABLE ... AS f(...)` need no lazy branch: a database with `lazy_load_tables = 1` deliberately loads them eagerly as a `StorageTableFunctionProxy` (see `DatabaseOrdinary::shouldLazyLoad`), and that load registers the dependency via `ITableFunction::getUsedNamedCollectionName`. - The cleanup of stale *active* dependencies (leftovers of a failed `CREATE TABLE`, pre-existing from #96181) no longer treats the table's absence from `DatabaseCatalog` alone as a proof of staleness: the dependency of an in-flight `CREATE`/`ATTACH` is registered while the engine arguments are resolved, before the table is committed to the catalog, and a concurrent `DROP NAMED COLLECTION` could prune it and drop the collection while the create later succeeds — recreating the broken metadata this PR fixes. The creating query holds the `DDLGuard` of the table name for the whole window between the registration and the commit, so the drop re-checks the table's existence under that guard before pruning: once the guard is acquired, no create is in flight, and the table's absence proves the entry is stale. (Entries with an empty database name come from dictionaries defined in the configuration files, which are not created through DDL; they are pruned as before.) The pruning removes only the exact stale entry (the collection and the recorded database, table and UUID): `CREATE TABLE ... UUID` can reuse the UUID of a failed create under a different table name, which the guard of the recorded name does not synchronize with, and removing everything under the UUID would erase the live dependency of such an in-flight create — the collection it uses could then be dropped from under the committed table. A new `create_table_pause_before_commit` failpoint keeps a create inside the window for the tests. ### Documented behavior impact `DROP NAMED COLLECTION` now rejects a collection that a detached table or a table in a detached database references, where it previously succeeded (and left a server that could not start). This is what `check_named_collection_dependencies` already promises - \"Check that DROP NAMED COLLECTION will not break tables that depend on it\" - so the documented behavior of the setting does not change, and setting it to `0` still allows the drop. No documentation update is needed. ### Verified locally (release build) - Before: `CREATE NAMED COLLECTION` + `CREATE TABLE ... ENGINE = URL(nc)` + `DETACH TABLE` + `DROP NAMED COLLECTION` succeeds, and the server then fails to start with the exact error chain above (exit code 210). Same with `DETACH DATABASE`. - After: the drop is refused, the collection stays, the table attaches back, and a restart of a server with the detached table present succeeds. - `04660_drop_named_collection_detached_table`, `04698_drop_named_collection_detached_after_rename`, and `04823_drop_named_collection_broken_attach` (all new, covering plain `DETACH TABLE` (blocks the drop), `DETACH TABLE ... PERMANENTLY` (does not block; the later `ATTACH` fails cleanly), `DETACH DATABASE`, `Ordinary` databases, renames of the table and of the database before and after the detach, stale dependencies of failed `CREATE TABLE`, and an `ATTACH` that fails after the dependencies were registered — the drop stays refused), and `04836_drop_named_collection_inflight_create` (new, runs `DROP NAMED COLLECTION` against a `CREATE TABLE` and an `ATTACH TABLE` paused between the dependency registration and the commit to the catalog: the drop blocks on the `DDLGuard` and is refused), and `04840_drop_named_collection_cross_engine_rename` (new, moves a table between an `Ordinary` and an `Atomic` database in both directions and then detaches the table and its database), and `04848_drop_named_collection_reused_uuid` (new, prunes the stale entry of a failed `CREATE TABLE ... UUID` while a create of a different table reusing the UUID is paused inside the window: the drop of the old collection succeeds, and the drop of the collection the new table uses stays refused; verified to fail without the fix), plus `03822_named_collection_drop_dependency_check` and `04003_named_collection_drop_dependency_check_dict` pass. - `test_named_collections/test.py::test_drop_while_used_by_lazily_loaded_table` (new integration test: a table using a named collection in an `Atomic` database with `lazy_load_tables = 1`, a server restart so the table comes back as a never-accessed proxy, and the drop refused both while it is attached and after `DETACH TABLE`). It is an integration test because the hole is only reachable once the in-memory list is empty, i.e. after a restart: without one, the `DETACH DATABASE` that precedes `ATTACH DATABASE` leaves its own entry behind and that entry refuses the drop on its own. Verified that it fails without the fix (the drop succeeds) and passes with it. - `test_drop_collection_recreated_under_lazily_loaded_table` (new integration test: the collection is dropped with `check_named_collection_dependencies = 0`, the server restarts while it is missing, and the drop of the recreated collection is refused; verified to fail without the fix), and `test_drop_not_used_by_lazily_loaded_distributed_table` (new integration test: a collection named after the cluster of a lazily loaded `Distributed` table is droppable; passes before and after, pinning the behavior). - `test_drop_not_used_by_lazily_loaded_remote_table` (new integration test: a collection named after the cluster of a lazily loaded `ENGINE = Remote(cluster, system, one)` table, created after the table, is droppable after a restart; verified to fail without the fix - the drop was refused with `NAMED_COLLECTION_IS_USED` by the unrelated table), and `test_drop_while_used_by_lazily_loaded_table_function` (new integration test pinning that a `CREATE TABLE ... AS bigquery(collection)` table in a database with `lazy_load_tables = 1` keeps blocking the drop after a restart, before the first access and after a `DETACH TABLE`: such tables are loaded eagerly as a `StorageTableFunctionProxy`, which re-registers the dependency; passes without any code change, confirming no lazy table-function branch is needed). ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the server failing to start after a named collection was dropped while a detached table (or a table in a detached database) still referenced it. `DROP NAMED COLLECTION` now counts detached tables as dependents and is refused with `NAMED_COLLECTION_IS_USED`, as `check_named_collection_dependencies` already does for attached tables.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112805",
        "createdAt": "2026-07-31T19:56:49Z",
        "updatedAt": "2026-08-13T17:56:31Z",
        "timestamp": "2026-08-13T17:56:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112816",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allow running queries detached from client session",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New setting `run_query_in_background`. The server accepts the query, immediately returns an empty result, and runs it to completion regardless of what happens to the connection. The result is discarded. Track the query by its `query_id` in `system.processes` and `system.query_log`. Intended for long queries like `INSERT ... SELECT`, `CREATE TABLE … AS SELECT`, or `CREATE MATERIALIZED VIEW … POPULATE` that must not die with a dropped connection.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112816",
        "createdAt": "2026-07-31T20:49:40Z",
        "updatedAt": "2026-08-13T17:37:14Z",
        "timestamp": "2026-08-13T17:37:14Z",
        "metrics": {
          "reactions": 4,
          "comments": 2
        },
        "labels": [
          "pr-feature"
        ],
        "author": "mstetsyuk",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112824",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cross-compile ClickHouse for Windows",
        "text": "Draft. Cross-compiles `clickhouse.exe` for `x86_64-w64-windows-gnu` (mingw-w64 + clang + lld → a native PE, no emulation layer and no MSVC licence). The goal is `clickhouse-client` and `clickhouse-local` running natively on Windows. **State: it links, it has never been run.** The development host is aarch64 and its Wine has no x86-on-ARM emulator, so everything here is compiled and linked but not executed. That is the next step and needs an x86-64 Windows machine. ``` programs/clickhouse.exe: PE32+ executable for MS Windows 6.00 (console), x86-64 520 MB · imports ADVAPI32, IPHLPAPI, KERNEL32, USER32, WS2_32, dbghelp, msvcrt ``` Every import ships with Windows, so the binary is a single self-contained file. Linux builds and links `clickhouse` unchanged after every commit in the branch. CI builds `clickhouse` for Windows on every pull request, so the port cannot silently regress. ## What is here The runtime (libc++/libc++abi/libunwind with SEH, compiler-rt builtins), 117 contribs behind a self-maintaining CI gate, all of Poco — whose Windows layer had been deleted from this fork and is restored from upstream 1.9.3 — every library under `src`, and `programs`. Roughly 900 of the changed lines are one mechanical class: `std::filesystem::path` used where a `String` is wanted, which compiles on POSIX through an implicit conversion that does not exist on Windows. Implemented natively rather than stubbed, because a client needs them: terminal size and console encoding, raw console mode and keystroke reading, Ctrl+C and Ctrl+Break, socket liveness, `statvfs` via `GetDiskFreeSpaceEx`, file mapping, `pread`, per-thread CPU time, `setThreadName`, `isLocalAddress`, `readpassphrase`, and a `WakeupFd` built on a loopback socket pair (a Windows pipe cannot be waited on alongside sockets). Compiled out with the reason recorded at each site, all server-side or POSIX-only: the sampling profiler and signal handlers (Windows reports faults through SEH), `ThreadFuzzer`, the `fork`-based watchdog, `ShellCommand` and everything built on it, the pseudo-terminal features, and the `su`/`docker-init`/`install` tools. ## Known gaps - Not executed, as above. - `-g0` on Windows: a PE image cannot exceed 4 GiB and the DWARF alone is several times that, and PE has no `.gnu_debuglink` equivalent to carry it separately. A crash symbolizes to module and offset, not file and line. - No `Epoll` backend, so nothing that polls sockets through it works yet. - Local syslog, archives (`libarchive` needs a hand-written Windows `config.h`), conditional writes to the local object storage (they need `flock` on a directory and sub-second modification times), and the web terminal report `NOT_IMPLEMENTED`. `docs/en/development/build-cross-windows.md` has the build instructions and a per-subsystem inventory of what remains. Related: https://github.com/ClickHouse/ClickHouse/pull/112185 Related: https://github.com/ClickHouse/ClickHouse/pull/112767 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a cross-compilation target for Windows (`x86_64-w64-windows-gnu`), producing a native `clickhouse.exe`. The build is not yet tested at runtime. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112824",
        "createdAt": "2026-07-31T21:57:07Z",
        "updatedAt": "2026-08-13T12:46:59Z",
        "timestamp": "2026-08-13T12:46:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 15
        },
        "labels": [
          "pr-build",
          "submodule changed"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112828",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Recover a NATS JetStream subscription closed by the broker",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/96651 Related: https://github.com/ClickHouse/ClickHouse/pull/103557 Related: https://github.com/ClickHouse/ClickHouse/pull/112464 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a `NATS` table with `nats_stream` set silently consuming nothing after the NATS server is restarted. The `JetStream` subscription is now re-established automatically instead of requiring `DETACH TABLE` and `ATTACH TABLE`. ### Description A `NATS` engine table reading from a `JetStream` stream stops consuming permanently once the NATS server is restarted. Nothing is logged, the connection reports healthy, and only `DETACH TABLE` plus `ATTACH TABLE` or a server restart recovers it. It is also how the flaky `test_nats_restore_failed_connection_without_losses_on_write` fails on master. Root cause: an asynchronous pull subscription renews its pull request only when a message is delivered, and a reconnect resends the `SUB` line without the outstanding pull request, so with nothing in flight when the server goes away the chain never restarts. The server does report this, answering the outstanding request with `409 Server Shutdown`, and the client then closes the subscription. ClickHouse missed it because `isSubscribed` only tests whether the subscription vector is non-empty, so the existing re-subscribe path was gated on a predicate that cannot see a dead subscription. This adds a per-subscription liveness check and consults it in the streaming task, which drops the subscriptions so the existing re-subscribe runs in the same iteration. Only `JetStream` consumers opt in: core NATS subscriptions are already restored by the client, and recovery drops buffered messages core NATS never redelivers. Validated with five new integration tests. Three restart the broker and fail on master, 9 of 9 repeats, before passing after the change; a fourth asserts a healthy consumer never re-subscribes, so it passes either way and exists to bound the cost. The fifth covers a restart of a table reading two subjects, which nothing covered before. Not covered: a broker loss leaving the subscription with no status at all, such as a hard kill, a partition, or a loss coinciding with a re-subscribe. That needs a local fetch timeout, which would also periodically tear down healthy subscriptions. #103557 targets the same defect from a connection-level reconnect counter, but no longer applies to this code and has no integration test. Close whichever you prefer. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1343` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112828",
        "createdAt": "2026-07-31T22:25:07Z",
        "updatedAt": "2026-08-13T16:41:13Z",
        "timestamp": "2026-08-13T16:41:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 17
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "antaljanosbenjamin",
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112847",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Make SQL SECURITY views an optimization barrier",
        "text": "### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A view with `SQL SECURITY DEFINER` or `SQL SECURITY NONE` is now an optimization barrier, so an expression in the query reading the view is never evaluated on rows that the view itself filters out. Previously such an expression could observe the hidden rows through an exception, and a view used to restrict which rows a user may see did not actually restrict them. --- ## The problem `SQL SECURITY DEFINER` is widely used to build a view that restricts which rows a user may see: ```sql CREATE VIEW user_query_log DEFINER = default SQL SECURITY DEFINER AS SELECT * FROM system.query_log WHERE user = currentUser(); GRANT SELECT ON user_query_log TO alice; ``` `alice` has no grant on the source table, only on the view. But the outer `WHERE` and the view's own `WHERE` are merged into a single filter over the source table, and nothing guarantees which of the two decides first. Any function that can signal through a side channel therefore observes the rows the view is supposed to hide: ```sql -- as alice, whose only grant is SELECT ON user_query_log SELECT event_time FROM user_query_log WHERE throwIf(query LIKE '%some probe%', 'DISCLOSED'); -- DB::Exception: DISCLOSED ... While executing MergeTreeSelect ``` That is a one-bit oracle per query. Exception messages carry the offending value, which turns it into a plain read of the hidden rows: ```sql SELECT * FROM v_secrets WHERE if(owner = 'alice', 1, toUInt8(secret)) = 1; -- DB::Exception: Cannot parse string 'TOP-SECRET-BOB-42' as UInt8 ``` Both work on `SQL SECURITY NONE` as well, and both work with `enable_analyzer = 0`. Row policies are *not* affected: a row policy is a separate `row_level_filter` that the reading step always applies before PREWHERE and before any pushed-down filter. This PR gives a view's own filtering the same standing. ## The fix `IQueryPlanStep` gets a `security_barrier` flag. After `StorageView::readImpl` builds the view's subplan, every step that can drop rows is marked, and optimizations that move a step down the plan refuse to cross a marked step: - `tryMergeExpressions`, `tryMergeFilters`, `tryPushDownFilter` and `tryMergeFilterIntoJoinCondition` refuse when the child is a barrier; - `optimizePrewhere` refuses to pull an outer filter into a barrier source — conditions are combined into the prewhere DAG with `and`, which gives no ordering guarantee — and transfers the barrier onto the source when it absorbs the view's own filter; - `trySplitFilter` moves the flag onto the new lower `FilterStep`, which is the one that still drops rows. - `tryPushDownLimit` refuses to move the invoker's `LimitStep` below a barrier step: once across the seal it would seed `DistinctStep::limit_hint` or a sorting limit inside the view's subplan, and `optimizeLimitForAggregationInOrder` walked through the seal to seed `AggregatingStep::limit_hint` the same way — hints that stop reading the source once enough visible rows are produced, so `read_rows`, progress and timing depended on the rows the view drops or collapses. Both walks, and `pushLimitByIntoSort`, fail closed on a barrier step. Blocking the steps that evaluate expressions is not enough on its own, because index analysis reaches the source by a different route and skips granules by the values of the rows the view hides — `read_rows` then tells the invoker whether such a row exists, with no exception needed. Eight walks are fenced as well: - `optimizePrimaryKeyConditionAndLimit` walks up from the reading step and hands every `FilterStep` it meets to the source's key condition; it now stops once it has consumed a barrier. The barrier's own condition still reaches the source — it is the definer's, and it is what decides which rows exist for the view — but nothing above it does. - `StorageView::readImpl` forwarded `query_info.filter_actions_dag` into the view's inner analyzer, where `Planner::collectFiltersForAnalysis` injects it into the inner plan and the filters it collects reach the inner tables' index analysis. It is no longer forwarded for a barrier view, which costs such a view over `Distributed` its shard skipping on the outer predicate. - `buildSortingDAG` in the read-in-order analysis descended through the view subplan and pulled outer `FilterStep` predicates into the fixed columns and the merged DAG, so an outer `ORDER BY` / `GROUP BY` / `DISTINCT` / `LIMIT BY` could still shape how the source below the barrier reads. It now reports a barrier in the chain and every consumer (sorting, aggregation, `DISTINCT`, `LIMIT BY`, the normal-projection choice, top-K, and the `Merge` child-plan walk) skips the analysis for that chain. Unlike the primary-key walk, the barrier step cannot be consumed here: the sort description belongs to the top of the chain, and a DAG missing the steps above the barrier could resolve a renamed column by name to the wrong source column and change the result order — so the analysis is skipped entirely, which is fail-closed and keeps correctness because the sorting step stays in the plan. - `tryOptimizeTopK` rewrites an `ORDER BY ... LIMIT` into a dynamic `__topKFilter` PREWHERE and minmax-skip-index granule pruning on the source, walking `LimitStep` → `SortingStep` → `ExpressionStep` → `FilterStep` → `ReadFromMergeTree` — and both rewrites are on by default. The walk now fails closed on the first barrier step it meets, including a reading step that carries the barrier after `optimizePrewhere` absorbed the view's own filter. - `tryTopKThroughJoin` peels the expression chain between the invoker's `Sorting` and a `Join` and grafts the invoker's `Sort + Limit` onto the join's preserved input, re-running the optimization passes on that subtree. For a barrier view whose inner query is a join, the graft landed below the seal — verified live pre-fix on both analyzers. The pass now fails closed on a marked `Limit`/`Sorting` (the pattern lies inside a view), a marked peeled expression (the seal), and a marked join (a join of a barrier view is always marked, not being row-preserving). Grafting above the seal of the invoker's own join input stays allowed: the inserted `Sort` consumes its whole input and the re-run passes are individually fenced. - `registerLeftSideIndexAnalysisSecondPass` of the join runtime filters walked from the `__applyFilter` step (which the fenced filter pushdown correctly keeps above the seal) down through every single-child expression or filter step — sealed or not — to the `ReadFromMergeTree` inside the view, and registered the invoker's build-side keys for granule pruning there. The walk now fails closed on the first barrier step. A live disclosure could not be constructed — the descriptor's key name has to match the read's namespace and the consumption path declined in every configuration tried — so this fence makes the contract structural rather than incidental. - The vector search rewrites — the vector-similarity-index pass and the quantized-codes shortlist — walk the same chain and prune the source to the top-N candidates of the invoker's `ORDER BY`. Both now fail closed on a barrier step the same way. - Projection planning: `QueryDAG::build` in `projectionsCommon.cpp` collects every filter of the chain below the aggregation — the invoker's predicates together with the view's own filtering — and both `optimizeUseAggregateProjections` (including `minmax_count_projection`) and `optimizeUseNormalProjections` prune parts and marks with it, so a projection-enabled table under a filtering barrier view still let the invoker's predicate shape the read. The DAG build now fails closed on the first barrier step, which makes the projection optimizations decline the read entirely. One more family reaches the source without going through its index analysis at all. `optimizeDistinctPerPartition`, `optimizeLimitByPerPartition` and `optimizeAggregationPerPartition` walk down through the sealing step and ask the reading to output each partition through a separate port, and `applyStreamDisjointness` carries the resulting partition disjointness back up across the seal, so the invoker's `DISTINCT` / `LIMIT BY` / `GROUP BY` skips its stream merging as well. Read scheduling, progress and resource consumption below the view then depend on how the rows the view drops are spread over the partitions, and all three `allow_*_partitions_independently` settings default to `1`. Both directions now fail closed on a barrier step. Four passes endanger the barrier by rebuilding steps rather than by walking past them, and each now fails closed on a barrier step: - lazy materialization splits every `Expression` / `Filter` step of the chain into a main and a lazy half, and the rebuilt steps do not carry the barrier flag, so the post-lazy `tryMergeExpressions` / `tryMergeFilters` passes saw an unmarked chain and could merge an invoker predicate into the view's own filtering — reopening the exception oracle itself, not just the read-shaping one; - `tryLiftUpUnion` rebuilds the `UnionStep` and clones the parent step into the branches as fresh unmarked steps, so a barrier view over `UNION ALL` lost its seal and `tryPushDownFilter` could then duplicate an invoker predicate into the branches; - `tryExecuteFunctionsAfterSorting` replaces the expression under a `SortingStep` with two new unmarked steps, which would strip the seal of a wrapper view and let the read-in-order and top-K walks descend through it again. - `tryLiftUpArrayJoin` splits the expression or filter above an `ArrayJoinStep`, moves one half below the `ARRAY JOIN`, and rebuilds both halves as fresh unmarked steps. When the parent is the seal of a view whose plan contains `ARRAY JOIN` (the seal is non-trivial whenever the view declares explicit column names or types), the invoker's predicate descended below the `ArrayJoinStep` and was evaluated on rows hidden by empty arrays — a live disclosure through the exception oracle, on both analyzers. Measured on a `DEFINER` view that exposes no row, over 100000 rows sorted by `key`, reading `WHERE key = <a key only a hidden row has>` against `WHERE key = <a key nothing has>`: 576 rows read against 0 without this, and 1000000 against 1000000 with it, on both `enable_analyzer = 1` and `enable_analyzer = 0`. `EXPLAIN SYNTAX` is left alone. It builds `InterpreterSelectQuery` with `only_analyze`, whose plan reads from `ReadNothingStep`, so no expression of the outer query is ever evaluated on a row and the view is still inlined for the diagnostic. Two paths substitute the view into the outer query before a plan exists, so a plan-level barrier cannot see them, and both are closed the same way — such a view is not inlined and is read through `StorageView::read`, which keeps the outer predicate in a step above it: - with `enable_analyzer = 0`, `InterpreterSelectQuery` replaces the view with a subquery and `TreeRewriter` merges the predicates; - with `analyzer_inline_views = 1`, `QueryAnalyzer::inlineViewSubqueryIfNeeded` does the same in the query tree. Without this, `SET enable_analyzer = 0` or `SET analyzer_inline_views = 1` would bypass the fix entirely. The flag also travels with a serialized query plan. A distributed worker deserializes a fragment and optimizes it again, so a barrier it does not know about is a barrier it will optimize across. `QueryPlan::serialize` writes the flag per step and fails closed when the negotiated query plan serialization version predates it (the version is bumped to 6), rather than sending a plan that silently loses its protection. ## Views that hide nothing are left alone Only a view that can actually drop rows becomes a barrier. A step is row-preserving if it is an `ExpressionStep`, a `SortingStep` without a limit, or a source step without PREWHERE; if none of the view's steps is anything else, nothing is marked and the plan is what it is today. The pre-plan paths make the same distinction. `StorageView::canHideRows` proves over the view's definition that the inner query preserves every row of a plainly readable source, and only a view for which that proof fails loses inlining and the forwarded outer filter. The proof fails closed: filters, limits, aggregation, `DISTINCT`, joins, `ARRAY JOIN`, `SAMPLE`/`FINAL`, multi-select unions, table functions, and a `FROM` that is itself a view or a view-wrapping engine (`Merge`, `Buffer`, anything remote) all count as able to hide rows. The proof classifies the storage that actually serves the read, not the object the name resolves to: proxy layers (a lazily loaded table of a database with `lazy_load_tables = 1`, or a table created from a table function) and `Alias` tables are unwrapped first, failing closed on a chain that cannot be resolved, and a storage that rewrites its own reads with `FINAL` and a `_sign` filter (`MaterializedPostgreSQL`) counts as able to hide rows even without another wrapper. So a projection-only `DEFINER` view produces exactly the plan of the same view declared `SQL SECURITY INVOKER` on every path, which the test pins byte-for-byte. Measured on 20M rows with `SELECT sum(length(payload)) FROM v WHERE tag = 'RARE'`, from `system.query_log`: | view | barrier | read_rows | read_bytes | ms | |---|---|---|---|---| | `DEFINER`, filters rows | on | 20 000 000 | 1.43 GiB | 35 | | `DEFINER`, filters rows | off | 1 638 400 | 82.84 MiB | 14 | | `DEFINER`, projection only | on | 1 638 400 | 82.84 MiB | 9 | | `DEFINER`, projection only | off | 1 638 400 | 82.84 MiB | 11 | A projection-only view is unaffected. A filtering view does pay: PREWHERE then holds only the view's own condition, so it no longer skips granules on the outer predicate. That is the inherent price of the guarantee — PostgreSQL's `security_barrier` views behave the same way — and it applies only to views that restrict rows, which are exactly the ones where it matters. The new server setting `sql_security_views_are_optimization_barriers` (default `1`) turns it off. It is deliberately a **server** setting and not a user setting: a user setting would be turned off by the very query that is trying to read the hidden rows. ## Testing `04758_sql_security_view_barrier_read_rows` covers the `read_rows` oracle through index analysis, on both analyzers, and prints `DISCLOSED` on both with the setting off. `04670_sql_security_view_barrier` covers the leak on `DEFINER` and on `NONE`, with `enable_analyzer = 1`, with `enable_analyzer = 0` and with `analyzer_inline_views = 1`, the value leak through a cast error message, the same oracle through a shard with `serialize_query_plan = 1`, that a projection-only `DEFINER` view and an `INVOKER` view still have the outer predicate merged into the view's own filter, and that a projection-only view keeps PREWHERE. `04813_sql_security_view_barrier_read_in_order` pins the read-in-order fence: the `INVOKER` twin and a projection-only `DEFINER` view read `InOrder`, a filtering `DEFINER` view does not, under both analyzers, with unchanged results. `04817_sql_security_view_barrier_top_k` pins the top-K fence: the `INVOKER` twin gets the `__topKFilter`, the filtering `DEFINER` view does not, and `read_rows` of an `ORDER BY ... LIMIT 1` over twin views is identical whether or not the hidden row holds the extreme minimum of the sort column that minmax pruning would rank first. `04818_sql_security_view_barrier_lazy_materialization` pins the lazy-materialization fence the same way and checks that an invoker predicate over the view is never evaluated on the hidden row. `04821_sql_security_view_barrier_projections` pins the projection fence: the `INVOKER` twin uses both a normal and an aggregate projection, the filtering `DEFINER` view uses neither, and `read_rows` of a predicate probe over twin views is identical whether or not the hidden row matches it. `04825_sql_security_view_barrier_union` pins that an outer predicate over a filtering `DEFINER` view on `UNION ALL` stays in a single filter above the union — before the `tryLiftUpUnion` fix it was duplicated into the branches — while the `INVOKER` twin keeps the pushdown. `04826_sql_security_view_barrier_functions_after_sorting` pins that an `ORDER BY ... LIMIT` over a wrapper `DEFINER` view (a `Merge` table over a nested filtering view) produces no in-order reading and no `__topKFilter` with `query_plan_execute_functions_after_sorting` on, while the `INVOKER` twin exploits the source order. `04827_sql_security_view_barrier_masked_wrappers` pins that the classification survives engine masking: a `DEFINER` view over a `Merge` wrapper behind a lazy `TableProxy` (re-masked before every round, since planning materializes the proxy) or behind an `Alias` table plans differently from its `INVOKER` twin on both analyzers and with `analyzer_inline_views = 1`. `04832_sql_security_view_barrier_limit_pushdown` pins the LIMIT fence: over a `DISTINCT` `DEFINER` view the invoker's `LimitStep` stays above the sealing step on both analyzers — before the fix it crossed the seal and sat directly on the `DistinctStep`, where it seeds the hint — and `read_rows` of an `ORDER BY ... LIMIT 1` over twin in-order `GROUP BY` views is identical whether the first group holds one raw row or almost all of them. `04837_sql_security_view_barrier_per_partition` pins the per-partition fence, with the `allow_*_partitions_independently` settings left at their defaults: an outer `DISTINCT` / `LIMIT BY` over a filtering `DEFINER` view produces none of the `Skip stream merging` / `Read each partition through separate port` markers its `INVOKER` twin gets, and the disjointness of a view whose own inner `DISTINCT` legitimately requests per-partition reading does not propagate across the seal into the invoker's `GROUP BY` / `LIMIT BY` — before the fix every `DEFINER` case was identical to its twin. `04840_sql_security_view_barrier_array_join` pins the `ARRAY JOIN` lift-up fence with an exception oracle on both analyzers: the `INVOKER` twin's predicate legitimately descends below the `ARRAY JOIN` and throws on the row an empty array hides, while the `DEFINER` view counts without throwing and its plan keeps every `throwIf` line above the `ArrayJoin` step — before the fix the `DEFINER` view threw as well. `04891_sql_security_view_barrier_top_k_through_join` pins the top-K-through-join fence: the `INVOKER` twin of a view over a `LEFT JOIN` gets the preserved-side `Sort + Limit` graft below the join, the `DEFINER` twin keeps its join input untouched — before the fix the `DEFINER` plan got the graft below the seal. `04892_sql_security_view_barrier_join_runtime_filter` pins the join-runtime-filter contract: with `enable_join_runtime_filters_index_analysis = 1`, twin filtering `DEFINER` views over tables identical except for the hidden row's primary-key value read exactly the same number of rows, while the `INVOKER` control is pruned by the build-side key. Every setting the plan shape depends on is pinned, because the test also runs with randomized settings. With `sql_security_views_are_optimization_barriers = 0` every one of those lines changes, so none of them passes vacuously. Ran 542 existing tests matching `view`, `prewhere`, `push_down`, `pushdown`, `row_policy`, `sql_security` and `definer`. 60 fail in my local environment, and the identical 60 fail with the barrier disabled on the same binary, so this introduces no regressions among them. ## Not covered here - The serialized-plan fence is defensive. With the serialization change reverted I could not make the barrier loss observable — neither the `throwIf` nor the failing-cast oracle leaks through `serialize_query_plan = 1`, `make_distributed_plan = 1`, or a `Distributed` table, because the initiator optimizes the fragment before shipping it. The flag is serialized so that the guarantee does not depend on that. - A separate hole remains: `additional_table_filters` and `additional_result_filter` from the invoker are copied verbatim into the definer's context by `StorageInMemoryMetadata::getSQLSecurityOverriddenContext`. Keyed on the view's *inner* table, the filter lands inside the view and its expression — which may contain scalar subqueries and table functions — is evaluated with the definer's privileges. That is a privilege escalation rather than a row disclosure, it is unaffected by this PR, and it needs its own fix. Until then it can be mitigated with a constraint on the definer's profile (`CREATE SETTINGS PROFILE p SETTINGS additional_table_filters = '' CONST TO <definer>`), which does not help for `SQL SECURITY NONE`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112847",
        "createdAt": "2026-08-01T02:18:25Z",
        "updatedAt": "2026-08-13T12:30:44Z",
        "timestamp": "2026-08-13T12:30:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-must-backport",
          "pr-critical-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "Algunenano"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112873",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Azure: log batch-delete events when SubmitBatch itself fails",
        "text": "> **Series**: #112871 -> **#112873** (this), #112872, #112874, #112875, #112876 **Problem.** When the batch `SubmitBatch` delete request itself fails, the per-object response loop is skipped, so zero `system.blob_storage_log` Delete events are recorded for the whole batch. Failure scenario: ``` removeObjectsBatchIfExists: SubmitBatch() throws (e.g. 403 on the batch endpoint) -> per-object GetResponse() loop skipped -> 0 Delete events logged ``` **Fix.** Record a Delete event per object on batch-level failure before rethrowing. **Changes.** - `Disks/…/AzureBlobStorage/AzureObjectStorage::removeObjectsBatchIfExists`: emit per-object `blob_storage_log` Delete events on `SubmitBatch` failure. - `tests/integration/test_azure_403_handling`: batch-delete-logging test + `configs/blob_log.xml`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed Azure batch object deletion recording no `system.blob_storage_log` Delete events when the batch request itself failed, leaving the whole batch unlogged. A Delete event is now recorded for each object on a batch-level failure.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112873",
        "createdAt": "2026-08-01T06:53:30Z",
        "updatedAt": "2026-08-13T15:58:20Z",
        "timestamp": "2026-08-13T15:58:20Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "arsenmuk",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112874",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "MergeTree: rethrow retryable errors from checkDataPart instead of returning empty checksums",
        "text": "> **Series**: #112871 -> **#112874** (this), #112872, #112873, #112875, #112876 **Problem.** `checkDataPart` swallows a retryable error and returns empty checksums, so `CHECK TABLE` and the fetch path treat a transient failure as \"verified\" and can persist an empty integrity baseline. Failure scenario: ``` checkDataPart(): catch retryable -> return {} CHECK TABLE : empty -> reports OK downloadPartToDisk: empty -> accepts an unverified packed part ``` **Fix.** Rethrow retryable errors so callers' retry/skip guards fire; remove the now-dead empty-checksums guard on the fetch path. **Changes.** - `Storages/MergeTree/checkDataPart`: `return {}` -> `throw` on a retryable error. - `Storages/MergeTree/DataPartsExchange`: delete the now-dead empty-checksums fetch guard. - `tests/integration/test_azure_403_handling`: CHECK-TABLE-surfaces-transient test. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `checkDataPart` swallowing a retryable error and returning empty checksums, so `CHECK TABLE` could report a part as OK and a fetch could accept an unverified part after a transient failure. Retryable errors are now rethrown so the check is retried later instead of masking a transient failure as verified.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112874",
        "createdAt": "2026-08-01T06:53:40Z",
        "updatedAt": "2026-08-13T06:49:01Z",
        "timestamp": "2026-08-13T06:49:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "arsenmuk",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112875",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Azure: retry transient authentication failures on the remaining object-storage call sites",
        "text": "> **Series:** **#112871** → #112875 (this), #112872, #112873, #112874, #112876 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Retry a transient Azure `AuthenticationException` (credential/token-acquisition failure, e.g. during the managed-identity/RBAC propagation window) on object-storage call sites that previously had no retry at all: blob existence/metadata/size probes, single and batch deletes, the post-upload verification, the container existence check, and ADLS Gen2 writes. Previously a single transient credential failure at any of these points failed the whole query or background job. --- Follow-up to #112871, which handles the HTTP half of the problem: the Azure SDK `RetryPolicy` now retries a transient 403 for every client (`retry_options.StatusCodes.insert(Forbidden)` in `AzureBlobStorageCommon.cpp`). What the SDK cannot retry is `Azure::Core::Credentials::AuthenticationException`: it is thrown by the credential/token layer *around* the transport, so the retry policy — which only sees HTTP responses and `TransportException` — never observes it. The read/download/write buffer loops already absorb it via their `catch (...)` retry, but these call sites had no retry of any kind: - `AzureObjectStorage::exists` and `getObjectMetadata` / `getObjectMetadataIfExists` — one-shot `GetProperties()` - `AzureObjectStorage::removeObjectImpl` (blob and ADLS) and the batch `SubmitBatch` - `ReadBufferFromAzureBlobStorage::tryGetFileSize` / `getRemoteFileMetadata` - `WriteBufferFromAzureBlobStorage::finalizeImpl` post-upload verification — the upload already succeeded, a transient credential failure must not turn a good write into a failure - `containerExists` in the client factory — the first request many flows issue - `WriteBufferFromAzureDataLakeStorage::runWithRetries` — retried HTTP errors but let `AuthenticationException` escape on the first attempt A single transient credential failure at any of them surfaced as a query/job failure. Changes: - New `retryAzureOnAuthError()` (`src/IO/AzureBlobStorage/retryAzureOnAuthError.h`): bounded exponential backoff around `AuthenticationException` **only**. HTTP-status retries deliberately stay in the SDK retry policy (#112871), so retry layers don't stack multiplicatively. `getBlobPropertiesWithRetry()` lives in the same header for the `GetProperties()` sites. - Route the call sites above through it; widen their `Azure::Storage::StorageException` catches to the base `Azure::Core::RequestFailedException` (a behavior-preserving superset; NotFound handling unchanged). - ADLS Gen2 `runWithRetries` additionally retries `AuthenticationException` within the same budget (shared give-up/backoff tail for both failure kinds). - `readBigAt`: guard against a null body stream in the download response (mirrors the existing guard in `initialize()`). - Tests (`test_azure_403_handling`): a one-shot injected `AuthenticationException` on the direct metadata path (no ClickHouse-level retry loop anywhere) must be absorbed — this test fails without this PR; a permanent one must still fail with the real auth error after the bounded budget. Compared to the previous revision of this PR: all 403/HTTP retrying has been removed from the helper. After #112871 that is the SDK retry policy's job; doing it here as well would stack a second retry loop on top of the SDK's (up to 3 × 11 attempts on a permanent 403). This PR is now strictly about the auth exception the SDK cannot see. The per-object batch re-issue loop is gone for the same reason (an RBAC 403/auth failure is per-principal, i.e. all-or-nothing across a batch; batch-level retry covers it). Not covered by tests: the ADLS Gen2 write path (azurite cannot emulate ADLS/OneLake endpoints). Stacked on #112871 (branch base); to be merged after it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112875",
        "createdAt": "2026-08-01T06:53:55Z",
        "updatedAt": "2026-08-13T01:40:18Z",
        "timestamp": "2026-08-13T01:40:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "arsenmuk",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112876",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Azure: fix ranged copy corruption and harden the copy path",
        "text": "> **Series**: #112871 -> **#112876** (this), #112872, #112873, #112874, #112875 **Problem.** A ranged Azure copy corrupts the destination — native `CopyFromUri` copies the whole source blob (ignoring offset/size) and the single-part read+write fallback reads from byte 0 — and a transient 403/auth after the server-side copy starts spawns a second writer to the same blob. Failure scenario: ``` copyFile(src, offset=X>0, size=S): native CopyFromUri -> copies the ENTIRE src (ignores X,S) -> corrupt single-part fallback -> reads [0,S) not [X,X+S) -> corrupt transient 403 after StartCopyFromUri -> read+write fallback -> 2 writers to dest ``` **Fix.** Gate native copy on `offset==0`; seek the fallback via `LimitSeekableReadBuffer(offset, total_size)`; only *start* the copy in the guarded block and poll the same operation on transient errors (no mechanism switch); retry the read+write upload writes through the shared retry helper. **Changes.** - `IO/AzureBlobStorage/copyAzureBlobStorageFile`: `offset==0` native-copy gate; seek the single-part fallback; poll-same-operation on transient errors; retry upload writes. - `tests/integration/{test_backup_restore_azure_blob_storage,test_backup_restore_s3}`: incremental Log-family / ranged-copy regression tests. ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a ranged Azure Blob Storage copy corrupting the destination during incremental backups: server-side native copy ignored the requested offset and size, and the single-part read-and-write fallback read from the start of the source. Native copy is now used only for a full-object copy, the fallback seeks to the correct offset, and a transient error after a server-side copy has started no longer starts a second writer to the destination.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112876",
        "createdAt": "2026-08-01T06:54:19Z",
        "updatedAt": "2026-08-13T01:11:55Z",
        "timestamp": "2026-08-13T01:11:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-critical-bugfix"
        ],
        "author": "arsenmuk",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112878",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keep the part minmax index over `_block_number` / `_block_offset` across a reload of a mutated part",
        "text": "A mutation that does not rewrite the whole part (the usual case for a `Wide` part: `ALTER UPDATE` of a single column, `RENAME COLUMN`, `DROP COLUMN`) only hardlinks the source part's files and copies its in-memory minmax index into the new part. That is enough for the partition-key columns, whose minmax files the source part always has on disk, but not for `_block_number` / `_block_offset` under `part_minmax_index_columns = 'with_block_number_offset'`: a level-0 unmutated part has no minmax file for them, because `MinMaxIndex::load` synthesizes their ranges from the part name and the row count. The mutated part is no longer eligible for that synthesis — its `mutation` is not zero, and a mutation may drop rows, so the offset range is no longer `[0, rows_count - 1]` — so the very next reload of the table (`DETACH`/`ATTACH`, or a server restart) read back the whole universe and the index stopped pruning. `finalizeMutatedPart` now materializes the inherited minmax index for the files that were not carried over, and `MinMaxIndex::store` skips a file the caller already recorded in the checksums, so a hardlinked file is never written through into the source part. Reproducer (before this change the second `SELECT` reports `(NULL,NULL)`): ```sql CREATE TABLE t (d Date, s String) ENGINE = MergeTree ORDER BY tuple() SETTINGS enable_block_number_column = 1, enable_block_offset_column = 1, part_minmax_index_columns = 'with_block_number_offset', min_bytes_for_wide_part = 0; INSERT INTO t SELECT toDate('2018-10-01') + number % 3, toString(number) FROM numbers(9); ALTER TABLE t UPDATE s = 'x' WHERE 1 SETTINGS mutations_sync = 2; SELECT DISTINCT part_name, minmax__block_number FROM mergeTreeIndex(currentDatabase(), 't', with_minmax = 1); DETACH TABLE t SYNC; ATTACH TABLE t; SELECT DISTINCT part_name, minmax__block_number FROM mergeTreeIndex(currentDatabase(), 't', with_minmax = 1); ``` Found by the `DETACH`/`ATTACH` randomization of https://github.com/ClickHouse/ClickHouse/pull/96130, which made `04652_part_minmax_block_columns_mutation` fail in [this CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=96130&sha=c4401c88193b6b28480990e9363f21ab7523b061&name_0=PR&name_1=Stateless%20tests%20%28arm_asan_ubsan%2C%20targeted%29). Related: https://github.com/ClickHouse/ClickHouse/pull/96130 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix the per-part minmax index over `_block_number` and `_block_offset` being lost after a table reload when the part was produced by a mutation that does not rewrite the whole part. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112878",
        "timestamp": "2026-08-12T22:14:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112890",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix transform_null_in=1 for a non-Nullable key vs a Nullable IN set",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/111340 Closes: https://github.com/ClickHouse/ClickHouse/issues/112905 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes several problems with a non-`Nullable` primary key column compared against a subquery or array whose values are `Nullable`, for example `s IN (SELECT s FROM t UNION ALL SELECT NULL)` under `transform_null_in = 1`. Such a query no longer fails with `Cannot convert NULL value to non-Nullable type`. `NOT IN` and `NOT has` no longer drop rows whose key value happens to equal the key type's default (`''` for `String`, `0` for numbers), which previously happened because the `NULL` was folded into that default and then used to prune partitions. And `IN` / `NOT IN` no longer return wrong results for three specific conversions that map two distinct key values onto one set value: between different text types, across a loss of `Decimal` / `DateTime64` / `Time64` scale, and where a temporal element cannot represent one of the key's components. Other conversions that can collapse are not addressed here and are unchanged. ### Description Three defects in MergeTree `KeyCondition` set-index analysis (`tryPrepareSetColumnsForIndex`). **1. Exception.** `canBeSafelyCast(Nullable(X), T)` returned `true` for a non-`Nullable` `String` target, where a `NULL` has no representation, so the strict `castColumn` branch threw error 349 instead of the NULL-safe fallback. Fixed in the predicate. **2. Rows silently dropped.** A source-`NULL` position carried the key type's DEFAULT, injecting a value the query never wrote into the pruning set. That weakens `IN` and, after negation, STRENGTHENS `NOT IN`: ```sql CREATE TABLE t (s String) ENGINE = MergeTree ORDER BY s PARTITION BY s; INSERT INTO t VALUES ('a'), ('b'), (''); SELECT s FROM t WHERE s NOT IN (SELECT 'a' UNION ALL SELECT NULL) ORDER BY s SETTINGS transform_null_in = 1; -- returned only 'b': the '' row was pruned, because the dropped NULL had been folded to '' ``` The row is now dropped instead, which is result-neutral: such a `NULL` can never match a key reaching this block, whose outer type is never `Nullable` there. **3. Exactness was not gated on the conversion.** Index preparation casts set values INTO the key type while runtime membership casts the KEY into the set's type, and `castColumnAccurateOrNull` only proves nothing overflowed. Both are now checked: a set-to-key cast that is not equality-preserving marks the atom relaxed, and a key-to-set cast that can collapse two distinct keys onto one set value makes it DECLINE, since relaxation only forces `can_be_false`. **Scope.** That predicate detects the three classes above, not every possible one, as its header comment says. Not closed here: a key type with no strict round-trip check at all, because `accurate::convertNumeric(strict = true)` needs an integer or float SOURCE and here the source is the KEY, so a `DateTime` key against a `UInt8` element still collapses as on master. Closing it by inverting the default was measured and rejected: it costs three of this PR's own liveness controls for zero correctness gain, since collapse depends on the conversion DIRECTION rather than the key type. Four residual shapes are tracked separately; none is regressed here. Validated by 42 labelled cases in `04545_transform_null_in_non_nullable_key`, each asserting the result plus, via `EXPLAIN indexes = 1`, the set size and pruned part count; seven are controls. Two `03733` reference lines move as the smaller-but-still-exact set is built.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112890",
        "createdAt": "2026-08-01T09:23:38Z",
        "updatedAt": "2026-08-13T12:41:51Z",
        "timestamp": "2026-08-13T12:41:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "v26.5-must-backport"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112921",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Evaluate randomHadamardTransform once for a constant vector",
        "text": "`randomHadamardTransform` did not use the default implementation for constant arguments, so it called `convertToFullColumnIfConst` on its first argument and ran the transform for every row of the block even when the vector was a constant. It also returned a plain `ColumnArray` rather than a `ColumnConst`, so the analyzer could not fold the call to a literal either (`resolveFunction.cpp` only folds when the executed column is a `ColumnConst`). This matters for the intended vector search usage, where the query vector has to be rotated the same way as the stored vectors: ```sql WITH randomHadamardTransform([...]) AS target SELECT id FROM t ORDER BY cosineDistanceTransposedQuantized(vec_quantized, target, 2, 384) LIMIT 100 ``` On 2 million 768-dimensional vectors (`QBit(Int8, 768, 128)`, 2 bits and 384 dimensions, 64-core machine, warm cache): | | before | after | |---|---|---| | transform written inline in the query | 2.744 s | **0.144 s** | | transform hidden in a scalar subquery (computed once) | 0.148 s | 0.126 s | The whole difference was the transform of the constant query vector being repeated for every row; after the change the inline form costs the same as the hand-rolled workaround. The fix sets `useDefaultImplementationForConstants` and declares `seed` and `output_dims` in `getArgumentsThatAreAlwaysConstant`, so an all-constant call is executed on a single row and wrapped in a `ColumnConst`. The constant check for `seed` and `output_dims` now comes from the base class, with the same `ILLEGAL_COLUMN` error code as before. Related: https://github.com/ClickHouse/ClickHouse/issues/103466 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `randomHadamardTransform` of a constant vector is now evaluated once instead of once per row. This speeds up vector search queries that rotate the query vector with `randomHadamardTransform` before comparing it against a `QBit` column.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112921",
        "createdAt": "2026-08-01T15:44:42Z",
        "updatedAt": "2026-08-13T01:17:17Z",
        "timestamp": "2026-08-13T01:17:17Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-performance",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced",
          "v26.7-must-backport"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112930",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Improve canceling queries with the `url` function",
        "text": "Resubmission of https://github.com/ClickHouse/ClickHouse/pull/104089 by Roman Vasin (@rvasin) into a branch of the main repository, with the review feedback from @Algunenano addressed. The original pull request is closed as superseded. Related: https://github.com/ClickHouse/ClickHouse/pull/104089 Related: https://github.com/ClickHouse/ClickHouse/pull/102801 Related: https://github.com/ClickHouse/ClickHouse/pull/104014 A query which reads over HTTP does not react to cancellation while `ReadWriteBufferFromHTTP` is retrying a request: `KILL QUERY`, `max_execution_time` or a disconnected client take effect only after all `http_max_tries` attempts and the backoffs between them are over, which is minutes with the default settings. ### What the change does {#what-the-change-does} - `doWithRetries` waits for the backoff on a cancellation flag instead of sleeping, so a cancellation interrupts the wait instead of being noticed after it has expired. The flag (`ReadWriteBufferFromHTTP::Cancellation`, a one-shot flag with a condition variable) is owned by `StorageURLSource` and set from its `cancel`. - After the backoff, and at the terminal exit of the loop - the attempt after which there is nothing left to retry, or one which failed with a non-retriable error - the loop asks the query status whether the query is still alive, via `CurrentThread::checkIfNotCancelled`, exactly as `Client::HeadObject` already does it for S3. A killed or timed out query is therefore reported with its own proper error (`QUERY_WAS_CANCELLED`, `TIMEOUT_EXCEEDED`, or the exception recorded for a disconnected client) instead of the network error that happens to be at hand. This part needs no plumbing, so it works for every user of the buffer, including those that pass no cancellation flag (the `web` disk, HTTP dictionaries, the data lake catalogs). - When the read is cancelled for another reason - the pipeline is being torn down because something else in the query has already failed, or the client has disconnected - the loop stops retrying and rethrows the error of its last attempt, which is the very exception the caller would have got once the attempts were exhausted. - `StorageURLSource::cancel` wakes the backoff for every cancellation reason - no one is left to wait for the remaining attempts. A cancellation after which the query must still succeed with what it has already read - a soft `max_execution_time` with `timeout_overflow_mode = 'break'` (`CancelledByTimeout`, which only ever comes from `PipelineExecutor::checkTimeLimitSoft`; a timeout with the `throw` overflow mode kills the query and arrives as `CancelledByUser`), or a consumer that has enough data (`PartialResult`) - is remembered, and `StorageURLSource::generate` then discards the error of the interrupted read and ends the stream, so the query returns its partial result instead of failing with the HTTP error. Only the interruption of the read is discarded: the sites which report an error *because* of the cancellation - the retry loop when it stops retrying, the failover loop when it stops probing the options - mark it (`Cancellation::markReadInterrupted`), and an error the cancellation has nothing to do with, for example a parse error of the data that had already been downloaded when it arrived, fails the query the same way it would with no cancellation at all. What is remembered is the *effective* kind of the cancellation: `ExecutingGraph::cancel` upgrades a `PartialResult` cancellation to the reason of a later hard one - for example, the first cancel of a client with `partial_result_on_first_cancel` followed by a `KILL QUERY` - and after the upgrade the read is not soft anymore, so nothing is discarded or synthesized - and nothing built under the soft state escapes either: an error of the interrupted read that was already in flight when the upgrade landed is suppressed rather than rethrown, because under a hard cancellation the query fails for a reason of its own - the error of the failed peer, the kill reported by the process list, or a disconnected client with no one left to report to - and the stale error must not mask it. - A cancellation stops `StorageURLSource::initialize` even where its helpers swallow the errors of the requests they make: `getFirstAvailableURIAndReadBuffer` rethrows the error of a cancelled read instead of probing the next failover option (and no longer probes them for a killed query at all, including a query killed right after its last option has failed, which is reported as cancelled instead of with the aggregate network error, and a cancellation which lands where no request is in flight for it to interrupt - between the options, or after the last one has failed - stops the choosing of the URI with the same two outcomes, a discarded cancellation error for a soft one and no buffer at all for a hard teardown, so the aggregate error can never stand in for the failure that really happened), and `ReadWriteBufferFromHTTP::tryGetFileSize` / `tryGetLastModificationTime` rethrow it instead of treating the interrupted `HEAD` request as a file without metadata (the rethrow is keyed off the mark the interrupted request leaves - `Cancellation::markReadInterrupted` - not off the cancellation flag at the moment the fallback runs, so an error that had already happened when the cancellation arrived stays swallowed, and a soft cancellation cannot turn a metadata failure the query would have survived into a failure of the query). `generate` re-checks `isCancelled` after `initialize`, so it never pulls a chunk no one needs. `initialize` itself ends the stream when the cancellation arrives after the URI has been chosen but before the metadata of the file - the modification time and the size - is requested, so a cancelled source does not start a fresh metadata `HEAD` request. The check is repeated between the two metadata probes: when the shared `HEAD` request has failed with a network error before the cancellation arrived, the probe of the modification time has nothing to remember, and the probe of the file size would otherwise send a fresh `HEAD` of its own. ### How the review feedback is addressed {#how-the-review-feedback-is-addressed} - *\"it is really easy to misuse, because you don't get any feedback on the thread that was cancelled, and exit normally. A simple look at `ReadWriteBufferFromHTTP::readBigAt` shows how trivial is now to hit asserts due to a pipeline cancellation.\"* - there is no silent `break` anymore: `doWithRetries` either does its work or throws, so a caller cannot mistake a cancellation for a success. The `chassert` in `readBigAt` stays untouched, `nextImpl` cannot report a cancellation as the end of the stream, and the row count cache cannot be filled from an interrupted read. Consequently `StorageURLSource` needs none of the `isCancelled()` checks the original pull request added after each buffer operation - which also answers *\"I don't understand the pattern of doing the HTTP request, and then checking if the pipeline is cancelled. Shouldn't it be the other way around?\"* and *\"Why not check at the start of the loop?\"*. - *\"part of the confusion comes from dealing with query cancellation and pipeline cancellation as the same thing, when they aren't, and how to report this properly both in logs and to the end user.\"* - the two are now separate. Query cancellation is taken from the query status and reported as a cancellation. Pipeline cancellation only stops the retrying and reports the HTTP error that really happened, so it cannot race the exception of a peer source with a wrong message and URI - which is what broke the two reverted attempts. Both cases are logged with the reason at the point where the loop gives up. - *\"We don't propagate any errors here, so I don't understand why we skip the first attempt.\"* - the `attempt > 1` condition is gone. The check now sits where the loop decides to wait and try again, which by construction can only be reached after an attempt has failed, so the error of a genuine first failure is never masked. - *\"we are sleeping without checks, that is, a cancel does not wake up the thread, which means we wait just to exit. We should use a condition variable and wait on it instead.\"* - done, see above. - *\"slow request won't be dealt with, we'll wait until the request either finishes or times out\"* - still not addressed, as agreed in that review. A request that neither answers nor times out is not interruptible; that needs closing the socket from the outside and is left for a separate change. ### Test {#test} `04615_kill_query_url_function` starts a server that always answers `503`, so that the query really is inside the retry loop - the test in the original pull request always answered `200` and never reached a second attempt. With 30 attempts and a backoff of 1 to 2 seconds, for both the `HEAD` and the `GET` request, the query would retry for minutes; the test kills it and checks that the client stops within seconds and reports an error. It also checks that the same query, when it is not killed, still reports the `503` of the server as before. Verified locally against a build of this branch: the killed query ends 3 ms after the `KILL` and is reported as `Code: 394. DB::Exception: Query was cancelled`. On an unpatched server the same query stays in `system.processes` with `is_cancelled = 1` and keeps retrying for more than five minutes. `04691_url_function_partial_result_on_break_timeout` covers the soft timeout: a glob over a file that the server serves completely and a URL that it always answers with `503`, with a backoff of 1 to 2 seconds over 30 attempts. A query with `max_execution_time = 1` and `timeout_overflow_mode = 'break'` succeeds with the rows of the first file in about a second, while the same query without the timeout still reports the `503` of the server. The partial result of the `break` mode consists of the rows already streamed to the consumer - a cancelled aggregation returns nothing even for regular tables - so the test asserts the streamed rows. `04759_url_function_cancel_during_initialize` covers a cancellation during initialization, whose helpers swallow the errors of the requests they make. Its server counts the requests to each path: after a soft `break` timeout interrupts the retries of the metadata `HEAD` request the data is never downloaded, and after it interrupts the retries of the first failover option of an `a|b` URL the second option is never probed - while an uninterrupted query still treats the failing metadata request as non-fatal and reads the data. Both assertions fail on the code before the fix. `04811_url_function_no_count_cache_poisoning_on_break_timeout` covers the row count cache (`use_cache_for_count_from_files`): a read of a slowly streaming URL interrupted by a soft `break` timeout leaves no entry in `system.schema_inference_cache`, while a complete read still caches the correct row count. `generate` records the count only when the read genuinely reached the end of the file - it checks the final status of its `PullingPipelineExecutor` and `isCancelled` on the source, as `StorageMemory` mutations do - so a cancelled read can never record the rows it happened to read as the row count of the file. `04812_url_function_no_fallback_probe_after_soft_cancel` covers a cancellation which lands *between* the failover options, where there is no request in flight for it to interrupt: an `empty|data` URL with `engine_url_skip_empty_files = 1`, where the empty file takes longer to be served than the soft `break` timeout of the query. Once the empty file is skipped, the loop checks the cancellation flag before constructing and probing the next option, so the query succeeds with no rows and the data URL receives no requests - the assertion fails on the code before the fix. An uninterrupted query still skips the empty file and reads the next option. This between-options check is reason-aware: it reports the interruption with a cancellation error - which `generate` then discards - only for the soft cancellations, after which the query must still succeed (`CancelledByTimeout` of the `break` overflow mode, or `PartialResult`). A pipeline torn down because something else in the query has already failed, or a disconnected client, must not be reported with a fabricated cancellation that could reach the user in place of the failure that really happened - and needs no synthetic error at all: `getFirstAvailableURIAndReadBuffer` returns no buffer and `initialize` ends the stream, since nobody is left who needs the data. `04824_url_function_kill_after_partial_result_cancel` covers a hard cancellation arriving after a soft one: the first cancel of a client with `partial_result_on_first_cancel` - after which the query must still succeed with its partial result - is followed by a `KILL QUERY`, both landing while the source is blocked in a request that a cancellation cannot interrupt (an empty file ahead of a failover option, whose response the test server withholds until the test releases it after the `KILL` has returned - the kill delivers the cancellation to the processors synchronously, so no timing can release the source early). The killed query must fail with the cancellation error instead of discarding it as if its result were partial, and must not probe the next failover option - the discard assertion fails on the code before the fix. The reason the executor delivers may understate a kill: its `checkTimeLimitSoft` poll observes a killed query as a soft timeout, and when that poll wins the race for the one-shot cancel reason of `ExecutingGraph`, the kill's own hard broadcast never reaches the source - so before discarding, the source additionally asks the process list, and a killed query fails with the proper cancellation error under either delivery order. `04825_url_function_no_next_option_after_disconnect` covers a hard teardown which does not kill the query, landing between the failover options: the client of a query blocked in the request for a held empty first option disconnects (`kill -9`), the test waits until the delivery of the resulting cancellation to the source is visible in the log (`StorageURLSource::cancel` leaves a debug trace of every delivered cancellation and its reason), and only then lets the server answer the held request. The source finds the empty file, and must end the stream instead of probing the next option, which would succeed - the assertion fails on the code before the fix. The tests which kill their query or assert on its log pin `parallel_replicas_for_cluster_engines` off: the rewrite of `url` to `urlCluster` would move the source into remote queries with their own query ids. `04829_url_function_no_metadata_requests_after_cancelled_head` covers the fallback for servers which do not support `HEAD`: `ReadWriteBufferFromHTTP::getFileInfo` treats a non-retriable 4xx response as \"the server cannot answer this\" and reports no metadata - and used to do so even when the request had been interrupted by a cancellation, so the initialization completed as if the file simply had no metadata instead of failing with the error of the interrupted read, which `generate` discards or fails with depending on the kind of the cancellation. The test answers the metadata `HEAD` with `400` only after a soft `break` timeout has been delivered to the source (the request-counting server withholds the response until the delivery is visible in the log): the query succeeds with its empty partial result, no request follows the cancelled `HEAD`, and the log records that the interrupted read was reported and its error discarded - the last assertion fails on the code before the fix. `04837_url_function_killed_query_after_last_attempt` covers the exit of the retry loop after the attempt which has nothing left to retry. Its reader is the schema inference of the `url` table function, which passes no cancellation flag, so the query status is the only thing that can tell the read that the query is gone. The server withholds its `503` response until the test has killed the query, and `http_max_tries` is 1, so the read is interrupted exactly where the retrying ends: the query must fail with `QUERY_WAS_CANCELLED`, while the same query which is not killed must still report the `503` of the server. On the code before the fix the killed query is reported with the `503` instead - which is the assertion the test makes. `04843_url_function_no_metadata_probe_after_soft_cancel` covers the window after the URI has been chosen and before the metadata of the file is requested: the server withholds the response of the first failover option until a soft `break` timeout has been delivered to the source (visible in the log), and then serves the file without a `Content-Length`, so the modification time and the size could only come from a `HEAD` request. The cancelled source must end the stream instead of probing the metadata no one is left to read: the query succeeds with its empty partial result and no `HEAD` request follows - the assertion fails on the code before the fix. `04844_url_function_parse_error_after_soft_cancel` covers an error which the cancellation has nothing to do with: the server streams a few good rows, withholds the rest until a soft `break` timeout has been delivered to the source, and only then sends a malformed row. The parse error must fail the query even though the soft cancellation - after which the query would otherwise succeed with its partial result - has already been latched: the file is malformed no matter when the query stopped wanting more of it. On the code before the fix the query succeeds, and the malformed row is discarded together with the interruption of the read. `04846_url_function_no_second_metadata_probe_after_cancel` covers a cancellation which lands between the two metadata probes of `initialize`: the server tears down the metadata `HEAD` request without a response - a network error which leaves the probe of the modification time with nothing to remember - and the source is held in the window between the probes with a failpoint (`storage_url_pause_between_metadata_probes`, added for the test: the window is a few instructions wide) until a soft cancellation has been delivered to it - the client is cancelled with SIGINT under `partial_result_on_first_cancel = 1`, so the delivery is under the test's control and cannot race the query the way a short `max_execution_time` did (which made the first version of the test flaky under the sanitizer builds) - and the delivery is visible in the log. The released source must end the stream instead of letting the probe of the file size send a second `HEAD` request: the query succeeds with its empty partial result and the server sees exactly one `HEAD` - the assertion fails on the code before the fix. `04869_url_function_stale_metadata_error_after_soft_cancel` covers the opposite side of the same fallbacks: a metadata `HEAD` request which fails on its own *before* any cancellation arrives must stay non-fatal even when a soft cancellation lands while its failure is being unwound. The server tears down the `HEAD` request without a response, the source is held inside the fallback with a failpoint (`http_read_buffer_pause_before_metadata_fallback`, added for the test: the window between the failed request and the fallback is a few instructions wide), the test delivers the soft cancellation (SIGINT with `partial_result_on_first_cancel = 1`, visible in the log) and only then releases the source. The fallback must swallow the error of the request the cancellation did not interrupt: the query succeeds with its empty partial result and the data is never requested - on the code before the fix the query fails with the stale network error of the `HEAD` request. `04871_url_function_hard_cancel_upgrade_after_interrupted_read` covers the upgrade of a soft cancellation to a hard one landing in the last window: after the error of the interrupted read has been thrown under the soft state, but before `generate` has handled it. One source retries an always-failing URL and its backoff is woken by a soft cancellation (SIGINT with `partial_result_on_first_cancel = 1`); the thrown error is held in the window with a failpoint (`storage_url_pause_before_handling_interrupted_read_error`, added for the test: the window is a few instructions wide). A second source, held by the test server until now, is then released into a parse error - a real failure the cancellation has nothing to do with - which cancels the pipeline hard and upgrades the paused source's reason to `Exception`; the test waits until both deliveries are visible in the log, so the upgrade deterministically lands inside the window, and only then releases the source. The stale error of the interrupted read must be suppressed and the query fails with the parse error of the peer - on the code before the fix it fails with the stale HTTP error of the read its own cancellation interrupted. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `KILL QUERY`, query timeouts and client disconnects now stop a query that reads over HTTP (for example over the `url` table function) while it is retrying a request, instead of waiting until all the retry attempts are exhausted.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112930",
        "createdAt": "2026-08-01T16:30:32Z",
        "updatedAt": "2026-08-13T10:22:51Z",
        "timestamp": "2026-08-13T10:22:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 22
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112932",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reimplement the KQL (Kusto) dialect on a lexer, an AST and AST translation",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/61742 --> ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reimplemented the experimental KQL (Kusto) dialect. It now has its own lexer and parser and builds an AST directly, instead of translating a token stream into SQL text and reparsing it. This fixes an expression-injection hole in the string operators (`contains`, `has`, ...), several results that silently disagreed with Kusto (`7 / 2`, `substring` with a negative start, `bin`), and 28 functions that were registered but did nothing. The supported subset is smaller and documented; anything outside it is now rejected by name instead of being mistranslated. ### Documentation entry for user-facing changes `docs/guides/clickhouse/kusto-query-language.mdx` --- ## Why The dialect was contributed in 2022 and abandoned by its authors; the rewrite they promised in #62668 was closed unmerged. Since then it has been emergency-disabled once (#59305) and gated behind `allow_experimental_kusto_dialect` (#74224), and #61742 (\"Remove KQL support if this code will not be fixed\") has stayed open. The implementation translated a KQL token stream directly into ClickHouse SQL **text** and reparsed it — 298 `fmt::format` sites, 41 places that re-lexed generated text, and no AST anywhere in the function layer. Everything below follows from that: | | before | after | |---|---|---| | `T \\| where s contains \"x') OR 1 = 1 OR ilike(s, 'y\"` | injected the expression | needle is an `ASTLiteral`; no injection is representable | | `'50x' contains '50%'` | `true` — needle was pasted into a LIKE pattern | `false` | | `substring('abcdefg', -3, 2)` | `'ab'` | `'ef'` | | `7 / 2` | `3.5` | `3` | | `format_timespan(col, 'hh:mm')` | `stoi: no conversion` — `std::stoi` on generated SQL | rejected by name | | `series_fir(...)` | `Function series_fir does not exist` (one of 28 stubs) | rejected by name | | `search 'x'` | `Unknown table expression identifier 'search'` | `'search' is not a supported KQL operator` | | `let` bindings | `static thread_local`, leaked between queries | scoped to one parse | ## What is here Two commits: the removal, then the reimplementation. ``` src/Parsers/Kusto/ KQLLexer KQL's own tokens: !in, =~, .., timespans (2.5h), verbatim @'...', datetime(...). A bad literal is an Error token carrying a reason, so nothing downstream asks isValidKQLPos() — the core of #61742. KQLAST the tabular level only. Scalar expressions are ClickHouse AST directly; a parallel expression hierarchy would add only conversions. KQLParser recursive descent. `let` bindings are an ordinary member. KQLTranslator each operator fills a still-empty clause of the select being built, or wraps what exists so far in a subquery and starts a new one. KQLFunctions name -> builder returning an ASTPtr. src/Functions/Kusto/ kqlDivide `7 / 2` is 3 in KQL and 3.5 in SQL — the choice depends on operand kqlBin types, so it is made by an IFunctionOverloadResolver during analysis rather than guessed from how the argument was spelled. ``` KQL now has its own entry point (`parseKQLQuery`) and never touches the SQL tokenizer. Because `ClientBase` already catches, the parser can simply throw — so there is no `tryParseKQLQuery`, and no `catch` anywhere in the new code. **`src/Parsers/Kusto` no longer needs its carve-out from check 19** (\"do not catch exceptions in src/Parsers\"), and this PR deletes it. `src/Parsers/Kusto` goes from 13017 lines to 3967, plus 337 lines of runtime functions. ## Scope Deliberately smaller than before, and written down in the guide. A construct is either translated with the semantics Kusto documents, or rejected by name. Rejected in this PR: `search`, `parse`, `mv-apply`, `lookup`, `evaluate`, `invoke`, `facet`, `top-nested`, `make-series` and friends; the `series_*`, `bag_*`/`pack_*` and `ipv4_*` families; `parse_url`, `parse_csv`, `parse_json`, `toscalar`, `format_timespan`; and `dynamic` **objects** (arrays map onto ClickHouse `Array`). Three known divergences are documented rather than papered over: subtracting two datetimes yields seconds instead of a timespan, `project-rename` moves the renamed column to the end, and `union` needs compatible schemas. ## Testing `04670`–`04673` cover the pipeline, the string operators (including the injection cases), the scalar semantics that used to be wrong, and 47 constructs that must be rejected. A 2500-query random fuzz found no crashes or logical errors. The 302 conformance cases that were deleted with the old implementation are being reintroduced separately, on top of this branch.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112932",
        "createdAt": "2026-08-01T16:56:06Z",
        "updatedAt": "2026-08-13T03:08:13Z",
        "timestamp": "2026-08-13T03:08:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 17
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112940",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reapply \"Make `postgresql` and `PostgreSQL` engine work against a ClickHouse instance\", with a fix for the cancel-request logical error",
        "text": "Reapply #110760, reverted in #112935, together with a fix for the bug that motivated the revert. Related: https://github.com/ClickHouse/ClickHouse/pull/110760 Related: https://github.com/ClickHouse/ClickHouse/pull/112935 Related: https://github.com/ClickHouse/ClickHouse/issues/84085 ### Why it was reverted, and why the bug is not in the reverted change Every `Stress test` job on master started failing with ``` Logical error: 'Query context must be created after authentication' DB::Session::makeQueryContextImpl @ src/Interpreters/Session.cpp:690 DB::PostgreSQLHandler::cancelRequest @ src/Server/PostgreSQLHandler.cpp:867 DB::PostgreSQLHandler::startup @ src/Server/PostgreSQLHandler.cpp:737 ``` CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=cd6fd72e2b29fa607cffe0595e347452f8dcba9b&name_0=MasterCI&name_1=Stress%20test%20%28amd_debug%29 A PostgreSQL client cancels a running statement by opening a *second* connection and sending a `CancelRequest` on it, carrying the `(process id, secret key)` pair the server handed out in `BackendKeyData`. By the protocol that connection never authenticates - the secret key is the credential - so `PostgreSQLHandler::cancelRequest` has no authenticated session, and the `session->makeQueryContext()` it called throws `LOGICAL_ERROR`. That has been true since sessions were introduced (`51ffc334573`), and it is reachable by any PostgreSQL client that cancels a statement: `psql` on Ctrl-C, or the `pgx` driver in issue #84085, which reports this very error. #110760 did not introduce it, it only made CI reach it: with ClickHouse acting as a libpq/pqxx client against itself, `PostgreSQLSource` cancels the remote statement through `PQcancel`, so a stress run that cancels a query now sends a cancel request to the server's own PostgreSQL port. ### The fix (second commit) - `cancelRequest` cancels the query through the process list directly, which needs no session, via a new `ProcessList::sendCancelToQueryOfAnyUser`. Only queries whose id has the `postgres:<connection id>:<secret key>` shape the server itself assigns can be named this way, and the secret key is what makes such an id unguessable - the same credential PostgreSQL relies on. - The raw `KILL QUERY` result the old code wrote straight into the client socket is gone with it. The protocol expects no answer at all to a cancel request, so that was protocol garbage. - The secret key never matched anything either: `BackendKeyData` was sent while `secret_key` was still zero, and every statement then re-randomized it, so the query id a cancel request resolved to was never the id of a running query - cancellation over the PostgreSQL protocol has never worked. The key now belongs to the connection, as in PostgreSQL, where it identifies the backend rather than one statement, and every statement of the connection runs under `postgres:<connection id>:<secret key>`. Statements of one connection run one after another, so reusing the id is safe. - New test `04669_postgresql_protocol_cancel_request`, which drives a real `psql` and cancels it with `SIGINT` - that is what makes `psql` send a `CancelRequest`. The `-- ping` half of issue #84085 (a comment-only statement is not recognized as an empty query) is not addressed here. ### Privileges A self-connect needs nothing beyond access to the table being read. `system.databases`, `system.tables` and `system.columns`, which the emulated `pg_namespace`, `pg_class` and `pg_attribute` are views over, are readable by every user even with `access_control_improvements.select_from_system_db_requires_grant` enabled, and their rows are filtered by the reader's own grants - so the catalog shows a user exactly the relations that user may read, as `pg_catalog` does in PostgreSQL. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `postgresql` table function and `PostgreSQL` table engine can now be used to connect to another ClickHouse server over the PostgreSQL protocol (when a table name is used; the `query(...)` variant is not supported yet). Added the PostgreSQL-compatibility functions `format_type` and `current_setting`. Fixed cancellation over the PostgreSQL protocol: a cancel request used to log a logical error and never cancelled anything.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112940",
        "createdAt": "2026-08-01T18:51:27Z",
        "updatedAt": "2026-08-13T02:34:26Z",
        "timestamp": "2026-08-13T02:34:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 20
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112945",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use `pread` when `preadv2` with `RWF_NOWAIT` cannot be used, and recognize `EPERM` from it",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/104634 Related: https://github.com/ClickHouse/ClickHouse/issues/49149 Related: https://github.com/ClickHouse/ClickHouse/issues/39753 The `pread_threadpool` read method hands every read off to a thread pool, unless the data is already in the page cache, which it checks with the `preadv2` system call and the `RWF_NOWAIT` flag. Two things can go wrong with that check, and both were handled badly. **`EPERM` was not recognized.** It is what a `seccomp` profile of a container runtime answers for a system call that is not in its allow list. `ThreadPoolReader::submit` handed the read off to the thread pool for `ENOSYS` and `EOPNOTSUPP`, but let `EPERM` through to the throw, failing the query with `CANNOT_READ_FROM_FILE_DESCRIPTOR`. This is the signature reported in #49149. **The check was never verified in advance.** `hasBugInPreadV2` only compared the kernel version, and nothing else was checked, so on a system where the check cannot work `pread_threadpool` kept paying for a thread pool hand-off on every read - including the reads that only had to copy the data from the page cache. In #104634, on Amazon Linux 2 (kernel 5.10), the profile shows `ThreadPoolReaderPageCacheMiss` 1,971,514 out of `LocalThreadPoolJobs` 1,972,484 - every read went to the pool - while the device only moved ~22 GB of the 258 GB the file descriptor delivered, i.e. ~90% of the data was in the page cache and still paid for the hand-off. The system call is now probed once, before it is used, by `preadNoWaitUnavailableReason`. The probe passes an invalid file descriptor on purpose: `seccomp` filters and the system call table are consulted before the descriptor is looked up, so an available system call answers `EBADF` without reading anything, while a blocked one answers `EPERM` or `ENOSYS`. When it says the check cannot be used, `applySettingsQuirks` switches the default value of `local_filesystem_read_method` from `pread_threadpool` to `pread` at start time, and says why in the server log. Nothing downstream has to know: the reader, the userspace page cache eligibility in `DiskLocal::prepareRead` and the prefetched read pool all see a plain `pread` setting. As with the other settings quirks, a value set explicitly - in the configuration, or with `SET` at runtime - is left alone. Such a session keeps `pread_threadpool` and keeps paying for the hand-off, which is what happens today. The switch is a property of the host, so the setting is left marked as unchanged: only the changed settings are serialized into the query the initiator sends to the remote shards, and a host that cannot use the system call must not impose `pread` on the shards that can. An explicitly requested value stays changed and is still sent. The per-read `errno` handling in `ThreadPoolReader::submit` is kept, with `EPERM` added to it: the probe answers for the system call, but a particular filesystem can still reject the flag (`tmpfs` answers `EOPNOTSUPP`, for example), and such a read is handed off to the thread pool instead of failing the query. ### How it was tested `preadv2` was rejected the way a container runtime does it, with a `seccomp` filter installed by a small wrapper (`SECCOMP_RET_ERRNO`), and an old kernel was simulated with `setarch --uname-2.6`. Reading a 2 million row `MergeTree` table with the default `local_filesystem_read_method`, before (the released 26.7.1 binary) and after: | | before | after | |---|---|---| | no filter | `ThreadPoolReaderPageCacheHit` 39, `LocalThreadPoolJobs` 90 | `ThreadPoolReaderPageCacheHit` 34, `ThreadPoolReaderPageCacheMiss` 3, `LocalThreadPoolJobs` 93 - the read method stays `pread_threadpool` | | `preadv2` → `EPERM` | `Code: 74 ... errno: 1, Operation not permitted (CANNOT_READ_FROM_FILE_DESCRIPTOR)`, already while attaching the table | the read method is `pread`, the query succeeds, no `ThreadPoolReader` events | | `preadv2` → `ENOSYS` | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 - every read to the pool | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 80 | | kernel reported as older than 5.11 | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 81 | In all three rejected cases the reason is in the log, for example: ``` <Warning> SettingsQuirks: The default value of local_filesystem_read_method has been switched from 'pread_threadpool' to 'pread' (you can explicitly set it back still), because the `preadv2` system call is not available (the probe with an invalid file descriptor answered errno: 1, strerror: Operation not permitted instead of `EBADF`); it is typically rejected by a `seccomp` profile of a container runtime, and can be allowed in the runtime configuration. ... ``` An explicitly requested `local_filesystem_read_method = 'pread_threadpool'` is kept, on every one of those systems, and this is where the per-read `EPERM` handling earns its place: under the `EPERM` filter the same query fails on the released binary and succeeds here, with `ThreadPoolReaderPageCacheMiss` 45 and `LocalThreadPoolJobs` 135 - every read handed off to the pool, which is the documented cost of asking for it there. Under the same filters, `system.settings` reports `local_filesystem_read_method = 'pread'` with `changed = 0`, so nothing is forwarded to the remote shards, while an explicitly requested `pread_threadpool` reports `changed = 1` and is still sent. Unit tests cover the `errno` classification, the probe's `EBADF` contract, and the quirk itself (the default is switched exactly when the probe says the system call is unusable, the switched value stays out of `Settings::changes()`, and an explicitly set value is never switched). An automated end-to-end test would need an instance with a restrictive `seccomp` profile, which the integration test framework cannot express today - it starts every instance with `seccomp:unconfined`. ### Documented behavior impact On a system where the page cache cannot be checked without waiting for the disk - Linux older than 5.11, a sandbox that rejects `preadv2` with an error code, and systems other than Linux, where `preadv2` does not exist and the method never checked the page cache in the first place - the default value of `local_filesystem_read_method` becomes `pread` instead of `pread_threadpool`, so local reads are performed in the calling thread. This includes the reads with `O_DIRECT`, which never look at the page cache and do not need the check; the read method is now resolved once, for the whole server, so they follow the same value. Nothing changes on a supported system, and nothing changes for an explicitly configured read method. The setting description in `src/Core/Settings.cpp`, from which the docs are generated, is updated in this PR. A `seccomp` profile that terminates the process instead of rejecting the system call with an error code cannot be detected: the startup probe is itself a `preadv2` call, so such a profile kills the server there. On `master` it kills it at the first read instead, under the same default read method. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `pread_threadpool` read method needs the `preadv2` system call with the `RWF_NOWAIT` flag to read the data that is already in the page cache without handing the read off to a thread pool. It is now checked at start time whether that system call can be used, and if it cannot - the Linux kernel is older than 5.11, or a `seccomp` profile of a container runtime rejects the system call - the default value of `local_filesystem_read_method` is switched to `pread`, and the reason is reported in the server log. Previously, every read paid for a thread pool hand-off on such systems, and a `seccomp` profile that answers `EPERM` made queries fail with `CANNOT_READ_FROM_FILE_DESCRIPTOR`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112945",
        "createdAt": "2026-08-01T20:09:06Z",
        "updatedAt": "2026-08-13T17:32:00Z",
        "timestamp": "2026-08-13T17:32:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 26
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112950",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support the Vortex file format",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/87327 Related: https://github.com/ClickHouse/rust_vendor/pull/74 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added support for reading and writing the [Vortex](https://github.com/vortex-data/vortex) columnar file format (the `Vortex` input and output format). This closes [#87327](https://github.com/ClickHouse/ClickHouse/issues/87327). ### Documentation entry for user-facing changes The implementation uses the Rust `vortex` crate (v0.83.0) through a new C FFI crate `rust/workspace/vortex` (`_ch_rust_vortex`), following the same pattern as `prql` and `polyglot`. Data crosses the FFI boundary through the Arrow C Data Interface and is converted with the same `ArrowColumnToCHColumn`/`CHColumnToArrowColumn` code as the `Arrow` format. IO is delegated back to ClickHouse through callbacks: reads go through ClickHouse's own read buffers (range reads for seekable inputs, whole-file buffering otherwise), and the produced file is streamed into the output buffer. All work is driven by a single-threaded runtime on the calling thread — the library spawns no threads, and Rust panics are caught at the FFI boundary and turned into exceptions. Features: - Reading with projection pushdown: only the columns used by the query are read from the file. - Schema inference and `count()`-only queries answered from file metadata without reading data. - Writing with the library's default adaptive compression (BtrBlocks-style cascading encodings + zstd), including a valid empty file for empty results. - Graceful errors on malformed and truncated files (fuzzer-friendly: no aborts, Rust panics become exceptions). Limitations (documented in `docs/reference/formats/Vortex.mdx`): - `Map`, `Int128`/`UInt128`/`Int256`/`UInt256`, `IPv6`, and `Interval` columns cannot be written (no corresponding Vortex type). - `String` and `FixedString` are written as Vortex `Binary` (ClickHouse strings are arbitrary bytes, while Vortex requires `Utf8` to be valid UTF-8). - The format is disabled in MSan builds: the MSan-instrumented library (with origin tracking) is so large that linking `unit_tests_dbms` overflows the 2 GiB `R_X86_64_PC32` relocation range (same approach as `wasmtime` and `delta-kernel-rs`). - Reading and writing are single-threaded in this first version: the whole scan (I/O, decompression, and decoding) runs on one thread, so on ClickBench reads are significantly slower than `Parquet`, which ClickHouse decodes with multiple threads (see the [benchmark results](https://github.com/ClickHouse/ClickHouse/pull/112950#issuecomment-5274283519) and [the explanation](https://github.com/ClickHouse/ClickHouse/pull/112950#issuecomment-5274483899)). Filter pushdown (added in https://github.com/ClickHouse/ClickHouse/pull/114373, `input_format_vortex_filter_push_down`, on by default) reduces the amount of data decoded by selective queries, but does not parallelize the scan. Writes are also slower than `Parquet` (the adaptive compressor samples many encodings per column) — parallelism can be added later. The 127 new vendored Rust crates are added in https://github.com/ClickHouse/rust_vendor/pull/74 (the `contrib/rust_vendor` submodule is bumped to that branch).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112950",
        "createdAt": "2026-08-01T21:39:00Z",
        "updatedAt": "2026-08-13T14:10:30Z",
        "timestamp": "2026-08-13T14:10:30Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "pr-feature",
          "submodule changed",
          "pr-autogenerated-docs"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112968",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix \"Not-ready Set\" for GLOBAL IN over an object-storage _path / _file filter",
        "text": "`SELECT ... FROM s3(...) WHERE _path GLOBAL IN (SELECT ...)` over a path without globs threw ``` Logical error: 'Not-ready Set is passed as the second argument for function 'globalIn'' ``` `ReadFromObjectStorageStep::applyFilters` intentionally leaves the sets of `globalIn` / `globalNotIn` unbuilt, so that `ReadFromRemote` can attach an external table to them first. Plan optimization then moves the subquery plan of such a set into a `CreatingSetsStep` (`DelayedCreatingSetsStep::makePlansForSets`), which means the set is only created once the pipeline runs. `StorageObjectStorageSource::createFileIterator` prunes an explicit key list while the pipeline is being built, and its `buildSetsForDAG` call is a no-op for a set whose subquery plan is already gone, so the pruning expression was executed with a set that was not ready yet. `GlobIterator` is unaffected: it applies the same filter while listing objects, when the set is already built. Pruning the keys is an optimization - the same predicate is applied by the `Filter` step above the source - so skip it when a set is not ready. `buildSetsForDAG` now reports whether every set in the DAG ended up ready. Reproducer (deterministic, fails on `master`, passes with this change): ```sql SELECT * FROM s3(s3_conn, filename = 'x', format = CSV, structure = 'x UInt64') WHERE _path GLOBAL IN (SELECT 'no such path'); ``` The stress test hit it through an AST fuzzer query over `test_s3_race`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=1c3181d592a4aaeec15077107e15b7edab5bca8d&name_0=MasterCI&name_1=Stress%20test%20%28arm_release%29 Related: https://github.com/ClickHouse/ClickHouse/issues/107619 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix the `Not-ready Set is passed as the second argument for function 'globalIn'` exception for a `GLOBAL IN` subquery over the `_path` or `_file` virtual column of an object storage table (`s3`, `azureBlobStorage`, ...) whose path has no globs.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112968",
        "timestamp": "2026-08-12T22:51:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:112973",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `sorted_merge` and `parallel_sorted_merge` join algorithms",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109005 Implements the two algorithms planned in [this discussion](https://github.com/ClickHouse/ClickHouse/pull/109005#discussion_r3694159098): `full_sorting_merge` and `parallel_full_sorting_merge` are always supported, so anything listed after them in `join_algorithm` is unreachable — listing them is an unconditional choice, not a preference. The new `sorted_merge` and `parallel_sorted_merge` algorithms execute the same merge join, but are **available only when both join inputs can be efficiently read in the order of the join keys** (e.g. MergeTree tables whose primary key starts with the join keys), so the pre-join sorts become cheap `FinishSorting` or disappear. When the tables' order cannot be exploited, the selection falls through to the next algorithm in the list. That makes them meaningful as a high-priority preference: `join_algorithm = 'sorted_merge,parallel_hash'` uses the streaming, low-memory merge join exactly when it is certainly beneficial, and a hash join otherwise. `sorted_merge` runs a single in-order merge join. `parallel_sorted_merge` additionally shards the join by ranges of the tables' common primary-key prefix into independent per-shard merge joins running in parallel — the same source-side sharding `query_plan_join_shard_by_pk_ranges` applies, enabled for this join by the algorithm itself; the in-order reads stay intact (no scatter, no re-sort). When the sharding cannot apply (e.g. an `ASOF` join), it degrades to a single `sorted_merge`. Implementation notes: - Eligibility is decided during plan physicalization, where the input subplans are visible: `JoinStepLogical::inputsCanBeReadInJoinKeyOrder` finds the `ReadFromMergeTree` below each input (mirroring `findReadingStep` of `optimizeReadInOrder`, without descending through nested joins) and probes the actual read-in-order matcher (`wouldReadInOrderBeUseful`, the same side-effect-free probe `topKThroughJoin` uses) with the join-key sort description. The predicted decision degrades gracefully in both directions: a false positive runs like `full_sorting_merge` (with a full sort), a false negative falls through to the next algorithm. - The same memoized predicate makes `tryAddJoinRuntimeFilter` keep its hands off an eligible join listed before the first hash-family algorithm. Without this, planting a runtime filter erases the merge algorithms from the list (a merge join reads both sides concurrently and cannot use a runtime filter), silently overriding the priority order — the defining feature of these algorithms. For non-eligible joins the runtime filter (and the erasure) stays, because those algorithms are not selectable anyway. When `applyParallelReplicas` later breaks the eligibility (a join input becomes a distributed read), the filter pass is re-run for exactly the joins it had skipped, so the `hash` fall-through gets its runtime filter back. - `FullSortingMergeJoin` now carries the selected algorithm (`getSelectedAlgorithm`) instead of an `is_parallel` flag; the hash-scatter rewrite (`optimizeParallelFullSortingMergeJoin`) stays exclusive to `parallel_full_sorting_merge`, and `optimizeJoinByShards` runs in a restricted mode (only `parallel_sorted_merge`-selected joins, with a cheap pre-scan bail-out) when `query_plan_join_shard_by_pk_ranges` is off. When the sharded stream counts diverge at pipeline-building time (e.g. a data-dependent `PREWHERE` prunes one side to a single empty stream), `JoinStep` merges each side's per-shard sorted streams back into one sorted stream and runs the single-stream merge join instead of failing (the same approach as #109393, applied at the sharding fallback). - The `CreateSetAndFilterOnTheFlyStep` pair (`max_rows_in_set_to_optimize_join`) is not added for sorted-merge joins: it sits between the read and the sort and would defeat the in-order read the algorithm was selected for. - The old analyzer has no query plan at selection time, so there the algorithms are never selected and the list falls through (documented). With only `sorted_merge` listed and no exploitable order, the query fails with `NOT_IMPLEMENTED`, like other unsupported single-algorithm configurations. - The known lower-priority-fallback side effects of listing merge algorithms (stricter `USING` key-type inference, `topKThroughJoin` deferral) extend to the new values and are documented in the `join_algorithm` setting description. Tests: `04669_sorted_merge_join_selection` pins the selection gating via `EXPLAIN PIPELINE` (selected on a primary-key join with no re-sort, falls through on non-key joins / disabled read-in-order / old analyzer, priority order respected, error when listed alone without exploitable order) and correctness against `hash` for `INNER`/`LEFT`/`RIGHT`/`FULL`/`ANY`/`join_use_nulls`. `04670_parallel_sorted_merge_join` pins the primary-key-range sharding (`Sharding:` marker with `query_plan_join_shard_by_pk_ranges = 0`, no `ScatterByPartitionTransform`, no `MergeSortingTransform`), the `ASOF` degradation, and correctness. `04824_sorted_merge_join_parallel_replicas_fallthrough` and `04894_sorted_merge_join_parallel_replicas_runtime_filter` pin the parallel-replicas edge (fall-through to `hash` with the runtime filter restored), `04893_parallel_sorted_merge_join_shard_stream_divergence` pins the diverged-shard degradation, and `04760`/`04850` pin the `join_use_nulls` and `query_plan_join_shard_by_pk_ranges` contracts. The PR #109005 regression tests and the join runtime filter tests pass unchanged. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added new `join_algorithm` values `sorted_merge` and `parallel_sorted_merge`: merge-join algorithms that are available only when both join inputs can be efficiently read in the order of the join keys (so the join benefits from the tables' order instead of sorting), and otherwise fall through to the next algorithm in the list. `parallel_sorted_merge` additionally shards the join by primary-key ranges into independent per-shard merge joins running in parallel. Listing them first, e.g. `join_algorithm = 'sorted_merge,parallel_hash'`, uses the streaming low-memory merge join exactly when it is certainly beneficial. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/112973",
        "createdAt": "2026-08-02T03:03:03Z",
        "updatedAt": "2026-08-13T15:24:35Z",
        "timestamp": "2026-08-13T15:24:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113021",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Build the `arrayIntersect` hash map from the smallest argument",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/2120 --> `arrayIntersect` filled its hash map from every argument and then rescanned the first one. For `arrayIntersect(a, b)` with 30 million and 5 million elements that is 35 million insertions into a map sized for the union of both, followed by 30 million lookups. A value that is missing from any one of the arguments cannot be in the intersection, so it is enough to fill the map from the smallest argument and to only look the other ones up. The map then stays as small as the smallest argument, which is what decides the speed once it no longer fits in cache. Which argument seeds the map does not affect the result, so it is chosen once for the whole column and short arrays pay nothing per row. Measured with a baseline built from unmodified sources in the same build directory (release, aarch64), best of three: | query | baseline | this PR | | |---|---|---|---| | `length(arrayIntersect(aa, bb))`, 30M vs 5M elements | 5.25 s, 1.94 GiB | 3.00 s, 1.26 GiB | 1.75x | | `length(arrayIntersect(bb, aa))`, 5M vs 30M elements | 3.95 s, 1.94 GiB | 2.14 s, 1.26 GiB | 1.85x | | 2 x 8M-element arrays | 1.59 s, 1.03 GiB | 1.26 s, 0.66 GiB | 1.26x | | 3 x 4M-element arrays | 0.86 s, 0.62 GiB | 0.75 s, 0.62 GiB | 1.15x | | 5M rows x 8-element arrays | 1.01 s | 1.05 s | 0.96x | | 1M rows x 4-element `String` arrays | 0.26 s | 0.26 s | 1.00x | | `arrayUnion`, 2 x 8M-element arrays | 1.41 s | 1.44 s | 0.98x | | `arraySymmetricDifference`, 2 x 8M-element arrays | 1.39 s | 1.42 s | 0.98x | The small cases lose 2-4%, reproducibly across best-of-seven runs, even though the new code executes fewer instructions for them (9.97 G against 10.38 G on the 5M x 8 case) - it looks like code layout, the element loop is now instantiated for both the filling and the looking-up argument. I left it as it is: a few percent on arrays that fit in cache in exchange for 1.75x on the arrays where this function actually gets slow. The performance tests run on quieter machines than the one I measured on, so they are the better judge of the small cases. Reordering the arguments requires the counter to be exact, and that also fixes a wrong result. A value repeated in a later argument used to be counted twice, so it could reach the \"present in every argument\" count while being absent from an argument in between: ```sql SELECT arrayIntersect([1], [2], [1, 1]); -- was [1], now [] SELECT arrayIntersect([1, 2], [2], [1, 1, 2]); -- was [1,2], now [2] SELECT arraySymmetricDifference([1], [2], [1, 1]); -- was [2], now [2,1] SELECT arraySymmetricDifference([1], [2, 2]); -- was [1], now [2,1] ``` For `arraySymmetricDifference` two arguments are enough, because it reads the counter directly instead of rescanning the first argument: the two copies of `2` took the counter to 2, which the old code read as \"present in both arguments\" and left out of the result. A value is now counted for an argument only when it was present in every argument before it, so a counter equal to the number of arguments means exactly \"present everywhere\". `arrayUnion` is unaffected, it only ever asked whether the counter was non-zero. The counter used to be exact - `if (*value == arg_num) ++(*value);`, written together with the comment above it in 54407002993. It was relaxed to `<= arg_num` in d29f0d4c966, the commit that added `arrayUnion` in https://github.com/ClickHouse/ClickHouse/pull/68989, because that mode takes every key whose counter is non-zero, and with the exact condition a value that first appears in argument `k > 0` is inserted as `0`, never incremented, and dropped from the union. What the relaxation also allowed - a second increment inside the same array whenever the counter lags behind `arg_num` - is the wrong result above. This PR does not restore the exact condition in place, it removes the need for the relaxation: `arrayUnion` no longer looks at the counter at all, because every key of the map is present in at least one of the arguments by construction. While in the same loop, `arrayUnion` and `arraySymmetricDifference` no longer look every key up again through `map.find` while iterating the map itself, which was one redundant lookup per key. Related: https://github.com/ClickHouse/ClickHouse/issues/2120 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `arrayIntersect` and `arraySymmetricDifference` no longer treat a value repeated inside a single argument as if it appeared in several arguments. Queries relying on the previous behavior can now return different results: `arrayIntersect([1], [2], [1, 1])` returns `[]` instead of `[1]`, `arrayIntersect([1, 2], [2], [1, 1, 2])` returns `[2]` instead of `[1, 2]`, and `arraySymmetricDifference([1], [2], [1, 1])` returns `[2, 1]` instead of `[2]`. For `arraySymmetricDifference` two arguments are already enough: `arraySymmetricDifference([1], [2, 2])` returns `[2, 1]` instead of `[1]`. A value is now counted for an argument only when it was present in every argument before it, so the result contains exactly the values present in all of the arguments. As part of the same change, `arrayIntersect` builds its hash table from the smallest argument rather than from all of them, which makes it up to 1.85x faster and use a third less memory when the arguments differ a lot in size. `arrayUnion` is not affected. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1295` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113021",
        "createdAt": "2026-08-02T18:10:30Z",
        "updatedAt": "2026-08-13T03:45:12Z",
        "timestamp": "2026-08-13T03:45:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "pr-backward-incompatible",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113022",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add a type-aware Bloom filter index for JSON",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 The original design used `JSONAllValues` as the input to a Bloom filter. `JSONAllValues` serializes each value as text. It does not preserve the runtime type. This loss of type information is important for `Dynamic` values. ClickHouse can compare JSON values with different runtime types. Some type pairs can match after conversion. Other type pairs can return an exception. A Bloom filter that stores only text cannot safely model these rules. It can skip a granule that contains a match. It can also hide an exception. `JSONAllValues` remains useful for text search, but it is not a safe base for this index. This PR replaces that design with `jsonbf_v1`. The new index creates tokens from these components: - JSON path - Container role - Runtime type - Binary value The container role separates scalar values, array elements, and map values. The index also processes nested JSON objects and named tuple fields. For `Dynamic` values, the index stores type-presence and complex-value presence tokens. Query analysis uses exact value tokens only when the comparison is safe. If analysis cannot prove safety, ClickHouse reads the granule. Container roles are preserved through nested tuples, and nested casts are handled conservatively. The index supports: - Equality and typed `IN` conditions - `has`, `hasAny`, and `hasAll` for arrays - Typed map values by key - Nested JSON objects and tuples The index does not optimize range conditions or whole-container equality. It also rejects unsafe comparison paths, such as Decimal-to-Float comparisons. An unsupported `Dynamic` runtime type disables skipping for its granule. In a one-million-row JSONBench test, the index reduced reads from 123 granules to 12–15 granules. Selected string equality queries were 1.6–2.3 times faster. Indexed inserts were approximately 2.9 times slower in the checked-in performance test. A local performance test produced these median results: | Operation | Without index | With `jsonbf_v1` | Difference | |---|---:|---:|---:| | Numeric equality, matching value | 13.05 ms | 10.99 ms | 1.19 times faster | | Numeric equality, missing value | 13.25 ms | 10.80 ms | 1.23 times faster | | String equality | 24.49 ms | 12.90 ms | 1.90 times faster | | Array `has` | 24.71 ms | 12.78 ms | 1.93 times faster | | Insert | 20.51 s | 40.87 s | 1.99 times slower | The small performance test shows limited benefit for numeric scalar equality and larger improvements for string equality and array membership. The larger JSONBench data set benefits more because the index skips more granules. Token generation and Bloom-filter construction increase insert time. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds the `jsonbf_v1` data-skipping index for type-aware equality and array membership on `JSON` values.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113022",
        "createdAt": "2026-08-02T18:31:45Z",
        "updatedAt": "2026-08-13T17:35:49Z",
        "timestamp": "2026-08-13T17:35:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "rorylshanks",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113023",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Rewrite `length(arrayFilter(f, arr))` to `arrayCount(f, arr)`",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/2120 --> `arrayFilter` builds an array of the elements that pass the predicate, and `length` then throws that array away and keeps only its size. `arrayCount` computes the same number without materializing anything. ```sql SELECT length(arrayFilter(x -> (x >= 2), arr)); -- becomes SELECT arrayCount(x -> (x >= 2), arr); ``` The rewrite is a new query tree pass in the spirit of the existing `RewriteArrayExistsToHasPass`, controlled by the new `optimize_rewrite_array_filter_length_to_array_count` setting (on by default). A new node is built rather than rewriting the `arrayFilter` node in place, because the filtered array can be referenced elsewhere in the query tree, where it is still an array. `arrayCount` counted into `UInt32`, which silently wrapped for a row whose array has more than `4294967295` matching elements, while `length(arrayFilter(...))` counts through `ColumnArray::Offset` and is exact. It now counts into `UInt64` and returns `UInt64`: this fixes the overflow of `arrayCount` itself, and makes the rewrite value-preserving for arrays of any size. The result type of `arrayCount` therefore changes from `UInt32` to `UInt64`; the values are unchanged. Because `length` and `arrayCount` now have the same result type, the rewritten node needs no `CAST`, and the pass only rewrites when the two result types match. Since this changes the stable public signature of an existing function, the previous behavior is kept behind the new compatibility setting `array_count_legacy_uint32_result` (default `false`): setting it to `true` restores the `UInt32` result type, and it is wired into the settings changes history, so `compatibility = '26.7'` (or older) restores it automatically on the servers where it is set. For distributed queries, the setting (like any setting) is forwarded from the initiator to the shards, so a query initiated by a server that has it set behaves exactly as before on every shard. The one case it does not cover is a rolling upgrade with a *not-yet-upgraded initiator*: an old server does not know the setting and cannot forward it, so type-sensitive expressions evaluated locally on already-upgraded shards (for example, `byteSize(arrayCount(...))`) observe `UInt64` there, while the initiator still converts the top-level result to its own `UInt32` header. To keep such queries fully unchanged during the upgrade, set `array_count_legacy_uint32_result = 1` on the upgraded servers for the users under which shard-side queries execute (with an interserver `secret` configured that is the initiator's current user; otherwise it is the user from the cluster definition or the `remote` table function - the simplest robust approach is to enable it for all users of the upgraded servers), and remove it once the whole cluster is upgraded. This is documented in the `arrayCount` documentation and in the setting's description, and both directions of the mixed-version scenario are pinned by the integration test `test_backward_compatibility/test_array_count_return_type.py`. In legacy mode the rewrite does not fire, because the pass requires the result types to match. Measured with the setting toggled on the same binary (release, aarch64), best of five, before the `CAST` was removed: | query | setting off | setting on | | |---|---|---|---| | one 30M-element array, half the elements pass | 0.38 s | 0.33 s | 1.15x | | one 30M-element array, all elements pass | 0.40 s | 0.31 s | 1.29x | | one 10M-element `String` array | 0.40 s | 0.37 s | 1.08x | | 5M rows x 32-element arrays, unpredictable predicate | 1.42 s | 1.29 s | 1.10x | | 1M rows x 8-element `String` arrays | 0.25 s | 0.22 s | 1.14x | | 5M rows x 32-element arrays, half the elements pass | 1.41 s | 1.48 s | 0.95x | The gain grows with the array length, which is where the copy that `arrayFilter` does starts to cost something. The last row loses 5%: for short arrays the copy is cheap, `ArrayCountImpl` counts the filter with a scalar loop while `arrayFilter` copies with vectorized code, and the cast to `UInt64` added a column of its own - that cast is no longer generated. I did try replacing that scalar loop with `countBytesInFilter`, but its vectorized path is `__SSE2__` only, so on the aarch64 machine I measured on it made no difference and I dropped it rather than commit something I could not verify. It may be worth doing separately, measured on x86. Found while profiling https://github.com/ClickHouse/ClickHouse/issues/2120, where `length(arrayFilter(x -> (x >= 2), arrayEnumerateUniq(...)))` is applied to a single array of 60 million elements. Related: https://github.com/ClickHouse/ClickHouse/issues/2120 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `arrayCount` now returns `UInt64` instead of `UInt32`, so that it is exact for arrays with more than 4294967295 matching elements; set the new setting `array_count_legacy_uint32_result` to `true` (or use `compatibility = '26.7'`) to restore the previous result type. During a rolling upgrade, set `array_count_legacy_uint32_result = 1` on the upgraded servers for the users under which shard-side queries execute (the simplest robust approach is to enable it for all users of the upgraded servers), so that distributed queries initiated by not-yet-upgraded servers (which cannot forward the setting) keep the previous behavior on upgraded shards; remove it after the upgrade is complete. In addition, `length(arrayFilter(func, arr))` is now rewritten to `arrayCount(func, arr)`, which counts the matching elements instead of building an array of them only to take its size - up to 1.29x faster on long arrays, controlled by the `optimize_rewrite_array_filter_length_to_array_count` setting.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113023",
        "createdAt": "2026-08-02T18:39:16Z",
        "updatedAt": "2026-08-13T02:22:02Z",
        "timestamp": "2026-08-13T02:22:02Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "pr-performance",
          "pr-backward-incompatible"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113024",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Run the documentation examples in CI",
        "text": "#108556 made the SQL examples embedded in `system.documentation` runnable, but nothing runs them, so they go stale again as soon as behaviour changes: since then a steady stream of one-off fixes has been needed (#109421, #109965, #110459, #112287, ...). This adds a runner and a CI job that execute every one of them, and brings the examples and their documented responses back in line with what the server actually does. ### The runner `tests/docs_examples/runner.py` reads every example out of `system.documentation` on a running server and executes it. The examples of one entity run **in order, in a single session, in a database of their own**: a documentation page is written to be followed from top to bottom, so one example commonly creates the table a later one queries. It is plain Python and can be pointed at any server: ``` python3 tests/docs_examples/runner.py --port 8123 python3 tests/docs_examples/runner.py --port 8123 --filter '^argM' --verbose ``` Each example gets one of three outcomes: * `ok` — it ran, and its output matches the documented response (or it documents no response, or it documents an exception and indeed threw it); * `error` — it failed to run (or documents an exception and did not throw it); * `output` — it ran, but its output differs from the documented response. Everything that is not `ok` is listed in `tests/docs_examples/known_failures.txt` with the reason for it. The run fails if an example that is not on the list fails, and also if a listed one starts passing, so the list can only shrink. A few entries are marked `unstable`, for the examples whose output is random enough to sometimes match the documented one. An example that documents an exception is expected to throw it from its **last** statement — what comes before is setup that has to succeed — and when the documented response names an error code, the exception has to carry the same one. The message text is not compared: it holds a version number and a query pipeline description that are not part of what the example teaches. The server the job starts runs in `Etc/UTC`. Many examples convert between a date with time and a number, or hash a `DateTime`, without naming a time zone, and the documented responses are the ones of a `UTC` server; without pinning it, the run would only reproduce on a machine whose time zone matches. The comparison pins the Pretty rendering to the plain form the documented responses are written in — no row numbers, no colour, long column names spelled out in full, no readable-number tip, a named tuple printed as a tuple — so that it is about the data and the shape of the result rather than about a rendering default that changed after a page was written. ### The job `ci/jobs/docs_examples_job.py` starts a server from the shipped configuration plus the fragments in `programs/server/config.d` (macros, the legacy geobase, the natural language processing data, Keeper, the test clusters) and `programs/server/users.d` (the localhost-only network of the `default` user, access management, query logging), and two fragments of its own in `tests/docs_examples/config.d` and `tests/docs_examples/users.d`, so the features the examples demonstrate are actually configured, and configured the way the shipped server configures them. The one thing the examples add to the shipped user configuration is stated explicitly: the `queryID`, `initialQueryID` and `initialQueryStartTime` examples read from three shards at `127.0.0.{1..3}`, so `default` is allowed in from the whole loopback network rather than from `127.0.0.1` alone. It publishes an HTML report naming every failing example with its source file, its query and both responses. The job runs on pull requests and on master, next to the other jobs that run a corpus of queries against a server. ### What it found **Examples that did not run.** All of these are fixed here: * examples that were never a query: a bare expression (`factorial(10)`), a syntax template with placeholders sitting in an `Examples` block (the `iceberg*`, `paimon*`, `deltaLake*`, `oss`, `cosn` and `mergeTree*` table functions — moved to the `syntax` field, where they belong), a leaked C++ string literal, a MySQL session transcript, an unbalanced parenthesis, a stray `\\G` left over from a `clickhouse-client` session; * an example calling the wrong function: `YYYYMMDDhhmmssToDateTime` demonstrated `YYYYMMDDToDateTime`; * examples reading a table nothing creates: `salary`, `Employees`, `t`, `key_val`, `encryption_test`, `student_ttest`, `points`, `example_table`, ... — each page now creates its own; * pages that cannot be followed top to bottom, because every example re-creates the same table and the second one hits `TABLE_ALREADY_EXISTS`; * arguments the function rejects: an H3 index that is not a valid directed edge, S2 cell ids that do not form a valid rectangle, a signed weight for a weighted quantile, an `INSERT ... FORMAT JSONEachRow` without the semicolon that ends its data, a cipher the build does not provide; * examples of an error whose response was prose (\"Raises a `NO_COMMON_TYPE` exception\") or a paste from version 19.14, now written as the exception the server prints — the runner treats a documented exception as an expectation to fail. **Responses that no longer match.** 621 of them are regenerated from what the server prints, after checking that the output is identical across two runs on a fresh server. Most were stale column names (`avg(x)` for a column named `t`, `argMax(a, tuple(b, a))` for what is now printed as `argMax(a, (b, a))`) or hand-typed values that never came from a server (`[4, 3, 2, 1]` where the server prints `[4,3,2,1]`, `'dcba'` where it prints `dcba`), but some were genuinely wrong results. **Examples that needed a dataset nobody has.** `anyHeavy`, `categoricalInformationValue`, `topK`, `IPv4NumToStringClassC`, `IPv6NumToString`, `arrayEnumerateUniq`, `bar`, `indexHint`, `transform` and `evalMLMethod` demonstrated themselves on `ontime`, `metrica.hits`, `test.hits`, `hits_all`, `test.visits` or `trips`. Each of them now builds a small table of its own, so the example is one a reader can run. `naiveBayesClassifier` and its two variants train a dictionary on an inline set of token counts instead of naming one that does not exist. The four `flameGraph` \"examples\" were not examples at all — their documented response was a `clickhouse client ... | flamegraph.pl` command line — so they are recipes in the description now, and the function has one example that runs. **What is left in the known-failures file** is what a test server cannot produce: examples that call an external model provider, need a trained CatBoost model, a TLS certificate or a geobase hierarchy the test configuration does not carry; responses that describe the machine or the build (a host name, a source path, a version, a disk size); and outputs that are random or depend on the time. ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113024",
        "createdAt": "2026-08-02T19:33:37Z",
        "updatedAt": "2026-08-13T11:34:16Z",
        "timestamp": "2026-08-13T11:34:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 19
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113045",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Perform the `Too many parts` check once per INSERT query",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109000 `MergeTreeSink::onStart` and `ReplicatedMergeTreeSink::onStart` perform the `parts_to_throw_insert` / `parts_to_delay_insert` check, and `onStart` is the only place where an `INSERT` may be rejected with `TOO_MANY_PARTS` - rejecting a query that has already written a part is not acceptable, as the comment right above the call says. Since #109000 a plain `INSERT` writes through up to `max_insert_threads` sinks running in parallel, and every one of them runs `onStart` (`ExceptionKeepingTransform::prepare` returns `Ready` for the `Start` stage of every sink, whether or not it ever receives a chunk). The `onStart` calls of the different sinks are not ordered with respect to each other's writes: a sink that the executor schedules late runs its check *after* another sink of the same query has already committed a part, counts that part, and rejects the query in the middle of it. The `INSERT` fails with `TOO_MANY_PARTS` even though the table was below the threshold when the query started - and even when the table was empty and the only part counted is the one the query wrote itself. Note that #109000 exposed the flaw rather than introduced it. The parallel sink fan-out itself (`sink_stream_size = max_insert_threads` in `InsertDependenciesBuilder`) has been in place for the `INSERT SELECT` path since `031a9d5a528`, and the default of `max_insert_threads` changed from `1` to `0` (auto, the number of CPU cores) in 26.8 - so an `INSERT SELECT` into a MergeTree table could already be rejected by a stream that counted a part the query had written itself. Before #109000 the plain-`INSERT` pipeline passed a hardcoded `/*max_insert_threads*/ 1` to the builder and always had a single sink, which is why the flaky tests below - all of them plain `INSERT ... VALUES` - only started failing when that changed. The check is now shared by all the sinks of one query: the first sink to reach `onStart` runs it and the others wait until it is finished, so it runs exactly once and strictly before any of the sinks writes anything. This restores the behaviour that the single-sink pipeline had. Only the rejection is shared: the `parts_to_delay_insert` backpressure is applied by every sink before each of its blocks, including the first one, so a parallel `INSERT` keeps the same per-block throttling as a single-sink one (a sink that receives no data does not delay). The gates live in a small per-query registry (`InsertStartGates`) keyed by the physical destination table id. `InsertDependenciesBuilder` creates the registry and takes each sink's gate from it next to `setRuntimeData` / `setHasDependentMaterializedViews`, so the sinks of all the streams that write into the same table share one gate - including the branches of different materialized views converging on the same target table. An `Alias` destination forwards the write through a nested `INSERT` per parallel branch and the real check runs inside those nested inserts, so `AliasSink` receives the registry and threads it into the nested `InterpreterInsertQuery`, whose builder inherits it instead of creating its own. `Distributed` and `Buffer` destinations are always kept single-stream by #109000, so their own pre-write checks are unaffected - but their in-query nested `INSERT`s into the underlying tables are not. The *local* writes of a `Distributed` destination run through nested `INSERT`s into the underlying table: one per incoming block on the `writeToLocal` path (the direct local write of a background-send insert with `prefer_localhost_replica`) and one per local replica job on the foreground path. Each of them used to create its own gates, so the second block of a multi-block `INSERT` counted the part the first block had already committed on the local target and failed the query mid-way; `DistributedSink` also receives the registry and threads it into both nested `INSERT` paths (the new test `04826_insert_distributed_local_too_many_parts_gate` covers a two-block insert into a local target with `parts_to_throw_insert = 1` on both paths). The `parallel_distributed_insert_select` rewrite of an `INSERT ... SELECT` between two `Distributed` tables writes the local shards the same way - one nested `InterpreterInsertQuery` per destination shard this server belongs to - and a server can be local for several destination shards of one cluster, so those sibling nested `INSERT`s write into the same underlying table and now share one registry as well (the new test `04847_insert_parallel_distributed_insert_select_local_shards` covers a two-local-shard cluster with `parts_to_throw_insert = 1` on the target, plus the non-parallel quorum variant of the same topology). The direct writes of a `Buffer` destination run through nested `INSERT`s into its destination table as well: one per block that bypasses the buffer by exceeding the max thresholds, and one per flush the buffer runs by threshold from within the query - so a later direct write of a multi-block `INSERT` used to count the part an earlier one had committed on the destination; `BufferSink` also receives the registry and threads it through `flushBuffer` / `writeBlockToDestination` into those nested `INSERT`s, while the background flush keeps creating fresh gates per flush as before (the new test `04827_insert_buffer_direct_write_too_many_parts_gate` covers both paths with `parts_to_throw_insert = 1` on the destination). A `TimeSeries` destination forwards the write through nested `INSERT`s into its inner target tables, so `TimeSeriesSink` also receives the registry and threads it into them, sharing the check across sibling branches converging on the same `TimeSeries` table. Those nested pipelines are now created in `onStart` rather than in the sink's constructor: the registry reaches the sink only after `StorageTimeSeries::write` has returned it, so building them in the constructor handed them an empty registry and the sharing held only by the accident of every sink of the query being constructed before any of them writes anything. A `WindowView` forwards every chunk of every branch into a fresh sink of its inner table, so `PushingToWindowViewSink` also receives the registry and threads it into `StorageWindowView::writeIntoWindowView`, which sets the query's gate of the inner table on the sinks it creates - sharing the check across the branches and across the successive chunks of one branch. This path cannot be covered by a test yet: a window view currently cannot receive inserted data at all - the view dependency of the source table is never registered and a direct `INSERT` into a window view fails with `std::out_of_range` - both since before this change (they reproduce on 26.7), see https://github.com/ClickHouse/ClickHouse/issues/113493. A related single-in-flight-part contract exists one level up: a non-parallel quorum insert (`insert_quorum >= 2` or `'auto'`, with `insert_quorum_parallel = 0`) permits a single in-flight quorum part per table, which is why its `max_insert_threads` fan-out is kept single-stream. But the single sink stream is still duplicated into one branch per dependent materialized view, and with `parallel_view_processing = 1` those branches ran concurrently - two views converging on one `ReplicatedMergeTree` target raced two in-flight quorum parts of one `INSERT` against each other, so a sibling branch could fail with `UNSATISFIED_QUORUM_FOR_PREVIOUS_WRITE` or conflict on the `/quorum/status` node. The serialization such an insert needs is derived from the collected sink graph, not from the settings alone (`InsertDependenciesBuilder::computeQuorumStreamRequirements`): only writes reaching a `ReplicatedMergeTree` table are quorum writes, and only two of them racing on the same table conflict. The `INSERT SELECT` fan-out - which had no quorum guard at all - is kept single-stream only when a branch of the write may produce a quorum part: a reachable `ReplicatedMergeTree` target, or a forwarding storage (an `Alias`, a `Distributed`, a `Buffer`, a `WindowView`, a `TimeSeries`) that hides its physical destination, in which case the probe fails closed. An insert whose write graph never reaches a replicated table keeps its `max_insert_threads` fan-out even under a global quorum profile. Only the branches reachable in the executable graph are counted: a view branch pruned from the graph - e.g. a view whose dropped target table is ignored by `ignore_materialized_views_with_dropped_target_table` - never creates a sink, so it does not force the serialization (the new test `04825_insert_quorum_ignored_broken_view_fanout` covers a broken ignored view attached to a plain `MergeTree` destination). The dependent views of such an insert are pushed sequentially only when two branches converge on the same replicated table (a sequential branch blocks in `commitPart` until the quorum of its part is satisfied, so the next branch starts with the quorum node already gone) or when a hidden write target makes such a convergence impossible to rule out - branches writing to distinct replicated tables keep running concurrently, since they do not share a `/quorum/status` node. An `Alias` destination hides its target's views behind the nested `INSERT` its sink runs, and that nested `INSERT` observes the same settings and derives the same requirements. The new test `04828_insert_timeseries_too_many_parts_gate` exercises non-parallel quorum inserts through two views converging on one `TimeSeries` table whose data table is replicated - the shape that raced `UNSATISFIED_QUORUM_FOR_PREVIOUS_WRITE` before `TimeSeries` was included in the hidden-target probes. The new test `04817_insert_quorum_sequential_dependent_views` covers the `INSERT SELECT` fan-out deterministically via `EXPLAIN PIPELINE` and exercises quorum inserts through two materialized views converging on one replicated target with `parallel_view_processing = 1`; the new test `04823_insert_quorum_graph_derived_serialization` covers the fan-out kept for a plain `MergeTree` destination, the fan-out dropped when a dependent materialized view targets a replicated table, and quorum inserts through two views writing to two distinct replicated tables. `parallel_view_processing = 0` gets the same treatment on the `INSERT SELECT` path: `buildInsertPipeline` already kept a plain `INSERT` single-stream when the destination is a forwarding storage that hides its target's dependent-view graph behind the nested `INSERT` its sinks run (`serial_hidden_views`), but `addInsertToSelectPipeline` still passed the raw `max_insert_threads` to the builder - which sees no views for such a destination - so an `INSERT SELECT` into an `Alias` fanned out into several `AliasSink`s whose hidden materialized views ran concurrently across sibling branches despite the setting. The guard is now mirrored there; an `Alias` in front of a target with dependent views is the only topology whose fan-out survived to that point, since `Distributed`, `Buffer` and `TimeSeries` destinations are already collapsed to one stream by the builder's `supportsParallelInsert` check. The new test `04869_insert_select_alias_hidden_views_serial` covers the single stream with a hidden dependent view, and the fan-out kept both without dependent views and with `parallel_view_processing = 1`. This is the cause of a family of flaky tests that started failing on `master` on 2026-08-01, right after #109000 was merged, each of them rejecting an `INSERT` that counts the part it has written itself: | test | failing query | reported parts | | --- | --- | --- | | `02458_relax_too_many_parts` | `INSERT INTO test VALUES (6, 'a')` | 3 with `parts_to_throw_insert = 3` | | `02280_add_query_level_settings` | `INSERT INTO table_for_alter VALUES (1, '1')` into an empty table | 1 with `parts_to_throw_insert = 1` | | `00980_merge_alter_settings` | `INSERT INTO table_for_alter VALUES (1, '1')` into an empty table | 1 with `parts_to_throw_insert = 1` | | `02015_async_inserts_5` | `INSERT INTO async_inserts VALUES` | 1 with `parts_to_throw_insert = 1` | The server log of the `02458_relax_too_many_parts` failure below shows the mechanism directly: 26 sinks are created for the query, its part `all_5_5_0` is renamed into place at `01:46:59.792095`, and the query is rejected with `Too many parts (3 ...)` at `01:46:59.798713` - six milliseconds after it committed the part it is being counted for. Report of that failure (`MergeQueueCI`, `Fast test`): https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-112777-508dbed54a148ce6ae2dabe8ff91f4c7842bb1a0&sha=c71ec1f59e4b52ffe87c96e12743488cec72a7e0&name_0=MergeQueueCI&name_1=Fast%20test Corresponding pull request: https://github.com/ClickHouse/ClickHouse/pull/112777 The new test `04692_insert_too_many_parts_check_once_per_query` covers both halves: an `INSERT` with a large `max_insert_threads` that brings a table from one part up to `parts_to_throw_insert = 2` has to be accepted, and `ProfileEvents['DelayedInserts']` of a single-row `INSERT` has to match the number of blocks the query actually writes (one; two when the query writes through two materialized views converging on the same target table) rather than the number of streams it writes through - also when it writes through an `Alias`. `TimeSeries` is also classified as a forwarding storage in the deduplication-safety probes (`storageDeduplicatesBlocksOnInsert`, `storageRebuildsDeduplicationIdsOnInsert`, `forwardedInsertReachesDependentView`, `forwardedInsertHidesDependentView`, `forwardedInsertHidesDependentViewForwardingToSeparateContext`), since `TimeSeriesSink` forwards a plain block into its nested `INSERT`s and drops the outer chunk's `DeduplicationInfo`. This is fail-close classification only, not a reachable fix: `StorageTimeSeries` does not override `IStorage::supportsParallelInsert`, so an insert whose graph reaches a `TimeSeries` table is already single-stream and the probes cannot change the outcome - and writes arriving through `TimeSeriesSink` turn out not to deduplicate on the inner tables at all, so no rows are lost when two views converge on one `TimeSeries` table. The clauses keep the classification correct if `StorageTimeSeries` ever starts supporting parallel inserts. The new test `04846_insert_mv_timeseries_dedup_parallel_views` pins the no-row-loss half end to end: two identical materialized views converging on one `TimeSeries` table with a deduplicating data table (`non_replicated_deduplication_window`) under `parallel_view_processing = 1` land both branches in full - nothing is deduplicated between the sibling branches of one query. #### Relation to #114016 While this pull request was open, #114016 landed on `master` with a narrower fix for the same root cause: it evaluates the check on the query thread in the `MergeTreeSink` / `ReplicatedMergeTreeSink` constructor and rethrows it from `onStart`. Merging `master` resolves that in favour of the shared gate, which subsumes it - a sink whose `onStart` runs late skips the check instead of re-running it - and additionally covers the sinks that are created *during* execution, which the constructor placement cannot: the nested `INSERT`s of `Alias`, `Distributed`, `Buffer`, `WindowView` and `TimeSeries`. Everything else from #114016 is kept, including the `merge_tree_sink_on_start_random_sleep` failpoint and its test `04826_parallel_insert_sinks_too_many_parts_self_race`, which the shared gate also makes pass. Related: https://github.com/ClickHouse/ClickHouse/pull/114016 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a spurious `TOO_MANY_PARTS` error for an `INSERT` executed with `max_insert_threads` greater than one. The `Too many parts` check was performed by every parallel writing stream, so a stream that started after another one had written a part counted that part and rejected the query, even when the table was below the threshold - or empty - when the query started.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113045",
        "createdAt": "2026-08-03T01:08:11Z",
        "updatedAt": "2026-08-13T01:30:59Z",
        "timestamp": "2026-08-13T01:30:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113059",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Collect SQL stacktraces on the hung-check and server-died abort paths",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/42701 Related: https://github.com/ClickHouse/ClickHouse/pull/112265 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ### Description Related: #42701, #112265 On an ASan build, a stateless run aborting on the hung check records no stack of the hung server: ``` Hung check failed: server is not responding Cannot collect C stacktraces under ASan: debugger attach is disabled. ``` Two things combine. `print_c_stacktraces` declines to attach lldb on ASan builds, because ptrace disables LeakSanitizer; that refusal is correct and stays. And `print_sql_stacktraces`, needing no debugger, was unreachable: its pre-check `check_server_liveness` probes **HTTP** (`http_port`, default 8123), while it collects over native **TCP** (`args.client --port=`, default 9000). Different listeners, different ports, so this signature (alive, not answering HTTP, TCP still serving) failed the gate. The hung-check abort site did not call it at all. This drops the mismatched pre-check and lets the collector be its own liveness test: it is already bounded (`timeout=30`) and reports failure as one trimmed line, so a dead socket costs at most 30 s and cannot re-emit the `Code: 210` tracebacks that motivated the pre-check. The dump is added to the three abort sites that had only the C path: hung check, server died, and the startup check. The stateless job keeps attaching the dump to its result and additionally clears any left by a previous job in the same workspace, so an aborted run cannot upload a stale dump as its own. #114143 has since added the same attachment upstream; this replaces it with the equivalent helper rather than attaching twice. It fixes no hang and does not restore C++ stacks on ASan. A server alive but not answering HTTP now yields the full `system.stack_trace` view with per-thread `query_id`, identifying the wedged query; one dead on both transports records \"tried, got nothing\" instead of silence. Validated against a live server: with HTTP dead and TCP live, master skips and writes nothing, while this branch writes a `sql_stacktraces.log` carrying `thread_name` and `query_id`. Green runs are unaffected. [Prompting report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=1ce152efafc5b00bf31eb7a0a64757fecbbc6e4a&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20distributed%20plan%2C%20parallel%29).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113059",
        "createdAt": "2026-08-03T06:03:46Z",
        "updatedAt": "2026-08-13T17:12:07Z",
        "timestamp": "2026-08-13T17:12:07Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "manual approve",
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113076",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Populate query_log views column for inserts through Alias",
        "text": "Share query access info with the forwarded target insert so materialized views triggered on the target are recorded in the `views` column of `system.query_log`. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `views` column in `query_log` for inserts to an `Alias` table whose target triggers materialized views, which was always empty.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113076",
        "createdAt": "2026-08-03T10:16:33Z",
        "updatedAt": "2026-08-13T15:20:13Z",
        "timestamp": "2026-08-13T15:20:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "eclbg",
        "state": "open",
        "assignees": [
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113107",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Disable the sampling query profiler under Memory Sanitizer",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/100242 Related: https://github.com/ClickHouse/ClickHouse/issues/106591 Every `MemorySanitizer` report produced in CI since 2026-05-28 is truncated to a single line: ``` ==117==WARNING: MemorySanitizer: use-of-uninitialized-value MemorySanitizer: nested bug in the same thread, aborting. ``` No stack trace, no `SUMMARY`, no origin — so no MSan failure can be located. In the CI database, `msan` checks recorded 1009 reports with a usable stack in 2026-04 and 0 in 2026-07, while `asan`, `tsan` and `ubsan` reports are unaffected and still carry full stacks. The regression window matches #100242 (merged 2026-05-28), which enabled the sampling query profiler for sanitizer builds. Printing an MSan report takes the sanitizer runtime seconds, because it symbolizes every frame through an `llvm-symbolizer` subprocess, and it holds `ScopedErrorReportLock` for the whole time. A profiler signal delivered to that thread meanwhile runs instrumented code in the handler and trips a second MSan check; compiler-rt treats a same-thread re-entry into the report lock as unrecoverable, writes `nested bug in the same thread, aborting.` with `CatastrophicErrorWrite` and calls `internal__exit`, so the report that was already being written to `MSAN_OPTIONS=log_path` stops after its header line. Reproduced with the `arm_msan` binary built from master `3885f7b7bbf1`, running the same query against the same server three times and forcing a sanitizer report with `max_allocation_size_mb=32 allocator_may_return_null=0`: | configuration | result | | --- | --- | | profiler on (default `global_profiler_real_time_period_ns`) | `nested bug in the same thread, aborting.`, 1-line report | | `<trace_log remove=\"1\"/>` (no profiler at all) | full 31-line report with stack | | `trace_log` on, `global_profiler_*_period_ns = 0`, per-query profiler off | full 31-line report with stack | So it is the profiler timer signals, not the trace collector. This turns `QUERY_PROFILER_SUPPORTED` off under MSan, next to the existing TSan-on-macOS exclusion. The trace collector stays enabled: the memory profiler, `trace_profile_events` and `SYSTEM INSTRUMENT` all feed it from ordinary code rather than from a signal handler. Other sanitizers keep the profiler, because their checks fire only on genuinely invalid accesses, which the handler does not perform. `00974_query_profiler` and `01569_query_profiler_big_query_id` assert that samples reach `system.trace_log`, so they get `no-msan` back. The other profiler and `trace_log` tests either already carry `no-msan`, exercise the memory profiler / `trace_profile_events` / `SYSTEM INSTRUMENT` (unaffected), or make no assertion on sample counts. This does not fix the uninitialized read that the `BuzzHouse (amd_msan)` run below hit — that one is unidentifiable as reported. It makes the next occurrence diagnosable. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=3885f7b7bbf1b8791021e1c97ca2e4eea132de04&name_0=MasterCI&name_1=BuzzHouse%20%28amd_msan%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113107",
        "createdAt": "2026-08-03T14:09:30Z",
        "updatedAt": "2026-08-13T03:10:52Z",
        "timestamp": "2026-08-13T03:10:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113140",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix cubic complexity of planning a JOIN with a `merge` table",
        "text": "Planning a query that joins a `Merge` table with something else could take minutes and hundreds of megabytes of memory, and it was interruptible neither by `max_execution_time` nor by `KILL QUERY`, because all of it happens inside `QueryPlan::optimize`. `ReadFromMerge::createChildrenPlans` builds a separate child query tree for every source table. When the `Merge` table takes part in a JOIN, this goes through `replaceTableExpressionAndRemoveJoin`, which rebuilds the projection list from the required column names. Every column there was resolved by its own `QueryAnalysisPass` run, and every such run rebuilt `AnalysisTableExpressionData` for all the columns of the `Merge` table. The structure of a `Merge` table is the union of the structures of its source tables, so the cost of planning was cubic in the number of source tables. All the required columns are now resolved with a single `QueryAnalysisPass` run. For 40 tables with 20 distinct columns each, a `SELECT *` joined with `merge` goes from 18 seconds down to about a second in a release build; the results are byte-identical before and after. This is what made the `Hung check` fail in a stress test: the AST fuzzer produced `SELECT * FROM (SELECT number AS a FROM numbers(11)) AS t1 PASTE JOIN merge('system', '') AS t2 LIMIT 294` for `02933_paste_join.sql`, and the query spent more than 16 minutes inside `QueryPlan::optimize` under TSan. https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=dd1e53dfe2715c22d4b7e63c2459bc9eec994581&name_0=MasterCI&name_1=Stress%20test%20%28arm_tsan%29 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed cubic complexity of query planning for a JOIN with a `Merge` table over many tables with different structures. Such queries could previously spend minutes in query planning without responding to `max_execution_time` or `KILL QUERY`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113140",
        "createdAt": "2026-08-03T15:52:16Z",
        "updatedAt": "2026-08-13T02:01:14Z",
        "timestamp": "2026-08-13T02:01:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "novikd"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113181",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: Split large function reference pages",
        "text": "Split the four regular-function families containing more than 100 functions into individual reference pages with compact searchable overview indexes. This draft tests a different information architecture for the largest Mintlify reference pages, reducing the amount of interactive code-block content rendered on a single page while preserving generated documentation and legacy fragment navigation. The generator now emits function pages, family navigation, manifests, and shared searchable index components for array, date and time, other, and type-conversion functions. It includes 546 individual function pages and a focused generator regression test. The result was validated with the regular-function and session-settings generator tests, generated-route and anchor coverage checks, `git diff --check`, and local Mintlify preview inspection. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): N/A",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113181",
        "createdAt": "2026-08-03T20:17:11Z",
        "updatedAt": "2026-08-13T14:52:47Z",
        "timestamp": "2026-08-13T14:52:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-documentation"
        ],
        "author": "dhtclk",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113192",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject shorthand setting changes carrying a value for the whole query tree",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113025 The valueless `SETTINGS name` form stands for `name = true`, and the SQL parser always writes Bool `true` for it, so the `shorthand` flag paired with any other value is a parser-impossible shape that can only arrive from the AST JSON dialect. #113025 made `BaseSettings::checkShorthandChange` reject it, but that only covers the settings applied through `BaseSettings`. `SettingsChanges` are also consumed raw — the `Join` engine reads `persistent` and its other settings directly, `EXPLAIN` settings never consult a `BaseSettings` schema, and dictionary and data-lake settings have similar readers — and every such reader would execute the carried value for a change that claims to be valueless. Instead of duplicating the check in each raw reader, reject the shape once for the whole query tree in `executeQueryImpl`, for ASTs deserialized from the JSON dialect. The check deliberately runs after `query_for_logging` is prepared, so the exception is logged with the AST masked rather than with the raw JSON text — the same reason the shape is not rejected at deserialization. Because the tree-wide check has no settings schema, it fires before the per-setting type check, so the crafted payload in `04665_valueless_setting_ast_json_and_secret_parts` now reports `BAD_ARGUMENTS` (carries a value despite claiming the valueless form) instead of `TYPE_MISMATCH`; the genuine valueless forms and their error codes are unchanged, and the masked-logging assertion still holds. This addresses the remaining review finding on #113025. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reject a setting change that is marked as the valueless `SETTINGS name` form but carries a value other than `true` for every consumer of settings (e.g. the `Join` engine or `EXPLAIN` settings), not only for the settings applied through `BaseSettings`. Such a change can only be produced by the experimental AST JSON dialect.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113192",
        "createdAt": "2026-08-03T21:58:16Z",
        "updatedAt": "2026-08-13T05:51:59Z",
        "timestamp": "2026-08-13T05:51:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113207",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add the `logsql` dialect: LogsQL, the query language of VictoriaLogs",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Added an experimental `logsql` dialect: LogsQL, the log query language of VictoriaLogs, can now be used to query ClickHouse tables with logs. Enable it with `SET allow_experimental_logsql_dialect = 1, logsql_table = '<table>', dialect = 'logsql'` (plus `logsql_database`, `logsql_time_column`, and `logsql_message_column` as needed). LogsQL queries are translated at parse time into ordinary `SELECT` queries, so text filters expect `String`-backed log columns, while the `math` pipe applies raw arithmetic and needs numeric operand columns (`String` field values are not coerced to numbers). Documentation entry for user-facing changes. - [ ] Documentation is written (mandatory for new features) Implementation notes: - Nearly the whole language is supported: the filters (words, phrases, prefixes, `field:value`, `:=`, `~`, comparisons, `range`, `in` with subqueries, `contains_any`/`contains_all` with subqueries, `seq`, `len_range`, `string_range`, `ipv4_range`, `ipv6_range`, `pattern_match*`, `contains_common_case`/`equals_common_case`, `json_array_contains_any`, `i()`, full `_time:` filters with `day_range`/`week_range`, `_stream:{...}` selectors), and the pipes `fields`, `delete`, `copy`, `rename`, `limit`, `offset`, `sort`/`first`/`last` (including `rank` and `partition by`), `stats` (with `by`-buckets, per-function `if`, `switch`, `rate`), `where`, `uniq`, `top`, `math`, `extract`, `extract_regexp`, `format`, `unpack_json`, `unpack_logfmt`, `join`, `union`, `unroll`, `running_stats`, `total_stats`, `len`, `hash`, `coalesce`, `decolorize`, `split`, `unpack_words`, `time_add`, `sample`, `generate_sequence`, `field_values`, `json_array_len`, `json_array_concat`, `replace`, `replace_regexp`, `pack_json`, `pack_logfmt`. On typed (non-`String`) columns, filters that emit string functions fail with a type error instead of applying VictoriaLogs' schemaless per-value matching. The numeric stats functions (`sum`, `avg`, `median`, `quantile`, `stddev`, `rate_sum`) parse the numeric value out of every field value and skip the values that are not numbers, like VictoriaLogs, which computes its stats in `float64`; a numeric column is therefore aggregated with `Float64` precision and not exactly. - The dialect is modeled on the existing `kusto`/`prql`/`promql`/`polyglot` dialects: a parser under `src/Parsers/LogsQL/` builds a ClickHouse `ASTSelectQuery` directly, so all downstream machinery (analyzer, distributed queries, query log) works as usual. Relational pipes map to real SQL constructs: `join` to `LEFT`/`INNER JOIN ... USING`, `union` to `UNION ALL`, `running_stats`/`total_stats` and the `rank`/`partition by` clauses to window functions. - The lexer and the grammar mirror the reference implementation in VictoriaLogs (`lib/logstorage` of `VictoriaMetrics/VictoriaLogs`, Apache 2.0), including Go-style string literals, compound tokens (`foo-bar.com:123/x`), comments with `#`, and the exact filter/pipe syntax. The test queries are reused from the VictoriaLogs parser tests: 675 of its 878 valid parser-test queries run end-to-end against a ClickHouse table, and 605 of its 612 invalid queries are rejected. - The implicit table is configured by the `logsql_database`/`logsql_table` settings (like `promql_database`/`promql_table`). The special fields `_time` and `_msg` are mapped to columns via `logsql_time_column`/`logsql_message_column`, so real tables (e.g. `system.text_log` with `event_time`/`message`) can be queried after setting those. - Word and phrase filters translate to `hasToken`/`match` with token boundaries, so tables with `tokenbf_v1` indexes on the message column benefit from index analysis. - Documented deviations from VictoriaLogs: word boundaries follow ClickHouse tokenization (ASCII alphanumerics; VictoriaLogs also treats `_` and non-ASCII letters as word characters); numeric comparison filters parse `String` field values per row (rows with non-numeric values do not match) and compare typed numeric columns exactly, but the per-row parsing of field values understands plain numeric text only (decimal integers, floats, scientific notation, `inf`/`nan`) - LogsQL-only spellings such as `10KiB`, `1h`, `0x10`, or `1_000` are supported in query literals but are treated as non-numeric when stored in a field value, whereas VictoriaLogs parses them there too; equality with a non-numeric value compares with the native column type instead of VictoriaLogs' string semantics; `count(field)` counts non-NULL values instead of non-empty strings; `pattern_match*` placeholders are approximated with regular expressions (VictoriaLogs uses a greedy matcher with extra boundary rules); `seq()` matches substrings in order without word-boundary checks. - The only constructs that remain `NOT_IMPLEMENTED` (with clear errors naming the construct) are the ones with no ClickHouse meaning: VictoriaLogs storage introspection (`block_stats`, `blocks_count`, `query_stats`, `value_type`, `histogram` buckets, `set_stream_fields`, `stream_context`), features that require the dynamic set of fields of a schemaless store (`facets`, `field_names`, wildcard field selectors like `foo*:filter`, stats over all fields like `sum(*)`, `unpack_json` without a `fields` list), `unpack_syslog` (a stateful wall-clock-dependent parser), and `collapse_nums` (its number-boundary rules require lookaround which RE2 lacks).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113207",
        "createdAt": "2026-08-04T00:00:57Z",
        "updatedAt": "2026-08-13T01:27:34Z",
        "timestamp": "2026-08-13T01:27:34Z",
        "metrics": {
          "reactions": 5,
          "comments": 5
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113208",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix array membership over an erased element holding one concrete type",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `has`, `indexOf`, `indexOfAssumeSorted`, `countEqual`, `mapContainsKey` and `mapContainsValue` missing rows that `=` matches when the array element type is type-erasing (`Dynamic` or `Variant`, including nested in `Tuple`): `has([1::UInt64]::Array(Dynamic), 1::UInt8)` returned 0 while `1::UInt64::Dynamic = 1::UInt8` is 1. The fix applies to a row whose own elements resolve to one concrete alternative and `equals` over that peeled pair succeeds; other rows keep the previous behaviour. ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/112953 At default settings these functions miss rows that scalar `=` matches, and `countEqual` breaks the `arrayCount(elem -> elem = x, arr)` equivalence its docs state. **Root cause.** `FunctionArrayIndex` tests equality with `IColumn::compareAt(...) == 0`, but for `ColumnDynamic`/`ColumnVariant` `compareAt` compares the variant type name (or discriminator) before the value: a sort order, not equality. Equal values under different variants never compare equal. **The change.** One dispatcher at both array entry points covers all six `FunctionArrayIndex` instantiations as constant, materialized and `LowCardinality`. For a type-erasing common type it peels an operand resolving to one concrete alternative and compares with the registered `equals`, folding each wrapper level's nullness at its own level so an outer NULL stays distinguishable from a nested one. `Tuple` recursion stops at `Array`/`Map`, also `compareAt`-based. **Scope, deliberately narrow, decided per ROW.** A block shares one flattened element column, so the decision is taken per row group: a row's answer follows from its own elements and the needle, not from which rows share its block. It declines, before any behaviour change, for a row mixing several concrete types, shared-variant rows, container alternatives, `LowCardinality` elements, and any pair `equals` rejects (the condition is `equals` succeeding, not the types, since comparability is partly value-dependent). Declined rows keep master's answer bit-for-bit. A `NULL` needle against a materialized erased array holding a NULL now matches: `has([NULL], NULL)` -> 1. No setting is added. **Validation.** New parallel-safe test `04706`: every cell asserted against an oracle in the same row, controls pinning each declined shape, and a group asserting three block partitions agree. A/B against pristine master over the 229 stateless tests reaching these functions: identical failure sets. 50/50: 100 OK. <details> <summary>Validation detail</summary> Every measurement pairs `SELECT lower(buildId())` against `readelf -n` on the serving binary. `#112953` touches this file but neither entry point. **Left for separate PRs**, each measured unchanged on both arms in a run where `has` itself moved 0 -> 1 on the same fixture: `hasAny`/`hasAll` (separate GatherUtils implementation); `arrayCompact`; container equality itself (`[1::UInt64::Dynamic] = [1::UInt8::Dynamic]` is 0 while the `Tuple` twin is 1, which is why this fix stops there); and the heterogeneous-row miss. **Mutations**, each rebuilt with the Build ID confirmed to move, the whole test re-run, then restored with it confirmed to return. Each reddens only what is named: | mutation | reddens | |---|---| | remove the new call | 15 lines; controls green | | guard on `Dynamic` only | the 2 `Variant` cells | | drop the NULL-vs-NULL match | the NULL cells | | bypass the dispatcher on the `Map` path | direct `has(map, key)` | | non-recursive, then one-level, `Tuple` recursion | the 3 `Tuple` cells; then `Tuple(Tuple(Dynamic))` | | decline constant arrays | the erased constant cell | | accept a row of several concrete types | the heterogeneous controls: the decline protects them | | never unwrap `Nullable` | the `Nullable(Tuple(Dynamic))` cells | | remove the `Map` cardinality normalisation | the 2 `Map` NULL-needle cells | | ignore a column's own nullness in the fold | the `LowCardinality`-needle cell | | decide per block, not per row group | the block-partition cells | **Performance** (debug, `max_threads=1`, median of 5): a 1000-element constant `Array(Dynamic)` over 2000 rows is unchanged at 0.023 s; a materialized 200k x 50 one goes 0.287 s -> **0.135 s**. **The updated reference** (`04338`, 2 lines) now enforces its own comment, \"`has(m, k)` must agree with `has(mapKeys(m), k)` for every row\": `has(map(NULL::Dynamic, 1), CAST(NULL, 'LowCardinality(Nullable(String))'))` was 0, the array path 1. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113208",
        "createdAt": "2026-08-04T00:05:50Z",
        "updatedAt": "2026-08-13T17:10:35Z",
        "timestamp": "2026-08-13T17:10:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113244",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix set index on an expression with a Nullable operand",
        "text": "A `SELECT` filtered on an expression that matches a `set` skip index raises `LOGICAL_ERROR`: ```sql CREATE TABLE events (id UInt32, INDEX b_set intDiv(id, 1000) TYPE set(100) GRANULARITY 1) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO events SELECT number FROM numbers(5000); SELECT count() FROM events WHERE intDiv(id, (SELECT 1000)) = 3; -- Code: 49. Unexpected return type from equals. Expected Nullable(UInt8). Got UInt8. ``` `SELECT 1000` is Nullable(UInt16), while the underlying column type of the index is plain UInt32, so when MergeTreeIndexConditionSet rewrites the predicate it loses the Nullable type and just saves that it is a UInt. However the equals function still expects a Nullable, and this mismatch throws an error in `ExpressionActions.cpp:executeAction`. ============================================ No `Nullable` appears anywhere in the schema or the query. A scalar subquery is typed `Nullable` because it may return no rows, and that is enough. In practice this shows up when the divisor or bucket size comes from a lookup or settings table instead of a literal; the typing is identical. Explicit wrappers reach the same failure: `nullIf(1000, 0)`, `toNullable(1000)`, `materialize(toNullable(1000))`. `MergeTreeIndexConditionSet` matched a query subexpression to an index key column by name only. Names are computed from constant-folded arguments and never include types, so `intDiv(id, (SELECT 1000))` renders exactly like the index expression `intDiv(id, 1000)` while carrying `Nullable(UInt32)`. The whole subtree was then replaced by the granule column while the enclosing `equals` was re-added with its already-resolved `IFunctionBase`, keeping the return type it was resolved with. `ExpressionActions::execute` binds inputs by name without checking types, so the granule supplied a non-`Nullable` column and the declared return type no longer matched what execution produced. The rebuilt condition is now typed from the granule side throughout. `atomFromDAG` binds the key column input to the type the granule block holds instead of the query-side type, and re-resolves the atom's function against the granule-side argument types instead of reusing the query-side `IFunctionBase`. The index therefore keeps pruning on these queries rather than being skipped: `EXPLAIN indexes = 1` reports the same 5/50 granules for `t % nullIf(19, 0) = 16` as for the plain `t % 19 = 16`. Two cases still fall back to `UNKNOWN_FIELD`, which leaves the granule unpruned and sends the query through the regular filter path. A function that cannot be re-resolved from `FunctionFactory` by name — internal casts, lambdas, parametric functions — and argument types the function rejects. And a key column that is already an `INPUT` of the filter DAG, which cannot be re-typed, because a second input under the same name would be left unbound where `ExpressionActions::execute` maps each name to one block column. The regression test covers both granule filtering paths, since `secondary_indices_enable_bulk_filtering` selects between `getPossibleGranules` and `mayBeTrueOnGranule`, and each reaches the same actions independently. An index whose expression is itself `Nullable` is unaffected in either direction, because the types agree and nothing is re-typed. Reproduces on the official release `26.8.1.727`. In builds with assertions enabled it aborts the server, which is how it surfaced in CI — 14 occurrences in 90 days across unrelated pull requests and every build configuration, through `tests/queries/0_stateless/01786_explain_merge_tree.sh` once the AST fuzzer wraps the modulus in a `Nullable` expression. Closes: https://github.com/ClickHouse/ClickHouse/issues/113234 Related: https://github.com/ClickHouse/ClickHouse/issues/113233 Related: https://github.com/ClickHouse/ClickHouse/issues/89802 Related: https://github.com/ClickHouse/ClickHouse/pull/111830 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `LOGICAL_ERROR` when reading a table with a `set` skip index defined on an expression and the query repeats that expression with a `Nullable` operand, for example `WHERE intDiv(id, (SELECT 1000)) = 3` over `INDEX b_set intDiv(id, 1000) TYPE set(100)`. A scalar subquery is enough to trigger it, since it is typed `Nullable`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113244",
        "createdAt": "2026-08-04T07:24:53Z",
        "updatedAt": "2026-08-13T01:52:52Z",
        "timestamp": "2026-08-13T01:52:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "george-larionov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113248",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not run the server-side AST fuzzer on the stress harness's own queries",
        "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) The stress test aborts with `Test script failed` before running a single test: ``` Error on processing query: Timeout exceeded while receiving data from server. Waited for 15 seconds, timeout is 15 seconds. (query: SELECT value FROM system.server_settings WHERE name = 'cannot_allocate_thread_fault_injection_probability') ... subprocess.CalledProcessError: ... returned non-zero exit status 159. ``` The fail-close verification in `install_thread_pool_fault_injection` (added in #104782) ran a single `clickhouse client` query with `--receive_timeout=15` and no retry, so one slow answer killed the whole job. The first commit retries it, mirroring `call_with_retry`, keeping the fail-close semantics: persistent failure or a zero probability still aborts the run. Retrying alone is not enough, because the server is not slow. The stress profile (`stress_tests.lib`) enables the server-side AST fuzzer for the `default` user with `ast_fuzzer_runs=5` and `ast_fuzzer_any_query=true`, and that also applies to the maintenance queries of the harness itself. The fuzzer runs as a query-finish callback, so the connection thread executes the five mutated follow-up queries before the response completes. In [Stress test (arm_release) on #113224](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113224&sha=297082adba654b8bafac191c6c61b5023bcc76ef&name_0=PR&name_1=Stress%20test%20%28arm%5Frelease%29) (https://github.com/ClickHouse/ClickHouse/pull/113224) the verification query read its 439 rows in 67 ms and the handler returned after 16.3 s, past the client's 15-second `receive_timeout`: ``` 21:12:45.393 executeQuery: Read 439 rows, 131.72 KiB in 0.067 sec. 21:12:45.395 ASTFuzzer: Fuzzed query: SELECT value FROM system.server_settings WHERE name = ... ... 5 fuzzed follow-up queries ... 21:13:01.597 TCPHandler: Processed in 16.274 sec. ``` Every one of the five queries the fuzzer fired on in that run was the harness's own: the `SYSTEM RELOAD CONFIG` and the verification query of `stress.py` (11.3 s and 16.3 s in the connection thread), the two `SELECT 1` readiness probes and the `SYSTEM STOP DISTRIBUTED SENDS` of `stress_tests.lib`. No test query was fuzzed, because `clickhouse-test` runs are already started with `ast_fuzzer_runs=0`. The second commit pins `ast_fuzzer_runs=0` on the harness's own queries, in `stress.py` and in `stress_tests.lib`. That is what `clickhouse-test` already does for its infrastructure queries (`clickhouse_execute_http`) and what `stress.py` already did for the smoke check and the hung check. An explicit value on the command line is marked `changed` even though it equals the default, so the client sends it and it overrides the profile. Besides the timeouts this also stops two hazards that `ast_fuzzer_any_query` made possible: * A fuzzed `DETACH DATABASE` / `KILL QUERY` from `prepare_for_hung_check`, or a fuzzed `SYSTEM STOP DISTRIBUTED SENDS` during shutdown, running some other statement. * Fuzzed copies of a `system.processes` query showing up in the very processlist the hung check is about to inspect. A fuzzed readiness probe can also burn most of the 30-second `receive_timeout` of `start_server`, which reports `Cannot start clickhouse-server` for a server that is up. Related: https://github.com/ClickHouse/ClickHouse/pull/109496 Related: https://github.com/ClickHouse/ClickHouse/pull/113224 Related: https://github.com/ClickHouse/ClickHouse/pull/104782",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113248",
        "createdAt": "2026-08-04T07:30:34Z",
        "updatedAt": "2026-08-13T11:42:36Z",
        "timestamp": "2026-08-13T11:42:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113266",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not inject random ORDER BY into queries planned to an intermediate stage",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110188 With `inject_random_order_for_select_without_order_by = 1`, `InjectRandomOrderIfNoOrderByPass` wrapped every top-level query into `SELECT * FROM (...) ORDER BY rand()`, including queries that are planned only up to an intermediate stage. For a `Merge` table with a `Distributed` child, the other children are planned to `WithMergeableState`; the wrapper made such a child plan complete its aggregation, so blocks without `AggregatedChunkInfo` reached `MergingAggregatedTransform`, and the query failed with a logical error (exception) `Chunk info was not set for chunk in MergingAggregatedTransform`. Now the injection pass is only added when the query is processed to stage `Complete`, so intermediate-stage plans (`Merge`/`Distributed` children) are left untouched, while the injection still applies to the user-facing query. Minimal reproducer (found by BuzzHouse on [this report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110188&sha=cdc23648d36878b136bf2c8d87ee9b0a29651412&name_0=PR&name_1=BuzzHouse%20%28arm_asan_ubsan%29), unrelated to that PR's changes — it reproduces on master): ```sql CREATE TABLE t_local (x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t_local SELECT number FROM numbers(100); CREATE TABLE t_dist (x UInt64) ENGINE = Distributed('test_shard_localhost', currentDatabase(), 't_local'); SET inject_random_order_for_select_without_order_by = 1; SELECT count() FROM merge(currentDatabase(), '^t_(local|dist)$'); -- LOGICAL_ERROR before this fix ``` ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a logical error (exception) `Chunk info was not set for chunk in MergingAggregatedTransform` when the setting `inject_random_order_for_select_without_order_by` is enabled and an aggregation query reads from a `Merge` table containing a `Distributed` child: the random `ORDER BY rand()` wrapper is no longer injected into queries planned only up to an intermediate stage.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113266",
        "createdAt": "2026-08-04T08:30:54Z",
        "updatedAt": "2026-08-13T01:01:26Z",
        "timestamp": "2026-08-13T01:01:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113289",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/113003 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a query plan optimization stall when a `WHERE` clause contains a large string constant with many dots and the table carries a `bloom_filter`, `tokenbf_v1`, `ngrambf_v1` or `text` skip index. Matching a filter column name against `JSONAllPaths(...)` index columns enumerated every dot split of the name, which made skip-index condition building quadratic in the constant's length. Closes #113003. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/113003 **What breaks.** A `SELECT` whose `WHERE` contains a large dotted string constant, over a table with a skip index, spends unbounded time in query plan optimization. The reporter measured 9.2 hours at 100% of one core with `max_execution_time = 300` set, on 26.4.3.37 and 26.7.1.1315. Nothing on that path observes cancellation, so `max_execution_time` fires and `KILL QUERY` is inert, no `QueryStart` row is written, and the handler thread is leaked until restart. Workaround was `SETTINGS use_skip_indexes = 0`. **Root cause.** `tryMatchJSONSubcolumnToIndex` reads a filter column name as `<json_column>.<path>`. Not knowing where the split is, it enumerated every dot split of the name and per split formatted a lookup key and scanned the index columns. The name embeds a folded constant verbatim (`position('a.a.a...', s)`), so its length is user-controlled, giving O(length^2). **The change.** Scan the index columns instead: keep entries shaped `JSONAllPaths(X)` and test each `X` against the name by prefix-plus-dot. Cost no longer depends on the name. Selection is deliberately unchanged (shortest matching `X`, ties to the first entry), which reproduces what walking dot positions left to right did. This is not `bloom_filter`-specific: `tokenbf_v1`, `ngrambf_v1` and `text` reach the same helper, so all four are fixed at once. **Validation.** The reporter's repro goes from 42.7s to 0.16s on a debug build. All 80 existing `EXPLAIN indexes = 1` cases in `04024_json_skip_index_*` pass unchanged; 3 new cases pin the previously untested ambiguous case where two index columns match one name. The new cost test asserts allocated bytes rather than wall clock, and was verified to fail on the pre-fix binary. The cancellation gap is real and separate; it ships as a follow-up.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113289",
        "createdAt": "2026-08-04T09:31:46Z",
        "updatedAt": "2026-08-13T08:03:04Z",
        "timestamp": "2026-08-13T08:03:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "can be tested",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113333",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Skip the aggregation hash-table stats cache key without GROUP BY keys",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/113037 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a performance regression where an aggregation without `GROUP BY` keys serialized its whole input plan subtree, including the full contents of constant-folded literals, on every execution to compute a hash-table-size cache key it can never use. A query such as `SELECT quantileMerge(arrayJoin(arrayMap(x -> state, range(5000000))))` spent about 63 ms per execution in query-plan optimization and now spends 0.15 ms. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/113037 Related: #107643 (introduced the pass) `setAggregationHashTableCacheKeys` admitted every `AggregatingStep` and serialized its input subtree to derive a cache key. `ActionsDAG::serialize` writes the full binary content of constant columns, so a folded `range(5000000)` literal materialized about 40 MB into a buffer on every execution. A keyless aggregation cannot use that key. `AggregatedDataVariants::chooseMethod` starts with `if (keys_size == 0) return Type::without_key;`, `init` discards the `size_hint` for `without_key`, and `without_key` is absent from `APPLY_FOR_VARIANTS_CONVERTIBLE_TO_TWO_LEVEL`. Such aggregations also wrote a meaningless entry (`median_size` = 1) into the shared statistics cache, evicting useful ones. The fix admits an `AggregatingStep` only when `!getParams().keys.empty()`. The early return when no aggregation is present, the per-node `try/catch` and the compute-all-then-stamp-all ordering are unchanged, and the key stays an optional flagged field in `AggregatingStep::serialize`, so no format or setting default changes. Validated on x86_64 (the issue reports aarch64), release, against pristine master with distinct build IDs. `QueryPlanOptimizeMicroseconds` for the reported query, median of 9: 63000 to 154. Keyed aggregations keep preallocating, and the existing keyed tests are the regression guard here; a keyless assertion would need a process-global asynchronous metric or a timing counter, both flaky. <details><summary>Validation matrix</summary> - `04509_hash_table_sizes_stats_table_functions` still reports 650000 and 520000 on both branches. - `HashTableStatsCacheEntries` grows by 12 for 12 distinct keyed aggregations on both branches; for 12 keyless ones it grows by 12 on master and by 0 with this change. - `hash_table_sizes_stats`, `hash_table_sizes_stats_small` and `group_by_consecutive_keys` queries: ratios 0.93 to 1.02. - Full `aggregat`/`group_by` stateless suite, 460 tests: identical per-test verdicts. - 50 runs each of `04506` and `04507`: 100 OK, 0 FAIL on both branches. - Wall clock for the reported query, min of 15: 275 ms to 214 ms. </details> `optimizeJoin` and `considerEnablingParallelReplicas` reuse these subtree hashes and run after this pass, so with a keyless aggregation under a join their key values shift: the flag bit recording whether a key was stamped now differs. `HashTablesStatistics` is process-local, so the cost is one un-preallocated execution. Constant serialization is still costly for keyed aggregations, about 62 ms here, and is left to a separate fix. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1275` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113333",
        "createdAt": "2026-08-04T14:04:48Z",
        "updatedAt": "2026-08-13T13:32:43Z",
        "timestamp": "2026-08-13T13:32:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-performance",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "nickitat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113334",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Introduce zk leader metrics to Keeper mntr",
        "text": "Adds leader-only Keeper `mntr` metrics: - `zk_leader_uptime` - `zk_sum_election_time` - `zk_cnt_election_time` - `zk_sum_leader_unavailable_time` - `zk_cnt_leader_unavailable_time` `zk_leader_uptime` starts at NuRaft `BecomeLeader`. The cumulative election and leader-unavailability metrics sample `isLeaderAlive` once per `heart_beat_interval_ms`; election completion is recorded at `BecomeLeader`, while leader-unavailability completes when polling observes a live local leader. `srst` resets all four cumulative values. The intended alternative was exact NuRaft lifecycle callbacks: an election-start callback before pre-vote or vote, plus a leader-ready callback for leader-unavailability completion. That requires changing vendored NuRaft and an upstream contribution, whose merge timeline could delay these Keeper monitoring metrics. This PR therefore delivers the metrics now through existing NuRaft state, while documenting the sampling limitation: boundaries can differ by up to one heartbeat interval and short no-leader windows can be missed. Also added keeper-only metrics - `KeeperLastLeaderElectionTime` - `KeeperLastLeaderUnavailableTime` ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Introduce emulated ZooKeeper metrics in Keeper `mntr` command: `zk_leader_uptime`, `zk_sum_election_time`, `zk_cnt_election_time`, `zk_sum_leader_unavailable_time`, `zk_cnt_leader_unavailable_time`; Also introduce related leader-oriented metrics for Keeper only: `KeeperLastLeaderElectionTime`, `KeeperLastLeaderUnavailableTime`",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113334",
        "createdAt": "2026-08-04T14:08:25Z",
        "updatedAt": "2026-08-13T10:50:45Z",
        "timestamp": "2026-08-13T10:50:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested",
          "pr-autogenerated-docs"
        ],
        "author": "UberDever",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113347",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Write wide integers in Parquet as Decimal",
        "text": "Previously UInt128/UInt256/Int128/Int256 were written as `FIXED_LEN_BYTE_ARRAY`, but as little-endian. Since little-endian is not lexicographically sortable, there were no statistics and no way to prune row groups and pages. This was not intentional, but rather an accidental artifact from when Parquet serialization was first introduced in ClickHouse. However, we cannot just change the endianness and break existing code/data. Here instead we switch to `DECIMAL` encoding in Parquet, which is big-endian. This enables row group and page pruning for querying for queries that use filtering on wide integer columns. The encoding is a little bit akward with 17 or 33 bytes per integer, but it enables other data consumers to have a correct data interpretation hint. Since the new data cannot be read by older versions, this new serialization mode is behind `output_format_parquet_wide_integer_as_decimal` setting, which defaults to `0` to preserver backward compatibility. The new version can still read old data (but not the other way around). <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add opt-in Parquet serialization of UInt128/UInt256/Int128/Int256 as DECIMAL to enable row group and page pruning.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113347",
        "createdAt": "2026-08-04T15:40:06Z",
        "updatedAt": "2026-08-13T05:18:16Z",
        "timestamp": "2026-08-13T05:18:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "bobrik",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113357",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cache the finalized cardinality in uniq statistics",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113038 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Speed up query planning when column statistics are used. `uniq` and `uniq_v2` statistics now cache the estimated number of distinct values instead of recomputing the sketch on every request, which join order optimization issues many times per query. This is most visible on aarch64. Closes #113038. ### Description `IStatistics::estimateCardinality` is a pure function of the aggregate state, but both implementations recomputed it on every call, and plan optimization calls it many times per query: once per column with statistics in `ConditionSelectivityEstimator::estimateRelationProfileImpl`, and once or twice per equality atom. For `uniq_v2` that recomputation is expensive. `uniq_v2` is `uniqCombined64` with `K = 12`, which selects the fully functional `Denominator` specialization whose `get` (`src/Common/HyperLogLogCounter.h:180-189`) walks all 54 rank buckets in `long double`. On aarch64 that is software binary128, so every step compiles to `__multf3` / `__floatunsitf` / `__addtf3` libcalls; on x86-64 it is native x87 arithmetic. That is why the regression is aarch64 only. Since 26.7 the default `auto_statistics_types` includes `uniq_v2`. The fix memoizes the finalized value in the object owning the state, as `cardinality + 1` so `0` means \"not computed yet\" without reserving a representable cardinality as a sentinel (0 distinct values is legal for an all NULL column). The member is a `mutable std::atomic<UInt64>` because the estimator is shared between concurrent queries; relaxed ordering suffices as all racers compute the identical value. No invalidation is needed: every caller finishes building and merging before any cardinality is read, which I measured across 10387 mutator entries without a single live memo. Both implementations are fixed: `findUniqStats` prefers `Uniq` over `UniqV2`, so `StatisticsUniq` is a live carrier too. On this PR's own arm_release Performance Comparison, `JoinOptimizeMicroseconds` drops 73% on TPC-DS Q14 and 63% on TPC-H Q20, the two queries in the report, with `server_time` -31% and -51%. Estimates and the chosen plan are unchanged (1206 `system.parts_columns` estimate rows and the `EXPLAIN indexes=1` output of a 6 way join are byte identical). The analysis, patch direction and aarch64 measurements are @ egor-click's. This differs from the patch in the issue in replacing the `std::numeric_limits<UInt64>::max()` sentinel with the `+1` encoding, and in fixing `StatisticsUniq` too.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113357",
        "createdAt": "2026-08-04T16:52:54Z",
        "updatedAt": "2026-08-13T15:11:13Z",
        "timestamp": "2026-08-13T15:11:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "hanfei1991"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113376",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `JSONAllValues` text index probe coercion",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113022 Fix incorrect result filtering when a `JSONAllValues` text index serializes a comparison constant using a representation that differs from the JSON subcolumn. This includes value-changing coercions such as `IPv4` to `UInt32`, types such as `Bool` whose semantic equality does not imply identical text, and `DateTime` representations that depend on time zones. Equality and `has` predicates now use the index only when the probe representation is compatible with the statically typed JSON subcolumn or array element type. Equality and `IN` predicates on runtime-typed paths, casts from `Dynamic` paths, and values with session-dependent serialization are evaluated without this index. Safe direct access and identity casts remain accelerated. Casts to `String` remain accelerated for statically typed values with stable serialization. Other value-changing casts decline index use because their stored representation cannot be inferred safely. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix incorrect results from `JSONAllValues` text indexes when comparison values require type coercion or have session-dependent or runtime types.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113376",
        "createdAt": "2026-08-04T19:05:43Z",
        "updatedAt": "2026-08-13T09:36:54Z",
        "timestamp": "2026-08-13T09:36:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "manual approve",
          "can be tested"
        ],
        "author": "rorylshanks",
        "state": "open",
        "assignees": [
          "CurtizJ"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113382",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix division by zero in `ReadFromMergeTree` when `StorageMerge` truncates the number of streams",
        "text": "The `AST fuzzer (amd_debug)` crashed the server with `Integer divide by zero` on ```sql SELECT id FROM merge(currentDatabase(), '^t$') ORDER BY id DESC SETTINGS max_streams_to_max_threads_ratio = 1073741824, max_threads = 4 ``` The planner computes `max_streams = max_threads * max_streams_to_max_threads_ratio = 2^32`, which passes the existing overflow check (it only rejects values that do not fit into `size_t`). `ReadFromMerge::createPlanForTable` then passed `UInt32(streams_num)` to `storage->read`, truncating `2^32` to `0`. The child `ReadFromMergeTree` divides by the requested number of streams in `spreadMarkRangesAmongStreamsWithOrder` (`(info.sum_marks - 1) / num_streams`), so debug and sanitizer builds crash with SIGFPE (release builds happen to eliminate the dead division, since its only use is in a loop that runs zero times). `IStorage::read` takes `size_t num_streams`, so the cast was a pointless leftover — this change removes it and adds a regression test that crashes debug builds without the fix. Additionally (per review), `StorageMerge` multiplied the requested number of streams by `max_streams_multiplier_for_merge_tables` (clamped to the number of selected tables) with unchecked `Float64` -> `size_t` casts, while the planner only bounds-checks `max_threads * max_streams_to_max_threads_ratio`. Both multiplications now go through a shared helper that throws `PARAMETER_OUT_OF_BOUND` (like the planner check) when the product does not fit into `size_t`, with a second regression test. Also (per review), `spreadMarkRangesAmongStreams` and `spreadMarkRangesAmongStreamsWithOrder` clamp an unnecessarily large number of streams down to the amount of data, but the check itself computed `num_streams * min_marks_for_concurrent_read`, which wraps around for huge stream counts. The clamp was then skipped and the ordered path went on to `split_parts_and_ranges.reserve(num_streams)`, throwing `std::length_error` in every build — reachable both directly on a `MergeTree` table and, now that the stream count is no longer truncated, through a `Merge` table. The comparison now uses a division, which is equivalent (`min_marks_for_concurrent_read` is always at least one) and cannot overflow, with a third regression test. Also (per review, and confirmed by the AST fuzzer on this PR), a streaming read (`FROM ... STREAM`) bypasses the `spreadMarkRanges*` helpers entirely: `groupPartitionsByStreams` created one `MergeTreeCommitOrderSequentialSource` per requested stream with no clamp, so a huge `max_threads * max_streams_to_max_threads_ratio` product threw `std::length_error` from `pipes.reserve` (or would exhaust memory for merely absurd values). Since a streaming read has no marks by which the stream count could be clamped, such values are now rejected with `PARAMETER_OUT_OF_BOUND` (with a cap of one million, mirroring the limit on the number of threads in `MergeTreeReadPool`), with a fourth regression test. Found by the AST fuzzer on an unrelated PR: [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113010&sha=59507af89c33f62942e855b50b3e66704cbcf2e9&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29) Related: https://github.com/ClickHouse/ClickHouse/pull/113010 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a crash (arithmetic exception in debug and sanitizer builds) when reading from a `Merge` table with a very large `max_streams_to_max_threads_ratio`: the number of streams was truncated to 32 bits, so multiples of `2^32` became zero. Also, throw an exception instead of undefined behavior when the product of the number of streams and `max_streams_multiplier_for_merge_tables` exceeds the range of `size_t`, fix a `std::length_error` when reading a `MergeTree` table with a huge `max_streams_to_max_threads_ratio`, and reject huge stream counts on streaming reads (`FROM ... STREAM`) instead of trying to create a source per stream.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113382",
        "createdAt": "2026-08-04T20:03:50Z",
        "updatedAt": "2026-08-13T01:18:16Z",
        "timestamp": "2026-08-13T01:18:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113383",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Push tuple element predicates into Parquet and ORC subcolumn reads",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/112575 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Filter pushdown now works for `Tuple` subcolumns in Parquet and ORC files. A predicate such as `WHERE tup.1 = 555555` over `file()`, `s3()` or `url()` now prunes row groups and row index strides using the tuple element's own statistics instead of reading the whole file. ### Description Closes: #112575 Three independent defects, all needed for the reported query to prune. Parquet's reader itself was fine. **The analyzer never produced a subcolumn.** `StorageFile`/`StorageURL`/`StorageObjectStorage` return `false` from `supportsOptimizationToSubcolumns`, so `tupleElement(tup, 1)` was never rewritten to `tup.1`; the pass also accepted only `TableNode`, skipping table functions. That blanket `false` keeps #106147 fixed (`NOT_FOUND_COLUMN_IN_BLOCK` on `.null` in PREWHERE), so rather than flipping it this adds a narrow `supportsOptimizationToTupleElementSubcolumns` virtual defaulting to the existing one, with a `{Tuple, tupleElement}`-only allow-list. `04303_object_storage_prewhere_isnotnull_subcolumn` passes unmodified. Source identity also moves to the table-expression node: every `file()` resolves to the same `_table_function.file` ID, so two such sources shared a key. Accepting table functions generally also makes `format()`, `values()` and `view()` eligible for the other transformers; those storages already answer `supportsSubcolumns()`, so the default covers them. **ORC's search argument builder resolved top-level names only,** so any dotted name emitted `YES_NO_NULL` while the read path in the same file resolved them recursively. Resolving recursively also reaches the flattened-`Nested` descent, which rewrites the type it is given, so the builder keeps the key's own type: an array-typed predicate over a flattened `Nested` leaf is not pushed, since scalar element statistics cannot decide it. **ORC built its KeyCondition from the reader header,** which carries only the parent column, so the predicate degraded to `unknown`. ORC now passes `initKeyConditionOnce` a local copy extended with the tuple element paths the filter references; `FormatFilterInfo`, the Parquet call site and the reader header are untouched. Admission requires a named tuple at every level (unwrapping `Nullable`/`LowCardinality`/`Array`), which refuses Map `.keys`/`.values`: they use `SubstreamType::TupleElement` but have no per-element statistics. <details> <summary>Measurements (100k rows, one row group / stride per 10k)</summary> | Arm | master | this PR | |---|---|---| | ORC `WHERE tup.1 = 55555` | 200000 rows read | **20000** | | ORC `WHERE id = 55555` (control) | 20000 | 20000 | | Parquet `WHERE tup.1 = 55555` | 7 row groups / 0 pruned | **1 / 6** | | Parquet `WHERE id = 55555` (control) | 1 / 6 | 1 / 6 | Results identical in every arm. Refusal arms (Map `.keys`/`.values`, `.null`, `.size0`, unnamed tuple, `Array(Tuple)`, type-mismatch structure hint) return correct results with pruning off. Multi-level `tup.2.1` stays unpruned: correct, and a separate optimization. Regression sweep over `*functions_to_subcolumns*`, `*tuple_element*`, `*_parquet_*`, `*_orc_*` and the named pushdown tests, run on this build and on an unmodified master build for attribution: no regression attributable to this change. New tests are 50/50 green under randomized settings. </details> Also noticed, not touched here: reading a standalone dotted ORC tuple element with its inferred type returns column defaults, because `Nested::flatten` does not descend a `Nullable(Tuple)`. #109741 (open) rewrites that helper for the `Arrow` spelling.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113383",
        "createdAt": "2026-08-04T20:17:22Z",
        "updatedAt": "2026-08-13T15:12:26Z",
        "timestamp": "2026-08-13T15:12:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 18
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113401",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix paimon timestamp precision",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> fix: https://github.com/ClickHouse/ClickHouse/issues/112768 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables partitioned by a `TIMESTAMP` or `TIMESTAMP WITH LOCAL TIME ZONE` column of precision above 3. Such tables previously failed with `scale 6 is not supported, only support scale <= 3` before returning any row, which affected every timestamp-partitioned table written by Spark, since Spark maps both `TIMESTAMP` and `TIMESTAMP_NTZ` to Paimon `TIMESTAMP(6)`. Partition pruning on such a column now uses the full precision as well.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113401",
        "createdAt": "2026-08-05T01:12:47Z",
        "updatedAt": "2026-08-13T17:09:40Z",
        "timestamp": "2026-08-13T17:09:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "JiaQiTang98",
        "state": "open",
        "assignees": [
          "hanfei1991"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113443",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add a failpoint inside the Paimon incremental-read at-most-once window",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/102343 Paimon incremental reads have documented at-most-once semantics: the Keeper watermark advances at file-collection time, before the collected batch is delivered, so a crash inside that window loses the batch. This window has been untestable — nothing can deterministically crash a server between two statements inside one query's execution. This adds a pauseable failpoint, `paimon_incremental_read_pause_after_watermark_commit`, exactly between the watermark commit and delivery. Integration tests can enable it, observe the committed watermark in Keeper while the read is paused, restart the server, and assert the batch is lost — pinning the at-most-once contract deterministically. The test flips the day delivery becomes at-least-once. The failpoint is inert unless enabled via `SYSTEM ENABLE FAILPOINT`, like every other failpoint. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113443",
        "createdAt": "2026-08-05T08:43:23Z",
        "updatedAt": "2026-08-13T08:24:23Z",
        "timestamp": "2026-08-13T08:24:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-ci"
        ],
        "author": "zlareb1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113448",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Honor `use_statistics_cache` when loading per-part statistics",
        "text": "`IMergeTreeDataPart::getEstimates` (introduced in #87241) builds the per-part statistics estimates used by part pruning and by `system.parts_columns`. Before this change it always consulted a per-part estimates cache and ignored the `use_statistics_cache` setting (introduced in #88670): a session that sets `use_statistics_cache = 0` still got the cached values on this path. That made the opt-out inconsistent with the selectivity-estimator path, which already honors `use_statistics_cache`. This change makes `IMergeTreeDataPart::getEstimates(bool use_cache)` honor the setting: when `use_statistics_cache = 0` it loads statistics directly from disk and neither reads nor populates the per-part cache, matching the selectivity-estimator path. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `use_statistics_cache = 0` now also bypasses the per-part statistics estimates cache used by part pruning and `system.parts_columns`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113448",
        "createdAt": "2026-08-05T10:03:58Z",
        "updatedAt": "2026-08-13T03:20:42Z",
        "timestamp": "2026-08-13T03:20:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "zoomxi",
        "state": "closed",
        "assignees": [
          "hanfei1991"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113450",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix reading Paimon tables with a nullable ARRAY or MAP column",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113337 Related: https://github.com/ClickHouse/ClickHouse/pull/113425 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables that contain a nullable `ARRAY` or `MAP` column. Such a table could not be read at all, because the schema mapper wrapped the composite type in `Nullable`, which ClickHouse forbids, so both `DESC` and `SELECT` failed with `Nested type Array(Nullable(Int32)) cannot be inside Nullable type`. A nullable composite column is now mapped to a non-`Nullable` composite type and a `NULL` value is read as an empty one. Closes #113337. ### Description This takes over #113425 by @zlareb1, who closed it and asked me to carry it forward. The diagnosis and the fixture are his; the source change uses the in-tree capability gate rather than deleting the wrap. **What breaks.** Paimon columns are nullable by default, so a plain `CREATE TABLE paimon.default.t (f ARRAY<INT>)` from Spark produces an unreadable table. `Paimon::DataType::parse` applied its `if (nullable)` wrap in the composite branch as well as the scalar one, and `DataTypeArray`/`DataTypeMap` return `canBeInsideNullable() == false`, so `DataTypeNullable`'s constructor threw. This happens while parsing the schema, so it takes out the whole table rather than one column. Affects `paimonS3`/`paimonLocal`/`paimonAzure`, their `*Cluster` variants, the `Paimon*` engines and the REST catalog. The engines need `allow_experimental_paimon_storage_engine`; the table functions do not. **The change.** The two inner wraps become a single `makeNullableSafe`, which wraps only when the type permits it. Neither the Iceberg nor the DeltaLake schema processor wraps a composite in `Nullable` (Iceberg gates on `canBeInsideNullable()`; DeltaLake keeps the wrap in its scalar branches only), so this aligns Paimon with them. The scalar wrap is untouched, so inner nullability survives: the fixture reads as `Array(Nullable(Int32))` and `Map(String, Nullable(Int32))`. The gate, rather than deletion, keeps the wrap available for a future `ROW`, whose `DataTypeTuple` does permit it. A `NULL` composite reads as an empty one, the Parquet reader's documented behaviour, so the two become indistinguishable. That is forced by the type system and matches Iceberg, DeltaLake, Arrow, ORC and Avro. **Validation.** New test `04757_paimon_nullable_composite_types` over @zlareb1's fixture fails on master with the error above and passes with the fix. The ten existing Paimon tests are identical on both binaries. Restoring either wrap individually reddens the new test on its own message. 50/50 randomized runs pass. `Types.h` is absent on 25.8, so 26.3 through 26.7 are affected; `must-backport` labels look appropriate.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113450",
        "createdAt": "2026-08-05T10:05:17Z",
        "updatedAt": "2026-08-13T16:40:49Z",
        "timestamp": "2026-08-13T16:40:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced",
          "v26.5-must-backport"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "JiaQiTang98"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113482",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #109932 to 26.6: Fix crash in direct JOIN over MergeTree with PREWHERE (shared PrewhereInfo corrupted by column pruning)",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/109932 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31008588537/job/92314616034)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113482",
        "createdAt": "2026-08-05T13:25:04Z",
        "updatedAt": "2026-08-13T13:07:48Z",
        "timestamp": "2026-08-13T13:07:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "diegomestre2",
          "antaljanosbenjamin"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113505",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "S3 tables engine",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): S3 tables engine catalog for datalakes. Same as https://github.com/ClickHouse/ClickHouse/pull/103220, but with working INSERT",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113505",
        "createdAt": "2026-08-05T14:34:19Z",
        "updatedAt": "2026-08-13T17:52:39Z",
        "timestamp": "2026-08-13T17:52:39Z",
        "metrics": {
          "reactions": 3,
          "comments": 2
        },
        "labels": [
          "pr-feature"
        ],
        "author": "scanhex12",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113512",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "[WIP] Seal-gated reading: gate the probe side of a hash JOIN on the runtime filter and prune read ranges by it",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> <details> <summary>Claude</summary> ## Motivation JOIN runtime filters (`enable_join_runtime_filters`) filter probe-side rows only after they were read: the row-level `__applyFilter` conjunct discards non-matching rows, and the read-time index analysis from #109085 (`enable_join_runtime_filters_index_analysis`) skips granules inside already-created read tasks. In both cases the probe side starts reading concurrently with the build side, before the filter exists, so the early tasks are read unpruned. This PR adds a stronger, structural variant for the case when the probe-side join key is a prefix of the table's primary key: the probe side does not read anything until the build side completes, and the completed filter then prunes whole mark ranges by the primary key *before* read tasks are created. ## Approach The gating is expressed as an edge of the query pipeline, not as a waiting state: - `FillingRightJoinSideTransform` gets an optional \"seal\" output port. Exactly one of the concurrent filling transforms — the one that completes the build — emits a single seal chunk carrying the completed runtime filter in its chunk info; the rest just finish. - On the probe side, the sources of a gated `ReadFromMergeTree` are replaced by `SealGatedReadTransform`s: source-like processors whose only input is the seal. Until the seal arrives, the executor has nothing to schedule below them, so the read pools never cut a task. On the seal, the filter is handed to a `RuntimeFilterReadRangesRefiner` (installed on the read pool), which turns it into a primary key `KeyCondition` — an exact `IN`-set, or the `[min, max]` envelope when the exact set overflowed into a bloom filter — and drops non-matching mark ranges at task-cut time, reusing the refiner contract of the MergeTree read pools. - The plan-level pass `markSealGatedReading` finds hash joins whose pushed-down `__applyFilter` conjunct references a probe-side primary key column, and marks the join step and the reading step. `JoinStep::updatePipeline` then wires the seal to the pending seal inputs collected by the `Pipe`. - Everything fails open: a gated read whose seal cannot be wired (the build side of a join, `YShaped`/by-shards pipelines, cancellation) is fed from a `NullSource` and reads ungated with row-level filtering only. FINAL, parallel replicas, and joins sharded by primary key ranges are not gated. Both the default multi-threaded pool and the in-order reading paths are gated (including `max_threads = 1` and reads in primary key order). On a gated read, the read-time index analysis by the same runtime filter is skipped as redundant — the refiner is already granule-exact through the primary key; filters of other joins are kept. Enabled by the experimental setting `enable_join_seal_gated_reading` (default off) on top of `enable_join_runtime_filters`. This is also groundwork for epoch-based (punctuated) collocated joins, where the same reader will consume a stream of per-epoch seals. ## Results On a 90M-row probe table (`ORDER BY k`, warm cache; `tests/performance/join_seal_gated_reading.xml`, CI perf host numbers): - 100 sparse build keys (exact `IN`-set path): 172 ms → 12 ms - 1M build keys in a narrow band (bloom overflow, `[min, max]` envelope path): 645 ms → 54 ms Compared with the read-time granule pruning by the same runtime filter (`enable_join_runtime_filters_index_analysis` + `use_skip_indexes_on_data_read`), the single-query latency on local storage is on par (the probe side cannot run far ahead of the build even ungated: the join does not consume it until the hash table is ready, so port backpressure stalls it after about a chunk per stream). The structural difference of gating is that no read or prefetch is issued for pruned ranges at all, and that it is the seam for the per-epoch seals of collocated joins. The stateless test `04653_join_seal_gated_reading` asserts result equality with ungated execution, the pipeline structure, `ReadPoolRangeRefinerDroppedMarks`/`read_rows` contrast, the fail-open shapes, and the suppression of the redundant read-time analysis. </details> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental setting `enable_join_seal_gated_reading` (default off). When enabled together with `enable_join_runtime_filters` and the probe-side join key is a prefix of the table's primary key, the probe side of a hash JOIN starts reading only after the build side completes, and the completed runtime filter prunes whole mark ranges by the primary key before read tasks are created, instead of only filtering rows after they were read.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113512",
        "createdAt": "2026-08-05T15:13:35Z",
        "updatedAt": "2026-08-13T17:43:34Z",
        "timestamp": "2026-08-13T17:43:34Z",
        "metrics": {
          "reactions": 1,
          "comments": 3
        },
        "labels": [
          "pr-improvement",
          "pr-performance"
        ],
        "author": "KochetovNicolai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113534",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a mixed JOIN ON condition evaluated over mismatched column types for a dictionary",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112831 A mixed join condition is a cross-side non-equi residual in the `ON` clause, for example `ON (t.key = d.key) AND (t.a * 10 < d.a)`. Only the hash family evaluates one: `HashJoin` passes `TableJoin::getMixedJoinExpression` into `AddedColumns::additional_filter_expression` and applies it while matching, resolving the right columns it needs **by name** against the stored right blocks. `buildPhysicalJoinImpl` drops the `join_use_nulls` conversion of the right columns from the right-side expression when the right side is a prepared storage, because such a storage delivers its columns already converted, and for a key-value storage it merges the conversion back into the right-side DAG under the original column names. But `DirectKeyValueJoin` declines a mixed condition, so a join onto a dictionary that carries one runs an ordinary algorithm which reads the dictionary as a stream - and is handed a right side whose `a` is `Nullable(UInt32)` under the very name the mixed condition declared as `UInt32`. `buildAdditionalFilter` creates the column from the declared type and fills it with `insertFrom` from the stored one, so a `ColumnUInt32` was filled from a `ColumnNullable`. ```sql CREATE TABLE dsrc (key UInt64, a UInt32) ENGINE = Memory; INSERT INTO dsrc VALUES (1, 100), (2, 20); CREATE DICTIONARY dict (key UInt64, a UInt32) PRIMARY KEY key SOURCE(CLICKHOUSE(TABLE 'dsrc')) LIFETIME(0) LAYOUT(FLAT()); CREATE TABLE t (key UInt64, a UInt32) ENGINE = Memory; INSERT INTO t VALUES (1, 1), (2, 2), (3, 3); SELECT count(), sum(d.a) FROM t LEFT ANY JOIN dict AS d ON (t.key = d.key) AND (t.a * 10 < d.a) SETTINGS allow_experimental_join_condition = 1, join_use_nulls = 1, join_algorithm = 'hash'; ``` Only key 1 satisfies the residual, so the answer is `3, 100`. A release build returned `3, 120`, having matched key 2 as well; a sanitizer build aborted on the `IColumn::insertFrom` type assertion. The fix keeps the conversion in the right-side expression when a key-value storage carries a mixed condition, so the join applies `join_use_nulls` itself exactly as it does for an ordinary table and the mixed condition keeps the non-`Nullable` columns it was built on. `StorageJoin` is unaffected - it rejects a mixed condition outright with `INCOMPATIBLE_TYPE_OF_JOIN`. `HashJoin` now also compares the type and not only the name when it resolves those columns, so a future divergence throws where both types are still known instead of reading a column through a mismatched `IColumn` interface. The shape has been reachable with an explicit `join_algorithm = 'hash'` for as long as the mixed condition and the prepared-storage handling have coexisted. https://github.com/ClickHouse/ClickHouse/pull/112831 made `DirectKeyValueJoin` decline a mixed condition, which put it on the default `join_algorithm` path and into the new `04667_mixed_join_condition_algorithm_selection` test, so the assertion started firing in most stress runs from 2026-08-04 - STID `2508-30f6` and `2508-2e0c`, which differ only by the concurrent-hash-join wrapper in the stack. No tracking issue was open for either STID. CI report the fix was written from: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=0f1b923e6247ef3d8409635a69738c7f3f02b84a&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results and, in a build with assertions enabled, an aborted assertion for a `JOIN` onto a dictionary whose `ON` clause has a non-equi condition over both tables, such as `ON (t.key = d.key) AND (t.a * 10 < d.a)`, when `join_use_nulls` is enabled. The condition was evaluated over a column read through a mismatched type, so rows could match arbitrarily.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113534",
        "createdAt": "2026-08-05T17:17:26Z",
        "updatedAt": "2026-08-13T00:57:05Z",
        "timestamp": "2026-08-13T00:57:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced",
          "v26.7-must-backport"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113553",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix data race when tracing profile events",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a theoretical data race in profile events when using the `trace_profile_events_list` setting to write stack traces of certain profile events to `system.trace_log`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113553",
        "createdAt": "2026-08-05T18:58:10Z",
        "updatedAt": "2026-08-13T00:30:12Z",
        "timestamp": "2026-08-13T00:30:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-must-backport",
          "pr-critical-bugfix"
        ],
        "author": "mstetsyuk",
        "state": "open",
        "assignees": [
          "Michicosun"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113557",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check table name length on RENAME DATABASE unconditionally",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/101747 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/101747 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `RENAME DATABASE` accepting a target name long enough to make the database's tables impossible to drop. The table name length check now runs regardless of `check_table_dependencies`, and it also covers detached tables, which were never checked and so were affected even at default settings. Closes #101747. ### Description `RENAME DATABASE` could leave every table in the database permanently undroppable: `DROP TABLE` then failed with `Code: 1001 ... in rename: File name too long`, recoverable only by renaming the database back to a shorter name. Two independent holes caused it. The length check sat inside the dependency-check guard, so `SET check_table_dependencies = 0` skipped it. The two invariants are unrelated: the check enforces a filesystem limit on `metadata_dropped/{db}.{table}.{uuid}.sql`, while the guard expresses a preference about dependency validation. Detached tables were never checked at all, and this half needed **no** non-default setting. `DETACH TABLE` moves the table out of the attached-tables map, so the existing loop could not see it; `ATTACH TABLE` does not re-check the length and re-attaching binds the table to the new database name. A database whose only table is detached has an empty attached-tables map, so the loop had nothing to iterate even at default settings. The check now runs in its own loop over both containers, before the symlink removal, the metadata `moveFile` and `updateDatabaseName`, so a rejection leaves no partial rename. `DatabaseReplicated::renameDatabase` delegates here before writing to ZooKeeper and is covered; `DatabaseOrdinary` has no override and throws `NOT_IMPLEMENTED`. A `RENAME DATABASE` that previously succeeded under `check_table_dependencies = 0` can now return `ARGUMENT_OUT_OF_BOUND`. The newly rejected renames are exactly those that would have produced undroppable tables (proof below), and renaming a database *shorter* only raises the limit, so the documented recovery path is unaffected. No setting default changes, so `SettingsChangesHistory.cpp` is not updated. New test `04700_rename_database_name_length_gate` has 13 arms: 5 that flip, and 8 controls including an accept/reject pair on the same database length that differ only by table name length. <details><summary>Why making the check unconditional does not over-reject</summary> `computeMaxTableNameLength` returns `min(N - 13, max(0, N - 42 - esc_db))` where `N` is the filesystem name limit, `esc_db` the escaped database name length. Since `42 > 13`, the second term always binds, so ``` reject <=> esc_db + esc_tbl > N - 42 ``` The dropped-metadata filename built by `DatabaseCatalog::getPathForDroppedMetadata` is `{db}.{table}.{uuid}.sql`, i.e. `esc_db + 1 + esc_tbl + 1 + 36 + 4` bytes, so ``` does not fit <=> esc_db + esc_tbl + 42 > N <=> esc_db + esc_tbl > N - 42 ``` The same inequality. The check rejects exactly the pairs whose table would be undroppable, and no others. Measured at `N = 255`: | esc_db | limit | esc_tbl | verdict | drop segment | droppable? | |---|---|---|---|---|---| | 200 | 13 | 2 | accept | 244 | yes | | 200 | 13 | 20 | reject | 262 | no | | 211 | 2 | 2 | accept | 255 | yes | | 214 | 0 | 2 | reject | 258 | no | Rows 1 and 2 are the discriminating pair in the test: same database length, opposite verdicts, so no length-blind rule satisfies both. Row 3 is the exact-boundary control, and the test drops that table afterwards to prove an accepted rename really leaves it droppable. `max_to_drop` is unconditionally the smaller term, so splitting the limit per engine would return the same value and ship dead code. Warning instead of refusing was also considered and rejected: the rename would still create undroppable tables, and the three other call sites of the check all throw `ARGUMENT_OUT_OF_BOUND`. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113557",
        "createdAt": "2026-08-05T19:30:13Z",
        "updatedAt": "2026-08-13T02:12:34Z",
        "timestamp": "2026-08-13T02:12:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "diegomestre2",
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113558",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix missing materialized-CTE gate when a `Merge` table has several children",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/113184 Related: https://github.com/ClickHouse/ClickHouse/pull/113489 Related: https://github.com/ClickHouse/ClickHouse/pull/111194 Related: https://github.com/ClickHouse/ClickHouse/pull/113043 Related: https://github.com/ClickHouse/ClickHouse/pull/108924 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the `LOGICAL_ERROR` `Reading from materialized CTE '...' before its materialization completed - DelayedPortsProcessor gate is missing in the query plan` raised when a `Merge` table with more than one child reads a materialized CTE that the outer query references. ### Description `ReadFromMerge::createChildrenPlans` optimizes every child plan on its own, and `resolveMaterializingCTEs` claims a materialized CTE globally through `MaterializedCTE::is_materialization_planned`. The first child plan to be optimized therefore moved the CTE's plan into *its own* tree, and every other `DelayedMaterializingCTEsStep` for that CTE - in the sibling children and in the outer plan - degenerated into a gate-less `MaterializingCTEsStep`. The CTE's writer then sat in one child's pipeline while readers sat in another, with no `DelayedPortsProcessor` between them, so a sibling's in-place `IN`-set build read the storage while it was still empty and `ReadFromMemoryStorageStep` raised the exception (in debug and sanitizer builds it aborts the server). Reproducer on master, deterministic (20/20), also with `max_threads = 1`, and a plain `EXPLAIN` is enough because the set is built during plan optimization: ```sql SET enable_analyzer = 1, enable_materialized_cte = 1; CREATE TABLE t (x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t SELECT number FROM numbers(10); CREATE TABLE tdist AS t ENGINE = Distributed(test_shard_localhost, currentDatabase(), t); CREATE VIEW tconst AS SELECT toUInt64(1) AS x; WITH t AS MATERIALIZED (SELECT number AS c FROM numbers(2)) SELECT count() FROM merge(currentDatabase(), '^(tconst|tdist)$') WHERE (x IN (t)) AND (x NOT IN (t)); ``` Children are visited in table-name order, so `tconst` is planned first and claims the CTE, and the `Distributed` child that follows builds the set in place while the CTE is unbuilt. Renaming so the `Distributed` child sorts first makes the same query pass, which is what pins the mechanism. **Fix.** A child plan no longer claims a CTE that the outer query references as well. `removeDelayedMaterializingCTEsStepFor` strips those steps from the child plan before it is optimized, leaving the outer plan - whose `MaterializingCTEsStep` sits above the whole merge - as the single owner that gates every child. This is the same reasoning `DelayedCreatingSetsStep::makePlansForSets` already applies to pre-built `IN`-subquery plans. The set of CTEs to strip comes from walking the outer `query_info.query_tree`, so a CTE defined *inside* one child (a `View` with its own `WITH ... AS MATERIALIZED`) is left owned by that child - it is the only reader, and stripping it unconditionally would leave it with no materialization at all. **Validation.** New `04811_materialized_cte_merge_child_gate` covers the failing child order, the explicit-subquery form, a satisfiable predicate that pins the data rather than only the absence of the exception, the `EXPLAIN` route, the reverse child order that always worked, and the view-owned-CTE case that must keep materializing inside the child. Every failing arm reproduces 5/5 on a master binary and passes 5/5 after the change. The `materialized_cte` suite is green. This is one shape of a recurring family - the same assertion is also reported in #113184 and addressed for other shapes by #113489, #111194 and #113043 - so the underlying `is_materialization_planned` claim being global while the gate is per-plan is worth revisiting separately. It surfaces constantly in the AST fuzzer; found via https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=9d0b1a25ba7aa4579c95a65baca002d1dd7a1e47&name_0=MasterCI&name_1=AST%20fuzzer%20%28amd_debug%29",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113558",
        "createdAt": "2026-08-05T19:32:08Z",
        "updatedAt": "2026-08-13T12:24:43Z",
        "timestamp": "2026-08-13T12:24:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "novikd"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113573",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: add `NOSIGN` to public `s3` documentation queries",
        "text": "Update anonymous public S3 documentation examples to pass `NOSIGN` explicitly following the ClickHouse 26.7 server credential behavior change. This prevents public reads from attempting to use server-managed credentials. The audit also found and repairs two stale public paths: the LAION guide now uses the surviving 10-million-row shard, and the S3 brace-expansion example references the four files that currently exist. Authenticated, requester-pays, write, placeholder, and Foursquare examples are intentionally outside this PR. Related: https://linear.app/clickhouse/issue/DOC-945/update-public-s3-docs-examples-to-use-nosign ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Updated public S3 documentation examples to use `NOSIGN` and repaired stale public dataset paths.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113573",
        "createdAt": "2026-08-05T20:49:36Z",
        "updatedAt": "2026-08-13T16:23:29Z",
        "timestamp": "2026-08-13T16:23:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-documentation",
          "pr-autogenerated-docs"
        ],
        "author": "dhtclk",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113609",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a logical error comparing arrays whose element type is Nothing",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Comparing two arrays whose element types have no least supertype aborted the server with `Bad cast from type DB::ColumnNothing to DB::ColumnVector<char8_t>` when an aligned element position held a bare `Nothing`, for example `SELECT CAST([], 'Array(Nullable(Nothing))') > [[1]]`, or over real columns of `Array(Tuple(Nothing, UInt64))` and `Array(Tuple(Array(UInt8), Int64))`. The abort happened during constant folding as well as at execution, so analysis alone could kill the server. Root cause: `getReturnTypeImpl`'s array-no-supertype branch declared a plain `UInt8`, while the element comparator it later builds declares `Nothing` for a value-less position. Execution then `assert_cast`-ed that `ColumnNothing` to `ColumnUInt8`. A bare `Nothing` position is genuinely undecidable: it holds no values, a `Tuple` member has no null map covering it, and there is no length to tie-break on. The fix makes the declared type honest. A new per-side `containsUndecidableNothing` predicate rejects such a pair during analysis with `ILLEGAL_TYPE_OF_ARGUMENT`, and `compareGatheredElements` keeps a fail-closed guard for defence in depth. Each side is classified independently, because one side's null map must never decide a position on the other side. The predicate deliberately mirrors `FunctionsNullSafeCmp::containsNothing` and does not descend into `Array`/`Map`; deeper positions stay covered because the caller recurses once per array level. Two `Nothing` shapes remain decidable and keep answering, now correctly rather than crashing: `Nullable(Nothing)` (decided by its own null map) and `Array(Nothing)` (shares a supertype with `Array(T)`). Those answers were checked against the same comparison on a pair that does have a supertype, and match exactly in both operand orders. Introduced by #110245, which is on `master` only, so no released version is affected and no backport is needed. Tracked in #113640 (CI signature STID 1499-2747), which also collects the `Equality` sibling at `FunctionsComparison.h:1485`. <details> <summary>Validation</summary> Both directions, one command per binary (pre-fix and fixed builds of the same branch): | query | pre-fix | fixed | |---|---|---| | `SELECT CAST([],'Array(Nullable(Nothing))') = [[1]]` | abort, `Bad cast ... ColumnNothing ...` | `0` | | `SELECT CAST([],'Array(Nullable(Nothing))') > [[1]]` | abort, same message | `0` | | `SELECT [NULL] = [[1]]` | `Code: 44` | `0` | The values the fixed build returns were compared against the supertype path (`[1]` in place of `[[1]]`) and match exactly: `0 1 0 0 1 1` for the six operators, `0 1 1 1 0 0` for an empty aligned prefix, `0 1` for `isNotDistinctFrom`/`isDistinctFrom`, and `0 0` in the reversed operand order. Coverage: all eight operators the introducing PR added; `Nothing` direct, under `Nullable`, nested in `Tuple`, and under 1-, 2- and 3-deep `Array` wrappers; `Map(k, Nothing)` and `Array(Nothing)` as must-still-answer controls; empty and non-empty aligned prefixes. An empty range is rejected too, so validity never depends on the data. No existing assertion was weakened: all 47 pre-existing reference lines of `04549_array_comparison_bigger_types_nullable` are unchanged. 200/200 runs pass across `--test-runs 50 --order random` and `--test-runs 50 --no-random-settings`, and the feature's own eight-test suite has no failures. Nine mutation arms were run against the new tests; eight are caught, and the ninth only makes the secondary fail-closed guard unreachable, which no input can pin because the analysis-time check rejects every such pair first. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1319` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113609",
        "createdAt": "2026-08-06T03:27:50Z",
        "updatedAt": "2026-08-13T11:48:45Z",
        "timestamp": "2026-08-13T11:48:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "diegomestre2"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113636",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not run distributed-plan rewrites on the child plans of a Merge table",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix a logical error (`Block structure mismatch`) when reading from a `Merge` table over children with internal set operations (e.g. a `Buffer` table) with `make_distributed_plan` enabled. A child plan of `ReadFromMerge` that acquired no exchange steps is united into the parent pipeline in the same process and is never serialized as a distributed fragment. With `make_distributed_plan = 1`, the plan optimization still ran `materializeConstantsForSetOperationBranches` on every plan: it materialized the constants of the `Union` inside a `Buffer` child (destination table + buffers), while a sibling child kept them const. That broke the equal-headers invariant across the children of `ReadFromMerge` and failed with a logical error at pipeline assembly (`Block structure mismatch in Pipe stream: different columns`), or with `Cannot convert column ... because it is non constant in source stream but must be constant in result` for a single `Buffer` child. The fix gates `materializeConstantsForSetOperationBranches` on the plan actually containing logical exchange steps. The rewrite exists because plan serialization stores only column names and types, so constness is re-derived per step after a fragment is deserialized — and a plan with no exchanges is never split into serialized fragments (`convertToDistributed` keeps it as a single stage executed in the same process), so it must keep its constants. Distributed planning of the child plans themselves is untouched: a `Merge` table over `Distributed` tables intentionally converts each child plan to a distributed plan on its own (see https://github.com/ClickHouse/ClickHouse/pull/108401 and `04367_distributed_plan_merge_scatter_multishard`). Minimal reproducer: ```sql CREATE TABLE t1 (k UInt64) ENGINE = EmbeddedRocksDB PRIMARY KEY (k); CREATE TABLE t2 ENGINE = Buffer(currentDatabase(), 't1', 1, 1, 10, 10000, 1000000, 10000000, 100000000); SET make_distributed_plan = 1, distributed_plan_execute_locally = 1; SELECT DISTINCT 1 FROM merge('^t') QUALIFY materialize(1); -- LOGICAL_ERROR before this fix ``` Found by the AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/100173: [AST fuzzer (amd_debug, targeted) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=100173&sha=22dee0374690f4c09727a81023ef38fe4e515785&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%29). Additionally, the review surfaced an adjacent pre-existing bug, fixed here as well: reading the destination table of a `Buffer` table can return a full column where the plan over the buffers keeps it constant (constants come back materialized from a `Distributed` destination), and a full column cannot be converted back to a constant, so such reads failed with `ILLEGAL_COLUMN` (`Cannot convert column ... because it is non constant in source stream but must be constant in result`) — with or without `make_distributed_plan`, and with or without a `Merge` table on top: ```sql CREATE TABLE dest (k UInt64) ENGINE = MergeTree ORDER BY k; CREATE TABLE dist (k UInt64) ENGINE = Distributed(test_shard_localhost, currentDatabase(), 'dest'); CREATE TABLE buf (k UInt64) ENGINE = Buffer(currentDatabase(), 'dist', 1, 1, 10, 10000, 1000000, 10000000, 100000000); SELECT DISTINCT 42 FROM buf QUALIFY materialize(42); -- ILLEGAL_COLUMN before this fix ``` `StorageBuffer::read` now materializes such constants in the buffers branch and unites the branches on the materialized header. This also covers the mixed local/distributed `ReadFromMerge` child shape raised in the review (a `Buffer` child over a `Distributed` destination next to a non-distributed sibling), which failed at child plan creation inside `StorageBuffer::read` — before the plan-level guard is ever reached. Related: https://github.com/ClickHouse/ClickHouse/pull/100173",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113636",
        "timestamp": "2026-08-12T22:10:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113651",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Wait for container boot before installing packages in `Install packages`",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Install packages (amd_release)` intermittently fails its `Install server rpm` substep with one line and nothing else: ``` + yum localinstall '--disablerepo=*' --allowerasing -y /packages/clickhouse-server-...rpm ... [Errno 2] No such file or directory: '/var/cache/dnf/metadata_lock.pid' ``` Example: [PR #109299](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109299&sha=3c7aad076ea307634a39a3089615b6a73e208621&name_0=PR&name_1=Install%20packages%20%28amd_release%29). **Root cause.** The image boots systemd, and the Dockerfile deliberately keeps `systemd-tmpfiles-setup.service`, which runs `systemd-tmpfiles --create --remove --boot`. centos:8 ships `/usr/lib/tmpfiles.d/dnf.conf`, whose entire content is remove directives for the dnf lock files, `metadata_lock.pid` among them. `test_install` starts the container `--detach` and `docker exec`s `install.sh` immediately, so `yum` and that lock wipe run concurrently. dnf takes the metadata lock in `Base.fill_sack` before it even opens the rpm files and does not guard against it vanishing, so the substep aborts before any package work: install, start and the smoke test never run. **Change.** Wait for the boot transaction before the first `docker exec`, in the shared `test_install` helper, so every substep of both images is covered. Only the rpm image is affected in practice: ubuntu:22.04 ships no `dnf.conf`, and the dpkg and apt locks survive the same run. Waiting for D-Bus first is the load-bearing half: in the first milliseconds `systemctl` cannot reach systemd at all, so a gate built on it alone silently does nothing in the window it guards. A bare `systemctl start` returned `Failed to connect to bus` in 5 of 8 tries; with the bus wait in front, `rc=0` in 8 of 8. **Validation.** Built the real image locally and amplified the race with 24 concurrent containers, no fault injection: **6 of 144** runs reproduced the exact CI line without the wait, **0 of 144** with it, gate `rc=0` in 144/144. Cost is ~0.17 s per container, so ~3.1 s over the 18 containers a job starts, against a job whose 30-day median is ~240 s. No related open issue found.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113651",
        "createdAt": "2026-08-06T11:22:40Z",
        "updatedAt": "2026-08-13T15:12:31Z",
        "timestamp": "2026-08-13T15:12:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113681",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Replace the per-bucket hash map in `timeSeries*ToGrid` with a sorted-append sample array",
        "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replaced the per-bucket hash map inside the `timeSeries*ToGrid` aggregate functions with a flat sorted array of samples: sample ingestion becomes an O(1) append for in-order inputs (the overwhelmingly common case) and the per-bucket copy-and-sort at finalization is gone. ### Description The `timeSeries*ToGrid` functions kept each bucket's samples in an `absl::flat_hash_map<timestamp, value>`: every `add()` paid a hash-map emplace, and the order-dependent functions (`rate`, `increase`, `delta`, `changes`, `resets`) copied and sorted every bucket at finalization. Yet the input is almost perfectly ordered — samples come from MergeTree tables sorted by `(id, timestamp)`; an instrumented probe on a 32-thread read of a 62.5-billion-sample table counted **1 out-of-order add in 1,474,559,998**. The bucket is now a flat, memory-tracked vector of `(timestamp, value)` pairs: O(1) append while timestamps ascend, in-place max on an equal timestamp, and a rare out-of-order add just clears a `sorted` flag — normalization (sort + max-dedup) runs lazily, only for buckets that actually saw disorder. `merge()` is a linear merge of sorted runs with an append fast path for disjoint time ranges. `forEachSample` now guarantees ascending order, so the copy-and-sort buffers are deleted from the rate/delta/changes aggregators. The wire format and `FORMAT_VERSION`s are unchanged; `deserialize()` assumes no order of incoming pairs (old peers send hash-map iteration order), so mixed-version clusters interoperate — verified in both directions. Duplicate timestamps keep the larger value with the old `std::max` argument order. Measured on the 62.5-billion-sample table: a 30-day `sum by(...)(rate(...))` over 25,600 series drops ~6% of total query CPU (1106 s -> 1037 s); wall time and peak memory move little (the scan dominates the critical path, and raw sample storage dominates the state either way). The structural point is what this enables: the sorted buffer is the prerequisite for O(1) per-segment summaries in the rate family (follow-up), which is where the ~44 GiB state peaks of such queries actually go away. All query results are fingerprint-identical. Tests: a new stateless test (shuffled and duplicate-timestamp inputs in both orders, NaN at duplicated timestamps, `-Merge` of unsorted in-memory states, interleaved parts, two-level merges, serialized-state merges through `remote('127.0.0.{1,2}', ...)`, an `AggregatingMergeTree` roundtrip, a fixed state literal in old-peer wire order); all 40 existing timeseries/PromQL stateless tests pass byte-identically; the perf test gains an ingestion-heavy scenario (50M rows, 10k series). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113681",
        "createdAt": "2026-08-06T13:46:38Z",
        "updatedAt": "2026-08-13T17:58:34Z",
        "timestamp": "2026-08-13T17:58:34Z",
        "metrics": {
          "reactions": 1,
          "comments": 9
        },
        "labels": [
          "pr-performance",
          "comp-promql"
        ],
        "author": "nikitamikhaylov",
        "state": "open",
        "assignees": [
          "vitlibar"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113691",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `theilsU` window state returning noise when the frame's first argument is constant",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/80373 Related: https://github.com/ClickHouse/ClickHouse/pull/93384 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `theilsU` over a window frame returning an arbitrary value instead of 0 when the first argument is constant within the frame. ### Description `TheilsUWindowData::getResult` (the window-optimized state introduced in https://github.com/ClickHouse/ClickHouse/pull/93384) computes the entropy `H(A)` from cached incremental `Σ n·log n` sums. When the first argument is constant within the frame, the true `H(A)` is zero, and the computed value is pure rounding noise from the incremental updates. The code compared it against exact zero, so a tiny positive noise value passed the check, and `1 - H(A|B) / H(A)` then divided noise by noise: in debug builds this tripped the sanity check as the exception `Logical error: 'res < 1.0 + 1e-4'`, and in release builds the function could return an arbitrary value in $[0, 1]$ instead of 0. The exact (non-window) code path recomputes the entropies from the count maps, where a constant column gives `log(1) = 0` exactly, so it is not affected. The fix compares `H(A)` against an error bound proportional to `N · ε · log N` instead of exact zero, and widens the sanity-check tolerance by the same relative amount so that near-threshold frames do not trip it either. Found by the AST fuzzer on an unrelated PR (it hit https://github.com/ClickHouse/ClickHouse/pull/80373 and https://github.com/ClickHouse/ClickHouse/pull/107667): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=80373&sha=6c271049214aa5a94bd9a5f12fb27a9ffa75648f&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113691",
        "createdAt": "2026-08-06T15:21:02Z",
        "updatedAt": "2026-08-13T17:48:39Z",
        "timestamp": "2026-08-13T17:48:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "nihalzp"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113707",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: add Erathos connector to data ingestion docs",
        "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds Erathos, an ELT platform, to the data ingestion integrations list.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113707",
        "createdAt": "2026-08-06T18:04:58Z",
        "updatedAt": "2026-08-13T13:10:44Z",
        "timestamp": "2026-08-13T13:10:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-documentation",
          "can be tested"
        ],
        "author": "gelsonbagetti",
        "state": "open",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113742",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `CANNOT_CONVERT_TYPE` and an exception in `GroupingAggregatedTransform` when reading a `Merge` table with one child being a `Distributed` table under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_sampling'` or `'custom_key_range'`). Closes [#113741](https://github.com/ClickHouse/ClickHouse/issues/113741). ### Documentation entry for user-facing changes The custom-key parallel replicas branch of the planner replaces the plan of a table expression with a remote read at the fixed stage `WithMergeableStateAfterAggregationAndLimit`, ignoring the stage the plan was requested up to. A `Merge` table over a `Distributed` child plans all of its children up to `WithMergeableState` through an interpreter, so a `MergeTree` child's plan produced finalized (post-aggregation, post-`LIMIT`) data where the parent `ReadFromMerge` expected partial aggregation states: `CANNOT_CONVERT_TYPE` for `count`, and the exception `Chunk should have AggregatedChunkInfo/ChunkInfoWithAllocatedBytes in GroupingAggregatedTransform` when the finalized type structurally coincides with the state type. Only the analyzer path is affected. The fix allows the custom-key read only when the requested stage is `Complete` or `WithMergeableStateAfterAggregationAndLimit` itself; a child planned to a partial stage now runs as a plain local read, as it does when parallel replicas are off. Found by the targeted AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/110972, where it is unrelated: the failure reproduces on master without `parallel_replicas_allow_merge_tables` (verified on a binary with that PR's changes swapped out to the merge base). Fuzzer report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110972&sha=ac0a584ce316ace31f4dbc16b38e1262a2344751&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29 (STID `3970-479a`). Closes: https://github.com/ClickHouse/ClickHouse/issues/113741 Related: https://github.com/ClickHouse/ClickHouse/pull/110972 <!-- ch-version-info:start --> ### Version info - Backported to: `26.7.4.29` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113742",
        "createdAt": "2026-08-06T23:22:57Z",
        "updatedAt": "2026-08-13T16:22:58Z",
        "timestamp": "2026-08-13T16:22:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113754",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject STREAM with parallel replicas at plan-build time",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/110144 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `LOGICAL_ERROR` when a `SELECT` used the `STREAM` modifier on a named table together with parallel replicas in the read-tasks mode. With a materialized CTE used as an `IN` set in `PREWHERE`, the server raised `Reading from materialized CTE ... DelayedPortsProcessor gate is missing in the query plan` (a server abort in debug and sanitizer builds). A table carrying `STREAM` is now rejected when parallel replicas are requested: at `enable_parallel_replicas = 2` the query fails with `SUPPORT_IS_DISABLED`, at `1` it runs without them, exactly as already happens for `FINAL`. ### Description `MergeTreeDataSelectExecutor::read` already refuses `STREAM` with parallel replicas, but its `enable_parallel_reading` argument is `canUseParallelReplicasOnFollower` (`StorageMergeTree.cpp:385`, `StorageReplicatedMergeTree.cpp:6305`), so the refusal only fires on a follower: the initiator builds the plan it should have refused and runs it. `STREAM` also suppresses the eager set build: `ReadFromMergeTree::applyFilters` early-returns for a streaming read (`ReadFromMergeTree.cpp:2593`), and that inplace build is what materializes a CTE, so a materialized CTE used as an `IN` set stays unmaterialized and its ungated `StorageMemory` read raises the error at `ReadFromMemoryStorageStep.cpp:87`. The fix extends the existing `FINAL` check in `Planner::buildPlanForQueryNode` (`Planner.cpp:2306` on master) to `STREAM`. It runs before `buildJoinTreeQueryPlan`, so no storage read is reached, and it clears `allow_experimental_parallel_reading_from_replicas` rather than skipping one branch, covering `parallel_replicas_plan_based` too. For `STREAM` it applies on an initiator only. The enclosing `canUseTaskBasedParallelReplicas` is role-blind, and a follower that cleared the setting here would lose its own read-side refusal, so an old initiator plus a new follower silently did a full local `STREAM` read instead of failing; measured on a two-server cluster at `enable_parallel_replicas = 1`, it now fails with `ILLEGAL_STREAM` as it does on two old servers. `FINAL` keeps its unconditional handling, having no read-side refusal to preserve. Scope, measured: this covers the read-tasks mode on a named table, the reported carrier. The custom-key and sampling-key modes never reach the guard's enclosing `canUseTaskBasedParallelReplicas` branch, and `STREAM` with `parallel_replicas_mode = 'custom_key_sampling'` hangs identically before and after this change; that hang stays open. A `TableFunctionNode` carrying `STREAM` is not covered either. An earlier revision covered both and was reduced on review. #110972 rewrites this same loop; a merge has to keep both sides. Found by the AST fuzzer on #110144",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113754",
        "createdAt": "2026-08-07T02:36:15Z",
        "updatedAt": "2026-08-13T13:33:49Z",
        "timestamp": "2026-08-13T13:33:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "Michicosun"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113796",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Bump `silk` to the latest `clickhouse-public`",
        "text": "Automated bump of `contrib/silk` from `52d2039c428` to `78aab9d603d`, the tip of [`clickhouse-public`](https://github.com/ClickHouse/silk/tree/clickhouse-public) in [ClickHouse/silk](https://github.com/ClickHouse/silk). Changes in `silk` ([compare](https://github.com/ClickHouse/silk/compare/52d2039c42830601013d8654ddf750be971481fa...78aab9d603d43247fa8f05f86c34ecdfe2ab4095)): ``` 78aab9d Remove all submodules from contrib c095ed0 Revert \"Migrate to bundled fcontext code\" fe1144b Revert \"Migrate to bundled cxxopts library\" f561e7c Rebuild `clickhouse-public` automatically after `main` passes ea5b51e pageSize should be retrieved via ::sysconf(_SC_PAGESIZE), not compiled into the binary 47057a9 Force migrate fibers under sanitizers 6c1e8f2 Add mmap accounting hooks 3e5eb01 Fix fiber-list RUNNING detection on aarch64 b7d0594 Install gdb in runner images (#113) 8f517ec Add waitCycles to waitForMultiple and waitWithTimeout 21b21aa Isolate the blocking-queue futexes on their own cache lines c24c9c5 Move values through BoundedQueue slots 3dd82b7 Add FiberSequencer::reset 10a9c99 Patch vendored Poco Format.cpp to resolve std::format ambiguity c879caa Remove nodiscard requirement from FiberSequencer::advance 2d33bc9 Cut the sequencer litmus iteration count under TSan e9f2261 Add a benchmark for injection from a reserved core dcdf454 Test the scheduler active CPU set 50fcb5b Restrict the scheduler to an active CPU set d041773 Size per-CPU state by the configured processor count 55d697e Add claude.md ``` This pull request self-updates until it is merged. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1277` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113796",
        "timestamp": "2026-08-12T20:14:23Z",
        "metrics": {
          "reactions": 2,
          "comments": 7
        },
        "labels": [
          "pr-not-for-changelog",
          "submodule changed",
          "pr-synced-to-cloud"
        ],
        "author": "clickhouse-gh[bot]",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113833",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add JOIN observability columns to system.query_log",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111352 Related: https://github.com/ClickHouse/ClickHouse/issues/111748 ## Motivation `system.query_log` says almost nothing about what the JOINs in a query actually did. Answering questions like \"which queries ran a `CROSS` join we didn't expect\", \"which joins fell back to `grace_hash` and spilled to disk\", or \"which algorithm was really chosen when `join_algorithm = 'auto'`\" currently requires re-running `EXPLAIN PIPELINE` (impossible post-mortem) or scraping `ProfileEvents` query by query. ## Changes Four columns are added to `system.query_log`, with the matching fields in `QueryLogElement`: - `used_number_of_joins` (`UInt64`) — the number of physical joins executed by the query. It is collected from the query pipeline, so it reflects the joins that really ran after all optimizations, not the number of `JOIN` clauses in the query text. - `used_join_algorithms` (`Array(LowCardinality(String))`) — the algorithms that were actually used, e.g. `hash`, `parallel_hash`, `grace_hash`, `direct`, `full_sorting_merge`, `partial_merge`. The `join_algorithm` setting only lists the allowed algorithms; the choice among them happens at runtime, and an algorithm can even be replaced mid-execution (a `hash` join switching to `grace_hash` under memory pressure). - `used_join_kinds` (`Array(LowCardinality(String))`) — Kind of the joins present in the query. - `used_join_strictness` (`Array(LowCardinality(String))`) — Strictness of the joins present in the query. - `join_spilled_to_disk` (`UInt8`) — whether any of the joins spilled to disk. This PR currently adds the schema only. The fields are declared but nothing populates them yet, so the columns read as `0` and `[]`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added four columns to `system.query_log` describing the JOINs a query executed: `used_number_of_joins` (the number of physical joins in the executed pipeline), `used_join_algorithms` (the algorithms actually used at runtime, which can differ from the `join_algorithm` setting), `used_join_kinds` (`INNER`, `LEFT`, `CROSS`, `ASOF` and so on), and `join_spilled_to_disk` (whether any join wrote temporary data to disk). This makes it possible to find problematic JOIN patterns across a fleet without reproducing each query.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113833",
        "createdAt": "2026-08-07T14:11:36Z",
        "updatedAt": "2026-08-13T17:36:09Z",
        "timestamp": "2026-08-13T17:36:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "Manerone",
        "state": "open",
        "assignees": [
          "Fgrtue"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113868",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"Add aggregate function `gini`\"",
        "text": "Reverts ClickHouse/ClickHouse#112280 Closes: https://github.com/ClickHouse/ClickHouse/issues/113763 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Besides adding `gini`, #112280 added nine `getArgumentsThatCanBeOnlyNull` overrides across eight combinator files and made `-If` union the nested set, so those combinators forward that set from the function they wrap. The `Null` combinator uses that set to decide whether an argument that can only be `NULL` folds the aggregate into `AggregateFunctionNothing`. Forwarding it removes the fold one level up, so expressions chaining one of them on top of `-If` changed type and value. Measured on debug builds of this branch and its parent: | expression | before #112280 | after | |---|---|---| | `countIfOrNull(number, NULL)` | `UInt64` `0` | `Nullable(UInt64)` `NULL` | | `sumIfResample(0, 2, 1)(number, NULL, number % 2)` | `Nullable(Nothing)` `NULL` | `Array(UInt64)` `[0, 0]` | | `sumIfState(number, NULL)` | folded `Nullable(Nothing)` | `AggregateFunction(sumIf, UInt64, Nullable(Nothing))` | Two of the new overrides sit outside the `-If` family, reaching expressions with no `-If` at all. A partial revert is not available: `gini` declares argument 0 in `getArgumentsThatCanBeOnlyNull` itself, and one forwarding path carries both that declaration and the `-If` filter index, so keeping `gini` without the propagation contradicts its own test. `gini` is unreleased (the merge is not an ancestor of 26.7, 26.6, 26.5, 26.3 or 25.8, and never reached `CHANGELOG.md`), so this withdraws no released behaviour. A behaviour-neutral re-land following the `sum` family, as @Manerone specified, comes separately. The docs page is removed here too. It did not come from the reverted merge: `2d4afadac21e7b4` moved it into the live Mintlify tree, so it arrived via the master merge. Its navigation entry, legacy redirect and slug-map row go with it, or those would point at a missing target. Regenerating is not an option: the aggregate family in `autogenerate_docs.py` lists that page directory and carries `skip_if_empty`, so an unregistered function's page is never revisited. Verified by byte identity rather than a new test: of the 14 source and test paths the merge touched, 13 match their pre-merge blobs, and `registerAggregateFunctions.cpp` differs only by unrelated `MergedJSONPatch` lines master added in `e3698631b023165`. cc @Manerone <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1317` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113868",
        "createdAt": "2026-08-07T18:25:03Z",
        "updatedAt": "2026-08-13T10:57:09Z",
        "timestamp": "2026-08-13T10:57:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-not-for-changelog",
          "manual approve",
          "can be tested",
          "pr-synced-to-cloud",
          "pr-autogenerated-docs"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "Manerone"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113890",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support GCS vended credentials in Unity Catalog",
        "text": "Unity Catalog on GCS vends an OAuth token in `gcp_oauth_token`, but `UnityCatalog` only parsed AWS temporary credentials for S3-compatible locations. As a result, table metadata could be listed while reads failed because no storage credentials were created. Parse the GCP OAuth token as `GCSCredentials` for both initial credential resolution and refresh. Add focused coverage for GCS parsing and refresh while preserving the existing S3 behavior, and update the user-facing support descriptions. Closes: https://github.com/ClickHouse/ClickHouse/issues/110149 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support Unity Catalog vended credentials for tables stored in GCS.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113890",
        "timestamp": "2026-08-12T22:46:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "KyriosGN0",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113894",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #110019 to 26.6: Support some settings alter for OneLake catalog",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/110019 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/112549 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31221192737/job/93005936464)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113894",
        "createdAt": "2026-08-07T22:03:13Z",
        "updatedAt": "2026-08-13T15:56:43Z",
        "timestamp": "2026-08-13T15:56:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-backport"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "alesapin",
          "scanhex12"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113895",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Strip the cosmetic parenthesized flag before comparing stored definitions",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/92340 Related: https://github.com/ClickHouse/ClickHouse/pull/110833 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed comparison of stored table definitions that were written with redundant parentheses (`PARTITION BY (a)`, `PRIMARY KEY (key)`). Since #92340 the formatter preserves those parentheses, so `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` rejected two otherwise identical tables with `Tables have different partition key`, and a `KeeperMap` table created by 26.5 or 26.6 could not be opened by a server of another version. ### Description #92340 started preserving the parentheses a user writes around a definition expression. They are cosmetic, but stored table metadata is compared as text against a form that may have been written by another server version, so two identical definitions began to differ as strings. Two user-visible consequences: - **`ATTACH PARTITION FROM`.** A table declared `PARTITION BY (a)` no longer matches one declared `PARTITION BY a`. This requires neither an upgrade nor a mixed-version cluster: the comparison is between two in-memory ASTs of which only one carries the flag, so both tables created by the same binary already fail. Measured on released binaries, 26.3 and 26.4 accept the pair and 26.5 and later reject it. ```sql CREATE TABLE src (a UInt32, b UInt32) ENGINE=MergeTree PARTITION BY (a) ORDER BY b; CREATE TABLE dst (a UInt32, b UInt32) ENGINE=MergeTree PARTITION BY a ORDER BY b; INSERT INTO src VALUES (1, 1); ALTER TABLE dst ATTACH PARTITION 1 FROM src; -- 26.4: ok -- 26.6: Code: 36. DB::Exception: Tables have different partition key. (BAD_ARGUMENTS) ``` - **`KeeperMap`.** The primary key is serialized into Keeper and compared there against the text written by whichever version created the table. 26.5 and 26.6 write `primary key: (key)`, every other version writes `primary key: key`, so a server of the other version refuses to open the table: ``` Path ... is already used but the stored primary key definition doesn't match. Stored metadata: ... primary key: (key) local metadata: ... primary key: key ``` On a multi-replica setup the replicas that cannot apply the definition never finish startup. ### Implementation `ReplicatedMergeTreeTableMetadata` already stripped the flag, but only on the top level of an expression list, and the helper was private to that file, so the other two comparison sites never got it. This promotes it to `Parsers/stripArtificialParens.h`, makes it walk the whole tree, and applies it in the two places that were missed. The walk also reaches members that are not in `children` and would otherwise be skipped: the `GROUP BY` keys, the `GROUP BY` assignments and the recompression codec of a TTL element, and the `parameters` and `lambda` of a projection's `APPLY` transformer. `StorageKeeperMap` additionally accepts a stored primary key that still carries the parentheses, so tables already created by 26.5 or 26.6 keep working after this change. If that stored text cannot be parsed the comparison stays strict and the mismatch is reported, rather than being silently accepted. Only the comparison changes. There are no parser or formatter changes: what the user wrote is still what is stored and what `SHOW CREATE` and `system.tables` report, and genuinely different definitions still differ (covered by negative cases in both tests). This also fixes two transposed format arguments in the `KeeperMap` mismatch message, which rendered the path and the field name in the wrong order (`Path columns is already used but the stored /keeper_map_tables/... definition doesn't match`). ### Relationship to #110833 This is the compatibility part of #110833, extracted so it can be reviewed and backported on its own, as requested in https://github.com/ClickHouse/ClickHouse/pull/110833#issuecomment-5058455151. The `getTreeHash` work from that PR is a separate, larger change and is not included here. Credit for the original diagnosis and the wider fix goes to @groeneai. ### Tests - `04821_parenthesized_key_attach_partition_from` covers `ATTACH PARTITION FROM` across the two spellings, asserts that `system.tables` still reports the parentheses the user wrote, and that a genuinely different partition key is still rejected. - `04822_keeper_map_parenthesized_primary_key` asserts that the primary key stored in Keeper does not depend on the spelling, that a second table on the same path with the other spelling opens, and that a different primary key is still rejected.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113895",
        "createdAt": "2026-08-07T22:24:44Z",
        "updatedAt": "2026-08-13T15:22:29Z",
        "timestamp": "2026-08-13T15:22:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "v26.6-must-backport"
        ],
        "author": "fm4v",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113899",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Avoid scans for constant sort keys",
        "text": "## What `isAlreadySorted` now returns early, right after the sort descriptors are resolved, when every sort key is a `ColumnConst`. Writing a `MergeTree` part whose sorting keys are all constant no longer walks the block doing adjacent-row comparisons to confirm an ordering that is constant by construction. Collation validation still runs before the early return, and any block that mixes constant and non-constant keys keeps the existing comparator path — only the all-constant case takes the new path. ## Why it helps When the sorting-key columns are constant across a block — a common shape when a leading `ORDER BY` column is fixed per part (time-ordered or batched ingestion, per-source or per-partition writes) — the sortedness check was doing a full comparison pass to reach a foregone conclusion. Returning as soon as the keys are known-constant removes that pass. Measured on `MergeTree inserts with constant and mixed sorting keys`, 64K/256K/1M rows (paired medians, co-measured on both trees): | Metric (1M rows) | Before | After | Δ | | --- | ---: | ---: | ---: | | Constant-key sortedness check, 1 key | 1106 µs | 2 µs | **−99.8%** | | Constant-key sortedness check, 4 keys | 4449 µs | 4 µs | **−99.9%** | | End-to-end insert latency, 4 keys | 24548 µs | 20201 µs | **−17.7%** | | Insert CPU, 4 keys | 20732 µs | 16301 µs | **−21.4%** | Smaller block sizes land in the same range (e.g. 1-key 64K: 66 µs → 2 µs). The sortedness check collapses to a near-constant cost, and that saving carries into a full insert as the end-to-end latency and CPU gains. Blocks that aren't all-constant take the unchanged comparator path. ## Testing Stateless coverage exercises single and multiple constant keys plus a non-constant suffix (the mixed case that must keep comparing). The declared correctness check passed. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Skip the redundant sortedness scan when writing MergeTree parts whose sorting keys are all constant. --- Contributed by [Perfloop](https://app.perfloop.ai): the numbers above were co-measured on both trees and independently re-verified before submission — the full public record is at [case_8ebwekrder](https://app.perfloop.ai/t/oss/case_8ebwekrder). Replies from this account are human-approved, and a human operator is accountable for this contribution. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1346` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113899",
        "createdAt": "2026-08-07T23:19:06Z",
        "updatedAt": "2026-08-13T16:40:53Z",
        "timestamp": "2026-08-13T16:40:53Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-performance",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "perfloop-agent",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113902",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: `break` does not always return a partial result",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112483 --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Both `timeout_overflow_mode` descriptions promise unconditionally that `break` returns \"the partial result, as if the source data ran out\", so a user picking `break` to avoid errors can still get one. `QueryStatus::checkTimeLimit` returns false rather than throwing under `break`, meaning \"stop and return what you have\". That works where a truncated output is a smaller valid one, not where it would be a wrong one, so several sites stop instead of truncating. A function computing one scalar value (`arrayFold`, `geohashesInBox`, string replace) stops without producing it, surfacing `TIMEOUT_EXCEEDED` only when not interrupted mid-pipeline, since a pipeline absorbs the throw into a clean cancellation. An incomplete `Memory`-table mutation (`StorageMemory::mutate`) leaves the table unchanged and reports it. A quorum write cannot report partial success at all, so it ignores the false return and does not stop at `max_execution_time`: it waits out `insert_quorum_timeout`, then reports `UNKNOWN_STATUS_OF_INSERT`. The mutation and dictionary-load waits are the same, hence \"some operations\". The `INSERT` code differs by executor: `QUERY_WAS_CANCELLED` when the caller pushes data block by block, as over the native protocol; \"may\", because an in-pipeline throw pre-empts that translation, and a pipeline completed with its own input source raises nothing. Retention is an engine property, so none is promised. That sentence also made `max_estimated_execution_time` a subject of `break`, but `ExecutionSpeedLimits.cpp:82` gates that check on `throw`, so the estimate is never evaluated. That gate predates this change, so only the text is corrected. The leaf setting has no estimate and reaches no `INSERT` or quorum write, so those clauses are omitted there. The sentence is shared boilerplate appearing 6 times in `Settings.cpp`; only these 2 are on the elapsed-time path, so the other 4 threshold modes are unchanged. Hedging is required in both directions: `FillingTransform` and `sleep` do return partial output. No behaviour changes. Raised by `clickhouse-gh[bot]` in https://github.com/ClickHouse/ClickHouse/pull/112483#discussion_r3686684259. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1300` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113902",
        "createdAt": "2026-08-07T23:52:27Z",
        "updatedAt": "2026-08-13T02:12:43Z",
        "timestamp": "2026-08-13T02:12:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-documentation",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113903",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add file IO and integrate new keeper storage",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Keeper can store data on disk now, in a custom LSM tree. It has similar performance to the previous storage. Use coordnation setting `use_new_storage = true` to enable, `storage_memory_only = false` to store files on disk, add `data_storage_path` or `data_storage_disk` to config to specify where to put the files. To write to S3, point `data_storage_disk` to a disk of type `s3_plain`. --- https://github.com/ClickHouse/ClickHouse/pull/107261 added new storage with memory-only mode. This PR * Adds ability to store data in files. The storage is not actually persistent, the directory is wiped on startup and re-created from snapshot, just like with previous rocksdb storage. So the files don't have things like headers and version numbers yet, we can add that later if we want faster startup. * Integrates the new storage into keeper. (This PR replaces https://github.com/ClickHouse/ClickHouse/pull/112378 )",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113903",
        "createdAt": "2026-08-08T00:05:18Z",
        "updatedAt": "2026-08-13T13:51:35Z",
        "timestamp": "2026-08-13T13:51:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-feature"
        ],
        "author": "al13n321",
        "state": "open",
        "assignees": [
          "antonio2368"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113909",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check the query cancellation while filling `system.parts` and its siblings",
        "text": "The tables based on `StorageSystemPartsBase` (`system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`) build the whole result eagerly in `initializePipeline`, so a cancelled or timed out query kept building rows over every storage and part until the very end. In a stress test with ThreadFuzzer this took minutes and tripped the hung check: a `SELECT` over `system.parts_columns` (with a filter matching every active part on the server) stayed in the process list for 252 seconds with `is_cancelled = 1` and `max_execution_time = 10`. Check the query status per storage and per part, following the pattern of `system.zookeeper` and `system.remote_data_paths`. The check also covers the storage-discovery prepass in `StoragesInfoStream` (the eager enumeration of all databases and tables), the analogous prepass in `StoragesDroppedInfoStream` (so `system.dropped_tables_parts` is interruptible as well), the skip loop in `StoragesInfoStreamBase::next`, and the per-storage part enumeration itself: the `MergeTreeData` helpers (`getDataPartsVectorForInternalUsage`, `getAllDataPartsVector`, `getProjectionPartsVectorForInternalUsage`, `getAllProjectionPartsVector`) take an optional `need_stop` callback that is checked periodically while the parts snapshot is being built. The two column-oriented tables additionally check the query status inside their column-enumeration loops and their metadata prepass, so a very wide table does not create a long uninterruptible stretch inside a single storage. The return value of `checkTimeLimit` is honored, so with `timeout_overflow_mode = 'break'` the eager build stops at the soft deadline and returns the rows collected so far. The table-lock acquisition in `StoragesInfoStreamBase::tryLockTable` is interruptible as well: instead of a single wait inside `RWLockImpl::getLock` for the whole `lock_acquire_timeout`, the lock is acquired in 100 ms slices with a query-status poll between the attempts (the total timeout and the `DEADLOCK_AVOIDED` semantics are preserved), so a killed or soft-timed-out query does not sit in the lock wait while a concurrent DDL query holds the drop lock. The test `04869_system_parts_lock_wait_cancellation` pins this with a share lock held by a long `SELECT` and a `DROP TABLE` in an `Ordinary` database queued behind it. The test uses the new `slowdown_system_parts_enumeration` failpoint, which only affects the specially named test tables (so concurrently running tests are unaffected). It sleeps 500 ms on every enumerated part, so building the full result for a 20-part table takes at least 10 seconds, and it sleeps 1 second per `COLUMNS_CANCELLATION_CHECK_PERIOD` (128) enumerated columns of a part, so building the full `system.parts_columns` / `system.projection_parts_columns` result over a single part with 1301 columns (and a projection over all of them) also takes at least 10 seconds. The test asserts that queries with a 1 second deadline in the `break` mode finish well under that, which is only possible by stopping at the per-part and per-column cancellation checkpoints. Timed assertions are needed because a plain row-count assertion cannot distinguish a build with the fix from one without: in the `break` mode the executor drops the eagerly built result after the deadline in both cases. The test also asserts partial row counts under a pre-expired deadline for all five tables, including `system.dropped_tables_parts` over a deterministic dropped-table fixture. The pre-expired-deadline checks also run under the failpoint, so the fewer-rows assertion is deterministic even on a machine fast enough to build the whole result in under a millisecond. For tables with the `_snap` name marker, the failpoint additionally slows down the parts-snapshot walks inside `MergeTreeData` (500 ms per enumerated part) and makes them poll the stop callback on every element, and for tables with the `_meta` name marker it slows down the column-metadata prepass of the column-oriented tables (1 second per 128 enumerated metadata columns), so the timed checks also prove that the snapshot materialization and the prepass are interruptible: all six of these checks fail against a binary with the `need_stop` polls and the prepass checkpoints disabled. Caught by `Stress test (arm_asan_ubsan, s3)`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113722&sha=ceedb772f6aa36e221d25ce48f9af71f032a7b08&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Related: https://github.com/ClickHouse/ClickHouse/pull/113722 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Queries over `system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`, and `system.dropped_tables_parts` now react to cancellation and `max_execution_time` while the result is being built.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113909",
        "createdAt": "2026-08-08T01:18:22Z",
        "updatedAt": "2026-08-13T16:41:37Z",
        "timestamp": "2026-08-13T16:41:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113912",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix UB in avg over Date/Time types at the Int64 boundary",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fix undefined behavior and wrong results in `avg` over `Date`/`DateTime`/`DateTime64`/`Time`/`Time64`: the average is now computed exactly in integer space, fixing both the `Int64`-boundary overflow (UB, wrong result on x86) and `Float64` precision loss above 2^53 (visible at nanosecond scale). `avgResultToValue` cast the rounded `Float64` average straight to the result's native integer type. The exact average always lies within the range of the inputs, but the `Float64` computation is inexact: for ticks near the bounds of `Int64` it can land on 2^63 exactly, and the cast is undefined behavior. UBSan caught it in a stress run as `9.22337e+18 is outside the range of representable values of type 'long'` at `AggregateFunctionAvg.h:39` (`avg` over `DateTime64`). On x86 the cast wraps to `INT64_MIN`, turning the average of values near the upper bound of `DateTime64` into `1677-09-21 00:12:43.145224192`; on AArch64 `fcvtzs` saturates silently, hiding the problem. A saturating cast alone is not enough (review finding): `Float64` cannot distinguish the last 1024 ticks of `Int64`, so clamping off the rounded `Float64` would still corrupt valid values just inside the boundary, and more generally the `Float64` division loses precision for any tick count above 2^53. Instead, for integer-backed Date/Time result types the exact accumulated numerator is now divided by the denominator in integer space, rounding half to even (matching the `nearbyint` semantics of the previous path), with saturation only when the accumulated sum itself has overflowed. The `Float64` path keeps a saturating cast for non-exact numerators. This also fixes a pre-existing precision artifact: `avg` of `2020-01-01 00:00:00.000000000` and `2020-01-01 00:00:00.000000002` at scale 9 now returns `.000000001` instead of `.000000000` (reference of `03799_avg_date_time_types` updated). CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109453&sha=2218c5bc60fe31decb1e21d9545bd9c9f9fc7a56&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%29 Related: https://github.com/ClickHouse/ClickHouse/pull/109453 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1198` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113912",
        "createdAt": "2026-08-08T02:35:56Z",
        "updatedAt": "2026-08-13T14:24:56Z",
        "timestamp": "2026-08-13T14:24:56Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113924",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check for cancellation in the h3 array-expanding functions",
        "text": "Check for cancellation in the h3 array-expanding functions ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `h3kRing`, `h3HexRing`, `h3Line` and `h3ToChildren` ignoring `max_execution_time` and `KILL QUERY` while expanding a block of rows. The query is resolved from the executing thread rather than captured, so a call stored in a key, index or TTL is bounded by the query that runs it. ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/112273 Each of these four functions expands every row of a block inside one `executeImpl` call, and the pipeline evaluates `max_execution_time` and `KILL QUERY` only between blocks, so a cancelled query kept a server thread busy until the whole block was done. Measured on a debug build against a one second limit, unfixed: `h3kRing` 8.1s, `h3HexRing` 4.7s, `h3Line` 12.2s, `h3ToChildren` 8.1s. The per-row caps these functions already have bound one row, not the sum over the block. The fix polls the query's `QueryStatus` inside each expanding loop, throttled on accumulated output items with every row counted as at least one, as in the merged fix for `geohashesInBox`. The element is resolved per call from `CurrentThread::tryGetQueryContext` rather than captured in the constructor, because these instances can live in table metadata and be run by unrelated later queries. Under `break` `checkTimeLimit` returns false rather than throwing; that becomes a throw here, a half-built array being a wrong value. `h3Line` needs the check in **both** its loops: `gridPathCellsSize` walks the grid even for a distance of zero, so at one item per row its sizing pass alone overshoots by 37.8x. `h3HexRing`'s sizing loop is `6 * k` arithmetic plus a cell validation and is left alone: reaching even 1.6x took eighty million rows. Every other array-returning `h3*`/`s2*`/`geohash*` row loop was checked for the discriminator, whether per-row output is driven by argument values. The seven left alone are not carriers, each row being bounded by a library constant instead. `Nullable` and `LowCardinality` arguments are unwrapped before `executeImpl`, so the checkpoint covers them. Tests removed at review request, being timing based. They did validate on the previous head before removal: Bugfix validation reported the bug reproduced on master HEAD and fixed here on [amd64](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113924&sha=f2d6b8ab4a01cd8e26efb86380aeec8d51a85096&name_0=PR&name_1=Bugfix%20validation%20%28functional%20tests%2C%20amd64%29), [aarch64](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113924&sha=f2d6b8ab4a01cd8e26efb86380aeec8d51a85096&name_0=PR&name_1=Bugfix%20validation%20%28functional%20tests%2C%20aarch64%29) and [unit tests](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113924&sha=f2d6b8ab4a01cd8e26efb86380aeec8d51a85096&name_0=PR&name_1=Bugfix%20validation%20%28unit%20tests%29). The fix now ships without regression coverage.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113924",
        "timestamp": "2026-08-12T20:45:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113937",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: update gui.mdx with CHOPs tool",
        "text": "Added CHOPs UI and Admin tool in the docs ### Changelog category (leave one): - Documentation (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113937",
        "createdAt": "2026-08-08T10:13:10Z",
        "updatedAt": "2026-08-13T15:35:43Z",
        "timestamp": "2026-08-13T15:35:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-documentation",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "rva-quantrail",
        "state": "closed",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113947",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "MySQL: reject an empty TLS contents override also when the collection stores contents",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112070 Related: https://github.com/ClickHouse/ClickHouse/pull/110615 The empty-override rejection in `StorageMySQL::getSSLParams` only ran when the base named collection stored a credential *path*: with the credential stored in the contents form (`ssl_ca_pem` / `ssl_cert_pem` / `ssl_key_pem`), `get_path` returned at the empty-path fast path before looking at `isQueryOverridden`, so a query could pass `ssl_ca_pem = ''` and silently strip the collection-provided CA or client certificate — e.g. disable the verification of the server certificate — on every MySQL surface using that collection. The check is now hoisted above the fast path, so an empty contents override is rejected regardless of the form the stored credential has. The same defect in the PostgreSQL counterpart was found by the AI review in #110615, where it is fixed the same way; this is the twin fix for the MySQL surface introduced in #112070. The existing empty-override test gains the contents-storing-collection case. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a defect in the `MySQL` integrations where overriding a TLS credential of a named collection with an empty `ssl_ca_pem`/`ssl_cert_pem`/`ssl_key_pem` value was accepted when the collection stored the credential in the contents form, silently dropping the configured CA or client certificate instead of rejecting the override. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1289` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113947",
        "createdAt": "2026-08-08T14:20:06Z",
        "updatedAt": "2026-08-13T00:19:22Z",
        "timestamp": "2026-08-13T00:19:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113983",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allowlist the expected FileLog bad-path reattach error in the upgrade check",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113781 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Upgrade check (amd_release)` intermittently fails its `Error message in clickhouse-server.log` sub-test on one benign line: ``` <Error> StorageFileLog (test_1.filelog_bad_path_attach): The absolute data path should be inside `user_files_path`(/var/lib/clickhouse/user_files/) ``` No product defect: the server starts, nothing crashes, no data is affected. Root cause. `04202_filelog_attach_path_outside_user_files` ATTACHes a FileLog table whose path is outside `user_files_path`. `ATTACH` is `LoadingStrictnessLevel::ATTACH` (2), which is `>= SECONDARY_CREATE` (1), so the constructor takes the relaxed branch at `src/Storages/FileLog/StorageFileLog.cpp:195-198`: it logs at `<Error>` and returns instead of throwing `BAD_ARGUMENTS`. That branch is deliberate and is what the test covers, since refusing to load at reattach time would break server startup. The table then outlives the test: stress threads run with a fixed `--database=test_N` (`ci/jobs/scripts/stress/stress.py`), and `clickhouse-test` skips its per-test teardown whenever `--database` is set (`need_cleanup = not args.database`), so that shared database is never dropped. The upgrade restart re-attaches the table, the relaxed branch fires again, and the line lands in the scanned log, where the post-restart scrub in `tests/docker_scripts/upgrade_runner.sh` had no entry for it. Hence the intermittency: `04202` must land on a fixed-database thread. Change. One `grep -av` entry in that scrub's existing secondary pipe, plus a short rationale comment next to the sibling entries. The pattern requires the fixture table name and the message together, and (bare parens are literals in BRE) the `StorageFileLog (db.table):` prefix shape. No source change, no test change. Validation. The scan pipeline, extracted verbatim from the runner, was run under GNU grep 3.11 against the failing run's own 19.9 MB `clickhouse-server.upgrade.log`. With the entry the artifact is empty; with it deleted the output is byte-identical to the 189-byte `upgrade_error_messages.txt` CI produced, so the sub-test flips `FAIL` to `OK`. Eight negative controls still surface, including a table whose name merely ends with the fixture name (`prod.other_filelog_bad_path_attach`), which the required `.` separator keeps visible.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113983",
        "createdAt": "2026-08-08T21:17:46Z",
        "updatedAt": "2026-08-13T17:50:47Z",
        "timestamp": "2026-08-13T17:50:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "manual approve",
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113984",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix logical error on filter push-down with a differently-typed same-name column",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix a logical error (`Unexpected return type`) when a filter over a view whose column type differs from the underlying storage (e.g. `information_schema.tables`, where `engine` is `Nullable(String)`, over `system.tables`, where it is `String`) is pushed down to the storage. Minimal reproducer: ```sql SELECT 1 FROM (SELECT 1 FROM information_schema.tables WHERE indexHint(toString(engine))); -- Code: 49. DB::Exception: Unexpected return type from toString. Expected Nullable(String). Got String. (LOGICAL_ERROR) ``` The predicate is typed against the view header, where `engine` is `Nullable(String)`. When `SourceStepWithFilter::applyFilters` rebuilds the filter with `ActionsDAG::buildFilterActionsDAG`, the name-based input replacement substituted the storage column `engine String` for the input while the parent `FUNCTION` nodes were rebuilt with their existing `function_base` — leaving the rebuilt DAG (including the DAG captured inside `indexHint`) internally inconsistent: `toString` still declared `Nullable(String)` while returning `String`. Evaluating it over the candidate-tables block in `getFilteredTables` then failed the return-type assertion in `ExpressionActions`. Two fixes: - `ActionsDAG::buildFilterActionsDAG`: skip a name-based input replacement when it would change the input type. Keeping the original input is safe — the subtree is then simply not evaluated over the storage columns, and the filter is still applied upstream. - `canEvaluateSubtree` in `VirtualColumnUtils`: match the allowed inputs by type as well as by name, as a fail-close guard at the evaluation site, so a mismatched predicate subtree is not pushed down. The exact-type check exposed a lying push-down sample in `StorageSystemTables`: both `detail::getFilteredTables` and `ReadFromSystemTables::applyFilters` declared `uuid` as `String` while the real column of `system.tables` / `system.detached_tables` is `UUID`, which would have silently disabled the `uuid` prefilter fast path. The samples now declare `UUID`, and a test pins the fast path via the `SelectedRows` profile event. Found by the AST fuzzer on two unrelated PRs (pre-existing on `master`; the same error also appears in `master` stress tests going back months): [AST fuzzer (amd_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113636&sha=bea222ed980b6db96146b4fe56dc968fa2280dd1&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29). Closes: https://github.com/ClickHouse/ClickHouse/issues/113982 Related: https://github.com/ClickHouse/ClickHouse/pull/113636 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1195` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113984",
        "createdAt": "2026-08-08T21:39:30Z",
        "updatedAt": "2026-08-13T04:46:36Z",
        "timestamp": "2026-08-13T04:46:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113996",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Observe the query time limit while collecting typo-correction hints",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/86768 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `max_execution_time` being ignored while a query was still being analyzed. Referencing an unresolved column of a deeply nested type sent typo correction into an enumeration of every subcolumn of every candidate column, exponential in the nesting depth and observing no time limit, so the query could keep running for minutes past its limit and after being cancelled. ### Description Reported on #86768, in the `Stress test (arm_tsan)` hung check ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=86768&sha=399ef4e99734b6a8c5cef4fd1bfbc2e724d33b8a&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29)): a query with `is_cancelled: 1` had run 1379 s against `max_execution_time: 10`, stack top in `SerializationObjectPool::getOrCreate` under an `enumerateStreams` recursion, `read_rows: 0`. Root cause is in the analyzer, not the serialization pool. An unresolvable identifier takes the typo-correction path, where `TypoCorrection::collectCompoundExpressionValidIdentifiers` enumerates every subcolumn of the candidate column looking for a suggestion. That substream tree doubles per nesting level, the walk repeats per column, table expression and scope, and nothing on it observes the time limit. Two changes in `TypoCorrection.cpp`: - Return before entering the walk when it cannot contribute: the insert needs prefix plus subcolumn name to have exactly as many parts as the unresolved identifier, and a subcolumn name has at least one part, so once the prefix alone is that long nothing can come out. That holds at every call site, so the hint set is unchanged by construction. `collectScopeValidIdentifiers` already guards its walk this way. - Poll `QueryStatus::checkTimeLimit` in the surviving walk, and before the early return so the limit does not depend on entering it. It reads the query's own watch rather than waiting to be told, so the limit also holds under `timeout_overflow_mode = 'break'` and in `clickhouse-local`, neither of which ever marks the query cancelled. A `KILL QUERY` still reports its own cause; with no query attached it is a no-op. Reproduced without the fuzzer: `Array(Map(String, Tuple(...)))` nested N deep plus `SELECT alias.nosuchcol FROM t AS alias`. At depth 12 that goes from 44 s to 2.0 s, and under a 1 s limit reports `TIMEOUT_EXCEEDED` instead of running 116 s. Hints are unchanged over a 15-shape matrix. `IDataType` and `ISerialization` are untouched, so other `forEachSubcolumn` callers are unaffected. Only this carrier is closed, so the note is amended, not dropped. <details><summary>Measurements</summary> Subcolumns enumerated per candidate column: 50, 106, 218, 442 at nesting depth 3, 4, 5, 6. Debug build, `Array(Map(String, Tuple(a T, b T)))` nested N deep, `SELECT alias.nosuchcol FROM t AS alias`: | depth | 7 | 9 | 10 | 11 | 12 | |---|---|---|---|---|---| | before | 0.75 s | 3.21 s | 7.58 s | 18.09 s | 43.97 s | | after | 0.68 s | 0.50 s | 0.63 s | 1.09 s | 2.01 s | At depth 10 with `max_execution_time = 1`: before 116 s and `UNKNOWN_IDENTIFIER`, after `TIMEOUT_EXCEEDED` in 1.2 s. A one-part identifier, which no walk can ever answer, went from 44 s to 0.09 s; with a second table expression the same shape went from 35 s to 0.84 s. The two modes that never mark a query cancelled, depth 12 under a 0.001 s limit: `timeout_overflow_mode = 'break'` returned `UNKNOWN_IDENTIFIER` after 0.67 s with the limit ignored, and now returns `TIMEOUT_EXCEEDED` in 0.20 s; `clickhouse-local` went from 1.92 s with the limit ignored to `TIMEOUT_EXCEEDED` in 1.26 s. A `KILL QUERY` at depth 15 reports `QUERY_WAS_CANCELLED`, never a timeout, 8 of 8. Pre-fix the one-part shape took 348 s under a sanitizer build and 23 s on a release build, so the committed test's limits hold on every build type rather than only on debug. Hint set compared over 15 shapes (bare, alias-qualified, table-qualified and database-qualified prefixes; `Tuple`, `Nullable`, `LowCardinality`, `LowCardinality(Nullable)`, `Array(Tuple)`, `Map`, `Variant`, `Dynamic`, `JSON`, nested `Tuple`): byte-identical before and after, 12 of the 15 carrying a real hint. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113996",
        "createdAt": "2026-08-09T00:12:46Z",
        "updatedAt": "2026-08-13T16:33:43Z",
        "timestamp": "2026-08-13T16:33:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:113998",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not merge patch parts across a pending mutation version",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/98898 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a permanently stalled mutation on `ReplicatedMergeTree` tables with lightweight updates enabled. A merge of patch parts could produce a patch whose source data versions span the version of a mutation that is still queued, which no mutation can then apply, so the `MUTATE_PART` entry failed and retried forever (`Found patch part ... that intersects mutation with version ...`). Such merges are now postponed until the mutation completes. ### Description Related: https://github.com/ClickHouse/ClickHouse/issues/98898 A patch part records the data-version range of what it patches, and applies whole or not at all. Merging patch parts unions those ranges, so merging patches for versions 1 and 6 yields one spanning 1..6. A mutation cutting at version 2 cannot apply it, `patchHasHigherDataVersion` throws, and the `MUTATE_PART` entry retries forever. Debug and sanitizer builds abort, hence the stress failures. Root cause: the merge predicate already refuses parts whose current mutation versions differ, but derives them from `mutations_by_partition`, where `getCurrentMutationVersion` returns 0 for an absent partition. The finished-mutation cleaner removes the `/mutations` znodes once every replica's `mutation_pointer` passed them, while the `MUTATE_PART` entries survive. Every patch then maps to 0 and every pair looks mergeable. The logs show `There are no mutations for partition ID all` just before the abort. The fix takes the pending versions from the replication queue instead, which is durable: a queued `MUTATE_PART` entry's `new_part_name` encodes its target version. A patch `MERGE_PARTS` whose sources would span one is postponed, so it proceeds once the mutation completes. No assertion is relaxed, and the cost is one walk of a queue the neighbouring guards already walk. The repro is deterministic: pristine master aborts with the exact message, the fix passes, and deleting only the new guard call brings the abort back. The test asserts the mutation completes, the queue drains and the data is fully patched, so no deadlock or dropped patch passes. Two limits. Like the guards beside it, this is an execution-side per-replica check: it stops a replica merging across a mutation queued there, not every route by which a spanning patch could reach one. And it does not repair a patch already spanning a version on disk, so such a table stays stuck; that needs per-row `_data_version` filtering when applying patches.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/113998",
        "createdAt": "2026-08-09T01:04:11Z",
        "updatedAt": "2026-08-13T10:02:32Z",
        "timestamp": "2026-08-13T10:02:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "CurtizJ"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114001",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Normalize open range bounds only for integer key types",
        "text": "Closes https://github.com/ClickHouse/ClickHouse/issues/113993 `Range` previously tightened every `Int64` or `UInt64` endpoint without knowing the key type it would be compared against. Move that normalization into `KeyCondition` after checking the key type is represented by an integer, preserving integer pruning while keeping fractional and future non-integer domains open. Related: https://github.com/ClickHouse/ClickHouse/issues/113993 Caused by: https://github.com/ClickHouse/ClickHouse/pull/98410 <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect MergeTree index pruning for strict comparisons between fractional types and integer constants.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114001",
        "createdAt": "2026-08-09T01:50:42Z",
        "updatedAt": "2026-08-13T14:39:25Z",
        "timestamp": "2026-08-13T14:39:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "v26.4-must-backport"
        ],
        "author": "EmeraldShift",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114003",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix Keeper CORRUPTED_DATA after cross-segment writeAt crash (#112101)",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix ClickHouse Keeper refusing to start with `CORRUPTED_DATA` after a crash during cross-segment Raft log truncation (`writeAt`), which could leave an acknowledged log entry stranded behind stale changelog segments. Closes #112101 ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Detail When `Changelog::writeAt` rewrote an entry in an earlier changelog segment, later segments were removed asynchronously and the rewrite could be fsynced/acked before those unlinks finished. A crash in that window left a duplicated index in the earlier file plus stale higher-index files; startup treated the gap as unrecoverable corruption. This change waits for those `RemoveChangelog` operations **outside** `writer_mutex` before appending the rewrite (durability ordering; no lock cycle with the remove thread). Startup still reports `CORRUPTED_DATA` for changelog gaps so recovery stays manual. ### Test plan - [x] `gtest_coordination_changelog` / `ChangelogTestWriteAtPreviousFile` — after cross-segment `write_at`, superseded changelog files are already gone before the rewrite is acknowledged",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114003",
        "createdAt": "2026-08-09T02:44:31Z",
        "updatedAt": "2026-08-13T13:34:31Z",
        "timestamp": "2026-08-13T13:34:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "nishant-uxs",
        "state": "open",
        "assignees": [
          "antonio2368"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114009",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Observe the deadline inside one string search",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/112203 Related: https://github.com/ClickHouse/ClickHouse/pull/113369 --> Related: https://github.com/ClickHouse/ClickHouse/issues/112203 Related: https://github.com/ClickHouse/ClickHouse/pull/113369 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `countSubstringsCaseInsensitive` and `countSubstringsCaseInsensitiveUTF8` now respect `max_execution_time` and `KILL QUERY` while searching one value. Previously a single substring search ran to completion inside one function call, so a query over a large value, or over many rows in which the needle does not occur, could keep running for tens of seconds past its deadline. ### Description Follow-up to #113369, whose checkpoint is charged *between* searches, so one `searcher.search` call still ran uninterruptibly. **What breaks.** A one-second deadline observed after 13.4 s, no settings overrides: ```sql SELECT countSubstringsCaseInsensitiveUTF8(materialize(repeat(repeat('Ж', 200), 250000)), repeat('Ж', 16) || 'Щ') FORMAT Null ``` Not only large values: 260000 ordinary 200-byte rows reproduce it at 13.8 s, since `vectorConstant` searches a whole block at once. `KILL QUERY SYNC` blocked 53 s on 400 MB. The ASCII case-insensitive searcher is affected too: a needle failing on its last byte takes 46.4 s on a 1 GiB row against a 1 s deadline. **Root cause.** Nothing under `src/Common` checked cancellation. Candidate comparison, not bytes scanned, dominates: a candidate walks the whole needle before failing. **The change.** The scanning loops take an optional `CancellationBudget *` and charge it per step; `CountSubstringsImpl` passes the budget it already builds per call. The budget header moves to `src/Common` for that (type unchanged). Both case-insensitive searchers charge; the case-sensitive one hands its whole range to StringZilla in one call, so it has no point to charge from. An uncharged scan is a separate instantiation and does not regress. No returned value changes. **Why Improvement and not Bug Fix.** The only observable is elapsed time, so no test can fail on master HEAD without bounding a duration, and #114061 removed seven tests of that shape as flaky. #108192 and #107929 fixed this same class as Improvement. `04829` asserts counts only; cancellation promptness stays unpinned. **Still unfixed.** `position*` and `multiSearch*` share the defect: only `countSubstrings*` passes a charger, and `MultiVolnitskyBase` takes none. No comparison has a checkpoint inside it, so a needle mismatching at its last byte overruns by one comparison: a bounded 0.85-1.03 s at the largest needle `repeat` allows.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114009",
        "timestamp": "2026-08-12T20:45:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114010",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not create unused aggregate states in aggregation in order with a partial GROUP BY key",
        "text": "Do not create unused aggregate states in aggregation in order with a partial GROUP BY key <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114000 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114000 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a memory leak with `optimize_aggregation_in_order = 1` when the table sorting key is a strict prefix of the `GROUP BY` key and an aggregate function whose state owns heap memory is used, such as `quantileDD`. Server memory grows with every such query until restart. ### Description `AggregatingInOrderTransform` has two output modes. When the sorting prefix covers the whole `GROUP BY` key, each state created by `createStatesAndFillKeyColumnsWithSingleKey` (stored in `variants.without_key`) is handed to a `ColumnAggregateFunction` by `addSingleKeyToAggregateColumns`, which also clears the pointer. In the *partial* key mode (`group_by_key == true`, e.g. table `ORDER BY parent_key` with `GROUP BY parent_key, child_key`) the result comes from the hash table via `prepareChunkAndFillSingleLevel`, which never reads `without_key`, and both `addSingleKeyToAggregateColumns` call sites are guarded by `if (!group_by_key)`. So that state is write-only: the next call overwrites the pointer and orphans the previous state, and the following `variants.invalidate()` sets the type to `EMPTY`, making `destroyAllAggregateStates` return early. Arena bytes are freed; the state's owned allocations are not. The producer was left unconditional when the mode and all its `!group_by_key` guards were added in `3931dbd848786e8` (2022-03-06), so the guard pair has been asymmetric since then. It is invisible for arena-only states, which is why the in-tree test of this plan shape is green with `count()`. This completes the guard pair: the key-column fill is split into `fillKeyColumnsWithSingleKey`, used by the two partial-key call sites, so no state is created where none is consumed. The `!group_by_key` path is unchanged. `overflow_row` cannot be set in this mode, so the key-only path needs no overflow-row state. The test runs its queries through `clickhouse local`, whose at-exit LeakSanitizer check aborts on a leaked state, so the leak is an ordinary test failure on any ASan build instead of something only a stress job sees. It also pins the plan shape and the results; on a non-ASan build only those are checked. Verified with `clickhouse-test`: FAIL on pristine master (1600 bytes leaked in 30 allocations), OK here, 20/20 randomized. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1276` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114010",
        "createdAt": "2026-08-09T05:37:58Z",
        "updatedAt": "2026-08-13T09:29:26Z",
        "timestamp": "2026-08-13T09:29:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "nickitat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114027",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Drop stale totals and extremes ports in MergingAggregatedStep",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/issues/113708 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required. The only shape found to reach this logical error needs `inject_random_order_for_select_without_order_by`, documented as only useful for testing and development, so no released configuration is known to be affected. It is fuzzer-reachable (BuzzHouse randomizes it), which is how the abort was found. ### Description Found 2026-08-08 by `AST fuzzer (amd_debug)` (STID 0993-250f, [report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112384&sha=32e19e343e8bbf3cc8bb41f5bdaad2471caa7c05&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)) on PR 112384 as a bystander: `Logical error: Block structure mismatch in function connect between Limit and MergingAggregatedBucketTransform`. It aborts in `Port.cpp:22` `connect` from `MergingAggregatedMemoryEfficientTransform.cpp:577`, so at pipeline construction, not execution. `MergingAggregatedStep::transformPipeline` inherited the pipeline's totals and extremes ports, then attached its merging transforms to them. `addMergingAggregatedMemoryEfficientTransform` uses the `StreamType`-less `Pipe::addSimpleTransform` overload, so the getter also runs on `totals_port`. The transform it creates has an empty input header, while that port still carries the aggregate-state header, hence the mismatch. Where the transform is not attached to those ports, the stale port survives into `TotalsHavingStep`, which requires it to be null. The fix mirrors `AggregatingStep.cpp:398-399`, the other half of two-stage aggregation: drop the current totals and extremes, which this step invalidates and are recalculated afterwards. One call before the branch point covers all three merge branches, the extremes stream and all five construction sites. The other caller needs no drop: it builds pipes from `SourceFromNativeStream`, so neither port exists. Reachability decides the category: two `Merge` children must reach the step still carrying a totals port, and the setting named above is the only producer found among candidates measured with an instrumented `hasTotals`. The guard is worth adding anyway, since the state is constructible today. The test asserts the build-time shapes and each arm's merge branch with `EXPLAIN PIPELINE`, plus values against a hand-derivable ground truth. Two unrelated defects, identical on both sides of this change: `Chunk info was not set for chunk in GroupingAggregatedTransform`, owned by #113266, and the same setting over `merge()` loses a child's rows. An earlier revision of this line attributed the first to `MaterializingTransform` dropping chunk info; that is wrong, since it only calls `detachColumns`/`setColumns` and the `ISimpleTransform` default swaps `chunk_infos` through. Issue 113708 reports this guard firing over `merge()` through a set-operation plan with no `MergingAggregatedStep`; that did not reproduce, so this PR does not close it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114027",
        "createdAt": "2026-08-09T10:34:40Z",
        "updatedAt": "2026-08-13T05:53:27Z",
        "timestamp": "2026-08-13T05:53:27Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114035",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not throw when comparing an Enum column with a non-member string literal",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111545 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `UNKNOWN_ELEMENT_OF_ENUM` being thrown when an `Enum` column is compared with a string literal that is not one of its members, even though `validate_enum_literals_in_operators` is disabled by default. Previously the query was rejected if the column was in the primary key or partition key, or merely carried a `minmax`, `set` or `bloom_filter` index, and `!=` was rejected for any `Enum` column. Closes [#111545](https://github.com/ClickHouse/ClickHouse/issues/111545). ### Description Closes: #111545 `validate_enum_literals_in_operators` is off by default and documents that non-member enum literals are then not validated, yet comparing an `Enum` column with one threw, and whether it threw depended on physical layout. Two independent causes: 1. **Index analysis.** `KeyCondition::extractAtomFromTree` converts the literal with the throwing `convertFieldToType`, leaving the `isNull()` handling right below it unreachable. Since `minmax` and `set` build a `KeyCondition` over the index expression, that one line served the primary key, the partition key and both indexes. `bloom_filter` repeated the defect at four of its own sites, as did the JSON-subcolumn helper shared by `bloom_filter` and `tokenbf_v1`. The conversions are now wrapped in a catch scoped with `isParseError` (following `ConditionSelectivityEstimator`), so a non-representable literal declines the atom (full scan, never a wrong answer) while fatal errors such as `MEMORY_LIMIT_EXCEEDED` still propagate. 2. **An inverted guard.** In `FunctionsComparison.h`, `!equals && not_equals` selects `notEquals` alone, yet the value it guards, `IsOperation<Op>::not_equals`, was written to serve `notEquals`, so the tolerant branch could only ever return 0. Deleting the guard is the fix: any rewrite that keeps a condition there breaks the four ordered operators, for which both traits are false. Two corrections to the report, both re-measured. `e < '4'` does not throw on a non-key column (the ordered operators all return 0), so there only `!=` was broken; and the column need not be in a key, since a `minmax`, `set` or `bloom_filter` index suffices. `e <=> '4'` and `isDistinctFrom(e, '4')` were affected too. No query that previously succeeded changes its answer, and declining loses no pruning: `LIKE` over an `Enum` key never pruned before, since the converted field is an `Int8` that the `like` atom handler rejects. The tests assert that each index is still used for representable literals, and that the setting still rejects every operator when enabled. <details><summary>Validation</summary> Every carrier measured on a debug build, before and after, each fixture carrying a positive control: | Carrier | Before | After | Control | |---|---|---|---| | primary key `ORDER BY e` | 691 | 0 | `= 'a'` -> 1 | | partition key, `LIKE '%Beta%'` | 691 | 500 | `use_partition_pruning = 0` -> 500 | | `minmax` index, no key | 691 | 0 | `use_skip_indexes = 0` -> 0 | | `set(0)` index, no key | 691 | 0 | `use_skip_indexes = 0` -> 0 | | `bloom_filter` index, no key | 691 | 0 | `use_skip_indexes = 0` -> 0 | | `bloom_filter` over `Array(Enum)`, `has` / `hasAny` / `hasAll` | 691 | 0 | `has(a, 'a')` -> 1 | | `bloom_filter` over `Map(Enum, ...)` keys, `mapContains` | 691 | 0 | `mapContains(m, 'a')` -> 1 | | `JSONAllPaths` `bloom_filter` / `tokenbf_v1` | 691 | 0 | `= 'y'` -> 1, index used | | `!=` on any table | 691 | row count | `NOT (e = 'zzz')` agrees | | `<=>`, `isDistinctFrom` | 691 | 0 / row count | member-literal arms | | `< <= > >=` on a non-key column | 0 | 0 | unchanged | | `IN ('4', 'a')` | 1 | 1 | unchanged (#72686) | Each of the four hunks was reverted independently: every one reddens its own carrier and only its own carrier. The tests assert index *usability* through `force_data_skipping_indices` rather than row counts alone, so neither an over-broad decline nor a missing one can pass silently; four mutations were run and each reddens a new assertion. `Nullable(Enum8)`, `Nullable(Enum16)` and `Enum16` keys covered, including a real `NULL` row (`NULL != '4'` stays `NULL`; `NULL <=> '4'` is 0). `LowCardinality(Enum)` is not covered because `DataTypeEnum` does not override `canBeInsideLowCardinality()`, so the type cannot be constructed. 50 randomized runs of both new tests passed 100/100; `01310_enum_comparison`, `03278_enum_in_unknown_value`, `04049_statistics_enum_invalid_value` and `04539_enum_string_search_dictionary` stay green. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114035",
        "createdAt": "2026-08-09T13:50:23Z",
        "updatedAt": "2026-08-13T15:59:32Z",
        "timestamp": "2026-08-13T15:59:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "nihalzp"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114043",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add a full-featured AI agent to the client (the `?` command)",
        "text": "Turns the primitive `??` \"generate one SQL query\" helper in the client into a full interactive AI agent, available with a single `?` (`??` is kept as an alias). The agent works over a live connection: it is given the context of the recent queries in the session (with results truncated to the first and last lines to save tokens, plus error messages) and can use tools to read the query history from `system.user_query_log`, inspect the schema (`SHOW`/`DESCRIBE`/`SHOW CREATE`), consult the embedded documentation (`system.documentation`, the same source as the `help` command), run read-only queries without confirmation under a sandbox (`readonly = 1`, a 30-second and a 10 GiB limit, no table functions reaching outside of the server's tables; if the session forbids applying these limits, the query fails instead of running without them), and run any other query after asking the user for confirmation. Queries it runs are echoed and executed on the user's connection and displayed exactly as if the user had typed them; it prints its thoughts and tool calls as it works, and the conversation keeps its context across `?` invocations within a session. When no client-side AI provider is configured (neither the `ai` section of the client configuration nor the `OPENAI_API_KEY`/`ANTHROPIC_API_KEY` environment variables), the agent falls back to the server-side `aiGenerate` function of the connected server (or of `clickhouse-local`) when it has default credentials configured for it (`ai_function_text_default_credentials`), so it works with no client-side setup when the administrator has already enabled AI functions. Implementation lives in `src/Client/AI/`: the agent loop (`AIAgent`), two model backends (`AIAgentTransport`: native tool-calling via `ai-sdk-cpp`, and a text tool-call protocol on top of `aiGenerate`), the tools (`AIAgentTools`), the read-only statement allowlist (`AIQueryValidation`), and the recent-query context buffer (`QueryContextBuffer`). Unit tests cover the tool-call protocol parsing, the read-only validation, and the context buffer; the end-to-end loop (both backends and the confirmation gate) was validated by driving `clickhouse-local` against a mock OpenAI endpoint. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The embedded AI assistant of `clickhouse-client` and `clickhouse-local` is now a full agent, invoked with a single `?`. It sees the recent queries and their results, explores the schema and the documentation, runs read-only queries on its own (and other queries with confirmation) displayed as if you typed them, and can use the server-side `aiGenerate` function when no client-side AI provider is configured. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114043",
        "createdAt": "2026-08-09T15:33:43Z",
        "updatedAt": "2026-08-13T11:57:26Z",
        "timestamp": "2026-08-13T11:57:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114053",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not run the LeakSanitizer check on the forced exit path",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/112846 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description On ASan builds the server's forced-shutdown branch calls `safeExit(0)`, which runs `__lsan_do_leak_check()`. That branch is taken precisely because handlers or refresh tasks missed the drain timeout, so the check stops the world while those threads are mid-query and classifies chunks they still own, producing unattributable reports. Two were reported on 2026-08-08 as a `uniqExact` hash-table \"direct leak\" (2 MiB on #112846, [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112846&sha=2e6a38112ad4dacef62f5cc4cdb396cc297282eb&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29); 4 MiB on #101264): each a single allocation with zero indirect leaks, while two green `asan_ubsan` jobs on the same commit also force-shutdown with connections live and report nothing. Since #113608 made such reports nameable, they compete with real ones. `safeExit` takes a second parameter, a `LeakCheck` enum defaulted to `Run`, so existing callers are unchanged. The three callers reached while other threads still run pass a skip: `Server.cpp` writes one stderr line recording that coverage was skipped rather than passed; `clickhouse-local` and the client's SIGINT/SIGQUIT handler skip **silently**, since neither redirects fd 2, so a notice there would be program output read as a test failure. Keeper's identical-looking branch is unchanged: it calls `server_pool.joinAll()` first. Clean shutdowns never reach `safeExit`. That notice is allow-listed in the one scanner that can see it, `sanitizer_hits` in `ci/jobs/scripts/clickhouse_proc.py`, matched as a whole line so a report sharing a line with it is still blamed; a new `ci/tests` test pins that. `tests/clickhouse-test` needs none: its `IGNORED_SANITIZER_ERRORS` filters only the sanitizer runtime's `log_path` files, and the carriers that runner sees skip quietly. Validated on an ASan build: the notice appears on the forced branch and no leak report does, while the same tree with the guard removed runs the check. Side effect: `_exit()` drops the async logger queue, so `Will shutdown forcefully.` is now missing on ASan builds, as it already is where no check delays the exit. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1312` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114053",
        "createdAt": "2026-08-09T17:34:15Z",
        "updatedAt": "2026-08-13T09:31:29Z",
        "timestamp": "2026-08-13T09:31:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "manual approve",
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexbakharew"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114054",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"Use libdeflate for gzip/zlib/deflate compression and decompression\"",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114045 Related: https://github.com/ClickHouse/ClickHouse/pull/108074 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Revert #108074 (use libdeflate for gzip/zlib/deflate): its streaming decompressor suspends only at DEFLATE block boundaries, so decompressing an HTTP request body whose single DEFLATE block spans the whole stream (the shape zlib-ng produces at level 1, which is the default in the official .NET SDK) was `O(n^2)` in the compressed block size and buffered the whole block in memory. A 22 MB gzip body that 26.6 ingested in 0.3 s took 15 s on 26.7. This returns the gzip/zlib/deflate paths to zlib-ng, restoring 26.6 behavior. The `gzip`/`deflate` compression levels `10-12` that 26.7 introduced stay accepted and are clamped to `9`, the maximum level supported by `zlib`, so configurations written against 26.7 keep working. ### Details This is a plain `git revert` of the merge commit of #108074, with the following manual adjustments: - `tests/queries/0_stateless/04230_iceberg_optimize_metadata_not_initialized_104711.sh` and `04240_iceberg_supports_parallel_insert_mv_104891.sh` keep their current (post-#108074) form: they now corrupt Iceberg metadata files directly instead of relying on `output_format_compression_level=11` being rejected, which works regardless of the compression backend. - The relocated doc `docs/reference/statements/select/into-outfile.mdx` documents the compression level range as `1-12` for `gzip`/`deflate` with the clamp, since the old `docs/en` tree no longer exists for the revert to apply to. The text for `http_zlib_compression_level` is changed in the `DECLARE` docstring in `src/Core/Settings.cpp`, because `docs/reference/settings/session-settings/http.mdx` is inside an `AUTOGENERATED` region and hand-editing it fails the `No direct edits to generated or read-only docs` check. **Backward compatibility.** 26.7 shipped `libdeflate` and, with it, `gzip`/`deflate` compression levels `10-12` on every output surface that routes through `CompressionMethod`: `INTO OUTFILE ... LEVEL`, `output_format_compression_level`, `http_zlib_compression_level` and the gRPC `output_compression_level`. A plain revert would narrow the accepted range back to `1-9`, so existing queries and settings profiles would start failing with `Invalid compression level` or `deflateInit2 failed: stream error`. To avoid that, `getCompressionLevelRange` keeps returning `1-12` for `gzip`/`zlib` and `createWriteCompressedWrapper` clamps levels above `9` to zlib's maximum. Levels above `12` are still rejected, as in 26.7. Two tests are restored in a backend-agnostic form instead of being deleted with the feature: - `04842_parquet_gzip_compression_level` round-trips a mixed-compressibility dataset through Parquet with `output_format_parquet_compression_method='gzip'` at levels `1, 3, 6, 9, 12`, keeping non-default-level coverage of `output_format_compression_level` for Parquet. - `04843_outfile_compression_level_range` covers the accepted `INTO OUTFILE ... LEVEL` range for `gzip`/`deflate`, the clamp of `10-12`, and the rejection of `13`. A follow-up PR will re-introduce libdeflate with the streaming decompressor fixed to suspend at symbol granularity (linear time, bounded memory).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114054",
        "createdAt": "2026-08-09T17:34:24Z",
        "updatedAt": "2026-08-13T00:24:29Z",
        "timestamp": "2026-08-13T00:24:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "submodule changed",
          "v26.7-must-backport"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114055",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix flaky test_zookeeper_fallback_session by waiting for zoo1 to catch up before asserting fallback",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113575 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `test_fallback_session` fails as `assert 'zoo2' == 'zoo1'` at `test.py:112`: 5 hits on 5 unrelated PRs in 30 days, 7 on 7 over 60 days, on four build flavours. No master hits; one further 60-day row is on the `26.5` release branch. Root cause: the test asserts a convergence the contract does not promise. Since `c39cac20ff8` Keeper refuses a reconnecting client whose `last_zxid_seen` exceeds that host's applied zxid (\"The client should try another server\", `KeeperTCPHandler.cpp:352`), and the client keeps that watermark across session restarts, so `connect` falls through to the next host. In the final phase zoo1 is a follower applying the log asynchronously (`async_replication=1`), so a node reconnecting in that window lands on zoo2 and stays: `in_order` has `hasOptimalNode() == false`, so the `ZKReconnect` task that would move it back is never armed. A longer retry budget cannot help. Fix: before blocking zoo3 and asserting re-convergence, wait until zoo1 has applied everything the nodes have seen, closing the multi-second catch-up window these failures came from. Not absolute: the cluster keeps committing, so a watermark can still advance between wait and handshake. On timeout the wait reports both zxids instead of the bare `'zoo2' == 'zoo1'`. `srvr`'s `Zxid` is the right oracle: it and the handler's check both read `KeeperStorage::getZXID()`. Validated by freezing zoo1's inbound Raft traffic. Without the wait it fails with the original signature and zoo1 logs `Refusing session as the client has seen zxid 161 while our last processed zxid is 146`; with it it passes, no refusals. A control arm keeping the helper but not awaiting it fails again. Sequential runs green: 50 on the previous revision, 10 on this one. A production `in_order` deployment can be pinned the same way; that Keeper change is for its owners. Related: #113575, where a separate fix PR was requested. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1278` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114055",
        "timestamp": "2026-08-12T21:14:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114057",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Convert test 04630_merge_over_stale_packed_tmp_dir to an integration test",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113978 Related: https://github.com/ClickHouse/ClickHouse/pull/111652 Test `04630_merge_over_stale_packed_tmp_dir` (added in https://github.com/ClickHouse/ClickHouse/pull/111652) copies a packed part directory with plain `cp` to simulate a stale `tmp_merge_` directory left by an interrupted merge. Modifying table data on disk is not allowed in stateless tests at all: the stateless suite runs against arbitrary server configurations (object storage, shared merge tree, encrypted disks), where direct filesystem manipulation is either meaningless or destructive — this test corrupted shared S3 blob reference counts in stress runs (https://github.com/ClickHouse/ClickHouse/pull/113978#pullrequestreview-4891978370) before it was pinned to the local disk in https://github.com/ClickHouse/ClickHouse/pull/113978. This PR moves the scenario to the integration test `test_packed_io::test_merge_over_stale_packed_tmp_dir`, where the cluster environment is fully controlled and the part directory can be copied safely inside the container. The test logic and all assertions are unchanged: the merge must reclaim the leftover `tmp_merge_all_1_2_1` directory (verified via the `Removing stale temporary directory` log message and the directory being gone), must not seed any data from it into the new part, and the resulting packed part must pass `CHECK TABLE`. Verified locally: the new integration test passes. A style check preventing new stateless tests from modifying the server's data directory is added separately in https://github.com/ClickHouse/ClickHouse/pull/114056. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114057",
        "createdAt": "2026-08-09T17:59:57Z",
        "updatedAt": "2026-08-13T00:20:37Z",
        "timestamp": "2026-08-13T00:20:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114059",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix quadratic gzip/zlib streaming decompression of single-block streams",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114045 Related: https://github.com/ClickHouse/ClickHouse/pull/108074 Related: https://github.com/ClickHouse/ClickHouse/pull/114054 Related: https://github.com/ClickHouse/libdeflate/pull/6 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `O(n^2)` decompression of gzip/zlib/deflate streams whose single DEFLATE block spans the whole stream, the shape zlib-ng produces at compression level 1 (the default of the official .NET SDK): a 24 MiB gzip HTTP body that took 11 s to ingest in 26.7 now takes 0.12 s, identical to a multi-block body, with memory bounded by the buffer size instead of the block size. ### Details The root cause: the streaming decompressor added to `contrib/libdeflate` kept its resume checkpoint only at DEFLATE block boundaries. On input exhaustion or output-buffer overflow it rolled back to the start of the current block, so for a stream that is one giant block, every socket refill re-decoded the block from its beginning — quadratic time — and the whole block's output had to be buffered before any byte was exposed — unbounded memory. The work happens in `copyDataImpl` while assembling the query, before the query starts, so it was invisible in `system.processes` and `query_duration_ms`. The fix (https://github.com/ClickHouse/libdeflate/pull/6, vendored here as a submodule bump) re-takes the checkpoint at every symbol boundary in the generic decode loop, plus two new decompressor fields (`in_block`, `block_is_final`) that let a resumed call jump straight back into symbol decoding — the litlen/offset decode tables already persist in the decompressor between calls. The fastloop is untouched: suspensions can only trigger from the generic loop, which always runs between the fastloop and input exhaustion. A suspension now loses at most one partially decoded symbol, so decompression is linear-time with bounded memory regardless of block structure. `LibdeflateInflatingReadBuffer` needed no logic changes, only comment updates: its \"grow the output buffer\" path is now reachable only when a single match (≤ 258 bytes) or stored block (≤ 64 KiB) exceeds the free output region. Verification: - Standalone harness: 5000+ randomized round-trips (zlib streams at all levels with random flush points; crafted static- and dynamic-Huffman single-block streams with matches reaching the full 32 KiB window; input chunking down to 1 byte; random output regions), zlib inflate as decode oracle, one-shot decoder as cross-check, and a progress assertion on every suspension. - End-to-end with the issue's reproducer shape (24 MiB single-block gzip body, `Transfer-Encoding: chunked`, local release build): **11.25 s before → 0.12 s after**, now identical to a multi-block body of the same content (0.12 s). Verified the vendored zlib-ng really emits this shape: at level 1 it produces 1 block for a 4 MiB input (stock zlib: 80 blocks). - No throughput regression on normal multi-block streams: 685 → 697 MB/s on a 256 MB zlib-level-6 stream. New tests: - `LibdeflateSingleBlock.StreamingConsumesInputWithinBlock` pins the linearity contract (a suspension inside a Huffman block leaves at most a few bytes unconsumed). Against the pre-fix library it fails immediately, with 35 MB left unconsumed on a 32 MiB single-block stream. - `LibdeflateInflateTest.RoundTripFromZlibNgLevelOneLarge` decodes a large stream produced by the vendored zlib-ng at level 1 — the real-encoder shape of the issue. - `LibdeflateInflateTest.SingleDeflateBlockSpanningWholeMember` decodes crafted single-block gzip/zlib members through `LibdeflateInflatingReadBuffer`. - `04836_http_gzip_single_deflate_block` ingests a crafted single-block gzip body over HTTP with chunked transfer encoding, end to end. Because the pre-fix decoder is correct, just quadratic, the assertion is on time: the body is one 48 MB block and the request is capped at 15 s. Ingest time for a single-block body of a given size, release build, `master`'s `contrib/libdeflate` vs this branch's: 10 MB - 1.8 s vs 0.05 s; 20 MB - 6.9 s; 40 MB - 27.3 s vs 0.20 s; 48 MB - 39.1 s vs 0.25 s. So the limit sits 2.6x below the pre-fix time and 60x above the fixed one, and the clean quadratic growth means the pre-fix margin does not depend on the machine speed. The payload is 120 lines of 400 KB so that line parsing and the `MergeTree` write do not dominate. Note: `Bugfix validation (unit tests)` reports `ERROR` (inconclusive, non-blocking) here for a structural reason: the fix itself lives in `contrib/libdeflate`, and the job cannot populate its merge-base \"before\" worktree at the merge-base submodule revision, so it refuses to build a before-binary that would validate the wrong submodule code. The functional test above is what validates the fix. Note: this PR and the revert #114054 are alternatives on master — if the revert merges first, this branch will be updated to re-apply the integration together with the fix; if this merges first, the revert can be closed (the 26.7 backport of the revert can proceed independently).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114059",
        "createdAt": "2026-08-09T18:13:29Z",
        "updatedAt": "2026-08-12T23:54:43Z",
        "timestamp": "2026-08-12T23:54:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "submodule changed"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114067",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "fix(Analyzer): skip rerunFunctionResolve for 'exists' nodes created by rewrite_in_to_join",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114026 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed `Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION)` thrown when a `PREWHERE` clause contains an `IN (subquery)` predicate and `rewrite_in_to_join = 1` (or `make_distributed_plan = 1`, which force-enables it) is set. The same query spelled with `WHERE` worked correctly. --- ### Problem With `rewrite_in_to_join = 1`, the analyzer rewrites `x IN (subquery)` into an `exists(...)` `FunctionNode` resolved via `FunctionExists`, a special function that is not registered in `FunctionFactory`. When `PREWHERE` is resolved, `ReplaceColumnsVisitor` calls `rerunFunctionResolve` on every `FunctionNode` in the predicate. `rerunFunctionResolve` (`src/Analyzer/Utils.cpp`) already special-cases `grouping` (also resolved outside the factory), but not `exists`, so it called `FunctionFactory::instance().get(\"exists\", context)` and threw `UNKNOWN_FUNCTION`. ### Fix Two parts: 1. `PREWHERE` is evaluated by the reading step and cannot execute a correlated subquery — the planner rejects one with `ILLEGAL_PREWHERE`. So the `rewrite_in_to_join` rewrite is now skipped while resolving a `PREWHERE` expression, and the plain `IN` is kept there. `PREWHERE x IN (subquery)` then returns the same result as its `WHERE` spelling, which is what the issue asks for. Subqueries nested inside `PREWHERE` still rewrite their own `IN` predicates. 2. `exists` is added to the special-case early return in `rerunFunctionResolve`, next to `grouping`. This matters for an explicitly written `PREWHERE EXISTS (correlated subquery)`, which is genuinely unsupported: it is now reported honestly as `ILLEGAL_PREWHERE` instead of `Unknown function exists`. ### Test `tests/queries/0_stateless/04820_rewrite_in_to_join_prewhere_exists.sql` covers `PREWHERE ... IN (subquery)` against the `WHERE` control arm, `NOT IN`, tuple `IN`, an `IN` nested in a subquery inside `PREWHERE`, and the explicit `PREWHERE EXISTS (...)` case asserting `ILLEGAL_PREWHERE`. ### Reproduction ```sql CREATE TABLE t (k UInt64, s String) ENGINE = MergeTree ORDER BY k; INSERT INTO t SELECT number, toString(number % 2) FROM numbers(1000); -- Threw: Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION) SELECT count() FROM t PREWHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; -- Worked (control) SELECT count() FROM t WHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; ```",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114067",
        "createdAt": "2026-08-09T20:05:33Z",
        "updatedAt": "2026-08-13T16:40:58Z",
        "timestamp": "2026-08-13T16:40:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "RohithPariki",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114070",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Convert stateless tests that modify the server's data on disk to integration tests",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113978 Related: https://github.com/ClickHouse/ClickHouse/pull/114056 Related: https://github.com/ClickHouse/ClickHouse/pull/114057 Stateless tests are not allowed to modify the server's data on disk at all (https://github.com/ClickHouse/ClickHouse/pull/113978#pullrequestreview-4891978370): the stateless suite runs against arbitrary server configurations — object storage, shared merge tree, encrypted disks — where the local part layout either does not exist or does not mean what the test assumes, and modifying it corrupts shared state. This PR converts all stateless tests that do such manipulation into integration tests, grouped by scenario, and adds a style check that rejects new ones: **Extended existing modules:** - `test_broken_projections` ← `02254_projection_broken_part`, `04511_check_table_broken_projection_columns` - `test_broken_tmp_txn_version_startup` ← `04104_transaction_version_metadata_dummy_tid_load`, `04492_attach_as_replicated_clears_tmp_txn_version`, `04493_attach_part_clears_tmp_txn_version`, `04507_packed_part_stale_txn_version_guard` - `test_lost_part` ← `02369_lost_part_intersecting_merges`, `02370_lost_part_intersecting_merges`, `04215_replicated_missing_covered_part_on_start` **New modules:** - `test_corrupted_part_files` (files inside part directories damaged/removed; loading, fetching, and `CHECK TABLE` reactions) ← `02253_empty_part_checksums`, `02255_broken_parts_chain_on_start`, `02444_async_broken_outdated_part_loading`, `04151_unique_key_sst_rebuild_on_load`, `04235_corrupted_columns_substreams_detection`, `04323_text_index_marks_empty_part`, `04506_packed_part_fetch_checksum` - `test_mutations_with_tampered_parts` (mutations over parts with old-version emulation or corrupted index files) ← `04401_dynamic_mutation_old_part_no_substreams_file`, `04412_mutation_old_part_basic_map_partial`, `04426_mutate_repair_corrupted_missing_idx_checksums`, `04427_mutate_some_columns_drop_index_corrupted_idx`, `04428_mutate_corrupted_text_index_multistream`, `04431_mutate_corrupted_index_sibling_owns_file` - `test_attach_tampered_detached_parts` (detached part directories manipulated before `ATTACH` / `DROP DETACHED`) ← `04063_drop_detached_part_with_try_n_suffix`, `04246_materialize_index_force_recalc` (the canned `part_25.8.tar.gz` moves into the module), `04402_mutate_all_columns_preserve_legacy_idx_minmax`, `04403_mutate_preserve_legacy_idx_packed_minmax`, `04404_mutate_rebuild_legacy_idx_minmax`, `04425_mutate_mixed_legacy_idx_minmax` - `test_backups_from_disk` (backups placed or tampered with directly on the `backups` disk) ← `04054_backup_restore_validate_entry_paths`, `04495_backup_metadata_version_overflow`, `02864_restore_table_with_broken_part`, `03001_restore_from_old_backup_with_matview_inner_table_metadata`, `03214_backup_and_clear_old_temporary_directories`, `03231_old_backup_without_access_entities_dependents`. The canned backup zips move from `tests/queries/0_stateless/backups/` into the module, and `helpers/install_predefined_backup.sh` is replaced by a module helper. - `test_server_metadata_files` (direct manipulation of on-disk table metadata `.sql` and the `flags/` directory) ← `03001_matview_columns_after_modify_query`, `04329_create_or_replace_force_drop_flag`, `04545_attach_projection_part_offset_setting_disabled` The conversions are faithful: every `.reference` output became an exact assertion, the originals' explanatory comments are preserved, and `DETACH`/`ATTACH` stayed `DETACH`/`ATTACH` except where the original explicitly emulated a server restart (`02255`, `04215` — \"on start\" scenarios now use a real server restart). Notable finds along the way: - `02444_async_broken_outdated_part_loading` contained a quoting bug since its introduction: `rm -f \"$path/*.bin\"` never expanded the glob, so the \"broken\" outdated part was never actually broken. The converted test applies the corruption for real, and the assertions still hold. - `02255_broken_parts_chain_on_start` dropped a nonexistent table (`projection_broken_parts_1`) in cleanup instead of its own tables. Verified locally against a fresh master binary: all new tests pass under the integration runner, and the tampered-parts modules were additionally cross-checked against a pre-fix binary where the bug-detecting assertions fire as intended. The style check (`server_data_manipulation_in_stateless_tests` in `ci/jobs/check_style.py`, folded in from https://github.com/ClickHouse/ClickHouse/pull/114056) flags stateless `.sh` tests that fetch a server-side filesystem path from a system table (`system.parts`, `system.detached_parts`, `system.projection_parts`, `system.tables`, `system.disks`, `system.server_settings` with `name = 'path'`) and also run file-modifying shell commands (`rm`/`cp`/`mv`/`dd`/`truncate`/`ln`/`chmod`/`touch`/`mkdir`/`tar`, `sed -i`, output redirection into a variable-derived path, `clickhouse-disks write`/`remove`/...). With all offenders converted here, the exclusion list holds only one documented false positive (`04326_disks_app_read_checksums`, which writes an `mktemp` scratch file under `CLICKHOUSE_TMP`) and `04630_merge_over_stale_packed_tmp_dir`, whose conversion is pending in https://github.com/ClickHouse/ClickHouse/pull/114057. Verified: the check is clean on this branch and catches each converted test if its stateless original is restored. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114070",
        "createdAt": "2026-08-09T20:47:29Z",
        "updatedAt": "2026-08-13T02:15:28Z",
        "timestamp": "2026-08-13T02:15:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114073",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `ORDER BY ... LIMIT` returning too few rows under a row policy",
        "text": "The top-K `ORDER BY <column> LIMIT n` optimization decides whether a query is filtered by looking at the plan-visible filters only: `where_clause = filter_step || getPrewhereInfo()`. A row policy is a reader-side filter that is not visible there, so a query filtered only by a policy took the unfiltered fast path: `MergeTreeDataSelectExecutor` enabled `perform_top_k_optimization` and narrowed the read to the marks holding the smallest values of the sort key, and the row policy then discarded all rows in those marks. The query returned fewer rows than the `LIMIT` - possibly none - even though later marks hold rows the policy keeps. ```sql CREATE TABLE t (key UInt64, INDEX mm_key key TYPE minmax GRANULARITY 1) ENGINE = MergeTree ORDER BY tuple() SETTINGS index_granularity = 8; INSERT INTO t SELECT number FROM numbers(300); CREATE ROW POLICY rp ON t FOR SELECT USING key >= 100 TO ALL; SELECT key FROM t ORDER BY key LIMIT 3; -- returned nothing, expected 100, 101, 102 SELECT count() FROM t; -- 200, correct ``` The optimization now counts `getRowLevelFilter()` as a filter, exactly like a visible `WHERE` or `PREWHERE`, so such a query takes the filtered path. The wrong result is reproducible with default settings since 25.12, when `use_skip_indexes_for_top_k` was enabled by default. Related: https://github.com/ClickHouse/ClickHouse/pull/110188 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `ORDER BY ... LIMIT` returning fewer rows than requested, possibly none, when a row policy was the only filter of the query and the sort column had a `minmax` skip index. The top-K optimization narrowed the read before the row policy was applied. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114073",
        "createdAt": "2026-08-09T21:19:31Z",
        "updatedAt": "2026-08-13T15:43:48Z",
        "timestamp": "2026-08-13T15:43:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114074",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix vector search with `arrayJoin` below the sort and with a row policy",
        "text": "Fixes for the vector search `ORDER BY <distance> LIMIT` optimization. **`arrayJoin` below the sort.** `arrayJoin` changes the number of rows - an empty array drops its base row - so the rewrite must not shortlist base rows before it runs. `optimizeTopK` rejects `arrayJoin` for this reason (https://github.com/ClickHouse/ClickHouse/issues/82279), but the vector search optimization did not, so the rows that a later base row would have contributed were never produced: ```sql CREATE TABLE t (id UInt32, tags Array(UInt32), vec Array(Float32), INDEX idx vec TYPE vector_similarity('hnsw', 'cosineDistance', 2)) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 4; INSERT INTO t SELECT number, if(number < 16, [], [number]), [toFloat32(number), toFloat32(number + 1)] FROM numbers(64); SELECT arrayJoin(tags) FROM t ORDER BY cosineDistance(vec, [0., 1.]) LIMIT 1; -- returned nothing; with `use_skip_indexes = 0` it returns 16 ``` **A row policy is an additional filter.** A row policy restricts rows inside the reader just like a `WHERE` or a `PREWHERE`, but it did not participate in `additional_filters_present`, so a query filtered only by a policy was treated as unfiltered: the `vector_search_filter_strategy = 'prefilter'` bailout was skipped, so an explicit request for exact search was ignored, and the index fetched only `LIMIT` neighbours without the `vector_search_index_fetch_multiplier` compensation, which the policy can then discard. **Filters that read the vector column broke the non-rescoring rewrite** (found during review). The rewrite drops the physical vector column from the read list and replaces it with the virtual `_distance` column, so any filter that still needs the column was left without its input and the query failed with `NOT_FOUND_COLUMN_IN_BLOCK` (an exception, pre-existing on master). The rewrite is now skipped - the same treatment as the vector column in `SELECT` - when the column is read by: - a row policy (including a policy carried in the deferred filter under `FINAL` with `apply_row_policy_after_final`), - a plain `WHERE` filter below the sort, e.g. `WHERE length(vec) > 0` (reachable with default settings, because the implicit `PREWHERE` optimization is disabled for vector-search candidates). The quantized-codes rewrite (`useVectorSearchWithQuantizedCodes`) is not affected by these problems: it splices the shortlist above the whole `Expression`/`Filter` chain, so the expansion and the filters run below it, and its reader-side filters prefilter the approximate ranking. This was checked with the same queries. `FINAL` with an explicit `PREWHERE` is covered by a separate regression test. `FINAL` can defer a `PREWHERE` until after the merge, but the deferral copies the filter and clears `query_info.prewhere_info` only when the pipeline is built, i.e. past every query plan optimization, so the existing `getPrewhereInfo` bailout still fires and the physical vector column survives for the deferred filter. Related: https://github.com/ClickHouse/ClickHouse/issues/82279 Related: https://github.com/ClickHouse/ClickHouse/pull/110188 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed vector search returning too few rows for a query with `arrayJoin` below `ORDER BY <distance> LIMIT`. A row policy now counts as an additional filter for vector search, so `vector_search_filter_strategy = 'prefilter'` is honored and the index fetch multiplier is applied for a query filtered by a row policy. Fixed `NOT_FOUND_COLUMN_IN_BLOCK` for a non-rescoring vector search query when a `WHERE` filter or a row policy reads the vector column. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114074",
        "createdAt": "2026-08-09T21:37:59Z",
        "updatedAt": "2026-08-13T11:23:41Z",
        "timestamp": "2026-08-13T11:23:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114083",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix local builds compiling vendored Rust code with `opt-level = 0`",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113575 Related: https://github.com/ClickHouse/ClickHouse/pull/110613 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix local (non-CI) builds compiling all vendored Rust code (the `PCO` codec, Delta Lake reads via `delta-kernel-rs`, `prqlc`, `chdig`, `wasmtime`, the client's fuzzy history search via `skim`) with `opt-level = 0`. The `SANITIZE` CMake variable was declared with `option`, so its default value was the BOOL `OFF`, which failed the `STREQUAL \"\"` \"no sanitizer\" checks, and every build configured without an explicit `-DSANITIZE=...` took the sanitizer workaround path that disables Rust optimization. ### Details `cmake/sanitize.cmake` declared `SANITIZE` with `option`, which creates a BOOL cache entry whose empty default is normalized to `OFF`. The string `OFF` is not equal to `\"\"`, so the `NOT SANITIZE STREQUAL \"\"` checks in `contrib/corrosion-cmake/CMakeLists.txt` took the \"sanitizer enabled\" branch in every default-configured build: the Cargo release profile was lowered to `opt-level = 0` and `codegen-units = 4` (a workaround for LLVM memory usage when sanitizer instrumentation is combined with `opt-level = 3` on huge crates), `--cfg rustix_use_libc` was applied, and the crate metadata was set to `with-sanitizer-OFF`. The build log even says so on every crate: `Finished 'release' profile [unoptimized] target(s)`. CI is unaffected by accident: `ci/jobs/build_clickhouse.py` always passes `-DSANITIZE=` explicitly, which creates a STRING cache entry with the empty value, so the comparison behaved as intended there. Only local builds that omit the flag — i.e. the default developer configuration — got unoptimized Rust. The practical impact is that any local performance measurement of a Rust-backed feature underestimated production performance by up to an order of magnitude, and Rust frames dominated local profiles that would not in production. This was found while benchmarking codecs for https://github.com/ClickHouse/ClickHouse/pull/113575: inserting 22 million `Float32` values with the `PCO` codec took 11.3 s in-engine, while the standalone `pcodec` CLI compressed the very same data in 0.88 s even at the same 16384-value chunking that ClickHouse blocks impose. The fix declares `SANITIZE` as a STRING cache variable (its legal values — `address`, `thread`, ... — are strings, so `option` was the wrong primitive) and tests it with CMake truthiness: `if (SANITIZE)` is false for `OFF`, `\"\"` and undefined, so build directories configured before this change, whose caches retain the stale `SANITIZE:BOOL=OFF` entry, produce optimized Rust as well. The effective Rust release profile is now printed at configure time to make future regressions visible. Verified by configuring a scratch build directory in three states: a fresh configure without the flag and a configure with a stale `SANITIZE:BOOL=OFF` cache entry both now generate an optimized release profile, `metadata=no-sanitizer`, and no `rustix_use_libc`; a `-DSANITIZE=address` configure keeps the `opt-level = 0` / `codegen-units = 4` workaround, `rustix_use_libc`, and `metadata=with-sanitizer-address` exactly as before.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114083",
        "timestamp": "2026-08-12T22:14:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-build",
          "submodule changed",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114085",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Delay background mutations by a bounded random amount in stress tests",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113925 Related: https://github.com/ClickHouse/ClickHouse/pull/113225 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Description A part written before an `ALTER` keeps its own older type (or misses the column) until the background mutation rewrites it. Reads over such parts are a distinct code path: columns resolve with the part's type while metadata and subcolumn entries carry the new one. Our tests pin their results with `mutations_sync = 2`, so the corpus systematically avoids this state, and in ordinary runs background mutations close the window in milliseconds. That is how the regression fixed by #113925 passed the full pre-merge CI of #113225: the bug lived only in the pending-mutation state, which nothing exercised until an unrelated test added the fixture two days later. This PR makes the stress test hold that window open routinely: - New failpoint `mutate_task_random_sleep_in_prepare`: a random 0-3 s sleep at the top of `MutateTask::prepare`, covering plain, replicated and shared `MergeTree` mutations (all funnel through `MutateTask`). - `stress.py` enables it via `SYSTEM ENABLE FAILPOINT` right after the smoke check (never in upgrade check, where the old binary may not know the failpoint) and disables it in `prepare_for_hung_check` next to `SYSTEM STOP THREAD FUZZER`, so pending mutations drain at full speed before the hung check. Disabling a registered failpoint that was never enabled is a documented no-op, so the cleanup is idempotent. With the failpoint on, every test that ALTERs without waiting reads unrewritten parts under the stress runner's random settings and the server-side AST fuzzer's query mutations - the exact combination that found the #113925 crash within four hours once a fixture existed, applied now to the whole corpus on every stress run. **Validation.** Built and measured on a `RelWithDebInfo` build: a `mutations_sync = 2` `UPDATE` takes 0.12 s with the failpoint off, 0.65-1.67 s with it on, and is fast again after `SYSTEM DISABLE FAILPOINT` (including a second, no-op disable). `stress.py` passes `python3 -m py_compile`. The delay is bounded, so `mutations_sync = 2` tests finish at most a few seconds later per mutation, and `prepare_for_hung_check` removes the delay entirely before the hung check. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1288` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114085",
        "createdAt": "2026-08-10T00:07:13Z",
        "updatedAt": "2026-08-13T00:19:34Z",
        "timestamp": "2026-08-13T00:19:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114087",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check column structure against the declared type in `collectOffsetsColumns`",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113925 Related: https://github.com/ClickHouse/ClickHouse/pull/113225 Related: https://github.com/ClickHouse/ClickHouse/issues/113891 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a column's declared type and its data diverge during a `MergeTree` read (mixed type provenance, e.g. after an `ALTER TABLE ... MODIFY COLUMN` whose mutation has not finished), the server now reports a clear exception naming the column and both structures, instead of an unchecked cast: debug and sanitizer builds previously aborted with a bare `Bad cast from type A to B` naming no column, and release builds walked mismatched memory silently. ### Description The type-directed `enumerateStreams` walk in `collectOffsetsColumns` (`src/Interpreters/inplaceBlockConversions.cpp`) pairs each available column's declared type with its data. When an entry's type was resolved from the table's metadata while the column was read from a data part with an older type - the mixed type provenance behind #113925 - the walk `assert_cast`s the column to the wrong class at the first diverging wrapper. In debug/sanitizer builds that is an abort whose message names two column classes and no column (the crash-classifier issue #113891, STID 4256-3fc1, took a cross-file hunt to attribute); in release builds `assert_cast` does not check at all, so the walk misbehaves on memory-unsafe reinterpretation. The new `columnMatchesTypeStructure` check runs right before the walk, only on the missing-columns path where the walk already happens. It descends exactly the levels whose serializations pair the declared type's structure with the column's: - `Array`, `Nullable`, `Tuple`, `Map` - the wrappers whose serializations `assert_cast` the column; - `Variant` - `SerializationVariant::enumerateStreams` walks every alternative declared by the type and pairs it with the column's variant of the same global discriminator, so the alternative list has to match in length and element-wise; - typed paths of `Object` - `SerializationObject::enumerateStreams` walks every typed path declared by the type and looks it up in the column, so the typed paths have to match by name and structure. Dynamic paths and shared data are taken from the column itself and need no check; - `ColumnReplicated` is unwrapped. `Dynamic` is checked by class only, because its `enumerateStreams` takes both the type and the column of its variant from the column itself (`column_dynamic->getVariantInfo().variant_type`), so the two cannot diverge. Every leaf the check does not know is accepted - leaf divergence, such as a part storing `UInt32` for a column widened to `UInt64`, is legitimate. Because the checked structure mirrors exactly what the `enumerateStreams` implementations themselves assert, the check cannot fire on any pairing that debug CI does not already abort on (or, for `Object`, fail with a raw typed-path lookup) - it only converts that failure into a diagnosable exception and closes the release-build hole. With the #113925 reproducer on a `RelWithDebInfo` build of master before that fix, the witness query fails with: ``` Code: 49. DB::Exception: Column `arr.n` is listed with type Array(Nullable(String)) among available columns, but its data has incompatible structure Array(size = 1, UInt64(size = 1), String(size = 2)). It is likely that a type resolved from the table's metadata was combined with a column read from a data part with an older type: (while reading from part .../all_1_1_0/ ...) ``` instead of silent unchecked-cast behavior. #113925 (which fixes the type selection) has since merged and is included here, so this check now guards the invariant against future regressions of the same family. **Validation.** Witness reproduced as above on a local `RelWithDebInfo` build. False-positive sweep: all 381 stateless tests matching `nested`/`subcolumn` run against that binary - 294 passed, 87 failed for documented bare-server environmental reasons (no Keeper, no clusters, no `protoc`, no `/var/lib/clickhouse`), and the check's message appears zero times in the whole run. After the `Variant`/`Object` extension, all 826 stateless `.sql` tests matching `variant`/`dynamic`/`json`/`nested`/`subcolumn`/`object` were run again - the check's message and `Bad cast` both appear zero times - plus a targeted check that reads old parts through newly added `Variant`, `JSON`, `Dynamic`, `Tuple`, `Map` and `Nested` columns and through unfinished `ALTER TABLE ... MODIFY COLUMN` mutations of `Variant` and `JSON`. The `Memory`-engine caller of `fillMissingColumns` always casts columns to the requested types first (`tryGetColumnFromBlock`), so it cannot trip the check either. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114087",
        "createdAt": "2026-08-10T00:15:02Z",
        "updatedAt": "2026-08-13T14:59:16Z",
        "timestamp": "2026-08-13T14:59:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114089",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `bloom_filter` index skipping granules for a `FixedString` constant",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112693 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `hasAny`, `hasAll`, `has`, `indexOf`, `mapContainsKey`, `mapContainsValue` and `mapContains` returning too few rows, or failing with `TOO_LARGE_STRING_SIZE`, when a `bloom_filter` index is queried with a `FixedString` constant. The index hashed the padded form of the constant while the function compares the unpadded one, so a matching granule was skipped. ### Description A `bloom_filter` skip index silently returns too few rows when the constant is a `FixedString`, and in four cases raises `TOO_LARGE_STRING_SIZE` on a query that succeeds without the index: ```sql CREATE TABLE k (id UInt64, v Array(String), INDEX idx v TYPE bloom_filter GRANULARITY 1) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 1; INSERT INTO k VALUES (0,['V0']),(1,['V0\\0']),(2,['X']); SELECT count() FROM k WHERE hasAny(v, [toFixedString('V0',3)]); -- 0, wrong SELECT count() FROM k WHERE hasAny(v, [toFixedString('V0',3)]) SETTINGS use_skip_indexes=0; -- 1 ``` Root cause: the array-search functions coerce the constant with CAST before comparing it to the elements (`hasAny`/`hasAll`, and `has`/`indexOf` over a `FixedString` element, cast both sides to the least supertype; `has`/`indexOf` over a `LowCardinality` element cast the constant straight to the dictionary type), and a CAST of `FixedString` to `String` strips the trailing zero padding. The index instead converted the constant with `convertFieldToType`, which keeps the padding, so it hashed a value the function never compares and pruned granules that match. The error cases are the same cause: an oversized `Field` passed through unchanged and reached a column `insert` during index analysis. The fix replicates the coercion at the `Field` level, which is exact because the string casts involved are fully predictable: strip the padding of a `FixedString` constant, re-pad it to the width of the element type (the stored form of every element the function can match), and decline the index when the constant has no stored representation, so a runtime `TOO_LARGE_STRING_SIZE` stays reachable instead of becoming a silent empty result. The direct cast to a dictionary type rejects an over-wide `FixedString` constant before stripping, while the supertype cast strips first - that one ordering difference is the only divergence between the two coercions, so it is a `bool` rather than a second code path. `has`/`indexOf` over a plain `String` element compare the constant's raw padded bytes (`executeString`) and keep the old conversion, as does `has(<constant array>, <indexed scalar>)`, whose runtime compares `Field`s directly. #### `Map` predicates `mapContainsKey`, `mapContainsValue`, `mapContains` and `has` over a `Map` are adapters of the same `arrayIndex.h` machinery, so a `bloom_filter` index on `mapKeys`/`mapValues` needs the same coercion - and it needs to pick the mode from the `Map` type rather than from the index header, because the `mapKeys`/`mapValues` index expression strips `LowCardinality` from the key/value type while the `mapContains*` adapters run over the keys/values subcolumn, which keeps it. Reading the mode from the stripped header alone would still prune a matching granule on `Map(LowCardinality(String), ...)`, which is broken on `master` today independently of the constant's width. `has` over a `Map` is the one exception, and it goes the other way: `FunctionArrayIndex::executeMap` rewrites the map to an array of its keys and calls `recursiveRemoveLowCardinality` on both arguments before `executeArrayImpl`, so it compares the raw padded bytes exactly like `has` over an `Array(String)`, unlike `mapContainsKey` on the same column. Its element type is therefore stripped of `LowCardinality` before the mode is read. The direct array spellings `has(mapKeys(m), ...)` / `indexOf(mapValues(m), ...)` need no special handling: `mapKeys`/`mapValues` return a full `Array` with `LowCardinality` stripped from the type, so the index header type is already the type the runtime compares against. This supersedes https://github.com/ClickHouse/ClickHouse/pull/112693 with a smaller implementation of the same semantics (one `Field`-level helper instead of a two-mode column-cast round-trip through `getLeastSupertype` plus a batched clone of `createColumnFromConstantArray`). The array test - 115 assertions over the full element-type/function/constant-width matrix, including granule-pruning and `TOO_LARGE_STRING_SIZE` reachability assertions - is taken unchanged from that pull request and passes byte-identically, so the two implementations are behaviorally equivalent on the covered surface. Three further tests cover the `Map` predicates, the direct `mapKeys`/`mapValues` array spellings, and `has` over a `Map` with `LowCardinality` keys; each compares the indexed answer against an unindexed oracle.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114089",
        "createdAt": "2026-08-10T00:59:11Z",
        "updatedAt": "2026-08-13T08:57:11Z",
        "timestamp": "2026-08-13T08:57:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114090",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add a regression test for duplicate parallel replicas announcements from a self-matching merge() child",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Add a regression test for the parallel replicas coordinator guard trips (`Initiator received more initial requests than there are replicas: replica_num=1`, `Duplicate announcement received for replica number N`) that fired when a `merge()` table function's regex matched the table the outer query reads. #112849 fixed this in code — the `enable_parallel_replicas` clear in `ReadFromMerge::createChildrenPlans` — but its test `04665_merge_table_parallel_replicas_child_plan` never reaches the announcement guards on the pre-fix code: its fixture aborts earlier with `Cannot serialize FutureSetFromSubquery with no query plan`, and it uses the default index granularity, at which the initiator's replica claims every mark range and the followers are cancelled before they announce. So if the `enable_parallel_replicas` clear were ever narrowed (the way the `make_distributed_plan` clear was narrowed to `queryHasSubquerySets`), the signature would regress with `04665` still green. The fixture is the one @groeneai verified against three binaries (see https://github.com/ClickHouse/ClickHouse/pull/102192#issuecomment-5232296945): the outer read and the `merge()` child read of the same table derive the identical `stream_id`, so without the clear one follower builds several read pools under one `replica_num` and announces more than once into the same coordinator. A small `index_granularity` gives the followers enough marks to survive to the announcement; the committed fixture uses 16384 rows with `index_granularity = 16` (~1024 marks). Heavier shapes (500000/128 and 62500/16, both ~3900 marks) exceeded the 180s per-test limit of the flaky check, which runs 50 copies concurrently on a debug build: `system.query_log` from the failed run shows the `parallel_replicas_local_plan = 0` SELECT stalling up to 333s wall at 12.5s CPU with 834 thread-seconds in `NetworkReceiveElapsedMicroseconds` and only 9 coordinator round-trips, followers idle after finishing their physical read in seconds — the round-trips of that mode degrade sharply on an oversubscribed server, so the mark count is kept as low as the reproduction allows (at ~512 marks the local-plan mode stops reproducing, so ~1024 keeps a 2x margin). Verified locally on a 3-replica server: reddens in both `parallel_replicas_local_plan` modes 3/3 runs with the two `enable_parallel_replicas` clears in `StorageMerge.cpp` disabled, passes with them in place, also under 8 concurrent runs. The test asserts the successful result rather than a guard message, since the pre-fix failure surfaces as either of the two adjacent guard messages depending on follower interleaving. To make sure it cannot silently stop exercising the announcement path, it runs with `allow_experimental_parallel_reading_from_replicas = 2` (an unsupported-shape fallback to a plain local read is an error, not a silent success) and additionally asserts `ProfileEvents['ParallelReplicasHandleRequestMicroseconds'] > 0` on the initiator's `system.query_log` entry, like `04545_parallel_replicas_projection_short_circuit_unknown_stream.sql` — a follower must survive past the announcement into the coordinator's request path for it to fire, so the initiator claiming every range and cancelling the followers early fails the test instead of passing it (a coordinator-creation log line alone could not distinguish that). The regex is anchored to the table name so concurrent tests cannot leak into the `merge()`, and the table lives in `default` because a single-argument `merge()` resolves against the default database on each hop. Because `default` is shared by every concurrently running test, the test is a `.sh` so the table name can carry `$CLICKHOUSE_DATABASE`: the first revision used a fixed literal name, and two runs of the test then raced on `CREATE`/`DROP` of the same table (`UNKNOWN_TABLE` / `TABLE_ALREADY_EXISTS`) - the flaky check runs each new test 50 times with `--jobs nproc-1`, and the ordinary parallel jobs repeat newly modified tests. Related: https://github.com/ClickHouse/ClickHouse/pull/112849 Related: https://github.com/ClickHouse/ClickHouse/pull/110972",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114090",
        "createdAt": "2026-08-10T01:00:49Z",
        "updatedAt": "2026-08-13T14:31:28Z",
        "timestamp": "2026-08-13T14:31:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114100",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Charge an empty match for the distance it may have scanned in countMatches",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113886 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114088 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `countMatches` and `countMatchesCaseInsensitive` ignoring `max_execution_time` and `KILL QUERY` for a pattern that matches empty but whose match is only found after scanning far into the value. ### Description Follow-up to #113886, which made the match loop observe cancellation but left the default `count_matches_stop_at_empty_match = 0` path under-charged. That branch charged one unit, the byte `pos` advances, while the search that produced the match can have scanned to the end of the value. The loop continues instead of breaking, so one such search happens per byte and the whole value stayed under a single unit against a 65536-unit check interval. An empty match arrives from two places and only one hides the distance. A trivial empty pattern reports the match at the position the search started from, having scanned nothing; re2 reports a zero-length whole match as `offset = std::string::npos`, so its distance is bounded only by the rest of the value. The charge is now that remaining distance, capped at a 64th of a check interval for re2 and nothing for a trivial pattern, decided once before the loop since it is fixed for the call. The cap is what makes the re2 side safe: charging the bare remainder makes every iteration exceed a whole interval, so a 60 MB `x*` value takes one check per byte, a 23 percent regression. Measured against `max_execution_time = 2`, pattern `\\b(?:[^ ]{8}|[^a]{9}|[^b]{11}|[^c]{13})*` over a run of spaces followed by `a`: | value | before | after | |---|---|---| | 200 KB | 28371 ms | 2026 ms | | 1 MB | 165029 ms | 2028 ms | | 4 MB | 686679 ms | 2048 ms | | `KILL QUERY ... SYNC` | 26324 ms | 185 ms | Throughput without a deadline is unchanged: 1e9 trivial empty-match positions run 9.14 s before this PR and 9.40 s here, and a re2 `x*` carrier is 30.1 to 30.4 s across both. Results are byte-identical on the #113886 correctness suite. <details> <summary>Why the cap is 64, and no test is added</summary> Sweep on the 60 MB `x*` value, whose baseline is 9.32 / 9.22 / 9.16 s, against the 200 KB carrier above at a 2 s deadline. `units_per_check` is shared by five files, so it is capped locally here rather than lowered globally. | iterations per interval | 60 MB `x*` | 200 KB carrier | |---|---|---| | 1 (the uncapped form) | 11.42 / 11.34 / 11.25 s | 2002 ms | | 4 | 9.79 / 9.77 / 9.64 s | 2000 ms | | 64 (this PR) | 9.23 / 9.17 / 9.23 s | 2003 ms | I wrote a case for `04822_count_matches_cancellation.sh` and confirmed it both directions: it passes here, and on the parent exactly one line flips to `empty rescan: OVERSHOT 28493 ms` while the other 17 stay green. #114061 has since merged and deleted that file, so no test is included: re-adding it conflicts with the deletion. The only observable this fix changes is elapsed time. Both binaries raise `TIMEOUT_EXCEEDED` with the identical message, under `timeout_overflow_mode = 'break'` too, and this path has no ProfileEvent, so any guard is the wall-clock shape #114061 removed as flaky. Tell me which shape you would accept and I will write it. The bug is reachable on a stable release: on 26.7.4.12 the 200 KB carrier above runs 51861 ms against `max_execution_time = 2`. CI report for this PR: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114100&sha=ace49715030202a86a29a99f19a2a2ba311f99c6&name_0=PR&name_1=Stateless%20tests%20%28amd_tsan%2C%20sequential%2C%202/2%29 </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114100",
        "timestamp": "2026-08-12T20:13:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114106",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add table schema to SQLInsert output",
        "text": "Adds the `output_format_sql_insert_include_table_schema` setting for the `SQLInsert` output format. When enabled, `SQLInsert` prepends a `CREATE TABLE` statement derived from the query result schema, making the generated SQL directly replayable into an empty database. The table definition uses the result column names and types with `MergeTree` and `ORDER BY tuple()`. Source-table metadata such as keys, default expressions, codecs, TTLs, and indexes is not preserved. The setting is disabled by default and cannot be enabled together with `output_format_sql_insert_use_replace`. The change includes embedded documentation and stateless coverage for replaying the generated SQL, complex types, identifiers, empty results, batching, and incompatible settings. Closes: https://github.com/ClickHouse/ClickHouse/issues/84736 ### Changelog category: - Improvement ### Changelog entry: The `SQLInsert` output format can now prepend a `CREATE TABLE` statement with result column names and types. Enable `output_format_sql_insert_include_table_schema` to produce a ClickHouse SQL script that recreates the table and loads the exported data.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114106",
        "createdAt": "2026-08-10T04:55:57Z",
        "updatedAt": "2026-08-13T04:31:01Z",
        "timestamp": "2026-08-13T04:31:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "mishok2503",
        "state": "open",
        "assignees": [
          "nikitamikhaylov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114107",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Count a rejected insert once per query in `RejectedInserts`",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Count a rejected `INSERT` once, not once per parallel sink stream, in the `RejectedInserts` profile event. Follow-up to the review of https://github.com/ClickHouse/ClickHouse/pull/114016 (unresolved review thread on `src/Storages/MergeTree/MergeTreeSink.cpp`). Since `max_insert_threads` started to apply to plain `INSERT`s (https://github.com/ClickHouse/ClickHouse/pull/109000), an insert fans out to `max_insert_threads` parallel sinks, and since https://github.com/ClickHouse/ClickHouse/pull/114016 every one of those sinks evaluates the \"too many parts\" check in its constructor, on the query thread, while the insert chain is being built. `delayInsertOrThrowIfNeeded` is not a pure predicate: it increments `ProfileEvents::RejectedInserts` before it throws. So one rejected `INSERT` recorded one rejected insert per sink stream — measured 16 for `max_insert_threads = 16` — even though there is a single rejected input block, while the event is documented as \"Number of times the INSERT of a block to a MergeTree table was rejected\". This routes the four throw paths of `delayInsertOrThrowIfNeeded` through the new `MergeTreeData::countRejectedInsert`, which bumps the event at most once per query and table. The \"already counted\" memo lives on the query's `QueryStatus` (`tryCountRejectedInsert`, a set keyed by table), so concurrent rejected inserts into the same table are each counted exactly once regardless of how their sink constructors interleave, a retry of a query gets a fresh `QueryStatus` and is still counted, and one query whose materialized views make several target tables reject (under `materialized_views_ignore_errors`) counts every rejecting table. When there is no process list element to identify the query by, every rejection is counted as before. Verification: the new test `04836_rejected_inserts_counted_once_per_insert` reads the rejected query's `ProfileEvents['RejectedInserts']` from the query log — it prints `16` without this change and `1` with it (verified as a negative control by disabling the flag check); it also runs four concurrent rejected inserts and checks each recorded exactly one event, and an insert into a table with two materialized views whose target tables are both at the throw threshold, under `materialized_views_ignore_errors`, checking the query recorded exactly two rejected inserts (one per rejecting table). `04826_parallel_insert_sinks_too_many_parts_self_race`, `02458_relax_too_many_parts`, `00940_max_parts_in_total`, `01603_insert_select_too_many_parts`, `02280_add_query_level_settings`, `04202_move_partition_to_table_too_many_parts` and `04492_max_insert_threads_auto_default` pass locally. Related: https://github.com/ClickHouse/ClickHouse/pull/114016 Related: https://github.com/ClickHouse/ClickHouse/pull/109000",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114107",
        "timestamp": "2026-08-12T21:27:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114117",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113654 to 26.3: Fix non-atomic Keeper Raft state persistence",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113654 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31365209776/job/93382009349)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114117",
        "createdAt": "2026-08-10T07:35:52Z",
        "updatedAt": "2026-08-13T08:02:47Z",
        "timestamp": "2026-08-13T08:02:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "do not test",
          "pr-cherrypick",
          "pr-critical-bugfix"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "antonio2368",
          "alexbakharew"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114118",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113654 to 26.5: Fix non-atomic Keeper Raft state persistence",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113654 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31365209776/job/93382009349)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114118",
        "createdAt": "2026-08-10T07:36:27Z",
        "updatedAt": "2026-08-13T08:02:49Z",
        "timestamp": "2026-08-13T08:02:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "do not test",
          "pr-cherrypick",
          "pr-critical-bugfix"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "antonio2368",
          "alexbakharew"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114119",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113654 to 26.6: Fix non-atomic Keeper Raft state persistence",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113654 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31365209776/job/93382009349)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114119",
        "createdAt": "2026-08-10T07:37:01Z",
        "updatedAt": "2026-08-13T08:02:50Z",
        "timestamp": "2026-08-13T08:02:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "do not test",
          "pr-cherrypick",
          "pr-critical-bugfix"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "antonio2368",
          "alexbakharew"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114131",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use a continuous primary-key range for whole-metric PromQL selectors of TimeSeries tables",
        "text": "A PromQL selector over a `TimeSeries` table filters the samples table with `id IN (SELECT id FROM tags WHERE <matchers>)`. For a metric with tens of thousands of series, `KeyCondition` runs its single-threaded generic exclusion search with the whole set: 284 ms per selector on a 62-billion-row part with 1.9M marks (503 ms at 8.1M marks), and rule-style queries evaluate up to 5 selectors. With the two-component id layout `Tuple(UInt64, UUID)` the canonical id generator derives the first component from the metric name alone, so all series of one metric form one continuous primary-key range. When a selector matches a whole metric — verified by metadata checks plus one `LIMIT 1` probe on the tags table, which also detects out-of-range ids left by an earlier `ALTER ... MODIFY SETTING id_generator` — the generated WHERE additionally carries `id >= tuple(hash(name), min) AND id <= tuple(hash(name), max)` and the inner query sets `use_index_for_in_with_subqueries_max_values = 1`. Index analysis uses the range; the `IN` stays for exact row filtering, so both emissions return identical rows on any data. Any failed check emits today's SQL unchanged. The range can select a few extra boundary granules (+5 of 135,005 marks on a 30-day scan). ## Measured effect tsbench PromQL suite: 62.455B samples / 361,432 series, 1.9M-mark part; Ryzen 9950X (16C/32T); baseline = clean master 9b6a2d7346f. Cold medians of 3 interleaved rounds: | query | master | this PR | delta | |---|--:|--:|--:| | s07 (30m range) | 2.22 s | 1.23 s | −44.8% | | r03 (rule, 3 selectors) | 5.02 s | 3.07 s | −38.8% | | s11 (24h range) | 4.30 s | 2.98 s | −30.8% | | r02 (25.6k-series instant) | 3.04 s | 2.17 s | −28.7% | | s06 (24h instant) | 5.04 s | 3.61 s | −28.4% | | full 24-query suite, cold geomean | 950 ms | 847 ms | **−10.8%** | Selectors that do not match a whole metric fall back and are unaffected (r05, s05: ±0.3%). Probe cost on non-firing selectors: ~2–4 ms each (r07: 66 → 73 ms); single-component id layouts never reach the probe. The removed cost grows with mark count, so the effect is larger at `index_granularity_bytes = 262144`. ## Tests `04836_time_series_selector_whole_metric_pk_range`: fires for whole-metric selectors (plan carries the range, `IN` retained), falls back byte-identically for label-filtered, regex, custom-generator, and ALTERed-`id_generator` history cases; both tuple layouts; full `prometheusQuery`/`prometheusQueryRange` results compared. `tests/performance/promql_selector_pk_range.xml`: 20,000-series metric, range path plus fallback control. Related: #113768 (open) — removes no-op casts in the same generated SELECT; complementary, each stands alone. --- ### Changelog category (leave one): - Performance Improvement ### Changelog entry: PromQL selectors that match all series of one metric now filter the samples table of a `TimeSeries` table with a continuous primary-key range on `id` during index analysis instead of a large `id IN <set>` condition, when the id layout is a two-component tuple with the canonical id generator. Removes the dominant single-threaded index-analysis cost of selector-heavy PromQL queries: up to −45% cold latency on dashboard and rule query shapes, −11% cold geomean over the full suite on a 62-billion-sample table. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114131",
        "createdAt": "2026-08-10T10:07:19Z",
        "updatedAt": "2026-08-13T15:55:30Z",
        "timestamp": "2026-08-13T15:55:30Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-performance",
          "comp-promql"
        ],
        "author": "nikitamikhaylov",
        "state": "open",
        "assignees": [
          "vitlibar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114150",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Disable the map element to subcolumn rewrite by default and revert the PREWHERE grouping change",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111879 Related: https://github.com/ClickHouse/ClickHouse/pull/99200 Related: https://github.com/ClickHouse/ClickHouse/pull/107988 Related: https://github.com/ClickHouse/ClickHouse/pull/111954 Related: https://github.com/ClickHouse/ClickHouse/issues/107912 ### Description In 26.3 we added bucketed serialization for the `Map` data type together with an analyzer rewrite of `m['key']` (internally `arrayElement(m, 'key')`) into the Map key subcolumn `m.key_<key>`, so that only the relevant bucket is read from disk. The bucketed serialization turned out to be imperfect and is planned to be reimplemented as a separate type. For regular (non-bucketed) `Map` serialization this rewrite can be harmful (for example, it complicates subcolumn size estimation) and gives no benefit for single-key lookups. **1. New setting.** `optimize_map_element_to_subcolumn` (default `false`) gates only the `{Map, arrayElement}` transformer in `FunctionToSubcolumnsPass`. All other `Map` subcolumn optimizations (`mapKeys`, `mapValues`, `length`, `empty`) keep following `optimize_functions_to_subcolumns`. The new setting takes effect only when `optimize_functions_to_subcolumns` is also enabled. This only disables the automatic rewrite; reading the `m.key_<key>` subcolumn explicitly still works, and query results are unchanged either way. **2. Revert of the PREWHERE grouping from #107988.** That change made `MergeTreeWhereOptimizer` and `tryBuildPrewhereSteps` group conditions by physical storage column instead of by the exact column set, so that several `m['kN']` subcolumn reads share one read step. Its only purpose was to compensate for the rewrite: with the rewrite disabled, all `m['kN']` conditions reference the same column `m` and the original exact-column-set grouping already places them into one step. The grouping also introduced a correctness regression, #111879: a guard condition and a throwing expression over subcolumns of the same column (for example `payload.longitude IS NOT NULL` and a `CAST(tuple(...), 'Point')` over `payload.latitude`) became adjacent and were merged into one read step. All conditions of a step are evaluated on the same unfiltered block, so the throwing expression ran on rows the guard had rejected. Measured with the repro from #111879 on pre-built binaries: | binary | result | |---|---| | master before #107988 (`cd3c529cd3a`) | `66666` | | master with #107988 | `CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN` | | this branch | `66666` | Reverting instead of adding a new heuristic keeps this change small enough to backport. Improving subcolumn read sharing in PREWHERE will be done separately on `master` and does not need backporting; #111954 is superseded by this change. The serialization part of #107988 (`ISerialization` cache keys, whole-map caching in `SerializationMapKeyValue`, the `SerializationSparse` refactor) is kept — it is unrelated to #111879 and still shares reads between subcolumns of the same column inside one read step. **Expected performance comparison results.** By default `m['key']` again reads the whole `Map` column, as before 26.3, so query shapes that read map elements are slower than on `master`, where the rewrite is enabled: `get_map_value` and the shapes in `tests/performance/map_subcolumns_prewhere.xml` (the setting opt-in was removed from that test so it measures the default path). This is the intended consequence of disabling the rewrite; enabling `optimize_map_element_to_subcolumn` restores the previous behavior. **Out of scope.** With `enable_multiple_prewhere_read_steps = 0` the whole PREWHERE is a single step, so no guard can protect a throwing expression and the repro still throws. This is pre-existing and unrelated to the rewrite: master before #107988 (`cd3c529cd3a`) already throws in that configuration, and it involves `JSON` subcolumns, not `Map`. The new test pins the setting to `1` because it is randomized in CI. **Tests.** `04815_prewhere_guard_throwing_expression_subcolumns` is a regression test for #111879. `04814_optimize_map_element_to_subcolumn_setting` covers the new setting and pins its value, because `clickhouse-test` randomizes it. The tests that assert the rewrite (`04000`, `04513`, `03989`, `04040`, `04207`) enable the setting explicitly. All 177 `prewhere`-named stateless tests were compared against a pre-revert binary and none changes behavior. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new setting `optimize_map_element_to_subcolumn` (disabled by default) that controls rewriting `m['key']` into the `Map` key subcolumn `m.key_<key>`. It is disabled by default because the bucketed `Map` serialization it relies on is being reworked. Also reverted the PREWHERE condition grouping that was introduced for that rewrite, which caused `CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN` when a guard condition and a potentially throwing expression, such as a `CAST`, read subcolumns of the same column, for example a `JSON` or `Map` column.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114150",
        "createdAt": "2026-08-10T12:25:51Z",
        "updatedAt": "2026-08-13T17:45:22Z",
        "timestamp": "2026-08-13T17:45:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "Avogar",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114152",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "[WIP]Add incremental rmv core",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add incremental refreshable materialized views managed via refresh_incremental, which when used each refresh reads only the data committed to the single source table since the previous run and persists the advanced cursor in the RMV's Keeper CoordinationZnode for at-least-once resumption. depends on https://github.com/ClickHouse/ClickHouse/pull/111794 cc @alesapin @Michicosun",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114152",
        "createdAt": "2026-08-10T12:45:59Z",
        "updatedAt": "2026-08-13T15:21:13Z",
        "timestamp": "2026-08-13T15:21:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-feature"
        ],
        "author": "SmitaRKulkarni",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114155",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Text index: fix out-of-bounds write in front-coding deserialization",
        "text": "Fixes https://github.com/ClickHouse/clickhouse-private/issues/65753. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reject corrupted/malicious dictionary where `lcp` exceeds the previous token or `lcp + data_size` overflows, before the buffer write. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1320` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114155",
        "createdAt": "2026-08-10T13:09:24Z",
        "updatedAt": "2026-08-13T12:39:10Z",
        "timestamp": "2026-08-13T12:39:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "ahmadov",
        "state": "closed",
        "assignees": [
          "CurtizJ"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114171",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Hash the AST members that `getTreeHash` did not see",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110833 This is the `getTreeHash` part of #110833, split off because it stands on its own: it fixes wrong results that have nothing to do with that pull request's motivation (comparing stored table definitions). Only the `src/Parsers` changes were taken; the `sameAST` helper and the metadata comparison stay in #110833. `IAST::getTreeHash` identifies an AST subtree, and several callers use it as the identity of an expression or of a query: `Context::executeTableFunction` caches the result of a table function under it, `ExecuteScalarSubqueriesVisitor` caches the value of a scalar subquery, `ActionsVisitor` keys prepared sets, `ComparisonGraph` and `WhereConstraintsOptimizer` decide whether two expressions are the same one. The default implementation hashes only `getID` and `children`, so anything a node keeps outside `children` is invisible to it, and a number of nodes keep meaningful state there. Two ASTs that mean different things then get the same hash, and a caller that treats the hash as an identity silently substitutes one for the other. For example, on `master`: ```sql SELECT (SELECT count() FROM view(SELECT 1 AS x UNION ALL SELECT 1)) AS union_all, (SELECT count() FROM view(SELECT 1 AS x UNION DISTINCT SELECT 1)) AS union_distinct ``` ``` ┌─union_all─┬─union_distinct─┐ 1. │ 2 │ 2 │ └───────────┴────────────────┘ ``` The second `view` returns `2` instead of `1`: `union_mode` is not a child of `ASTSelectWithUnionQuery` and was not hashed, so the two calls got the same cache key. The same happens for `INTERSECT` against `EXCEPT`, for `LIMIT 2` against `OFFSET 2` (the same literal in a different role), for `WITH FILL ... TO 5` against `WITH FILL ... STEP 5`, for a window frame bounded from below against one bounded from above, and for `APPLY(quantile(0.1))` against `APPLY(quantile(0.9))`. This change hashes the members that are part of an element's meaning and are not children: * `ASTSelectQuery`, `ASTProjectionSelectQuery`, `ASTOrderByElement` - the role each child plays (`WHERE` against `HAVING`, the limit against the offset, the `WITH FILL` upper bound against the step), which is recorded only in `positions`. `ASTSelectQuery::group_by_with_grouping_sets` was missing as well. * `ASTSelectWithUnionQuery` - `union_mode` and `list_of_modes`; `ASTSelectIntersectExceptQuery` - `final_operator`. * `ASTWindowDefinition` - the frame and the parent window name; `ASTWindowListElement` - the name the window is bound to. * `ASTWithElement` - the name the CTE is bound to, `MATERIALIZED`, and the column aliases. * `ASTQueryWithOutput` - the `INTO OUTFILE` modifiers; `ASTQueryWithTableAndOutput` - `TEMPORARY` and an explicit `UUID`. * `ASTSetQuery` - `is_standalone`, the settings reset with `= DEFAULT`, the query parameters, and the node's own identity, which this override skipped altogether. * `ASTTTLElement` - the mode, the destination, the `GROUP BY` key and assignments, and the recompression codec. * `ASTIndexDeclaration`, `ASTConstraintDeclaration`, `ASTProjectionDeclaration` - the declared name (and the index granularity, and `CHECK` against `ASSUME`). * `ASTCollation` - the collation name; `ASTStreamSettings` - the cursor tree and the watermark column and idle timeout. Follow-up fixes in the same area, found by the debug build's format+parse round-trip check and by review: * A per-column `PRIMARY KEY` is normalized by `ParserCreateQuery` into the storage definition, but `primary_key_specifier` stayed `true` on the column declarations while formatting never printed it, so once the flag was hashed, such a `CREATE` no longer round-tripped format+parse to the same tree hash and the debug build reported the `Inconsistent AST formatting` logical error. The parser now clears the flag when its meaning is transferred; where it does survive - `ALTER TABLE ... ADD/MODIFY COLUMN` - formatting now prints `PRIMARY KEY` instead of silently dropping it. * `ParserQueryWithOutput` canonicalizes the output-option children to the end of `children`, and most parsers add the database/table children first, but several `clone` implementations rebuilt them in a different order, so a clone of `CHECK TABLE t FORMAT JSONEachRow` or `SHOW CREATE TABLE t INTO OUTFILE 'x'` hashed differently than the original. The clones now rebuild `children` in the parser's order. Along the way: `ASTDropQuery::clone` and `ASTUndropQuery::clone` did not clear the copied `children` (the clone kept the source's children and appended the cloned ones on top), `ASTOptimizeQuery::clone` pushed `deduplicate_by_columns` into `children` while the parser keeps it member-only, and `ASTCheckTableQuery::partition`, `ASTWatchQuery::limit_length`, and the `where_expression` / `limit_length` of `SHOW COLUMNS` / `SHOW INDEXES` were left shared with the source instead of being cloned. * The `static_assert`s pinning the size of the AST nodes are checked only on 64-bit targets: the wasm32 parser build has a different layout. Two more fixes in the same area: * `ASTColumnsApplyTransformer` reached its non-child `parameters` and `lambda` through `updateTreeHashImpl` rather than `updateTreeHash`, which stops at the node itself and never descends into it, so the `0.5` of `APPLY(quantile(0.5))` was not hashed. * `ASTWithAlias` hashed the alias without its length, so the alias ran into whatever the node writes next and `foo` followed by `Identifier_bar` produced the same byte stream as the alias `bar` on `Identifier_foo`. `ASTCollation::readJSON` and `ASTWithElement::readJSON` put a member into `children` that neither the parser nor `clone` puts there, so an AST restored from JSON had a shape - and a hash - that no parsed AST has, and `clone` of it dropped the child again. They now reproduce the parser's shape. `ASTTTLElement::clone` left the recompression codec shared with the source, which the new unit test surfaced. New tests: `04836_tree_hash_ast_identity` covers the wrong results above, and `gtest_tree_hash_completeness` covers the members that no pair of queries can differ in on their own, by editing the JSON serialization of the AST, plus the JSON round-trip shape. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results for a query that calls the same table function twice with arguments that differ only in a part of the query that was not taken into account by the AST hash, such as `UNION ALL` against `UNION DISTINCT`, `INTERSECT` against `EXCEPT`, `LIMIT` against `OFFSET`, or the bounds of a window frame. The second call reused the result of the first one.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114171",
        "createdAt": "2026-08-10T15:12:26Z",
        "updatedAt": "2026-08-13T08:31:05Z",
        "timestamp": "2026-08-13T08:31:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114177",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a crash when a primary-key range layer produces an empty pipe",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed a server crash when reading a `MergeTree` table by primary-key range layers and one of the layers produced an empty pipe. `readByLayers` in `PartsSplitter` applies the per-layer border filter to every pipe returned by the reading step getter, and such a pipe can be empty. `applyRangeFilterFromAST` then calls `Pipe::getHeader`, which dereferences the pipe's null `header`. The ASan+UBSan build reports `reference binding to null pointer of type 'element_type' (aka 'const DB::Block')`, the TSan builds a segmentation fault. An empty pipe carries nothing to filter and is dropped by the consumers anyway (`Pipe::unitePipes` starts with `removeEmptyPipes`), so it is skipped now, exactly like a null filter AST. Found by the AST fuzzer in CI on two unrelated pull requests, so it is not caused by either of them: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=99639&sha=e3675c072cc6c7cb1184a060d9c872ac519c78a2&name_0=PR&name_1=Stress%20test%20%28amd_asan_ubsan%29 (see https://github.com/ClickHouse/ClickHouse/pull/99639) and https://github.com/ClickHouse/ClickHouse/pull/103182. There is no deterministic reproducer - the fuzzed query is not recoverable from the logs, because the server dies before it is logged - so no test is added. Closes: https://github.com/ClickHouse/ClickHouse/issues/114176",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114177",
        "createdAt": "2026-08-10T15:31:49Z",
        "updatedAt": "2026-08-13T10:47:24Z",
        "timestamp": "2026-08-13T10:47:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114178",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"NATS: add inline credentials setting\"",
        "text": "Reverts https://github.com/ClickHouse/ClickHouse/pull/110733 (merge commit `48ebef16af656664bab51cd5bb605a65e9b8ec9e`), which added the `nats_credentials` setting to the `NATS` table engine. Removed by this pull request: * the `nats_credentials` setting and its use in `NATSConnection` (`natsOptions_SetUserCredentialsFromMemory`), so the engine again takes credentials only as a path through `nats_credential_file`; * the mutual-exclusion check between `nats_credential_file` and `nats_credentials`, including the named-collection provenance handling in `registerStorageNATS`; * the runtime handling of `nats_credentials` only — the logging-only masking of the old spelling is deliberately kept. `nats_credentials` stays in both mask lists (`NATS::SETTINGS_TO_HIDE` and `FunctionSecretArgumentsFinder::nats_secret_keys`), because a query is formatted for logging before the storage settings are validated, so the old spelling must not leak the raw JWT/seed into `query_log` even though the server then rejects it. The masking branch (`findNATSTableEngineSecretArguments`) is kept as well: it also masks the still-supported secret keys (`nats_password`, `nats_token`, `nats_credential_file`) and the `nats_url` userinfo password in the table-engine argument form `ENGINE = NATS(collection, key = ...)`. The `ParserCreateQuery.MaskNATS*` unit tests are kept, and a new `MaskNATSTableEngineRemovedCredentialsSetting` test covers both spellings of the removed setting; * the test `04665_nats_credentials_named_collection`; * the `nats_credentials` lines from the `Documentation` block of `registerStorageNATS`. Two notes on the mechanics of the revert: * `src/Parsers/FunctionSecretArgumentsFinder.{h,cpp}` are resolved by hand: `findNATSTableEngineSecretArguments` and the full `nats_secret_keys` list (including `nats_credentials`, for masking only) are kept. Every other file comes out either byte-identical to its pre-`#110733` state or equal to it plus unrelated later changes (the message-broker schedule pool now returns a `shared_ptr`). * `docs/reference/engines/table-engines/integrations/nats.mdx` still mentions `nats_credentials`. That page is generated from the `Documentation` block this pull request updates, and direct edits of an `{/*AUTOGENERATED_START*/}` region are rejected by the docs check, so the page is left to the nightly documentation autogeneration. The setting was merged into `master` for the unreleased `26.8` and is not part of any release branch (the newest is `26.7`), so no changelog entry is needed. Related: https://github.com/ClickHouse/ClickHouse/pull/110733 Related: https://github.com/ClickHouse/ClickHouse/issues/85213 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1332` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114178",
        "createdAt": "2026-08-10T15:32:57Z",
        "updatedAt": "2026-08-13T14:32:50Z",
        "timestamp": "2026-08-13T14:32:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114182",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix top-K dynamic filtering for empty Tuple columns",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/113419 Related: https://github.com/ClickHouse/ClickHouse/pull/113406 `Tuple()` can be sorted, but the scalar comparison functions used by `__topKFilter` reject zero-sized tuples. This change skips top-K dynamic filtering when the sort column is directly `Tuple()`. A nested empty `Tuple`, including through `Nullable`, remains eligible for the optimization and is routed through the general column-comparison path. Composite types with their own comparison implementation, such as `Array(Tuple())`, are not misclassified and keep vectorized filtering. Local validation used the final SQL test blob against the existing debug binary: all 10 Praktika stateless runs passed, `check_cpp` passed, and the produced output matched the reference with an empty diff. These were local checks, not upstream CI. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `ORDER BY ... LIMIT` queries on `Tuple()` columns raising a `NOT_IMPLEMENTED` exception when top-K dynamic filtering is enabled. Direct empty tuples now skip the optimization, while supported nested composite types continue to use top-K filtering.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114182",
        "createdAt": "2026-08-10T15:50:01Z",
        "updatedAt": "2026-08-13T13:05:26Z",
        "timestamp": "2026-08-13T13:05:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "Boulea7",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114184",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Implementing generic block nested loop join",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Implemented generic block nested loop join",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114184",
        "createdAt": "2026-08-10T16:04:43Z",
        "updatedAt": "2026-08-13T18:00:39Z",
        "timestamp": "2026-08-13T18:00:39Z",
        "metrics": {
          "reactions": 2,
          "comments": 3
        },
        "labels": [
          "pr-feature"
        ],
        "author": "vdimir",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114187",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Rework the comments in the Web UI",
        "text": "Make the comments in `programs/server/play.html` concise and on point: state the constraint and the failure it prevents in simple terms, and drop meta-narrative and repeated details (the file shrinks from 13416 to 11810 lines). Stale comments are fixed along the way: the syntax highlighting no longer uses `mix-blend-mode: difference` (the textarea text is transparent over the backdrop), the active tab no longer erases the strip's border with a box-shadow (its `::after` does), and the roadmap listed already-implemented items. A few code simplifications found during the pass (verified with a comment-stripped diff to be the only code changes): - Define `@keyframes hourglass-animation` once and use it everywhere: the databases-panel hourglasses referenced it while only `tab-hourglass-animation` was defined, so they never animated. - Remove a redundant `position: relative` in `.tab.active` and a pointless local alias of `setQueryOrRun`. - Put visible characters directly into string literals (📌, 🌈, №, Σ, ↻, ⧗) instead of `\\u` escapes. The `test_play_reconcile_startup` harness (which runs the real script extracted from `play.html`) passes all scenarios. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Web UI: the hourglass loading indicator in the databases panel is animated again.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114187",
        "createdAt": "2026-08-10T16:50:15Z",
        "updatedAt": "2026-08-13T10:47:52Z",
        "timestamp": "2026-08-13T10:47:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114188",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Ignore redundant parentheses in stored table definitions",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/110833 Related: https://github.com/ClickHouse/ClickHouse/pull/92340 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `ATTACH PARTITION FROM`, `REPLACE PARTITION`, `MOVE PARTITION TO TABLE` and adding a `ReplicatedMergeTree` replica failing with `Tables have different ...`, `METADATA_MISMATCH` or `INCOMPATIBLE_COLUMNS` for tables whose definitions were written with redundant parentheses, such as `PARTITION BY (a)`, `ORDER BY (b)`, `INDEX ix (b * c) TYPE minmax`, `PROJECTION p (SELECT (b) ...)`, `CONSTRAINT c CHECK (a > 0)`, `TTL (d + INTERVAL 10 YEAR)` or `DEFAULT (a + 1)`. ### Description This is an alternative to https://github.com/ClickHouse/ClickHouse/pull/110833 that contains only the parentheses normalization, so that it is easy to backport. It does not touch `getTreeHash` and does not change how definitions are compared - they are still compared as text. The definition expressions of a table (keys, `TTL`, indices, projections, constraints, column defaults) are stored in ZooKeeper as text and compared as text. #92340 made the AST remember whether the user wrote redundant parentheses around an expression, so the same definition now has two spellings, and every comparison of a stored definition started to reject one against the other: ```sql CREATE TABLE dst (x UInt64, y UInt64) ENGINE = MergeTree PARTITION BY x ORDER BY y; CREATE TABLE src (x UInt64, y UInt64) ENGINE = MergeTree PARTITION BY (x) ORDER BY y; INSERT INTO src VALUES (1, 1); ALTER TABLE dst ATTACH PARTITION 1 FROM src; -- Code: 36. DB::Exception: Tables have different partition key. (BAD_ARGUMENTS) ``` An upgrade or a mixed-version cluster is not needed: both tables above are created by the same binary. Measured on released binaries, 25.8, 26.3 and 26.4 accept this and 26.5, 26.6 and 26.7 reject it, which matches the branches that contain `b38a892dfbfab6`. The fix is a single formatting mode. `IAST::FormatSettings::ignore_redundant_parentheses` suppresses the parentheses of the `parenthesized` flag - the one place that emits them, `decideParensEmission`, is skipped - and `IAST::formatIgnoringRedundantParentheses` is the one-line formatter built on it. Because the formatter visits everything it prints, this covers nested parentheses (`PARTITION BY (x + (1))`) and every kind of definition, not just the top level of a key. It is used where a stored definition is serialized or compared: - `ReplicatedMergeTreeTableMetadata::formattedAST` / `formattedASTNormalized` (the ZooKeeper `/metadata` payload and the replica-join comparison). This replaces the partial `stripArtificialParens` workaround added by #92340, which only cleared the flag on an expression list and its immediate children, and it no longer has to clone the AST. - `IndicesDescription::explicitToString` / `allToString`, `ProjectionsDescription::toString`, `ConstraintsDescription::toString` and `ColumnDescription::writeText` - the serializers of the `indices`, `projections`, `constraints` and `columns` parts of the replicated metadata. `ColumnsDescription::operator==` compares two sets of columns through the same serializer. - `MergeTreeData::checkStructureAndGetMergeTreeData` - the `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` structure gate (sorting/partition/primary keys, secondary indices, projections). - `StorageReplicatedMergeTree::alter` - the definitions an `ALTER` writes back into Keeper, so an `ALTER` never publishes a parenthesized definition either. - `AlterCommand::isTTLAlter` (whether restating a `TTL` schedules a `MATERIALIZE TTL` mutation) and the `MODIFY ORDER BY` no-op detection in `AlterCommands::prepare`. The text these produce is exactly what every server produced before #92340, so a new replica and an old one agree in both directions. Nothing else changes: the parser and the formatter are untouched, the table metadata keeps what the user wrote, and `SHOW CREATE` still prints `PARTITION BY (x + (1))`. Definitions that differ in more than the parentheses are still rejected (`a` vs `b`, a different index expression, a different column default). The `columns` payload of `StorageKeeperMap` and `ObjectStorageQueue` also contains these expressions, but both storages now compare it structurally, so a table created by any version keeps working. `StorageKeeperMap` already parses both sides and re-serializes them with the current serializer before comparing. `ObjectStorageQueueTableMetadata` compared the stored string verbatim; it now falls back to comparing the parsed `ColumnsDescription` when the strings differ, so a queue table created by 26.5-26.7 with redundant parentheses in a column `DEFAULT`, `CODEC` or `TTL` is accepted in both directions, and this also un-breaks such a table created by 25.8 and earlier (which is broken on 26.5-26.7 today). Retrying the creation of a replica that failed between the ZooKeeper transaction and saving the local metadata is also covered: `createReplicaAttempt` recognizes an already-created empty replica by comparing its `/metadata` and `/columns`, and now falls back to a structural comparison (`ReplicatedMergeTreeTableMetadata::checkEquals`, parsed `ColumnsDescription`) when the raw strings differ, so such a retry works across the old and new spellings instead of failing with `REPLICA_ALREADY_EXISTS`. Tests: `04836_parenthesized_definitions_attach_partition_from`, `04837_parenthesized_definitions_replicated_metadata` and `04850_parenthesized_definitions_replica_recovery`. Each case in them fails on master with the error it is named after. The unit test `ObjectStorageQueueTableMetadata.ColumnsComparisonIgnoresRedundantParentheses` covers the stored-JSON compatibility of the queue metadata. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1335` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114188",
        "createdAt": "2026-08-10T16:56:26Z",
        "updatedAt": "2026-08-13T15:35:57Z",
        "timestamp": "2026-08-13T15:35:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "pr-must-backport",
          "pr-synced-to-cloud",
          "pr-must-backport-synced"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114204",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Apply optimize_functions_to_subcolumns in WHERE/PREWHERE on shards of distributed queries",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114174 Related: https://github.com/ClickHouse/ClickHouse/pull/109188 Related: https://github.com/ClickHouse/ClickHouse/pull/99200 Related: https://github.com/ClickHouse/ClickHouse/pull/107988 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Applies `optimize_functions_to_subcolumns` inside `WHERE`/`PREWHERE` on the shards of distributed queries, so that e.g. `map['key']` reads a single key subcolumn instead of the whole `Map` column. Previously the optimization never applied when querying through a `Distributed` table. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Description `FunctionToSubcolumnsPass` never applied through a `Distributed` table: - The initiator declines it: `StorageDistributed::supportsOptimizationToSubcolumns` returns `false` since #109188, deliberately — the initiator cannot see the shards' real schema (#108299) or their skipping indexes (#108874), so rewriting on the initiator is unsafe. - The shard never compensates, contrary to the assumption in #109188: shards run only the resolve passes (`runOnlyResolve` in `buildQueryTreeAndRunPasses`) because optimization passes could change the block header expected by the initiator. As a result, `SELECT count() FROM t_dist WHERE tags['k0'] != ''` read the whole `Map` on every shard (~20x more bytes than the direct query in the issue's repro; the `notEmpty(arr)` → `arr.size0` rewrite loss is ~136x). The fix applies the rewrite on the shard, restricted to `WHERE`/`PREWHERE`: - `FunctionToSubcolumnsPass` gets an `only_filter_clauses` mode that demotes all qualifying identifiers to the existing `filter_only` bucket, whose rewrites the second visitor already restricts to `WHERE`/`PREWHERE` (both the direct and the chained form, e.g. `json.a[1].b`). Rewrites there cannot change the header returned to the initiator. Rewrites in other clauses (projection, `GROUP BY`) are not applied on shards — they could leak into the intermediate-stage header. - `buildQueryTreeAndRunPasses` runs this mode after the resolve passes for queries executed on shards of distributed queries: remote shards (`query_kind == SECONDARY_QUERY`), the `prefer_localhost_replica` local plan and the serialized query plan path (`is_local_plan_for_distributed_query` / `build_logical_plan`). `CREATE VIEW` definitions and projection analysis (the other users of the resolve-only path) are unaffected — the former is excluded explicitly, the latter has no `WHERE` clause. Doing this on the shard side is what makes it safe where the initiator-side rewrite was not: the pass's existing index guard sees the real `MergeTree` metadata (so a column required by a skipping index is not rewritten — the #108874 scenario), and the subcolumn is resolved against the shard's actual schema (the #108299 scenario). Running the pass on shards exposed a pre-existing hole in the pass's index guard (caught by `02346_text_index_bug108874`): the guard collected primary key and secondary index columns from the resolution-time storage snapshot, which for a table in a database with `lazy_load_tables = 1` is the `StorageTableProxy` stub — columns only, no keys and no indices — so no column was index-protected and the rewrite could break index analysis. The bug pre-exists on master for a direct query on a freshly attached lazy table (verified on 26.8.1.1235: `INDEX_NOT_USED` under `force_data_skipping_indices`). The guard now reads the storage's current metadata, which the proxy forwards once the nested storage is materialized by query resolution; a new stateless test covers the direct-query case. The stateless test checks result correctness and header stability across the local-replica, remote and serialized-plan paths, and asserts via `system.query_log` that the `Distributed` query reads several times less data than with the optimization disabled, on par with the direct query. Before the fix the read-amplification assertions fail (`0` instead of `1`); the rest of the test passes, confirming behavior is otherwise unchanged. Also extends `tests/performance/map_subcolumns_prewhere.xml` with a `remote()` case.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114204",
        "timestamp": "2026-08-12T23:25:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "andyzzhao",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114210",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use higher quality hash for nullable fixed-width keys in external aggregation",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): 64-bit hash function is used now for external nullable fixed-width aggregation methods (avoids collision disaster for very high cardinality aggregation). --- External aggregation merges partitions whose combined key count can far exceed 4 billion, which is why `mergeBlocks` re-aggregates the spilled stream with a method hashing over more than 32 bits rather than `HashCRC32`. `key64`, `keys128`, `keys256` and the serialized methods all have such a counterpart and are remapped to it; the nullable fixed-width methods never got one, so a `GROUP BY` on a nullable key that spills is still re-aggregated with a 32-bit hash and takes the collision blowup the remap exists to avoid. Add `nullable_key64_hash64`, `nullable_keys128_hash64` and `nullable_keys256_hash64`, and remap to them. The packed forms need no new data type: a nullable packed key carries its null map inside the key, so they reuse `AggregatedDataWithKeys128Hash64` / `…256Hash64` with `has_nullable_keys`. Only the single nullable key needs one, mirroring how `nullable_key64` wraps the `HashCRC32` map in `AggregationDataWithNullKey`. `nullable_key32` is deliberately left out, matching the existing choice not to remap `key32`: a 32-bit key space cannot exceed 4 billion distinct values.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114210",
        "createdAt": "2026-08-10T19:33:14Z",
        "updatedAt": "2026-08-13T11:45:24Z",
        "timestamp": "2026-08-13T11:45:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "nickitat",
        "state": "open",
        "assignees": [
          "nihalzp"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114212",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Explain the required column order when a dictionary QUERY returns columns in the wrong order",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a dictionary with a PostgreSQL source and a custom `QUERY` returns its columns in an order that does not match the dictionary structure, the error now names the destination dictionary column and the column order the query must return, instead of only reporting that a value could not be parsed. ### Description A `RANGE_HASHED` dictionary over `SOURCE(POSTGRESQL(... QUERY '...'))` failed with `Cannot parse LocalDate: 97955`, where `97955` was an unrelated `id` column's value: the DDL was valid, only the column order was wrong, and the message pointed at the data. A dictionary's expected source structure is keys-first: key(s), `RANGE MIN`, `RANGE MAX`, then the attributes. ClickHouse aliases every column when it builds the `SELECT` itself, but a user `QUERY` passes through verbatim and the result is read strictly by position, so `id` was deserialized into the `Date` attribute `contract_time`. On a conversion failure `PostgreSQLSource` now names the destination column and its position, and for a custom `QUERY` states the required order, derived from the dictionary's own structure so it suits every layout (the `RANGE` wording appears only for range dictionaries). Successful loads are untouched, and `BAD_ARGUMENTS` is preserved because `MaterializedPostgreSQL` relies on it. <details><summary>The message for the reported case</summary> ``` Cannot parse PostgreSQL value '97955' as Date: Cannot parse LocalDate: 97955: while reading column 2 of the result into `contract_time`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `meter_no`, `contract_time`, `end_date`, `id` (the key column(s) first, then the RANGE MIN and RANGE MAX columns, then the remaining attributes) ``` For a `FLAT` dictionary the same failure omits the `RANGE` clause: ``` Cannot parse PostgreSQL value 'abc' as Int64: Could not convert string to l: 'abc': while reading column 2 of the result into `num`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `id`, `num`, `name` ``` </details> A new integration test covers the range case, the documented order, and a `FLAT` dictionary asserting no `RANGE` wording appears; it fails on master. `04401_composite_key_dict_non_leading_key` gains the range shape over the ClickHouse source, pinning its by-name reordering against regression. Scope, since the mechanism is wider than the fix. The MySQL dictionary source (and likely XDBC and Cassandra) maps positionally too, but its parse errors do not funnel through one place, so folding it in grows the diff; say the word and I will extend it. Two cases stay untouched: an unnamed-column query still matches positionally and can silently return a wrong value, and a short row silently defaults trailing columns. Matching by name is possible, not ruled out: a `LIMIT 0` probe, or projecting the expected names inside the streamed query. I avoided it because it silently changes behaviour for every dictionary relying on positional matching, and arbitrary queries can return duplicate or unnamed columns.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114212",
        "createdAt": "2026-08-10T19:51:47Z",
        "updatedAt": "2026-08-13T17:39:53Z",
        "timestamp": "2026-08-13T17:39:53Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114216",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Wait for the RemovePart part_log row in 02950 and 02491",
        "text": "### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02950_part_log_bytes_uncompressed` and `02491_part_log_has_table_uuid` read `system.part_log` immediately after a DDL that removes a part and assert the `RemovePart` row is already there. That row is written asynchronously, so the assertion is unsound and the tests fail with just that line missing. `RemovePart` has a single emitter, `MergeTreeData::removePartsFinally`, which on both DDL paths here is reached only through `grabOldParts()`. Both DDLs (`DROP PART`, `TRUNCATE`) do run cleanup in the query thread, so usually the row lands first. But that grab selects nothing if it loses `try_lock` on `grab_old_parts_mutex`, or if the part's `shared_ptr` is not unique, e.g. while a concurrent read holds a reference. The part then stays `Outdated` and a later cleanup pass writes the row: late, not lost, so the engine is correct and only the tests need fixing. `01686_event_time_microseconds_part_log.sh` already asserts the same shape behind a bounded poll. Both tests become `.sh`, since a bounded poll is not expressible in `.sql`, and wait for the row under a 60 s bound. Each iteration issues `SYSTEM START CLEANUP <table>` to schedule a pass, because a cleanup thread that found nothing to do backs off up to `max_cleanup_delay_period`. Assertion queries, tags and both `.reference` files are unchanged; no `CREATE TABLE` setting is added here (02491's `old_parts_lifetime` pin and 02950's tags are pre-existing, kept verbatim); no `no-parallel` and no blanket `no-random-*`. Validation: holding a reference across the DDL reproduces the non-unique-ownership path deterministically. With that lever and no poll both tests fail 8/8; with the poll they pass 10/10. Dropping only the `SYSTEM START CLEANUP` line fails at the deadline once the cleanup thread is backed off. Deleting the DDL makes the poll time out rather than pass, so it is not satisfied by a stale row. Then 150/150 green per test across default, `-j 8` and randomized-order batches.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114216",
        "createdAt": "2026-08-10T20:03:37Z",
        "updatedAt": "2026-08-13T16:41:22Z",
        "timestamp": "2026-08-13T16:41:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114219",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113291 to 26.5: Fix for virtual row is not being applied in some cases",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114219",
        "createdAt": "2026-08-10T20:06:41Z",
        "updatedAt": "2026-08-13T17:15:21Z",
        "timestamp": "2026-08-13T17:15:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-2",
        "state": "open",
        "assignees": [
          "vdimir",
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114220",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113291 to 26.6: Fix for virtual row is not being applied in some cases",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114220",
        "createdAt": "2026-08-10T20:07:09Z",
        "updatedAt": "2026-08-13T17:54:03Z",
        "timestamp": "2026-08-13T17:54:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-2",
        "state": "open",
        "assignees": [
          "vdimir",
          "Avogar"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114226",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112601 to 26.6: Iterate ColumnObject subcolumns in sorted path order",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112601 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.32` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114226",
        "createdAt": "2026-08-10T20:10:17Z",
        "updatedAt": "2026-08-13T12:21:42Z",
        "timestamp": "2026-08-13T12:21:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-2",
        "state": "closed",
        "assignees": [
          "Avogar",
          "kssenii"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114247",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Clear plain LIMIT/OFFSET in the window-view backfill source query",
        "text": "Follow-up to https://github.com/ClickHouse/ClickHouse/pull/113759: the last AI review finding on that PR landed after the PR had been added to the merge queue, so the branch could no longer be updated. This addresses it. `StorageWindowView::getSourceTableSelectQuery` builds the raw-source backfill query for `CREATE WINDOW VIEW ... POPULATE`. The rows it produces are inserted into the window view, where `writeIntoWindowView` executes the mergeable view query over them, so the backfill query must deliver the raw source rows and leave all row transformations of the original query to the view query — otherwise the initialized state diverges from live behavior. The PR originally made the helper clear the leftovers of the original `SELECT` that violated this invariant one by one — `LIMIT`/`OFFSET` (and `WITH TIES`), plain `DISTINCT`, `ARRAY JOIN`, `WHERE`/`PREWHERE`, table-expression `SAMPLE`/`FINAL` — on top of the `JOIN`/`GROUP BY`/`ORDER BY`/`LIMIT BY`/`WINDOW`/`QUALIFY`/`INTERPOLATE` handling it already had. Review then found that this strip-the-leftovers approach misses wrapped sources: the same constructs inside a `FROM (SELECT ...)` subquery or a CTE definition survived the rewrite, and covering them would have required recursing the rewrite into every nested select. So the helper now builds the backfill query from scratch instead of stripping a clone of the view query. The contract makes this valid: `writeIntoWindowView` always receives raw source-table blocks (`getInputHeader` is the source table header no matter how the view query wraps or transforms the table — `PushingToWindowViewSink` is created with exactly that header), so the correct backfill query is exactly `SELECT <source columns> FROM <source table>`, plus `ORDER BY` on the timestamp column so the watermark is initialized from the earliest record. Wrapped sources, joins, and every row-shaping clause are covered by construction because the user query is no longer cloned at all. The helper shrinks by ~100 lines. Note: the divergence is currently unobservable because `CREATE WINDOW VIEW ... POPULATE` fails before writing any rows (https://github.com/ClickHouse/ClickHouse/issues/113493), so no regression test is possible yet; this keeps the rewrite invariant consistent for when `POPULATE` is fixed. Verified with `clickhouse-local` probes that `POPULATE` over wrapped-subquery/CTE/`JOIN`/`WHERE`/`FINAL` sources now analyzes the backfill query cleanly and proceeds to the pre-existing #113493 sink error. Related: https://github.com/ClickHouse/ClickHouse/pull/113759 Related: https://github.com/ClickHouse/ClickHouse/issues/113493 ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114247",
        "createdAt": "2026-08-10T23:40:02Z",
        "updatedAt": "2026-08-13T17:46:37Z",
        "timestamp": "2026-08-13T17:46:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114262",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Lazy materialization for reading local Parquet files (`file` / `File`)",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110970 Follow-up to the lazy materialization for object storage (#110970), as discussed in https://github.com/ClickHouse/ClickHouse/pull/110970#issuecomment-5247858385: implement it for plain local Parquet files read through `StorageFile` — the `file` table function and the `File` table engine. For `ORDER BY ... LIMIT n` queries, the columns that are not needed for sorting and filtering are read only for the `n` rows that survive the `LIMIT`. The format side (row-selective Parquet reads via `FormatFilterInfo::rows_to_read`) and the plan split (`JoinLazyColumnsStep`, `LazyMaterializingTransform`) from #110970 are reused as-is; this PR adds the `StorageFile` counterpart of the two branches: - The main pass appends a `__global_row_index` column (file index in a per-query `LazyFileRegistry` + physical row numbers from `ChunkInfoRowNumbers`) in `StorageFileSource::generate`. - The lazy branch (`LazilyReadFromFile` → `LazyReadFromFileSource` → `StorageFileLazyRowsSource`) reopens only the surviving files (at most `LIMIT n` of them) with a per-file set of rows to read and `parquet.preserve_order`. - The storage-agnostic column-split logic of `ReadFromObjectStorageStep::keepOnlyRequiredColumnsAndCreateLazyReadStep` (which columns can be deferred: `DEFAULT` expression dependencies, PREWHERE inputs, hive partition and virtual columns) is extracted into the shared `splitLazilyReadColumnsFromFormatInfo` and reused by both storages. **Generation safety.** POSIX has no conditional read, so the reread cannot be pinned the way `If-Match` pins it on S3. Instead the reread fails close with the new `FILE_CHANGED_DURING_READ` error when the file's generation token — sub-second mtime + inode + size, reusing `computeFileCacheVersionToken` from the query condition cache integration — no longer matches the one captured when the main pass opened the file. The token is validated at file registration time in the main pass, and both before and right after the reopen in the lazy pass. Replace-by-rename (the common atomic-update pattern) is always caught since it changes the inode; an in-place rewrite is caught up to the filesystem timestamp tick (and already tears a single-pass read today). **Gates** (`ReadFromFile::canUseLazyMaterialization`): Parquet format only, no file descriptor reads (stdin cannot be reopened), no archive entries, no `distributed_processing`, no `--rename_files_after_processing`, uncompressed files only. Controlled by the new setting `query_plan_optimize_lazy_materialization_for_file` (enabled by default), on top of `query_plan_optimize_lazy_materialization`. **Bug fix for #110970 shared via the helper:** deferring a requested subcolumn (e.g. of a `JSON` column) left the lazy branch's format header empty and the query failed with `Not found column or subcolumn ... in block`, because the format header contains the parent column of a requested subcolumn while the split filtered it by subcolumn names. The split now maps requested columns to their storage-level names — the parent of a deferred subcolumn moves to the lazy branch, and stays in the main branch as well when a sort key still needs another subcolumn of it. Covered for both storages by the new tests. On a 1 GB local Parquet file (2M rows × 105 columns equivalent shape: two 300-byte strings), `SELECT * ... ORDER BY k DESC LIMIT 5` runs 3× faster (0.15 s vs 0.45 s, page-cache warm), with results verified identical with the optimization on and off. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Lazy materialization for `ORDER BY ... LIMIT n` queries (#110970) now also applies to local Parquet files read with the `file` table function and the `File` table engine: the columns that are not needed for sorting and filtering are read only for the `n` rows that survive the `LIMIT`. The second read of a surviving file fails close with the new `FILE_CHANGED_DURING_READ` error if the file was modified between the two passes. Controlled by the new setting `query_plan_optimize_lazy_materialization_for_file` (enabled by default). Also fixes lazy materialization for object storage failing with `Not found column or subcolumn ... in block` when a requested subcolumn (e.g. of a `JSON` column) is deferred.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114262",
        "createdAt": "2026-08-11T03:20:36Z",
        "updatedAt": "2026-08-13T14:30:55Z",
        "timestamp": "2026-08-13T14:30:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114273",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add test: A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested",
        "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _This is a test-only PR — no source code changes. Please review: test quality, whether the claimed coverage gaps are real, and whether test output makes sense._ Adds test coverage for 1 untested code path, found during automated review of [PR #109744](https://github.com/ClickHouse/ClickHouse/pull/109744). That PR (1) The PR changes `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:403-421): the whole-column `input_format_null_as_default` short-circuit (`SerializationNullable::deserializeNullAsDefaultOrNestedTextCSV`) is now skipped for a bare, non-empty `Tuple` whose `tuple_ **1. A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested** `src/Processors/Formats/Impl/CSVRowInputFormat.cpp:413`, `src/DataTypes/Serializations/SerializationTuple.cpp:683` **Risk:** `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:409-413) now skips the whole-column `null_as_default` short-circuit for a bare `Tuple`, so the leading field is handed to `SerializationTuple::deserializeTextCSV`, whose per-element arm at src/DataTypes/Serializations/SerializationTuple.cpp:683 applies `null_as_default` to the FIRST ELEMENT — and when that element is itself a `Tuple`, one field stands for the whole nested element. This is exactly the sentence … **What's unique vs PR tests:** The PR's 04405_csv_tuple_leading_null_null_as_default.sql covers a leading `\\N` for a SCALAR first element and, for nested tuples, only `Tuple(Nullable(Int32), Tuple(Int32, Int32))` with `\\N,2,3` where the nested tuple's own fields are present; its comment states `a null inside a nested tuple is read as that whole nested element and is not covered`. This test covers precisely that: the `\\N` field is the field of a nested `Tuple` element (and of a `Nullable(Tuple)` element), the resulting one- … [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/e4ec131d-5685-4047-bd72-b61c347ba13b) cc @groeneai (author of #109744), @Avogar (merged/approved #109744) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114273",
        "createdAt": "2026-08-11T05:31:50Z",
        "updatedAt": "2026-08-13T17:59:33Z",
        "timestamp": "2026-08-13T17:59:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested"
        ],
        "author": "clickgapai",
        "state": "open",
        "assignees": [
          "PedroTadim"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114274",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Wait for replication before asserting a row count in test_rename_distributed",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `test_rename_column/test.py::test_rename_distributed` asserts an exact row count right after inserting into a `Distributed` table, with no barrier - only a 30 second poll in `select()`. Two things must finish first, and neither is awaited: 1. The INSERT is background-spooled: `insert_sync` at `StorageDistributed.cpp:1143` is false here (`distributed_foreground_insert` defaults false, table is persistent), so `DistributedSink` takes `writeAsync` and the INSERT returns while a block is still on the initiator's disk. 2. `internal_replication` is true, so a shard's peer replica must FETCH every part (~100 tiny parts per INSERT per shard given `PARTITION BY num % 100`), and `load_balancing` defaults to `RANDOM`, so a shard leg can be served by a replica still fetching. Found by inspection plus a local repro; public CI has not lost this race yet. Forcing the assert onto its first read (`poll=0`) fails reproducibly with `assert '1473\\n' == '1998\\n'`. There, `system.distribution_queue` held 2 un-flushed files while all four `system.replication_queue`s were empty, so a bare `SYSTEM SYNC REPLICA` returns vacuously. The spool drains first. The fix adds a helper running `SYSTEM FLUSH DISTRIBUTED` then `SYSTEM SYNC REPLICA` on all four nodes, after each insert. The first insert needs it too: the spool replays stored query text naming the columns, so a block outliving the `num2 -> foo2` rename replays `INSERT ... (num, num2)` and fails permanently with `NO_SUCH_COLUMN_IN_TABLE`, reproduced separately. `poll=30` is unchanged; raising it only lengthens the bet. Validated with the same `poll=0` probe in both arms on one binary, so the barrier does the work and not the poll: 2/2 fail without it, 3/3 pass with it. Then 50/50 green with `poll=30` restored, whole file green. The sibling `test_rename_distributed_parallel_insert_and_select` also reads a `Distributed` table, but its `select()` calls pass no `expected_result`, so they assert nothing about counts and are left alone. The other eight tests read plain `ReplicatedMergeTree`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114274",
        "createdAt": "2026-08-11T05:45:22Z",
        "updatedAt": "2026-08-13T13:33:32Z",
        "timestamp": "2026-08-13T13:33:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114275",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a Keeper server never joining the cluster when background snapshot IO is enabled",
        "text": "Related: https://github.com/ClickHouse/NuRaft/pull/123 ## Problem With `nuraft_use_bg_thread_for_snapshot_io` enabled, a Keeper server added to the cluster never joins. The leader's catch-up loop waits indefinitely for a response that cannot arrive, so the configuration change is never committed. The setting is off by default, which is why this has gone unnoticed. ## Root cause `sync_log_to_new_srv` builds one of two requests for the joining server — a snapshot request when the peer is below the retained log, or a log-pack `sync_log_request` otherwise — and then dispatches on the snapshot IO mode: ```cpp if (!params->use_bg_thread_for_snapshot_io_) { srv_to_join_->send_req(srv_to_join_, req, ex_resp_handler_); } else { snapshot_io_mgr::instance().invoke(); } ``` Only the *snapshot* branch hands its work to the IO thread. Nothing ever enqueues a log pack with `snapshot_io_mgr` — `create_sync_snapshot_req` is its only producer — so with background snapshot IO on, invoking the thread does not send the log-pack request. It is dropped. Observed on the unfixed build as the new server absent from the cluster configuration and stuck at log index 1 while the leader was at 12. ## Solution Dispatch on what was actually built rather than on the IO mode. The snapshot branch leaves the request null exactly when it has queued the work with the IO thread; in every other case — including the log-pack branch in *both* IO modes — there is a request and it must be sent. Covered by `log_sync_with_async_snapshot_io_test` in NuRaft's `learner_new_joiner_test`, verified to fail against the exact pre-patch code while the four existing tests in that binary keep passing. The submodule is pinned to the pull request branch so that the Keeper integration tests, `nightly_keeper` and Jepsen exercise it here. **The pointer must be moved to the merge commit before this pull request is merged.** ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a Keeper server never joining the cluster when `nuraft_use_bg_thread_for_snapshot_io` is enabled: the log batch sent to a joining server was silently dropped, so the membership change never completed.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114275",
        "createdAt": "2026-08-11T05:52:43Z",
        "updatedAt": "2026-08-13T01:24:03Z",
        "timestamp": "2026-08-13T01:24:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "submodule changed"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": [
          "antonio2368"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114278",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix `XDG_CACHE_HOME` being read from `XDG_STATE_HOME`",
        "text": "A copy-paste slip in `XDGBaseDirectories`: `ENV_XDG_CACHE_HOME` was defined as `\"XDG_STATE_HOME\"`, so `XDGBaseDirectories::getCacheHome` ignored the `XDG_CACHE_HOME` environment variable and obeyed `XDG_STATE_HOME` instead. `getCacheHome` currently has no in-tree callers, so this does not change observable behavior yet; it fixes the helper before callers appear. Found while reviewing https://github.com/ClickHouse/ClickHouse/pull/112824. Related: https://github.com/ClickHouse/ClickHouse/pull/112824 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114278",
        "createdAt": "2026-08-11T06:33:24Z",
        "updatedAt": "2026-08-13T14:31:43Z",
        "timestamp": "2026-08-13T14:31:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114283",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add pre-hook to insert CI links into PR body",
        "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) -- Adds a `ci_links.py` pre-hook to the `PR` workflow that, on upstream `ClickHouse/ClickHouse` pull request runs, appends a `:ci_links:` block to the PR description with: - a link to the workflow report, and - a link to a GitHub search for the corresponding sync PR (`sync-upstream/pr/<number>`). The block is added only when it is not already present, so subsequent runs do not re-edit the PR body. Non-upstream / non-PR runs are skipped, and any failure is caught so it can never break the workflow. <!-- CI automatic block start :ci_links: --> --- Workflow [[PR](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114283&sha=latest&name_0=PR)] Sync PR [[sync-upstream/pr/114283](https://github.com/search?q=head%3Async-upstream%2Fpr%2F114283+org%3AClickHouse+type%3Apr&type=pullrequests)] <!-- CI automatic block end :ci_links: -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114283",
        "createdAt": "2026-08-11T07:48:23Z",
        "updatedAt": "2026-08-13T17:58:00Z",
        "timestamp": "2026-08-13T17:58:00Z",
        "metrics": {
          "reactions": 1,
          "comments": 3
        },
        "labels": [
          "pr-ci"
        ],
        "author": "maxknv",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114285",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
        "text": "Files written first turn the namespace into a plain directory, which a catalog sharing the storage view (SeaweedFS) rejects with HTTP 500; the swallowed error left an orphaned metadata file that broke retries. Ensure the namespace before the first write; propagate failures except 404-then-create and 409 (REST) / AlreadyExists (Glue). Example: https://pastila.nl/?01f91a25/620869a81af815ba6927860efbc72af4#NhE2Nbzh3AHhUHhA5Fdkzw==GCM ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114285",
        "createdAt": "2026-08-11T08:20:50Z",
        "updatedAt": "2026-08-13T17:56:38Z",
        "timestamp": "2026-08-13T17:56:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "azat",
        "state": "open",
        "assignees": [
          "alesapin"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114298",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113509 to 26.6: Add a dedicated thread pool for lightweight snapshot creation",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113509 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114240 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31479296882/job/93740282532) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.34` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114298",
        "createdAt": "2026-08-11T10:05:22Z",
        "updatedAt": "2026-08-13T15:24:31Z",
        "timestamp": "2026-08-13T15:24:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "closed",
        "assignees": [
          "Diskein",
          "hanfei1991"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114300",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "TimeSeries: store all tags in the `tags` column",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): TimeSeries: store all tags in the `tags` column The `tags` column of the tags target table now contains all the tags, including the `__name__` tag with the metric name and the tags with dedicated columns from the `tags_to_columns` setting, so the map alone fully identifies a time series. The default id generator now hashes just `tags`, i.e. now it looks like `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(tags)))` (instead of `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(metric_name, all_tags)))` )",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114300",
        "createdAt": "2026-08-11T10:29:30Z",
        "updatedAt": "2026-08-13T17:41:11Z",
        "timestamp": "2026-08-13T17:41:11Z",
        "metrics": {
          "reactions": 1,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "comp-promql"
        ],
        "author": "vitlibar",
        "state": "closed",
        "assignees": [
          "nikitamikhaylov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114305",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add coverage tests for four untested MergeTree and index branches",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/101781 Related: https://github.com/ClickHouse/ClickHouse/pull/101750 Related: https://github.com/ClickHouse/ClickHouse/pull/103984 Related: https://github.com/ClickHouse/ClickHouse/pull/108884 Related: https://github.com/ClickHouse/ClickHouse/pull/108815 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Re-files four of the seven MergeTree/index coverage tests @ zlareb1 asked me to pick up on #101781. Tests only, no production code. Each pins a guard no existing test observes: - `04489_part_starting_offset_in_subquery` - a `_part_starting_offset + _part_offset` predicate written as `IN` (subquery and literal list) must still prune marks. Existing tests cover equality and ranges. (#101750) - `04489_virtual_row_disabled_with_final` - `SELECT ... FINAL` must not lose rows when the read-in-order virtual row is enabled. (#103984) - `04489_aggregation_in_order_limit_with_ties` - `LIMIT n WITH TIES` must not be pushed into aggregation-in-order, which would truncate the tie group. (#108884) - `04401_low_cardinality_fixed_string_partition_pruning` gains two rows for a `String` literal too long for a `LowCardinality(FixedString)` key: no pruning condition may be built, and the result stays correct. Extends my existing test instead of adding a near-duplicate file. (#108815) I adopted each test only after proving it observes its branch: inject a regression at the cited line, rebuild, require the test to fail, then revert and require it to pass. All four reddened, each with a positive and a negative control. `04489_virtual_row_disabled_with_final` initially did not: instrumentation showed the guard was never reached, because the outer `count()` eliminated the subquery's `ORDER BY` and the plan had no sorting step. I repaired the query shape so the read-in-order plan is built; it then reddens. Validation: 200 randomized runs and 200 with randomization off, no failures, also run 8-way concurrently. I am not filing the other three: #101724's oracle is satisfied by a type error raised upstream of the guard it claims to cover, so it would pass with that guard deleted; #103889's disk combination is already exercised by `tests/integration/test_attach_backup_from_s3_plain/test.py`; #103893 duplicates `03203_count_with_non_deterministic_function.sql`, which is the guard's own regression test and asserts more precisely. I am posting these reasons on those PRs. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1283` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114305",
        "timestamp": "2026-08-12T22:14:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114310",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113563 to 26.6: Do not analyze the shared row policy AST in place in `Merge`",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113563 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31486886135/job/93764217930)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114310",
        "createdAt": "2026-08-11T11:50:56Z",
        "updatedAt": "2026-08-13T04:26:26Z",
        "timestamp": "2026-08-13T04:26:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-backport",
          "pr-critical-bugfix"
        ],
        "author": "robot-clickhouse",
        "state": "open",
        "assignees": [
          "KochetovNicolai",
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114312",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix nested disk S3 I/O ignoring inner RESOURCE throttler at CachedObjectStorage delegation points",
        "text": "Fix nested disk S3 I/O ignoring inner RESOURCE throttler at DiskObjectStorage entry points ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fix nested disk RESOURCE bandwidth limiting: when a cache disk wraps an S3 disk, S3 I/O is now throttled by the inner disk's RESOURCE instead of being ignored. ### Description When a cache disk (e.g. `cached_s3`) wraps an S3 disk (e.g. `s3_inner`), creating a RESOURCE on the inner S3 disk had no effect — S3 I/O was never throttled. This happened because `DiskObjectStorage` always used the outer disk's resource name (`getReadResourceName()`), but the outer cache disk's local I/O never enters `IOSchedulingScope` (which only exists in `ReadBufferFromS3` and `WriteBufferFromS3`). For example, with this disk configuration: ```xml <s3_inner> <type>s3</type> <endpoint>http://minio:9001/root/data/</endpoint> </s3_inner> <cached_s3> <type>cache</type> <disk>s3_inner</disk> </cached_s3> ``` The user creates a RESOURCE on the inner S3 disk and a table on the cache policy: ```sql CREATE RESOURCE io_s3 (WRITE DISK s3_inner, READ DISK s3_inner); CREATE WORKLOAD all SETTINGS max_bytes_per_second = 10000000 FOR io_s3; CREATE TABLE t ... SETTINGS storage_policy = 'cached_s3'; ``` Before the fix, `INSERT INTO t ... SETTINGS workload = 'all'` would bypass `io_s3` entirely — `system.scheduler` showed zero `dequeued_cost` for the `io_s3` resource. After the fix, S3 I/O is correctly metered through the inner disk's RESOURCE. The fix pushes the inner (wrapped) disk's resource name at three `DiskObjectStorage` entry points: - `createObjectStorageTransaction` - `createObjectStorageTransactionToAnotherDisk` - `prepareRead` Each now uses `wrapped_disk->getReadResourceName()` when a wrapped disk is present, falling back to `getReadResourceName()` for non-nested disks. Only one source file is modified. Non-nested disks are unaffected — the ternary falls through to the same `getReadResourceName()` call as before.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114312",
        "createdAt": "2026-08-11T11:53:05Z",
        "updatedAt": "2026-08-13T03:33:33Z",
        "timestamp": "2026-08-13T03:33:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "Binnn-MX",
        "state": "open",
        "assignees": [
          "Michicosun"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114316",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Use the vector similarity index for integer reference vectors",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112233 Related: https://github.com/ClickHouse/ClickHouse/issues/114291 The reference vector of an ANN query is extracted only when its array type is `Float64`, `Float32` or `BFloat16` and every element is a `Float64` field. An integer literal such as `[1, 2]` is typed `Array(UInt8)`, so `tryUseVectorSearch` bails out and the query silently falls back to a brute-force scan over the whole table, although `[1, 2]` denotes the same point as `[1.0, 2.0]` and `L2Distance` accepts it. `EXPLAIN indexes = 1` shows no `vector_similarity` entry and gives no hint why, so a one-character difference in a literal becomes a sharp performance cliff on large tables. Native integer arrays are now accepted as reference vectors and their elements are converted to `Float64`, which is the type the reference vector is stored in anyway. Added `02354_vector_search_bug112233`, covering unsigned, signed, mixed integer/float, and not-exactly-representable reference vectors, plus an equality check between the integer and float spellings of the same query. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Vector search queries now use the `vector_similarity` index when the reference vector is written as an integer array literal, e.g. `ORDER BY L2Distance(vec, [1, 2])`. Previously such queries silently fell back to a brute-force scan.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114316",
        "createdAt": "2026-08-11T12:52:06Z",
        "updatedAt": "2026-08-13T17:48:13Z",
        "timestamp": "2026-08-13T17:48:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "hamidr",
        "state": "open",
        "assignees": [
          "rschu1ze"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114320",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "feat(clickpipes, docs): document clickpipe DDL DEFAULT propagation logic",
        "text": "- https://github.com/PeerDB-io/peerdb/pull/4635 - https://github.com/PeerDB-io/peerdb/pull/4632 ### Changelog category - Documentation (changelog entry is not required) ### Changelog entry Document `DEFAULT ...` DDL propagation for mysql and postgres ClickPipes. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1323` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114320",
        "createdAt": "2026-08-11T13:05:16Z",
        "updatedAt": "2026-08-13T13:33:41Z",
        "timestamp": "2026-08-13T13:33:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-documentation",
          "pr-synced-to-cloud"
        ],
        "author": "dtunikov",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114323",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: require canonical internal links",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114230 This is a one-off cleanup of repository-authored documentation links that use legacy redirect aliases. It updates the current English documentation and source-embedded reference documentation to use routes relative to the docs root, so the automated translation PR can parse and localize them without producing missing locale routes. This PR intentionally adds no permanent CI checks or ongoing enforcement. Its scope is limited to the current link corrections needed to get the automated translation PR parsing successfully. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114230&sha=e599855a281a4dd841be20794036908a58acbf63&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114323&sha=86d2d7e194d7fbf75cc84616a8fa0da1ed84802e&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114323",
        "createdAt": "2026-08-11T13:26:47Z",
        "updatedAt": "2026-08-13T17:58:20Z",
        "timestamp": "2026-08-13T17:58:20Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-ci",
          "pr-autogenerated-docs"
        ],
        "author": "Blargian",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114328",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Enhance MetadataStorageFromMemory",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Pure refactoring change, doesn't affect any working part of code.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114328",
        "createdAt": "2026-08-11T13:44:39Z",
        "updatedAt": "2026-08-13T17:44:07Z",
        "timestamp": "2026-08-13T17:44:07Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "comp-object-storage-disks"
        ],
        "author": "alesapin",
        "state": "open",
        "assignees": [
          "Michicosun"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114340",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Populate the submodule working trees in parallel in the build job",
        "text": "The build job populated the submodule working trees with a bare `git submodule update` on a submodule cache hit, which walks the submodules one at a time on a single core. In `Build (arm_tidy)` on master that step takes 39 s at ~3% CPU of a 32-core `m8g.8xlarge` — an otherwise idle machine waiting on one `git` process, right before the equally single-threaded cmake configuration and long before `ninja` can saturate the box. The first 253 s of that job produce no compilation at all; this is one of the three serial phases in it. Report the numbers come from: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=03ad5002ff1501ceb5d584b6d0bb028674b276e7&name_0=MasterCI&name_1=Build%20%28arm_tidy%29 Fan the checkouts out over the idle cores instead, the same way the cache-miss branch right below it (`contrib/update-submodules.sh --max-procs 10`) and `ci/jobs/fast_test.py` already do: feed the submodule paths from `.gitmodules` to `xargs --max-procs`. The submodule list is read with `get_output_or_raise` and asserted non-empty, so failing to enumerate `.gitmodules` fails the step rather than silently checking out nothing — `xargs --no-run-if-empty` would otherwise exit 0 on an empty pipeline. Measured on a synthetic superproject with 40 submodules and a populated `.git/modules` (the CI cache-hit state): **2.13 s serial vs 0.16 s** with `--max-procs=20`, both producing an identical working tree and a clean `git status`. In CI the win is bounded by the largest submodule rather than by the sum of all of them, so expect the step to shrink to roughly the cost of `contrib/llvm-project` alone; the `Checkout Submodules` duration in this PR's build reports against 39 s on master is the real measurement. ### Changelog category (leave one): - CI Fix or Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114340",
        "createdAt": "2026-08-11T14:18:28Z",
        "updatedAt": "2026-08-13T11:17:16Z",
        "timestamp": "2026-08-13T11:17:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "maxknv"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114373",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Enhancement for Vortex format",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> TODO ### Changelog category (leave one): - New Feature - Experimental Feature - Improvement - Performance Improvement - Backward Incompatible Change - Build/Testing/Packaging Improvement - Documentation (changelog entry is not required) - Critical Bug Fix (crash, data loss, RBAC) - Bug Fix (user-visible misbehavior in an official stable release) - CI Fix or Improvement (changelog entry is not required) - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114373",
        "timestamp": "2026-08-12T23:06:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [],
        "author": "m7kss1",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114380",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Clarify `aiSimilarity` cosine similarity range in its documentation",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110777 ### Changelog category (leave one): - Documentation (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1306` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114380",
        "createdAt": "2026-08-11T20:21:22Z",
        "updatedAt": "2026-08-13T03:54:15Z",
        "timestamp": "2026-08-13T03:54:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-documentation",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "davidmenggx",
        "state": "closed",
        "assignees": [
          "george-larionov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114383",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #102033 to 26.3: Fix LOGICAL_ERROR crash in IcebergMetadata::iterate when datalake_table_state is missing",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/102033 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31534495936/job/93922243414)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114383",
        "timestamp": "2026-08-12T21:09:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-2",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114392",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #111930 to 26.6: Fix Block structure mismatch when the split filter column name clashes with an input",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/111930 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114290 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31542213094/job/93946983899)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114392",
        "timestamp": "2026-08-12T20:04:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll4",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114394",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject a data lake schema whose column name is empty",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/114350 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114350 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a crash when reading a data lake table whose schema declares a column with an empty name. Malformed Iceberg metadata is rejected with `ICEBERG_SPECIFICATION_VIOLATION`; a schema supplied by a catalog, Delta Lake or Paimon is rejected with `AMBIGUOUS_COLUMN_NAME`. The check is unconditional, so an Iceberg table that merely retains an unused historical schema with an empty field name also becomes unreadable instead of aborting once that schema is read. ### Description Every lake reader copies the provider's field name into a `NamesAndTypesList` unvalidated. That becomes the table's column list, so an empty name reached `ASTIdentifier`, whose constructor asserts no identifier part is empty. `SELECT *` converts the query tree back to an AST during planning, so it aborted the server on debug and sanitizer builds, while `DESCRIBE` and `SELECT count()` succeeded and made the table look readable. In release the assert compiles out and the name surfaces one layer down. An empty column name is unrepresentable in ClickHouse, but nothing validated the point where an external lake schema becomes the table structure. The check goes there rather than into each provider's parser, so it also covers readers added later, and reuses the code `Block::insert` already returns. Three call sites. `tryGetTableStructureFromMetadata` and `buildStorageMetadataFromState` are members of one template class instantiated over every `IDataLakeMetadata` subclass, covering Iceberg, both Delta readers, Paimon and Hudi; the second is needed because the schema-reload path skips the first. `DatabaseDataLake` builds columns straight from the catalog and reaches neither, so `TableMetadata::setSchema` covers the REST, Glue, Unity, Hive and PaimonRest catalogs. Iceberg keeps the earlier check in its schema processor, which fires first when the name comes from metadata.json; a catalog that builds the column list itself, such as Glue, does not reach it, so an Iceberg table read that way reports `AMBIGUOUS_COLUMN_NAME`. Validation: both Delta readers and Paimon aborted before the change and now report the error, measured under `allow_experimental_delta_kernel_rs` 0 and 1. Well-named lakes still read on every arm, and the test fails on a build without this change. The catalog site is covered by a Glue integration test asserting on `SHOW CREATE TABLE`: a read is not a usable oracle there, because the empty name also reaches `Block::insert` and yields the same code without the check.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114394",
        "createdAt": "2026-08-11T22:59:21Z",
        "updatedAt": "2026-08-13T09:09:41Z",
        "timestamp": "2026-08-13T09:09:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114398",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Accept signed PromQL @ timestamps",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Accept signed numeric timestamps in PromQL @ modifiers. ### Summary PromQL allows an optional `+` or `-` sign before a numeric timestamp in an `@` modifier. ClickHouse previously accepted only unsigned timestamps. This change accepts signed timestamps and preserves negative values in the PromQL query tree. It also keeps negative evaluation times safe for TimeSeries tables with non-negative timestamp types. For `DateTime` and `UInt32` storage, selector ranges are clipped at the Unix epoch instead of allowing negative timestamps to wrap to a large positive value. `DateTime64` continues to support pre-epoch timestamps. Examples: - `http_requests_total @ -100` - `http_requests_total @ +3.3e1` - `http_requests_total @ 1700000000` ### Changes - Allow an optional sign before numeric timestamps in the PromQL grammar. - Preserve negative timestamps in the query tree. - Clip selector ranges to the supported timestamp range for `DateTime` and `UInt32` TimeSeries storage. - Keep negative timestamps available for signed `DateTime64` storage. ### Tests - Added parser coverage for negative timestamps. - Added parser coverage for explicitly positive scientific timestamps. - Added integration coverage for `UInt32`, `DateTime`, and `DateTime64` TimeSeries timestamp types. - Covered negative timestamps before the Unix epoch and positive timestamps whose lookback range crosses the epoch. - Regenerated the ANTLR parser artifacts. References: - [Prometheus query basics](https://prometheus.io/docs/prometheus/latest/querying/basics/)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114398",
        "createdAt": "2026-08-11T23:40:48Z",
        "updatedAt": "2026-08-13T15:35:58Z",
        "timestamp": "2026-08-13T15:35:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "submodule changed",
          "manual approve",
          "can be tested",
          "comp-promql"
        ],
        "author": "fallintoplace",
        "state": "open",
        "assignees": [
          "vitlibar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114400",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not let artifact-collection rows discard a bugfix-validation verdict",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113397 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description @ nikitamikhaylov asked me to investigate why `Bugfix validation` fails on #113397 and suggest improvements ([comment](https://github.com/ClickHouse/ClickHouse/pull/113397#issuecomment-5258844349)). This fixes the harness defect I found while answering that. Related: #113397, @ tiandiwonder's PR, which this one does not close. There, both functional bugfix-validation jobs report FAIL even though the bug reproduced on both arches: the regression-test row is `OK`, and the only FAIL row is `Scraping system tables`. An ordering defect. `invert_bugfix_validation_status` decides the verdict, then the COLLECT_LOGS stage appended artifact-collection rows with a bare `extend_sub_results`, which re-derives the parent status from its children (`praktika/result.py`), so one failed system-table dump overwrote the verdict with `FAIL`. `new_tests_check.py` reads that status with strict `is_success`, so `any_bugfix_validation_passed` found no validating arch and blocked the PR with \"No per-arch Bugfix Validation job validated the bug\". Those rows come from `prepare_logs`, after the verdict, so they can never be part of it. The fix attaches them through a helper that restores the captured status on a labelled bugfix-validation job. The rows stay visible in the report, they just stop voting. Restoring the captured status rather than forcing `OK` keeps all three verdicts: reproduction `OK`, no-repro `SKIPPED`, inconclusive `ERROR`. An ordinary job still reddens. The defect reproduces in-process, and six new cases in `ci/tests/test_bugfix_validation_inverter.py` pin an ordering that had no coverage. Each was checked against a mutated fix so its redness is attributable: dropping the restore fails the three verdict cases, restoring the bare call fails the AST case, and applying the restore unconditionally fails the ordinary-job control. One records why the rows must not be labelled `LOG_CHECK` instead, which would count a broken dump as a reproduction. Only the ordering is fixed here. The dump itself failed because a graceful-stop timeout left the server holding its status-file lock, which is separate. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1299` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114400",
        "createdAt": "2026-08-12T00:02:39Z",
        "updatedAt": "2026-08-13T01:17:21Z",
        "timestamp": "2026-08-13T01:17:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114401",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keeper: do not lose a session request when the Raft leader changes",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/78474 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a bug where a client connecting to ClickHouse Keeper during a Raft leader change could be held for the whole `session_timeout_ms` (30 seconds by default) before its connection was rejected, instead of being rejected as soon as the in-flight request was dropped. The connecting client can now reconnect to another replica sooner. ### Description A client's `Connect` makes Keeper submit an internal `SessionID` request. If the Raft append stream breaks while it is in flight, during a leader election say, that request is lost silently. **Root cause.** Such a request carries `session_id = -1` and no xid; its identity lives in `(server_id, internal_id)`. Two places used the wrong key: * `KeeperRequestDispatcher::onCommit` correlated a commit with its in-flight head by `(session_id, xid)`. Every `SessionID` request shares `(-1, 0)`, and `onCommit` runs on every node for every commit, so a `SessionID` committed for another server retired a still-uncommitted local request. * The error path queued the response for lookup by `session_id`, which `-1` has no callback for, so it was discarded and `getSessionID` timed out. The real waiter is a promise keyed by `internal_id`. **The change.** `onCommit` additionally requires `(server_id, internal_id)` to match for `OpNum::SessionID`. Error responses go through one `SessionID`-aware routing helper on `KeeperDispatcher`, wired into **both** dispatchers: `use_new_dispatcher` is a setting and the old one shares the defect. The request now fails at the in-flight drain bound rather than at `session_timeout_ms`. `KeeperTCPHandler` does not branch on the error code, so that earlier rejection is the whole user-visible gain; the accurate `ZCONNECTIONLOSS` only improves the server log. `test_keeper_force_recovery` also gets a retry around the connect after the election, since dropping in-flight appends is deliberate. The correlation fix is scoped to `OpNum::SessionID`. The garbage collectors also use negative session ids, but their `TryRemove` is idempotent and unwaited, so they are unaffected. **Validation.** New gtests cover both the routing decision and the production wiring behind it, each verified to go red under a mutation of the change it covers; the old dispatcher was exercised with `use_new_dispatcher = false`. An injected fault at the connect under test reddens the integration test without the retry and passes with it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114401",
        "createdAt": "2026-08-12T00:10:11Z",
        "updatedAt": "2026-08-13T17:47:09Z",
        "timestamp": "2026-08-13T17:47:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "antonio2368"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114408",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not throw a logical error when an ephemeral node is held by someone else",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/113913 Related: https://github.com/ClickHouse/ClickHouse/issues/86434 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a logical error reported as `Ephemeral node ... still exists after ...s, probably it's owned by someone else` when activating a replica whose `is_active`/`active` node in Keeper is still held by another session. This is expected runtime state, so it is now reported as `ABORTED` instead of `LOGICAL_ERROR`, and no longer aborts the server in debug and sanitizer builds. ### Description Related: #113913, #86434 `deleteEphemeralNodeIfContentMatches` finds the znode held by a foreign owner, logs `isn't owned by us. Will wait until it disappears`, waits `3 * session_timeout_ms`, and on timeout threw `LOGICAL_ERROR`. That is expected runtime state. Another server may legitimately hold the node, and the wait can be shorter than the holder's session: Keeper drops a dead session's ephemerals only at expiry, and the timeout negotiated in the handshake is stored in `Coordination::ZooKeeper`'s own `args` copy, never propagated to the `zkutil::ZooKeeper::args` this wait reads. The message names that case and then called it a logical error anyway. Fixed at the shared site: `LOGICAL_ERROR` -> `ABORTED`. A caller-side `catch` cannot fix it, because `LOGICAL_ERROR` aborts from the `Exception` constructor, before unwinding. The wait and control flow are unchanged, so the operation still fails and activation still refuses to proceed. The message now names the foreign owner as the primary explanation instead of a config mismatch or a bug. Follows #113913, which chose `ABORTED` for the sibling condition in `StorageKafka2::assertActive` (direct-read path; this is the activation path). All six call sites already treat a failed activation as retryable and none branches on the code. Both directions verified with a `TestKeeper` gtest, and end-to-end on a live server via `ReplicatedMergeTreeRestartingThread`: before the change the server aborts, after it the error is retried with backoff and the replica stays read-only. Two latent issues nearby are left alone: the negotiated `session_timeout_ms` is never synced back into `zkutil::ZooKeeper::args`, and `StorageKafka2` writes `active_node_identifier` as a bare UUID where the MergeTree restarting thread uses the quoted `generateActiveNodeIdentifier()` form.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114408",
        "createdAt": "2026-08-12T01:15:24Z",
        "updatedAt": "2026-08-13T08:18:50Z",
        "timestamp": "2026-08-13T08:18:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114409",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"Revert the PromQL topk/limitk streaming plan and its shared-subquery materialization\"",
        "text": "Reverts ClickHouse/ClickHouse#114326 Depends on https://github.com/ClickHouse/ClickHouse/pull/113397",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114409",
        "createdAt": "2026-08-12T01:15:49Z",
        "updatedAt": "2026-08-13T18:00:37Z",
        "timestamp": "2026-08-13T18:00:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114410",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix quadratic insert into Poco::ListMap for repeated keys",
        "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Parsing an HTTP request that carries many query parameters or headers with the same name (for example thousands of repeated `role=` parameters) no longer takes quadratic time. ### Details `HTMLForm` and `MessageHeader` keep their entries in `Poco::Net::NameValueCollection`, which is a `Poco::ListMap` over `std::list`. `ListMap::insert` keeps all entries with an equal key in one contiguous block, but it located the end of that block by scanning from the front of the list, so inserting *n* values under one key cost O(n²) case-insensitive comparisons. This is visible when a client selects a large set of roles per request via `?role=a&role=b&…`: with ~10k `role` parameters roughly 0.2–0.3 s per request was spent inside `HTMLForm` before the query started (server-side `query_duration_ms` stays flat while client-observed latency grows super-linearly). Because equal keys always form a single block (`insert` and `operator[]` are the only ways to add entries), the end of the block is simply one past the *last* matching entry, so search from the back instead. Appending under the most recently added key — the repeated-parameter case — becomes O(1); in general an insert now costs the distance from the back of the list to that key's block rather than the distance from the front. A key that is not present still costs one full scan, exactly as before; iteration order and `find()` (first of block) are unchanged. Micro-benchmark of `Poco::ListMap<std::string, std::string>::insert`, *n* values under the same key (clang, `-O2`): | n | before | after | |---:|---:|---:| | 1 000 | 2.4 ms | 0.09 ms | | 10 000 | 226 ms | 0.85 ms | | 30 000 | 2.2 s | 3.3 ms | A unit test is added for the block-contiguity / `find` / `erase` / `operator[]` semantics and for the repeated-key ordering; it passes both before and after this change (no timing assertions). Inserting many *distinct* keys is still O(n²) overall since each miss scans the list; that path is already bounded by `http_max_fields` and is left as is here. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1279` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114410",
        "timestamp": "2026-08-12T21:14:29Z",
        "metrics": {
          "reactions": 2,
          "comments": 2
        },
        "labels": [
          "pr-performance",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "sanjams2",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114413",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix loss of primary-key pruning for `DateTime64` columns under a `toUnixTimestamp` sorting key",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114407 Related: https://github.com/ClickHouse/ClickHouse/issues/79977 Related: https://github.com/ClickHouse/ClickHouse/pull/101814 For a `MergeTree` table with `ORDER BY toUnixTimestamp(ts)` where `ts` is `DateTime64`, a plain range filter on `ts` (e.g. `WHERE ts >= '2026-06-15'`) stopped using the primary key in 26.7: `EXPLAIN indexes = 1` shows `PrimaryKey Condition: true` and every granule of the matched parts is read. The regression came from #101814, which correctly changed the monotonicity gate in `KeyCondition::extractMonotonicFunctionsChainFromKey` to ask `getMonotonicityForRange` over the function's argument type instead of its result type. That exposed a gap: `ToNumberMonotonicity` (the monotonicity implementation behind `toUnixTimestamp` and the `toInt*`/`toUInt*` family) did not support `DateTime64` arguments at all and answered \"unknown\", so the monotonic chain was rejected and the pruning was silently lost. Keys like `toDate(ts)` or `toStartOfHour(ts)` were unaffected because their monotonicity classes handle `DateTime64` explicitly, and `toUnixTimestamp` over a plain `DateTime` column was unaffected because `DateTime` is on the whitelist of `ToNumberMonotonicity`. This PR teaches `ToNumberMonotonicity` to answer for `DateTime64` arguments. The conversion takes the whole number of seconds, truncating the fractional part toward zero, and throws a `DECIMAL_OVERFLOW` exception instead of wrapping around (see `DecimalUtils::convertTo`), so it preserves order everywhere it is defined: - For a signed target of at least 64 bits (e.g. `toInt64`), the conversion is total — the whole number of seconds of any `DateTime64` fits — so it is reported as `is_always_monotonic`, for concrete ranges as well. As a bonus, the mirror-image filter `toInt64(ts) >= c` over `ORDER BY ts` now prunes too (the batched application over columns of index values can never throw), and stays consistent with the exact-ranges `count()` optimization. - For narrower or unsigned targets (e.g. `toUnixTimestamp`, whose result is `UInt32`), the conversion throws for a part of the domain, so it must not claim `is_always_monotonic`: `matchesExactContinuousRange` takes that claim as a promise that the per-range analysis can confirm every granule of a found range, while a part may hold out-of-range values, for which the per-range analysis has to answer \"unknown\". The first version of this PR made exactly that inconsistent claim, and the AST fuzzer found the debug assertion \"Inconsistent `KeyCondition` behavior\" (a logical error, so the stateless runs on the debug build also showed \"Server died\"). Instead, the conversion now reports a new flag `Monotonicity::is_always_monotonic_where_defined` — monotonic over the subset of the domain where the evaluation succeeds — which is consumed only by the constant-pushdown gate in `canConstantBeWrappedByMonotonicFunctions`. It is sound there: stored keys cannot correspond to out-of-range values, because computing the sorting key at insert time would have thrown, and an unrepresentable constant is rejected gracefully by the guards in `applyFunctionChainToColumn`. Per the review findings, those guards are also fixed: the pre-execution range checks are extended from `Date`/`DateTime`/`UInt32` to the whole native integer family, so an out-of-range constant pushed through e.g. `ORDER BY toUInt8(ts)` is rejected gracefully instead of throwing a `DECIMAL_OVERFLOW` exception during index analysis, and the negative-value fast reject is now unsigned-specific, so signed integer keys (e.g. `ORDER BY toInt64(ts)`) keep pruning for pre-`1970-01-01` filters. For partial conversions like `toUnixTimestamp`, the mirror-image case #79977 (`WHERE toUnixTimestamp(ts) >= c` over `ORDER BY ts`, a full scan since 23.1) remains out of scope: index analysis applies chain functions to whole columns of index values (see `applyFunction` in `KeyCondition.cpp`), and a part may contain out-of-range values next to the checked range, for which the batched conversion would throw an exception. For total conversions like `toInt64` it is fixed here. The test covers the restored pruning (with `force_primary_key`), sub-second bounds of the relaxed atom (positive and negative), constants outside the `UInt32` range, `toInt64` sorting keys including pre-`1970-01-01` filters, out-of-range constants over a narrow `toUInt8` key, the mixed-part case where index analysis must not throw an exception, and the exact-ranges `count()` optimization over a total conversion. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a regression in 26.7: for a `MergeTree` table ordered by `toUnixTimestamp` (or another integer conversion) of a `DateTime64` column, a plain range filter on that column no longer used the primary key and read all granules of the matched parts. Additionally, a filter like `toInt64(ts) >= c` over a table ordered by the raw `DateTime64` column now uses the primary key.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114413",
        "createdAt": "2026-08-12T02:15:56Z",
        "updatedAt": "2026-08-13T15:15:36Z",
        "timestamp": "2026-08-13T15:15:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114414",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add SHARED REGEXP path placement policy to JSON",
        "text": "Adds `SHARED REGEXP 'pattern'` to `JSON` type declarations so matching root-relative flattened paths are always stored in shared data instead of competing for dedicated dynamic-path subcolumns. This is useful for high-cardinality path families whose promotion would displace more useful paths. Matching is partial by default, like `SKIP REGEXP`. Set the persisted `JSON` type parameter `shared_regexp_use_partial_match=0` to require full-string matching. Typed paths take precedence over `SHARED REGEXP`; `SKIP` and `SKIP REGEXP` continue to discard matching data. Rules are evaluated against the complete path from the original `JSON` root, including derived sub-objects. The rules are compiled once into an immutable `RE2` matcher, with bounded rule count and pattern sizes. A single rule uses direct `RE2` matching and multiple rules use `RE2::Set`. `JSON` columns without rules keep a null matcher, their existing binary type encoding, and their existing row-placement path. Metadata snapshots are copied lazily only when retained placement provenance actually changes a type. Placement provenance is stored in part column metadata and retained by default through horizontal and vertical merges, wide and compact mutations, column renames, lightweight-update patch materialization, and `Array`/`Nullable`/`Tuple`/`Map` wrappers. Removing a rule therefore does not unexpectedly promote already-shared paths during a later rewrite. Set the `MergeTree` table setting `allow_json_shared_data_paths_repromotion=1` to opt into reconsidering those paths. Policy-only metadata changes inside `Variant` are documented as unsupported. The `Native` binary type encoding uses `JSON` encoding version 1 only when a policy is present; ordinary `JSON` types remain byte-identical version 0. The canonical `JSON` documentation and `Native` format specification are updated accordingly. Testing includes 18 focused unit tests and 9 stateless scenarios covering syntax, partial/full matching, root-relative sub-objects, all supported wrappers, flattened `Native` and `RowBinary` paths, horizontal/vertical merges, compact/wide mutations, renames, statistics, and both lightweight-patch application paths. An 88.6 MB real Elasticsearch `indices_stats` document produced 143,555 flattened paths; all 143,290 paths targeted by `^indices[.]` remained in shared data and none became dynamic. A one-run end-to-end smoke comparison was 1.66 s without the policy and 1.69 s with it, with a 0.055% max-RSS difference; these small differences are noise-level, not a statistical benchmark. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `SHARED REGEXP` rules to the `JSON` data type to keep matching paths in shared data, with configurable full-string matching and explicit control over re-promoting paths after rules are removed.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114414",
        "createdAt": "2026-08-12T02:35:41Z",
        "updatedAt": "2026-08-13T17:11:36Z",
        "timestamp": "2026-08-13T17:11:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "valerypetrov",
        "state": "open",
        "assignees": [
          "Avogar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114417",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add the `Cluster` database engine",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114411 Related: https://github.com/ClickHouse/ClickHouse/issues/59304 Related: https://github.com/ClickHouse/ClickHouse/pull/110975 Implement the `Cluster` database engine, a follow-up to the `Remote` database engine (#110975). It provides real-time access to the tables of a database on a cluster from the server configuration — the named-cluster counterpart of `Remote`, exactly as the `cluster` table function relates to the `remote` table function: ```sql CREATE DATABASE db ENGINE = Cluster('cluster_name', 'database'); ``` The engine shares the whole metadata and query machinery with the `Remote` database engine (the list of tables and their structure are fetched from the cluster on demand, each table is a `Distributed` proxy forwarding `SELECT` and `INSERT`, the same local-shard visibility rules and remote-replica fallback). The differences: - The cluster is resolved from the server configuration by name on every access, so the database follows configuration reloads and cluster auto-discovery, like a `Distributed` table does. Macros such as `{cluster}` are supported. The cluster must exist at `CREATE`, but a database whose cluster later disappears from the configuration does not prevent the server from starting — it reports the missing cluster until the configuration brings it back. - Connections use the per-replica settings of the cluster configuration (credentials, secure connections, compression, the inter-server secret), so the engine takes no credential arguments and stores no secrets. - `SHOW CREATE TABLE` prints a re-executable `Distributed('cluster_name', 'database', 'table')` definition (including the implicit `rand()` sharding key of a multi-shard database). The only exception is a table currently served through the remote-replica fallback: no equivalent re-executable definition exists for that transient state (a `Distributed` table over the whole cluster performs no such fallback, and the per-replica settings of the configuration cannot be carried by an explicit address list), so `SHOW CREATE TABLE` reports `THERE_IS_NO_QUERY` until the local replica has the objects again. Notes: - A chain of `Remote`/`Cluster` databases on the same server that refers back to itself is rejected at `CREATE DATABASE` of the database that would complete it (`INFINITE_LOOP`), because a lazily-reported cycle would fail every whole-server scan (`system.tables` and the like) for every user. A cycle that comes into existence later (configuration reload, explicit `ATTACH`) is skipped in listings and reported only by table resolution, and does not prevent the server from starting. - The remote-only metadata-lookup fallback (the cluster with the local replicas stripped from their shards) is derived by the new `Cluster::tryGetClusterWithoutLocalReplicas`, which preserves the per-replica settings and the inter-server secret — the string-based rebuild used by the `Remote` engine could not do that for configured clusters, and the `Remote` engine now uses the new method as well. - The follow-up `ddl_mode` setting (forwarding DDL to the cluster) is tracked separately in https://github.com/ClickHouse/ClickHouse/issues/114412. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added the `Cluster` database engine that provides real-time access to the tables of a database on a cluster from the server configuration, forwarding `SELECT` and `INSERT` queries to it. It is the named-cluster counterpart of the `Remote` database engine, as the `cluster` table function relates to the `remote` table function. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114417",
        "createdAt": "2026-08-12T03:48:20Z",
        "updatedAt": "2026-08-13T12:13:40Z",
        "timestamp": "2026-08-13T12:13:40Z",
        "metrics": {
          "reactions": 1,
          "comments": 4
        },
        "labels": [
          "pr-feature"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114420",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not apply DROP fault injection to CREATE OR REPLACE internal DROPs",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/110893 Related: https://github.com/ClickHouse/ClickHouse/pull/110971 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `CREATE OR REPLACE` leaking an internal `_tmp_replace_*` table, and `CREATE OR REPLACE VIEW` failing with `NOT_IMPLEMENTED`, when `ignore_drop_queries_probability` is set. The DROPs it issues internally are steps of one user statement, so DROP fault injection no longer applies to them. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/110893 Related: https://github.com/ClickHouse/ClickHouse/pull/110971 `CREATE OR REPLACE` builds the replacement under a temporary `_tmp_replace_*` name and publishes it by rename. It issues two DROPs internally: cleaning up that temporary table if the statement fails, and dropping the replaced table after the swap. Both took their context from `make_drop_context`, which did not mark it DDL-internal, so `InterpreterDropQuery` treated them as *user* DROPs and injected faults. Two things follow. If the storage keeps data on disk the DROP is skipped outright and a populated `_tmp_replace_*` table is stranded: listed by `SHOW TABLES`, holding its data, surviving a restart, one per statement. Otherwise (a materialized view, say) the DROP is rewritten to `TRUNCATE`, whose branch takes an exclusive lock under the outer statement's query id while that id already holds read locks, raising `RWLockImpl::getLock(): Cannot acquire exclusive lock while RWLock is already locked`. That fired twice on master, under [`Stress test (amd_tsan)`](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=d134c6da364233695c456f9c51fbbc521a4fbd5b&name_0=MasterCI&name_1=Stress%20test%20%28amd_tsan%29) and [`Stress test (arm_debug)`](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=1033651ae6f23b6c10f0023dd1247c6feb8bde4f&name_0=MasterCI&name_1=Stress%20test%20%28arm_debug%29). A third internal DROP is misclassified the same way. `CREATE OR REPLACE VIEW` uses that machinery only on `Atomic` or `Replicated`; other engines drop the view in place through `doCreateTable`. `StorageView` implements no `TRUNCATE`, so there the rewrite fails the statement with `NOT_IMPLEMENTED` and the stale view survives. The fix marks both contexts DDL-internal, as the sibling helper in the same function and `InterpreterDropQuery::executeDropQuery` already do; that asymmetry is why plain `CREATE ... AS SELECT` was immune and only REPLACE was exposed. User DROPs are still skipped, which the new test pins. The setting defaults to 0, so a default configuration is unaffected. The read lock the assertion reports is not identified; the fix does not depend on it, since without the rewrite nothing requests an exclusive lock. <details> <summary>Validation</summary> - New test `04796_create_or_replace_internal_drop_not_ignored` fails on master and passes with the fix, on every shape it covers. - All three sites covered by shape: `CREATE OR REPLACE TABLE ... AS SELECT` over MergeTree takes the skip path (stray `_tmp_replace_*` count accumulates 1, 2, 5, 6 on master); `CREATE OR REPLACE MATERIALIZED VIEW ... POPULATE` takes the rewrite-to-`TRUNCATE` path; `CREATE OR REPLACE VIEW` on a non-Atomic database fails on master with `Truncate is not supported by storage View`, leaving the stale definition readable. Failure path, success path (which leaks the *replaced* table), and repeated replaces all covered. - 50 randomized runs each, on the final binary, of the new test, `03013_ignore_drop_queries_probability`, `00916_create_or_replace_view` and `01866_view_persist_settings`: 200/200. Same for `04492`, `04524`, `04326`, `04328` on the first binary: 200/200. - A/B sweep of the 41 tests matching `create_or_replace`/`replace_table`/`tmp_replace`/`create_as_select`, plus a focused sweep of the 14 that reach the view site: each removes exactly one failure (the new test) and introduces none. - All 8 `Common.RWLock*` unit tests pass. - Replicated databases measured on both binaries across every reachable `CREATE OR REPLACE` shape: no change, no `INCORRECT_QUERY`. That path already marks the context internal (`DatabaseReplicatedTask::makeQueryContext`), so the fix is a no-op there. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1308` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114420",
        "createdAt": "2026-08-12T04:39:37Z",
        "updatedAt": "2026-08-13T11:44:46Z",
        "timestamp": "2026-08-13T11:44:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114422",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Convert only SEMI JOIN to IN in the convertJoinToIn optimization",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/101698 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results with `query_plan_convert_join_to_in = 1`. The optimization replaced a JOIN with `key IN (subquery)` for `ALL` and `ANY` strictness, but `IN` only tests membership: `ALL INNER JOIN` lost rows on a duplicated right key, and `ANY INNER JOIN` returned extra rows on a duplicated left key. Only `SEMI JOIN` is converted now. Closes #101698. ### Description `tryConvertJoinToIn` rewrites a JOIN into `key IN (subquery)` over the left side, emitting each matching left row exactly once: semi-join semantics. The gate excluded strictnesses by name rather than requiring that property, so it admitted two that break it in opposite directions. - **`ALL`**, the default and the reported bug, emits `left_count(k) * right_count(k)` rows, so a duplicated right key multiplies left rows and `IN` loses them: 3 rows off, 2 on. - **`ANY`** deduplicates the *left* side at the default `any_join_distinct_right_table_keys = 0`, so `IN` instead adds rows: 2 off, 5 on. The gate now admits only `Semi`; `Left` joins the kind gate since `SEMI` needs `LEFT`/`RIGHT`. Six more declines, each a divergence I measured off vs on: - a `Join` engine right side, whose declared kind and strictness the rewrite stops validating: on `Join(ANY, LEFT, id)` it threw `INCOMPATIBLE_TYPE_OF_JOIN` off, returned rows on; - an active `max_rows_in_join`, `max_bytes_in_join`, `max_rows_to_transfer` or `max_bytes_to_transfer`: the join bounds its stored right side, the set only its hash table, so a limit could stop being enforced; - a key whose type has dynamic structure, which `IN` rejects; - a key-value prepared right side (dictionary, `EmbeddedRocksDB`, any `IKeyValueEntity`), which the planner probes by key: `DirectKeyValueJoin` off, a 2000000-element set on. The `Join` engine check above tests the whole `PreparedJoinStorage`, covering both; - an algorithm the planner would pick ahead of hash, since being enabled is not being chosen: `partial_merge`, `prefer_partial_merge`, `grace_hash` or `auto` listed first streams or spills the right side the set materializes (1 off, 0 on each). `hash,partial_merge` and the stock `direct,parallel_hash,hash` still convert, so position decides; `full_sorting_merge` rejects `SEMI` and already falls through to hash. - a key expression consuming a left column the projection needs, e.g. `ON arrayJoin(l.tags) = r.tag` selecting `l.tags`, throwing `NOT_FOUND_COLUMN_IN_BLOCK`. #104809 fixes that hole from the `ALL` side, so I reuse its predicate. Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/101698#issuecomment-5254250810",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114422",
        "createdAt": "2026-08-12T05:30:34Z",
        "updatedAt": "2026-08-13T15:37:31Z",
        "timestamp": "2026-08-13T15:37:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114423",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113534 to 26.7: Fix a mixed JOIN ON condition evaluated over mismatched column types for a dictionary",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113534 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31567499495/job/94022263699)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114423",
        "createdAt": "2026-08-12T06:02:32Z",
        "updatedAt": "2026-08-13T00:26:07Z",
        "timestamp": "2026-08-13T00:26:07Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-clickhouse",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114434",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Preserve the RabbitMQ broker log in integration tests",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114415 Related: https://github.com/ClickHouse/ClickHouse/pull/113610 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description When a RabbitMQ container never becomes available, `wait_rabbitmq_to_start` raises and every test in the module fails from that one event. On master @ `e4d3698282346fcd7e496ef668ff517106c84ac5` that cost all 21 tests of `test_storage_rabbitmq/test_system_stop.py` from one fixture failure ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=e4d3698282346fcd7e496ef668ff517106c84ac5&name_0=MasterCI&name_1=Integration%20tests%20%28arm_binary%2C%20distributed%20plan%2C%201%2F4%29)). Four such events in 21 days, two on master. The artifact that would say why does not survive the run. `rabbitmq.conf` sets `log.file = /rabbitmq_logs/rabbit.log` and no `log.console`, so everything goes to that file and nothing to the console, which is also why the `docker logs` capture is empty. `/rabbitmq_logs` was declared `tmpfs`, so the log died with the container. `cluster.py` has always exported `RABBITMQ_LOGS` and `RABBITMQ_LOGS_FS=\"bind\"` and created the host directory, but no compose file consumed them after 545bb5bae91233313868ee4d98d7e6dc6efe20db replaced the bind mount with tmpfs. RabbitMQ was the only broker whose `*_LOGS` export was dead. `rabbit.log` has zero hits in the failing run's whole artifact bundle. This restores the mount in the same `${VAR_FS:-tmpfs}` / `${VAR:-}` form the siblings use, so an unset `RABBITMQ_LOGS` still degrades to tmpfs. `/var/lib/rabbitmq` stays a tmpfs: it holds Mnesia state, and persisting it would change test semantics. `test_system_stop.py` passes 21/21 and `test.py` 58/58 with the mount, each producing a `rabbit.log` of over 100 KB under `_instances*/rabbitmq/logs/`, which the integration job already tars. The new `test_rabbitmq_broker_log_is_collected` asserts that log exists and is non-empty; reverting the hunk alone leaves the rest of the module passing and reddens only that test, on an empty log directory. This neither stops the abort nor explains why the Erlang node failed to register with epmd. #114415 recreates a hung container; this makes the next occurrence diagnosable. #113610 fixed only the second diagnostics loop of `wait_rabbitmq_to_start` and is not a container-start fix.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114434",
        "createdAt": "2026-08-12T08:11:44Z",
        "updatedAt": "2026-08-13T01:17:12Z",
        "timestamp": "2026-08-13T01:17:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim",
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114437",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Refuse Iceberg data compaction until it can publish its result",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/114194 Related: https://github.com/ClickHouse/ClickHouse/pull/114324 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `OPTIMIZE TABLE` on an Iceberg table (data compaction, `allow_experimental_iceberg_compaction = 1`) deleted files that retained snapshots still referenced, including `version-hint.text` and every `vN.metadata.json`, which left the table unreadable and could permanently lose an acknowledged snapshot's rows. In the open-source build it now reports `NOT_IMPLEMENTED` instead of running. `OPTIMIZE TABLE ... MANIFEST`, `expire_snapshots` and `remove_orphan_files` are unaffected. ### Description Related: #114194 (the same stale-hint root reaching `remove_orphan_files`). Measured on master (`26.8.1.1`), v2 table with position deletes: `OPTIMIZE TABLE` returns success silently, drops the object count 18 to 10, and deletes `metadata/version-hint.text` with `v1..v5.metadata.json`; every later read fails `Code: 107 FILE_DOESNT_EXIST`. No stale hint is needed. When the hint is behind the newest metadata, the rewrite is rooted at the older version, so an acknowledged snapshot's data files are not carried forward yet are still deleted: `groupArray(x)` returns `[1]` where `[1,99]` was committed. Root cause: `getOldFiles` takes a raw listing of `metadata/` and `data/` with no reachability filter and no age gate, and `clearOldFiles` deletes all of it. Those files are reachable by the engine's own definition - `collectSnapshotReferencedFiles` walks every entry of the `snapshots` array - so deleting them is wrong at any age, under any guard. The rewrite is also never published atomically: `getPlan` never calls `setVersion`, so it always writes `v0.metadata.json`, and `writeMetadataFiles` commits with a bare `WriteMode::Rewrite`, no hint and no catalog commit. Since the old files are listed before it runs, cleanup deletes the previous `vN` and leaves `v0`, so a reader resolves whichever generation survived - in the measured run, the truncated one. Publishing correctly needs the current-snapshot replacement model and catalog plumbing that do not exist here, so this refuses the path rather than repairing it, following the existing format-version 3 refusal for `OPTIMIZE ... MANIFEST`. The data-compaction success contracts in the tests were deleted rather than inverted into \"the feature is absent\" assertions, which would have to be deleted again once publication lands. The new integration test asserts the property that holds either way: no pre-existing file is removed and the committed rows survive. A stateless test pins the refusal. `OPTIMIZE ... MANIFEST` coverage is untouched.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114437",
        "createdAt": "2026-08-12T08:24:40Z",
        "updatedAt": "2026-08-13T11:38:19Z",
        "timestamp": "2026-08-13T11:38:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114441",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112479 to 26.6: RabbitMQ related fix",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112479 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31577531496/job/94052923751) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.31` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114441",
        "createdAt": "2026-08-12T08:36:35Z",
        "updatedAt": "2026-08-13T12:21:44Z",
        "timestamp": "2026-08-13T12:21:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "closed",
        "assignees": [
          "kssenii",
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114442",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112479 to 26.7: RabbitMQ related fix",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112479 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31577531496/job/94052923751) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.26` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114442",
        "createdAt": "2026-08-12T08:37:03Z",
        "updatedAt": "2026-08-13T12:21:46Z",
        "timestamp": "2026-08-13T12:21:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "closed",
        "assignees": [
          "kssenii",
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114451",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Show the real per-query verdict in the performance comparison report",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/111992 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description A `Performance Comparison` shard that PASSED renders a list of red rows, and every affected query is listed twice. Reported in #111992: shard `arm_release, master_head, 1/6` was `OK`, its own message said `11 unstable`, and the report showed 22 red rows that had to be hand-verified. Two causes, both in the block that builds the `Check Results` rows: - The rows exist only to carry a per-query \"query history\" link, and each was built with a hardcoded `FAIL`, discarding the `slower`/`unstable` verdict into `info`. The renderer already classifies both natively, so passing the real one through is enough. - `compare.sh` emits one row per side (`array join map('old', left, 'new', right)`, named `<query> #<idx>::old` / `::new`), so 11 queries became 22 rows. Only the candidate side is now listed, and only its displayed name is stripped. `compare.sh` is unchanged: its per-side rows are what goes into the CI database, so changing the emission would fork that history and break every existing history link. Nothing is dropped, only displayed once. This cannot turn a passing shard red: `Check Results` is created with an explicit status, and praktika aggregates child statuses only when none is given. A test pins that guard, and the slower-query gate that decides the job verdict is untouched. One more line in `json.html`: `getStatusClass` had no `slower` case and fell through to gray, while the status filter and the row ordering both already treat it as a failure. Tests: `ci/tests/test_perf_check_results_children.py`, driven by two trimmed real shard artifacts - the one from #111992, and a failing `release_base` shard as the negative control, which must stay red. The parent's own message states the query count, an independent oracle for the row count. The CI database rows are byte-identical before and after, and the history links still resolve.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114451",
        "createdAt": "2026-08-12T09:11:17Z",
        "updatedAt": "2026-08-13T04:42:48Z",
        "timestamp": "2026-08-13T04:42:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "maxknv",
          "egor-click"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114457",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix header column order after the no-rescoring vector search rewrite",
        "text": "<!--- A technical description of your changes with a motivation --> The second-pass vector search optimization (`vector_search_with_rescoring = 0`) removed the distance function from the `ExpressionStep` outputs and re-appended the rewritten `_distance` alias at the end, changing the header column order. Steps created above the `Sorting` before that rewrite runs — the local top-N `Limit` and the exchange steps of a distributed plan — kept the original column order, and `makeDistributedPlan` failed to rebuild the plan fragments with a logical error: `Cannot add step Limit to QueryPlan because it has incompatible header with root step Sorting`. Non-distributed plans never re-validate step headers after optimization, which is why this only surfaced with `make_distributed_plan = 1`. The fix reinserts the rewritten output node at the original position of the distance column, so the header column order is preserved and all previously created parent steps stay consistent. Found by AST fuzzer on [#42701](https://github.com/ClickHouse/ClickHouse/pull/42701) ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=019244e0d46b4d9029fc261a549e1c5cd79b1bbb&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)); the same signature also hit the unrelated #98789 on 2026-08-08. The regression test asserts only the row count: the distributed plan currently returns an incorrect top-N for vector search queries regardless of the rescoring mode — a separate, pre-existing bug. Related: https://github.com/ClickHouse/ClickHouse/issues/114456 Related: https://github.com/ClickHouse/ClickHouse/pull/42701 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a logical error `Cannot add step Limit to QueryPlan because it has incompatible header with root step Sorting` when a vector search query with `vector_search_with_rescoring = 0` was executed with the experimental `make_distributed_plan` setting.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114457",
        "createdAt": "2026-08-12T09:43:36Z",
        "updatedAt": "2026-08-13T14:55:29Z",
        "timestamp": "2026-08-13T14:55:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114460",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Text index: fix accidentally removed columns when the direct read is enabled",
        "text": "When the same column is used in both `PREWHERE` and `WHERE` conditions, it must not be removed, since `PREWHERE` is executed before `WHERE`. Even if a `WHERE` filter column is rewritten into a text index virtual column, it still has to be read to apply the `PREWHERE` condition. To fix it, now every column still required by the `PREWHERE` is kept in the read set instead of removing it. Fixes #113320. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Do not remove columns when it's referenced in both `PREWHERE` and `WHERE` filter conditions even if it's rewritten into a text index virtual column.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114460",
        "timestamp": "2026-08-12T23:12:53Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "ahmadov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114461",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "CI: route documentation review requests",
        "text": "Enable documentation review requests now that the relevant teams have repository access. ClickPipes documentation changes request reviews from both `docs` and `clickpipes`; language-client and connector documentation changes request reviews from both `docs` and `integrations-ecosystem`. Documentation reviews are skipped when a PR also changes files under `src/`. GitHub App installation tokens cannot resolve private teams when requesting reviews, even with the documented repository permission. Use the existing robot credential for internal PRs and for a `pull_request_target` workflow that checks out only the trusted base revision. Fork PRs must have the `can be tested` label before the trusted workflow runs. This internal draft exercises the credential and API path from the PR branch. It includes a one-line Python language-client documentation edit, which should request reviews from both `docs` and `integrations-ecosystem`. Related: https://github.com/ClickHouse/ClickHouse/pull/114458 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable automatic documentation team review requests through a trusted workflow.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114461",
        "createdAt": "2026-08-12T09:57:54Z",
        "updatedAt": "2026-08-13T12:47:29Z",
        "timestamp": "2026-08-13T12:47:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-ci"
        ],
        "author": "Blargian",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114463",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Report mutation cancellation instead of stopping nested pipelines silently",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107619 Related: https://github.com/ClickHouse/ClickHouse/pull/112152 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `Not-ready Set is passed as the second argument for function 'in'` error when a mutation whose predicate contains `IN (subquery)` is cancelled, for example by `KILL MUTATION` or `DETACH DATABASE`, while the subquery's set is still being built. Related: #107619. ### Description Reported by @ alexey-milovidov on https://github.com/ClickHouse/ClickHouse/pull/112152#issuecomment-5180281674 after a `Stress test (amd_debug)` run aborted this way. `buildSetInplace` materializes the right side of `IN (subquery)` through a nested `CompletedPipelineExecutor` and polls the mutation's interactive-cancel callback. That callback returned a bare `bool`, so on cancellation the executor called `cancel()` and the poll loop returned normally: `Set::finishInsert` never ran, while `build()` had already moved the source plan out. `buildSetsForDAG` returns `void`, so `getMinMaxCountProjectionBlock` evaluated the filter in the same call and `FunctionIn` threw. A constant left-hand side (`1 IN (...)`) makes this the first materialization attempt: it maps to no key column, so primary-key analysis returns first. The callback now throws `ABORTED`, mirroring `MergeTask::checkOperationIsNotCanceled`, which reports the merge path the same way and is likewise used as a nested-pipeline callback. `MergeTreeBackgroundExecutor` already treats `ABORTED` as a normal outcome and logs it at DEBUG. Only the mutation installs this callback shape, so the installers in `TCPHandler` and `LocalConnection` are untouched and client cancellation still returns `QUERY_WAS_CANCELLED` quietly. Validation: `KILL MUTATION` and `DETACH DATABASE` mid-build abort the server before the change and are clean after it, and the table stays mutatable. New test `04865_cancel_mutation_in_subquery_minmax_projection`, 50 of 50 runs. The other cancellation routes and overflow modes checked are listed in the validation gate comment below. Out of scope: a caller that executes a filter DAG synchronously still does not check set readiness, and `buildSetInplace` still leaves the set unbuildable if some future path stops it silently. I found no input reaching either independently of this cancellation. Note for review: #113939 rewrites the same lambda for the S3 read path but keeps the interactive callback non-throwing, so it does not fix this. The two conflict textually; a resolution must keep both the throw and that PR's cancellation persistence.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114463",
        "createdAt": "2026-08-12T10:03:39Z",
        "updatedAt": "2026-08-13T15:58:55Z",
        "timestamp": "2026-08-13T15:58:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114465",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix an error reading an Iceberg table whose `current-snapshot-id` is JSON null",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109739 ## Problem The Iceberg spec lets `current-snapshot-id` be absent, `-1`, or JSON null, and all three mean the same thing: the table has no current snapshot. External writers do emit `\"current-snapshot-id\": null`. Reading such a table, several commands fail with an unrelated Poco conversion error instead of taking the no-snapshot path: - `SELECT ... FROM system.iceberg_history` reports the table as broken and skips it, so its history is silently missing from the result: `Ignoring broken table <db>.<t>: Poco::Exception. Code: 1000, Invalid access: Can not convert empty value.` - `ALTER TABLE ... EXECUTE expire_snapshots(...)` fails. - `DELETE` / `UPDATE` fail. ## Root cause `Poco::JSON::Object::has` returns true for a key whose value is JSON null — the key is present, the value is an empty `Poco::Dynamic::Var`. A reader guarded only by `has` therefore proceeds to `getValue<Int64>`, and `Var::convert<Int64>` throws `Poco::InvalidAccessException`: ``` Poco::Dynamic::Var::convert<long>() DB::IcebergMetadata::getHistory(std::shared_ptr<DB::Context const>) const DB::StorageSystemIcebergHistory::fillData(...) ``` Each affected site is immediately followed by a `< 0` / `>= 0` test, so treating null like a negative id is what the surrounding code already intends; it simply never handled that spelling. ## Fix Three reads now also check `isNull`, matching the idiom already used by their neighbours in the same files (`IcebergMetadata.cpp:581`, `Mutations.cpp:601`, `IcebergWrites.cpp:1120`): | Site | Reached by | |---|---| | `IcebergMetadata::getHistory` | `system.iceberg_history` | | `expireSnapshots` | `ALTER TABLE ... EXECUTE expire_snapshots(...)` | | `mutate` | `DELETE` / `UPDATE` | Regression test: `tests/queries/0_stateless/04846_iceberg_null_current_snapshot_id.sh`, which rewrites `current-snapshot-id` to JSON null and exercises all three. Verified closed-loop — with the three guards reverted and rebuilt, the test fails with the error above; with them restored it passes. Each guard was attributed individually rather than only as a group. Two related reads are deliberately **not** changed here: - `writeMetadataFiles` (`Mutations.cpp`) has the same pattern, but no repro could be constructed: it runs only after a mutation has matched rows, and a null `current-snapshot-id` makes the table read as empty, so the two conditions are mutually exclusive. Reverting only that line leaves the new test green. - The two reads in `Compaction.cpp` are fixed by #109739, which also adds the manifest-compaction validation and null coverage there. Both pull requests target `master` independently; their changes are complementary and can merge in either order. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an error when reading an Iceberg table whose `current-snapshot-id` is JSON null, which some writers emit for a table with no current snapshot.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114465",
        "createdAt": "2026-08-12T10:37:46Z",
        "updatedAt": "2026-08-13T11:32:44Z",
        "timestamp": "2026-08-13T11:32:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114466",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: internationalize master",
        "text": "### Changelog category (leave one): - Documentation (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114466",
        "createdAt": "2026-08-12T10:49:59Z",
        "updatedAt": "2026-08-13T17:50:49Z",
        "timestamp": "2026-08-13T17:50:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 47
        },
        "labels": [
          "pr-documentation"
        ],
        "author": "locadex-agent[bot]",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114468",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "fix: error message for ALTER DROP COLUMN of a key column",
        "text": "<!-- PR title: Fix error message for ALTER DROP COLUMN of a column used in a key expression --> Closes: https://github.com/ClickHouse/ClickHouse/issues/71776 Related follow-up issues: #114481, #114181 `ALTER DROP COLUMN` and `ALTER CLEAR COLUMN` behaved inconsistently on a column that belongs to `ORDER BY`, `PRIMARY KEY` or `PARTITION BY` An `ALTER` is validated by building the new table schema first and inspecting it afterwards. Once the column is gone the sorting key cannot be resolved, so building the schema fails before the check that knows the real reason. That check is now done first, on the current schema, where the key columns are already known. It understands all the ways a key can refer to a column: the column itself, a column inside a key expression, a subcolumn (`ORDER BY a.x` for `a Tuple(...)`) and a whole nested group dropped by its common prefix Examples: 1. Basic ```sql CREATE TABLE t (a UInt64, b UInt64) ENGINE = MergeTree ORDER BY a; ALTER TABLE t DROP COLUMN a; -- before: Code: 47. Missing columns: 'a' while processing: 'a' ... (UNKNOWN_IDENTIFIER) -- after: Code: 524. Trying to ALTER DROP key a column which is a part of key expression. (ALTER_OF_COLUMN_IS_FORBIDDEN) ``` 2. Dropping a column that is used only in a TTL expression now also reports which TTL expression it breaks Note: `ALTER_OF_COLUMN_IS_FORBIDDEN` would be wrong here. Dropping a column used in a `TTL` is not always forbidden - `ALTER TABLE t DROP COLUMN d, MODIFY TTL a + INTERVAL 1 DAY` is valid and passes. And this is a wrapper not a check: the same step returns `ILLEGAL_TYPE_OF_ARGUMENT` for `MODIFY COLUMN d Array(UInt8)`, which `ALTER_OF_COLUMN_IS_FORBIDDEN` would mislabel ```sql CREATE TABLE t (d Date, a UInt64) ENGINE = MergeTree ORDER BY a TTL d + INTERVAL 1 DAY; ALTER TABLE t DROP COLUMN d; -- before: Code: 47. Missing columns: 'd' ... (UNKNOWN_IDENTIFIER) -- after: Code: 47. Cannot apply ALTER because it breaks the TTL of the table: Missing columns: 'd' ... (UNKNOWN_IDENTIFIER) ``` 3. For a key column the command is rejected whatever the partition is, so the old error only sent the user to fix something that would not help ```sql CREATE TABLE t (a UInt64, b UInt64, c UInt64) ENGINE = MergeTree PARTITION BY b ORDER BY a; ALTER TABLE t CLEAR COLUMN a IN PARTITION 'nonsense'; -- before: Code: 53. Cannot convert string 'nonsense' to type UInt64. (TYPE_MISMATCH) -- after: Code: 524. Trying to ALTER CLEAR key a column ... (ALTER_OF_COLUMN_IS_FORBIDDEN) ``` 4. With `share_nested_offsets` enabled, a name that is not a column of the table denotes the whole nested group `<name>.*`, and `DROP`/`CLEAR` of the group is now rejected the same way when any column of the group is used in a key. Before, this was worse than a confusing message: the checks compare names exactly and did not see the group, so `DROP COLUMN IF EXISTS n` destroyed the group's data with a mutation that then failed halfway and `CLEAR COLUMN n` on a non-empty table left a mutation that can never finish (a key column cannot be rewritten) in `system.mutations` until `KILL MUTATION` ```sql CREATE TABLE t (`n.a` UInt64, `n.b` UInt64, x UInt64) ENGINE = MergeTree ORDER BY `n.a`; ALTER TABLE t DROP COLUMN n; -- before: Code: 47. Missing columns: 'n.a' while processing: '`n.a`' ... (UNKNOWN_IDENTIFIER) -- after: Code: 524. Trying to ALTER DROP whole Nested group n whose columns (`n.a`) are part of key expression. (ALTER_OF_COLUMN_IS_FORBIDDEN) ALTER TABLE t CLEAR COLUMN n; -- before: accepted, data of `n.a` and `n.b` is zeroed and the mutation stays in system.mutations forever -- after: Code: 524. Trying to ALTER CLEAR whole Nested group n whose columns (`n.a`) are part of key expression. (ALTER_OF_COLUMN_IS_FORBIDDEN) ``` Groups with no key columns inside are dropped as before, and with `share_nested_offsets = 0` the name does not denote the group and the check does not apply. The remaining group/exact-name mismatches (silent data destruction for groups without key columns, dependent views, races with background mutations) predate this PR and are reported separately with reproductions ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... DROP COLUMN` of a column that is used in the sorting, primary or partition key now fails with `ALTER_OF_COLUMN_IS_FORBIDDEN` and an explanation, the same as `ALTER TABLE ... CLEAR COLUMN` does. Previously it failed with a confusing `UNKNOWN_IDENTIFIER: Missing columns` error coming from the recalculation of the key expressions. The same applies to dropping or clearing a whole `Nested` group by its common prefix when a column of the group is used in a key; previously that could silently destroy the group's data or leave a mutation that never finishes. Dropping a column that is used only in a `TTL` expression now also reports which TTL expression it breaks",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114468",
        "createdAt": "2026-08-12T11:43:38Z",
        "updatedAt": "2026-08-13T11:04:29Z",
        "timestamp": "2026-08-13T11:04:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-improvement",
          "can be tested"
        ],
        "author": "m7kss1",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114469",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keep heavy check in `ColumnArray` only in debug build",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) Fixes #114105. Introduced in #112504. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1291` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114469",
        "createdAt": "2026-08-12T11:47:45Z",
        "updatedAt": "2026-08-13T00:19:13Z",
        "timestamp": "2026-08-13T00:19:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "CurtizJ",
        "state": "closed",
        "assignees": [
          "scanhex12"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114472",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keep the patch release version bump increasing across recoveries",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113834 Related: https://github.com/ClickHouse/ClickHouse/pull/113528 --> Scheduled patch releases stopped incrementing the patch number — successive releases on a branch reused the same `vX.Y.P.*` line (e.g. `26.6.2.81`, `26.6.2.158`, `26.6.2.160`), because the post-release version bump was lost whenever a release was interrupted after the tag push, and every later recovery skipped it too. Prepare now classifies a run from the ref and the branch-tip version file into two flags: `is_recovery` (this run re-publishes an existing release rather than creating one) and `is_late_recovery` (the branch has already advanced to a newer release). The deferred bump is gated on `not is_late_recovery`, so a normal run and a current-release recovery complete the interrupted bump, while a superseded recovery never rewrites the version backwards. `stage_bump` refuses to write a version older than the branch tip, and asserts `is_recovery` on an empty bump — a fresh release must advance the version, a recovery may legitimately find it already done. Split out of #113834 — this is the version-bump half with a single push. Handling a non-fast-forward push (a backport moving the branch mid-release) is left to #113834. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114472",
        "createdAt": "2026-08-12T12:04:19Z",
        "updatedAt": "2026-08-13T17:12:19Z",
        "timestamp": "2026-08-13T17:12:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "can be tested",
          "pr-ci"
        ],
        "author": "leshikus",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114473",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113289 to 25.8: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114473",
        "createdAt": "2026-08-12T12:04:52Z",
        "updatedAt": "2026-08-13T07:17:25Z",
        "timestamp": "2026-08-13T07:17:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll3",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "Avogar",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114474",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113289 to 26.3: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114474",
        "createdAt": "2026-08-12T12:05:35Z",
        "updatedAt": "2026-08-13T07:18:32Z",
        "timestamp": "2026-08-13T07:18:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll3",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "Avogar",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114475",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113289 to 26.5: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114475",
        "createdAt": "2026-08-12T12:06:09Z",
        "updatedAt": "2026-08-13T17:59:14Z",
        "timestamp": "2026-08-13T17:59:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll3",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "Avogar",
          "groeneai"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114476",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113289 to 26.6: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114476",
        "createdAt": "2026-08-12T12:06:37Z",
        "updatedAt": "2026-08-13T11:32:09Z",
        "timestamp": "2026-08-13T11:32:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll3",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "Avogar",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114477",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allow to create table without providing credentials if catalog specified",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow to create table without providing credentials if catalog specified. Earlier the query to create table looked like `create table catalog.table enigine = IcebergS3(url, key1, key2)`, but now we can create it via `create table catalog.table enigine = IcebergS3`",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114477",
        "createdAt": "2026-08-12T12:06:45Z",
        "updatedAt": "2026-08-13T17:28:10Z",
        "timestamp": "2026-08-13T17:28:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "scanhex12",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114478",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113289 to 26.7: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114478",
        "createdAt": "2026-08-12T12:07:07Z",
        "updatedAt": "2026-08-13T15:38:46Z",
        "timestamp": "2026-08-13T15:38:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll3",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "Avogar",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114479",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Enable reading in reverse order with FINAL for ReplacingMergeTree",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/58035 Related: https://github.com/ClickHouse/ClickHouse/pull/58361 Related: https://github.com/ClickHouse/ClickHouse/pull/111609 When a query with `FINAL` sorts in reverse order of the sorting key, for example `ORDER BY key DESC LIMIT n`, the read-in-order optimization was disabled and the query read the whole table. It now applies for `ReplacingMergeTree`. `ReplacingSortedAlgorithm` learns a `read_in_reverse` mode. A row with a strictly higher version always replaces the selected one; among rows with equal (or absent) versions, the previously selected row is kept unless the current row comes from a newer data part. That mirrors the \"last written row wins\" rule of the direct reading order, because in the reverse reading order rows within one part arrive backwards while the parts still arrive from the oldest to the newest one. This is what makes the reverse read select the same row of a duplicate key group as a direct read, which is the correctness concern that stopped https://github.com/ClickHouse/ClickHouse/pull/58361. The other engines keep the previous behavior, since their merging algorithms depend on the direct order of rows: the sequence of sign rows in `CollapsingMergeTree`, the order of rows fed to order-dependent aggregate functions in `AggregatingMergeTree`, and so on. On a 110 million row `ReplacingMergeTree` table with two overlapping parts, `SELECT x FROM t FINAL ORDER BY x DESC LIMIT 1` (measured on a `RelWithDebInfo` build): | | before | after | |---|---|---| | Rows read | 110.1 million | 1.71 million | | Elapsed | 4.5 s | 0.22 s | | Peak memory | 41 MB | 23 MB | Trade-off: as with the already existing direct-order in-order reads with `FINAL`, an in-order plan disables vertical `FINAL` and the splitting of parts ranges into intersecting and non-intersecting ones. A query that reads the full result with `ORDER BY key DESC` and no small `LIMIT` may therefore become slower on a wide or mostly merged table. The new setting `optimize_read_in_reverse_order_final` (default enabled) turns the optimization off, and `compatibility` with an earlier version restores the previous plans. Out of scope, to keep this change reviewable: `Merge` tables over `ReplacingMergeTree`, and the interaction with `topKThroughJoin`, which keeps its own optimization for `... FINAL LEFT JOIN ... ORDER BY key DESC LIMIT n`. Both can follow separately. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable the read-in-order optimization for queries with the `FINAL` modifier that sort in reverse order of the sorting key on `ReplacingMergeTree` tables, so that queries such as `SELECT ... FROM t FINAL ORDER BY key DESC LIMIT n` read only the relevant tail of the data instead of the whole table. Can be disabled with the new setting `optimize_read_in_reverse_order_final`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114479",
        "createdAt": "2026-08-12T12:07:10Z",
        "updatedAt": "2026-08-13T14:33:14Z",
        "timestamp": "2026-08-13T14:33:14Z",
        "metrics": {
          "reactions": 1,
          "comments": 3
        },
        "labels": [
          "pr-performance"
        ],
        "author": "cwurm",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114484",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Document that PREWHERE filters one join input before the JOIN",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required for a documentation change. ### Description `prewhere.mdx` did not mention `JOIN` at all. This is the documentation KochetovNicolai asked for when he closed issue 89097 as not-a-bug: \"We need to document this.\" One `<Note>`, next to the existing note that documents the same class of fact for `FINAL`. It states the rule and gives the equivalent explicit spelling as a filtered subquery. The `SELECT` clause list is annotated to agree with it. Measured on the example as it appears on the page: `PREWHERE b.y > 50` returns `[1,2,3,4]`, the filtered-subquery spelling the same, the `WHERE` spelling `[1,4]`. The `WHERE` result is unchanged whether `query_plan_filter_push_down` is on or off, which is why the note credits `WHERE` to the join result rather than to a fixed position, per PedroTadim's review. The note claims nothing about individual join kinds or strictnesses, since `FULL JOIN` and `ASOF INNER JOIN` both differ too, and it attributes the fill value to [join_use_nulls](https://clickhouse.com/docs/reference/settings/session-settings/join#join_use_nulls) rather than naming one. Related: https://github.com/ClickHouse/ClickHouse/issues/89097 Related: https://github.com/ClickHouse/ClickHouse/issues/114206 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1341` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114484",
        "createdAt": "2026-08-12T13:01:26Z",
        "updatedAt": "2026-08-13T16:41:08Z",
        "timestamp": "2026-08-13T16:41:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-documentation",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114490",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "ci: remove stale cache status files",
        "text": "Though note that the server was still active, not sure why (since it got SIGTRAP, maybe some deadlock in sanitizers build). Fixes: https://pastila.nl/?004f7797/8607b82ba8c324e4f5081e6c747ad583#ad4mwd7GRSHyUgNa46pJHw==GCM ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1327` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114490",
        "createdAt": "2026-08-12T14:22:41Z",
        "updatedAt": "2026-08-13T13:35:08Z",
        "timestamp": "2026-08-13T13:35:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "azat",
        "state": "closed",
        "assignees": [
          "maxknv"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114492",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113484 to 26.6: Push down plan level constants from joins",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113484 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114488 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31605886263/job/94144581099) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.33` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114492",
        "createdAt": "2026-08-12T14:35:23Z",
        "updatedAt": "2026-08-13T14:31:05Z",
        "timestamp": "2026-08-13T14:31:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse",
        "state": "closed",
        "assignees": [
          "diegomestre2",
          "vdimir"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114495",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Targeted tests: skip previously failed tests that no longer exist",
        "text": "The targeted functional-test jobs replay tests that failed in previous CI runs of the PR (a 30-day CIDB window). When such a test is deleted or renamed on master in the meantime — e.g. a flaky test removed instead of deflaked — its name still comes back from CIDB, `clickhouse-test` matches zero tests, prints `No tests were run.` and exits with code 1, turning the whole job red for weeks on every PR that ever saw the test fail. The same failure mode existed for orphan data files and was fixed in #104097; this is the deleted-test variant. Seen on #110958: `Stateless tests (arm_asan_ubsan, targeted)` kept replaying `04648_geohashes_in_box_cancellation`, which was deleted from master on 2026-08-09 (d5bc95515d9e \"Remove recently added cancellation tests, they are flaky\"). CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110958&sha=7acb5ba0dddf5862fe1cd118b17899a272708202&name_0=PR&name_1=Stateless%20tests%20%28arm_asan_ubsan%2C%20targeted%29 The fix filters the CIDB result in `Targeting.get_previously_failed_tests` to tests that still exist in the checkout: a file with a known test extension under `tests/queries/0_stateless/` for stateless jobs, the test directory under `tests/integration/` for integration jobs. Skipped names are logged, not silently dropped. Unit tests added in `ci/tests/test_find_tests.py`. Related: https://github.com/ClickHouse/ClickHouse/pull/110958 Related: https://github.com/ClickHouse/ClickHouse/pull/104097 ### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1290` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114495",
        "createdAt": "2026-08-12T14:44:06Z",
        "updatedAt": "2026-08-13T00:19:18Z",
        "timestamp": "2026-08-13T00:19:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114496",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not rewrite arrayExists to has when the element and needle string types differ",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `arrayExists(x -> x = needle, arr)` returning a wrong result, or failing with `TOO_LARGE_STRING_SIZE`, when the array element and the needle are different string types (for example a `String` needle against an `Array(FixedString(N))` element). The `optimize_rewrite_array_exists_to_has` optimization, enabled by default, rewrote such a call to `has`, which does not compare zero-padded the way `=` does. ### Description `optimize_rewrite_array_exists_to_has` (on by default) rewrites `arrayExists(x -> x = c, arr)` into `has(arr, c)`. The two do not agree for a string-family pair whose types differ, so the same query returns different answers depending on the setting. On released 26.7.4.21 and on master, with `v = ['V0']` of type `Array(FixedString(3))`, `arrayExists(x -> x = 'V0\\0', v)` is `1` with the setting off and `0` with it on, while `toFixedString('V0', 3) = 'V0\\0'` is `1`. Over an `Array(LowCardinality(FixedString(N)))` element the rewrite turns a working query into `Code: 131 TOO_LARGE_STRING_SIZE`, rejecting valid input. <details> <summary>Measured on 26.7.4.21 and master</summary> | query on `v = ['V0']`, `Array(FixedString(3))` | setting off | setting on (default) | |---|---|---| | `arrayExists(x -> x = 'V0\\0', v)` | 1 | **0** | | `arrayExists(x -> x = 'V0\\0\\0', v)` | 1 | **0** | | `arrayExists(x -> 'V0\\0' = x, v)` | 1 | **0** | | `arrayExists(x -> x = 'V0abc', v)` (control) | 0 | 0 | | `arrayExists(x -> x = 'V0', v)` (control) | 1 | 1 | | `arrayExists(x -> x = tuple('V0\\0'), v)` on `Array(Tuple(FixedString(3)))` | 1 | **0** | | `arrayExists(x -> x = toFixedString('V0',4), [toFixedString('V0',3)])` | 1 | **0** | | `arrayExists(x -> x = 'V0\\0\\0', v)` on `Array(LowCardinality(FixedString(3)))` | 1 | **Code: 131** | | `arrayExists(x -> x = ['V0\\0'], [[toFixedString('V0',3)]])` | 0 | **1** | | `arrayExists(x -> x = map('k','V0\\0'), [map('k',toFixedString('V0',3))])` | 0 | **1** | </details> Root cause: the pass admits the rewrite whenever the element and needle have a common supertype. For `FixedString(N)` plus `String` that supertype exists, but reaching it casts the element to `String`, stripping its trailing NULs, while `equals` compares the pair zero-padded. A supertype is not sufficient for interchangeability, which is why this pass already declines for `NULL`, and `HasToInPass` for Date, Enum, IPv4 and floats. The fix declines the rewrite when the element and needle are both in the string family but not the same type, recursing through `Tuple`, `Array` and `Map` members because the divergence reproduces at any nesting depth. Over a constant array the container cases diverge the other way (`0` off, `1` on): there `has` compares raw `Field`s, which are not padded either. Identical types cannot diverge, so `String`/`String` and equal-width `FixedString` pairs keep the optimization, as do numeric, Date and Enum elements. It takes no position on what `has` should mean here; that belongs to #112953. Validated by an A/B over two binaries: every divergent cell goes from disagreeing to agreeing, the controls do not move, and a 298-test array/membership/FixedString sweep shows no regression against pristine master. The test's pinned direct `has`/`indexOf`/`countEqual` row is today's behaviour, held as a control; it moves when #112953 lands.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114496",
        "createdAt": "2026-08-12T14:44:50Z",
        "updatedAt": "2026-08-13T06:41:13Z",
        "timestamp": "2026-08-13T06:41:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114504",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Imply make_distributed_plan from distributed_plan_workers_num",
        "text": "Leasing Stateless Workers only makes sense for a distributed query plan, so `distributed_plan_workers_num` no longer has to be paired with `make_distributed_plan`: a non-zero Worker count enables the plan on its own. An explicit `make_distributed_plan` still wins if provided. The implication is applied before the adjustments added in https://github.com/ClickHouse/ClickHouse/pull/112463, so an implied plan still turns off the features it does not support yet. Closes: https://github.com/ClickHouse/ClickHouse/issues/114501 Related: https://github.com/ClickHouse/ClickHouse/pull/112463 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Setting `distributed_plan_workers_num` to a non-zero value now enables `make_distributed_plan` automatically, unless that setting is set explicitly.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114504",
        "createdAt": "2026-08-12T15:12:42Z",
        "updatedAt": "2026-08-13T08:14:16Z",
        "timestamp": "2026-08-13T08:14:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "andreev-io",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114505",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject duplicate names in CTE and table expression column alias lists",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a bug where the query gave incorrect result on duplicated column alias name in CTE and table expression (must be rejected). Closes https://github.com/ClickHouse/ClickHouse/issues/114427",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114505",
        "timestamp": "2026-08-12T20:18:30Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yariks5s",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114507",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Column IDs for MergeTree: the data structure",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ## Background First part of [#99754](https://github.com/ClickHouse/ClickHouse/pull/99754) This one adds the data structure — `ColumnId`, `ColumnIdMapping`, its durable store, and the schema fields that carry a mapping: - `ColumnId`: a strong typed struct to replace the `std::string column_name` in most places - plug in `NameAndTypePair` & `ColumnDescription`, which are used in table schema - `ColumnIdMapping`: in-memory structure to store the `name <-> id` mapping - plug in `StorageInMemoryMetadata`: so that it is captured atomically with the schema - `ColumnIdMappingStore`: persistence store for the mapping ## Not in this PR Stamping ids on columns, id-keyed reads and writes, the experimental setting that turns the feature on, metadata-only `RENAME` / `DROP COLUMN`, and `ReplicatedMergeTree` / `SharedMergeTree` support. Each is planned as a follow-up on top of this one. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114507",
        "createdAt": "2026-08-12T15:18:12Z",
        "updatedAt": "2026-08-13T11:14:51Z",
        "timestamp": "2026-08-13T11:14:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "murphy-4o",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114508",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: add ProbeDeck SQL client integration",
        "text": "Adds a community-maintained ProbeDeck integration guide for iOS and iPadOS. The guide covers ClickHouse Cloud and self-hosted setup, database authentication, mTLS, SSH bastions, monitoring sources, a reproducible query, limits, and troubleshooting. It also adds ProbeDeck to the SQL client navigation and overview. The screenshots use deterministic synthetic data and contain no production credentials or infrastructure. Related: https://github.com/ClickHouse/clickhouse-docs/pull/6625 ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114508",
        "createdAt": "2026-08-12T15:25:38Z",
        "updatedAt": "2026-08-13T15:35:48Z",
        "timestamp": "2026-08-13T15:35:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-documentation",
          "manual approve",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "chamav",
        "state": "closed",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114509",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: Remove private preview banner and stale private preview references fr…",
        "text": "…om SCIM docs SCIM for CHC is GA <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114509",
        "createdAt": "2026-08-12T15:45:37Z",
        "updatedAt": "2026-08-13T13:39:14Z",
        "timestamp": "2026-08-13T13:39:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-documentation",
          "can be tested"
        ],
        "author": "ohlookadollar",
        "state": "open",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114510",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: Improve supported regions page layout",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Improves the [supported cloud regions page](https://clickhouse.com/docs/products/cloud/reference/supported-regions) by grouping regions into provider tabs and presenting Private Region, HIPAA, and PCI availability as comparison flags. This makes the information easier to scan without duplicating each provider name or maintaining separate compliance inventories. Linear issue: [DOC-967](https://linear.app/clickhouse/issue/DOC-967/improve-supported-regions-page-layout) ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved the layout of the supported cloud regions reference page. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1321` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114510",
        "createdAt": "2026-08-12T15:51:33Z",
        "updatedAt": "2026-08-13T12:39:06Z",
        "timestamp": "2026-08-13T12:39:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-documentation",
          "pr-synced-to-cloud"
        ],
        "author": "dhtclk",
        "state": "closed",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114518",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Raise a catchable error instead of LOGICAL_ERROR on spec-violating Iceberg metadata",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114487 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reading an Iceberg table whose metadata describes a schema evolution that the Iceberg specification forbids no longer aborts the server. Four spec violations in `IcebergSchemaProcessor` were reported as `LOGICAL_ERROR`, which is treated as a failed assertion, or reached a fatal assertion unchecked; they now raise `ICEBERG_SPECIFICATION_VIOLATION`. ### Description Validators in `IcebergSchemaProcessor` reject genuinely invalid Iceberg metadata, but did so with `ErrorCodes::LOGICAL_ERROR`, which `Exception::handleErrorCode` treats as a failed assertion (`src/Common/Exception.cpp:112-116`). Metadata content alone took the server down: no `ALTER`, no write path, a corrupt or crafted `.metadata.json` is enough. The three existing detections are correct, so only their error code changes. A fourth shape had no check. Sites fixed, all reachable from metadata content: - `getSchemaTransformationDag`: an old struct/list/map field becomes a primitive. Its message named the opposite direction and is corrected. - `getSchemaTransformationDag`: a new field with no old counterpart is `required` with no default. The frame in the reported stack. - `registerSnapshotWithSchemaId`: one `snapshot-id` bound to two `schema-id`s. Not named in the issue, but `IcebergMetadata.cpp:384-389` registers every entry of the `snapshots` array, both values read straight out of the JSON. - `getSchemaTransformationDag`: the mirror shape, an old primitive becoming a struct, list or map. The complex branch keyed off the new field's type alone, so it built `EvolutionFunctionStruct` and aborted in its `lazyInitialize` `chassert` instead of reporting anything. Now rejected before the transform is built. This file uses both conventions. `Utils.cpp:1196-1199` picks 743 over `BAD_ARGUMENTS` for a versioned specification rule broken by a file being read, which is what all four sites are, matching `SchemaProcessor.cpp:419`; `MetadataGenerator.cpp:452` keeps `BAD_ARGUMENTS` for the write path. The exposure is wider than the issue states: `abort_on_logical_error` (`ServerSettings.cpp:1433`) is linked by `tests/config/install.sh:244-246` for every non-fast-test stateless run, so release-build CI aborts here too. Not touched: `SchemaProcessor.cpp:189/194` fire when stringifying an already-parsed object, a genuine internal invariant. Manifest-file sites need their own proof. A new stateless test covers all four sites. On master it kills the server (`[ FAIL ] Reason: server died`); with the fix each returns `Code: 743`, and 50/50 randomized runs pass.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114518",
        "createdAt": "2026-08-12T16:14:30Z",
        "updatedAt": "2026-08-13T09:05:31Z",
        "timestamp": "2026-08-13T09:05:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114520",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Require full field consumption on every Regexp escaping rule",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/108091 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the `Regexp` input format silently discarding the unparsed rest of a matched field, which stored a truncated or fabricated value instead of reporting an error: `v=-1` into `UInt64` read `0`, and `v=2020-01-01junk` into `Date` read `1975-07-14` under the `JSON` rule. Malformed fields are now rejected under the `Escaped`, `CSV` and `JSON` rules, matching the formats those rules are documented as behaving like. Under `Raw`, a matched field containing a tab is now read as a whole instead of being truncated at the tab. ### Description **The problem.** A matched capture group is one complete field, and `Regexp` has no delimiter after it, so when a type's parser stopped early the bytes it left behind were discarded and the partial parse reported as success. `v=-1` into `UInt64` gave `0`, `v=1abc` gave `1`, and `v=2020-01-01junk` into `Date` under the `JSON` rule gave `1975-07-14`, which is not even a prefix of the input. `TSV`, `TSKV`, `CSV`, `JSONEachRow` and `CustomSeparated` all reject these same bytes, because there the leftover fails the delimiter that follows. **Why the existing check was not enough.** #108091 added this check for the `Quoted` rule only; it now applies to every rule. `CSV` and `JSON` first skip trailing whitespace, because the formats they mirror accept it, so `v=1 ` still reads `1` while `v=1 abc` is rejected. **A second, more reachable mechanism.** `Raw` is the default rule, so it needs no setting at all, and it fails differently: its reader stops at a tab, so `v=abc<TAB>junk` into `String` read `abc`. A tab is legal inside a `String`, so rejecting that field would be a new bug; it is now read as a whole instead, while a tab that cannot belong to the value (`v=1<TAB>junk` into `UInt64`) is rejected. A field equal to the null representation keeps the old reader wherever a null-aware one is selected, so `Raw`'s null tokens are unchanged. The check stays inside `Regexp`: the two other formats sharing this deserialization helper have delimiters, so for them a leftover is normal. Validation: run unchanged against unpatched master the new test fails, nearly every error assertion reporting a missing error; the exceptions are the pre-existing `Quoted` and `Raw` controls. It passes 50/50 under randomized settings. Also noticed, not changed here: `Values` accepts `v=1 ` while `Regexp`+`Quoted` rejects it, a divergence predating this PR.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114520",
        "createdAt": "2026-08-12T16:38:11Z",
        "updatedAt": "2026-08-13T02:12:21Z",
        "timestamp": "2026-08-13T02:12:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114521",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix Iceberg query failure after MODIFY COLUMN to Nullable",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/85029 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `LOGICAL_ERROR` when filtering an Iceberg column that `ALTER TABLE ... MODIFY COLUMN` made `Nullable`. Closes #85029. ### Description Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/85029#issuecomment-5264517370; his analysis comment there carries the reproducer. ```sql CREATE TABLE t (id Int64, s String) ENGINE = IcebergLocal('lake/'); INSERT INTO t SELECT number, toString(number) FROM numbers(5); ALTER TABLE t MODIFY COLUMN id Nullable(Int64); SELECT * FROM t WHERE id > 3; -- Unexpected return type from greater. Expected Nullable(UInt8). Got UInt8 ``` Root cause: in `IcebergSchemaProcessor::getSchemaTransformationDag` the node emitted for an existing field id is chosen from the Iceberg `type` string alone. Nullability lives in the separate `required` key, which `getFieldType` folds into ClickHouse nullability. `MODIFY COLUMN ... Nullable` flips only `required`, leaving type string and name unchanged, so neither the cast nor the alias branch fired and the transform forwarded the input node built from the old flag. Its output header said `Int64` while consumers were told `Nullable(Int64)`. Only PREWHERE observes this: the fallback `FilterTransform` is built against that header, so `greater` yields `UInt8` where the planner expects `Nullable(UInt8)`. Without the pushdown the filter sits above the reader, past the cast in `getColumnFromBlock`, so `SELECT *` and `count()` were correct. The fix emits the missing cast, gated twice: inside the equal-type comparison, leaving the `allowPrimitiveTypeConversion` allow-list and spec-prohibited conversions untouched; and only for the legal relaxation of required to optional, so an externally written reverse pair keeps its current passthrough instead of gaining a `Nullable(T) -> T` cast that would reject rows holding NULL. Both excluded directions measure byte-identical. Validated on unpatched and patched binaries for `Int64`, `Int32`, `String`, `Float64`, `Date32` and `DateTime64`, with a rename composed on top and mixed pre/post-`ALTER` files. The new test is 50/50 green; the Iceberg and Delta suites show no failure absent unpatched.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114521",
        "createdAt": "2026-08-12T16:38:44Z",
        "updatedAt": "2026-08-13T01:42:40Z",
        "timestamp": "2026-08-13T01:42:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114522",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "QueryRunner follow-up: do not occupy threads eagerly",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): QueryRunner tables now start worker threads on demand and release them once idle, instead of occupying threads for the table's whole lifetime.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114522",
        "createdAt": "2026-08-12T16:42:37Z",
        "updatedAt": "2026-08-13T17:27:23Z",
        "timestamp": "2026-08-13T17:27:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "mstetsyuk",
        "state": "open",
        "assignees": [
          "azat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114523",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix wrong results of the partial aggregation strategy in distributed query plans",
        "text": "The partial+merge aggregation strategy of `make_distributed_plan` rewrites a final aggregation into a partial `AggregatingStep` plus a memory-efficient `MergingAggregatedStep`. The memory-efficient merge consumes each input as a stream of two-level buckets in ascending order, but the rewrite kept `should_produce_results_in_order_of_bucket_number = false` on the partial step, so its multi-stream output was combined into the single exchange stream in arbitrary order. When a bucket arrived after the merge had already emitted it, the merge emitted it a second time: duplicated `GROUP BY` keys with split aggregate states (for ClickBench Q08, a wrong `COUNT(DISTINCT UserID)`). The fix, one commit each: - Make the row count estimation look through `LogicalExchangeStep`. The `#if CLICKHOUSE_CLOUD` guard around it was needed when the exchange steps existed only in the private repo; since they are in the public repo, the guard only made every aggregation over a distributed read fall back to the Shuffle strategy in non-cloud builds (and masked this bug there). - Validate the bucket delivery order in `GroupingAggregatedTransform`: an unannounced late bucket now fails with a `LOGICAL_ERROR` exception instead of silently duplicating groups. - Wait for delayed buckets in `GroupingAggregatedTransform` when the consumer needs the output in bucket order, and release the announcements of finished inputs. Covered by a gtest. - Build the partial step with `should_produce_results_in_order_of_bucket_number = true` when the merge is memory-efficient (`AggregatingStep::cloneAsPartial`). The partial aggregation then produces a single bucket-ordered stream per worker, the same contract the classic distributed path establishes on shards. Covered by a stateless test that reproduces the duplicated keys on the code without the fix (about 80% of single runs, 8 repetitions, 50k rows, ~0.5s). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix duplicated `GROUP BY` keys in the partial aggregation strategy of `make_distributed_plan`, and allow this strategy for aggregations over a distributed read. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114523",
        "createdAt": "2026-08-12T16:51:05Z",
        "updatedAt": "2026-08-13T14:08:11Z",
        "timestamp": "2026-08-13T14:08:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "davenger",
        "state": "open",
        "assignees": [
          "nickitat"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114525",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Optimize merges of the text index",
        "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved performance of merges of text indexes. ### Additional context A few optimizations: - The main one: reducing overhead on deserialization of embedded and small postings caused by the allocation of the bitmap - Removed unneeded conversion to roaring bitmap on build of the output posting list - Used specialized sort cursor and batch sorting strategy for merging of text index segments",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114525",
        "createdAt": "2026-08-12T17:15:12Z",
        "updatedAt": "2026-08-13T17:52:32Z",
        "timestamp": "2026-08-13T17:52:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-performance"
        ],
        "author": "CurtizJ",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114527",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix flaky `test_url_reconnect`",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110130 `test_redirect_url_storage/test.py::test_url_reconnect` blocks the outgoing connections to the HDFS datanode web port with a silent `DROP` and heals the network only after the client has demonstrably hit a connect timeout, waiting up to 30 seconds for a fresh `connect timed out` line in the server log. That budget is not always enough, because the query does not necessarily open a new connection at all: the earlier tests in the module leave keep-alive connections to `hdfs1:50075` in the HTTP connection pool, and a silent `DROP` kills an already established connection only by the receive timeout - `http_receive_timeout`, 30 seconds by default - rather than by the one-second connect timeout. The first try then dies after 30 seconds with a bare `Timeout`, and the first `connect timed out` line appears only on the next try, about 31 seconds into the query, just past the budget: ``` 16:57:35.001 executeQuery: select sum(cityHash64(id)) from url('http://hdfs1:50075/... 16:58:05.013 ReadWriteBufferFromHTTP: ... Error: Timeout. Failed at try 2/10. 16:58:06.115 ReadWriteBufferFromHTTP: ... Error: Timeout: connect timed out: 172.16.1.2:50075. Failed at try 3/10. ``` The fix drops the connection cache before installing the rule, so the query has to open a fresh connection and the first try fails by the connect timeout, as the test expects. The wait budget is also raised to 60 seconds: the loop exits as soon as the line shows up, so a generous budget costs nothing on a healthy run. Seen in an unrelated pull request: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110130&sha=7acc19de7fb0c97f7ea0fb7c6c234930d188ac66&name_0=PR&name_1=Integration%20tests%20%28amd_msan%2C%205%2F8%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1294` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114527",
        "createdAt": "2026-08-12T17:49:11Z",
        "updatedAt": "2026-08-13T01:17:29Z",
        "timestamp": "2026-08-13T01:17:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114528",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #102033 to 26.5: Fix LOGICAL_ERROR crash in IcebergMetadata::iterate when datalake_table_state is missing",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/102033 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114385 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31624414958/job/94207002824) <!-- ch-version-info:start --> ### Version info - Merged into: `26.5.7.46` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114528",
        "createdAt": "2026-08-12T18:07:36Z",
        "updatedAt": "2026-08-13T01:35:58Z",
        "timestamp": "2026-08-13T01:35:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll",
        "state": "closed",
        "assignees": [
          "alexey-milovidov",
          "SmitaRKulkarni",
          "groeneai"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114529",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #102033 to 26.6: Fix LOGICAL_ERROR crash in IcebergMetadata::iterate when datalake_table_state is missing",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/102033 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114386 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31624414958/job/94207002824)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114529",
        "timestamp": "2026-08-12T20:49:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114530",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject a Parquet offset index whose first page does not start at row 0",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/114464 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114464 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Parquet files whose offset index does not start at row 0. Such a file could make the native Parquet reader return wrong rows, or report `LOGICAL_ERROR` instead of `INCORRECT_DATA`. The offset index is now validated when it is read. ### Description Reported by @ PedroTadim in #114464 ([his request to take it](https://github.com/ClickHouse/ClickHouse/issues/114464#issuecomment-5265506353)): a single flipped bit in a Parquet offset index made the v3 reader raise `Row passes filters but its page was not selected for reading. This is a bug.` That `LOGICAL_ERROR` aborts on debug and sanitizer builds, and in release blames ClickHouse for a defect in the file. Reachable from untrusted file content on defaults. Root cause: `decodeOffsetIndex`, the only place the offset index is deserialized, validated byte ranges, monotonicity and `first_row_index < num_rows`, but never that the *first* page starts at row 0, as the spec requires since a chunk covers every row of its row group. Page ends come from the next page's `first_row_index` while the row sweep covers `[0, num_rows)`, so a nonzero anchor leaves rows `[0, anchor)` described by no page. The reported abort is the mildest of three symptoms. For a V1 data page in an array column `num_rows_in_page` stays unset, so the row-count cross-check is skipped, the corrupt anchor seeds the row cursor and the read returns wrong rows with no error: on a ClickHouse-written file, `WHERE id = 10` returned the row-6 payload. The anchor is now validated in `decodeOffsetIndex`, which runs before every consumer, so all three become a clear `INCORRECT_DATA`. `Page doesn't contain requested row` is reclassified too (a page the offset index lists can turn out to be an index or dictionary page, which the reader skips). `Row passes filters ...` stays `LOGICAL_ERROR`: with a validated anchor it is unreachable from file content, so it remains the page-selection tripwire @ PedroTadim asked to keep. Such a file is now rejected rather than read, but it already aborts or returns wrong rows. ClickHouse's writer always anchors page 0 at row 0, and all 124 offset-index fixtures in `tests/` conform.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114530",
        "createdAt": "2026-08-12T18:14:39Z",
        "updatedAt": "2026-08-13T09:24:07Z",
        "timestamp": "2026-08-13T09:24:07Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114531",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject a lossy codec on columns backing keys and indexes",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114406 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A lossy codec such as `SZ3` is now rejected at DDL time on any column that backs the sorting key, primary key, partition key, a secondary index or the unique key. Because such a codec does not return the value that was written, a merged part could be stored out of physical order and index analysis would skip rows matching the query. Column statistics are no longer built for, or used to prune parts by, a lossily compressed column, which returned too few rows on default settings. Existing tables stay loadable. ### Description `ORDER BY` makes a physical promise: rows are stored inside a part in sorting-key order and the primary index samples those stored values. A lossy codec breaks `read(write(v)) == v`, so the merge sorts pre-compression values while the part stores post-compression ones. `SZ3` is not monotonic, so the stored sequence is not sorted. The same applies to anything else computed from the pre-write block: skip-index granules, the unique-key index, partition values. The issue's reproducer gives `min(i - prev) = -0.2436889648437699` after `OPTIMIZE TABLE t FINAL` and 0 after one `INSERT`: the disorder appears only at the merge, where a debug build aborts in `CheckSortedTransform`. Lossiness comes from the existing `ICompressionCodec::isLossyCompression()`, not a codec-name list. `CompressionCodecMultiple` did not override it, so a stacked `CODEC(SZ3(...), LZ4)` reported itself lossless; it now ORs over its children. Classification is per serialized substream, mirroring `MergeTreeDataPartWriterWide::addStreams`, so `ORDER BY arr.size0` stays allowed while `arraySum(arr)` is not. The check sits at the two user-facing entry points: `registerStorageMergeTree` for CREATE and full-definition ATTACH, `checkAlterIsPossible` for ALTER on the initiating execution, following `4a29ef847411256`. A replica replaying a durable DDL entry is not re-checked, which would wedge its DDL worker. Column statistics, found during review, needed a different remedy: they are built pre-write too, but `auto_statistics_types` attaches `basic` to every numeric column, so a DDL rejection would ban the codec on any table that does not opt out. Instead none are built for such a column, and the pruner ignores any an earlier version wrote. With no setting changed, a predicate matching 2 rows returned 0. `system.columns` still lists them, as the two settings' descriptions now note. Seen in CI on one AST-fuzzer run, on #113575: [report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113575&sha=84df30db17165b65e3c78bc11513ed26534187d7&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%29). <details> <summary>Related cases deliberately left out of this PR</summary> - A **projection** over a lossily compressed column answers differently than the base part: `sum(i)` is `249957173.79` read from the base part and `249998750` via the projection. The mechanism is a separate post-write consumer, which writes the base part lossily and then computes the projection from the unchanged pre-compression block, so it follows in its own PR rather than being folded in here. - A legacy table can still gain an implicit minmax index over such a column through server configuration on the load path. Reachable, but a boundary sweep over 300 values produced no wrong result. - A part written by an earlier version keeps its statistics if the codec is then replaced by a lossless one without any merge or mutation rewriting that part. Deciding this needs per-part codec provenance, which parts do not record. `ALTER TABLE ... MATERIALIZE STATISTICS` or `OPTIMIZE TABLE ... FINAL` clears it. - `ORDER BY length(arr)` is now rejected although its value is exact. A key expression records which columns it needs and not which of their streams, so at this layer `length(arr)` and `arraySum(arr)` are indistinguishable, and `arraySum` is a genuine wrong-results carrier. The syntactic form `ORDER BY arr.size0` stays allowed because it names the stream. - Alias-dependent index rebuilds look incorrect independently of codecs: an alias-only `MODIFY COLUMN` requires no mutation while the index expression is rebuilt, so existing parts keep a same-named index built from the old expression. No claim is made about it here. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114531",
        "createdAt": "2026-08-12T18:16:56Z",
        "updatedAt": "2026-08-13T17:58:38Z",
        "timestamp": "2026-08-13T17:58:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114533",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix wrong results for non-boolean conditions taken out of `and`",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112236 `and` implicitly converts its arguments to booleans, so any non-zero value is true. When a query plan optimization takes a part of a conjunction away and a single conjunct is left as the new predicate, that conjunct was converted with a cast to the type of the original predicate. A cast is not a boolean conversion: it maps values like 256 or 0.1 to 0, so rows whose condition value is a non-zero multiple of 256 silently disappeared. ```sql CREATE TABLE t (id UInt32) ENGINE = MergeTree ORDER BY id; INSERT INTO t SELECT number FROM numbers(600); SELECT count() FROM t AS l LEFT JOIN t AS r ON l.id = r.id WHERE r.id AND l.id = r.id; ``` returned 597 instead of 599, the rows with `id = 256` and `id = 512` were dropped. `mergeFilterIntoJoinCondition` moves `l.id = r.id` into the JOIN and leaves `CAST(r.id, 'UInt8')` as the filter: ``` Filter column: CAST(id AS UInt8) ``` The same happens in `ActionsDAG::removeUnusedConjunctions` when a conjunct is pushed down and the filter column is still needed in the result. That one is reachable without a JOIN, and the value of the condition was wrong there as well (the raw value instead of a boolean): ```sql SELECT count(), sum(f) FROM ( SELECT id, (id != 1000 AND s) AS f FROM (SELECT id, sum(id) AS s FROM t GROUP BY id) WHERE id != 1000 AND s ); ``` returned `597 597` instead of `599 599`. It only converted floating point types, now every non-boolean type is converted. Both places now wrap the remaining conjunct into `and(x, true)`, the same way `toBoolIfNeeded` does it in `JoinStepLogical.cpp`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results when a non-boolean condition, such as a bare integer column in `WHERE`, is left alone after the other conditions are merged into the JOIN condition or pushed down. Rows whose condition value was a non-zero multiple of 256 were skipped. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114533",
        "createdAt": "2026-08-12T18:40:57Z",
        "updatedAt": "2026-08-13T17:24:11Z",
        "timestamp": "2026-08-13T17:24:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "vdimir",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114534",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: fix absolute documentation component links",
        "text": "The settings explorers and documentation badges rendered root-relative links that already contained the production `/docs` mount. Mintlify prepended the mount again, producing broken `/docs/docs/...` destinations. Use fully qualified ClickHouse documentation URLs for the generated settings explorers and all localized badge variants. Example: https://clickhouse.com/docs/reference/settings/session-settings Related: https://github.com/ClickHouse/ClickHouse/pull/114432 ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed broken settings explorer and documentation badge links that resolved to `/docs/docs/...`. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1282` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114534",
        "timestamp": "2026-08-12T22:14:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-documentation",
          "pr-synced-to-cloud"
        ],
        "author": "Blargian",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114535",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Estimate compact-part input bytes from the part's measured ratio",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): todo",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114535",
        "timestamp": "2026-08-12T22:59:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "nickitat",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114536",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: add ClickStack guide for isolating read and write workloads",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/113284 ### Changelog category (leave one): - Documentation (changelog entry is not required) --- ClickStack's docs claim the ability to \"independently isolate read and write workloads with Warehouses\" in five places without documenting how. The only existing guidance (in `deployment/managed`, added in ClickHouse/clickhouse-docs#5452) covers one half of it: that the ClickStack UI binds to whichever Cloud service it is launched from. This adds a guide for the full topology and links it from the places that previously mentioned the capability without explaining it. **New page** — `clickstack/managing/isolating-read-write.mdx`: - Why isolate: read-only services run no background merges, so their compute is dedicated to queries; ingestion is insulated from expensive queries; each side scales independently against the ingest and query figures from the sizing model. - Recommended topology: one read-write service for ingestion, one read-only service for ClickStack, plus the planning constraints — the first service is always read-write, type is fixed at creation, and merge assignment crosses read-write services. - Setup steps: prepare the read-write service and ingestion user, add the read-only service, point the collector at the read-write endpoint, point the UI at read-only compute (managed and self-hosted), then verify the split with `system.query_log` — including the `all_groups.default` note, since `system` tables are per-service. - Where DDL and materialized views execute, and how alert evaluation follows the connection of the source it is attached to. - Advanced: separating merges from ingestion onto a dedicated merge service, flagged as requiring a support request, with the caveats that come with it (mutation tracking, TTL deletion, auto-idling, keeping queries off both read-write services). **Cross-links**: `managing/overview` (admin guides table), `managing/production` (new subsection), `managing/estimating-resources` (ties the ingest and query vCPU split to separate services), `deployment/managed` (extends the existing read-only-compute section), and the previously unlinked bullets in `overview` and `architecture`. `getting-started/managed.mdx` carries the same bullet but is deliberately left untouched — #112118 rewrites that section and already links the concept, so editing it here would only create a conflict.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114536",
        "createdAt": "2026-08-12T19:46:18Z",
        "updatedAt": "2026-08-13T13:59:25Z",
        "timestamp": "2026-08-13T13:59:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-documentation"
        ],
        "author": "andremm",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114537",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"Resolve the query status per call in functions that check for cancellation\"",
        "text": "Reverts ClickHouse/ClickHouse#113456 Needs a better solution",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114537",
        "createdAt": "2026-08-12T19:52:40Z",
        "updatedAt": "2026-08-13T00:56:46Z",
        "timestamp": "2026-08-13T00:56:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "PedroTadim",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114538",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject Delta Lake partition column absent from the table schema",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114462 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `LOGICAL_ERROR` when reading a Delta Lake table whose `metaData.partitionColumns` names a column that `metaData.schemaString` does not declare. Such metadata is now rejected with `BAD_ARGUMENTS` on the delta-kernel reader and `INCORRECT_DATA` on the legacy reader, whenever the table or its schema is read from a snapshot, not only when the query carries a predicate. ### Description Requested by @ PedroTadim in #114462. `partitionColumns` and `schemaString` are independent fields of externally supplied metadata and nothing cross-checked them, so a mismatch was reported as `LOGICAL_ERROR`, which aborts on debug and sanitizer builds. Paimon already threw `BAD_ARGUMENTS` here. - delta-kernel reader: validate in `TableSnapshot::initOrUpdateSchemaIfChanged`, where both values first exist, beside the existing empty-schema rejection. Previously an unfiltered `SELECT *` succeeded and only a predicate threw, since `PartitionPruner` is built only when a filter exists. That site keeps `LOGICAL_ERROR`, now a genuine internal invariant. - legacy reader: `LOGICAL_ERROR` -> `INCORRECT_DATA` where an `add` action's `partitionValues` names an undeclared key, in the JSON-log and checkpoint branches. It also validates `partitionColumns` where `metaData` is loaded, in both branches, so a snapshot with no `add` action to resolve is rejected too. That check compares logical names: a parsed legacy schema is keyed by column-mapping physical names, so comparing against it would reject well formed column-mapped tables. Iceberg is unaffected: it resolves partitions by numeric `source_id` field ids, not by name. Change-data-feed reads go through `TableChanges`, which resolves no partition name, and are out of scope. Behaviour change: such a table was readable without a predicate, and on the legacy reader also with no `add` action; it is now rejected for every query that reads it. It is malformed per the Delta protocol. One case is left alone, unchanged from master: on a table with an explicit column list a bare `count()` is answered from transaction-log row counts, which resolve no column name, so it is not rejected, while reading any column from it is. Validated with the new stateless test `04877`: every failure arm fails on a targeted revert of the line it covers, including a well formed column-mapped control that must keep reading. Also 50/50 randomized runs and the `delta_lake` suite.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114538",
        "createdAt": "2026-08-12T19:55:56Z",
        "updatedAt": "2026-08-13T07:50:11Z",
        "timestamp": "2026-08-13T07:50:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114539",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Enable the query condition cache for `ORDER BY ... LIMIT n` queries by default",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/111492 Related: https://github.com/ClickHouse/ClickHouse/pull/104478 Related: https://github.com/ClickHouse/ClickHouse/pull/110507 `use_query_condition_cache_for_top_k` was introduced by #111492 and defaulted to `false` as a precaution while the soundness of the query condition cache entries written by a TopK (`ORDER BY <column> LIMIT n`) read was being established. Such entries are partitioned by the TopK plan parameters and by a snapshot of the set of parts read, so only a query with the same plan over the same parts reuses them. The gate is no longer needed, so the default becomes `true` and `ORDER BY ... LIMIT n` queries use the query condition cache again. This effectively reverts #111492 by flipping its setting, rather than by removing it: the setting and every gating point it drives are kept, so the cache can still be kept out of TopK reads with `use_query_condition_cache_for_top_k = 0`. The tests that cover that configuration are kept too — they already pinned the setting explicitly rather than relying on the default — while the tests of the feature itself no longer have to enable it. `04628_query_condition_cache_topk_default_off` is renamed to `04628_query_condition_cache_topk_gate_off` since it no longer describes the default. The settings-history entry keeps `previous_value = false`, because the gate was backported to 26.7. `compatibility` with 26.7 or earlier therefore still turns the query condition cache off for TopK reads, while `compatibility = '26.8'` keeps it on; `04631_query_condition_cache_topk_compatibility` pins this. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The [query condition cache](/operations/query-condition-cache) is now enabled by default for queries that use the `ORDER BY <column> LIMIT n` (TopK) optimization. It can be turned off again with the setting `use_query_condition_cache_for_top_k`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114539",
        "createdAt": "2026-08-12T20:09:36Z",
        "updatedAt": "2026-08-13T14:31:39Z",
        "timestamp": "2026-08-13T14:31:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-performance",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114540",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Reject a Variant whose ORC branches read back as one type",
        "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/114169 Related: https://github.com/ClickHouse/ClickHouse/pull/110085 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/114169 The native ORC writer accepted a `Variant` whose branches map to *different* ORC types but which the reader parses back to the *same* type. The file was written, and reads fine with an explicit structure, but inference on it failed, breaking the round-trip: ``` $ clickhouse local -q \"SELECT NULL::Variant(String, Int128) FORMAT ORC\" | clickhouse local --input-format ORC -q \"SELECT * FROM table\" Code: 50. ORC union type 'uniontype<binary,string>' has branches with identical types (UNKNOWN_TYPE) ``` Root cause: the writer's duplicate-branch guard keyed its dedup map on the ORC type (`child_type->toString()`), while the reader keys on the ClickHouse type each branch parses back to, mapping both ORC `binary` and ORC `string` to `String`. Two different equivalence relations, so the writer's guard passed where the reader's fired. This keys the guard on a new `orcTypeDedupKey` helper, a recursive rendering of the ORC type that folds `BINARY` into `STRING` and sorts nested `UNION` children (a nested union reads back as a `Variant`, which sorts its branches, whereas ORC keeps them positional). `LIST`, `MAP` and `STRUCT` stay positional, matching `Array`, `Map` and `Tuple`. The writer never emits `CHAR`/`VARCHAR`, so that is the only collapse it can produce. The reader was left alone: remapping ORC `binary` would change the inferred type of every existing ORC `binary` column, and `uniontype<binary,string>` has no type to infer *to* anyway. Validated on a debug build: 16 measured carriers (the 6 reported scalar pairings, `FixedString`, and the same collapse through `Array`/`Tuple`/`Map` and nested unions) wrote and then failed inference before, and are rejected at write time after. 7 positive controls still round-trip. Unreleased feature, so nothing to break and no backport. The `.cpp` adds 26 lines of code, within the 50 @ PedroTadim authorized. Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/114169#issuecomment-5265686354",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114540",
        "createdAt": "2026-08-12T20:34:35Z",
        "updatedAt": "2026-08-13T07:57:35Z",
        "timestamp": "2026-08-13T07:57:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114541",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add ProfileEvents and CurrentMetrics for fiber stacks",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/73510 Related: https://github.com/ClickHouse/ClickHouse/issues/72968 Fibers are used for asynchronous communication with remote replicas, and a stack is the only thing a fiber consists of, so allocating the stack is the whole cost of creating a fiber. Nothing observed it: neither the number of allocations, nor the amount of memory they hold, nor the time they take. When distributed queries became slow because every fiber went through `mmap`, `mprotect` and `munmap` (https://github.com/ClickHouse/ClickHouse/issues/72968), it had to be found with a profiler. `FiberStack::allocate` and `FiberStack::deallocate` now account for: - profile events `FiberStackAllocs` and `FiberStackAllocBytes`; - profile events `FiberStackAllocNanoseconds` and `FiberStackFreeNanoseconds` — in nanoseconds, because a single allocation normally takes less than a microsecond and would be truncated to zero; - metrics `FiberStacks` and `FiberStackBytes`. This is the observability part of https://github.com/ClickHouse/ClickHouse/pull/73510, which was closed because the allocator it proposed became obsolete after https://github.com/ClickHouse/ClickHouse/pull/79147. For a two-shard query, it looks like this: ``` SELECT count() FROM remote('127.0.0.{1,2}', system.one) SETTINGS async_socket_for_remote = 1, prefer_localhost_replica = 0 allocs: 6 bytes: 1966080 -- 6 stacks of 320 KiB alloc_ns: 26393 free_ns: 10700 ``` ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added observability for the stacks of fibers, which are used for asynchronous communication with remote replicas: profile events `FiberStackAllocs`, `FiberStackAllocBytes`, `FiberStackAllocNanoseconds`, `FiberStackFreeNanoseconds`, and metrics `FiberStacks`, `FiberStackBytes`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114541",
        "createdAt": "2026-08-12T20:47:15Z",
        "updatedAt": "2026-08-13T01:57:40Z",
        "timestamp": "2026-08-13T01:57:40Z",
        "metrics": {
          "reactions": 1,
          "comments": 2
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114542",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: improve data warehousing diagram rendering",
        "text": "Replaces the shared data warehousing diagram with an opaque dark-canvas version that includes internal padding. This keeps the artwork consistent in light and dark themes and prevents edge cropping without CSS or component changes. The existing asset path remains valid for the English page and all generated locale pages. Validated in the local Mintlify preview in both light and dark themes. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Updated the data warehousing architecture diagram for consistent theme rendering. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1292` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114542",
        "createdAt": "2026-08-12T21:01:08Z",
        "updatedAt": "2026-08-13T01:07:00Z",
        "timestamp": "2026-08-13T01:07:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-documentation",
          "pr-synced-to-cloud"
        ],
        "author": "dhtclk",
        "state": "closed",
        "assignees": [
          "Blargian"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114543",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix spelling of 'prefetches' and 'prefetched' in documentation",
        "text": "Corrected spelling of 'prefetches' and 'prefetched' in multiple sections. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1353` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114543",
        "createdAt": "2026-08-12T21:18:03Z",
        "updatedAt": "2026-08-13T17:55:27Z",
        "timestamp": "2026-08-13T17:55:27Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-documentation",
          "can be tested"
        ],
        "author": "linhgiang24",
        "state": "closed",
        "assignees": [
          "tiandiwonder"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114544",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Text index: fix read mode of LIKE/ILIKE with an index built on a container",
        "text": "Evaluating `LIKE`/`ILIKE` by scanning the text index dictionary always used an exact direct read, which removes the original condition from the query plan. That is only valid when a matching dictionary token proves the predicate, which is not the case when the index is built on a container while the predicate reads a single element out of it: with an index on `mapValues(m)`, `m['a'] LIKE '%foobar%'` also returned rows whose match is under another key. However, it may produce correct result due to the data, but it's not guaranteed. To ensure for producing correct results, we need to use `HINT` mode when a text index built on a container (Map or JSON). It is better to be safe than sorry. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix text index evaluation of `LIKE/ILIKE` operator built on Map or JSON containers. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1313` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114544",
        "createdAt": "2026-08-12T21:22:24Z",
        "updatedAt": "2026-08-13T10:01:06Z",
        "timestamp": "2026-08-13T10:01:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "pr-synced-to-cloud"
        ],
        "author": "ahmadov",
        "state": "closed",
        "assignees": [
          "rschu1ze"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114545",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support quoted PromQL grouping labels",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Support quoted label names in PromQL grouping modifiers. ### Summary Prometheus accepts quoted label names in grouping modifiers. ClickHouse previously only accepted identifier-style tokens there. Prometheus 3.x supports UTF-8 label names, and OpenTelemetry uses dotted resource attributes such as `service.name`, `deployment.environment`, and `k8s.namespace.name`. When a PromQL backend preserves these names, grouping queries need the quoted form. This comes up with OpenTelemetry-aware dashboards and backends that keep the original label names instead of converting dots to underscores. Examples: - `sum by (\"service.name\", \"k8s.namespace.name\") (http_requests_total)` - `max without (\"deployment.environment\") (http_request_duration_seconds)` - `http_requests_total + on (\"service.name\") group_left (\"pod.name\") target_info` - `http_requests_total / ignoring (\"cluster.name\") group_right (\"instance.name\") target_info` ### Changes - Allow quoted strings in `by`, `without`, `on`, `ignoring`, `group_left`, and `group_right` lists. - Unquote and validate the names before passing them to the PromQL tree. - Quote non-legacy names again when serializing the tree back to PromQL. - Use PromQL-specific escaping for serialized strings, including control bytes. ### Tests - Regenerated ANTLR output from the grammar. - Added `gtest_PromQLParser` coverage. - Added realistic quoted-label queries to `04131_prometheus_query_parser`. - Added invalid empty quoted-label coverage. - Added parse/serialization/parse round-trip coverage for an escaped NUL label. References: - [Prometheus UTF-8 in Prometheus](https://prometheus.io/docs/guides/utf8/) - [Using Prometheus as your OpenTelemetry backend](https://prometheus.io/docs/guides/opentelemetry/) - [Prometheus query language parser](https://github.com/prometheus/prometheus/tree/main/promql/parser)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114545",
        "timestamp": "2026-08-12T22:24:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "submodule changed",
          "manual approve",
          "can be tested",
          "comp-promql"
        ],
        "author": "fallintoplace",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114546",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix wrapped `Time64` values from an overflowing scale conversion in `convertFieldToType`",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/94537 Related: https://github.com/ClickHouse/ClickHouse/pull/111534 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes incorrect, often sign-flipped, `Time64` literals and `IN`-list constants produced when rescaling a lower-scale `Decimal64` overflowed `Int64`. Such a conversion now reports `DECIMAL_OVERFLOW`, matching the `DateTime64` branch and explicit `CAST`. ### Description `convertFieldToTypeImpl` rescales a `Decimal64`-backed field into a `Time64` column at `src/Interpreters/convertFieldToType.cpp:473-476`. The scale-increasing arm multiplied without a range check, so `value * scale_multiplier_diff` could exceed `Int64` and wrap. The wrapped product was then handed to `decimalFromComponentsWithMultiplier<Time64>(value, 0, 1)`, whose own `mulOverflow` check is a no-op at multiplier `1`, so the corrupted value became the field. Observed on master (`bb1bd307`): `Values 'x Time64(6)' (253402207200000::Decimal64(0))` returned `-999:59:59.722624`, and `INSERT` persisted it. The wrapped `Int64` of `253402207200000 * 1000000` is `-4852209831933722624`, which renders as exactly that; the negated input wraps to `+4852209831933722624`, so a negative input returned a positive time. Explicit `CAST(... AS Time64(6))` already reported `DECIMAL_OVERFLOW` here, so the literal path disagreed with `CAST`. The file carried this same statement twice, for `DateTime64` and `Time64`, both unguarded. The related PR above guarded the `DateTime64` one and added test `03797`; its `Time64` twin was left as it was. This change mirrors that guard onto the twin. The operand also becomes `Int64`, which the guard requires: `mulOverflow` on an unsigned operand reports overflow for every negative value, which would reject in-range negative times. Reporting rather than returning Null matches the `DateTime64` twin, which `03797` asserts, and explicit `CAST`. The neighbouring `Date32` branches keep their Null contract and are untouched. Found by a UBSan signed-overflow report on this line; there is no open issue for it. It keeps reproducing on `master`, in both `asan_ubsan` stress jobs, as `signed integer overflow: 253402239600000 * 1000000 cannot be represented in type 'long'` at `src/Interpreters/convertFieldToType.cpp:475`: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=86363e819f9cb704a892063da9998e4c1281eb76&name_0=MasterCI&name_1=Stress%20test%20%28amd_asan_ubsan%29 - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=895af217db902c1458fecc608c6f6b3a78cc985e&name_0=MasterCI&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Only inputs that were producing wrong values change: `9223372036854` at scale 6, the largest whose rescale still fits, is still accepted and returns `999:59:59.000000` identically. Verified by building both arms and diffing: every overflow arm goes from a wrong value to `DECIMAL_OVERFLOW`, while 30 in-range controls across scales 0/3/6/9, both signs and both scale directions are byte-identical. `03797` and `04837` still pass. New test `04883` fails on master and is green here over 100 randomized runs.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114546",
        "createdAt": "2026-08-12T21:31:34Z",
        "updatedAt": "2026-08-13T16:41:17Z",
        "timestamp": "2026-08-13T16:41:17Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114547",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Horizon Support in ClickHouse",
        "text": "Add support for Snowflake Horizon. It support read and write path. I tested it against my own catalog. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description] Add support for Snowflake Horizon. You can now query Iceberg table in Iceberg behind the horizon catalog. You can also write to the Iceberg table via the catalog.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114547",
        "createdAt": "2026-08-12T21:33:29Z",
        "updatedAt": "2026-08-13T08:19:48Z",
        "timestamp": "2026-08-13T08:19:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-feature"
        ],
        "author": "melvynator",
        "state": "open",
        "assignees": [
          "asya-ch"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114548",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support trailing commas in PromQL grouping labels",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Accept trailing commas in PromQL grouping label lists. ### Description The current Prometheus operators guide explicitly documents trailing commas in label lists. Both `(label1, label2)` and `(label1, label2,)` are valid syntax. Prometheus has supported this since 2.16.0. The upstream change was merged in [PromQL: Support trailing commas in grouping opts](https://github.com/prometheus/prometheus/pull/6480) and is listed in the [Prometheus changelog](https://github.com/prometheus/prometheus/blob/main/CHANGELOG.md). It followed an upstream request for consistency with trailing commas already supported in label matchers ([issue #6470](https://github.com/prometheus/prometheus/issues/6470)). This is useful for generated or templated queries where a label list can be assembled with a final comma. ClickHouse previously rejected these valid PromQL queries. The shared ANTLR `labelNameList` rule now accepts one optional trailing comma for `by`, `without`, `on`, `ignoring`, `group_left`, and `group_right`. Examples: - `sum by (job,) (up)` - `sum by (job, instance,) (up)` - `foo + on(job,) bar` - `foo + ignoring(instance,) bar` ### References - [Prometheus operators guide](https://prometheus.io/docs/prometheus/latest/querying/operators/) - [Prometheus implementation PR #6480](https://github.com/prometheus/prometheus/pull/6480) - [Prometheus issue #6470](https://github.com/prometheus/prometheus/issues/6470) - [Prometheus changelog](https://github.com/prometheus/prometheus/blob/main/CHANGELOG.md) ### Tests - Added `PromQLParser.TrailingCommasInGroupingLabelLists`. - Covered one-label and multi-label trailing commas. - Covered malformed empty and double-comma lists with rejection assertions. - Regenerated the ANTLR parser artifacts. - Ran parser smoke tests for all six forms and malformed empty/double-comma lists.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114548",
        "createdAt": "2026-08-12T21:35:23Z",
        "updatedAt": "2026-08-13T15:22:55Z",
        "timestamp": "2026-08-13T15:22:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "submodule changed",
          "can be tested",
          "comp-promql"
        ],
        "author": "fallintoplace",
        "state": "open",
        "assignees": [
          "vitlibar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114549",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Let functions answer `canThrow` instead of deriving it from short-circuit suitability",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/55656 `IExecutableFunction::canThrow` decides whether the rows that a query does not reference have to be removed from `ColumnReplicated` arguments before a function is executed. It was derived from `isSuitableForShortCircuitArgumentsExecution`, with a `TODO` in `IFunctionAdaptors.h` asking for a dedicated interface. Those two properties answer different questions, and they disagree in both directions: - a function that is expensive but cannot throw is reported as throwing - this only loses an optimization; - a function that is cheap but can throw is reported as not throwing - this is not safe. Comparisons are the second case. `FunctionComparison` declares itself unsuitable for lazy execution because a comparison is cheap, so it was reported as never throwing, while a comparison does throw while it is executed: ```sql SELECT count() FROM numbers(3) WHERE toDate('2020-01-01') = 'garbage'; -- Code: 38. DB::Exception: Cannot parse date: value is too short: while converting 'garbage' to Date (CANNOT_PARSE_DATE) ``` This adds `IFunction::canThrow` so that a function describes the property on its own, and describes it for comparisons: whenever both sides are compared the way they are stored, or only the narrower side is widened - numbers, strings, `Date` with `Date32`, equal decimal, `DateTime64` and `Time64` scales, and equal types compared by `IColumn::compareAt` - a comparison cannot throw, which keeps the fast path for the overwhelming majority of queries. The cases that have to interpret one of the sides report that they can throw: a string parsed as a date or a tuple, a string validated against an enum, or different decimal scales brought to a common scale. `isNotDistinctFrom` is described the same way and additionally requires the two types to be equal, because it casts both sides to their common type first. `repeat` now says explicitly that it can throw - it throws `TOO_LARGE_STRING_SIZE` depending on the data of a single row. Its answer does not change (it is suitable for lazy execution, so the old approximation already reported `true`), but the property is now stated where it belongs instead of following from an unrelated one. The default implementation still falls back to the old approximation, so no function other than the comparisons changes its behaviour, and the `TODO` is narrowed to the remaining work: annotate more functions, then make the default the conservative `true`. Verified locally by compiling the affected translation units (`equals.cpp`, `less.cpp`, `isNotDistinctFrom.cpp`, `repeat.cpp`, `IFunction.cpp`) and by checking the added test against a build of `master`, where its results are the same - the compaction of unreferenced rows is invisible in the results, only in whether an exception can escape. The rest is left to CI. ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114549",
        "timestamp": "2026-08-12T23:06:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114550",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Deduplicate the final DISTINCT in parallel",
        "text": "The final `DISTINCT` merged all of its input streams into one and deduplicated them on a single thread, so `SELECT DISTINCT` did not scale with `max_threads` while the equivalent `GROUP BY` did - aggregation parallelizes its merge over two-level hash tables, `DISTINCT` had nothing of the sort. This repartitions the input streams by the hash of the `DISTINCT` columns instead of merging them: equal values then land in the same stream, so every stream can be deduplicated on its own, and what is left to merge is already deduplicated. The repartitioning uses `scatterByPartition`, which `SortingStep` and `ShuffleSendStep` already use. Whether the order of the step's output is free to change is only known where the sorting properties are, so the decision is taken in `applyOrder`. The new setting `allow_parallel_distinct`, enabled by default, controls it. It is not applied when: - the input is already globally sorted, because then `DISTINCT` keeps the order of its input and the repartitioning would break it; - `max_rows_in_distinct` or `max_bytes_in_distinct` is set, because those limits are enforced by the single transform that sees the whole merged input; - there is a `LIMIT`, because the transform stops early anyway and the single stream keeps `LIMIT` returning the first values of the input rather than an arbitrary subset; - the streams are already disjoint - `allow_distinct_partitions_independently` covers that case and skips the merge entirely. Every input chunk is split across all partitions, so the number of partitions is capped rather than following `max_threads`. On `SELECT DISTINCT number FROM numbers_mt(4e7)` the total CPU time the scatter adds grows by 8% at 4 partitions, 20% at 16 and 96% at 96; without the cap that turned into a wall-clock regression at high `max_threads`. Wall clock, minimum of 3 runs, aarch64: | workload | `max_threads` | before | after | speedup | |---|---|---|---|---| | 20M distinct `UInt64`, `MergeTree` | 4 | 0.442 s | 0.205 s | 2.16x | | | 16 | 0.458 s | 0.148 s | 3.09x | | | 96 | 0.502 s | 0.165 s | 3.04x | | 40M distinct `UInt64`, `numbers_mt` | 4 | 0.965 s | 0.545 s | 1.77x | | | 16 | 0.929 s | 0.360 s | 2.58x | | | 96 | 0.886 s | 0.299 s | 2.96x | | 8M distinct `String`, `numbers_mt` | 4 | 0.515 s | 0.288 s | 1.79x | | | 16 | 0.517 s | 0.093 s | 5.56x | | | 96 | 0.518 s | 0.083 s | 6.24x | The measurements were taken on a busy machine, hence the minimum of several runs; the existing `distinct_in_order`, `distinct_high_cardinality` and `sorted_distinct` performance tests cover these shapes, so the performance comparison job is the authoritative check. `DISTINCT` without `ORDER BY` has no defined row order, and the repartitioning changes the order it happens to produce. Two existing tests relied on it: `04109_sort_propagation_apply_order` asserts that per-stream sortedness reaches `DISTINCT`, so it pins the setting off, and `00585_union_all_subquery_aggregation_column_removal` gets the explicit `ORDER BY` it was missing. Other tests that depend on the incidental order of a `DISTINCT` may surface in CI and should be fixed the same way. Closes: https://github.com/ClickHouse/ClickHouse/issues/114532 Related: https://github.com/ClickHouse/ClickHouse/pull/48988 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Evaluate the final `DISTINCT` on several threads by repartitioning its input by the hash of the `DISTINCT` columns, instead of merging everything into a single stream and deduplicating it on one thread. Up to 6x faster on high-cardinality `DISTINCT`. Controlled by the new setting `allow_parallel_distinct`, enabled by default. Note that this changes the order in which a `DISTINCT` without `ORDER BY` returns rows, which was never defined. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114550",
        "timestamp": "2026-08-12T23:25:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114551",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support quoted identifiers in PromQL selectors",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support quoted metric and label names in PromQL selectors. ### Description Prometheus 3 supports UTF-8 metric and label names. Prometheus's own end-to-end test uses selectors such as: - `{\"http.requests\", \"service.name\"=\"api-server\", instance=\"0\", group=\"canary\"}` AWS CloudWatch also documents quoted OpenTelemetry selectors such as: - `{\"http.server.active_requests\", \"@resource.service.name\"=\"myservice\"}` These names are also used in OpenTelemetry metrics and user-facing PromQL products. For example, SigNoz documents queries such as: - `{\"system.cpu.utilization\", \"service.name\"=\"frontend\"}` - `sum by (\"k8s.pod.name\") (rate({\"container.cpu.utilization\", \"k8s.namespace.name\"=\"ns\"}[5m]))` ClickHouse currently accepts only identifier-style tokens for selector names. A quoted selector identifier is rejected by the grammar before query evaluation. This change: - accepts quoted metric and label names in selectors - treats a standalone quoted selector name as an `__name__` matcher - validates and unquotes the identifier - keeps non-legacy names quoted when serializing the query tree References: - [Prometheus UTF-8 guide](https://prometheus.io/docs/guides/utf8/) - [Prometheus end-to-end UTF-8 selector test](https://github.com/prometheus/prometheus/blob/main/promql/promqltest/test_test.go#L166-L191) - [AWS CloudWatch PromQL examples](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-PromQL-Querying.html) - [OpenTelemetry HTTP metric conventions](https://opentelemetry.io/docs/specs/semconv/http/http-metrics/) - [OpenTelemetry Prometheus compatibility survey](https://opentelemetry.io/blog/2024/prometheus-compatibility-survey/) - [SigNoz PromQL UTF-8 guide](https://signoz.io/docs/userguide/write-a-prom-query-with-new-format/) ### Tests - Regenerated the ANTLR parser artifacts. - Added `gtest_PromQLParser` coverage for quoted metric names, quoted label names, and invalid identifiers. - Ran focused syntax checks for the generated parser and touched PromQL sources. - Ran focused grammar checks for quoted metric and label selectors.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114551",
        "createdAt": "2026-08-12T21:52:00Z",
        "updatedAt": "2026-08-13T15:33:58Z",
        "timestamp": "2026-08-13T15:33:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "submodule changed",
          "can be tested",
          "comp-promql"
        ],
        "author": "fallintoplace",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114552",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Raise the log-wait timeout in `test_keeper_session_loss_direct_read` to fix flakiness",
        "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Raise the log-wait timeout in the integration test `test_storage_kafka/test_keeper_session_loss_direct_read.py` to fix its flakiness. The test waits for the `StorageKafka2` activation task to notice the expired Keeper session after a network partition. The task's check period is one minute, and on an overloaded CI host (sanitizer builds, busy schedule pool) the log line has been seen missing the 180-second window, failing the test with `retry_failed`. `wait_for_log_line` returns as soon as the line appears, so a healthy run does not pay for the extra margin. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1305` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114552",
        "createdAt": "2026-08-12T22:48:38Z",
        "updatedAt": "2026-08-13T03:54:19Z",
        "timestamp": "2026-08-13T03:54:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114553",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #110283 to 26.5: Compose join-order statistics over parts surviving partition/PK pruning",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/110283 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31648223640/job/94286604401)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114553",
        "timestamp": "2026-08-12T23:05:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll3",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114554",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #110283 to 26.6: Compose join-order statistics over parts surviving partition/PK pruning",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/110283 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31648223640/job/94286604401)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114554",
        "timestamp": "2026-08-12T23:05:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll3",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114555",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #110283 to 26.7: Compose join-order statistics over parts surviving partition/PK pruning",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/110283 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31648223640/job/94286604401)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114555",
        "timestamp": "2026-08-12T23:06:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll3",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114556",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Ask `@oranjeai` instead of `@groeneai` in the private repository",
        "text": "The `continue-pr` skill hardcoded `@groeneai` as the person to ping about CI failures that are unrelated to the pull request being worked on. That handle is only correct for `ClickHouse/ClickHouse`; when the skill runs in `ClickHouse/clickhouse-private`, the request should go to `@oranjeai` instead. Step 1 now detects the repository (`gh repo view --json nameWithOwner`, falling back to `git remote get-url origin`), remembers it as `$REPO` for the `gh` commands, API URLs, and GraphQL arguments further down, and step 4 picks the reviewer from it. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1297` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114556",
        "createdAt": "2026-08-12T23:18:07Z",
        "updatedAt": "2026-08-13T01:17:25Z",
        "timestamp": "2026-08-13T01:17:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114557",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Unattended `continue-pr-auto` skill and the `continue-all-prs` driver",
        "text": "Adds the unattended `continue-pr-auto` skill and the `utils/continue-all-prs.sh` driver that runs it over the open pull requests, with a status bar renderer (`utils/continue-all-prs-status.py`) and an `utils/exclude-authors.txt` list. `continue-pr` keeps its interactive behavior (it may still ask about ambiguous conflicts and unclear reviewer suggestions); `continue-pr-auto` never asks, resolves conflicts autonomously, and always pushes, so `disable-model-invocation: true` keeps interactive sessions on `continue-pr`. Both skills now determine the repository they run in and pick whom to ping about CI failures unrelated to the pull request from it: `@groeneai` in `ClickHouse/ClickHouse`, `@oranjeai` in `ClickHouse/clickhouse-private`. Every comment the unattended skill posts is prefixed with 🕵 so automated comments are identifiable. The branch was stale, so it also merges current `master`, and it includes the `continue-pr` change from the pull request below (that change touches the same lines, so both are here to keep this branch based on `master` directly). Related: https://github.com/ClickHouse/ClickHouse/pull/114556 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1304` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114557",
        "createdAt": "2026-08-12T23:20:59Z",
        "updatedAt": "2026-08-13T03:03:28Z",
        "timestamp": "2026-08-13T03:03:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114558",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support PromQL @ start() and @ end() modifiers",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): PromQL: support `@ start()` and `@ end()` timestamp modifiers. ### Summary PromQL allows `start()` and `end()` as special values for the `@` modifier. They mean the fixed start and end of a range query. For an instant query, both resolve to the evaluation time. Without this support, valid PromQL such as `http_requests_total @ start()` and `rate(http_requests_total[5m] @ end())` cannot be used against ClickHouse. Fixed-`@` subqueries keep their complete inner time grid and remain step-invariant across the outer range query. Standalone `start()` and `end()` functions are not included. ### Why this matters This is standard PromQL surface used by major Prometheus-compatible systems: - VictoriaMetrics has executable range-query tests for both `time() @ start()` and `time() @ end()`. - Grafana Mimir, Thanos, and Cortex replace these modifiers with fixed timestamps before splitting range queries, including subqueries. - Elasticsearch includes `@ start()`, `@ end()`, and subquery forms in its valid PromQL grammar fixtures. These are public implementation and compatibility references, not a claim about private customer query volume. Supporting the syntax keeps dashboards and recording/alerting queries portable across Prometheus-compatible backends. References: - [Prometheus Querying basics: `@` modifier](https://prometheus.io/docs/prometheus/latest/querying/basics/) - [VictoriaMetrics executable tests](https://github.com/VictoriaMetrics/VictoriaMetrics/blob/master/app/vmselect/promql/exec_test.go#L1207-L1227) - [Grafana Mimir query splitting](https://github.com/grafana/mimir/blob/main/pkg/frontend/querymiddleware/split_and_cache.go#L2838-L2948) - [Thanos query splitting](https://github.com/thanos-io/thanos/blob/main/internal/cortex/querier/queryrange/split_by_interval.go#L593-L690) - [Cortex query splitting](https://github.com/cortexproject/cortex/blob/master/pkg/querier/tripperware/queryrange/split_by_interval.go#L1061-L1157) - [Elasticsearch valid PromQL grammar fixtures](https://github.com/elastic/elasticsearch/blob/main/x-pack/plugin/esql/qa/testFixtures/src/main/resources/promql/grammar/queries-valid.promql#L1112-L1122) ### Changes - Add `start()` and `end()` to the PromQL timestamp grammar. - Preserve the symbolic modifier in the Prometheus query tree. - Resolve the symbols against the outer query boundaries during evaluation. - Keep offset and range-selector handling consistent with numeric `@` timestamps. - Preserve the complete inner grid for fixed-`@` subqueries. ### Tests - Added parser coverage for both symbols, modifier order, literals, and subqueries. - Added instant and range query coverage, including a range selector under `@ end()`. - Added `last_over_time()` coverage for symbolic and numeric `@` on subqueries. - Compared the new range-query cases with Prometheus through the existing integration helpers. - Regenerated the ANTLR parser artifacts.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114558",
        "createdAt": "2026-08-12T23:36:48Z",
        "updatedAt": "2026-08-13T15:33:49Z",
        "timestamp": "2026-08-13T15:33:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-improvement",
          "submodule changed",
          "comp-promql"
        ],
        "author": "fallintoplace",
        "state": "open",
        "assignees": [
          "vitlibar"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114559",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Never cancel a query before its `max_execution_time` has elapsed",
        "text": "`CancellationChecker::appendTask` derived a query's cancellation deadline from the current time truncated to whole milliseconds, and then aligned it up to the 100 ms grid the worker batches deadlines on. The alignment normally hides the truncation, but when the deadline already sits on the grid - one query in a hundred - it adds no padding, and the deadline lands up to 1 ms before `now + max_execution_time`. The query is then cancelled ahead of its own timeout and fails with a self-contradictory message: ``` Code: 159. DB::Exception: Timeout exceeded: elapsed 999.672 ms, maximum: 1000 ms. (TIMEOUT_EXCEEDED) ``` The current time is now rounded up to a whole millisecond instead, which is the direction the code around it already promises: *\"Tasks may be cancelled slightly later than their exact timeout, but never before\"*. Only the `throw` overflow mode was affected: in `break` mode the checker re-reads the elapsed time before acting on the deadline, so an early wake-up does nothing. This is what makes `02122_join_group_by_timeout` flaky - it asserts that a `max_execution_time = 1` cancellation takes at least 1000 ms, and `query_duration_ms` is measured from a stopwatch that starts before the deadline is computed. The failure shows up as `query_duration 1` becoming `query_duration 0`, always on one of the two `throw`-mode queries and never on the `break`-mode ones. In the server logs of both `Fast test` runs below, the query was cancelled early - at 999.672 ms and at 999.630 ms of its 1000 ms timeout. The deadline computation moved into `CancellationChecker::taskDeadlineMs` so that the invariant can be tested directly: the new unit test checks every sub-millisecond phase of `now` against every phase of the grid, for a range of timeouts, and fails with the old truncation. CI reports of the flaky test: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-113021-93e2851fddb5add7bad5b28bcef780be2f6ff7bb&sha=bef88e3d96ecb0f57a0f8d81f6b7ee2117db19fd&name_0=MergeQueueCI&name_1=Fast%20test - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-110451-2f02abe52a79c02fb554728f14d8fa847dcbc1e6&sha=2eba84a8a12a21c36516158d4fbe9530d3f0af8c&name_0=MergeQueueCI&name_1=Fast%20test The truncation is as old as the checker itself. Before deadlines were aligned to a grid in https://github.com/ClickHouse/ClickHouse/pull/95563 the deadline was plainly `floor(now) + timeout`, i.e. early for every query rather than for one in a hundred; and the assertion that catches it was only sharpened from `round(query_duration_ms / 1000) BETWEEN 1 AND 60` to `query_duration_ms BETWEEN 1000 AND 60000` in daf1b4980b1, which is why the flake is recent. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A query could be cancelled up to 1 ms before its `max_execution_time` had passed, failing with a self-contradictory `Timeout exceeded: elapsed 999.672 ms, maximum: 1000 ms`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114559",
        "createdAt": "2026-08-13T00:25:45Z",
        "updatedAt": "2026-08-13T04:58:10Z",
        "timestamp": "2026-08-13T04:58:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "thevar1able"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114560",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113534 to 26.7: Fix a mixed JOIN ON condition evaluated over mismatched column types for a dictionary",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113534 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114423 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31653108878/job/94302376144)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114560",
        "createdAt": "2026-08-13T00:26:20Z",
        "updatedAt": "2026-08-13T03:23:47Z",
        "timestamp": "2026-08-13T03:23:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114561",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix MATERIALIZED columns frozen at their INSERT-time value after ALTER UPDATE",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/100613 Closes: https://github.com/ClickHouse/ClickHouse/issues/50237 Related: https://github.com/ClickHouse/ClickHouse/pull/99281 ## Problem A `MATERIALIZED` column computed from another `MATERIALIZED` column is never recalculated by `ALTER TABLE ... UPDATE`. It keeps the value `INSERT` computed, forever, no matter how many mutations run. ```sql CREATE TABLE t (x Int32, m1 Int32 MATERIALIZED x + 1, m2 Int32 MATERIALIZED m1 + 1) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t (x) VALUES (10); -- m1 = 11, m2 = 12 ALTER TABLE t UPDATE x = 20 WHERE 1; ALTER TABLE t UPDATE x = 30 WHERE 1; SELECT x, m1, m2 FROM t; -- 30 31 12 <- m2 frozen at its INSERT-time value -- expected: 30 31 32 ``` Concretely, on `master`: - `m2` and everything downstream of it are stale for the lifetime of the table. - A projection over a recalculated `MATERIALIZED` column keeps answering from pre-mutation data: `SELECT sum(m1)` returns `500500` through the projection versus `1000500500` without it. - A `minmax` skip index over such a column prunes every granule, so a filter that matches all rows returns `0`. - With `apply_mutations_on_fly = 1` and the mutation still pending, the same query returns different values depending on its select list: `SELECT m2` gives `12` while `SELECT x, m1, m2` gives `22`. - Also with a pending mutation, reading a `MATERIALIZED` column whose expression mentions an `EPHEMERAL` column fails outright with `Code: 47 ... Missing columns: 'e' while processing: 'x + e' ... (UNKNOWN_IDENTIFIER)`. ## Root cause `MutationsInterpreter::prepare` collected the `MATERIALIZED` columns to recalculate from **direct** dependencies only. It analysed each `MATERIALIZED` default expression and recorded an edge `dependency -> materialized column`, but only when `updated_columns` already contained the dependency: ```cpp for (const auto & dependency : required_columns) if (updated_columns.contains(dependency)) column_to_affected_materialized[dependency].push_back(column.name); ``` For `m2 MATERIALIZED m1 + 1` the dependency is `m1`, which the user did not update, so `m2` never entered `affected_materialized`, and the guard added by #99281 skipped it: ```cpp if (column.default_desc.kind == ColumnDefaultKind::Materialized && affected_materialized.contains(column.name) && column.default_desc.expression) ``` Before #99281 every `MATERIALIZED` column was recalculated in a single stage, so `m2` *was* rewritten — but from the pre-stage `m1`, i.e. always one mutation behind. That second half is #50237. Both halves come from the same place: the affected set is not transitive, and even a transitively dependent column was evaluated in the same stage as the column it reads. The remaining symptoms are the same \"direct dependencies only\" assumption in the neighbouring consumers. The skip-index, projection and dependency-analysis predicates only looked at `updated_columns`, `changed_columns` and `patch_updated_columns`, none of which ever names a recalculated `MATERIALIZED` column. `AlterConversions::addColumnsRequiredForMaterialized`, which decides what an on-fly mutation must read, walked one level and then required the dependency to be an updated column, so a read set of `{m2}` stayed `{m2}`, `filterMutationCommands` dropped the `UPDATE x` command because nothing overlapped, and the stale value was returned. That same function analysed default expressions against `getAllPhysical`, which excludes `EPHEMERAL` columns, hence the `UNKNOWN_IDENTIFIER`. ## Solution In `src/Interpreters/MutationsInterpreter.cpp`: 1. Record **every** edge of the `MATERIALIZED` dependency graph, including `MATERIALIZED -> MATERIALIZED` edges. 2. `getAffectedMaterializedByLevel` walks that graph from the columns a command updates and groups the reachable columns by **longest-path** dependency level. 3. Emit one mutation stage per level. Stage *N* reads the output of stage *N-1*, so an expression sees the recalculated value of the column it depends on instead of the stale on-disk one — this is what closes #50237, not merely refreshing the value. 4. `validateUpdateColumns` receives the same transitive closure, so an `UPDATE` reaching a `MATERIALIZED` **key** column through another `MATERIALIZED` column is rejected with `CANNOT_UPDATE_COLUMN`, as the direct case already was. Without this, step 3 would rewrite a sorting-key column and break the part's sort order. 5. Feed the recalculated columns into the skip-index, projection and dependency analysis. In `src/Storages/MergeTree/AlterConversions.cpp`: 6. `addColumnsRequiredForMaterialized` follows the chain, pulling in an intermediate `MATERIALIZED` column when an updated column is reachable through it. 7. `EPHEMERAL` columns enter its analysis set, as `MutationsInterpreter::prepare` already does, and a `MATERIALIZED` column reading one is excluded from the walk because it cannot be recalculated anyway. The pre-existing `EPHEMERAL` exclusion that #99281 added is preserved: such a column is skipped before its incoming edges are recorded, so the walk never reaches it. The exclusion removes that column from the graph, not the subgraph behind it — a column that also reads an updated column directly (`mc MATERIALIZED x + e`, `mc2 MATERIALIZED mc + x`) is still recalculated, from the stored `mc`, which is well defined because its own expression reads only stored columns. Items 5-7 are pre-existing defects rather than regressions of this change; item 5 is reproducible on `master` for a *directly* recalculated column too. They are fixed here because otherwise items 1-3 would turn a uniformly stale table into an inconsistent one: correct on disk, wrong through a projection, and different per select list. ### Tests `tests/queries/0_stateless/04869_transitive_materialized_column_update.sql` covers nine cases, one per clause of the fix: chain, diamond (longest-path placement), `EPHEMERAL` must-not-recalculate, `EPHEMERAL` with converging paths, `MATERIALIZED` key column, projection (direct and transitive), skip index, on-fly chain, on-fly `EPHEMERAL`. Each assertion was verified to be load-bearing by reverting only its own clause: the chain freezes at `12 13`, the diamond gives `5 6 60 26` when a re-reached column is placed at its first level instead of its deepest, the projections return `500500` / `501500`, the skip index returns `0`, and the on-fly reads return `12` / `13` or throw. `04044_mutation_ephemeral_materialized`, the test #99281 added, still passes. The projection and skip-index cases keep an untouched column `y`, because a mutation that rewrites every column goes through `MutateAllPartColumns`, which rebuilds everything and hides the bug. ### Notes for reviewers - **Behaviour change:** `UPDATE` of a column that transitively feeds a `MATERIALIZED` key column now throws `CANNOT_UPDATE_COLUMN` where it previously succeeded and left the key stale. The direct equivalent already throws on `master`, so this makes the two consistent. - Tables with `MATERIALIZED -> MATERIALIZED` chains now get one mutation stage per dependency level; everything else keeps the single stage that existed before. - More projections and skip indices are rebuilt than before. That is the point of item 5, but it does make such mutations more expensive. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `MATERIALIZED` columns computed from another `MATERIALIZED` column keeping their `INSERT`-time value forever after `ALTER TABLE ... UPDATE`. Projections and skip indices over a recalculated `MATERIALIZED` column were also left stale, returning wrong results.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114561",
        "createdAt": "2026-08-13T00:41:12Z",
        "updatedAt": "2026-08-13T05:23:32Z",
        "timestamp": "2026-08-13T05:23:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114562",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix NO_SUCH_COLUMN_IN_TABLE naming a column absent from the table when a part's columns were all renamed or all dropped",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/103654 Closes: https://github.com/ClickHouse/ClickHouse/issues/103116 Related: https://github.com/ClickHouse/ClickHouse/pull/96675 Related: https://github.com/ClickHouse/ClickHouse/issues/79110 ## Problem `NO_SUCH_COLUMN_IN_TABLE` naming a column that is **not** in the table, when reading or merging a part whose columns the metadata no longer knows under those names. Two reported shapes, both with a mutation queued but not yet applied: - every column renamed, and a merge has to read a column missing from the part (#103654) — the merge then fails and is retried forever, so the table stops compacting: ``` Code: 16. DB::Exception: There is no column h in table. (NO_SUCH_COLUMN_IN_TABLE) (in query: OPTIMIZE TABLE t_rename_vm FINAL;) ``` - every column dropped (#103116) — a plain `SELECT` fails, no merge needed: ``` Code: 16. DB::Exception: There is no column a in table. (NO_SUCH_COLUMN_IN_TABLE) (in query: SELECT c FROM t_all_dropped ORDER BY c;) ``` ## Root cause `injectRequiredColumns` injects the smallest physical column of a part when none of the requested columns is physically present, only to learn the number of rows. That column must be **both** readable from the part **and** resolvable by the caller, which looks names up in the metadata. The candidate set kept being taken from one side or the other — which guarantees only one of the two properties — and each fix patched the other side: | commit | candidate set | fixed | broke | |---|---|---|---| | `5c1db5fc661` | metadata columns | unresolvable names | part may have no file for any of them → `LOGICAL_ERROR` in `getColumnNameWithMinimumCompressedSize` | | `6d0b4dc9885` | part columns | that `LOGICAL_ERROR` | part name may be absent from the metadata → `NO_SUCH_COLUMN_IN_TABLE` (#96675) | | `1d2e2d1c6b2` (#96675) | intersection, else part columns | both, when non-empty | the empty case reverts to part-only names → **#103116**; the intersection compares *part* names against *metadata* names, so renamed columns drop out → **#103654** | Two things kept this going: - **Wrong name space.** The rest of the function works in metadata names and maps to part names only for the part lookup (`injectRequiredColumnsRecursively`, lines 105-107). This block alone enumerated part names and filtered them against the snapshot, which is exactly why renames fell through it. - **Wrong problem.** The block's own comment says the goal is to \"know number of rows\", implemented as a column read, while `getReadTaskColumns` already documents that a column is needed *only* for non-adaptive granularity. And `if (available_columns.empty()) available_columns = part_columns;` substituted an unusable value rather than answering whether that set can legitimately be empty — it can, and the answer differs by case. ## Solution Build the candidate set as the **pairing**, so both properties hold by construction: walk the metadata columns, map each to its name in the part through `AlterConversions` exactly as `injectRequiredColumnsRecursively` does, skip the ones a pending mutation drops, and keep those the part has files for. Sizes are looked up under the part's name; the metadata name is what gets injected. When the pairing is empty, decide on the merits instead of falling back, because the two ways of getting there are opposites: - **Everything the part holds is accounted for** by the structure or by a pending drop. Then each requested column is legitimately absent and the part reads as rows of defaults. Nothing needs to be injected under adaptive granularity; if the granularity is not adaptive, throw, since a column really must be read. - **The part holds a column that is neither in the structure nor being dropped** — a part attached after the schema moved on. Its rows exist on disk, so reporting defaults would hide data. Throw, naming the offending column. That second rule is not hypothetical: `04011_detach_rename_attach_column` (issue #79110) requires exactly this error for the renamed variant, and an earlier version of this patch that always injected nothing broke it. #79110's reporter asked whether such an attach should fail and the answer was yes. A drop is recorded under the name the command used, so when a column was also renamed the drop and the part-side name differ and both are consulted. A single `ALTER` can rename and then drop the same column, so this composition is reachable — two statements cannot produce it, because an `ALTER` issued while a `RENAME` is still pending blocks. ### Tests - `04870_vertical_merge_inject_column_after_rename.sh` — #103654. - `04871_inject_column_no_common_column_with_metadata.sh` — four cases: all columns dropped (#103116's repro); a stale re-attached part, which must **error**; renamed columns with `index_granularity_bytes = 0`; and a single-statement pending rename + drop. Both are shell tests so a `trap` always releases the server-global failpoint — their failure mode is a query throwing, which in a `.sql` test would leave mutations disabled for every later test on that server. Every clause was checked by reverting it alone: without the rename mapping the non-adaptive case fails; without the adaptive-granularity branch the dropped-columns case fails; without the second namespace in the drop check the rename+drop case fails. `03830_vertical_merge_inject_column_after_drop` (the test of #96675, which introduced this fallback) and `04011` both keep passing. ### Notes for reviewers - The candidate walk builds `getAllPhysical()` per part rather than taking a reference to the part's column list. The branch only runs when no requested column is physically present in the part, so it is not a hot path, but it is more work than before on a wide table read across many such parts. - Behaviour change: a part whose contents are fully accounted for is now read as rows of defaults where it previously threw. That is the point of #103116. - Two related defects found while verifying this are **not** fixed here and are independent of this function: - A part produced by a merge while a `RENAME COLUMN` is pending reads the renamed columns as their type default, because `StorageMergeTree::MutationsSnapshot::getOnFlyMutationCommandsForPart` decides applicability from the part's data version alone while the merge writes current metadata names. Values reappear once the mutation applies. Mirroring `ReplicatedMergeTreeQueue`'s metadata-version gate does *not* work — I tried it, and the source parts satisfy it too, so a plain `SELECT` regresses immediately. - `RENAME a TO b, DROP b, ADD b` in one statement makes the re-added `b` return the old `a` data (`min(b), max(b)` gave `100, 1099` instead of `7, 7`). The stale mapping is applied by `IMergeTreeReader`'s own name resolution, so it cannot be fixed in `injectRequiredColumns`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NO_SUCH_COLUMN_IN_TABLE` naming a column that is not in the table, when reading or merging a part whose columns had all been renamed or all dropped by a mutation that was not applied yet. The merge kept being retried, so the table stopped compacting.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114562",
        "createdAt": "2026-08-13T00:42:13Z",
        "updatedAt": "2026-08-13T10:46:30Z",
        "timestamp": "2026-08-13T10:46:30Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114563",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Order the /play startup test's load-window keystroke by construction",
        "text": "Order the /play startup test's load-window keystroke by construction <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `test_play_reconcile_startup` can fail one assertion of `dirty-startup-merge-entry-reowned` (\"the URL hash carries the live query, not the dropped blank tab stale hash\", actual `?tab=Scratch#U0VMRUNUIDExMQ==`) while the scenario's other five checks pass. No related open issue found. The Web UI is not at fault. The four `dirty-startup-*` scenarios must deliver their \"user typed while IndexedDB was still opening\" keystroke inside the load window, but the harness expressed that ordering as a duration: `await sleep(config.duringLoadDelayMs || 5)`. The fake `IndexedDB.open` registers its `openDelayMs: 30` timer *during* page-script evaluation, while the keystroke timer is registered only *after* evaluation returns, so the keystroke wins only while `eval_end - open_call` stays under the measured 26-28 ms slack. Evaluating the 647 KB page script takes 1-2 ms idle, up to 17 ms under CPU contention. When the keystroke loses, `markBootstrapDirty` short-circuits on `bootstrap_settled`, `bootstrap_dirty` stays false, and reconciliation runs the non-dirty named-URL path, for which the stale URL is the *correct* result. The test was asserting the dirty contract against the non-dirty path. Yielding only to microtasks before the interaction makes the ordering structural: the bootstrap is synchronous, the open can resolve only through a `setTimeout`. Three premise assertions, one of which is that the interaction precedes the fake open's completion, make a future ordering regression fail loudly instead of being misattributed to `play.html`. The now-unreachable `duringLoadDelayMs` key is removed. So the suite can catch a reintroduction of the timing dependency, `dirty-startup-merge-entry-reowned` opens IndexedDB with zero delay; the other three keep the 30 ms open so the timer path stays covered. `test.py` also pins all five `dirty-startup-*` names, as it already did for `run-marker-*`: dropping any of them used to still report \"All scenarios passed\". No new test file, and no assertion weakened. All 84 checks pass, and the fix repairs all four `duringLoad` scenarios. <details> <summary>Validation</summary> Simulating a slow host (burn N ms of synchronous time after `indexedDB.open()` registers its timer, still inside evaluation). The injection is **external**: each arm runs its own tree's scenarios exactly as committed, no knob overridden. | stall | before | after | |---|---|---| | 26 ms | 84 PASS | - | | 30 ms | **3 FAIL** | 84 PASS | | 40 ms | **3 FAIL** | 84 PASS | | 400 ms | - | 84 PASS | The boundary matches the independently measured 26-28 ms slack; after the change the ordering holds at 16x that margin. On the runtime the test actually uses (`clickhouse/mysql-js-client`, node v22.22.0, page over HTTP), stall 40 ms was 3/3 red with the reported signature verbatim and is 3/3 green after. The suite now detects its own regression: reverting just the microtask yield to the old `sleep(5)` fails with `HARNESS ERROR: duringLoad ran after the IndexedDB open completed`, on both runtimes, and at every post-registration stall from 0 to 400 ms (a stall makes both fake timers overdue, the one case the workspace-state premises alone let through). Deleting any one `dirty-startup-*` scenario now fails `test.py`, where each deletion previously reported \"All scenarios passed\". Other arms: 20/20 consecutive unperturbed runs green with an invariant 84 PASS; keystroke forced to 45 ms reproduces before and is green after; under 8 busy CPU loops 3/3 red before, 3/3 green after. Counting evaluations shows the ordering assertion runs in all four scenarios and fires in none of them unperturbed, so it is neither dead nor trigger-happy. All four `duringLoad` carriers hold their premise after the change, and the adjacent timing knobs (`stale-reload-run-race`'s `openDelayMs`, the `wasmInstantiateDelayMs`/`disableWasm` scenarios) are deliberately untouched and stay green. Also checked against the in-flight `play.html` rewrite in PR 114187: both globals the new assertions read survive there and all 17 dirty-startup checks pass against its `play.html`. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1310` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114563",
        "createdAt": "2026-08-13T00:43:38Z",
        "updatedAt": "2026-08-13T08:36:42Z",
        "timestamp": "2026-08-13T08:36:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114564",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112921 to 26.7: Evaluate randomHadamardTransform once for a constant vector",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112921 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31655073318/job/94307646536)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114564",
        "createdAt": "2026-08-13T00:58:05Z",
        "updatedAt": "2026-08-13T04:08:09Z",
        "timestamp": "2026-08-13T04:08:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114565",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Scale the lldb stacktrace budget by build flavor, keep timed-out dumps",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109455 --> Related: https://github.com/ClickHouse/ClickHouse/pull/109455 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `clickhouse-test` collects C-side backtraces with `lldb -o 'thread backtrace all'` under a hard 30s timeout. On a debug server the attach needs longer, the timeout fires, and the partial dump was discarded too, leaving no C stacks for the wedged process. CIDB, 90 days: 181 `suspiciously small stacktraces` rows over 118 PRs, 80 `timed out after 30 seconds` over 59 PRs, 78 carrying both. Two causes: 1. The 30s budget assumed the backtrace \"should finish in seconds even on a debug build\", so anything longer meant a wedged lldb. Measured on a 174-thread debug server in the CI stress image, a healthy attach takes 44-118s and returns a complete dump: the cost is symbolization. 2. `shell_get_output` honoured `keep_output_on_error` only for `CalledProcessError`; `TimeoutExpired` is not a subclass, so a timeout returned `\"\"`, discarding what lldb wrote. The per-PID budget now scales with the build flavor, reusing the `slow_build` predicate (debug, TSan/ASan/UBSan/MSan, coverage) that `MergeTreeSettingsRandomizer` uses for the same reason: 120s on a slow build, 30s unchanged on release (`Stress test (arm_debug)` is 51 of those 80 rows, so the driver is debug). The flavor comes from `args.build_flags`, not the module globals a spawned worker resets; the startup-failure caller runs before those flags exist, so there it is read from the binary via `clickhouse local`, as `is_asan_build` already does. A separate aggregate ceiling bounds the pid loop and reports skipped pids. The per-test timeout handler keeps 30s, since it runs inside its own fired alarm. A timed-out dump is now kept, decoded and marked truncated. Validated against a live debug server with real lldb: before, 30.8s and a 0-byte dump called suspiciously small; after, 63.6s and 834KB with 348 headers. Forcing the budget below the attach cost keeps 48.7KB where master keeps 0. Twelve new tests, each change reverted reddening one.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114565",
        "createdAt": "2026-08-13T01:06:44Z",
        "updatedAt": "2026-08-13T10:36:18Z",
        "timestamp": "2026-08-13T10:36:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "manual approve",
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114566",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Improve pruned statistics planning benchmark and comments",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110283 ### Summary Follow up on review feedback from #110283: - clarify the range-analysis, exact-empty, STREAM, and statistics-cache contracts; - document what `ConditionSelectivityEstimator::isStale()` compares; - remove cross-component implementation references from optimizer comments; - make the performance workload exercise statistics from 950 of 1000 parts instead of producing noisy 2–3 ms measurements. This PR does not change production behavior. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry: Not for changelog. ### Testing - `git diff --check` - XML validation with `xmllint` - local hostile diff review The release performance comparison is left to CI. 🕵️‍♂️",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114566",
        "createdAt": "2026-08-13T01:12:41Z",
        "updatedAt": "2026-08-13T05:44:39Z",
        "timestamp": "2026-08-13T05:44:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "skuznetsov-clickhouse",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114567",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Check the table name length in `DatabaseOverlay`",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113019 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `clickhouse-local` accepting a `CREATE TABLE` whose name is too long for the configured `--default_database`, which produced a table that could not be dropped. `DatabaseOverlay` now checks the table name length like the on-disk database it writes to. ### Description In `clickhouse-local` with a long `--default_database`, `CREATE TABLE` was accepted and the resulting table could then never be dropped: ``` $ clickhouse local --path=$D -q \"CREATE TABLE tc (a UInt8) ENGINE=MergeTree ORDER BY a; DROP TABLE tc\" \\ -- --default_database=<214 'd's> Code: 1001. std::filesystem::filesystem_error: filesystem error: in rename: File name too long [\".../store/<uuid>/tc.sql\"] [\".../metadata_dropped/<214 d's>.tc.<uuid>.sql\"] ``` The end state survives restarts and has no SQL escape hatch: `DROP`, `DROP ... SYNC` and `DETACH` all fail the same way and the table stays in `system.tables`. The threshold is 212 characters of `--default_database`. `IDatabase::checkTableNameLength` is a no-op default and only `DatabaseOnDisk` overrides it. `DatabaseOverlay` did not, so nothing was checked, while the overlay forwarded the create to a real `DatabaseAtomic` member carrying the same long name. `DROP` then builds `metadata_dropped/{db}.{table}.{uuid}.sql`, whose prefix alone exceeds `NAME_MAX`, so the per-table budget saturates to 0. The fix overrides `checkTableNameLength` in `DatabaseOverlay`, delegating to the first non-read-only member, which is the same member `createTable` writes to. It copies the loop of the adjacent `checkMetadataFilenameAvailability`. No new setting, constant or arithmetic. A real `DatabaseAtomic` with the same name already rejects this up front with `ARGUMENT_OUT_OF_BOUND`, so this makes the overlay consistent with the database that owns the file. The change only affects rejection at DDL time; a table already on disk still attaches and reads, because both `CREATE` call sites are gated by `mode <= LoadingStrictnessLevel::CREATE`. It does not rescue tables already stranded on disk, which would mean changing the `metadata_dropped` naming convention. Related: #113019, where this was found and measured. <details> <summary>Validation: zero regression window, and the object kinds covered</summary> The delegated limit rejects exactly the names that already fail. Longest table name that survives CREATE+DROP today, versus the limit, measured at four database sizes: | escaped db length | limit | longest working today | first failing today | | --- | --- | --- | --- | | 100 | 113 | 113 | 114 | | 150 | 63 | 63 | 64 | | 190 | 23 | 23 | 24 | | 200 | 13 | 13 | 14 | At every size, a name at the limit still works after the change and a name one over is now refused at CREATE instead of being accepted and stranded. `escapeForFileName` emits 3 bytes per non-word byte and the check measures the escaped name: 37 `-` characters (111 escaped bytes) work, 38 (114) do not, against the limit of 113. All six object kinds routed through `InterpreterCreateQuery` reproduced the accepted-then-undroppable behaviour and are covered by the single override: `MergeTree`, `Log` and `Memory` tables, `VIEW`, `MATERIALIZED VIEW` and `DICTIONARY`. `CREATE TEMPORARY TABLE` is not affected, since temporary tables never reach `metadata_dropped`. In the only construction site today the overlay and its writable member share a name, so the delegated limit is the intended one. If a future overlay had a differently named writable member, delegating still gives the correct answer, because the limit follows the database that owns the file. `DatabaseOverlay` is currently built only by `clickhouse-local`; #86768 would add server-side consumers. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114567",
        "createdAt": "2026-08-13T01:34:05Z",
        "updatedAt": "2026-08-13T07:10:12Z",
        "timestamp": "2026-08-13T07:10:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114568",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `/dialect`, `/lang`, `/language` client commands",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added client-side commands `/dialect`, `/lang`, `/language` to `clickhouse-client` and `clickhouse-local` — an equivalent of `SET dialect = '...'` that is handled before parsing, so it also works when the current dialect cannot express a `SET` query (for example, to leave `kusto` or `promql`). While a non-default dialect is active, it is displayed in the prompt after the server display name, e.g. `clickhouse-cloud (polyglot) :)`. The command without an argument prints the current dialect. A plain `SET dialect = '...'` also updates the prompt, since the prompt now renders the dialect from the client-side session settings on every line. Implementation notes: - The prompt is now rendered by `getPrompt` on every line from a template that keeps the `{display_name}` placeholder: the current dialect is appended to the display name in parentheses (with a separating space if the display name is non-empty), and the `:) ` smiley is appended at the end, preserving the previous behavior for the default dialect. - The command works in `clickhouse-client`, `clickhouse-local`, and the embedded client (SSH/web terminal). - Documentation: a new \"Switching the SQL dialect\" section in the client page. - Test: `04869_client_dialect_command.expect`, verified locally along with manual checks against a running server (prompt display, alias forms, escaping from `kusto` back to `clickhouse`, and that a Kusto query parses after `/dialect kusto`).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114568",
        "createdAt": "2026-08-13T01:59:34Z",
        "updatedAt": "2026-08-13T04:36:32Z",
        "timestamp": "2026-08-13T04:36:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114569",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add test: Truncated `Array` elements stream: `CANNOT_READ_ALL_DATA` branch has no test",
        "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #109212](https://github.com/ClickHouse/ClickHouse/pull/109212). That PR: (1) Moves the `Array` offsets monotonicity check in `SerializationArray::deserializeOffsetsBinaryBulk` from `settings.native_format` to `!settings.position_independent_encoding`, and rescopes the scan to values appended by the current call (starting one element early); adds a second consistency … **1. Truncated `Array` elements stream: `CANNOT_READ_ALL_DATA` branch has no test** `src/DataTypes/Serializations/SerializationArray.cpp:537`, `src/DataTypes/Serializations/SerializationArray.cpp:545` **Risk:** `SerializationArray::deserializeBinaryBulkWithMultipleStreams` restructured the offsets/elements consistency check into a nested `if`: the `!nested_column->empty()` arm (`SerializationArray.cpp:536-538`) throws `CANNOT_READ_ALL_DATA`, the new `!settings.position_independent_encoding` arm throws … **Unique vs PR tests:** The PR's gtest covers decreasing offsets and the empty-elements case by calling the serialization directly; nothing covers a short-but-non-empty elements column, and nothing exercises either new/restructured branch through `NativeReader` from SQL. cc @groeneai (author of #109212), @alexey-milovidov (merged/approved #109212) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1315` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114569",
        "createdAt": "2026-08-13T02:09:17Z",
        "updatedAt": "2026-08-13T10:01:03Z",
        "timestamp": "2026-08-13T10:01:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "clickgapai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114570",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix presentation URL in `clickhouse-git-import` help",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Documentation entry for user-facing changes The help message of `clickhouse-git-import` referenced the presentation at `https://presentations.clickhouse.com/matemarketing_2020/`. Replace it with `https://presentations.clickhouse.com/2020-matemarketing/`. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1328` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114570",
        "createdAt": "2026-08-13T02:12:26Z",
        "updatedAt": "2026-08-13T14:31:52Z",
        "timestamp": "2026-08-13T14:31:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114571",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Interpolate quantiles in the integer domain",
        "text": "Interpolate quantiles in the integer domain <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results from the interpolating quantile functions (`quantile`, `median`, `quantileExactWeightedInterpolated`, `quantileInterpolatedWeighted` and their `quantiles*` forms) over integer-backed types such as `DateTime64` and `Decimal`. The interpolated value was computed as a `Float64` and then narrowed, which loses ticks above magnitude 2^53 and is undefined near the top of the range: `quantile` over a `DateTime64(9)` column at 2262-04-11 returned a date in 1677. `quantileInterpolatedWeighted` also subtracted the endpoints in the column's own type, which overflows once they span more than it, so it could be wrong at any magnitude: over a full-range `Int8` column it returned `-128`. Related: https://github.com/ClickHouse/ClickHouse/pull/44067 ### Description Interpolating quantiles computed the result as a `Float64` and narrowed it to the native integer type. Three defects follow. `Int64::max` has no exact `Float64` representation (the nearest is 2^63, one past it), so the cast is undefined. On x86 it wrapped to `Int64::min`: ```sql SELECT quantile(x) FROM (SELECT fromUnixTimestamp64Nano(9223372036854775807, 'UTC') AS x FROM numbers(2)); -- 1677-09-21 00:12:43.145224192, was 2262-04-11 23:47:16.854775807 ``` On AArch64 it saturates, so only UBSan complained. Above magnitude 2^53, which a nanosecond timestamp passes by two orders of magnitude, the `Float64` spacing exceeds one, so the middle tick between two ordinary values was unrepresentable. Third, `quantileInterpolatedWeighted` subtracted the endpoints in the column's own type, so it was wrong far below 2^53 too: `-128` over a full-range `Int8` column. Each family loses the value inside its own callee, so a fix at the narrowing site cannot work. Each now keeps its own `Float64` expression while both endpoints are below magnitude 2^53 and, for the subtracting family, their difference fits, so results there are unchanged otherwise. Two equal endpoints, and a level equal to a sample's own percentile, return that endpoint directly rather than through the expression, so a result there can move by one, to the true value. Outside it the offset from the lower endpoint is exact in the unsigned domain and never exceeds the endpoint distance, so the result stays inside. Coincident endpoints on the `Float64` result path are also returned directly, so an infinite one no longer forms `inf * 0.0`. `NO_SANITIZE_UNDEFINED` is removed from `QuantileInterpolatedWeighted::interpolate`, whose suppression hid this. `QuantileExactWeighted` keeps its own, now confined to the endpoint selection whose position cast is defective near `UInt64::max` already.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114571",
        "createdAt": "2026-08-13T02:18:38Z",
        "updatedAt": "2026-08-13T10:36:19Z",
        "timestamp": "2026-08-13T10:36:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114572",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #112498 to 25.8: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114572",
        "createdAt": "2026-08-13T02:34:01Z",
        "updatedAt": "2026-08-13T02:34:09Z",
        "timestamp": "2026-08-13T02:34:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "Algunenano",
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114573",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #112498 to 26.3: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114573",
        "createdAt": "2026-08-13T02:34:42Z",
        "updatedAt": "2026-08-13T02:34:49Z",
        "timestamp": "2026-08-13T02:34:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "Algunenano",
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114574",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #112498 to 26.5: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114574",
        "createdAt": "2026-08-13T02:35:19Z",
        "updatedAt": "2026-08-13T02:35:26Z",
        "timestamp": "2026-08-13T02:35:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "Algunenano",
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114575",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112498 to 26.6: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.35` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114575",
        "createdAt": "2026-08-13T02:35:48Z",
        "updatedAt": "2026-08-13T16:21:33Z",
        "timestamp": "2026-08-13T16:21:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll2",
        "state": "closed",
        "assignees": [
          "Algunenano",
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114576",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112498 to 26.7: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.25` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114576",
        "createdAt": "2026-08-13T02:36:15Z",
        "updatedAt": "2026-08-13T06:34:50Z",
        "timestamp": "2026-08-13T06:34:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll2",
        "state": "closed",
        "assignees": [
          "Algunenano",
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114577",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Mongo queries: keep whole documents in a JSON column",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/68493 Related: https://github.com/ClickHouse/ClickHouse/issues/58394 Adds two ways to run MongoDB queries against ClickHouse: - a **wire protocol endpoint** (`mongo_port`), so MongoDB drivers and tools such as `pymongo` and `mongosh` can connect to ClickHouse as if it were a MongoDB server; - a **query dialect** (`SET dialect = 'mongo'`), so MongoDB shell syntax can be sent over the usual ClickHouse interfaces. This continues @scanhex12's https://github.com/ClickHouse/ClickHouse/pull/68493 - it holds all of its commits - and changes how a collection is stored, which is what the review of that pull request asked for (https://github.com/ClickHouse/ClickHouse/pull/68493#discussion_r3771132310): a collection the endpoint creates keeps whole documents in one `JSON` column rather than a column inferred per field of the first document. ```sql CREATE TABLE db.users (`_id` String, `json` JSON) ENGINE = MergeTree ORDER BY `_id` ``` The document goes into the `JSON` column, whose paths are its fields, and its object id into the `_id` column, which is the primary key. A document may then hold a field that no document before it had, may leave out a field another one has, and `$exists` answers what it means. A table that was created in ClickHouse keeps its own columns, and a field of a query names the column of the same name there, so an application can also be pointed at a table that already holds the data. A read of the documents as they are stored selects them next to the types of their paths, so a date reads back as a BSON date and an integer as an `int32`/`int64` of its width. **This is a draft**: the storage change is done and verified by hand against `pymongo` (insert, `find` with filters, projections, `$exists`, nested paths, sort, limit, `count`, `distinct`, `aggregate`, `delete`, `createCollection`, a read by `_id`, the round trip of a date), but the integration tests of `test_mongo_protocol` still pin the previous shape and have to be reworked, and four things of a document collection are an explicit error rather than an answer: an `update` of it, an index over a path, the `mongo` dialect over it, and the element-wise match of an array by equality or `$in`. They are the next commits; the limitations are listed in `docs/en/interfaces/mongo.md` in the meantime. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a MongoDB-compatible wire protocol endpoint (the `mongo_port` server setting) and a MongoDB query dialect (`SET dialect = 'mongo'`) for basic collection operations. A collection the endpoint creates keeps whole documents in a `JSON` column.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114577",
        "createdAt": "2026-08-13T02:49:29Z",
        "updatedAt": "2026-08-13T05:45:23Z",
        "timestamp": "2026-08-13T05:45:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114578",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not erase the source column of a filter deferred after FINAL",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114512 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `THERE_IS_NO_COLUMN` / `NOT_FOUND_COLUMN_IN_BLOCK` when a filter that runs after `FINAL` is a plain column reference. This affected an explicit `PREWHERE b` under `FINAL` and, with default settings, a row policy whose expression is a bare column. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/114512 `SELECT count() FROM t FINAL PREWHERE b` threw `Code: 8. Cannot find column 'b' in source stream`. No `SETTINGS` clause is needed: `apply_row_policy_after_final` defaults to 1, so a row policy defers `PREWHERE` past `FINAL`. Root cause: a filter deferred after `FINAL` runs in a `FilterTransform` built with `remove_filter_column = true`. When the predicate is a bare column reference, the predicate node *is* the DAG's own input node, so the transform erases the source column the stream carries. The existing `restoreDAGInputs` guard cannot help: it only re-adds an input that is not already an output, and a bare predicate is already one. A wrapped predicate (`b = 1`, `NOT b`) has a distinct result node, so erasing it leaves the input alone. Hence only the bare form failed. The fix clears the local `remove_column` flag in `add_deferred_filter` when the filter column is one of the DAG's own inputs, which is how `optimizePrewhere.cpp` already resolves the same case. That lambda is shared by the deferred row policy and the deferred `PREWHERE`, so one change covers both. The defect is broader than reported: a bare-column row policy with no `PREWHERE` anywhere in the query fails identically under default settings. That shape is in the test. Validated in both directions on a debug build. Every shape in the issue's matrix now returns what its `WHERE` equivalent returns, and the new test fails on an unfixed binary. Covered: all four projections, subcolumns, a stacked policy plus `PREWHERE`, the four FINAL engines, and `Nullable`/`LowCardinality`/`Float` predicates. `SELECT *` stays byte-identical to the `WHERE` equivalent. A versioned fixture whose predicate flips across deduplication pins that the filter runs after `FINAL`, on the source column. Pre-existing and unchanged here: a row policy on a bare subcolumn fails `NOT_FOUND_COLUMN_IN_BLOCK` on the non-deferred read path, with and without `FINAL`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114578",
        "createdAt": "2026-08-13T02:56:02Z",
        "updatedAt": "2026-08-13T12:43:49Z",
        "timestamp": "2026-08-13T12:43:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "yariks5s"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114580",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add test: Duplicate TLS argument rejection and positional-arity stripping untested",
        "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #110615](https://github.com/ClickHouse/ClickHouse/pull/110615). That PR: (1) Adds TLS/SSL to every PostgreSQL integration: `sslmode` plus certificate/key either as server-local paths (`sslrootcert`/`sslcert`/`sslkey`, config-only) or as literal contents (`*_pem`, accepted from SQL, materialized into `TemporarySecretFile` and masked as secrets). New code: … **1. Duplicate TLS argument rejection and positional-arity stripping untested** `src/Storages/StoragePostgreSQL.cpp:726`, `src/Databases/PostgreSQL/DatabasePostgreSQL.cpp:573` **Risk:** `StoragePostgreSQL::extractSSLParamsFromArguments` strips trailing TLS `key = value` pairs and rejects repeats at `StoragePostgreSQL.cpp:726-727`; the stripped list then feeds the arity check at `DatabasePostgreSQL.cpp:573`. Risk if broken: a repeated `sslmode = 'require', sslmode = 'disable'` … **Unique vs PR tests:** 04820 formats queries with distinct TLS keys and checks masking plus path rejection; 04846 checks query-tree masking; test_postgresql_ssl exercises real handshakes through named collections. None repeats a TLS key (the `specified more than once` branch) and none combines the maximum positional … **Tags:** `-- Tags: no-fasttest` — `no-fasttest`: the PostgreSQL integration is not built in the fast test build; the PR's own 04820 uses the same tag cc @alexey-milovidov (author of #110615) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114580",
        "createdAt": "2026-08-13T03:32:59Z",
        "updatedAt": "2026-08-13T16:45:10Z",
        "timestamp": "2026-08-13T16:45:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested"
        ],
        "author": "clickgapai",
        "state": "open",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114583",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #107028 to 25.8: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114583",
        "createdAt": "2026-08-13T03:40:10Z",
        "updatedAt": "2026-08-13T03:40:19Z",
        "timestamp": "2026-08-13T03:40:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "kssenii",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114584",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #107028 to 26.3: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114584",
        "createdAt": "2026-08-13T03:40:48Z",
        "updatedAt": "2026-08-13T03:40:57Z",
        "timestamp": "2026-08-13T03:40:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "kssenii",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114585",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #107028 to 26.5: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114585",
        "createdAt": "2026-08-13T03:41:21Z",
        "updatedAt": "2026-08-13T05:48:10Z",
        "timestamp": "2026-08-13T05:48:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "kssenii",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114586",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #107028 to 26.6: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114586",
        "createdAt": "2026-08-13T03:41:46Z",
        "updatedAt": "2026-08-13T07:33:42Z",
        "timestamp": "2026-08-13T07:33:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll2",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "kssenii",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114589",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "sync ErrorCodes.cpp from private",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1318` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114589",
        "createdAt": "2026-08-13T04:01:46Z",
        "updatedAt": "2026-08-13T10:57:05Z",
        "timestamp": "2026-08-13T10:57:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested",
          "pr-synced-to-cloud"
        ],
        "author": "chhetripradeep",
        "state": "closed",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114590",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add test: JSON shared data at the max 256 buckets is untested end-to-end",
        "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #112172](https://github.com/ClickHouse/ClickHouse/pull/112172). That PR: Rewrites the write path of bucketed `JSON` shared data. `splitSharedDataPathsToBuckets` and `flattenAndBucketSharedDataPaths` are replaced by a single `SharedDataBucketsSplitter` that computes each path's bucket once (cached in `PODArray<UInt8>`) plus per-bucket byte sizes, then hands out one … **1. JSON shared data at the max 256 buckets is untested end-to-end** `src/DataTypes/Serializations/SerializationObjectHelpers.cpp:88`, `src/DataTypes/Serializations/SerializationObjectHelpers.cpp:119` **Risk:** This PR replaces `splitSharedDataPathsToBuckets` (which kept the bucket index in a `size_t`) with `SharedDataBucketsSplitter`, which caches the bucket of every path in a `PODArray<UInt8>` via `static_cast<UInt8>(bucket)` (`SerializationObjectHelpers.cpp:88`). The value range now exactly saturates … **Unique vs PR tests:** The PR's own `04617_array_json_shared_data_buckets_insert_memory.sh` is a memory-limit test at the default bucket counts (8/32) that only asserts `count()` and is skipped on every sanitizer build. This test pins the maximum legal bucket count (256), which no test in either tree uses, and asserts … [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/4753df1e-4460-4e86-9262-9fdbe47e2cdb) cc @Avogar (author of #112172), @thevar1able (merged/approved #112172) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114590",
        "createdAt": "2026-08-13T04:14:22Z",
        "updatedAt": "2026-08-13T11:00:48Z",
        "timestamp": "2026-08-13T11:00:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested"
        ],
        "author": "clickgapai",
        "state": "open",
        "assignees": [
          "tiandiwonder"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114594",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "[Draft/Prototype] Add mutations_restrict session setting as a mutation safety catch",
        "text": "> ⚠️ **Experimental prototype — heavy LLM assistance.** > > This PR is an early prototype built with substantial LLM assistance > (GitHub Copilot CLI). Early draft, sharing for searchability and > to use CI support. ## Summary Adds a new session-level `UInt64` setting **`mutations_restrict`** that acts as a safety catch against accidental mutations. Modeled on the existing \"block all DDL\" switch `allow_ddl`, but scoped to mutation-producing statements and shaped as tiers like `readonly` / `mutations_sync`: | Value | Effect | |-------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | `0` | No restriction (default; existing behavior). | | `1` | Reject `ALTER TABLE` forms that produce a `system.mutations` entry: `UPDATE`, `DELETE`, `MATERIALIZE INDEX`/`PROJECTION`/`COLUMN`/`STATISTICS`/`TTL`, `APPLY DELETED MASK`, `APPLY PATCHES`, `DROP INDEX`/`PROJECTION`/`STATISTICS`, `RENAME COLUMN`, `DROP COLUMN` on a physical column, non-metadata `MODIFY COLUMN`, `CLEAR ... IN PARTITION`. Metadata-only `ALTER`s, `INSERT`, and `SELECT` still succeed. Lightweight `DELETE`/`UPDATE` still succeed. | | `2` | Additionally reject standalone lightweight `DELETE` and `UPDATE`. | The setting complements the server-level `disable_insertion_and_mutation` (which is global and also blocks `INSERT`). Unlike that setting, `mutations_restrict` is a normal `Settings` entry, so it can be: - set per session with `SET mutations_restrict = N;`, - set per user via profile, - pinned system-wide or per user with a `<readonly/>` constraint (same pattern as `allow_drop_detached`). Typical intended usage: users pin `1` (or `2`) as a soft default in their profile and lower it in-session when they need to mutate; admins pin it as `<readonly/>` for hard enforcement. ## Rationale There is currently no per-session safety catch for accidental mutations. Options in-tree today: - `disable_insertion_and_mutation` (server-scoped, blocks inserts too), - RBAC (`ALTER UPDATE`, `ALTER DELETE`, ... as separate grants — heavy for a \"guard against fat-fingering\" workflow), - `readonly = 1` (blocks everything, not just mutations). None of these fit \"let me use this session but refuse to run heavy mutations by mistake\". ## Where the checks live - **Tier 1** — `InterpreterAlterQuery::execute` after the `CommandSegments` are built and the existing `validateMutationsAllowed` runs. Uses `MutationCommands::hasNonEmptyMutationCommands`, so the schema-dependent cases (e.g. `MODIFY COLUMN` where the type change is a metadata-only conversion) are resolved correctly — the interpreter is the earliest point where the mutation-vs-metadata question has been answered. - **Tier 2** — `InterpreterDeleteQuery::execute` and `InterpreterUpdateQuery::execute`, at the top of `execute`. The alternative of gating in `ContextAccess::checkAccessImplHelper` alongside `allow_ddl` was considered and rejected: it would require introducing mutation-specific `AccessFlags`, is a larger refactor, and cannot distinguish schema-dependent cases. ## What's not covered (by design) - `SYSTEM ...` commands. - Partition manipulation (`ATTACH`/`DETACH`/`DROP PARTITION`). - `OPTIMIZE`, `TRUNCATE`, `CREATE`/`DROP TABLE`. - `INSERT`. These do not appear in `system.mutations`. Documented as such in the setting docstring. ## Tests - `tests/queries/0_stateless/04869_mutations_restrict_safety_catch.sql` — tier 0/1/2 behavior + in-session override + mutation vs metadata-only ALTER distinction. - `tests/integration/test_mutations_restrict/` — profile pin + `<readonly/>` enforcement, plus a \"soft default\" profile that demonstrates in-session override. ## Docs - Setting reference under `docs/reference/settings/session-settings/mutations.mdx` will regenerate from the `DECLARE(...)` docstring in `src/Core/Settings.cpp`. - Non-autogenerated section added to `docs/reference/statements/alter/index.mdx` cross-referencing the new setting. ## Known open questions (please review) 1. **Naming.** `mutations_restrict` vs `restrict_mutations` vs `mutation_safety_catch`. Chose `mutations_restrict` to group with the `mutations_*` family and match the \"higher = more restrictive\" direction of `readonly`. 2. **Tier semantics.** Currently tier 1 blocks `ALTER TABLE ... UPDATE` even when `alter_update_mode` would rewrite it to a lightweight update. Rationale: it's syntactically an `ALTER`. Alternative: move the tier-1 check to run after the lightweight rewrite decision, so it only fires on the heavy path. 3. **Standalone `DELETE FROM ...` heavy path.** For engines where `supportsDelete()` returns true (e.g. `KeeperMap`, `RocksDB`) the standalone `DELETE FROM` still calls `table->mutate` — this is currently only blocked at tier 2 (via the top-of-execute check). Should it also be blocked at tier 1? (Unclear whether the produced mutation is user-observable in the same \"system.mutations\" sense.) 4. **`SettingsChangesHistory.cpp`.** Added to the `26.8` block; the correct release for master at merge time may differ. ## Not tested locally The sandbox that authored this branch has no C++ toolchain, so I could not build, run stateless tests, or run the integration test locally. Relying on CI. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new session setting `mutations_restrict` as a safety catch against accidental mutations. Value `1` rejects `ALTER TABLE` forms that would create an entry in `system.mutations`; value `2` additionally rejects standalone lightweight `DELETE` and `UPDATE`. Can be set per session, per user profile, or pinned system-wide with a `<readonly/>` constraint.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114594",
        "createdAt": "2026-08-13T05:04:28Z",
        "updatedAt": "2026-08-13T05:47:48Z",
        "timestamp": "2026-08-13T05:47:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [],
        "author": "ringerc",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114596",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix target access checks for the Alias engine",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> `Alias` table now requires `SHOW COLUMNS` on its target when a new definition is submitted. This prevents users from using `Alias` to reveal a target table's schema. The trivial `count` optimization is disabled when the user has no `SELECT` privilege on the target, allowing the normal read path to enforce access checks. `rows` and `bytes` statistics are also hidden unless the user has `SHOW TABLES` on the target. Closes: https://github.com/ClickHouse/ClickHouse/issues/114511 ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an RBAC bypass in the `Alias` table engine that allowed users without privileges on the target table to reveal its schema, row count, size, and existence.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114596",
        "createdAt": "2026-08-13T05:30:42Z",
        "updatedAt": "2026-08-13T16:12:31Z",
        "timestamp": "2026-08-13T16:12:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-must-backport",
          "can be tested",
          "pr-critical-bugfix"
        ],
        "author": "nauu",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114597",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix zero-capacity hash-table statistics cache seeded by lazy FINAL",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/101647 Related: https://github.com/ClickHouse/ClickHouse/pull/113333 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed hash-table statistics — and with them aggregation hash-table preallocation — being silently and permanently disabled for the whole server when a query using the lazy `FINAL` optimization (`query_plan_optimize_lazy_final`) was the first aggregation to run after server startup. ### Description The `FINAL`-replacing dedup aggregation of the lazy `FINAL` optimization (`LazyReadReplacingFinalSource`, #101647) built its `AggregatingStep` with default-constructed `StatsCollectingParams`, i.e. `max_entries_for_hash_table_stats = 0`. The query-plan optimization pass `setAggregationHashTableCacheKeys` stamps a cache key onto every keyed `AggregatingStep`, which enables statistics collection. The process-wide statistics cache is created lazily by the first query that touches it, with `max_entries_for_hash_table_stats * sizeof(Entry)` as its capacity — so when a lazy `FINAL` query was the first aggregation in a freshly started server, the cache was created with zero capacity: every insertion was evicted immediately and every lookup missed, permanently disabling hash-table statistics (and preallocation) for the whole server lifetime. In the affected CI run's server log the aggregation statistics cache had 3515 `Statistics updated` entries and zero `found in cache` hits over the entire lifetime, while the join statistics cache (a separate instance created with proper parameters) worked normally. This stayed unnoticed while keyless aggregations, which carry properly constructed params, also touched the cache and almost always created it first. #113333 (merged 2026-08-12 19:00 UTC) stopped keyless aggregations from touching the cache, after which the lazy `FINAL` tests started to win the first-toucher race, and `04625_hash_table_sizes_stats_table_expression_modifiers` began to fail with `final-prealloc 0` / `no-final-prealloc 0` across unrelated jobs (`amd_msan`, `amd_asan_ubsan, distributed plan`, `Fast test`, `Fast test (arm_darwin)`) starting 2026-08-12 22:39 UTC — 6 failures, no failures of this kind before that date. Because the poisoned cache is server-instance state, the flaky-check diagnosis reproduced the failure 5/5 within the affected run while adjacent master commits passed. CI reports: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=2790ac4a68cf412d05db2d1573685a1a8467ecdb&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_msan%2C%20WasmEdge%2C%20parallel%2C%201%2F3%29 - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=86363e819f9cb704a892063da9998e4c1281eb76&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_msan%2C%20WasmEdge%2C%20parallel%2C%201%2F3%29 The fix: - `LazyReadReplacingFinalSource` now constructs real `StatsCollectingParams` (same pattern as `Planner`), so the dedup aggregation also benefits from size hints on repeated `FINAL` queries. - `setAggregationHashTableCacheKeys` skips steps whose builder left `stats_collecting_params` unconfigured (`max_entries_for_hash_table_stats == 0`), so a stamped key can never enable collection with a zero-capacity configuration; this also makes an admin-configured server-level `max_entries_for_hash_table_stats = 0` a clean disable for the aggregation path instead of a zero-capacity cache. - The new test `04869_lazy_final_hash_table_stats_cache` runs the first-toucher scenario deterministically in `clickhouse-local` (a fresh process owns a fresh statistics cache): the lazy `FINAL` query runs first, then a 650000-group victim aggregation twice; the second run must preallocate, observed through the `AggregationPreallocatedElementsInHashTables` event in `system.events`. On the unfixed binary this deterministically reports nothing; with the fix it reports 650000. Verified locally: the `clickhouse-local` repro flips from empty to 650000 with the fix; `04625_hash_table_sizes_stats_table_expression_modifiers`, the new `04869` test and the lazy `FINAL` family (`03988`, `03990`, `03991`, `04092`, `04093`) all pass against a patched server, including the exact CI scenario (fresh server, `03990_lazy_final_index_analysis` first, then `04625`). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114597",
        "createdAt": "2026-08-13T05:56:02Z",
        "updatedAt": "2026-08-13T13:10:04Z",
        "timestamp": "2026-08-13T13:10:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114599",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "AST fuzzer: do not create a view that duplicates a non-parallel sink",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110166 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/110166 Reported by @ alexey-milovidov there: a `Stress test (arm_release)` hung check on an `INSERT` stuck in `StorageLog::write`. https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110166&sha=c890bb8e35d0d9138c1fa770e1fd4a1db05f39bf&name_0=PR&name_1=Stress%20test%20%28arm_release%29 It is not lock contention under load. Fuzzing a `CREATE MATERIALIZED VIEW` renames the view but leaves its external `TO` untouched, so the clone shares source *and* target with the original and one `INSERT` builds two sinks for that target. `LogSink` holds the table's exclusive lock from pipeline build until `onFinish`, so the second sink waits out `lock_acquire_timeout` and throws `Code: 159`; the `INSERT` re-arms, the processlist never empties, and the hung check fails. Reproducible in 4 statements with no concurrency and no fuzzer, on three build flavours. This skips such a fuzzed query instead of creating the view. Whether it is hazardous is a sink-graph question, so it is answered in `InsertDependenciesBuilder`, which owns that graph, and it errs only towards skipping: safe shared-target fuzzing keeps running, and a sole writer of a non-parallel target is unaffected. `Buffer` and `Distributed` end a branch, since the write they forward runs as an `INSERT` of their own whose sinks this walk does not model. Anything the walk cannot decide executes normally and is counted, so an undecided answer never suppresses fuzzing. Two `ProfileEvents` make both outcomes attributable; the path is confined to `ast_fuzzer_runs > 0`. A proven duplicate survives a later branch that throws while materializing a lazy storage. And since the answer describes the catalog, where a view appears only once its `CREATE` has executed, the fuzzer reserves the (source, target) pair across the decision and the create; the contract states that boundary. Refusing two writers on one `Log`-family target at INSERT planning time is out of scope: that is user-visible, and on #74080 the same topology was answered as unsupported. <details><summary>Validation</summary> Ground truth measured on a pre-fix binary by building the duplication by hand, against the predicate's verdict on the same 9 topologies (agreement is 9/9): | topology | wedges pre-fix | skipped | |---|---|---| | `src -> mv -> tgt(TinyLog)` | yes | yes | | `src -> mvA -> mt(MergeTree) -> mvB -> lg(TinyLog)` | yes | yes | | `Alias`-hidden: `mv TO Alias(mt)`, `mt -> mv2 -> lg` | yes | yes | | deep cascade `mt1..mt8 -> lg` | yes | yes | | `Alias(MergeTree) -> view -> Memory` | no | no | | directly shared `Buffer` | no | no | | directly shared `Distributed` | no | no | | stale edge (`MODIFY QUERY` away from `mt`) | no | no | | shared `MergeTree` | no | no | Plus the hazard behind a lazily loaded table (`lazy_load_tables`, wedges pre-fix, skipped) and its `MergeTree` counterpart (skipped in neither). Through the real carrier (`--ast_fuzzer_runs=40 --ast_fuzzer_any_query=1`): pre-fix the `INSERT` fails `Code: 159`; with the fix the hazardous arms are skipped, the insert completes, and the safe arms are untouched. Removing the skip restores the failure; removing the name withdrawal makes 345 of 480 later fuzzed queries name a view that was never created (0 with it). `04876_ast_fuzzer_skips_duplicate_non_parallel_sink` states each claim against its own fixture's query_log rows rather than server-global counters, so a concurrent fuzzing process cannot move them: 30/30 green under a deliberate adversary, 50/50 randomized green, 32 consecutive runs in one database with nothing left behind, and eight mutations each flipping exactly their own assertion. It fails with the skip removed (`Code: 159`) and with the proxy hop reverted. </details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114599",
        "createdAt": "2026-08-13T07:00:18Z",
        "updatedAt": "2026-08-13T09:16:52Z",
        "timestamp": "2026-08-13T09:16:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114600",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add is_nullable to system.columns, use it in information_schema",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/44930 Related: https://github.com/ClickHouse/ClickHouse/pull/48560 `information_schema.columns` decided nullability by matching the printed type name (`type LIKE 'Nullable(%)'`), so a `LowCardinality(Nullable(T))` column — which accepts `NULL` — was reported as `is_nullable = 0`, and MySQL-compatible clients inferred it as `NOT NULL`. On #48560 @alexey-milovidov said about that approach: > I don't like either the modification in this PR or the original code. Because it is doing ad-hoc, > rough text parsing, like what you typically do with Perl. The right way to do it is to introduce > `is_nullable` column in the `system.columns` and reuse it here. So this does that: `system.columns` gains `is_nullable UInt8`, computed from the data type via the existing `isNullableOrLowCardinalityNullable()`, and the view selects it. The text parsing is removed rather than extended, matching how `numeric_precision`, `numeric_scale`, `datetime_precision` and `character_octet_length` are already computed in C++ and passed through. `IS_NULLABLE` is an alias of `is_nullable`, so it is fixed too. A column reports `is_nullable = 1` exactly when it accepts `NULL`: | type | accepts NULL | is_nullable | |---|---|---| | `Nullable(T)` | yes | 1 | | `LowCardinality(Nullable(T))` | yes | 1 | | `LowCardinality(T)` | no | 0 | | `Array(Nullable(T))`, `Tuple(Nullable(T))`, `Map(K, Nullable(V))` | no | 0 | Composite types stay `0` because the column itself cannot hold `NULL`, only its elements can. New test `04669_information_schema_is_nullable_lowcardinality` covers both `system.columns` and the view. Regenerated `02117_show_create_table_system`, `02206_information_schema_show_database` and `01602_temporary_table_in_system_tables` for the added column. `SimpleAggregateFunction(any, Nullable(T))`, `Variant(...)` and `Dynamic` also accept `NULL` while reporting `0`. Left for a follow-up — now a change in one place. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `information_schema.columns` reporting `is_nullable = 0` for `LowCardinality(Nullable(T))` columns. Add an `is_nullable` column to `system.columns`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114600",
        "createdAt": "2026-08-13T07:31:59Z",
        "updatedAt": "2026-08-13T09:58:54Z",
        "timestamp": "2026-08-13T09:58:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "zainulabidin302",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114601",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix a renamed and dropped column being read from the wrong file while the mutation is pending",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114562 ## Problem A column that is renamed, dropped and added again in one `ALTER` silently returns the **dropped** column's data instead of its own default, while the mutation is still pending: ```sql CREATE TABLE t (a UInt64, h UInt8 DEFAULT 0) ENGINE = MergeTree ORDER BY tuple() SETTINGS min_bytes_for_wide_part = 0, min_rows_for_wide_part = 0; INSERT INTO t SELECT number + 100, 0 FROM numbers(1000); SET alter_sync = 0; -- mutation left pending ALTER TABLE t RENAME COLUMN a TO b, DROP COLUMN b, ADD COLUMN b UInt64 DEFAULT 7; SELECT count(), min(b), max(b) FROM t; -- 1000 100 1099 <- the old `a` values -- expected: 1000 7 7 ``` No error and nothing in the log; it affects `Wide` and `Compact` parts alike. The added `b` is a different column from the one that was dropped, so its rows have to come from its `DEFAULT`. ## Root cause `AlterConversions` records a pending drop under the name the command used: ```cpp else if (command.type == DROP_COLUMN) { dropped_columns.emplace(command.column_name); } ``` while every consumer of `isColumnDropped` asks about a column under the name it has **in the part**: - `IMergeTreeReader::isColumnDroppedByPendingMutation` takes the name out of `columns_to_read`, which is built from `getColumnInPart` and therefore already has the rename mapping applied (`IMergeTreeReader.cpp:81`); - `injectRequiredColumnsRecursively` passes `column_name_in_part` (`MergeTreeBlockReadUtils.cpp:119`). Without a rename the two names are identical, so the hazard those checks exist for — a column dropped and re-added under the same name, whose data in the part is stale — behaves correctly. That is why this was never noticed. With a rename they diverge: the part holds `a`, the drop is recorded for `b`, the check asks about `a`, finds nothing, and the reader streams the `a` file into the new `b`. ## Solution Resolve the drop to the part-side name while the commands are still in order. Commands arrive as issued, so the renames recorded so far are exactly those preceding the drop: ```cpp auto name_in_part = command.column_name; for (const auto & entry : rename_map) { if (entry.rename_to == name_in_part) { name_in_part = entry.rename_from; break; } } dropped_columns.emplace(std::move(name_in_part)); ``` Every existing caller already passes a part-side name, so all of them become correct at once and none of them changes. Chained renames are collapsed into one entry by `addMutationCommand`, so `a -> b -> c` resolves `c` straight back to `a`. The same ordering has to retire a rename mapping whose source was dropped, once that name is taken over. `RENAME a TO b, DROP b, RENAME c TO b` left **two** entries pointing at `b` — from the dropped `a` and from the live `c` — and the lookup returns whichever comes first, so `b` read the wrong column: ```cpp std::erase_if(rename_map, [&](const RenamePair & entry) { return entry.rename_to == command.rename_to && dropped_columns.contains(entry.rename_from); }); ``` This one is pre-existing too, and both halves are needed to get it right: | variant | `min(b), max(b)`, want `5000 5999` | |---|---| | `master` | `100 1099` — the **dropped** column's data | | part-side drop only | `0 0` — the default | | both | `5000 5999` ✔ | ### Why not fix it in the reader Checking the renamed-to name in the reader as well looks equivalent and is wrong. The opposite command order produces exactly the same `AlterConversions` state, and `01281_alter_rename_and_other_renames` already pins the required behaviour for it: ```sql ALTER TABLE t DROP COLUMN value2_old, RENAME COLUMN value2 TO value2_old; ``` Here the drop hit a *different* column that merely carried that name, and `value2`'s data is live and must be readable as `value2_old`. The only thing separating the two situations is where the drop sits among the renames, which the readers cannot see. Measured with the reader-side variant applied, live data is replaced by defaults: ``` dropped, then the name reused by a rename -1000 100 1099 +1000 0 0 ``` Hence the ordering has to be consumed at construction time. ### Tests `tests/queries/0_stateless/04872_read_renamed_then_dropped_and_readded_column.sh` — six cases: the defect on `Wide` and on `Compact`, the `RENAME/DROP/RENAME` case above, plus three controls that must not move — drop-and-re-add without a rename (what the check was written for), a pending rename with no drop, and the mirror case. Because the two possible wrong answers fail in opposite directions, the result is checked three ways: | variant | the defect | the mirror case | |---|---|---| | `master` | `1000 100 1099` ✗ | `1000 100 1099` ✔ | | reader-side variant | `1000 7 7` ✔ | `1000 0 0` ✗ | | this change | `1000 7 7` ✔ | `1000 100 1099` ✔ | `01281_alter_rename_and_other_renames`, `01278_alter_rename_combination`, `03905_chained_rename_column_mutation`, `04053_alter_add_rename_column_in_single_query` and `04011_detach_rename_attach_column` all pass. ### Notes for reviewers - `isColumnDropped` now consistently answers about part-side names. That matches all four existing call sites, so it is a fix rather than a contract change for them, but a future caller holding a metadata name would have to resolve it first. - **Overlaps with #114562, in a way that is safe in either merge order.** That PR adds a second lookup in the empty-candidate accounting loop of `injectRequiredColumns`, mapping a part name forward to its renamed-to name and testing the drop under that, so that a single-statement `RENAME a TO b, DROP b` is recognised as accounted for. This change makes the *first* lookup answer that correctly, because the drop is then already recorded as `a`, so the second lookup becomes dead code once both are in. It is deliberately left in place here rather than removed pre-emptively: while only #114562 is merged the second lookup is what makes that case work. A small cleanup to delete it belongs after this merges. Checked on a scratch branch carrying both changes with that lookup deleted — `04870`, `04871`, `04872`, `03830`, `04011` and `01281_alter_rename_and_other_renames` all still pass, so the cleanup is safe. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a column that was renamed, dropped and added again in one `ALTER` returning the dropped column's values instead of its own default while the mutation was still pending.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114601",
        "createdAt": "2026-08-13T07:45:51Z",
        "updatedAt": "2026-08-13T12:25:54Z",
        "timestamp": "2026-08-13T12:25:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "tiandiwonder",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114602",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: regenerate reference documentation from source",
        "text": "This pull request is opened automatically by the nightly documentation autogeneration workflow. It regenerates settings, functions, table and database engines, data types, formats, table functions, window functions, system tables, and asynchronous metrics from the structured documentation embedded in the ClickHouse source and exposed through the corresponding `system.*` tables. The generator preserves page frontmatter and hand-written content outside the `{/*AUTOGENERATED_START*/}` / `{/*AUTOGENERATED_END*/}` regions. Some pages are fully generated below their frontmatter. Do not edit generated content by hand -- edit the structured documentation in the defining source code instead; the next nightly run regenerates the pages. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1348` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114602",
        "createdAt": "2026-08-13T08:02:58Z",
        "updatedAt": "2026-08-13T16:35:49Z",
        "timestamp": "2026-08-13T16:35:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud",
          "pr-autogenerated-docs"
        ],
        "author": "clickhouse-gh[bot]",
        "state": "closed",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114604",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix flaky 02999_scalar_subqueries_bug_2",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Failure report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=2485f1496f7dff85c6468cc59d1856cec0f92a9d&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_tsan%2C%20parallel%29 Related: https://github.com/ClickHouse/ClickHouse/pull/111983 Related: https://github.com/ClickHouse/ClickHouse/pull/91744 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02999_scalar_subqueries_bug_2` checked that the scalar subquery in a materialized view definition is not executed at CREATE time by racing a 2 second `max_execution_time` against a 3 second `sleepEachRow`. That makes the test depend on machine load rather than on the property under test, and it failed on master in `Stateless tests (amd_tsan, parallel)`: ``` Code: 159. DB::Exception: Timeout exceeded: maximum: 2000 ms. (TIMEOUT_EXCEEDED) (query 10, line 17) ``` The diagnostics rerun passed 43/43 with the same randomized settings. The scalar was not executed. The message carries no `elapsed ... ms` clause, while the in-sleep deadline check (`sleep.cpp:168`) goes through `checkTimeLimit`, which always reports a non-zero elapsed; the statement was cancelled while still pending, since every query including DDL is registered with the `CancellationChecker` watchdog unconditionally (`ProcessList.cpp:381`). The engine behaved correctly and only the clock failed. `max_execution_time` is not randomized by the runner, so there is no setting to pin, and widening the bound was already tried on the structural twin (#91744) without holding. I removed the timing oracle and assert the property directly with `throwIf(1)`, following @ alexey-milovidov's fix for that twin in #111983. If the scalar is not executed `throwIf` never fires; if it ever is, the statement fails `FUNCTION_THROW_IF_VALUE_IS_NON_ZERO`. `throwIf::isSuitableForConstantFolding` returns `false` (`throwIf.cpp:94`), the same guard `sleepEachRow` relies on, so the mechanism the old oracle depended on is preserved, and the check is now stronger: it detects execution at all, not only execution slower than 2 seconds. I also cover `CREATE TABLE ... EMPTY AS SELECT`, which reaches a different analysis site, and added an executed-position arm that must fail so the check has teeth. The reference file stays empty. Validation: the new test is green 50/50 under `enable_analyzer=1` and 50/50 under `=0`, green with 8 concurrent copies, and two mutations of the new oracle each redden it. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1338` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114604",
        "createdAt": "2026-08-13T08:49:21Z",
        "updatedAt": "2026-08-13T15:35:52Z",
        "timestamp": "2026-08-13T15:35:52Z",
        "metrics": {
          "reactions": 2,
          "comments": 6
        },
        "labels": [
          "can be tested",
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114606",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "PostgreSQL: allow an empty TLS contents override when the collection stores no credential",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113947 Related: https://github.com/ClickHouse/ClickHouse/pull/110615 Port of the #113947 narrowing (MySQL) to `StoragePostgreSQL::getSSLParams`. An empty `ssl*_pem` query override is rejected only when it would actually drop a TLS credential the named collection carries — a path (a query cannot override path keys, so the value read is the collection's own) or contents, read in the pre-override form via `NamedCollection::getValueBeforeQueryOverride`. On a collection with no TLS keys at all, the empty override stays the no-op it is for the direct arguments, instead of throwing `BAD_ARGUMENTS`. The rejection cases are unchanged and remain covered by `test_path_overrides_are_rejected` and `test_tls_credentials_in_sql_named_collection`; the new `test_empty_override_without_stored_credential_is_noop` covers the no-op case (it fails without the code change). ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix of the unreleased #110615: an empty `ssl*_pem` override on a PostgreSQL named collection without TLS credentials is a no-op again instead of an error.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114606",
        "createdAt": "2026-08-13T09:11:09Z",
        "updatedAt": "2026-08-13T13:52:08Z",
        "timestamp": "2026-08-13T13:52:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114607",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix an infinite uncancellable loop in `hop` and `windowID` on an interval whose span wraps modulo 2^32",
        "text": "<!-- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed an infinite, uncancellable loop in functions `hop` and `windowID` when the span of an interval argument in seconds is a multiple of 2^32 (for example, `toIntervalDay(2147483648)`): the wrapped subtraction dodged the time-overflow check, and with constant arguments the loop ran at analysis time, where the query could not even be killed. An interval like `toIntervalDay(2147483648)` spans `2^31 * 86400 = 43200 * 2^32` seconds, so subtracting it in the wrapping `UInt32` arithmetic of `AddTime` is a no-op: `wstart == wend`, which dodges the `wstart > wend` time-overflow guard added by https://github.com/ClickHouse/ClickHouse/pull/61523. The window-searching loop of `executeHop` then decrements `wend` by one hop at a time past zero, where it wraps back to ~2^32 while staying congruent to its starting point modulo `gcd(86400, 2^32) = 128`, never hits a value `<= time`, and spins forever. With constant arguments the loop runs during constant folding in `QueryAnalyzer::resolveFunction`, at analysis time, where nothing checks the cancellation. The same loop shape exists in `executeHopSlice` (`windowID`). The fix makes the guard `wstart >= wend` - subtracting a whole positive interval must change the time, and equality only happens on a wrap - and adds a wrap check to the loop itself (the new `wend` coming out greater than the old one throws the same `Time overflow` error), which terminates every wrapping loop regardless of how the pre-loop values were corrupted. The new test `04891_time_window_functions_wrapped_interval_overflow` covers both functions with the fuzzed arguments (both hung before the fix) and checks that a sane `hop`/`windowID` is unaffected. Found by the AST fuzzer in the stress test of a CI run of https://github.com/ClickHouse/ClickHouse/pull/112930 (hung check): [Stress test (arm_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112930&sha=1c5bf0b613146bd6666b6d2bd2be55f539f6c0d8&name_0=PR&name_1=Stress%20test%20%28arm_debug%29). Closes: https://github.com/ClickHouse/ClickHouse/issues/114605 Related: https://github.com/ClickHouse/ClickHouse/issues/61521 Related: https://github.com/ClickHouse/ClickHouse/pull/61523",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114607",
        "createdAt": "2026-08-13T09:11:49Z",
        "updatedAt": "2026-08-13T13:51:50Z",
        "timestamp": "2026-08-13T13:51:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114608",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Make the 04780 index-analysis allocation oracle the minimum of several runs",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/104948 `04780_json_subcolumn_index_match_not_quadratic` compares `MemoryAllocatedWithoutCheckBytes` of a dotted-constant `EXPLAIN indexes = 1` against a no-dots control with a 150% threshold, measuring each arm with a single query. A single run is not a stable oracle: whichever query happens to be the first to touch a cache or spin up a thread pool absorbs a transient multi-megabyte allocation, and the dotted arm always runs first in the loop, so such a one-off lands on it and inflates the ratio arbitrarily. Locally the effect is easy to see: the first query after a server start reports 8–22 MB in this counter against a ~2.8 MB steady state for the identical query. The test failed this way on at least 6 unrelated PRs since 2026-08-12 (e.g. `Stateless tests (amd_asan_ubsan, distributed plan, parallel)` on https://github.com/ClickHouse/ClickHouse/pull/104948 at commit 81300c20: `longidx index analysis over a constant with 100000 dots allocated 4216780 bytes, more than 150% of the no-dots control (22224 bytes)` — reruns passed; CIDB shows the same failure on #106011, #108522, #114475, #96130, #114476). Measure each arm three times and take the minimum: a genuine quadratic regression is deterministic and shows up in every run, so the oracle keeps discriminating (the minimum can only remove one-sided transient noise), while a one-off transient can no longer fail the test. Verified against the current master binary: the modified test passes repeatedly. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1329` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114608",
        "createdAt": "2026-08-13T09:13:17Z",
        "updatedAt": "2026-08-13T14:33:03Z",
        "timestamp": "2026-08-13T14:33:03Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-synced-to-cloud",
          "pr-ci"
        ],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114609",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix flaky 04780_json_subcolumn_index_match_not_quadratic",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113289 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Requested in https://github.com/ClickHouse/ClickHouse/pull/112705#issuecomment-5276141544. Related: https://github.com/ClickHouse/ClickHouse/pull/113289, which added the test. Two independent test defects. The engine fix itself is fine. **1. The oracle read a counter that does not exist on sanitizer builds.** `ProfileEvents['MemoryAllocatedWithoutCheckBytes']` is written only from the allocation paths that skip the memory-limit check, whose only non-Darwin producer is `src/Common/malloc.cpp`, compiled under `#if USE_JEMALLOC`. jemalloc is disabled for every sanitizer except UBSan, and the throwing `operator new` goes through `CurrentMemoryTracker::allocThrow` instead, so on `asan_ubsan`, `tsan` and `msan` the counter has no writer: on an ASAN build a query peaking at 76 MB reports 256 bytes through it. The test's `-gt 0` guard then aborted the first arm, which is the `no query_log row for plain dotted arm` failure. The test now reads `system.query_log.memory_usage`. Peak memory is tracked by `operator new` accounting, which is not sanitizer-gated, so it is defined on every flavour. The untracked-memory limit stays pinned to 0 so allocations reach the query tracker immediately, not in deferred batches whose flush points depend on concurrent load. **2. The last `EXPLAIN` asserted a master-only default.** `Parts: 0 | Granules: 0` comes from `describeActions`, gated by `actions`, which only master force-enables (`set_default_pretty_explain_settings`, absent from 26.5 and 26.6), so the two open backports of #113289 fail deterministically. Fixed by requesting `actions = 0` and dropping that line. `Granules: 0/1000`, the pruning under test, comes from `describeIndexes` and is unaffected. I could not confirm the 150% bound is too tight under sanitizers. It is sound on a quantity that measures the matcher, so it stays. Validation: 80 of 80 runs pass on debug and ASAN with randomization on and off, and 10 of 10 on ASAN under six concurrent load generators. On a pre-#113289 binary the oracle reddens on all four arms at 2.372x to 2.483x, against 1.0001x with the fix present.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114609",
        "createdAt": "2026-08-13T09:13:35Z",
        "updatedAt": "2026-08-13T16:17:19Z",
        "timestamp": "2026-08-13T16:17:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "shankar-iyer"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114610",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Lazy-load part-level column statistics",
        "text": "When loading a part's statistics, the part currently loads statistics for every column via `getEstimates`, even when only a few columns are needed. On wide tables this performs a lot of pointless deserialization. Related: https://github.com/ClickHouse/ClickHouse/pull/104691 This change makes per-part statistics load lazily, only for the requested columns: - `IMergeTreeDataPart::getEstimates` now takes the set of requested columns, loads only the missing ones, and caches the result per column. A deterministic miss (no statistics declared, file absent, or corrupted) is recorded as a `nullopt` entry so the column is not re-probed on every query. - Part pruning asks for only the filter columns: `StatisticsPartPruner` exposes `getCandidateColumns` and `filterPartsByStatistics` passes them to `getEstimates`. - `system.parts_columns` discovers which columns have statistics via the new `getColumnsWithStatistics` and loads only those. A new `LoadedStatisticsColumns` profile event counts the per-column statistics that are actually deserialized. The performance test measures a 100-column wide table with 50 parts and a query filtering on 3 columns; part pruning loads 150 columns instead of 5000. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Part-level column statistics are now loaded lazily, only for the columns that are actually needed, reducing statistics deserialization on wide tables.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114610",
        "createdAt": "2026-08-13T09:17:26Z",
        "updatedAt": "2026-08-13T09:33:44Z",
        "timestamp": "2026-08-13T09:33:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "hold"
        ],
        "author": "zoomxi",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114613",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `jsonPathValues` tokenizer for JSON text indexes",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 Related: https://github.com/ClickHouse/ClickHouse/pull/110757 I first tried to fix the edge cases between `JSONAllValues` and the text index. I kept finding more cases where the index could return an incorrect result. I therefore needed to disable more useful index paths to keep queries correct. The main problem is that `JSONAllValues` stores plain text. It does not retain the path or type of values from `Dynamic` JSON columns. PR #113376 documents several examples. A tokenizer made specifically for `JSON` seemed like a better fit. The new `jsonPathValues` tokenizer stores the JSON path, type, and value in each token. This gives the index enough information for safe typed comparisons. ```text +-------------------+-------+---------------------+-------+--------+------------------+ | escaped JSON path | 00 00 | binary-encoded type | 00 00 | 1-byte | payload | | | | | | kind | | +-------------------+-------+---------------------+-------+--------+------------------+ Payload: complete value : | full value | truncated value : | value prefix | SipHash-2-4 (8 bytes, LE) | map entry : | escaped key | 00 00 | complete/truncated value | validation : | empty | Kinds: 1/2 = scalar, 3/4 = array element, 5/6 = map entry (complete/truncated), 7 = dynamic validation Escaping: 00 -> 00 01 Component terminator: 00 00 ``` The path-value format also enables direct reads. This part was inspired by [ClickStack's use of text-index direct reads for dynamic map attributes](https://clickhouse.com/blog/making-clickstack-5x-faster-clickhouse-observability). `jsonPathValues` brings this model directly to `JSON` columns. It does not require user-defined alias columns or query rewrites. Long values use a bounded prefix and a hash. ClickHouse validates candidate rows when a token is truncated or a runtime type is unsafe. This prevents false-negative results. The tokenizer supports declared and dynamic paths, arrays, scalar leaves in `Array(JSON)`, and declared `Map(String, String)` paths. It supports equality, `IN`, `has`, map lookups, prefix and suffix searches, `LIKE`, `ILIKE`, regular expressions, and multi-search functions. `startsWith` uses an ordered dictionary lookup, so it does not require a full dictionary scan. Substring searches still scan the dictionary. ## Benchmark We tested `jsonPathValues(1024)` with 99,999,984 distinct events from 100 JSONBench files. Both target tables had the same 192-part layout. Each indexed query returned the same result as the query without an index. | Metric | No index | `jsonPathValues(1024)` | |---|---:|---:| | Build time | 214 s | 863 s | | Total storage | 9.29 GiB | 23.48 GiB | | Exact `cid` lookup | 1,926 ms | 13 ms | | Exact long-value lookup | 637 ms | 8 ms | | `startsWith` | 29 ms | 16 ms | | Substring `LIKE` | 145 ms | 93 ms | The index used 14.19 GiB, or approximately 152 bytes per JSON object. It made the table 2.53 times larger and made index construction 4.03 times slower. The exact lookups were 80 to 148 times faster. `startsWith` was 1.81 times faster. Query times are medians from ten warm runs. This PR also adds a checked-in performance test with 500,000 generated JSON rows. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `jsonPathValues` tokenizer for bounded, type-safe text indexes on `JSON` paths, arrays, and maps.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114613",
        "createdAt": "2026-08-13T09:35:07Z",
        "updatedAt": "2026-08-13T14:42:44Z",
        "timestamp": "2026-08-13T14:42:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "rorylshanks",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114614",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix join NDV propagation for distributed aggregation",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Preserve NDV statistics across optimized inner joins so distributed planning can choose partial aggregation for low-cardinality GROUP BY queries instead of shuffling all joined rows. ### Description An optimized `JoinStepLogical` has two children, so `estimateReadRowsCount` returned before reading its saved result-column statistics. Move the optimized-join handling before the generic single-child guard so row estimates and column statistics propagate to downstream planning. The regression test verifies that a low-NDV GROUP BY after an inner join selects partial aggregation. It pins the join-order limit and distributed aggregation mode to isolate the planner behavior from injected settings and exchange execution details. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114614",
        "createdAt": "2026-08-13T09:48:39Z",
        "updatedAt": "2026-08-13T10:10:05Z",
        "timestamp": "2026-08-13T10:10:05Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-performance"
        ],
        "author": "XuJia0210",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114615",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "diff-review skill second edition",
        "text": "Reworks the `diff-review` skill in `.claude/skills/diff-review` — the local in-browser diff review that Claude Code serves before a commit or a PR. Heads-up: I tailored this to how *I* like to review, so please treat it as a suggestion rather than a standard. Check the branch out, run it on one of your own diffs, and keep it only if you find it helpful. What changed: - **Multi-round review.** Sending a round no longer ends the review: the server stays up, the agent works on the comments it was handed while you keep reading, and each comment turns green in the page as it gets addressed. - **Comments are durable state.** They are written to the `--out` file as you type, so they survive a server restart or a page reload, and after the agent's fixes move the code they are relocated by their anchor line instead of pointing at the wrong place. Resolutions, replies and dismissals are part of that state. - **Explicit review range.** `--staged`, `--head <sha>` and `--committed` alongside `--base`, so a review of recorded work never picks up local edits; the header always states which range is on screen. - **Navigation.** Directory tree with per-file status, `+a −d` counts and open-comment badges, a path filter, an all-files page, whole-file mode with expandable folded regions, a Comments pane listing open and already-addressed comments, draggable sidebar and split, keyboard shortcuts with a `?` overlay, and light / dark / system themes. - **Tests.** `ui.html` is split into ES modules under `ui/`, covered by `ui_test.mjs`, and `persist_test.mjs` exercises the whole persistence round-trip against real servers on a throwaway repository. - Vendored `@pierre/diffs` bumped from 1.2.12 to 1.3.5. Local agent tooling only — nothing in the server, the build or the tests. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114615",
        "createdAt": "2026-08-13T09:54:07Z",
        "updatedAt": "2026-08-13T17:17:21Z",
        "timestamp": "2026-08-13T17:17:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "vdimir",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114617",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112573 to 26.5: Fix async bounded read buffer readbigat race",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114617",
        "createdAt": "2026-08-13T10:33:24Z",
        "updatedAt": "2026-08-13T11:26:27Z",
        "timestamp": "2026-08-13T11:26:27Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "kssenii",
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114618",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112573 to 26.6: Fix async bounded read buffer readbigat race",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114618",
        "createdAt": "2026-08-13T10:33:52Z",
        "updatedAt": "2026-08-13T11:39:38Z",
        "timestamp": "2026-08-13T11:39:38Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "kssenii",
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114619",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112573 to 26.7: Fix async bounded read buffer readbigat race",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.27` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114619",
        "createdAt": "2026-08-13T10:34:20Z",
        "updatedAt": "2026-08-13T14:31:07Z",
        "timestamp": "2026-08-13T14:31:07Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll",
        "state": "closed",
        "assignees": [
          "kssenii",
          "arsenmuk"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114620",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix serialization of Map-valued settings in access entities",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114591 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a bug where an access entity carrying a `Map`-valued setting, such as a settings profile with `http_response_headers` or `additional_table_filters`, was stored in a form that ClickHouse could not read back, so the entity became permanently unloadable after a restart. Closes #114591. ### Description Closes #114591. `CREATE SETTINGS PROFILE p SETTINGS http_response_headers = '{...}'` succeeds, but the entity is stored as `http_response_headers = [('k', 'v')]`, which nothing can parse back. No error appears at `CREATE` time, so the entity is **permanently unloadable** after a restart: ``` stored: ATTACH SETTINGS PROFILE `p` SETTINGS http_response_headers = [('a', 'b')] CONST; after restart: Code: 62. Syntax error: failed at position 65 ([) ... Could not parse <path>/access/<uuid>.sql ``` The list rebuild drops it silently, reading it back throws, so `SELECT` from `system.settings_profile_elements` fails while it is present, as does `RESTORE` of a backup holding it. Root cause: the value is cast to the setting's native type, so a Map setting holds a `Map` Field, which `FieldVisitorToString` renders as an array of tuples; but `ParserSettingsProfileElement` reads values with a scalar-only `ParserLiteral` and cannot open a `[`. That spelling is rejected everywhere, `SET http_response_headers = [('a','b')]` included, so the write side is wrong. Fix: when a profile element's value, MIN or MAX is a `Map` and the setting is builtin, emit the setting's canonical text as a quoted string. Write side only, no grammar change. Custom settings are excluded because `castValueUtil` returns their value unchanged, so a string would come back a `String` rather than a `Map`. Covers `CREATE USER`/`ROLE`, `ALTER ... SETTINGS` and MIN/MAX, and transitively BACKUP/RESTORE and both storages. `SHOW CREATE` now prints a quoted string rather than `[('k', 'v')]`, intended since the new form is copy-pasteable. Entities already stored in the broken form are not repaired, as they were never parseable; recreate them. Downgrade is safe: a pre-fix binary reads the new form correctly, so no versioning is needed. New test `04902_access_entity_map_setting_round_trip`: 11 of its 15 arms fail on pristine master and pass here, covering empty, multi-key and hostile maps plus a two-process on-disk reload.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114620",
        "createdAt": "2026-08-13T11:13:29Z",
        "updatedAt": "2026-08-13T16:50:47Z",
        "timestamp": "2026-08-13T16:50:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "pufit"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114621",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #94748 to 25.8: Fix invalid result of joining two `-Cluster` table functions",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/94748 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31694431101/job/94428840299)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114621",
        "createdAt": "2026-08-13T11:31:25Z",
        "updatedAt": "2026-08-13T11:31:52Z",
        "timestamp": "2026-08-13T11:31:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll4",
        "state": "open",
        "assignees": [
          "thevar1able",
          "KochetovNicolai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114622",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not apply DROP fault injection to a refreshable view's cleanup DROP",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/114420 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a refreshable materialized view leaking its rotated-out target table when `ignore_drop_queries_probability` is enabled. The `DROP` a refresh issues to clean up the previous target is a step of the refresh, not a `DROP` the user asked for, so the fault injection no longer applies to it. ### Description Follow-up to #114420, requested by @ tiandiwonder in https://github.com/ClickHouse/ClickHouse/pull/114420#discussion_r3772879738. `ignore_drop_queries_probability` makes a `DROP TABLE` silently do nothing (or become a `TRUNCATE`) so the stress suite exercises \"the table you dropped is still there\". It must only affect `DROP`s the user issued. **Root cause.** After a refresh swaps in a new target, `StorageMaterializedView::dropTempTable` drops the rotated-out one via `InterpreterDropQuery(drop_query, refresh_context)`. `createRefreshContext` never marks that context internal, so the gate treated the cleanup `DROP` as a user `DROP`. Its existing refreshable-view exemption does not cover this: that test inspects the table *being dropped*, which here is the inner target, not the view. Both refresh exits are affected. When it fires, the old target survives as `.tmp.inner_id.<uuid>` still holding a full copy of the view's data, outside the view's metadata and surviving restart. **The change.** `InterpreterDropQuery` gains an `internal` member mirroring the one `InterpreterCreateQuery` already has, `dropTempTable` sets it, and the shared `refresh_context` is untouched. Marking that context instead would break the refresh outright in a `Replicated` database: the publishing `RENAME` runs on it, and `DatabaseReplicated` rejects a non-initial query unless the interpreter also passes `flags.internal`, which neither the Rename nor the Drop interpreter did. That is why #114420's one-liner was safe there and is not here. `QueryFlags{ .internal = internal }` at the replicated enqueue is behaviour-neutral for every other `DROP`, since all its members default to false. The interpreter's other policy branches (`ON CLUSTER` dispatch, access and dependency checks) deliberately keep applying. **Validation.** New test `04887_refreshable_mv_cleanup_drop_not_ignored`: three positive arms (success path, failure path, `Replicated` database) leak on master and are clean with the fix; a fourth asserts a genuine user `DROP` is still skipped. 50 randomized runs plus 50 with `ignore_drop_queries_probability=0.2` were green, as were `04247`, `04796`, `04218` and `04327`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114622",
        "createdAt": "2026-08-13T11:43:28Z",
        "updatedAt": "2026-08-13T17:26:44Z",
        "timestamp": "2026-08-13T17:26:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114623",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "CI: Cache: restrict cross-branch reuse to pull_request workflows only",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112358 CI cache reuse in praktika keyed the reuse decision off the branch that produced a record. Branch, though, was only ever a proxy for how much a record can be trusted. This gates reuse on the producing **workflow event** instead and ignores the branch entirely in the reuse decision. Correctness is unaffected either way — a digest match already means identical inputs; the producing event only encodes how much we trust the result. Each cache record now carries the event that produced it. The policy: | Producing event → <br> Reusing event ↓ | `pull_request` | `push` / `schedule` / `dispatch` / `merge_queue` | |---|:---:|:---:| | `pull_request` | ✅ | ✅ | | `push` / `schedule` / `dispatch` / `merge_queue` | ❌ | ✅ | In words: a `pull_request` run reuses any record; every other (trusted) event reuses any record **except** one produced by a `pull_request`. The write side is the dual — a `pull_request` run only fills an empty `(job, digest)` slot (`if_not_exist=True`), while every trusted event overwrites it, so the shared slot always keeps a record reusable by every lane. This keeps the two properties the branch rule aimed at, more directly: - **Security.** A `pull_request` run executes untrusted (possibly fork) code, so a trusted lane must never reuse a record it produced. The trust boundary is stated as pull_request-vs-not rather than inferred from a branch name. - **Drift guard (#112358).** `merge_queue` never reuses a `pull_request` record, so a PR's green flaky-check result can't satisfy the queue's lookup and skip the re-run against the merge-group state — and this now holds independently of whether the PR and merge-queue digests happen to coincide. It also removes a clobbering hole: reuse no longer checks the branch, so a record produced by a release-branch push (or any trusted event) is reused whenever the digest matches, instead of being rejected on a branch mismatch. `CACHE_VERSION` is bumped to 2 so pre-existing records, which lack the event field, are not reused under the new rule. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114623",
        "createdAt": "2026-08-13T11:50:30Z",
        "updatedAt": "2026-08-13T16:55:54Z",
        "timestamp": "2026-08-13T16:55:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-ci"
        ],
        "author": "maxknv",
        "state": "open",
        "assignees": [
          "leshikus"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114624",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Find a LowCardinality needle equal to the type's default value",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `has`, `indexOf`, `countEqual`, `mapContainsKey`, `mapContainsValue` and `Map` subscript returning \"not found\" for a constant `LowCardinality` needle equal to the element type's default value, such as an empty `String` or a zero number. ### Description **The problem.** A constant needle equal to the element type's default value is never found in an `Array(LowCardinality(T))` or a `Map` with a `LowCardinality` key. No exception, so a `WHERE` on such a predicate silently drops rows: ```sql CREATE TABLE t (a Array(LowCardinality(String))) ENGINE = Memory; INSERT INTO t VALUES (['', 'a']); SELECT has(a, '') FROM t; -- 0, expected 1 ``` Likewise for `0` over the numeric types, a NUL-padded `FixedString` and the `1970-01-01` `Date`, while `m['']` returns `''` instead of the stored value. Any non-default needle is correct. Reproduces on 26.5 to 26.7. **Root cause.** A `LowCardinality` dictionary reserves prefix slots for the default and NULL values, and `ReverseIndex` is built with `num_prefix_rows_to_skip`, so the reserved slot is never indexed. `ColumnUnique::uniqueInsertData` compensates for that on the write path; the read path had no counterpart. **The change.** `ColumnUnique::getOrFindValueIndex` now performs the same default-slot match as the write path. `ReverseIndex` is untouched, so no write behaviour moves. That slot becoming reachable brings two equality details. The lookup casts the constant into the element type without reporting loss, so `UInt64(256)` arrived as `UInt8(0)` and would match the default; that slot now answers only if the element type can represent the constant. And a dictionary can hold `-0.0` and `0.0` as separate entries while the shortcut has room for one index, so a zero constant over a float element type is left to the value-comparing path, as `indexOfAssumeSorted` already is. The new test covers both call sites. Lookup timing is unchanged. **Overlap with #112953.** That open PR declines this same shortcut for every float element type, subsuming the float-zero decline here, and the two conflict textually; whichever merges second should keep the broader decline and drop this one, collapsing the duplicated representability predicate to one copy. Non-float types are independent of it.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114624",
        "createdAt": "2026-08-13T11:55:50Z",
        "updatedAt": "2026-08-13T17:39:52Z",
        "timestamp": "2026-08-13T17:39:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114625",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix the bit-sliced full adder in `groupNumericIndexedVector`",
        "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/106208 Related: https://github.com/ClickHouse/ClickHouse/pull/110072 `addValue` set a bit when the computed sum bit was 1 but never cleared it when the sum bit was 0, so every carry left the lower bit set and the write path was not addition: ```sql SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), val)) FROM (SELECT arrayJoin([toInt64(10), toInt64(10)]) AS val); -- {5:30}, expected {5:20} SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), toInt64(1))) FROM numbers(8); -- {5:255}, expected {5:8} ``` Any repeated index whose addends share a set bit is affected, on every index type and in both the small and the promoted representation. `n` rows of value `1` at one index accumulate to `2^n - 1`. Additions whose bits are disjoint need no carry and were already correct (`10 + 5` gives `15`), which is why the existing tests and the documented examples — all of which use distinct indexes — did not catch it. `merge` and `numericIndexedVectorPointwiseAdd` share `pointwiseAddInplace`, which computes whole-bitmap XORs and assigns the result, so clearing is implicit there and those paths were already correct. Only the per-row path was wrong, which is why `numericIndexedVectorAllValueSum` disagreed with `sum(value)` over the same rows. `RoaringBitmapWithSmallSet` had no way to clear an element, so this adds a `remove`. `SmallSet` has no erase and the small set holds at most `small_set_size` elements, so that path rebuilds it without the removed value. `zero_indexes` is now maintained too. It holds the present indexes whose value is zero, so it has to gain an index when an update drives the value to zero and lose it when the value becomes non-zero — the same invariant the pointwise operations restore when they finish: ```cpp /// For any of the total_indexes, if it is not in the non-zero index of the result, the result is 0. total_indexes->rb_andnot(*getAllNonZeroIndex()); zero_indexes = total_indexes; ``` With that, adding `5` and then `-5` row by row produces `{5:0}`, matching what merging the two values already produced, and `numericIndexedVectorGetValue` and `numericIndexedVectorCardinality` agree with the map. `numericIndexedVectorBuild` also goes through `addValue`, but from a map whose keys are unique, so it never carried and is unaffected. The test asserts `numericIndexedVectorAllValueSum` equals `sum(value)` over repeated indexes with negative and fractional values, that the row-by-row path agrees with the pointwise path, and that an index driven to zero is present with value zero. Eight of its eleven assertions fail without the fix; the three that pass are controls — an addition with disjoint bits, the pointwise reference path, and adding zero to an index that already holds a value. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `groupNumericIndexedVector` producing wrong values when the same index appears in more than one row. The bit-sliced adder never cleared a bit when the computed sum bit was zero, so every carry left the lower bit set: eight rows of value `1` at one index accumulated to `255` instead of `8`, and `10 + 10` produced `30` instead of `20`. Values that share no set bits were unaffected. An index whose value is driven to zero is now reported as present with value zero, consistent with `numericIndexedVectorPointwiseAdd` and with merging aggregate states. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114625",
        "createdAt": "2026-08-13T12:12:29Z",
        "updatedAt": "2026-08-13T16:57:26Z",
        "timestamp": "2026-08-13T16:57:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yakov-olkhovskiy",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114626",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not merge-sort a distributed gather whose sort description is all-constant",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 Related: https://github.com/ClickHouse/ClickHouse/issues/106237 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description With `make_distributed_plan = 1`, a window `PARTITION BY <constant>` builds a `GatherSend` fragment whose sort description is entirely constant. `GatherSendStep::updatePipeline` adds an order-preserving `MergingSortedTransform` for any non-empty description, and that merge waits for *every* input stream to have data (`IMergingTransformBase::prepareInitializeInputs`). Hashing a constant key sends all rows to one bucket, so the other branches are never fed and stay `NeedData` while the loaded branch is `PortFull`. Neither side can move: the query deadlocks at zero CPU until `receive_timeout` and then raises `Pipeline stuck` (a logical error, so the server aborts in debug and sanitizer builds). An all-constant description orders nothing, so this change takes the `pipeline.resize(1)` branch that already exists in that function. A `ResizeProcessor` pairs any waiting output with any ready input and has no all-inputs barrier, which is why the code before [4a1ab1e](https://github.com/ClickHouse/ClickHouse/commit/4a1ab1e3d94fb0d) - which replaced an unconditional `resize(1)` with this conditional merge - could not wedge this way. Every row compares equal under an all-constant description, so any interleaving is validly sorted and the contract `GatherReceiveStep` relies on still holds. A description with at least one real column keeps the merge. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 (where this failure was reported) Related: https://github.com/ClickHouse/ClickHouse/issues/106237 (context only, different mechanism) Found by `AST fuzzer (amd_debug, targeted, old_compatibility)`: [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113558&sha=b9fe1650232abf4fefc16992f3efcce44222bebc&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The `BufferedShardByHashTransform` deadlock tracked in #106237 is a different mechanism: that transform does not appear in the failing pipeline at all. The new test `04888` fails on master with `Pipeline stuck` for a constant key, a `LowCardinality` constant and a `Nullable` constant, and passes with this change. Its controls - a real column key, a mixed constant-plus-column key, and the non-distributed plan - pass both before and after, so the change is narrow. `04837_distributed_plan_window_partition_shuffle`, which pins the full distributed plan for a column-keyed window, still passes.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114626",
        "createdAt": "2026-08-13T12:23:35Z",
        "updatedAt": "2026-08-13T16:32:28Z",
        "timestamp": "2026-08-13T16:32:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "davenger"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114627",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #112217 to 26.7: Measure cancellation server-side in test_cancel_backup.py",
        "text": "Backport of https://github.com/ClickHouse/ClickHouse/pull/112217 to `26.7`. Related: https://github.com/ClickHouse/ClickHouse/pull/112217 Related: https://github.com/ClickHouse/ClickHouse/pull/114478 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Description `test_backup_restore_on_cluster/test_cancel_backup.py::test_cancel_backup` is flaky on the `26.7` branch with the same signature that #112217 fixed on `master`: the client-side `time_to_cancel` stopwatch measures the integration harness (several `docker exec` clickhouse-client launches plus `wait_status` poll quantum) rather than the cancellation itself, and the resulting body exception is masked in reports as `NoTrashChecker.__exit__` asserting `'QUERY_WAS_CANCELLED' in []`. It just failed this way on the backport PR https://github.com/ClickHouse/ClickHouse/pull/114478 (report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114478&sha=f5364950f96d60206363d2b2c1bff278a6976ee8&name_0=BackportPR&name_1=Integration%20tests%20%28amd_asan_ubsan%2C%20db%20disk%2C%20old%20analyzer%2C%201%2F6%29), and #112217 itself documents an occurrence on the `26.6` release branch, so release branches keep hitting it. This is a clean cherry-pick of the test-only fix (measure cancellation server-side via `system.backups` timestamps).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114627",
        "createdAt": "2026-08-13T12:29:22Z",
        "updatedAt": "2026-08-13T14:10:31Z",
        "timestamp": "2026-08-13T14:10:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [],
        "author": "alexey-milovidov",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114628",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Wait for DETACH DATABASE to release tables in three integration tests",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/93064 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/issues/93064 Three integration tests intermittently fail with `Code: 219 ... Database <db> cannot be detached, because some tables are still in use. Retry later.` thrown from `DatabaseAtomic::assertCanBeDetached` (`src/Databases/DatabaseAtomic.cpp:521`). Root cause: non-SYNC `DETACH DATABASE` is best-effort by contract, and these tests assume it is atomic. `assertCanBeDetached` throws if any table of the database still has a live refcount. The engine already has the deterministic wait, `waitDetachedTableNotInUse`, but `executeToDatabaseImpl` calls it only under `query.sync` (`src/Interpreters/InterpreterDropQuery.cpp:779-790`). Stateless CI installs `tests/config/users.d/database_atomic_drop_detach_sync.xml` (`database_atomic_wait_for_drop_and_detach_synchronously=1`); the integration helpers install no such profile and run at the server default of 0. Hence the failures are confined to the integration suite, and 16 of the 17 integration occurrences in 180 days are sanitizer builds, where the holder's window is wider. Code 219 is intended behaviour of the asynchronous form and is pinned in-tree: `01107_atomic_db_detach_attach.sh:19` sets the setting to 0 to provoke it and asserts it. So this is a test-side defect, and the change asks for the wait the engine already implements rather than altering it. I added `SYNC` to the five `DETACH DATABASE` statements with observed CI failures: `test_drop_replica` (11 hits/180d), `test_replicated_table_structure_alter` (4), and `test_drop_database_replica:192` (2, both stacks confirm that line is the thrower). All 21 non-SYNC sites under `tests/integration/` were enumerated; the other 16 are excluded because they have zero CIDB hits in 180 days, already pass `database_atomic_wait_for_drop_and_detach_synchronously`, or sit inside the existing `detach_database_with_retry` helper, whose docstring documents a different holder. Validation: with a holder injected the way `01107` does it, the DETACH returns 219 without `SYNC` and succeeds with it, 10/10 on `Atomic` and on `Replicated`. The three tests pass 42/42 runs. Each `DETACH` in `test_drop_replica` measures ~0.115 s over 25 measurements, so the wait costs nothing when no table is held, and it is cancellable and shutdown-aware. CI report for the master failure: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=91b700711adf0a07172e1046c8d8f013f1e0b82c&name_0=MasterCI&name_1=Integration%20tests%20%28amd_tsan%2C%206%2F6%29 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1356` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114628",
        "createdAt": "2026-08-13T12:36:13Z",
        "updatedAt": "2026-08-13T17:55:30Z",
        "timestamp": "2026-08-13T17:55:30Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "can be tested",
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "closed",
        "assignees": [
          "PedroTadim"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114629",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix DPsub join reordering silently dropping single-table ON-clause filters",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111898 The `dpsub` join-order algorithm silently dropped single-table filter and constant predicates that live in a `JOIN ... ON` clause (e.g. `t1.value = 'x'`), returning extra rows. `greedy` uses a different placement path and was unaffected. The predicates are placed in `collectJoinEdgesMask`. Two placement conditions there each silently dropped such a predicate: - `two_relations` required the whole join step to be exactly two relations, so the predicate was dropped whenever its relation was introduced against an already-multi-relation subplan (e.g. `t1` at the top of `t1 JOIN (t2 JOIN t3)`). - `fromLeft() || fromRight() || fromNone()` dropped the predicate for any single-table filter on a relation whose id is >= 2, because `fromLeft`/`fromRight` test relation ids 0 and 1 specifically (they describe the two inputs of a binary join step, not \"references a single relation\"). The predicate is now attached at the join that introduces its relation (the split whose one side is exactly that relation), and pure constants at the earliest two-relation join. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fixed `dpsub` join-order optimization (`query_plan_optimize_join_order_algorithm = 'dpsub'`) silently dropping single-table filter conditions from a `JOIN ... ON` clause, which could return extra rows.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114629",
        "createdAt": "2026-08-13T12:38:37Z",
        "updatedAt": "2026-08-13T16:57:48Z",
        "timestamp": "2026-08-13T16:57:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "fkastrati",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114631",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113742 to 25.8: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114631",
        "createdAt": "2026-08-13T12:48:05Z",
        "updatedAt": "2026-08-13T12:48:13Z",
        "timestamp": "2026-08-13T12:48:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll4",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114632",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113742 to 26.3: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114632",
        "createdAt": "2026-08-13T12:48:47Z",
        "updatedAt": "2026-08-13T12:48:55Z",
        "timestamp": "2026-08-13T12:48:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll4",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114633",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #113742 to 26.5: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114633",
        "createdAt": "2026-08-13T12:49:24Z",
        "updatedAt": "2026-08-13T12:49:32Z",
        "timestamp": "2026-08-13T12:49:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll4",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114634",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113742 to 26.6: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114634",
        "createdAt": "2026-08-13T12:49:55Z",
        "updatedAt": "2026-08-13T16:43:00Z",
        "timestamp": "2026-08-13T16:43:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll4",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114635",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113742 to 26.7: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.29` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114635",
        "createdAt": "2026-08-13T12:50:24Z",
        "updatedAt": "2026-08-13T16:21:35Z",
        "timestamp": "2026-08-13T16:21:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-ch-test-poll4",
        "state": "closed",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114636",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix logical error in text index lazy apply mode on cancelled queries",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114603 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed logical error `Multi-block postings must be compressed` in queries over tables with a text index in the lazy posting-list apply mode, when the query was canceled during the read.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114636",
        "createdAt": "2026-08-13T12:51:52Z",
        "updatedAt": "2026-08-13T15:39:26Z",
        "timestamp": "2026-08-13T15:39:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "CurtizJ",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114642",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: document filtering for `system.query_log` initial queries",
        "text": "Document the recommended filters for analyzing `system.query_log`. The page now explains that `is_initial_query = 1` selects top-level client queries, while `initial_query_id` correlates the full cascade of a distributed query across nodes. The existing basic example also filters for initial queries so child executions are not treated as separate client queries. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114642",
        "createdAt": "2026-08-13T13:10:53Z",
        "updatedAt": "2026-08-13T17:37:19Z",
        "timestamp": "2026-08-13T17:37:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-documentation",
          "pr-autogenerated-docs"
        ],
        "author": "Blargian",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114643",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Re-land aggregate function `gini` in the `sum` family",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/113763 Related: https://github.com/ClickHouse/ClickHouse/pull/113868 Related: https://github.com/ClickHouse/ClickHouse/pull/112280 --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New aggregate function `gini`, which calculates the [Gini coefficient](https://en.wikipedia.org/wiki/Gini_coefficient) of a column of finite, non-negative numeric values. The result ranges from `0` (all values equal) towards `1` as inequality grows; for a sample of `n` values the maximum is `(n - 1) / n`. `NaN` values are skipped and infinite values are rejected. The function returns `Float64` and consumes `O(n)` memory. ### Description Re-lands the `gini` function that #113868 reverted, following the option-2 spec @ Manerone gave in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5217284116 and confirmed in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5279278069. It re-adds #112280's function with exactly three items removed, and touches no file under `src/AggregateFunctions/Combinators/`: - the `getArgumentsThatCanBeOnlyNull` override, - `.returns_default_when_only_null = true` on registration, - the `argument_type->onlyNull()` branch in the creator. The property was doing the work. It made `AggregateFunctionFactory::getImpl` skip its only-null guard and build a real `gini` instance over `Nullable(Nothing)`, so `gini(NULL)` returned `Float64` `nan` where every `sum`-family function folds to `Nullable(Nothing)`. Without it the fold happens in the `Null` combinator before any other combinator is applied, and the creator's only-null branch becomes unreachable, which is also why `createAggregateFunctionSum` has no equivalent. `gini` now registers exactly like `sum`. Runtime `Nullable` handling is a separate axis, via `getOwnNullAdapter`, and is unchanged: `gini` over `[1, NULL, 3]` still returns `0.25`. Validated on three binaries: post-revert master (`gini` absent), a build carrying #112280's version, and this branch. Every literal-`NULL` cell on this branch equals the `sum` value measured on the same binary, and all non-`NULL` results are unchanged from #112280. No documentation files are touched. The page body is generated from the `FunctionDocumentation` block in `AggregateFunctionGini.cpp`, which this PR restores, so the docs autogeneration workflow fills the page from that source. cc @Manerone",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114643",
        "createdAt": "2026-08-13T13:34:59Z",
        "updatedAt": "2026-08-13T17:57:20Z",
        "timestamp": "2026-08-13T17:57:20Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [
          "Manerone"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114644",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"Revert \"NATS: add inline credentials setting\"\"",
        "text": "Reverts https://github.com/ClickHouse/ClickHouse/pull/114178, restoring the `nats_credentials` setting of the `NATS` table engine (originally added in https://github.com/ClickHouse/ClickHouse/pull/110733), and fixes the reason of the original revert: the possibility of referencing arbitrary server paths through `nats_credential_file`. `nats_credential_file` is a path on the server filesystem: the server opens it with its own privileges, and during authentication the credentials are sent to `nats_url`, which comes from the same query. So a path taken from SQL lets anyone who can define a `NATS` source probe the local filesystem (the error text distinguishes a missing file, a permission error, and a file without a seed), and exfiltrate files the server can read to a NATS server they control (the part of the file before the seed line is sent verbatim in the `CONNECT` frame's `jwt` field). See the analysis in https://github.com/ClickHouse/ClickHouse/pull/110733#discussion_r3750203599. This hole predates #110733 (`nats_credential_file` exists since 24.2), but with the inline `nats_credentials` setting restored, the path form is no longer needed in SQL at all. The restriction, modelled on `StorageMySQL::getSSLParams` (`e700bbec4c84c585`): `nats_credential_file` is accepted only from a named collection defined in the server configuration file, or as `nats.credential_file` in the server configuration itself. Every SQL spelling — the `SETTINGS` clause, an engine-argument override `NATS(collection, nats_credential_file = ...)`, or a named collection created by `CREATE NAMED COLLECTION` — throws `BAD_ARGUMENTS` with a message directing to `nats_credentials`. Replacing a configured path with inline `nats_credentials` from the query remains allowed (the path itself is not used then). Loading from previously-validated metadata (server startup, force-restore, and short-syntax `ATTACH`) is exempt via `isLoadingFromExistingMetadata`, so existing tables keep working after an upgrade; a user-issued full `ATTACH TABLE` query is still checked. `ALTER TABLE ... MODIFY SETTING` is not a bypass: the `NATS` engine does not support settings alter. Tests: a new stateless test `04891_nats_credential_file_path_restriction` covers all rejected SQL spellings and the accepted configuration-file sources, using a `NATS` named collection added to the stateless-test server configuration (`tests/config/config.d/named_collection.xml`, with `nats_url = '127.0.0.1:1'`, so passing the validation surfaces as `CANNOT_CONNECT_NATS`); `04665_nats_credentials_named_collection` is updated — the `nats_credential_file` query-override direction is now rejected. Related: https://github.com/ClickHouse/ClickHouse/pull/110733 Related: https://github.com/ClickHouse/ClickHouse/pull/114178 Related: https://github.com/ClickHouse/ClickHouse/issues/85213 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `NATS` table engine accepts credentials inline in the new `nats_credentials` setting (the same payload as a `.creds` file), and no longer accepts `nats_credential_file` from SQL: the path is a reference to a file on the server filesystem, which the server opens with its own privileges, so it can only be specified in a named collection defined in the server configuration file, or as `nats.credential_file` in the server configuration itself. Tables created before this restriction keep working after an upgrade.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114644",
        "createdAt": "2026-08-13T14:03:24Z",
        "updatedAt": "2026-08-13T17:38:08Z",
        "timestamp": "2026-08-13T17:38:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-backward-incompatible"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114645",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Speed up `IN (subquery)` set building by pre-deduplicating each `MergeTree` partition independently",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/108326 Related: https://github.com/ClickHouse/ClickHouse/pull/105126 The set for `IN (subquery)` is built by a single `CreatingSetsTransform`: all streams of the subquery are merged into one and every row is hashed serially, no matter how many threads read the data. If the partition expression of the subquery's table is a function of the subquery's output columns (the set is keyed on all of them), the reading will now emit each partition through a single port and each stream is deduplicated independently before the filling transform. Because a key then lives in exactly one stream, per-stream deduplication is complete, and the single filling transform only hashes unique rows — the serial part of the build shrinks from all rows to distinct rows, and the deduplication itself runs in parallel. ```sql CREATE TABLE t (a UInt64) ENGINE = MergeTree ORDER BY tuple() PARTITION BY a % 8; INSERT INTO t SELECT number % 1000000 FROM numbers(100000000); OPTIMIZE TABLE t FINAL; EXPLAIN PIPELINE SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t) SETTINGS allow_creating_set_partitions_independently = 1, max_threads = 8; ``` ```response (CreatingSets) DelayedPorts 9 → 8 (Expression) ExpressionTransform × 8 (Aggregating) Resize 1 → 8 AggregatingTransform (Expression) ExpressionTransform (Filter) FilterTransform (ReadFromSystemNumbers) NumbersRange 0 → 1 (CreatingSet) CreatingSetsTransform <- the single filling transform now hashes ~1M unique rows instead of 100M Resize 8 → 1 DistinctTransform × 8 <- new: parallel pre-deduplication on partition-disjoint streams (Expression) ExpressionTransform × 8 (ReadFromMergeTree) MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) × 8 0 → 1 <- per-partition reading (8 partitions → 8 streams) ``` On the table above (100M rows, 1M distinct keys, 8 balanced partitions; 64-core machine, average of 3 runs after 2 warm-ups): | query | off | on | speedup | |------------------------------------------------------------|--------|--------|---------| | `SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t)` | 0.643s | 0.113s | 5.7× | ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Speed up set building for `IN (subquery)` on partitioned `MergeTree` tables by keeping each partition's rows within a single stream and deduplicating each stream independently, so the single set-filling transform — previously hashing every row serially — only sees unique rows. This applies when the partition expression is a deterministic function of the subquery's output columns. The optimization is not applied when the largest partition holds more than twice the rows of the average partition; the new setting `force_creating_set_partitions_independently` (disabled by default) bypasses this check. Controlled by the new setting `allow_creating_set_partitions_independently` (enabled by default).",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114645",
        "createdAt": "2026-08-13T14:04:03Z",
        "updatedAt": "2026-08-13T17:50:48Z",
        "timestamp": "2026-08-13T17:50:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-performance"
        ],
        "author": "nihalzp",
        "state": "open",
        "assignees": [
          "yariks5s"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114646",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fail-closed allowlist of hypothetical index types for WHATIF",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114646",
        "createdAt": "2026-08-13T14:06:56Z",
        "updatedAt": "2026-08-13T17:07:49Z",
        "timestamp": "2026-08-13T17:07:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "yariks5s",
        "state": "open",
        "assignees": [
          "nihalzp"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114647",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add LIKE/NOT LIKE/ILIKE filtering to SHOW access entity statements",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111692 Implement `[NOT] [I]LIKE 'pattern'` for all `SHOW` statements that list access entities: - `SHOW USERS` - `SHOW ROLES` / `SHOW CURRENT ROLES` / `SHOW ENABLED ROLES` - `SHOW SETTINGS PROFILES` - `SHOW QUOTAS` - `SHOW ROW POLICIES` - `SHOW MASKING POLICIES` This allows filtering access entities by name pattern, matching the existing behavior of `SHOW TABLES LIKE` and `SHOW DATABASES LIKE`: ```sql SHOW USERS LIKE '%admin%' SHOW ROLES NOT LIKE '%-internal' SHOW USERS ILIKE '%ALICE%' ``` The implementation mirrors the existing `SHOW TABLES` approach: three fields on the AST (`like`, `not_like`, `case_insensitive_like`), parsing in the shared access entity parser, and query rewriting in the interpreter to append a `WHERE name [NOT] [I]LIKE '...'` clause to the `system.*` table query. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `[NOT] [I]LIKE 'pattern'` filtering to `SHOW USERS`, `SHOW ROLES`, `SHOW SETTINGS PROFILES`, `SHOW QUOTAS`, `SHOW ROW POLICIES`, and `SHOW MASKING POLICIES` statements. Made with [Cursor](https://cursor.com)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114647",
        "createdAt": "2026-08-13T14:20:37Z",
        "updatedAt": "2026-08-13T14:20:37Z",
        "timestamp": "2026-08-13T14:20:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "anandheritage",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114648",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: move ODBC driver documentation to clickhouse-odbc",
        "text": "## Summary - move the ODBC driver guide into a group-backed folder under `concepts/features/interfaces` - remove the duplicate ODBC connector page and its navigation entry - redirect the retired connector URL and its legacy alias directly to the ODBC table-engine reference - update remaining English documentation links to bypass the redirect ## Why The retired connector page duplicated the ODBC table-engine reference. The actual driver guide belongs with `ClickHouse/clickhouse-odbc`, where it can evolve alongside the driver and later be split into focused pages. ## Coordination Paired draft PR: https://github.com/ClickHouse/clickhouse-odbc/pull/581 The paired PR vendors the driver guide and adds verification and one-way documentation sync workflows. This PR prepares the corresponding target folder and navigation reference in `ClickHouse/ClickHouse`. ## Validation - scoped Mintlify validation passed for `concepts/features/interfaces/odbc` - snippet and component import checks passed - internal link validation passed with zero errors - redirect validation passed with zero errors and no redirect chains ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Move the ODBC driver guide to the driver repository and retire the duplicate connector page.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114648",
        "createdAt": "2026-08-13T14:23:22Z",
        "updatedAt": "2026-08-13T14:56:55Z",
        "timestamp": "2026-08-13T14:56:55Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-documentation"
        ],
        "author": "Blargian",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114650",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keep the projection's intermediates out of the WITH FILL header",
        "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114404 Caused by: https://github.com/ClickHouse/ClickHouse/pull/107700 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NUMBER_OF_COLUMNS_DOESNT_MATCH` for a distributed query combining an `ALIAS` column whose body is an expression with `ORDER BY ... WITH FILL ... INTERPOLATE`. ### Description `analyzeSort` passed every column *available* to the `Before INTERPOLATE` step through as an output. A step's available columns are all the nodes of the previous step's `ActionsDAG` — the intermediates of a computed expression included, since a later step is allowed to *reference* any of them. Turning \"referenceable\" into \"must be in the stream\" pinned those intermediates into the header of the `Filling` step. That made the header depend on something that should not matter. Two query trees that differ only by whether an `ALIAS` column was inlined into its defining expression produced different headers, because `v * 2` contributes its `2` and the un-inlined `a_v` does not: ``` un-inlined inlined __table1.k __table1.k __table1.a_v multiply(__table1.v, 2_UInt8) __table1.v __table1.v 2_UInt8 <- extra materialize(__table1.k) materialize(__table1.k) a_v a_v ``` A distributed query is planned from the un-inlined tree on the initiator and executed from the inlined one on the shard, so the two headers meet. `buildShardCollapseFanOut` only handles a *smaller* shard header, and the positional `makeConvertingActions` then throws `NUMBER_OF_COLUMNS_DOESNT_MATCH`. The fix holds back the projection's own intermediates and passes everything else through as before, including the sort columns materialized just above, which is what `Filling` fills by. They all remain *inputs* either way, so the step still asks the previous one to produce them. One exception has to be carved out: a column that an `INTERPOLATE` expression names. Those expressions become actions later, in `Planner`, against a dag built from this step's *header*, so whatever they reference must survive as a column even when the query does not select it — `INTERPOLATE (inter AS inter2 + inter)` where `inter2` is not in the `SELECT` list. Those names are collected by visiting the expressions into a throwaway dag. ### Scope Everything here is inside `if (query_node.hasInterpolate())`, so only queries with `INTERPOLATE` change. The shape was already broken over a `Distributed` table before #107700, since that path has always inlined `ALIAS` columns; #107700 extended the inlining to the parallel-replicas paths and so exposed it there too. Both are fixed. It also stops the reconciliation failure from masking a query error: `INTERPOLATE (k AS k)` on an `ORDER BY` column reports `INVALID_WITH_FILL_EXPRESSION` again instead of `NUMBER_OF_COLUMNS_DOESNT_MATCH`. ### Validation Built and run against a three-replica localhost cluster and a two-shard `Distributed` table, compared with the CI binaries of the commit before #107700 (`75b17ad`) and of its merge (`dd01d270e`): | | before #107700 | after #107700 | this | |---|---|---|---| | repro, parallel replicas shipping a plan | pass | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | repro, parallel replicas shipping SQL | pass | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | repro over a 2-shard `Distributed` table | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | `INTERPOLATE (k AS k)` under parallel replicas | `INVALID_WITH_FILL_EXPRESSION` | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | `INVALID_WITH_FILL_EXPRESSION` | The new test `04891_with_fill_interpolate_alias_column_header` reports four exceptions on the commit before #107700 and seven on the merge commit, and passes here. Every self-contained stateless test mentioning `WITH FILL` or `INTERPOLATE`, plus the `ALIAS`-shipping tests from #107700, passes: 94 of 94. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-dd669bf9-3d72-4334-961f-0871feaf9f98?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-dd669bf9-3d72-4334-961f-0871feaf9f98&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114650",
        "createdAt": "2026-08-13T14:39:41Z",
        "updatedAt": "2026-08-13T17:17:23Z",
        "timestamp": "2026-08-13T17:17:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "yakov-olkhovskiy",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114651",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #114188 to 25.8: Ignore redundant parentheses in stored table definitions",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114651",
        "createdAt": "2026-08-13T15:07:10Z",
        "updatedAt": "2026-08-13T15:07:19Z",
        "timestamp": "2026-08-13T15:07:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114652",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #114188 to 26.3: Ignore redundant parentheses in stored table definitions",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114652",
        "createdAt": "2026-08-13T15:07:56Z",
        "updatedAt": "2026-08-13T15:08:04Z",
        "timestamp": "2026-08-13T15:08:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114653",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #114188 to 26.5: Ignore redundant parentheses in stored table definitions",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114653",
        "createdAt": "2026-08-13T15:08:36Z",
        "updatedAt": "2026-08-13T15:08:44Z",
        "timestamp": "2026-08-13T15:08:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114654",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #114188 to 26.6: Ignore redundant parentheses in stored table definitions",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114654",
        "createdAt": "2026-08-13T15:09:14Z",
        "updatedAt": "2026-08-13T15:09:22Z",
        "timestamp": "2026-08-13T15:09:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114655",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cherry pick #114188 to 26.7: Ignore redundant parentheses in stored table definitions",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114655",
        "createdAt": "2026-08-13T15:09:46Z",
        "updatedAt": "2026-08-13T15:09:54Z",
        "timestamp": "2026-08-13T15:09:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "do not test",
          "pr-bugfix",
          "pr-cherrypick"
        ],
        "author": "robot-ch-test-poll",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114656",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "docs: convert merge loop code block to Steps component",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replace the numbered code block describing the background merge loop on the AWS performance page with a `<Steps>` component for clearer rendering. <!-- mintlify-agent-attribution --> --- Generated by Mintlify Agent. Requested by: shaun.struwig@clickhouse.com via Slack Mintlify session: slack_1785790244.316179_D0B0U0V3Z88",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114656",
        "createdAt": "2026-08-13T15:16:40Z",
        "updatedAt": "2026-08-13T17:41:13Z",
        "timestamp": "2026-08-13T17:41:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-documentation"
        ],
        "author": "mintlify[bot]",
        "state": "closed",
        "assignees": [
          "Blargian"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114658",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix inconsistent AST formatting for a subquery argument of the view table function",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114004 (auto-closes the issue when this PR is merged into the default branch) --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114004 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed formatting of a subquery argument of the `view` and `viewIfPermitted` table functions when the enclosing query has a trailing `SETTINGS` clause. Such a query was formatted as `view((SELECT ...))`, which cannot be parsed back, so the query failed the internal format-parse-format check and raised `Inconsistent AST formatting`. ### Description `ASTQueryWithOutput::formatImpl` sets `parent_has_trailing_settings` so an inner `ASTSelectWithUnionQuery` parenthesizes its individual SELECTs, keeping the re-parser from consuming the trailing `SETTINGS` into the last SELECT. The flag is inherited down the format frame, so it also reached the argument of `view` / `viewIfPermitted`. There the parentheses are both redundant and rejected: the closing paren of the table function already terminates the select, and `ViewLayer::parse` bails out on a parenthesized lone select (`ExpressionListParsers.cpp:2981-2991`). The formatted text therefore did not parse back. The trailing `SETTINGS` clause is the trigger. Without it nothing sets the flag and the output is already correct. This clears the flag at the two places that cross the `view` argument boundary, which is what three other boundary owners already do: `ASTSubquery.cpp:98`, `ASTCreateQuery.cpp:1113` and `ASTAlterQuery.cpp:1032`. `ViewLayer` is the only producer of a bare-select function argument and serves exactly these two functions, so the two sites are the complete set. The issue frames parser-versus-formatter as the fork. This takes the formatter side, on the grounds that the parentheses carry no meaning in this position and that the same reset is the established pattern; the parser side would widen the accepted grammar to fix an output bug. The call is his to close. One output change beyond the broken shape: a multi-select argument such as `view(SELECT 1 UNION ALL SELECT 2)` also loses its branch parentheses. That form parsed before and parses now. Validation: `DESCRIBE TABLE view(SELECT 1) SETTINGS input_format_orc_use_fast_decoder = 0` aborts a debug server with `Logical error: 'Inconsistent AST formatting'` before the change and returns `1 UInt8` after it. The new test fails on the unpatched binary and passes 50/50 under randomized settings on the patched one. The `formatQuery` / round-trip / `1941` stateless family was run on both binaries and the failure sets are byte-identical.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114658",
        "createdAt": "2026-08-13T15:46:03Z",
        "updatedAt": "2026-08-13T17:30:00Z",
        "timestamp": "2026-08-13T17:30:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114659",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Record TopK-filtered granules in the query condition cache",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Granules fully filtered by the dynamic threshold of an `ORDER BY ... LIMIT n` (TopK) read are now recorded in the query condition cache, so repeat runs of such queries skip them at the mark-selection stage instead of re-reading and re-filtering most of the table. On the ClickBench Q24/Q26 shapes (`SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10`), a warm run drops from ~95M rows / ~12000 granules read to ~1.3M rows / ~160 granules, and from ~15–30 ms to ~8 ms on a 96-core machine. The `ORDER BY ... LIMIT n` (TopK) optimization pushes a dynamic `__topKFilter` threshold into the `MergeTree` read as a PREWHERE, dropping rows that cannot beat the running top-N. But the granules it emptied were never recorded in the query condition cache: the PREWHERE write path in `MergeTreeSelectProcessor::read` rejects non-deterministic conditions, and `__topKFilter` is one. Only the downstream WHERE `FilterTransform` wrote entries, at whole-chunk granularity, which learns almost nothing (a single surviving row voids the attribution of the whole chunk) — measured on hits, warm runs still selected 11766 of 12348 granules and re-read ~95M of 100M rows. The original out-of-tree build of this optimization (PR #81944, commit `2e2f594308f4` plus the follow-up making the TopN dynamic filters deterministic and reusable for the cache) recorded them and skipped ~98% of granules on warm runs — that is the 2–3x Q24/Q26 gap measured in issue #114639. Recording these granules is sound: for a fixed plan and data, the running threshold only tightens, so a granule none of whose rows survive the filter contains no row that could reach the final top-N, regardless of the threshold trajectory. The entries are keyed with the TopK plan salt (`TopKFilterInfo::condition_hash`: sort column, type, `LIMIT`, direction, `NULLS` direction, collation locale, number of sort columns, and the part-set snapshot), mirroring what the WHERE write path already does since #104478/#110507 — only the same TopK plan over the same part set ever reuses them. Changes: - `isDeterministicAllowingTopKFilter` (two identical static copies in `updateQueryConditionCache.cpp` and `ReadFromMergeTree.cpp`) moved into `VirtualColumnUtils`. - The TopK salt is plumbed to the reader via `MergeTreeReaderSettings::query_condition_cache_top_k_salt`. - The PREWHERE write path accepts a `__topKFilter`-bearing condition when the salt is present, folding the salt into the cache key. Any other non-deterministic condition is still never cached. - The PREWHERE consult path (`filterPartsByQueryConditionCache`) applies the salt exactly when the PREWHERE contains `__topKFilter`, so the keys match the write side; deterministic user PREWHERE conditions keep using plain, unsalted entries shared with non-TopK queries. - `04217_query_condition_cache_topk` and `04242_query_condition_cache_topk_collate` pin exact cache entry counts, which double (each TopK plan now writes a WHERE entry and a PREWHERE entry per part). - New test `04891_query_condition_cache_topk_prewhere_granules` isolates the PREWHERE write path with a WHERE-less TopK query (on current master such a query writes no cache entries at all), asserts that the ASC and DESC plans do not share granule decisions, and that a warm run selects fewer marks. - New performance test `topk_query_condition_cache.xml` with the ClickBench Q24/Q26 shapes over `hits_100m_single`. Measured on the 100M-row single-part ClickBench `hits` (96-core aarch64, hot, interleaved runs): warm Q24/Q26 at 0.007–0.009 s, on par with the original bench-opt build (0.009–0.011 s), against 0.015–0.03 s for current master; warm runs read 159 of 12348 granules against master's 11766. A first run with an empty cache shows no measurable overhead (0.016–0.019 s with the mechanism on and off). Closes: https://github.com/ClickHouse/ClickHouse/issues/114639 Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/110507",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114659",
        "createdAt": "2026-08-13T15:46:11Z",
        "updatedAt": "2026-08-13T17:29:49Z",
        "timestamp": "2026-08-13T17:29:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114660",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Stop SeaweedFS deleting live object folders in stateless CI",
        "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113828 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Since 2026-08-12, `s3 storage` stateless jobs intermittently fail an ordinary INSERT with `Code: 499 ... Immediately after upload: Object ... suddenly disappeared (S3_ERROR)`. The object really was destroyed: SeaweedFS's asynchronous empty-folder cleaner counts a folder's entries and then deletes the folder without excluding a PUT arriving in between, so a write landing in that window is lost. From a failing master run's own artifacts: ``` 16:18:17.675651 EmptyFolderCleaner: deleting empty folder /buckets/test/test/jys 16:18:17.676290 PUT of the object SUCCEEDS, 282 bytes 16:18:17.678805 verifying HEAD -> 404 16:18:17.681060 WriteBufferFromS3: Nothing to abort (so ClickHouse did not remove it) ``` This started with #113828, which replaced MinIO with SeaweedFS. ClickHouse keys objects as `<prefix>/<3 chars>/<random>`, so the bucket holds thousands of shallow folders that empty and refill continuously, and that churn arms the race: in one job 34,819 of 45,578 cleaner log lines were deletions, peaking at 3,238 per minute. CIDB has 0 hits in the preceding 180 days, then hits on 2026-08-12 only. I set the filer's empty-folder cleanup delay past any job's lifetime, so no folder becomes eligible for deletion while the suite runs. The cleaner keeps running and keeps queueing; only its eligibility window moves, and the folders persist in an instance destroyed at job end. No `src/` change: `s3_check_objects_after_upload` caught a genuinely destroyed write. Retrying or relaxing it would have hidden a real lost object. Validated against the shipped script, arms differing only by these lines: without the change the cleaner deletes 13 folders and 0 of 12 survive; with it, 0 deletions and 12 of 12 survive while the queue still holds its 12 items past 3m03s, where the baseline drained at 2m03s. The race is upstream's, which already carves `.uploads` out of this same cleaner for the identical reason (`empty_folder_cleaner.go:247`); this only keeps CI out of its way.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114660",
        "createdAt": "2026-08-13T15:52:45Z",
        "updatedAt": "2026-08-13T17:13:04Z",
        "timestamp": "2026-08-13T17:13:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-ci"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114661",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert \"Document that PREWHERE filters one join input before the JOIN\"",
        "text": "Reverts ClickHouse/ClickHouse#114484 - it is too low-quality, sorry. CC @PedroTadim @dhtclk <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1350` (included in `26.8` and later) <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114661",
        "createdAt": "2026-08-13T15:53:35Z",
        "updatedAt": "2026-08-13T17:56:23Z",
        "timestamp": "2026-08-13T17:56:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-not-for-changelog",
          "pr-synced-to-cloud"
        ],
        "author": "rschu1ze",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114662",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113450 to 26.5: Fix reading Paimon tables with a nullable ARRAY or MAP column",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114662",
        "createdAt": "2026-08-13T16:08:00Z",
        "updatedAt": "2026-08-13T16:08:37Z",
        "timestamp": "2026-08-13T16:08:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "JiaQiTang98",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114663",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113450 to 26.6: Fix reading Paimon tables with a nullable ARRAY or MAP column",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114663",
        "createdAt": "2026-08-13T16:08:31Z",
        "updatedAt": "2026-08-13T17:32:57Z",
        "timestamp": "2026-08-13T17:32:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "JiaQiTang98",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114664",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Backport #113450 to 26.7: Fix reading Paimon tables with a nullable ARRAY or MAP column",
        "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114664",
        "createdAt": "2026-08-13T16:09:00Z",
        "updatedAt": "2026-08-13T16:09:46Z",
        "timestamp": "2026-08-13T16:09:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backport"
        ],
        "author": "robot-clickhouse-ci-1",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "JiaQiTang98",
          "groeneai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114665",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Allow overriding the HTTP method for SELECT through the url table function and URL engine",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/62352 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `http_method='POST'` — as a key-value argument of the `url` table function and the `URL` table engine, or through the `http_method`/`method` named collection keys — now applies to `SELECT` queries: reads use `POST` instead of the default `GET`, for servers that accept only `POST`. Schema inference follows the configured method. `PUT` keeps its write-only meaning: a `SELECT` through a configuration with `http_method='PUT'` still uses `GET`. ### Implementation notes - `IStorageURLBase::getReadMethod()` now returns `POST` when the configured `http_method` is `POST`. `PUT` still applies to writes only (pre-signed upload URLs, #44326), and `INSERT` behavior is unchanged: `POST` by default, `PUT` when configured. - The write path no longer mutates the shared `http_method` member when defaulting to `POST` — the storage instance is shared between queries, so persisting the write default would have flipped subsequent reads of the same table from `GET` to `POST` (and raced concurrent reads). - Schema inference (`getTableStructureAndFormatFromData` and the URL read-buffer iterator) uses the same effective read method instead of hardcoded `GET`, for both `url()`/`URL` and `urlCluster`. - The `http_method = '...'` (or `method = '...'`) key-value argument is parsed by `StorageURL::evalArgsAndCollectHeaders` alongside `headers(...)`, kept in the `CREATE` AST (so it survives `SHOW CREATE TABLE` and DETACH/ATTACH), validated to be `POST`/`PUT`, and rejected when the URL scheme dispatches to another backend (`file://`, `s3://`, ...), mirroring `headers(...)`. - In the analyzer, the argument's left-hand identifier is excluded from column resolution (`skipAnalysisForArguments`), like `headers(...)`. - Documentation for the `url` table function and the `URL` engine is updated. The new stateless test `04869_url_select_http_method` asserts the method actually used on the wire via `system.query_log.http_method`: default `GET`, `POST` override for both the data and the schema-inference requests, `PUT`-configured `SELECT` staying on `GET`, the `URL` engine path, rejection of unsupported methods, and persistence of the argument in the table DDL. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114665",
        "createdAt": "2026-08-13T16:17:38Z",
        "updatedAt": "2026-08-13T18:00:25Z",
        "timestamp": "2026-08-13T18:00:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-feature",
          "can be tested"
        ],
        "author": "valerypetrov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114666",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Promote ConstantValue to Core with Field-free value accessors",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113051 --> Continue removing `DB::Field` from constant handling by promoting the owning single-constant type to a shared primitive with `Field`-free value accessors. ### What - **Move `ConstantValue` from `Analyzer/` to `Core/`.** It bundles `{size-1 ColumnConst, DataTypePtr}`. It carries a `DataTypePtr`, and `DataTypes` depends on `Columns`, so `Core` is the correct layer (this also removes a lower layer having to reach up into `Analyzer` for the type). It stays deliberately distinct from `ColumnWithTypeAndName` (size-1 const invariant, no name, scalar accessors). - **Add value accessors** that read row 0 of the size-1 column without materializing a `Field`: `isNull`, `getUInt`, `getInt`, `getFloat64`, `getBool`, `getDataAt`. `ConstantNode` gains matching delegators, and `ConstantNode::getValue` now delegates to a single transitional `ConstantValue::getField`. - **`evaluateConstantExpressionAsColumn` now returns `ConstantValue`** instead of `std::pair<ColumnPtr, DataTypePtr>`. Callers updated: `numbers`/`primes`/`generateSeries`/`values` table functions, `ActionsVisitor`, `InterpreterSelectQuery` (LIMIT/OFFSET), prometheus/timeSeries selectors, and the gtest. The two `getStringConstArgument` helpers now read the value directly via `isNull`/`getDataAt`. ### Behavior No user-visible change: the value is the same size-1 const column with the same exact type, just bundled and readable without a `Field`. `ConstantValue`'s `Field` constructor and `getField` remain as the only `Field` entry/exit while the analyzer still folds constants into `Field`s; a later phase removes them. ### Verification `gtest_convert_column_to_type`, `gtest_evaluate_constant_expression`, and a `numbers`/`generate_series`/`values`/`LIMIT` (integer + fractional)/`IN` stateless spot-check. Related: https://github.com/ClickHouse/ClickHouse/pull/113051 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not for changelog: internal refactor toward removing `DB::Field`; no user-visible behavior change.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114666",
        "createdAt": "2026-08-13T16:25:52Z",
        "updatedAt": "2026-08-13T16:49:57Z",
        "timestamp": "2026-08-13T16:49:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "yakov-olkhovskiy",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114667",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix File row policy / PREWHERE breaking DEFAULT columns missing from the data file",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114616 When a file-backed table (`File`, and the same path for object storage) has a `DEFAULT` column that is not present in the data file, a row policy or `PREWHERE` could prune the inputs of that default expression before `AddingDefaultsTransform` ran. That led to `UNKNOWN_IDENTIFIER` on current master, and on 26.7 to silently wrong row-policy results (type defaults instead of real values). This keeps columns required by `DEFAULT` expressions in the read set through prewhere pruning, and applies filters that reference defaulted columns after defaults are computed. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fix queries over file-backed tables where a row policy or `PREWHERE` interacted with `DEFAULT` columns missing from the data file: they no longer fail with `UNKNOWN_IDENTIFIER`, and row policies are evaluated on real default values instead of type defaults.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114667",
        "createdAt": "2026-08-13T16:31:30Z",
        "updatedAt": "2026-08-13T16:52:46Z",
        "timestamp": "2026-08-13T16:52:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "Ria-K912",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114668",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Stop Parquet background reads before releasing the format's read buffer",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114612 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed heap memory corruption when reading Parquet through an input format that owns its read buffer, for example a dictionary with `SOURCE(FILE(... format 'Parquet'))`. Background prefetch and decode tasks could still read and write through the buffer after the pipeline released it, which could abort the server. ### Description Closes #114612, reported by @ PedroTadim, who asked me to go ahead in https://github.com/ClickHouse/ClickHouse/issues/114612#issuecomment-5278451991. A `ReadBuffer` handed to an input format via `addBuffer()` lives in the format's `owned_buffers`. `ISource::work()` calls `onFinish()` on both the clean and the exception path, reaching `IInputFormat::resetReadBuffer()`, which clears `owned_buffers` and destroys the buffer. `ParquetV3BlockInputFormat` did not override that hook, and `Parquet::Prefetcher` holds a non-owning `SeekableReadBuffer *` to the same buffer, so background tasks kept using a destroyed object. It is not only a bad read: the report on the issue is a 1 MiB **write** into freed heap through `ReadBuffer::next()`, silent on a release build. `~Prefetcher()` does the right handshake, but only at format destruction, later than `onFinish()`; that window is the bug. The fix overrides `resetReadBuffer()` to drain background tasks before the base class releases the buffers, mirroring `ParallelParsingInputFormat::onFinish()`. Three details are forced by the surrounding code: the hook is `resetReadBuffer()`, not `onFinish()`, since it frees the buffers and has a second caller in `StreamingFormatExecutor`; the reader is drained but kept alive, because `getMatchedBuckets()` reads row group metadata after exhaustion; and `ReadManager` is drained before `Prefetcher`, because decode tasks re-enter `readSync` inline via `getRangeData()`. Sibling teardown paths: `resetParser()` already destroys the reader before delegating, `onCancel()` cancels it, and the reuse route through `StreamingFormatExecutor` calls `resetParser()` on every exit, so a drained reader is never reused. No fuzzer needed: `ReadBufferFromFile` leaves `use_pread` false, so a local file takes `SeekAndRead`, and any `CREATE DICTIONARY ... SOURCE(FILE(... format 'Parquet'))` whose load throws reaches it. Validated on ASAN: the added test fails 7/7 unfixed, passes 11/11 fixed; a standalone reproducer 31/33 unfixed, 0/35 fixed; a negative control with only the override reverted reddens at the unfixed rate. Parquet and polygon-dictionary suites gave an identical failing set on both binaries (187 each); 50 randomized-settings repeats passed clean.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114668",
        "createdAt": "2026-08-13T16:50:49Z",
        "updatedAt": "2026-08-13T17:53:37Z",
        "timestamp": "2026-08-13T17:53:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114669",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: Clarify `PREWHERE` behavior with `JOIN`",
        "text": "This clarifies how `PREWHERE` behaves in queries with `JOIN`. It removes the misleading claim that `SELECT` clauses follow a single execution order, clearly states the single-table restriction, and adds a runnable minimal example with expected output. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required for a documentation change.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114669",
        "createdAt": "2026-08-13T16:57:35Z",
        "updatedAt": "2026-08-13T17:36:29Z",
        "timestamp": "2026-08-13T17:36:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "pr-documentation"
        ],
        "author": "dhtclk",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114670",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Docs: expand the Managed Postgres autoscaling documentation",
        "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ## Summary - Expand the one-line Autoscaling section in the Managed Postgres scaling docs with the behavior sourced from the Ubicloud codebase: the 85% storage notification, the 90% automatic scale-up, and the 95% maintenance-window bypass - Document what autoscaling means for the cutover and client connections: same process as a manual instance change, connections dropped and in-flight transactions rolled back, DNS repointed to the new primary under the same hostname - Document read-only mode: free-space trigger and recovery thresholds per disk size, the error writes receive, and automatic recovery after the scale-up - Add a worked example scenario for a 1024 GB instance scaling to 2048 GB",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114670",
        "createdAt": "2026-08-13T16:59:03Z",
        "updatedAt": "2026-08-13T17:52:13Z",
        "timestamp": "2026-08-13T17:52:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-documentation",
          "can be tested"
        ],
        "author": "amogiska",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114671",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Make release creation steps self-gating and drop fail-closed flag defaults",
        "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/114472 --> Makes the release-creation steps self-gating so an idempotent operation decides for itself whether there is work to do, instead of the orchestrator gating on `is_recovery` / `is_late_recovery` flags with fail-closed `True` defaults. - `ci/jobs/scripts/create_release.py`: `push_release_tag` self-skips when `Git.tag_exists(release_tag)` — a recovery finds the tag already published and does nothing. `update_version_and_contributors_list` self-skips a superseded (late) recovery for a patch release, so it never rewrites the branch version backwards. - `ci/jobs/release_job.py`: the tag-push step now runs unconditionally and the patch bump runs for every patch release. The `is_recovery = True` / `is_late_recovery = True` defaults are removed; `is_recovery` is read only under `ok` (remaining reads short-circuit on it), and `is_late_recovery` is no longer needed at the top level (only the docker floating-tag block reads it, from its own release info). Behavior-preserving: the tag-exists skip derives from the same local `Git.tag_exists` that `prepare` used to compute `is_recovery`, and the late-recovery skip mirrors the old `not is_late_recovery` bump gate. Addresses the review that an idempotent step should not need an `is_recovery` gate. Stacked in content on #114472; until it merges the diff shows its commits too. The only new commit here is `Make release creation steps self-gating and drop fail-closed flag defaults`. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114671",
        "createdAt": "2026-08-13T17:33:22Z",
        "updatedAt": "2026-08-13T17:43:01Z",
        "timestamp": "2026-08-13T17:43:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "can be tested",
          "pr-ci"
        ],
        "author": "leshikus",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114672",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Do not hold the DDL guard during TRUNCATE",
        "text": "`TRUNCATE TABLE` held the table's DDL guard while removing data, and truncate can wait for running merges or for other replicas to process `DROP_RANGE`. Any DDL on that name — including background threads that need it — was blocked for that whole time. Release the guard before `IStorage::truncate` for databases with UUIDs, matching `ALTER TABLE ... DROP PARTITION`, which never took it. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `TRUNCATE TABLE` no longer blocks concurrent `DROP`, `RENAME` and other DDL queries on the same table while the data is being removed.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114672",
        "createdAt": "2026-08-13T17:52:48Z",
        "updatedAt": "2026-08-13T17:53:50Z",
        "timestamp": "2026-08-13T17:53:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "evillique",
        "state": "open",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:114673",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Read a subcolumn of an `ALIAS` parent through a `Merge` table instead of a default",
        "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix reading a subcolumn (`.size0`, `.null`, a tuple element, `String.size`) through a `Merge` table when the underlying table declares the parent column as `ALIAS`. Such a read returned the type default, and when the parent was selected in the same query it could return the value of an unrelated column. ### Description Selecting `arr.size0` from a `Merge` table returns `0` when the child declares `arr` as an `ALIAS` column, while the same query against the child returns the correct `5`: ```sql CREATE TABLE ag (n UInt64, arr Array(UInt8) ALIAS [1,2,3,4,5]) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO ag (n) VALUES (77); CREATE TABLE mg (arr Array(UInt8), n UInt64) ENGINE = Merge(currentDatabase(), '^ag$'); SELECT arr.size0 FROM mg; -- 0, should be 5 ``` Root cause: `ColumnsDescription::add` registers no subcolumns for an `ALIAS` column, because the value has to be extracted after the expression is evaluated. So `arr.size0` does not resolve against the child, but does against the `Merge` table, where `arr` is ordinary. `ReadFromMerge::getModifiedQueryInfo` reads that as \"the child does not have this column\": it substitutes a constant default, and its `with_aliases` loop skips the column, so the child is never asked for the data. With the parent also selected, alias expansion put the alias's dependency column in the child read list, and the misaligned read surfaced that value under the subcolumn's name. The analyzer already implements the extract-after-evaluation contract, turning such a name into `getSubcolumn` over the alias expression. This teaches both guards to recognise the case and route it to the existing alias branch with the full identifier, adding no new carrier. Also wrong for `Tuple.a`, `String.size` and `Nullable.null` (a NULL read back as non-NULL), under `optimize_functions_to_subcolumns` 0 and 1, and it mis-scoped a row policy. `Map` subcolumns raise `NO_SUCH_COLUMN_IN_TABLE` before and after. The old analyzer raises `UNKNOWN_IDENTIFIER` and is unchanged, hence the `no-old-analyzer` tag. `MATERIALIZED` and `DEFAULT` parents were already correct, and the default substitution for a child genuinely lacking a column is preserved; both have test arms. A mistyped child (`String ALIAS` under an `Array(UInt8)` Merge column) stays as it is, because `String` has no `size0`. That conversion case is #112975.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/114673",
        "createdAt": "2026-08-13T17:59:07Z",
        "updatedAt": "2026-08-13T18:01:28Z",
        "timestamp": "2026-08-13T18:01:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "groeneai",
        "state": "open",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:42701",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add table function `obfuscate`",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/39067 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add table function `obfuscate` which applies the same transformation as the `clickhouse-obfuscator` tool to the result of an arbitrary query, e.g. `SELECT * FROM obfuscate(SELECT * FROM table)`. The seed can be controlled with the `obfuscate_seed` setting; with an empty seed a fresh random one is derived per execution.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/42701",
        "createdAt": "2022-10-26T13:41:41Z",
        "updatedAt": "2026-08-13T12:32:43Z",
        "timestamp": "2026-08-13T12:32:43Z",
        "metrics": {
          "reactions": 2,
          "comments": 44
        },
        "labels": [
          "pr-feature"
        ],
        "author": "evillique",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:63383",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Improve the performance of `MODIFY TTL`",
        "text": "`ALTER TABLE ... MODIFY TTL` currently rewrites every part of the table, which on a large table means reading and writing all of its data just to change when rows expire. Very often the new TTL is the old one shifted in time - the retention period is extended or shortened, e.g. `create_time + INTERVAL 300 DAY` becomes `create_time + INTERVAL 10 DAY`. In that case every row's expiry time moves by the same constant number of seconds, so the parts do not have to be rewritten at all: it is enough to shift the expiry timestamps ClickHouse already stores per part. This pull request adds that fast path, and the result for the user is that such a `MODIFY TTL` completes almost instantly instead of taking minutes or hours: ```sql CREATE TABLE test_fast_ttl (`id` UInt32, `name` String, `create_time` DateTime) ENGINE = MergeTree ORDER BY id TTL create_time + toIntervalDay(300); INSERT INTO test_fast_ttl SELECT number, 'AAA', date_sub(day, 100, now()) from numbers(100000000); -- Before ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 25.564 sec. -- After ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 0.046 sec. ``` There is nothing to enable and no new syntax: the optimization is applied automatically inside the `MATERIALIZE TTL` mutation that `MODIFY TTL` already produces, and a plain `ALTER TABLE ... MATERIALIZE TTL` benefits from it as well. The observable result is exactly the same as before - the same rows expire and the parts end up with the same TTL bounds - only the work is avoided. Per part, the mutation now does one of the following: - the part is fully expired under the new TTL - it is replaced with an empty part; - no row of the part is expired yet - the part is cloned (its data files hardlinked) and only its stored TTL bounds are shifted; - otherwise - the part is rewritten exactly as before. The fast path is only taken when it is provably equivalent to the rewrite. It requires that the unconditional rows TTL (`TTL <expr>`) is the only TTL of the table, and that the old and the new TTL are the same date/time column shifted by constant fixed-length intervals, so that `new_ttl(row) - old_ttl(row)` is one constant for every row. Calendar `MONTH`/`YEAR` intervals, `DAY`/`WEEK` intervals in a time zone with daylight saving time, and row-dependent expressions are all rejected. The proof is redone for each part against the TTL expression (and time zone) that the part's stored timestamps were actually computed under, so a part that lags the table metadata, or was written by an older server, falls back to the regular rewrite rather than being shifted unsoundly. The same applies to the boundary cases of the stored timestamps themselves: a part containing a row whose TTL timestamp is exactly `1970-01-01 00:00:00` UTC (which ClickHouse treats as \"no TTL\"), a shift that would move some timestamp onto that value, and a part whose stored TTL is already fully expired (which the regular rewrite drops wholesale, even when the new TTL is longer) all take the regular rewrite. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `ALTER TABLE ... MODIFY TTL` no longer rewrites the table's data when the new TTL is the old one shifted by a constant amount of time (the same date/time column plus fixed-length intervals), which is the common case of extending or shortening the retention period. Fully expired parts are replaced with empty ones and the rest are cloned with only their stored TTL metadata shifted, which makes such an `ALTER` nearly instant. Cases where the shift is not provably constant - calendar month/year intervals, day/week intervals in a time zone with daylight saving time, or row-dependent TTL expressions - fall back to the regular rewrite. `ALTER TABLE ... MATERIALIZE TTL` benefits from the same optimization. <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> > Information about CI checks: https://clickhouse.com/docs/en/development/continuous-integration/ <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration Tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASAN - [ ] <!---ci_include_tsan--> All with TSAN - [ ] <!---ci_include_analyzer--> All with Analyzer - [ ] <!---ci_include_azure --> All with Azure - [ ] <!---ci_include_KEYWORD--> Add your option here #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration Tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASAN - [ ] <!---ci_exclude_tsan--> All with TSAN - [ ] <!---ci_exclude_msan--> All with MSAN - [ ] <!---ci_exclude_ubsan--> All with UBSAN - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 - [ ] <!---ci_exclude_KEYWORD--> Add your option here #### Extra options: - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 <details>",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/63383",
        "createdAt": "2024-05-05T15:40:47Z",
        "updatedAt": "2026-08-13T17:02:37Z",
        "timestamp": "2026-08-13T17:02:37Z",
        "metrics": {
          "reactions": 1,
          "comments": 33
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "zhongyuankai",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:64184",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Adding storage pulsar",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category: - Experimental Feature ### Changelog entry: Added the experimental `Pulsar` table engine for reading from and writing to Apache Pulsar topics. Creating tables with the engine requires enabling the `allow_experimental_pulsar_storage_engine` setting. Closes: https://github.com/ClickHouse/ClickHouse/issues/4226",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/64184",
        "createdAt": "2024-05-21T10:18:47Z",
        "updatedAt": "2026-08-13T12:55:13Z",
        "timestamp": "2026-08-13T12:55:13Z",
        "metrics": {
          "reactions": 2,
          "comments": 7
        },
        "labels": [
          "submodule changed",
          "manual approve",
          "can be tested",
          "pr-experimental"
        ],
        "author": "SteveBalayanAKAMedian",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:68493",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Mongo queries (wire protocol + dialect)",
        "text": "Adds two ways to run MongoDB queries against ClickHouse: - a **wire protocol endpoint** (`mongo_port`), so MongoDB drivers and tools such as `pymongo` and `mongosh` can connect to ClickHouse as if it were a MongoDB server; - a **query dialect** (`SET dialect = 'mongo'`), so MongoDB shell syntax can be sent over the usual ClickHouse interfaces. A Mongo database maps onto a ClickHouse database, a collection onto a table, a document onto a row, and a nested field onto an `a.b` column. A collection created by the first `insert` gets one column per field of the first inserted document. The supported commands are `insert`, `find`, `count`, `update`, `delete`, `create`, `drop`, `createIndexes`, `listDatabases`, `listCollections`, `isMaster` and `saslStart`; only the `PLAIN` authentication mechanism is supported, because it is the only one that provides the cleartext password ClickHouse needs. Both are experimental and cover a subset of MongoDB. They are meant for pointing an existing MongoDB application at ClickHouse without rewriting its queries, not as a MongoDB replacement. The limitations are listed in `docs/en/interfaces/mongo.md`. Task: https://github.com/ClickHouse/ClickHouse/issues/58394 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a MongoDB-compatible wire protocol endpoint (the `mongo_port` server setting) and a MongoDB query dialect (`SET dialect = 'mongo'`) for basic collection operations.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/68493",
        "createdAt": "2024-08-17T07:09:19Z",
        "updatedAt": "2026-08-13T03:18:54Z",
        "timestamp": "2026-08-13T03:18:54Z",
        "metrics": {
          "reactions": 1,
          "comments": 25
        },
        "labels": [
          "can be tested",
          "pr-experimental"
        ],
        "author": "scanhex12",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:71028",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Randomize parallel_replicas_min_number_of_rows_per_replica",
        "text": "<!--- Disable AI PR formatting assistant: true --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Details The setting enables code execution, which can trigger hidden bugs, in particular in GLOBAL JOINs with parallel replicas. Discovered one while doing https://github.com/ClickHouse/ClickHouse/pull/70658 within `02967_parallel_replicas_joins_and_analyzer` test",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/71028",
        "createdAt": "2024-10-24T14:24:19Z",
        "updatedAt": "2026-08-13T17:23:21Z",
        "timestamp": "2026-08-13T17:23:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 24
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "devcrafter",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:73776",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Enable dynamic evaluation of whether a short-circuit function's argument should be lazily executed",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Lazy execution in short-circuit evaluation is not entirely cost-free. The `maskedExecute` method incurs additional overhead for filtering and expanding columns. In certain scenarios, this overhead can exceed the cost of fully evaluating the expression. In this PR, we have added runtime performance profiling for functions. Based on the collected runtime performance data, we can evaluate the cost of using lazy evaluation versus direct full computation for a function, and dynamically select the execution strategy with the lower cost. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> > Information about CI checks: https://clickhouse.com/docs/en/development/continuous-integration/ #### CI Settings (Only check the boxes if you know what you are doing): - [ ] <!---ci_set_required--> Allow: All Required Checks - [ ] <!---ci_include_stateless--> Allow: Stateless tests - [ ] <!---ci_include_stateful--> Allow: Stateful tests - [ ] <!---ci_include_integration--> Allow: Integration Tests - [ ] <!---ci_include_performance--> Allow: Performance tests - [ ] <!---ci_set_builds--> Allow: All Builds - [ ] <!---batch_0_1--> Allow: batch 1, 2 for multi-batch jobs - [ ] <!---batch_2_3--> Allow: batch 3, 4, 5, 6 for multi-batch jobs --- - [ ] <!---ci_exclude_style--> Exclude: Style check - [ ] <!---ci_exclude_fast--> Exclude: Fast test - [ ] <!---ci_exclude_asan--> Exclude: All with ASAN - [ ] <!---ci_exclude_tsan|msan|ubsan|coverage--> Exclude: All with TSAN, MSAN, UBSAN, Coverage - [ ] <!---ci_exclude_aarch64|release|debug--> Exclude: All with aarch64, release, debug --- - [ ] <!---ci_include_fuzzer--> Run only fuzzers related jobs (libFuzzer fuzzers, AST fuzzers, etc.) - [ ] <!---ci_exclude_ast--> Exclude: AST fuzzers --- - [ ] <!---do_not_test--> Do not test - [ ] <!---woolen_wolfdog--> Woolen Wolfdog - [ ] <!---upload_all--> Upload binaries for special builds - [ ] <!---no_merge_commit--> Disable merge-commit - [ ] <!---no_ci_cache--> Disable CI cache",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/73776",
        "createdAt": "2024-12-24T06:56:51Z",
        "updatedAt": "2026-08-13T00:57:54Z",
        "timestamp": "2026-08-13T00:57:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 13
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "lgbo-ustc",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:76595",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Parallel Replicas: a setting to disable for queries with multiple tables",
        "text": "Queries with a `JOIN` read the non-leftmost side in full on every replica, which can make parallel replicas slower than a plain single-node execution. This adds a kill switch so such queries can be excluded from parallel replicas without disabling parallel replicas altogether. New setting `parallel_replicas_for_queries_with_multiple_tables` (default `true`, i.e. the pre-existing behaviour). When set to `false`, parallel replicas are not used for a query that joins multiple tables: the decision is applied in `findParallelReplicasQuery` (query/table selection) and in `buildJoinTreeQueryPlan`, and it is propagated into subquery table expressions — including `UNION` table expressions and their branches, which are planned by independent `Planner` instances — and into the `IN` subqueries collected into prepared sets, which are also planned by independent `Planner` instances built from the subqueries' own contexts. The legacy (pre-analyzer) interpreter respects the setting as well: when `parallel_replicas_only_with_analyzer = 0` allows task-based parallel replicas there, `InterpreterSelectQuery` applies the same kill switch before the storage read. `ARRAY JOIN` does not count as a join between tables, and a `UNION` query without a `JOIN` is not affected: each `UNION` branch is an independent single-table read, so parallel replicas remain applicable to it. ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Added the setting `parallel_replicas_for_queries_with_multiple_tables` (default `true`) to control whether parallel replicas are used for queries joining multiple tables. When disabled, parallel replicas are not used for queries with `JOIN`, where the non-leftmost side is read in full on every replica. The setting does not affect a `UNION` query without a `JOIN` (each branch is an independent single-table read) nor `ARRAY JOIN`.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/76595",
        "createdAt": "2025-02-21T17:53:10Z",
        "updatedAt": "2026-08-13T00:28:54Z",
        "timestamp": "2026-08-13T00:28:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 14
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "devcrafter",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:76867",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Minmax indices by default",
        "text": "<!--- Disable AI PR formatting assistant: true --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): MergeTree tables will have `add_minmax_index_for_numeric_columns` by default. Resolves #70605 ### Details Resolves #70605 <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > Changes the default MergeTree storage behavior by enabling per-numeric-column min-max skipping indexes, which can affect ingest throughput, disk usage, and query plans across all newly created tables unless explicitly disabled. > > **Overview** > **Enables implicit min-max skipping indices on numeric columns by default** by flipping the MergeTree setting `add_minmax_index_for_numeric_columns` to `true` and recording the default change in settings history/compatibility notes. > > To avoid regressions in write-heavy/system tables and test baselines, the PR **forces `add_minmax_index_for_numeric_columns = 0`** when auto-building system log table engines, updates CI/stateful dataset setup and many integration/stateless tests to opt out explicitly, and adds new stateless coverage to verify both the new default and the opt-out behavior. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit e7a97c8b8c10055b2026c1269a674b2e2f73decf. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/76867",
        "createdAt": "2025-02-27T10:47:59Z",
        "updatedAt": "2026-08-13T07:58:42Z",
        "timestamp": "2026-08-13T07:58:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 79
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "devcrafter"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:79393",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Query result cache on disk",
        "text": "On-disk configuration for query result cache is provided in addition to the already existing in-memory configuration. When data is written to the cache, it will be written to both of them (write-through). When data is searched inside the cache, it is searched in memory first, then on disk. If the data is found on disk, it will be also put in memory. On-disk cache has independent configurations of max size, max elements and max element size. ``` <query_cache> <max_size_in_bytes>1073741824</max_size_in_bytes> <max_entries>1024</max_entries> <max_entry_size_in_bytes>1048576</max_entry_size_in_bytes> <max_entry_size_in_rows>30000000</max_entry_size_in_rows> <on_disk> <max_size_in_bytes>2147483648</max_size_in_bytes> <max_entries>2048</max_entries> <max_entry_size_in_bytes>2097152</max_entry_size_in_bytes> <max_entry_size_in_rows>60000000</max_entry_size_in_rows> </on_disk> </query_cache> ``` The data is stored in subdirectory `query_cache/` relative to the base directory (`<path>...</path>` in the server configuration). For each query result, a file is created, with the AST tree hash as file name. To support future changes of the persistence format, a file `query_cache/format_version.txt` stores the version of the data, e.g. `1`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an optional on-disk tier for the query result cache. Query results are written through to disk (under the `query_cache/` directory) in addition to memory, so they survive server restarts. The on-disk cache has its own `max_size_in_bytes`, `max_entries`, `max_entry_size_in_bytes`, and `max_entry_size_in_rows` limits, configured under `<query_cache><on_disk>` in the server configuration. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/79393",
        "timestamp": "2026-08-12T22:24:37Z",
        "metrics": {
          "reactions": 4,
          "comments": 21
        },
        "labels": [
          "pr-feature",
          "can be tested",
          "hold"
        ],
        "author": "nbarannik",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:79490",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Real cache size",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Added a new filesystem cache setting `use_real_disk_size` (disabled by default). When enabled, cached-file sizes are accounted in filesystem allocation units (block size) instead of the written byte count, so the `FilesystemCacheSize` metric, the eviction-size metrics, and `current_size` in `system.filesystem_cache_settings` use block-aligned accounting. Without it, a cached file smaller than one block (for example, a 1-byte file) is accounted as its written size while occupying a whole block on disk, which can make the reported cache size much smaller than the actual on-disk usage. The accounting is a block-aligned approximation of the physical size, not the exact allocated size. ### Documentation entry for user-facing changes - [x] Documentation is written Documented the `use_real_disk_size` filesystem cache setting in `docs/en/operations/storing-data.md`. It accounts cache usage in filesystem allocation units (block size), so `FilesystemCacheSize`, the eviction-size metrics, and `current_size` in `system.filesystem_cache_settings` use block-aligned accounting (a block-aligned approximation of the physical on-disk usage). The per-segment `size` in `system.filesystem_cache` is not affected by this setting and stays the segment's logical range size; for a partially downloaded segment, use `downloaded_size` for the bytes actually written. <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/79490",
        "createdAt": "2025-04-23T15:44:34Z",
        "updatedAt": "2026-08-13T00:59:05Z",
        "timestamp": "2026-08-13T00:59:05Z",
        "metrics": {
          "reactions": 1,
          "comments": 35
        },
        "labels": [
          "pr-improvement",
          "manual approve",
          "can be tested"
        ],
        "author": "codeworse",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:79509",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Async insert parallel parsing",
        "text": "Parses the data of one batch of asynchronous inserts with several threads when the batch is flushed, controlled by the new setting `async_insert_parse_threads` (`0` by default, which keeps the previous behaviour). Addresses https://github.com/ClickHouse/ClickHouse/issues/74162 ### Motivation With `wait_for_async_insert = 1` the client waits for the whole batch to be flushed, and for text formats most of that time is spent parsing the accumulated data in a single thread. That makes the observed `INSERT` latency proportional to the size of the batch, which is what makes `wait_for_async_insert = 1` unusable for some workloads. ### How it works The entries of a batch are split into `async_insert_parse_threads` contiguous ranges. Every range is parsed by its own `StreamingFormatExecutor` with its own input format (the format holds mutable parsing state and cannot be shared), on a thread pool bounded server-wide by the new `max_async_insert_parsing_thread_pool_size` (100 by default). The calling thread parses one of the ranges itself instead of only waiting. This is a pool of its own rather than the format parsing pool, because some input formats (`Parquet`, `ArrowStream`, ...) parallelize their own work on the format parsing pool: a range waiting for such a nested task while holding a thread of that very pool would deadlock once every thread of it is held by a range. Afterwards the ranges are concatenated back into a **single** chunk with a **single** `DeduplicationInfo`, and the per-entry bookkeeping (the `system.asynchronous_insert_log` elements, the deduplication tokens, the rows and bytes reported back to the waiting clients) is replayed by the flushing thread in the original order of the entries. So only parsing is parallelized: the flush still pushes one block into the insert pipeline, the number of parts written per flush does not change, and no shared state is written from more than one thread. The tasks run through `ThreadPoolCallbackRunnerLocal`, which attaches them to the thread group of the flush, so profile events, memory and CPU accounting of the parsing are attributed to the flush query instead of being lost. The number of ranges is clamped to the number of entries in the batch and to the size of the pool, so a batch of 3 inserts never uses more than 3 threads. `0` and `1` mean the flushing thread parses everything and the pool is not touched at all. ### Measurements 2000 asynchronous inserts of 500 `JSONEachRow` rows each, collected into one batch of one million rows and flushed explicitly; the numbers are the `AsyncInsertFlush` entry of `system.query_log`, two runs per value (aarch64, 96 cores, shared machine): | `async_insert_parse_threads` | flush `query_duration_ms` | rows per ms | `UserTimeMicroseconds` | | --- | --- | --- | --- | | `0` | 738, 680 | 1355, 1471 | 634 ms, 627 ms | | `2` | 457, 600 | 2188, 1667 | 629 ms, 643 ms | | `5` | 331, 323 | 3021, 3096 | 628 ms, 638 ms | | `10` | 293, 272 | 3413, 3677 | 632 ms, 625 ms | So the flush of a one-million-row batch goes from ~709 ms to ~327 ms with 5 threads (2.2x) and to ~283 ms with 10 (2.5x), and `UserTimeMicroseconds` stays at ~630 ms throughout - the CPU time of the parsing is accounted to the flush query no matter how many threads did it, and parallelizing it adds no measurable CPU overhead. An earlier measurement of the first version of this change, on a different data set, showed the same effect on latency (~2763 ms against ~663 ms for 5 threads): https://pastila.nl/?007f67ed/3ce4613dcb3ed6f2f8cfc00d6e8bc906#en58FLI4Ilg9Nk2Vp6b8XQ== - but there `UserTimeMicroseconds` collapsed from 2.3 s to 15 ms, because that version did not attach the parsing tasks to the thread group of the flush, so their profile events were lost. That is what the flat `UserTimeMicroseconds` column above verifies is fixed. ### Testing `tests/queries/0_stateless/04869_async_insert_parse_threads.sh` flushes explicitly built batches with 0, 1, 4 and 16 parse threads and checks that all the rows are inserted with their column defaults, that one flush still produces exactly one part, and that a row which fails to parse fails only its own insert (`ParsingError` in `system.asynchronous_insert_log`) while the other entries of the batch are inserted. `async_insert_parse_threads` is randomized in `tests/clickhouse-test`, so the `Stateless tests (AsyncInsert)` job exercises the parallel path across the whole suite. Locally, the 83 stateless tests whose name contains `async_insert` were run twice against the same binary, once with `async_insert_parse_threads = 4` forced for every query and once with `0`, and the two failure sets are identical - 12 failures in both, all of them gaps in the minimal local server config used for the run (no second shard on `127.0.0.2`, no `pandas` for the `.python` tests, and `Settings['async_insert']` not being recorded because the CI `users.d` default is absent). With those gaps filled, `02481_async_insert_dedup` and `02481_async_insert_dedup_token` - the coverage that matters for the rebuilt `DeduplicationInfo` - both pass with 4 parse threads. `02187_async_inserts_all_formats` (tagged `long`) exceeds the 600 s harness timeout on that machine with 4 threads and with 0 alike, so that one is the machine, not the setting; it sends one insert per batch anyway, which clamps to a single range and the plain serial path. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added the setting `async_insert_parse_threads` that specifies how many threads parse the data of one batch of asynchronous inserts when it is flushed. It reduces the latency of `INSERT`s that wait for the flush (`wait_for_async_insert = 1`). Disabled by default. ### Documentation entry for user-facing changes",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/79509",
        "createdAt": "2025-04-23T21:31:36Z",
        "updatedAt": "2026-08-13T01:26:55Z",
        "timestamp": "2026-08-13T01:26:55Z",
        "metrics": {
          "reactions": 2,
          "comments": 24
        },
        "labels": [
          "pr-performance",
          "manual approve",
          "can be tested"
        ],
        "author": "ilejn",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:79991",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Keeper log s3",
        "text": "Experimental support for storing the Keeper changelog directly on object storage (S3) instead of (or in addition to) a local disk, gated behind the coordination setting `s3_experimental_changelog`. Enables running Keeper without a local data volume. Configuration: - `keeper_server.coordination_settings.s3_experimental_changelog` — enable the experimental path. - `keeper_server.coordination_settings.s3_log_disk` — name of the disk from `<storage_configuration>` to use for the changelog. - `keeper_server.coordination_settings.s3_flush_interval` — flush interval in microseconds (default `500`). A background compaction thread merges adjacent S3 changelog files in the background to keep the file count bounded. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Experimental: support storing the Keeper changelog on S3, gated behind the coordination setting `s3_experimental_changelog`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds a new code path for Keeper Raft log persistence to object storage (S3) with background file merging; although gated behind a new setting, it touches log durability/rotation behavior and could affect recovery if enabled. > > **Overview** > Adds **experimental** support for writing Keeper changelogs directly to an S3 disk when `s3_experimental_changelog` is enabled, including selecting the S3 disk via `s3_log_disk`. > > Refactors the changelog writer behind a new `IChangelogWriter` interface and introduces `S3ChangelogWriter`, which writes log segments to S3 and runs a background compaction thread to merge adjacent small segments; local-disk log loading/moving logic is adjusted to avoid cross-disk moves in S3 mode. > > Extends coordination settings / `KeeperContext` to parse the new S3 changelog settings and adds an integration test scaffolding (MinIO-backed) plus a PR template checkbox prompting documentation for user-facing changes. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit fddb3bec925253858e79b944a59194964c39396e. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/79991",
        "timestamp": "2026-08-12T21:08:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 27
        },
        "labels": [
          "can be tested",
          "comp-keeper",
          "pr-experimental"
        ],
        "author": "trololo23",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:80353",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Redis-wire protocol",
        "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Add an opt-in Redis wire-protocol server backed by `Join` tables: a configured `redis.port` serves `GET`/`MGET` and `HGET`/`HMGET` point lookups (plus `AUTH`, `SELECT`, `PING`, `ECHO`) against ClickHouse tables mapped to Redis database numbers. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/80353",
        "createdAt": "2025-05-16T14:55:43Z",
        "updatedAt": "2026-08-13T17:03:46Z",
        "timestamp": "2026-08-13T17:03:46Z",
        "metrics": {
          "reactions": 4,
          "comments": 10
        },
        "labels": [
          "pr-feature",
          "comp-protocols",
          "manual approve",
          "can be tested"
        ],
        "author": "m4ttheux",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:81944",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "WIP some perf optimizations",
        "text": "This pull request bundles two query-execution changes: - `count_distinct_optimization` is enabled by default. The rewrite is applied only by the analyzer's `CountDistinctPass`, which skips `Nullable` / `LowCardinality(Nullable)` arguments and refuses to fire for remote storages or with `distributed_group_by_no_merge`. The legacy AST-level `RewriteCountDistinctFunctionMatcher` was removed: it never fired (its `table_expr->size() != 1` guard uses the recursive `IAST::size`, which is always `>= 2` for a real table expression) and it lacked all of those guards. - A new opt-in `query_plan_rewrite_order_by_limit` optimization that rewrites `ORDER BY ... LIMIT` over wide `MergeTree` reads into a row-offset form. It is **off by default**: it is not yet neutral for the `RuntimeDataflowStatisticsInputBytes` read-bytes estimation (the estimate diverges from `ReadCompressedBytes` by ~24x on wide `SELECT * ... ORDER BY ... LIMIT` reads). `RewriteOrderByLimitPass` rejects `FINAL`, `QUALIFY`, window functions, `WITH FILL`, `arrayJoin`, and non-deterministic `ORDER BY` expressions such as `rand`. The `ABStringRef` string-key aggregation method that this branch previously carried has been dropped: master merged https://github.com/ClickHouse/ClickHouse/pull/93271, which upstreams the same idea as `PackedStringRef` and makes `AggregationMethodPackedString` the default `key_string` / `key_string_two_level` method. Keeping the branch's variant would mean maintaining a duplicate hash method plus a `base/base/StringRef.h` compatibility shim that existed only to keep it compiling after master replaced `StringRef` with `std::string_view`. The trivial `GROUP BY ... LIMIT` optimization and the top-N threshold pushdown into `MergeTree` reading are no longer part of this branch either; the leftover, callerless scaffolding for the latter (`TopNFilterParameters`, `SortColumnDescription::column_name_in_storage`, `FilterTransform::updateQueryConditionHash`, the `getPrewhereInfo` accessors, and the `getFilterMask` rework in `PartialSortingTransform`) has been removed as well. Related: https://github.com/ClickHouse/ClickHouse/pull/93271 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Enable `count_distinct_optimization` by default for eligible local-table `countDistinct` / `uniqExact` queries, and add the opt-in `query_plan_rewrite_order_by_limit` optimization for wide `MergeTree` reads. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/81944",
        "createdAt": "2025-06-16T14:32:02Z",
        "updatedAt": "2026-08-13T12:56:58Z",
        "timestamp": "2026-08-13T12:56:58Z",
        "metrics": {
          "reactions": 4,
          "comments": 35
        },
        "labels": [
          "pr-performance"
        ],
        "author": "amosbird",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:83505",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add early short-circuit evaluation for OR/AND in the analyzer to prevent unnecessary scalar subquery execution",
        "text": "<!--- Disable AI PR formatting assistant: true --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add early short-circuit evaluation for logical OR/AND expressions in the new analyzer. When resolving OR/AND functions, if any argument is a decisive constant (truthy for OR, falsy for AND), the entire expression is replaced with the constant value before argument resolution, preventing scalar subqueries in non-decisive branches from being executed. Resolves #83017 ### Details This PR adds early short-circuit evaluation for logical OR/AND expressions in `QueryAnalyzer::resolveFunction()`. When resolving OR/AND functions, the analyzer now checks if any argument is a decisive constant: - For `OR`: if any argument is a truthy constant (1, true, \"true\"), replace the entire expression with 1 - For `AND`: if any argument is a falsy constant (0, false, \"false\"), replace the entire expression with 0 This optimization happens **before** argument resolution, preventing scalar subqueries in non-decisive branches from being executed at all. Resolves #83017",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/83505",
        "createdAt": "2025-07-09T04:57:56Z",
        "updatedAt": "2026-08-13T03:31:23Z",
        "timestamp": "2026-08-13T03:31:23Z",
        "metrics": {
          "reactions": 0,
          "comments": 12
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "fhw12345",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:86353",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Cascades cost-based optimizer for distributed query plans",
        "text": "A Cascades-style cost-based optimizer that chooses distribution strategies for the multi-stage distributed query plans of #106020. It explores alternatives in a memo (a shared store of equivalent plan fragments) with top-down, goal-directed search and picks the cheapest plan satisfying the required distribution and sorting properties, inserting exchange operators (plan steps that move rows between nodes) as needed. Implemented: - **Join strategies**: shuffle hash join, broadcast hash join (with `ReplicatedRead` — every worker repeats the same read of a small table instead of a network broadcast, assuming shared storage where all workers see the same data), replicated join (a small deterministic join is recomputed identically on every node over such reads, so its result never crosses the network; nested joins compose; `ANY` joins are excluded because the kept row depends on the build order), local join. - **Aggregation strategies**: two-phase (partial + merge), shuffle by group keys, local; `distributed_aggregation_memory_efficient` and `distributed_plan_force_shuffle_aggregation` are honored. - **Top-N**: two-stage distributed top-N (per-node bounded sort, sorted-merge gather, coordinator limit); disabled under `exact_rows_before_limit`, which needs the full row count. - **Read strategies**: parallel N-way read, replicated read, local read. For `FINAL`, #108148 (already in master) taught the rule-based distributed plan to split a `FINAL` read into disjoint primary-key-range buckets where that is safe; Cascades now reuses that machinery, so `FINAL` no longer forces a serial read here either. The coordinator ships each bucket's marks in the `read_bucket` task parameters. - **`IN (subquery)`**: follows the `rewrite_in_to_join` setting like the rest of the planner (the forced join form is removed). In the default set form the set-building subqueries are planned separately and distributed like any other query. - **Properties and enforcers**: distribution (node count, replication, partitioning columns with equivalence classes and the types the keys are cast to before hashing) and sorting; when a plan alternative lacks a required property, an enforcer inserts the step that provides it (`Gather`/`Shuffle`/`Broadcast`/`ScatterExchange`, `Sort`). - **Transformations**: join commutativity (only for semantics-preserving joins: `INNER ALL`, `CROSS`, `SEMI`/`ANY`/`ANTI`; never `ASOF`, and never `ANY` under `join_any_take_last_row`), two-phase aggregation split, two-stage top-N split. - **Cost model**: `work`, `network`, and `sequential` components, each priced as wall-clock per node: a shuffle moves 1/N of the data per node, a broadcast payload is ingested once by every receiver in parallel, and a gather funnels every row through one endpoint, so its transfer stays undivided and pays a per-row cost. A hash-table build counts as parallel work (`parallel_hash` shards it across threads); a fixed per-exchange overhead keeps small inputs local. A table read is priced on its scan volume - the rows the primary key keeps - not on its output estimate, so a filter off the sorting key cannot make a replicated re-read look free. Standalone filters (e.g. `HAVING`) are estimated from column NDVs with join-key equivalence classes; join estimates are clamped to join kind and strictness semantics; exchange costs use per-row byte widths measured from the parts' column sizes (followed through renames, not derived from types). All weights and calibration constants are overridable at query time. What this improves over the rule-based distributed planner, on TPC-H plans. The rule-based planner broadcasts a small table when its read is below `distributed_plan_max_rows_to_broadcast`, but it often cannot size the result of a join, so a small join result (`nation x region`, 5 rows after the region filter) is scattered across nodes, joined there, and shuffled again (repartitioned across nodes) by the next join key. It also often inserts a shuffle at join and aggregation boundaries even when the rows are already divided by the right key. Cascades estimates sizes through joins, knows which partitioning already holds, compares broadcast against shuffle by cost for each join, and recomputes a small deterministic join on every node when that is cheaper than moving its result. On TPC-H SF100 over 8 nodes (same binary, same run window; times are server-side means of the hot runs) the join-heavy queries improve: | Query | Rule-based -> Cascades | What changed in the plan | |---|---|---| | Q21 | 7.66 s -> 5.34 s | The four `supplier x nation` joins compute their small results once and broadcast them, so `lineitem` is not shuffled to meet them. | | Q09 | 4.07 s -> 2.47 s | `part` and `nation` are read in full by every node, so `lineitem` and `supplier` are not shuffled to meet them. | | Q08 | 2.39 s -> 0.99 s | `nation x region` (5 rows) is recomputed by every node; `part` is read in full per node, so `lineitem` is not shuffled to join it. | | Q02 | 2.14 s -> 0.84 s | The `supplier x nation x region` chain is recomputed by every node, so only `partsupp` and `supplier` are shuffled. The top-100 sort becomes two-stage, sending at most 100 rows per node. | | Q05 | 2.05 s -> 1.12 s | The whole dimension side (`orders x customer x nation x region`) is recomputed by every node over full local reads; `lineitem` joins it in place with no shuffle at all. | | Q17 | 2.72 s -> 1.78 s | The small per-part average is broadcast to every node, so the outer 600M-row `lineitem` read is not shuffled. | | Q11 | 0.49 s -> 0.26 s | `supplier x nation` is computed once and broadcast; `partsupp` joins it in place. | | Q12 | 0.76 s -> 0.59 s | The filtered `lineitem` rows (~30K of 600M) are gathered once and broadcast; nothing is shuffled. | The remaining queries change less. Summed over all 22 queries, hot server time drops from 36.8 s to 28.8 s (about 22% lower). The remaining regressions are the `IN (subquery)` queries Q18 (2.43 s -> 2.85 s) and Q20 (2.24 s -> 2.99 s). With the forced join rewrite removed, both run their `IN`s as sets in both modes, and the sets built are identical; the difference is in the plans around them. Cascades picks replicated-read shapes that rescan moderate tables on every node, which loses here for a structural reason: the main reads are filtered by `x IN <set>`, and the set contents do not exist at costing time, so the model cannot credit a shuffle-based plan for how few rows survive the set filter, while the replicated read's full scan is paid regardless. Two follow-ups: a cost-informed choice of the `IN` form, and set-filter selectivity from the subquery's output estimate. `EXPLAIN pretty = 1, estimates = 1` shows the chosen plan with a row estimate and the accumulated cost for each step. Also in this PR, two improvements to the shared bucketed-read machinery (they benefit the rule-based path too): the `FINAL` layer split no longer depends on the coordinator's core count, and a many-partition `FINAL` split groups its layers into the target task count instead of falling back to a serial read. Design, a worked example on TPC-H data (a simplified 3-table query traced through the memo), and current limitations are documented in `src/Processors/QueryPlan/Optimizations/Cascades/ARCHITECTURE.md`. Plan-shape tests cover the actual TPC-H queries (`03836_tpch_join_order_plans`), and focused tests pin the cost-model contracts (e.g. `04869_cascades_read_cost_granule_volume`, `04838_cascades_filter_selectivity`). Disabled by default. Requires the analyzer; remote execution requires the stateless-worker configuration, while `distributed_plan_execute_locally = 1` runs the stages in-process without it: ```sql SET enable_cascades_optimizer = 1, make_distributed_plan = 1; ``` For tests, `param__internal_cascades_cluster_node_count` overrides the cluster size, `param__internal_cascades_cost_config` overrides the cost model configuration, `param__internal_join_table_stat_hints` injects table statistics, `param__internal_cascades_task_limit` lowers the task budget (it can never raise it). Related: https://github.com/ClickHouse/ClickHouse/pull/106020 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental Cascades cost-based optimizer for distributed query plans, enabled by `enable_cascades_optimizer = 1` together with `make_distributed_plan = 1`. It chooses between shuffle, broadcast, replicated, and local join strategies, two-phase, shuffle, and local aggregation, two-stage distributed top-N, and parallel and replicated reads by estimated cost, inserting exchange operators as needed. ### Documentation entry for user-facing changes - [ ] Documentation written in [/docs](https://github.com/ClickHouse/ClickHouse/tree/master/docs)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/86353",
        "createdAt": "2025-08-28T11:27:09Z",
        "updatedAt": "2026-08-13T17:45:48Z",
        "timestamp": "2026-08-13T17:45:48Z",
        "metrics": {
          "reactions": 22,
          "comments": 7
        },
        "labels": [
          "pr-experimental"
        ],
        "author": "davenger",
        "state": "open",
        "assignees": [
          "novikd"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:86768",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Feature: Enable overlay databases for server.",
        "text": "Enables server-side `Overlay` databases. An `Overlay` database is a read-only facade that exposes the union of the tables of several underlying databases, resolving each table name through the listed sources in order (the first source that has the table wins). DDL on the facade is rejected — it has no storage of its own — while `SELECT` and pass-through `INSERT` resolve to the underlying source table. Reading or writing through the facade requires the corresponding grant on both the facade database and the underlying source, and the facade's row policies are combined with the source's. Previously an `Overlay` database existed only as the implicit default database of `clickhouse-local`. Closes: https://github.com/ClickHouse/ClickHouse/issues/52764 <!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow creating `Overlay` databases. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/86768",
        "createdAt": "2025-09-05T21:38:13Z",
        "updatedAt": "2026-08-13T16:33:47Z",
        "timestamp": "2026-08-13T16:33:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 108
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested",
          "pr-autogenerated-docs"
        ],
        "author": "AlyHKafoury",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:88234",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add Rewrite rules",
        "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental support for defining custom query rewrite rules. Closes: https://github.com/ClickHouse/ClickHouse/issues/80084 ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds a new query-rewriting/rejection pipeline step in `executeQuery` (gated by `query_rules`) plus persistent rule storage (local or ZooKeeper/Keeper), so mistakes could affect query correctness when enabled and introduce new background reload behavior. > > **Overview** > Adds **Query Rewrite Rules**: new `CREATE RULE`/`ALTER RULE`/`DROP RULE` statements that can rewrite matched queries or reject them with a message. > > Rules are persisted via a new `RewriteRules` subsystem with configurable storage (`query_rules_storage` on disk or ZooKeeper/Keeper, including reload/watch support), exposed through `system.query_rules` and `system.query_rules_log`, and enforced via new access grants (`CREATE_RULE`, `ALTER_RULE`, `DROP_RULE`). > > Query execution now optionally applies these rules by traversing the parsed AST before normal processing (enabled by the new `query_rules` setting), with new error codes and coverage via added stateless tests and documentation. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit fd232f5e9d1b6946b568a8402bba477ee7728c8a. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/88234",
        "createdAt": "2025-10-08T11:03:24Z",
        "updatedAt": "2026-08-13T13:08:12Z",
        "timestamp": "2026-08-13T13:08:12Z",
        "metrics": {
          "reactions": 1,
          "comments": 61
        },
        "labels": [
          "manual approve",
          "can be tested",
          "pr-experimental"
        ],
        "author": "hinata34",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:89350",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `secondary_indices_materialized` column to `system.parts`",
        "text": "### Changelog category: - New Feature ### Changelog entry [Users can now validate which secondary indices are fully materialized in a data part via the new “secondary_indices_materialized” column in the “system.parts” table.] ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) # secondary_indices_materialized Within the “/src/Storages/System/StorageSystemParts.cpp” code file, added a new column named “secondary_indices_materialized” to the “system.parts” system table to resolve issue #88576. This column provides an array of strings containing the specific names of all materialized secondary indices for each data part which was previously not possible if a given part had multiple secondary indices. The new column definition and its associated array output calculation are both placed in the correct order with respect to the existing column definition and calculation pairs. **Syntax** {\"secondary_indices_materialized\", std::make_shared<DataTypeArray>(std::make_shared<DataTypeString>()), \"The array of materialized secondary index names\"}, ... if (columns_mask[src_index++]) { Array materialized_indices; auto secondary_indices_descriptions = part->storage.getInMemoryMetadataPtr()->secondary_indices; for (const auto & index_description : secondary_indices_descriptions) { if (part->hasSecondaryIndex(index_description.name)) materialized_indices.push_back(index_description.name); } columns[res_index++]->insert(materialized_indices); } **Arguments** Not applicable (no function added) **Implementation Details** The “secondary_indices_materialized” column is first defined; If run by a user query, the code retrieves all secondary index definitions for all existing parts from the overall metadata table using “part->storage.getInMemoryMetadataPtr()->secondary_indices”, then iterates through each of these secondary index definitions and validates if it is actually materialized in the current part, with “part->hasSecondaryIndex(index_description.name)”. This way, index names are only added to the resulting array if they exist for the specified part, before the result array is inserted into the output column. **Returned value** - Returns {columns[res_index++]->insert(materialized_indices);}. [Array](../data-types/float.md) **Example** Query: \\```sql SELECT 'name, secondary_indices_materialized FROM system.parts WHERE database = ‘ex_database' AND table = ‘ex_table’; \\``` Response: \\```response ┌──────name─────────────secondary_indices_materialized───────────┐ │ [part_number] │ [‘indexStr1’, ‘indexStr2’, ‘…’] └─────────────────────────────────────────────── ┘ \\``` --- **Additional fix (2026-07-27)** While adding the regression test for the `escape_index_filenames` semantics of the new column, CI (under the randomized MergeTree setting `auto_statistics_types`) exposed an unrelated pre-existing bug, fixed here as well: In the settings-alter branch of `StorageMergeTree::alter` and `StorageReplicatedMergeTree::alter`, `changeSettings` is the sole writer of the setting-derived fields `StorageInMemoryMetadata::escape_index_filenames` and `IndexDescription::escape_filenames`. When the table carries an explicit `auto_statistics_types` setting, the subsequent `setInMemoryMetadata` call that re-commits the implicit statistics writes back a metadata copy that was taken *before* `changeSettings` (`commands.apply` never sets those fields), so `ALTER TABLE ... MODIFY SETTING escape_index_filenames = ...` did not take effect until the server was restarted. The committed escape fields are now carried into that copy, exactly as the `areNonReplicatedAlterCommands` branch of `StorageReplicatedMergeTree::alter` already does for the comment commit. New test: `04650_alter_escape_index_filenames_with_auto_statistics`. Closes: https://github.com/ClickHouse/ClickHouse/issues/88576",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/89350",
        "timestamp": "2026-08-12T22:14:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "pr-feature",
          "manual approve",
          "can be tested"
        ],
        "author": "aryan32agrawal",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:89945",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add a Chinese tokenizer (jieba) for the tokens function and text indexes",
        "text": "<!--- Disable AI PR formatting assistant: false --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a `chinese` tokenizer for the `tokens` function and `MergeTree` text indexes. It segments Chinese text into words using a dictionary and a Hidden Markov Model (the algorithm follows [jieba](https://github.com/fxsjy/jieba)), with `coarse_grained` (default) and `fine_grained` granularities. Continues #80174. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) ### Implementation notes - The tokenizer is a from-scratch C++ reimplementation following the fxsjy/jieba algorithm (dictionary-based maximum-probability segmentation with an HMM fallback). The embedded dictionary and HMM model are generated from a pinned [cppjieba](https://github.com/yanyiwu/cppjieba) commit (MIT) and verified by SHA-256; regenerator scripts are included. - The engine lives in a standalone repository, https://github.com/ClickHouse/jieba_cpp, consumed here as the `contrib/jieba_cpp` submodule with a `contrib/jieba_cpp-cmake` wrapper that builds it against ClickHouse's in-tree abseil/darts-clone/zstd. Built by default (`ENABLE_CHINESE_TOKENIZER`); the dictionary ships in little- and big-endian variants and is `#embed`-ed per host byte order, so it works on big-endian targets too. - `hasToken` now bypasses the text index for any non-`splitByNonAlpha` tokenizer (use `hasAnyTokens` / `hasAllTokens` for tokenizer-aware matching). This also fixes a pre-existing case where `hasToken` over `ngrams`/`array`/`splitByString`/`asciiCJK` indexes could return wrong rows.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/89945",
        "createdAt": "2025-11-12T16:27:31Z",
        "updatedAt": "2026-08-13T17:11:31Z",
        "timestamp": "2026-08-13T17:11:31Z",
        "metrics": {
          "reactions": 3,
          "comments": 36
        },
        "labels": [
          "pr-feature",
          "submodule changed",
          "can be tested",
          "hold"
        ],
        "author": "amosbird",
        "state": "open",
        "assignees": [
          "Ergus"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:91062",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add an experimental regex-free glob parser",
        "text": "Introduces `GlobAST`, a regex-free glob parser, and wires it into the file and object-storage listing paths. It is **off by default** (the experimental `use_glob_ast_parser` setting), so `master` behavior is unchanged — the goal is to land the parser and all the wiring now so that enabling it later is a one-setting switch. The legacy engine builds a regex (`makeRegexpPatternFromGlobs` + `re2::RE2`), which has noticeable downsides: - lists whole prefixes for enum globs (#73333) - blows up to megabyte regexes on numeric ranges (#43456) - has no formal grammar (#80950) - diverges from POSIX shell semantics in places (e.g. stripping single-element brace groups like `{a}`/`{-}`). `GlobAST` parses a pattern once into typed expressions (constant, `?`/`*`/`**`, `{M..N}` range, `{a,b,…}` enum) and matches/expands directly — ranges by numeric bounds checks, enum-only globs by expanding to concrete keys. A `GlobMatcher` front-end selects the new or legacy backend per the setting, so call sites are unchanged; the grammar is documented in `parseGlobs.h`. Tested by unit tests, a differential fuzzer against the legacy matcher (`GlobASTLegacyMatchFuzz`), and a stateless parity test. Related: https://github.com/ClickHouse/ClickHouse/issues/73333 Related: https://github.com/ClickHouse/ClickHouse/issues/43456 Related: https://github.com/ClickHouse/ClickHouse/issues/80950 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental setting `use_glob_ast_parser` (default `false`) that switches glob matching for the `file`/`s3`/object-storage listing paths to a new regex-free parser (`GlobAST`), avoiding regex blow-up on large numeric ranges and listing fewer keys for brace-enumeration globs.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/91062",
        "createdAt": "2025-11-27T22:07:46Z",
        "updatedAt": "2026-08-13T14:52:19Z",
        "timestamp": "2026-08-13T14:52:19Z",
        "metrics": {
          "reactions": 1,
          "comments": 14
        },
        "labels": [
          "pr-experimental",
          "hold"
        ],
        "author": "thevar1able",
        "state": "open",
        "assignees": [
          "alexey-milovidov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:91358",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Support binary for aggregate_function_input_format, unify implementation",
        "text": "Addresses issues in #88088 Support the `RowBinary` family of input formats for `aggregate_function_input_format = 'value'` and `aggregate_function_input_format = 'array'`, and unify the value/array deserialization across all input formats (text and binary) behind a single implementation that deserializes into a temporary column of the argument type (or `Array` of it) and builds the `AggregateFunction` state via `addBatchSinglePlace`. The setting `aggregate_function_input_format` remains production-tier: it was already released as such in `v25.12` and `v26.1`, so this PR does not change its feature tier and is not a backward-incompatible change. Related: https://github.com/ClickHouse/ClickHouse/pull/88088 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `aggregate_function_input_format = 'value'` and `aggregate_function_input_format = 'array'` for binary input formats such as `RowBinary` when inserting into `AggregateFunction` columns, and unify the deserialization implementation across input formats. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/91358",
        "timestamp": "2026-08-12T22:14:24Z",
        "metrics": {
          "reactions": 1,
          "comments": 28
        },
        "labels": [
          "pr-improvement"
        ],
        "author": "GrigoryPervakov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:93114",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "WIP: Projection Index Text",
        "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): TODO ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/93114",
        "createdAt": "2025-12-29T00:00:29Z",
        "updatedAt": "2026-08-13T17:12:47Z",
        "timestamp": "2026-08-13T17:12:47Z",
        "metrics": {
          "reactions": 7,
          "comments": 25
        },
        "labels": [
          "pr-improvement",
          "submodule changed"
        ],
        "author": "amosbird",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:93833",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "TimeSeries: enable dynamic remote-write routing via URL table prefix",
        "text": "This PR addresses https://github.com/ClickHouse/ClickHouse/issues/93831 ### Changelog category (leave one): - Improvement ### Changelog entry Support dynamic table routing for Prometheus `remote-write` based on the request URL, allowing a single handler to ingest data into multiple TimeSeries tables while preserving backward compatibility with fixed-table configuration. ### Documentation entry for user-facing changes - [x] Documentation is written #### Motivation In large-scale observability deployments, users often maintain many TimeSeries tables with similar schemas but different retention or ownership. The existing Prometheus `remote-write` integration requires one handler per table, with the table name hard-coded in `config.xml`, making configuration and operations difficult to scale. This change allows routing the target TimeSeries table dynamically from the request URL, enabling a single `remote-write` handler to serve multiple tables while keeping the previous behavior unchanged by default. #### Parameters - **`enable_table_name_url_routing`** (handler option, default: `false`) When enabled, the handler expects URLs in the form `/{database}/{table}/...` and extracts the target table from the first two path segments. - **`prometheus_remote_write_dynamic_routing_enabled`** (TimeSeries table setting, default: `false`) Enables dynamic table routing for the specified TimeSeries table. Requests targeting tables without this setting enabled will be rejected. #### Example usage ```xml <handlers> <remote_write> <url>regex:^/[^/]+/[^/]+/write$</url> <handler> <type>remote_write</type> <enable_table_name_url_routing>true</enable_table_name_url_routing> </handler> </remote_write> </handlers> ``` Request example: `POST /{db_name}/{table_name}/write` Enable dynamic routing on the target TimeSeries table: ```sql ALTER TABLE db.ts MODIFY SETTING prometheus_remote_write_dynamic_routing_enabled = 1; ``` <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds request-path-driven table selection to the Prometheus `remote_write` handler and gates writes with a new per-table setting, so misconfiguration or unexpected URLs could redirect writes or cause new 4xx/5xx failures. > > **Overview** > Enables Prometheus `remote_write` to **dynamically route writes to different `TimeSeries` tables** by extracting `{database}/{table}` from the request URL when `enable_table_name_url_routing` is set on the handler. > > Dynamic routing is **opt-in and additionally gated per target table** via new `TimeSeries` setting `prometheus_remote_write_dynamic_routing_enabled` (default `false`); requests to tables without it error out. Docs and integration tests were updated with the required URL regex and validation coverage for both enabled/disabled cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit d0e40205e663a54acd9edfa31186b6790c8f5443. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/93833",
        "createdAt": "2026-01-09T16:10:59Z",
        "updatedAt": "2026-08-13T00:37:35Z",
        "timestamp": "2026-08-13T00:37:35Z",
        "metrics": {
          "reactions": 2,
          "comments": 14
        },
        "labels": [
          "pr-improvement",
          "manual approve",
          "can be tested",
          "comp-promql"
        ],
        "author": "niyue",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:94148",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Introducing a new fuzzer check in the ClickHouse CI for my dear friend",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ---- New buzzhouse job <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds a new multi-node, docker-in-docker fuzzer job to CI and refactors shared fuzzer log analysis/parsing, which could affect how failures are detected and reported across fuzz/stress pipelines. Changes also touch integration test cluster process tracking (`exec_id` handling), so misclassification of failures or missed crashes is the main risk. > > **Overview** > **CI now runs a new fuzzer check, `La Casa Del Dolor`, in both `master` and `pull_request` workflows**, replacing the previous `BuzzHouse` job names/keys and wiring it into `finish_workflow` dependencies and job definitions. > > **Adds `ci/jobs/lacasadeldolor_job.py`** to execute the fuzzer via `tests/casa_del_dolor/dolor.py` in the integration-tests runner (docker-in-docker), collect multi-node logs/config artifacts, and reuse a new shared `analyze_job_logs()` routine. > > **Refactors failure detection/reporting**: `ast_fuzzer_job.py` extracts log analysis into `analyze_job_logs()` (including expanded accepted exit codes and improved OOM/sanitizer handling), and `FuzzerLogParser` is updated to search across multiple (including `.gz`) server/stderr logs and prefer the matched log when extracting stack traces/failed queries; `stress_job.py` and `clickhouse_proc.py` are updated to the new parser API. > > **Updates Casa del Dolor test harness and cluster helpers** to support the new CI mode: fixes imports under `tests.integration`, adjusts generator config/tempfile handling and exit-code validation, tightens disk/policy XML generation, and tracks ClickHouse container `exec_id` in `ClickHouseCluster/Instance` for more reliable shutdown/exit-code inspection. > > <sup>Written by [Cursor Bugbot](https://cursor.com/dashboard?tab=bugbot) for commit 82bd7ef20780c0d5b5bc648ace847418429573cb. This will update automatically on new commits. Configure [here](https://cursor.com/dashboard?tab=bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/94148",
        "createdAt": "2026-01-14T11:15:28Z",
        "updatedAt": "2026-08-13T17:37:11Z",
        "timestamp": "2026-08-13T17:37:11Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [
          "pr-not-for-changelog"
        ],
        "author": "maxknv",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:94748",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix invalid result of joining two `-Cluster` table functions",
        "text": "Assisted-by: Claude Sonnet 4.5 via GitHub Copilot ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix invalid result on joining multiple table expressions, when leftmost table expression is a `-Cluster` table function. Resolves https://github.com/ClickHouse/ClickHouse/issues/89996 <!-- ch-version-info:start --> ### Version info - Merged into: `26.1.1.899` <!-- ch-version-info:end -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/94748",
        "createdAt": "2026-01-21T18:10:36Z",
        "updatedAt": "2026-08-13T11:31:33Z",
        "timestamp": "2026-08-13T11:31:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "pr-bugfix",
          "pr-backports-created",
          "pr-synced-to-cloud",
          "pr-must-backport-synced",
          "v25.8-must-backport"
        ],
        "author": "thevar1able",
        "state": "closed",
        "assignees": [
          "KochetovNicolai"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:96130",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Randomize tests with DETACH/ATTACH table before query execution",
        "text": "Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that enable randomly detaching and reattaching tables used in a query before its execution. This is a testing-only feature designed to find bugs related to table reattachment. Before executing a query, the system collects all tables referenced in the AST, and for each eligible table (stores data on disk, supports detaching, has no action locks or dependencies), it performs a `DETACH` followed by `ATTACH`. Changes: - Add `supportsDetachingTables` virtual method to `IDatabase` (overridden to `false` for engines that do not support non-permanent `DETACH TABLE`: `DatabaseDictionary`, `DatabaseReplicated`, `DatabaseSQLite`, `DatabaseBackup`, `DatabaseFilesystem`, `DatabaseHDFS`, `DatabaseS3`, `DatabaseURL`, `DatabaseRemote`, `DatabaseDataLake`, `DatabaseMaterializedPostgreSQL`) - Add `has`/`hasAny` methods to `ActionLocksManager` for checking existing locks (skipping expired `weak_ptr` entries) - Add table collection visitor and reattach logic in `executeQuery` (runs after AST validations, process list admission, and external tables initialization; skips `EXPLAIN`, transactions, internal/non-initial queries, and CTE name collisions) - Fix off-by-one in `MergeTreeDeduplicationLog::dropOutdatedLogs` (don't drop the active log) and add `sync` call in shutdown - Add `no-random-detach` tag to tests incompatible with this feature - Add `--no-random-detach` and `--reattach-tables-probability` options to `clickhouse-test` - Add `02461_reattach_tables` test Continuation of #55943. Continuation of #42336 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that randomly `DETACH` and `ATTACH` tables used in a query before its execution. This is a testing-only feature that helps find reattachment-related bugs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Introduces new pre-execution mutations (internal `DETACH`/`ATTACH`) in `executeQuery`, which can affect table availability and concurrency behavior if enabled; guarded by new experimental settings but touches core query execution paths. > > **Overview** > Adds experimental settings `reattach_tables_before_query_execution` and `..._probability` to optionally **DETACH and ATTACH back** eligible tables referenced by a query immediately before execution, including AST table discovery that accounts for CTE scoping, privilege checks, dependency/lock checks, and safety skips (e.g. `system`, non-disk storages, dynamic-structure columns, transactions, `EXPLAIN`, internal/non-initial queries). > > Extends `IDatabase` with `supportsDetachingTables()` and marks multiple database engines as not supporting non-permanent detach; adds `ActionLocksManager::has/hasAny` helpers to avoid detaching tables with active action locks. Updates the test runner and stress tooling to randomize this behavior (with `--no-random-detach` and probability control), adds a new `02461_reattach_tables` test, and tags many existing tests to opt out where DETACH/ATTACH would add flakiness/overhead. Also fixes `MergeTreeDeduplicationLog` cleanup to avoid dropping the active log and ensures writer `sync()` on shutdown. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit a369371ff61cb1934815ea1e8caf7debd83c98c4. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/96130",
        "createdAt": "2026-02-05T23:27:47Z",
        "updatedAt": "2026-08-13T17:49:48Z",
        "timestamp": "2026-08-13T17:49:48Z",
        "metrics": {
          "reactions": 2,
          "comments": 91
        },
        "labels": [
          "pr-build"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:96487",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix LazilyReadFromMergeTree optimization with ALIAS columns (#96452)",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry: Fix missing lazy read and top-K read optimizations (skip-index top-K and dynamic top-K filtering) when selecting `ALIAS` columns with `WHERE ... ORDER BY ... LIMIT` on MergeTree tables. Closes: https://github.com/ClickHouse/ClickHouse/issues/96452 ### Documentation entry for user-facing changes N/A",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/96487",
        "createdAt": "2026-02-09T21:21:08Z",
        "updatedAt": "2026-08-13T17:45:34Z",
        "timestamp": "2026-08-13T17:45:34Z",
        "metrics": {
          "reactions": 1,
          "comments": 31
        },
        "labels": [
          "pr-bugfix",
          "can be tested"
        ],
        "author": "jayvenn21",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:96494",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix hang in `--queries-file` / `-q` when stdin is a pipe with no data",
        "text": "`ClientBase` eagerly called `isStdinNotEmptyAndValid` to check for trailing stdin data on every `INSERT`. That helper calls `!std_in.eof()`, which performs a blocking `read(0, ...)`. When stdin is a pipe that is open but has no data and no `EOF` (common in CI, Docker, stress tests, and `--queries-file` invocations), the read blocks indefinitely and the client hangs until it is killed by timeout. Concretely, this caused an exception (signal 11 after timeout) in the stress test `03538_analyzer_scalar_correlated_subquery_fix`: the process hung on `read(0, ...)` and was killed. The fix replaces the eager blocking probe with a non-blocking one: `poll(timeout=0)` first asks \"does stdin appear to have data right now?\", and only when it does we fall through to `isStdinNotEmptyAndValid` to confirm and consume it. This: - avoids the hang on open empty pipes — `poll` returns 0 immediately; - preserves deterministic rejection of `inlined data + stdin` mixed input — when stdin actually contains data (the case asserted by `04064_inline_insert_data`), `poll` returns `POLLIN` and we still throw `NOT_IMPLEMENTED`; - is best-effort by design: a slow producer pipe (e.g. `(sleep 1; echo) | clickhouse-client ...`) may not have written by the time `poll` runs, in which case the trailing stdin payload is ignored. This is preferable to the alternative of a hard hang, and is documented in the helper. Covered cases (`is_async_insert_with_inlined_data`, `is_inline_insert_data`, and the appending-stdin path in `sendData`) all go through the same probe so the hang is closed for `--queries-file`, `-q`, and `--inline-insert-data` modes. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b9e68f4b9b0b33c7db43b00afb3eff4ff2050694&name_0=MasterCI&name_1=Stress%20test%20%28amd_debug%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a hang in `clickhouse-client` for `INSERT` queries in `--queries-file`/`-q`/`--inline-insert-data` mode when stdin is an open pipe with no data and no `EOF` (common in CI, Docker, and stress tests). ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/96494",
        "timestamp": "2026-08-12T22:49:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 18
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:96630",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Optimization of `GROUP BY` in the presence of `ORDER BY` and `LIMIT`",
        "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Optimize `GROUP BY ... ORDER BY ... LIMIT` and `GROUP BY ... LIMIT` queries by maintaining a bounded heap during aggregation to prune groups that cannot appear in the result, significantly reducing memory usage and execution time for high-cardinality grouping. Controlled by a new experimental `enable_group_by_top_k_optimization` setting. Based on the initial implementation by [Dmitriy Terenichev](https://github.com/Dmitry909) ([#78553](https://github.com/ClickHouse/ClickHouse/pull/78553)). Related: https://github.com/ClickHouse/ClickHouse/issues/72610 Query shapes that benefit from this optimization: - `SELECT key, agg() FROM t GROUP BY key ORDER BY key ASC LIMIT N` -- single key, full ORDER BY match - `SELECT key, agg() FROM t GROUP BY key LIMIT N` -- single key, no ORDER BY - `SELECT k1, k2, agg() FROM t GROUP BY k1, k2 ORDER BY k1, k2 ASC LIMIT N` -- composite keys, full ORDER BY match - `SELECT k1, k2, k3, agg() FROM t GROUP BY k1, k2, k3 ORDER BY k1, k2, k3 ASC LIMIT N` -- 3+ keys, full match - `SELECT k1, k2, agg() FROM t GROUP BY k1, k2 ORDER BY k1 ASC LIMIT N` -- composite keys, ORDER BY is a prefix of GROUP BY - `SELECT k1, k2, k3, agg() FROM t GROUP BY k1, k2, k3 ORDER BY k1 ASC LIMIT N` -- prefix 1 of 3 - `SELECT k1, k2, k3, agg() FROM t GROUP BY k1, k2, k3 ORDER BY k1, k2 ASC LIMIT N` -- prefix 2 of 3 - `SELECT k1, k2, agg() FROM t GROUP BY k1, k2 LIMIT N` -- composite keys, no ORDER BY - `SELECT URL, count() FROM t GROUP BY URL ORDER BY URL ASC LIMIT N` -- String keys - `SELECT URL, count() FROM t GROUP BY URL LIMIT N` -- String keys, no ORDER BY Not supported (optimization is skipped): - `ORDER BY agg()` (ORDER BY aggregate function, not a GROUP BY key) (handled by [another PR](https://github.com/ClickHouse/ClickHouse/pull/98607)) - `WITH TOTALS`, `HAVING`, `GROUPING SETS`, `CUBE`, `ROLLUP` - `LIMIT WITH TIES` - ~~Non-final aggregation (distributed partial aggregation)~~ - `max_rows_to_group_by` is set - `rows_before_limit_at_least` mode (exact count) - queries having collators in ORDER BY - ORDER BY columns that are not a leading prefix of GROUP BY keys",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/96630",
        "timestamp": "2026-08-12T20:07:24Z",
        "metrics": {
          "reactions": 8,
          "comments": 13
        },
        "labels": [
          "pr-performance",
          "hold"
        ],
        "author": "thevar1able",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:96844",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Columns Cache",
        "text": "Implement the deserialized columns cache that keeps columns in memory. In comparison to the uncompressed cache, this saves time on deserialization and copying. Closes #82335. The cache identifies entries by table UUID, so it is only active for tables in databases that assign UUIDs (`Atomic` and `Replicated`); tables in `Ordinary` databases are silently excluded. The feature is gated behind the `EXPERIMENTAL` settings `use_columns_cache`, `enable_reads_from_columns_cache`, and `enable_writes_to_columns_cache`. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Implement an experimental columns cache that keeps deserialized columns in memory for `MergeTree` tables. In comparison to the uncompressed cache, this saves time on deserialization and copying. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/96844",
        "createdAt": "2026-02-13T21:00:45Z",
        "updatedAt": "2026-08-13T03:49:57Z",
        "timestamp": "2026-08-13T03:49:57Z",
        "metrics": {
          "reactions": 5,
          "comments": 77
        },
        "labels": [
          "pr-performance",
          "pr-experimental"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:96978",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Fix JSON/XML format statistics race condition with parallel replicas",
        "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `rows_read` and `rows_before_limit_at_least` reported as 0 or stale in JSON/XML format output when a query with `LIMIT` reads from remote connections (parallel replicas or the `remote` table function). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) No documentation changes needed — this is a bug fix with no user-facing API changes. --- ## Summary - When parallel replicas are used with `LIMIT`, `rows_read` in JSON/XML output is reported as 0 - Root cause: race condition where `finalizeImpl` writes statistics BEFORE `PipelineExecutor::finalizeExecution` collects remaining progress from connection draining - Fix: two-phase output format finalization — statistics are written AFTER all progress has been collected - The same drain also delivers late `ProfileInfo` packets that update the `rows_before_limit_at_least` / `rows_before_aggregation` counters, so the deferral covers the whole trailer, and the native-protocol `ProfileInfo` sent to the client is refreshed from the post-drain counters ## Approach Split the output format's finalization so statistics are written AFTER `finalizeExecution` collects all remaining progress: - **Phase 1** (during pipeline, in `finalizeImpl`): Write everything EXCEPT the trailer after the `rows` field (`rows_before_limit_at_least`, `rows_before_aggregation`, the `\"statistics\"` section) and closing delimiter - **Phase 2** (after `finalizeExecution`): Re-read the rows-before-* counters, write the trailer and close the document via a finalize callback ### Key changes: - `IOutputFormat`: Added `hasDeferredStatistics`, `writeDeferredStatisticsAndFinalize`, `completeDeferredStatistics` mechanism; `snapshotRowsBeforeCounters` re-reads the shared counters before the deferred trailer is written - `PipelineExecutor`: Added `setFinalizeCallback` called at end of `finalizeExecution` after progress collection - `CompletedPipelineExecutor`: Sets the finalize callback to invoke `completeDeferredStatistics` - `JSONRowOutputFormat`, `XMLRowOutputFormat`, `JSONColumnsWithMetadataBlockOutputFormat`: Override deferred statistics methods; the whole trailer (`rows_before_limit_at_least`, `rows_before_aggregation`, `statistics`) is deferred to phase 2 (`JSONUtils::writeRowsBeforeAndStatistics` factors out the shared writer) - `ParallelFormattingOutputFormat`: Propagates deferred statistics through the parallel formatting pipeline - `LazyOutputFormat` / `PullingOutputFormat`: `getProfileInfo` re-snapshots the rows-before-* counters, so the `ProfileInfo` packet sent to native-protocol clients (TCP `clickhouse-client`, gRPC, `LocalConnection`) carries the post-drain values — by the time it is sent, the executor thread has been joined and the counters are final - New failpoint `tcp_handler_sleep_before_secondary_query_trailing_packets` delays a secondary query's trailing `Totals` / `Extremes` / `ProfileInfo` / `Progress` / `EndOfStream`, making the race deterministic for tests Closes https://github.com/ClickHouse/ClickHouse/issues/85785 ## Test plan - [x] Build succeeds (Release) - [x] Test `00365_statistics_in_formats` passes - [x] JSON format tests pass (`00159_parallel_formatting_json_and_friends_1/2`, `00378_json_quote_64bit_integers`, `00685_output_format_json_escape_forward_slashes`, `01447_json_strings`, `01449_json_compact_strings`, `01486_json_array_output`, `02554_format_json_columns_for_empty`) - [x] XML format tests pass (`00307_format_xml`, `02122_parallel_formatting_XML`) - [x] `03918_statistics_in_formats_parallel_replicas` — forward guard: parallel-replicas `rows_read` matches the single-node value - [x] New deterministic reproducer `04603_rows_before_limit_parallel_replicas_late_packets` — fails without the fix (failpoint forces the trailing packets into the drain window); checks `rows_read` under parallel replicas and `rows_read` + `rows_before_limit_at_least` through the `remote` table function, over HTTP (server-side formatting) and native TCP (client-side formatting), across `JSON`, `JSONColumnsWithMetadata`, and `XML` - [ ] CI: Run with parallel replicas to verify `rows_read` is no longer 0 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches query pipeline execution, cancellation, and output formatting/finalization paths; while scoped to statistics/progress reporting, it changes timing and thread-synchronization and could affect query completion/cancellation behavior under parallel replicas. > > **Overview** > Fixes a race where JSON/XML output could report `rows_read=0` with parallel replicas + `LIMIT` by ensuring final `Progress` packets (drained after cancellation/early completion) are incorporated before statistics are emitted. > > This introduces **two-phase output finalization**: `IOutputFormat` can defer writing statistics/closing delimiters (`hasDeferredStatistics`/`writeDeferredStatisticsAndFinalize`), and `PipelineExecutor` now supports a `setFinalizeCallback` invoked at the end of `finalizeExecution()` after collecting remaining progress; `CompletedPipelineExecutor` wires this to `IOutputFormat::completeDeferredStatistics()`. > > Cancellation/draining logic is tightened to preserve trailing progress and avoid blocking hard cancels: `PipelineExecutor` triggers a `PartialResult` cancel on normal completion, `ISource::cancel` skips `onCancel` for `PartialResult` but allows later escalation, `RemoteSource` adds explicit cancel-reason upgrades and aborts drain via `RemoteQueryExecutor::abortDrain`, and `RemoteQueryExecutor::finish()` drains until all replica connections complete while swallowing expected cancel exceptions. Adds a stateless regression test `03918_statistics_in_formats_parallel_replicas` to assert `rows_read > 0` across JSON/XML formats over TCP/HTTP. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 5818266286a7645dcf1b57212bf9b3214cfd2ec8. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/96978",
        "createdAt": "2026-02-15T07:04:09Z",
        "updatedAt": "2026-08-13T09:34:49Z",
        "timestamp": "2026-08-13T09:34:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 53
        },
        "labels": [
          "pr-bugfix"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:97032",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Implement Prometheus /api/v1/series, /labels, /label/values endpoints",
        "text": "Implement the three remaining Prometheus HTTP API metadata endpoints for the `TimeSeries` engine, which previously threw `NOT_IMPLEMENTED`. These endpoints are required by Grafana's Prometheus data source for label autocomplete and the metric browser. - `/api/v1/series`: returns all series with their full label sets by querying the tags table, filtered by the `match[]` series selectors. Each row is normalized to its logical Prometheus label set (a `tags_to_columns` tag can live in the dedicated column or in the residual `tags` `Map` of a supported external table) with the same rules as `timeSeriesStoreTags`: both carriers are merged, exact duplicates collapse, empty values are dropped, and conflicting carrier values are rejected with `bad_data`, so deduplication is by series identity rather than by the raw table layout. - `/api/v1/labels`: returns all distinct label names from the tags `Map` column and from the `tags_to_columns` columns, always including `__name__` as a virtual label. An empty label value means the label is absent, so a map entry with an empty value (possible in a supported external `tags` table, e.g. `tags = {'env': ''}`) does not surface its key as a label name and does not count toward `limit`, consistently with the other endpoints. - `/api/v1/label/<name>/values`: returns the distinct values for a specific label, reading the `metric_name` column for `__name__`, the dedicated `tags_to_columns` column (falling back to the residual `tags` `Map`) for a configured tag, or the tags `Map` otherwise. Both of these endpoints validate the carriers of the labels they materialize the same way `/api/v1/series` and the query endpoints do, so a row holding different non-empty values for one tag in its dedicated column and in the residual `tags` `Map` is rejected with `bad_data` instead of being silently reported. The optional `limit` parameter is supported on all the endpoints, and a present-but-empty `limit=`, `start=`, or `end=` is rejected with `bad_data` rather than being treated as an omitted parameter. The limit is enforced in the emission layer rather than as a SQL `LIMIT`, so the carrier validation above sees every matched row and a corrupted row cannot be hidden by a small `limit` and a favorable row order. The optional `start`/`end` parameters filter by the `min_time`/`max_time` columns of the `tags` table and are accepted only when both `store_min_time_and_max_time` and `filter_by_min_time_and_max_time` are enabled - the same gate as the query path's tags-table prefilter. When either setting is disabled (including an external `tags` table that still physically has the bound columns while `store_min_time_and_max_time = 0`), a ranged metadata request is rejected instead of silently diverging from `/api/v1/query` and `/api/v1/query_range`. The `match[]` parameter accepts full Prometheus series selectors (a bare metric name or an instant selector with `=`, `!=`, `=~`, `!~` label matchers, e.g. `cpu_usage{host=\"server1\"}`), parsed with the same PromQL parser as `/api/v1/query` and translated into the same `tags`-table filter the query endpoints use, so all three metadata endpoints select exactly the series the query endpoints would read. A repeated `match[]` is the union of the selectors, as in Prometheus. A matcher on a `tags_to_columns` tag resolves the tag with the same carrier normalization as the endpoints themselves: the dedicated column wins when non-empty (NULL is normalized to the empty label value), the residual `tags` `Map` is used otherwise, and a row carrying different non-empty values in the two carriers is rejected with `bad_data` instead of silently preferring one of them. Regexp matchers are fully anchored (`^(?:...)$`), and non-legacy (UTF-8) label names are written as quoted string literals, e.g. `{\"http.status_code\"=\"200\"}`, both as in Prometheus. A selector whose matchers all match the empty label value (e.g. `{host=~\".*\"}`) is rejected with `bad_data`, as in Prometheus - at least one matcher must not match the empty string, so a `match[]` cannot degenerate into an unfiltered scan - and an invalid regexp anywhere in a selector is likewise rejected with `bad_data` at parse time. Label names and values are JSON-escaped on output and string literals interpolated into the generated SQL are quoted, so values containing quotes, backslashes, or control characters are handled correctly. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Implemented the Prometheus `/api/v1/series`, `/api/v1/labels`, and `/api/v1/label/<name>/values` metadata endpoints for the `TimeSeries` engine. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/97032",
        "createdAt": "2026-02-15T21:08:25Z",
        "updatedAt": "2026-08-13T10:47:46Z",
        "timestamp": "2026-08-13T10:47:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 49
        },
        "labels": [
          "pr-feature",
          "submodule changed",
          "manual approve",
          "can be tested",
          "comp-promql",
          "pr-autogenerated-docs"
        ],
        "author": "ajonkisz",
        "state": "open",
        "assignees": [
          "alexey-milovidov",
          "nikitamikhaylov"
        ]
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:97254",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Push join key filters into MergeTree index during recursive CTE evaluation",
        "text": "When a recursive CTE joins against a `MergeTree` table, each recursion step previously scanned the entire table, because the join condition (e.g. `ON e.from_id = t.current_id`) was not pushed into `MergeTree`'s key condition. This change analyzes the recursive query's join tree to find equi-join conditions between the CTE working table and real tables. Before each step it reads the join-key values from the working table and injects them as an `IN (...)` predicate directly into the `WHERE` clause of each recursive `QueryNode` that joins against the CTE (the original `WHERE`/`HAVING`/`QUALIFY` clauses are snapshotted at construction and restored after each step, so nothing accumulates across steps). The planner then lowers this predicate into `ReadFromMergeTree`'s key condition. The injected set is bounded by the new setting `recursive_cte_max_in_filter_cardinality`, and injection fails closed to a plain scan whenever it could otherwise change results — when the generated set would exceed `max_rows_in_set`/`max_bytes_in_set`, cannot be materialized under `max_memory_usage`, or the `IN` predicate cannot be resolved for the join-key type. Parallel replicas are disabled for the recursive step queries to avoid stale cached `GLOBAL JOIN` tables returning incorrect results; the forcing mode `allow_experimental_parallel_reading_from_replicas = 2` raises `SUPPORT_IS_DISABLED` for the recursive part rather than silently downgrading. For a table with ~1M rows and 10 recursion steps, `read_rows` drops from ~10M to ~120. Closes: https://github.com/ClickHouse/ClickHouse/issues/75026 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `WITH RECURSIVE` queries that use `ON` or comma equi-joins against `MergeTree` tables can now use the primary key index, reducing the number of rows read per recursion step. The optimization is controlled by the new setting `recursive_cte_max_in_filter_cardinality`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) No separate documentation page is needed — this is a transparent performance improvement. The only user-visible addition is the setting `recursive_cte_max_in_filter_cardinality`, which is documented at the source level in `src/Core/Settings.cpp`. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches recursive CTE execution and query settings, and dynamically injects `additional_table_filters`, which could affect correctness/performance for some JOIN patterns despite being scoped to recursive steps. > > **Overview** > Improves recursive CTE performance by detecting equi-join conditions between the CTE working table and real tables, then **pushing join-key values into `additional_table_filters`** (built as `IN (...)` predicates) on each recursive step so MergeTree reads can use the primary-key index; user-provided `additional_table_filters` are merged rather than overwritten. > > Also **disables parallel replicas** for recursive CTE step queries to avoid incorrect results from reused cached GLOBAL JOIN tables, and adds a stateless test (`03924_recursive_cte_join_index`) asserting both correct output and low `read_rows` for explicit `INNER JOIN` and comma-join forms. > > <sup>Written by [Cursor Bugbot](https://cursor.com/dashboard?tab=bugbot) for commit 37f26180bfa2189247b6862712551c7ecc41ed18. This will update automatically on new commits. Configure [here](https://cursor.com/dashboard?tab=bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/97254",
        "createdAt": "2026-02-18T07:38:36Z",
        "updatedAt": "2026-08-13T01:56:02Z",
        "timestamp": "2026-08-13T01:56:02Z",
        "metrics": {
          "reactions": 1,
          "comments": 77
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:97540",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Validate IN tuple/subquery column count mismatch in analyzer",
        "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/74442 The analyzer (`resolveFunction`) now validates that the number of elements on the left side of `IN` matches the number of columns in the right-side subquery. Previously, constant folding could optimize away the `IN` expression and silently hide the mismatch: for example, `(1, 1) IN (SELECT 1)` wrapped in a subquery with a constant-false outer `WHERE` succeeded silently. Now it correctly throws `NUMBER_OF_COLUMNS_DOESNT_MATCH` during analysis. A single `Tuple` column on the right side is still accepted, since `FunctionIn` compares the whole left-side value against it; `LowCardinality` and `Nullable` wrappers are unwrapped before the check. Covered by the new test `03380_in_tuple_subquery_column_count`, including the original reproducer from the issue. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Detect column count mismatch between a tuple and a subquery in the `IN` clause during analysis, preventing constant folding from silently hiding the error. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/97540",
        "createdAt": "2026-02-21T00:39:53Z",
        "updatedAt": "2026-08-13T08:11:20Z",
        "timestamp": "2026-08-13T08:11:20Z",
        "metrics": {
          "reactions": 0,
          "comments": 34
        },
        "labels": [
          "pr-bugfix",
          "submodule changed"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:98498",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Revert dangerous change in CreateUniqueArrayJoinAliasesVisitor",
        "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Revert the dangerous change from https://github.com/ClickHouse/ClickHouse/pull/98376, which breaks the invariant of the query tree. ColumnNode is supposed to always have a valid source. cc @alexey-milovidov ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/98498",
        "createdAt": "2026-03-02T14:17:32Z",
        "updatedAt": "2026-08-13T14:23:26Z",
        "timestamp": "2026-08-13T14:23:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "pr-not-for-changelog",
          "close in a month if not active"
        ],
        "author": "novikd",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:99495",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Add `GradualResizeProcessor` to limit effective parallelism for GROUP BY on small data volumes",
        "text": "When ClickHouse processes GROUP BY, it often overestimates the number of threads needed. With `max_threads = 64` but only a few thousand rows, all 64 `AggregatingTransform` instances get data, produce 64 partial hash tables, and the merge phase has to combine all of them — most nearly empty. This wastes time on merging overhead, which is especially noticeable for heavy aggregate states such as `uniq`, `uniqExact`, `groupArray`, etc. The new `GradualResizeProcessor` starts by pushing data to a single output port (or one port per split group when `min_outstreams_per_resize_after_split` applies), and activates all aggregation streams at once as soon as the configured row or byte threshold is crossed. For small datasets, only one aggregating thread receives data (or one per split group); for large datasets, all threads are used as before. New settings: - `min_rows_per_stream_for_gradual_resize` (default: `1000`) - `min_bytes_per_stream_for_gradual_resize` (default: `0`) When either threshold is non-zero, the pre-aggregation `StrictResize` is replaced with `GradualResize` in the pipeline. The optimization is enabled by default; set both `min_rows_per_stream_for_gradual_resize = 0` and `min_bytes_per_stream_for_gradual_resize = 0` to opt out. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve performance of GROUP BY on small data volumes.",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/99495",
        "createdAt": "2026-03-14T07:35:33Z",
        "updatedAt": "2026-08-13T17:57:39Z",
        "timestamp": "2026-08-13T17:57:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 30
        },
        "labels": [
          "pr-performance"
        ],
        "author": "alexey-milovidov",
        "state": "open",
        "assignees": [
          "nihalzp"
        ],
        "change": "updated"
      },
      {
        "id": "github:ClickHouse/ClickHouse:pull_request:99981",
        "source": "github",
        "group": "data-infrastructure",
        "project": "ClickHouse/ClickHouse",
        "kind": "pull_request",
        "title": "Iceberg: propagate table UUID from REST catalog to avoid metadata cac…",
        "text": "REST catalog inline responses already contain the table UUID and metadata location. Propagate the UUID through DataLakeSpecificProperties -> StorageObjectStorageConfiguration.catalog_uuid_hint -> initializePersistentTableComponents so the metadata cache is checked before fetching metadata.json from network storage. ### Performance mechanism **Before:** On first table initialization, `getMetadataJSONObject` is called without a known UUID. The cache probe is skipped (`table_uuid == nullopt`), so metadata.json is always fetched from remote storage as a cold read. Even if the same table is initialized again later, the cache was never populated on the first pass. **After:** The REST catalog provides the table UUID upfront via `catalog_uuid_hint`. `getMetadataJSONObject` probes the cache with `uuid:path` key before doing any remote I/O. If the metadata was previously cached (e.g., by a prior query that retroactively populated it), the cold read is eliminated entirely. If not, the cold read still happens but the result is cached with the correct UUID for future lookups. **Impact:** Deterministic reduction in remote `metadata.json` fetches for REST catalog tables. Each avoided fetch saves one round-trip to S3/GCS/ABFS plus deserialization overhead. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve Iceberg catalog metadata caching by using table UUID from catalog responses to warm the metadata cache, avoiding a redundant remote metadata.json read on table initialization. ### Documentation entry for user-facing changes Improve Iceberg catalog metadata caching. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Medium risk because it changes Iceberg metadata cache key composition and the metadata-loading path to use a new UUID hint, which could affect cache hit rates or correctness if UUIDs/paths are inconsistent. > > **Overview** > **Improves Iceberg REST catalog metadata caching** by propagating the table UUID from REST inline responses through `DataLakeSpecificProperties` into `StorageObjectStorageConfiguration` as `catalog_uuid_hint`, so the first metadata fetch can probe the metadata cache before doing extra remote IO. > > Updates metadata loading to accept an optional known UUID, capture raw metadata JSON for reuse, and retroactively populate the cache once the real UUID is discovered. Also fixes potential cache-key collisions by changing `IcebergMetadataFilesCache::getKey` to use a `uuid:path` delimiter, and adds gtest coverage for key uniqueness and cache hit/miss behavior. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 2e7400b8dc007641c702d41b3fd6cafdc10eb823. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
        "url": "https://github.com/ClickHouse/ClickHouse/pull/99981",
        "createdAt": "2026-03-18T21:44:06Z",
        "updatedAt": "2026-08-13T11:30:37Z",
        "timestamp": "2026-08-13T11:30:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 29
        },
        "labels": [
          "pr-performance",
          "can be tested"
        ],
        "author": "bacek",
        "state": "open",
        "assignees": [
          "SmitaRKulkarni"
        ]
      }
    ],
    "events": [
      {
        "id": "event:cd9a054b0ad2513119c3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114558",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114558",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support PromQL @ start() and @ end() modifiers",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): PromQL: support `@ start()` and `@ end()` timestamp modifiers. ### Summary PromQL allows `start()` and `end()` as special values for the `@` modifier. They mean the fixed start and end of a range query. For an instant query, both resolve to the evaluation time. Without this support, valid PromQL such as `http_requests_total @ start()` and `rate(http_requests_total[5m] @ end())` cannot be used against ClickHouse. Fixed-`@` subqueries keep their complete inner time grid and remain step-invariant across the outer range query. Standalone `start()` and `end()` functions are not included. ### Why this matters This is standard PromQL surface used by major Prometheus-compatible systems: - VictoriaMetrics has executable range-query tests for both `time() @ start()` and `time() @ end()`. - Grafana Mimir, Thanos, and Cortex replace these modifiers with fixed timestamps before splitting range queries, including subqueries. - Elasticsearch includes `@ start()`, `@ end()`, and subquery forms in its valid PromQL grammar fixtures. These are public implementation and compatibility references, not a claim about private customer query volume. Supporting the syntax keeps dashboards and recording/alerting queries portable across Prometheus-compatible backends. References: - [Prometheus Querying basics: `@` modifier](https://prometheus.io/docs/prometheus/latest/querying/basics/) - [VictoriaMetrics executable tests](https://github.com/VictoriaMetrics/VictoriaMetrics/blob/master/app/vmselect/promql/exec_test.go#L1207-L1227) - [Grafana Mimir query splitting](https://github.com/grafana/mimir/blob/main/pkg/frontend/querymiddleware/split_and_cache.go#L2838-L2948) - [Thanos query splitting](https://github.com/thanos-io/thanos/blob/main/internal/cortex/querier/queryrange/split_by_interval.go#L593-L690) - [Cortex query splitting](https://github.com/cortexproject/cortex/blob/master/pkg/querier/tripperware/queryrange/split_by_interval.go#L1061-L1157) - [Elasticsearch valid PromQL grammar fixtures](https://github.com/elastic/elasticsearch/blob/main/x-pack/plugin/esql/qa/testFixtures/src/main/resources/promql/grammar/queries-valid.promql#L1112-L1122) ### Changes - Add `start()` and `end()` to the PromQL timestamp grammar. - Preserve the symbolic modifier in the Prometheus query tree. - Resolve the symbols against the outer query boundaries during evaluation. - Keep offset and range-selector handling consistent with numeric `@` timestamps. - Preserve the complete inner grid for fixed-`@` subqueries. ### Tests - Added parser coverage for both symbols, modifier order, literals, and subqueries. - Added instant and range query coverage, including a range selector under `@ end()`. - Added `last_over_time()` coverage for symbolic and numeric `@` on subqueries. - Compared the new range-query cases with Prometheus through the existing integration helpers. - Regenerated the ANTLR parser artifacts.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114558",
          "createdAt": "2026-08-12T23:36:48Z",
          "updatedAt": "2026-08-13T13:47:41Z",
          "timestamp": "2026-08-13T13:47:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-improvement",
            "submodule changed"
          ],
          "author": "fallintoplace",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6524a35eb3cf9c8bc4f1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114531",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114531",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject a lossy codec on columns backing keys and indexes",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114406 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A lossy codec such as `SZ3` is now rejected at DDL time on any column that backs the sorting key, primary key, partition key, a secondary index or the unique key. Because such a codec does not return the value that was written, a merged part could be stored out of physical order and index analysis would skip rows matching the query. Existing tables stay loadable. ### Description `ORDER BY` makes a physical promise: rows are stored inside a part in sorting-key order and the primary index samples those stored values. A lossy codec breaks `read(write(v)) == v`, so the merge sorts pre-compression values while the part stores post-compression ones. `SZ3` is not monotonic, so the stored sequence is not sorted. The same applies to anything computed from the pre-write block: skip-index granules, the unique-key index, partition values. On master the issue's reproducer gives `descending_pairs = 1`, `min(i - prev) = -0.2436889648437699` after `OPTIMIZE TABLE t FINAL` and 0 after a single `INSERT`, so the disorder appears only at the merge, where a debug build aborts in `CheckSortedTransform`. Lossiness comes from the existing `ICompressionCodec::isLossyCompression()`, not a codec-name list, so any codec declaring itself lossy is covered. `CompressionCodecMultiple` did not override it, so a stacked `CODEC(SZ3(...), LZ4)` reported itself lossless; it now ORs over its children. Classification is per serialized substream using that substream's own type, mirroring `MergeTreeDataPartWriterWide::addStreams`, so `ORDER BY arr.size0` stays allowed while `ORDER BY arraySum(arr)` is refused. The check sits at the two user-facing entry points, `registerStorageMergeTree` for CREATE and full-definition ATTACH and `checkAlterIsPossible` for ALTER on the initiating execution, following `4a29ef847411256`. A replica re-executing an already durable DDL entry is not re-checked, which would wedge its DDL worker, and ALTER validates only the pairs it introduces, so unrelated ALTERs on such a table still work. Seen in CI on one AST-fuzzer run, on #113575, neighbours `Float64_12.415000000000006` and `Float64_12.168837890625007`: [report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113575&sha=84df30db17165b65e3c78bc11513ed26534187d7&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%29). <details> <summary>Related cases deliberately left out of this PR</summary> - A **projection** over a lossily compressed column answers differently than the base part: `sum(i)` is `249957173.79` read from the base part and `249998750` via the projection. The mechanism is a separate post-write consumer, which writes the base part lossily and then computes the projection from the unchanged pre-compression block, so it follows in its own PR rather than being folded in here. - A legacy table can still gain an implicit minmax index over such a column through server configuration on the load path. Reachable, but a boundary sweep over 300 values produced no wrong result. - `ORDER BY length(arr)` is now rejected although its value is exact. A key expression records which columns it needs and not which of their streams, so at this layer `length(arr)` and `arraySum(arr)` are indistinguishable, and `arraySum` is a genuine wrong-results carrier. The syntactic form `ORDER BY arr.size0` stays allowed because it names the stream. - Alias-dependent index rebuilds look incorrect independently of codecs: an alias-only `MODIFY COLUMN` requires no mutation while the index expression is rebuilt, so existing parts keep a same-named index built from the old expression. No claim is made about it here. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114531",
          "createdAt": "2026-08-12T18:16:56Z",
          "updatedAt": "2026-08-13T13:47:05Z",
          "timestamp": "2026-08-13T13:47:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8a4d31ebbec4343de592",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113983",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113983",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allowlist the expected FileLog bad-path reattach error in the upgrade check",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113781 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Upgrade check (amd_release)` intermittently fails its `Error message in clickhouse-server.log` sub-test on one benign line: ``` <Error> StorageFileLog (test_1.filelog_bad_path_attach): The absolute data path should be inside `user_files_path`(/var/lib/clickhouse/user_files/) ``` No product defect: the server starts, nothing crashes, no data is affected. Root cause. `04202_filelog_attach_path_outside_user_files` ATTACHes a FileLog table whose path is outside `user_files_path`. `ATTACH` is `LoadingStrictnessLevel::ATTACH` (2), which is `>= SECONDARY_CREATE` (1), so the constructor takes the relaxed branch at `src/Storages/FileLog/StorageFileLog.cpp:195-198`: it logs at `<Error>` and returns instead of throwing `BAD_ARGUMENTS`. That branch is deliberate and is what the test covers, since refusing to load at reattach time would break server startup. The table then outlives the test: stress threads run with a fixed `--database=test_N` (`ci/jobs/scripts/stress/stress.py`), and `clickhouse-test` skips its per-test teardown whenever `--database` is set (`need_cleanup = not args.database`), so that shared database is never dropped. The upgrade restart re-attaches the table, the relaxed branch fires again, and the line lands in the scanned log, where the post-restart scrub in `tests/docker_scripts/upgrade_runner.sh` had no entry for it. Hence the intermittency: `04202` must land on a fixed-database thread. Change. One `grep -av` entry in that scrub's existing secondary pipe, plus a short rationale comment next to the sibling entries. The pattern requires the fixture table name and the message together, and (bare parens are literals in BRE) the `StorageFileLog (db.table):` prefix shape. No source change, no test change. Validation. The scan pipeline, extracted verbatim from the runner, was run under GNU grep 3.11 against the failing run's own 19.9 MB `clickhouse-server.upgrade.log`. With the entry the artifact is empty; with it deleted the output is byte-identical to the 189-byte `upgrade_error_messages.txt` CI produced, so the sub-test flips `FAIL` to `OK`. Eight negative controls still surface, including a table whose name merely ends with the fixture name (`prod.other_filelog_bad_path_attach`), which the required `.` separator keeps visible.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113983",
          "createdAt": "2026-08-08T21:17:46Z",
          "updatedAt": "2026-08-13T13:46:16Z",
          "timestamp": "2026-08-13T13:46:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:55f2cae93fa69817720c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113022",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113022",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a type-aware Bloom filter index for JSON",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 The original design used `JSONAllValues` as the input to a Bloom filter. `JSONAllValues` serializes each value as text. It does not preserve the runtime type. This loss of type information is important for `Dynamic` values. ClickHouse can compare JSON values with different runtime types. Some type pairs can match after conversion. Other type pairs can return an exception. A Bloom filter that stores only text cannot safely model these rules. It can skip a granule that contains a match. It can also hide an exception. `JSONAllValues` remains useful for text search, but it is not a safe base for this index. This PR replaces that design with `jsonbf_v1`. The new index creates tokens from these components: - JSON path - Container role - Runtime type - Binary value The container role separates scalar values, array elements, and map values. The index also processes nested JSON objects and named tuple fields. For `Dynamic` values, the index stores type-presence and complex-value presence tokens. Query analysis uses exact value tokens only when the comparison is safe. If analysis cannot prove safety, ClickHouse reads the granule. Container roles are preserved through nested tuples, and nested casts are handled conservatively. The index supports: - Equality and typed `IN` conditions - `has`, `hasAny`, and `hasAll` for arrays - Typed map values by key - Nested JSON objects and tuples The index does not optimize range conditions or whole-container equality. It also rejects unsafe comparison paths, such as Decimal-to-Float comparisons. An unsupported `Dynamic` runtime type disables skipping for its granule. In a one-million-row JSONBench test, the index reduced reads from 123 granules to 12–15 granules. Selected string equality queries were 1.6–2.3 times faster. Indexed inserts were approximately 2.9 times slower in the checked-in performance test. A local performance test produced these median results: | Operation | Without index | With `jsonbf_v1` | Difference | |---|---:|---:|---:| | Numeric equality, matching value | 13.05 ms | 10.99 ms | 1.19 times faster | | Numeric equality, missing value | 13.25 ms | 10.80 ms | 1.23 times faster | | String equality | 24.49 ms | 12.90 ms | 1.90 times faster | | Array `has` | 24.71 ms | 12.78 ms | 1.93 times faster | | Insert | 20.51 s | 40.87 s | 1.99 times slower | The small performance test shows limited benefit for numeric scalar equality and larger improvements for string equality and array membership. The larger JSONBench data set benefits more because the index skips more granules. Token generation and Bloom-filter construction increase insert time. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds the `jsonbf_v1` data-skipping index for type-aware equality and array membership on `JSON` values.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113022",
          "createdAt": "2026-08-02T18:31:45Z",
          "updatedAt": "2026-08-13T13:46:07Z",
          "timestamp": "2026-08-13T13:46:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a0a0ca40b04ffbe25647",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114536",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114536",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add ClickStack guide for isolating read and write workloads",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/113284 ### Changelog category (leave one): - Documentation (changelog entry is not required) --- ClickStack's docs claim the ability to \"independently isolate read and write workloads with Warehouses\" in five places without documenting how. The only existing guidance (in `deployment/managed`, added in ClickHouse/clickhouse-docs#5452) covers one half of it: that the ClickStack UI binds to whichever Cloud service it is launched from. This adds a guide for the full topology and links it from the places that previously mentioned the capability without explaining it. **New page** — `clickstack/managing/isolating-read-write.mdx`: - Why isolate: read-only services run no background merges, so their compute is dedicated to queries; ingestion is insulated from expensive queries; each side scales independently against the ingest and query figures from the sizing model. - Recommended topology: one read-write service for ingestion, one read-only service for ClickStack, plus the planning constraints — the first service is always read-write, type is fixed at creation, and merge assignment crosses read-write services. - Setup steps: prepare the read-write service and ingestion user, add the read-only service, point the collector at the read-write endpoint, point the UI at read-only compute (managed and self-hosted), then verify the split with `system.query_log` — including the `all_groups.default` note, since `system` tables are per-service. - Where DDL and materialized views execute, and how alert evaluation follows the connection of the source it is attached to. - Advanced: separating merges from ingestion onto a dedicated merge service, flagged as requiring a support request, with the caveats that come with it (mutation tracking, TTL deletion, auto-idling, keeping queries off both read-write services). **Cross-links**: `managing/overview` (admin guides table), `managing/production` (new subsection), `managing/estimating-resources` (ties the ingest and query vCPU split to separate services), `deployment/managed` (extends the existing read-only-compute section), and the previously unlinked bullets in `overview` and `architecture`. `getting-started/managed.mdx` carries the same bullet but is deliberately left untouched — #112118 rewrites that section and already links the concept, so editing it here would only create a conflict.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114536",
          "createdAt": "2026-08-12T19:46:18Z",
          "updatedAt": "2026-08-13T13:45:59Z",
          "timestamp": "2026-08-13T13:45:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "andremm",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6fc55f1916f24a4b5be2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114178",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114178",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"NATS: add inline credentials setting\"",
          "text": "Reverts https://github.com/ClickHouse/ClickHouse/pull/110733 (merge commit `48ebef16af656664bab51cd5bb605a65e9b8ec9e`), which added the `nats_credentials` setting to the `NATS` table engine. Removed by this pull request: * the `nats_credentials` setting and its use in `NATSConnection` (`natsOptions_SetUserCredentialsFromMemory`), so the engine again takes credentials only as a path through `nats_credential_file`; * the mutual-exclusion check between `nats_credential_file` and `nats_credentials`, including the named-collection provenance handling in `registerStorageNATS`; * the runtime handling of `nats_credentials` only — the logging-only masking of the old spelling is deliberately kept. `nats_credentials` stays in both mask lists (`NATS::SETTINGS_TO_HIDE` and `FunctionSecretArgumentsFinder::nats_secret_keys`), because a query is formatted for logging before the storage settings are validated, so the old spelling must not leak the raw JWT/seed into `query_log` even though the server then rejects it. The masking branch (`findNATSTableEngineSecretArguments`) is kept as well: it also masks the still-supported secret keys (`nats_password`, `nats_token`, `nats_credential_file`) and the `nats_url` userinfo password in the table-engine argument form `ENGINE = NATS(collection, key = ...)`. The `ParserCreateQuery.MaskNATS*` unit tests are kept, and a new `MaskNATSTableEngineRemovedCredentialsSetting` test covers both spellings of the removed setting; * the test `04665_nats_credentials_named_collection`; * the `nats_credentials` lines from the `Documentation` block of `registerStorageNATS`. Two notes on the mechanics of the revert: * `src/Parsers/FunctionSecretArgumentsFinder.{h,cpp}` are resolved by hand: `findNATSTableEngineSecretArguments` and the full `nats_secret_keys` list (including `nats_credentials`, for masking only) are kept. Every other file comes out either byte-identical to its pre-`#110733` state or equal to it plus unrelated later changes (the message-broker schedule pool now returns a `shared_ptr`). * `docs/reference/engines/table-engines/integrations/nats.mdx` still mentions `nats_credentials`. That page is generated from the `Documentation` block this pull request updates, and direct edits of an `{/*AUTOGENERATED_START*/}` region are rejected by the docs check, so the page is left to the nightly documentation autogeneration. The setting was merged into `master` for the unreleased `26.8` and is not part of any release branch (the newest is `26.7`), so no changelog entry is needed. Related: https://github.com/ClickHouse/ClickHouse/pull/110733 Related: https://github.com/ClickHouse/ClickHouse/issues/85213 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114178",
          "createdAt": "2026-08-10T15:32:57Z",
          "updatedAt": "2026-08-13T13:45:30Z",
          "timestamp": "2026-08-13T13:45:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:256fab3e805edd7b6ee2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114539",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114539",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enable the query condition cache for `ORDER BY ... LIMIT n` queries by default",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/111492 Related: https://github.com/ClickHouse/ClickHouse/pull/104478 Related: https://github.com/ClickHouse/ClickHouse/pull/110507 `use_query_condition_cache_for_top_k` was introduced by #111492 and defaulted to `false` as a precaution while the soundness of the query condition cache entries written by a TopK (`ORDER BY <column> LIMIT n`) read was being established. Such entries are partitioned by the TopK plan parameters and by a snapshot of the set of parts read, so only a query with the same plan over the same parts reuses them. The gate is no longer needed, so the default becomes `true` and `ORDER BY ... LIMIT n` queries use the query condition cache again. This effectively reverts #111492 by flipping its setting, rather than by removing it: the setting and every gating point it drives are kept, so the cache can still be kept out of TopK reads with `use_query_condition_cache_for_top_k = 0`. The tests that cover that configuration are kept too — they already pinned the setting explicitly rather than relying on the default — while the tests of the feature itself no longer have to enable it. `04628_query_condition_cache_topk_default_off` is renamed to `04628_query_condition_cache_topk_gate_off` since it no longer describes the default. The settings-history entry keeps `previous_value = false`, because the gate was backported to 26.7. `compatibility` with 26.7 or earlier therefore still turns the query condition cache off for TopK reads, while `compatibility = '26.8'` keeps it on; `04631_query_condition_cache_topk_compatibility` pins this. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The [query condition cache](/operations/query-condition-cache) is now enabled by default for queries that use the `ORDER BY <column> LIMIT n` (TopK) optimization. It can be turned off again with the setting `use_query_condition_cache_for_top_k`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114539",
          "createdAt": "2026-08-12T20:09:36Z",
          "updatedAt": "2026-08-13T13:45:29Z",
          "timestamp": "2026-08-13T13:45:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:761e6643e3439b8ea457",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114278",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114278",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `XDG_CACHE_HOME` being read from `XDG_STATE_HOME`",
          "text": "A copy-paste slip in `XDGBaseDirectories`: `ENV_XDG_CACHE_HOME` was defined as `\"XDG_STATE_HOME\"`, so `XDGBaseDirectories::getCacheHome` ignored the `XDG_CACHE_HOME` environment variable and obeyed `XDG_STATE_HOME` instead. `getCacheHome` currently has no in-tree callers, so this does not change observable behavior yet; it fixes the helper before callers appear. Found while reviewing https://github.com/ClickHouse/ClickHouse/pull/112824. Related: https://github.com/ClickHouse/ClickHouse/pull/112824 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114278",
          "createdAt": "2026-08-11T06:33:24Z",
          "updatedAt": "2026-08-13T13:45:29Z",
          "timestamp": "2026-08-13T13:45:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f1db66fa7f8fa5b380e6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114643",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114643",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-land aggregate function `gini` in the `sum` family",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/113763 Related: https://github.com/ClickHouse/ClickHouse/pull/113868 Related: https://github.com/ClickHouse/ClickHouse/pull/112280 Related: https://github.com/ClickHouse/ClickHouse/pull/114314 --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New aggregate function `gini`, which calculates the [Gini coefficient](https://en.wikipedia.org/wiki/Gini_coefficient) of a column of finite, non-negative numeric values. The result ranges from `0` (all values equal) towards `1` as inequality grows; for a sample of `n` values the maximum is `(n - 1) / n`. `NaN` values are skipped and infinite values are rejected. The function returns `Float64` and consumes `O(n)` memory. ### Description Re-lands the `gini` function that #113868 reverted, following the option-2 spec @ Manerone gave in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5217284116 and confirmed in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5279278069. It re-adds #112280's function with exactly three items removed, and touches no file under `src/AggregateFunctions/Combinators/`: - the `getArgumentsThatCanBeOnlyNull` override, - `.returns_default_when_only_null = true` on registration, - the `argument_type->onlyNull()` branch in the creator. The property was doing the work. It made `AggregateFunctionFactory::getImpl` skip its only-null guard and build a real `gini` instance over `Nullable(Nothing)`, so `gini(NULL)` returned `Float64` `nan` where every `sum`-family function folds to `Nullable(Nothing)`. Without it the fold happens in the `Null` combinator before any other combinator is applied, and the creator's only-null branch becomes unreachable, which is also why `createAggregateFunctionSum` has no equivalent. `gini` now registers exactly like `sum`. Runtime `Nullable` handling is a separate axis, via `getOwnNullAdapter`, and is unchanged: `gini` over `[1, NULL, 3]` still returns `0.25`. Validated on three binaries: post-revert master (`gini` absent), a build carrying #112280's version, and this branch. Every literal-`NULL` cell on this branch equals the `sum` value measured on the same binary, and all non-`NULL` results are unchanged from #112280. The docs artifacts restore what #113868's wholesale revert removed from @ Blargian's already-merged Mintlify port #114314 (the `gini.mdx` page plus its slug-map, navigation and redirect entries), taken verbatim from that commit, so this is a restoration rather than new docs work. Since the restored page carries an `AUTOGENERATED` region, `check_autogenerated_regions.py` rejects it (verified locally, with a control). @ Manerone, could you apply `pr-autogenerated-docs`, as you did on #113868? cc @Manerone",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114643",
          "createdAt": "2026-08-13T13:34:59Z",
          "updatedAt": "2026-08-13T13:46:36Z",
          "timestamp": "2026-08-13T13:46:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "Manerone"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:40ccf5fa38131a5115ac",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114466",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114466",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: internationalize master",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114466",
          "createdAt": "2026-08-12T10:49:59Z",
          "updatedAt": "2026-08-13T13:45:02Z",
          "timestamp": "2026-08-13T13:45:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 41
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "locadex-agent[bot]",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:70972f6f1b3f294f677a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114522",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114522",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "QueryRunner follow-up: do not occupy threads eagerly",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): QueryRunner tables now start worker threads on demand and release them once idle, instead of occupying threads for the table's whole lifetime.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114522",
          "createdAt": "2026-08-12T16:42:37Z",
          "updatedAt": "2026-08-13T13:43:40Z",
          "timestamp": "2026-08-13T13:43:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:892df190daf0347773bb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114283",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114283",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add pre-hook to insert CI links into PR body",
          "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) -- Adds a `ci_links.py` pre-hook to the `PR` workflow that, on upstream `ClickHouse/ClickHouse` pull request runs, appends a `:ci_links:` block to the PR description with: - a link to the workflow report, and - a link to a GitHub search for the corresponding sync PR (`sync-upstream/pr/<number>`). The block is added only when it is not already present, so subsequent runs do not re-edit the PR body. Non-upstream / non-PR runs are skipped, and any failure is caught so it can never break the workflow. <!-- CI automatic block start :ci_links: --> Workflow [[PR](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114283&sha=latest&name_0=PR)] Sync PR [[sync-upstream/pr/114283](https://github.com/search?q=head%3Async-upstream%2Fpr%2F114283+org%3AClickHouse+type%3Apr&type=pullrequests)] <!-- CI automatic block end :ci_links: -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114283",
          "createdAt": "2026-08-11T07:48:23Z",
          "updatedAt": "2026-08-13T13:43:25Z",
          "timestamp": "2026-08-13T13:43:25Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:ed02ba7c1223cea93228",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111852",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111852",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix(Silk): honor O_NONBLOCK in the fiber TLS BIO",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... --- This is a latent-interaction bug that appears only when **three** components are combined — none is wrong on its own: 1. **Silk fiber sockets** (#107680): the fiber-aware OpenSSL BIO `silkBioRead`/`silkBioWrite` submits io_uring I/O and parks the caller for the socket's send/receive timeout. That is correct for a blocking read — but the BIO never consults the fd's `O_NONBLOCK` flag, unlike OpenSSL's default socket BIO, which returns `EAGAIN` immediately when the flag is set. 2. **TLS** (`USE_SSL`): the affected path is the TLS BIO. The plain-socket staleness check uses a raw `recv(MSG_PEEK | MSG_DONTWAIT)` and never touches this code, so the problem is TLS-only. 3. **The connection pool's `SSL_peek` staleness probe** (#110402): before reusing a pooled connection, `DB::getSocketState` flips the fd non-blocking via `ScopedNonBlocking` (a raw `fcntl(F_SETFL, O_NONBLOCK)`, behind Poco's back) and calls `SSL_peek`, expecting an immediate `EAGAIN` — its comment reads *\"The socket is non-blocking, so this never blocks.\"* #110402 introduced this probe, replacing the previous `poll`-based keep-alive disconnect check. Only with all three present does that non-blocking `SSL_peek` route through the silk BIO, which ignores `O_NONBLOCK` and blocks for the receive timeout left on the socket by the previous request. Each component is individually correct; the fix lands on the silk BIO because it is the one whose behavior diverges from OpenSSL's socket-BIO contract (honoring `O_NONBLOCK`), while the TLS layer and the `#110402` probe are behaving as intended. On `master` today the silk BIO has no wired production consumer — it is infrastructure — so this three-way combination is not yet reachable in a shipped server; it was reproduced with downstream work that routes object-storage-disk connections through silk fiber sockets. **Impact.** Every borrow of a pooled TLS connection to an object-storage disk pays a timeout it should not. A server loading tables from an HTTPS object-store disk at startup does many such borrows and stalls — a deterministic ~13.5 s in a local reproduction, and an unbounded boot hang (never reaching \"Ready for connections\", no error logged) with production timeouts or a zero/unset receive timeout, where the wait becomes a deadline-less `future.wait()`. **Root-cause evidence** (local TLS-MinIO reproduction): a server-side request trace showed each request arriving only *after* its wait expired; the ~13.5 s decomposed exactly into the adaptive per-method receive timeouts paid in sequence (GET 500 ms + PUT 3000 ms + DELETE 10000 ms, `ConnectionTimeouts.cpp`); `ss` showed frozen `bytes_sent` through each stall; and a live backtrace was parked at `silkBioRead` ← `SSL_peek` ← `getSslSocketState` ← `isStale` ← `getConnection`. A control with silk sockets disabled does the same step in <15 ms. The plain-HTTP staleness probe uses a raw `recv(MSG_PEEK|MSG_DONTWAIT)` and is unaffected — the bug is TLS-specific. **Fix.** When the fd is non-blocking, `silkBioRead`/`silkBioWrite` do a direct `recv`/`send` with `MSG_DONTWAIT` and set the BIO retry flags (immediate `EAGAIN` → `SSL_ERROR_WANT_READ`), matching OpenSSL's default BIO. The fiber/io_uring path is unchanged for blocking sockets. The non-blocking state is read fresh from the fd on each call via `fcntl(F_GETFL)`, because `Poco::Net::SocketImpl::getBlocking()` is a cached flag the raw-`fcntl` probe never updates (and silk sockets reject `setBlocking(false)` outright). Adds a regression test (`SilkFiberSecureSocketTest.NonBlockingPeekDoesNotBlockOnIdleConnection`) that drives the real `getSocketState` path against an idle TLS connection with a 5 s receive timeout and asserts it returns in under 500 ms; without the fix it blocks the full timeout. Not for changelog: the silk fiber BIO is infrastructure with no in-tree production consumer yet, so no released user is affected. Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111852",
          "createdAt": "2026-07-24T20:53:25Z",
          "updatedAt": "2026-08-13T13:43:09Z",
          "timestamp": "2026-08-13T13:43:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "mstetsyuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:db1bb61b9b8a81c31eea",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113833",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113833",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add JOIN observability columns to system.query_log",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111352 Related: https://github.com/ClickHouse/ClickHouse/issues/111748 ## Motivation `system.query_log` says almost nothing about what the JOINs in a query actually did. Answering questions like \"which queries ran a `CROSS` join we didn't expect\", \"which joins fell back to `grace_hash` and spilled to disk\", or \"which algorithm was really chosen when `join_algorithm = 'auto'`\" currently requires re-running `EXPLAIN PIPELINE` (impossible post-mortem) or scraping `ProfileEvents` query by query. ## Changes Four columns are added to `system.query_log`, with the matching fields in `QueryLogElement`: - `used_number_of_joins` (`UInt64`) — the number of physical joins executed by the query. It is collected from the query pipeline, so it reflects the joins that really ran after all optimizations, not the number of `JOIN` clauses in the query text. - `used_join_algorithms` (`Array(LowCardinality(String))`) — the algorithms that were actually used, e.g. `hash`, `parallel_hash`, `grace_hash`, `direct`, `full_sorting_merge`, `partial_merge`. The `join_algorithm` setting only lists the allowed algorithms; the choice among them happens at runtime, and an algorithm can even be replaced mid-execution (a `hash` join switching to `grace_hash` under memory pressure). - `used_join_kinds` (`Array(LowCardinality(String))`) — Kind of the joins present in the query. - `used_join_strictness` (`Array(LowCardinality(String))`) — Strictness of the joins present in the query. - `join_spilled_to_disk` (`UInt8`) — whether any of the joins spilled to disk. This PR currently adds the schema only. The fields are declared but nothing populates them yet, so the columns read as `0` and `[]`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added four columns to `system.query_log` describing the JOINs a query executed: `used_number_of_joins` (the number of physical joins in the executed pipeline), `used_join_algorithms` (the algorithms actually used at runtime, which can differ from the `join_algorithm` setting), `used_join_kinds` (`INNER`, `LEFT`, `CROSS`, `ASOF` and so on), and `join_spilled_to_disk` (whether any join wrote temporary data to disk). This makes it possible to find problematic JOIN patterns across a fleet without reproducing each query.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113833",
          "createdAt": "2026-08-07T14:11:36Z",
          "updatedAt": "2026-08-13T13:43:00Z",
          "timestamp": "2026-08-13T13:43:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-feature"
          ],
          "author": "Manerone",
          "state": "open",
          "assignees": [
            "Fgrtue"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3094d7e9a612c8354c85",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112873",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112873",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Azure: log batch-delete events when SubmitBatch itself fails",
          "text": "> **Series**: #112871 -> **#112873** (this), #112872, #112874, #112875, #112876 **Problem.** When the batch `SubmitBatch` delete request itself fails, the per-object response loop is skipped, so zero `system.blob_storage_log` Delete events are recorded for the whole batch. Failure scenario: ``` removeObjectsBatchIfExists: SubmitBatch() throws (e.g. 403 on the batch endpoint) -> per-object GetResponse() loop skipped -> 0 Delete events logged ``` **Fix.** Record a Delete event per object on batch-level failure before rethrowing. **Changes.** - `Disks/…/AzureBlobStorage/AzureObjectStorage::removeObjectsBatchIfExists`: emit per-object `blob_storage_log` Delete events on `SubmitBatch` failure. - `tests/integration/test_azure_403_handling`: batch-delete-logging test + `configs/blob_log.xml`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed Azure batch object deletion recording no `system.blob_storage_log` Delete events when the batch request itself failed, leaving the whole batch unlogged. A Delete event is now recorded for each object on a batch-level failure.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112873",
          "createdAt": "2026-08-01T06:53:30Z",
          "updatedAt": "2026-08-13T13:42:46Z",
          "timestamp": "2026-08-13T13:42:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "arsenmuk",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:38c20d9359bbec3e04de",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112705",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112705",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve canceling queries in nested expression functions in FilterTransform",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/103705 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve canceling queries by `KILL QUERY` and Ctrl+C in `clickhouse-client` while a long-running function is evaluated in a filter expression: `FilterTransform` (`WHERE`, including the `FilterSortedStreamByRange` path) and `TotalsHavingTransform` (`HAVING` with `WITH TOTALS`) now forward cancellation into nested expression functions. This supersedes https://github.com/ClickHouse/ClickHouse/pull/103705 by @rvasin (Roman Vasin), whose commits are preserved here — all credit for the feature goes to them. The original branch is in a fork with maintainer edits disabled, so the merge conflicts with `master` and the remaining review feedback could not be pushed there. On top of the original PR, this PR: - Merges `master` and resolves the conflicts (`FailPoint.cpp` failpoint list, `FilterTransform.cpp` includes). - Addresses the two unresolved review threads (the AI review \"Request changes\" verdict): - `FilterSortedStreamByRange` holds a private `FilterTransform` and calls its `transform` directly, but the pipeline cancels only the outer processor. Now `FilterSortedStreamByRange::onCancel` forwards the cancellation into the inner transform, so `cancelExecution` reaches functions of range-filter predicates from `PartsSplitter`. - `TotalsHavingTransform` evaluated the `HAVING` expression without a cancellation callback, so a long-running function in `HAVING` remained uninterruptible. It now mirrors `FilterTransform`: `onCancel` cancels function execution, `expression->execute` gets a `check_cancelled` callback, and the transform returns before the totals/filter bookkeeping when cancelled. - Adds a test `04658_kill_query_having_totals_pause` with a new `totals_having_transform_pause` failpoint, mirroring `04612_kill_query_filter_pause`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112705",
          "createdAt": "2026-07-31T05:28:48Z",
          "updatedAt": "2026-08-13T13:40:43Z",
          "timestamp": "2026-08-13T13:40:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 32
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:289d46a670820a5ac081",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114607",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114607",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix an infinite uncancellable loop in `hop` and `windowID` on an interval whose span wraps modulo 2^32",
          "text": "<!-- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed an infinite, uncancellable loop in functions `hop` and `windowID` when the span of an interval argument in seconds is a multiple of 2^32 (for example, `toIntervalDay(2147483648)`): the wrapped subtraction dodged the time-overflow check, and with constant arguments the loop ran at analysis time, where the query could not even be killed. An interval like `toIntervalDay(2147483648)` spans `2^31 * 86400 = 43200 * 2^32` seconds, so subtracting it in the wrapping `UInt32` arithmetic of `AddTime` is a no-op: `wstart == wend`, which dodges the `wstart > wend` time-overflow guard added by https://github.com/ClickHouse/ClickHouse/pull/61523. The window-searching loop of `executeHop` then decrements `wend` by one hop at a time past zero, where it wraps back to ~2^32 while staying congruent to its starting point modulo `gcd(86400, 2^32) = 128`, never hits a value `<= time`, and spins forever. With constant arguments the loop runs during constant folding in `QueryAnalyzer::resolveFunction`, at analysis time, where nothing checks the cancellation. The same loop shape exists in `executeHopSlice` (`windowID`). The fix makes the guard `wstart >= wend` - subtracting a whole positive interval must change the time, and equality only happens on a wrap - and adds a wrap check to the loop itself (the new `wend` coming out greater than the old one throws the same `Time overflow` error), which terminates every wrapping loop regardless of how the pre-loop values were corrupted. The new test `04891_time_window_functions_wrapped_interval_overflow` covers both functions with the fuzzed arguments (both hung before the fix) and checks that a sane `hop`/`windowID` is unaffected. Found by the AST fuzzer in the stress test of a CI run of https://github.com/ClickHouse/ClickHouse/pull/112930 (hung check): [Stress test (arm_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112930&sha=1c5bf0b613146bd6666b6d2bd2be55f539f6c0d8&name_0=PR&name_1=Stress%20test%20%28arm_debug%29). Closes: https://github.com/ClickHouse/ClickHouse/issues/114605 Related: https://github.com/ClickHouse/ClickHouse/issues/61521 Related: https://github.com/ClickHouse/ClickHouse/pull/61523",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114607",
          "createdAt": "2026-08-13T09:11:49Z",
          "updatedAt": "2026-08-13T13:40:40Z",
          "timestamp": "2026-08-13T13:40:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cd561d748827a5dabc16",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110321",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110321",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support INSERT ... VALUES in the polyglot SQL dialect",
          "text": "Enable `INSERT ... VALUES` with inline data in the polyglot SQL dialect (`dialect = 'polyglot'`). Previously, running e.g. `INSERT INTO t VALUES (1), (2), (3)` with `dialect = 'polyglot'` failed with `Multi-statement queries are not supported in polyglot dialect mode`. The underlying problem is that transpiling inside the parser cannot deliver inline data to the executor: the transpiled buffer is transient, and the executor overwrites `ASTInsertQuery::tail` with the external input stream, so the inline-data pointers (`data`/`end`) must reference a live query buffer. This transpiles the query up front instead of inside the parser: - The server (`executeQuery`) transpiles a foreign-dialect query to ClickHouse SQL before parsing, keeps the transpiled text alive on the query context, and parses it with the standard parser. Inline INSERT data then points into a live buffer and is processed by the normal machinery. `SET` queries are still parsed as-is so `dialect`/`polyglot_dialect` can always be changed back. - The client (`clickhouse-client`/`clickhouse-local`) parses a non-ClickHouse-dialect query into an AST — which, for a foreign dialect, means transpiling it locally only to drive client-side handling (statement classification, output format, INSERT detection) — but then sends the *original* query text verbatim, without splitting off inline data. The server performs the authoritative transpilation whose result is actually executed, so inline INSERT data lives in a server-owned buffer and survives parsing. The client-side transpilation is throwaway; note this means the transpiler must also be available on the client (a client built without `USE_POLYGLOT` fails locally with `SUPPORT_IS_DISABLED`), and the client and server transpilers are assumed to agree — acceptable for this experimental dialect. Every parse-time setting the query was parsed under (`dialect`, `allow_experimental_polyglot_dialect`, `polyglot_dialect`, `allow_settings_after_format_in_insert`, `implicit_select`, and the parse limits `max_query_size`, `max_parser_depth`, `max_parser_backtracks`) is pinned in the per-query settings sent along with the verbatim text, so the query's own `SETTINGS` clause cannot change how the server reparses that same text (it still applies to the query's execution, and a `SET` still takes effect for subsequent queries). All changes are gated on the dialect, so ordinary ClickHouse INSERTs are unaffected. Validated over the HTTP interface, the native client, and `clickhouse-local` (multi-row and single-row `VALUES`, `INSERT ... SELECT`, and PostgreSQL literal transpilation such as `true`/`false`); `SET` passthrough and multi-statement rejection are preserved. External insert data combined with a foreign-dialect `INSERT` is rejected with `NOT_IMPLEMENTED` instead of being silently dropped, on both surfaces: the client rejects piped stdin and `INFILE` (it sends the query verbatim and cannot forward a data tail), and the server rejects a non-empty HTTP request body appended to a streaming `INSERT` (`POST /?query=INSERT ... &dialect=polyglot` with a body). A foreign-dialect `INSERT` is transpiled as a whole, so the body would go through neither the transpiler nor the `max_query_size` guard, mixing two parsing rules in one `INSERT`. An empty body still works, which is the normal way to run a polyglot `INSERT` over HTTP. Limitations (scoped, experimental): because a foreign-dialect query is transpiled as a whole (the transpiler rewrites the inline data too and cannot know where the SQL header ends without parsing the dialect), the inline `INSERT ... VALUES` data counts towards `max_query_size` — unlike a native ClickHouse `INSERT`, whose inline data is streamed and is not bounded by `max_query_size`. An oversized payload fails-close with a dedicated, actionable error rather than silently changing the `INSERT` size contract; increase `max_query_size` to submit larger inline payloads. Only `INSERT ... VALUES` inline data is transpilable by the bundled dialects. `INSERT ... FORMAT ...` is not: `FORMAT` is a ClickHouse-only extension, so a foreign-dialect parser rejects the query at the inline data that follows (empirically, `postgresql`/`mysql`/`sqlite`/`duckdb`/`snowflake`/`bigquery` all fail at the first data row after `FORMAT`; a hypothetical identity transpiler even drops the raw `FORMAT` payload rather than re-emitting it). A foreign-dialect `INSERT ... FORMAT` therefore fails cleanly with a syntax error and inserts nothing — like `EXPLAIN INSERT ... VALUES`, which is also not transpilable by the bundled dialects (rejected at the `VALUES` token). The server-owned transpiled buffer that carries the inline data is itself format-agnostic and would handle `FORMAT` data if a transpiler ever produced such a query; the parser also defensively clears the inline-data pointers of an `EXPLAIN`-wrapped `INSERT` — the same way the client unwraps it — so both forms are safe if a future transpiler supports them. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `INSERT ... VALUES` with inline data when using the experimental `polyglot` SQL dialect. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110321",
          "createdAt": "2026-07-13T21:39:09Z",
          "updatedAt": "2026-08-13T13:40:32Z",
          "timestamp": "2026-08-13T13:40:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6287eed88da508fc8e77",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114543",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114543",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix spelling of 'prefetches' and 'prefetched' in documentation",
          "text": "Corrected spelling of 'prefetches' and 'prefetched' in multiple sections. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114543",
          "createdAt": "2026-08-12T21:18:03Z",
          "updatedAt": "2026-08-13T13:40:27Z",
          "timestamp": "2026-08-13T13:40:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "linhgiang24",
          "state": "open",
          "assignees": [
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0f0d985f04d033c91991",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110029",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110029",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Read-through filesystem cache for the experimental ReaderExecutor",
          "text": "First of two PRs adding a read-through cache to the experimental `use_reader_executor` read path (default off), split out of the `ReaderExecutor` series (#103706), after decryption (#109702). This PR adds the cache-provider interface and the filesystem cache tier; the page cache and the richer, coordinated driver follow in a second PR. Adds `ICacheProvider` and the FileCache-backed `DiskCacheProvider`, consulted and populated per read window. The interface and provider are adopted from #103706's redesigned head: a whole-range `resolve(object, offset, range)` returns the window's residency as an ordered list of `Resolution`s in one cache transaction — each hit carrying a `CacheReader`, each populating miss carrying its open `CacheWriter` (a read-only/bypass tier returns writer-less misses). The executor drives it with the simplest possible per-window loop: `resolve` the window start, serve a cache hit straight from the tier's buffer (zero-copy), or on a miss claim the covering cell(s), fetch from the source, populate, and serve one block. The miss read goes through the long connection, so a cold sequential scan streams from one held connection; the window is returned as a `ChainedBuffers` (block-chunked) and decrypted per node on an encrypted disk. Concurrent readers of the same cold cell elect a single downloader via a claim taken before the fetch, so only one populates each cell; a cell another reader is already downloading is fetched through from the source (its populate lands zero bytes) rather than waited on. Coordinated waiting arrives with the page cache PR. Scope of this PR: - The filesystem cache serves known-size sources only — an FS cache is never attached to an unknown-size object, so the cache path is never entered for one. The executor's handling of unknown-size sources is otherwise unchanged from `master`. - No page cache, no cross-window plan, no prefetch, and none of #103706's plan machinery (`ResidencyIterator`, `CoverageMap`, `MemoryPressureMonitor`). - `ChainedBuffers` is functionally unchanged from `master` (only a couple of over-long comments trimmed). - Everything is gated behind `use_reader_executor`; the executor also falls back for the distributed cache and async prefetch, which it does not implement. New settings: `reader_executor_window_size` (serve window, 4 MiB) and `reader_executor_block_size` (buffer chunk, 1 MiB), each at least 4 KiB (rejected at settings load otherwise). Tests: `04511_reader_executor_disk_cache` (an `s3_cache` MergeTree) asserts the executor engages and consults the filesystem cache; `04604_reader_executor_min_size` asserts the sub-4-KiB window/block rejection; IO gtests cover the executor, the provider, and the offset map. Related: https://github.com/ClickHouse/ClickHouse/pull/103706 Related: https://github.com/ClickHouse/ClickHouse/pull/109702 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a read-through filesystem cache to the experimental `ReaderExecutor` read path (`use_reader_executor`, disabled by default). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110029",
          "createdAt": "2026-07-10T16:50:28Z",
          "updatedAt": "2026-08-13T13:40:25Z",
          "timestamp": "2026-08-13T13:40:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b7a5df3f1da863a6a83d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114624",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114624",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Find a LowCardinality needle equal to the type's default value",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `has`, `indexOf`, `countEqual`, `mapContainsKey`, `mapContainsValue` and `Map` subscript returning \"not found\" for a constant `LowCardinality` needle equal to the element type's default value, such as an empty `String` or a zero number. ### Description **The problem.** A constant needle equal to the element type's default value is never found in an `Array(LowCardinality(T))` or a `Map` with a `LowCardinality` key. No exception, so a `WHERE` on such a predicate silently drops rows: ```sql CREATE TABLE t (a Array(LowCardinality(String))) ENGINE = Memory; INSERT INTO t VALUES (['', 'a']); SELECT has(a, '') FROM t; -- 0, expected 1 ``` Likewise for `0` over the numeric types, a NUL-padded `FixedString` and the `1970-01-01` `Date`, while `m['']` returns `''` instead of the stored value. Any non-default needle is correct. Reproduces on 26.5 to 26.7. **Root cause.** A `LowCardinality` dictionary reserves prefix slots for the default and NULL values, and `ReverseIndex` is built with `num_prefix_rows_to_skip`, so the reserved slot is never indexed. `ColumnUnique::uniqueInsertData` compensates for that on the write path; the read path had no counterpart. **The change.** `ColumnUnique::getOrFindValueIndex` now performs the same default-slot match as the write path. `ReverseIndex` is untouched, so no write behaviour moves. That slot becoming reachable brings two equality details. The lookup casts the constant into the element type without reporting loss, so `UInt64(256)` arrived as `UInt8(0)` and would match the default; that slot now answers only if the element type can represent the constant. And a dictionary can hold `-0.0` and `0.0` as separate entries while the shortcut has room for one index, so a zero constant over a float element type is left to the value-comparing path, as `indexOfAssumeSorted` already is. The new test covers both call sites. Lookup timing is unchanged. **Overlap with #112953.** That open PR declines this same shortcut for every float element type, subsuming the float-zero decline here, and the two conflict textually; whichever merges second should keep the broader decline and drop this one, collapsing the duplicated representability predicate to one copy. Non-float types are independent of it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114624",
          "createdAt": "2026-08-13T11:55:50Z",
          "updatedAt": "2026-08-13T13:40:23Z",
          "timestamp": "2026-08-13T13:40:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f0763dc0d6a7cccb32a7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101841",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101841",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add groupBloomFilter aggregate function and bloomFilterContains scalar function",
          "text": "Adds the `groupBloomFilter` aggregate function and the `bloomFilterContains` scalar function for memory-efficient probabilistic set-membership testing. This can be used to detect new values, perform approximate deduplication checks, and compare large datasets without materializing exact sets. Related: https://github.com/ClickHouse/ClickHouse/issues/11700 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `groupBloomFilter` aggregate function for building Bloom filter states and the `bloomFilterContains` scalar function for testing whether values are probably present. The functions support configurable expected element counts, false-positive rates, and seeds, and work with the `-State` and `-MergeState` combinators and `AggregatingMergeTree`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) Documentation is added in `docs/en/sql-reference/aggregate-functions/reference/groupbloomfilter.md` and `docs/en/sql-reference/functions/bloom-filter-functions.md`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101841",
          "createdAt": "2026-04-06T07:19:46Z",
          "updatedAt": "2026-08-13T13:39:51Z",
          "timestamp": "2026-08-13T13:39:51Z",
          "metrics": {
            "reactions": 13,
            "comments": 5
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested"
          ],
          "author": "otselnik",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:603033ef83ae0e54b518",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101273",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101273",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add geo aggregate functions #80186",
          "text": "### Functions for geospatial aggregation #80186 Add three SQL aggregate functions for polygon set operations and convex hull computation: - `groupPolygonUnion` — union of polygonal geometries in a group - `groupPolygonIntersection` — intersection of polygonal geometries in a group - `groupConvexHull` — convex hull of grouped point, linear, and polygonal geometries The functions support typed Geo inputs and `Geometry` variant columns, `State` / `Merge` combinators, binary aggregate-state serialization with writer/reader invariant checks and corruption guards, and well-defined empty-geometry semantics. Closes: https://github.com/ClickHouse/ClickHouse/issues/80186 ### Changelog category: - New Feature ### Changelog entry Added geospatial aggregate functions `groupPolygonUnion`, `groupPolygonIntersection`, and `groupConvexHull`. ### Documentation entry for user-facing changes - [x] Documentation is written in code",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101273",
          "createdAt": "2026-03-30T21:25:08Z",
          "updatedAt": "2026-08-13T13:39:42Z",
          "timestamp": "2026-08-13T13:39:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 36
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "zhemalb",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:afa1f16a1084e6671c6a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114509",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114509",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: Remove private preview banner and stale private preview references fr…",
          "text": "…om SCIM docs SCIM for CHC is GA <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114509",
          "createdAt": "2026-08-12T15:45:37Z",
          "updatedAt": "2026-08-13T13:39:14Z",
          "timestamp": "2026-08-13T13:39:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "ohlookadollar",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:90b49f074485c0ef4a59",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112386",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112386",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keeper: don't refuse an empty-data `Set` at the memory soft limit",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/109460 **Background.** `ZooKeeper::initSession` writes empty data to `<sessions_path>/<zookeeper_name>/<server-uuid>` purely for the side effect of bumping that znode's version, and records the new version. Critical writes then attach `Check(session_path, recorded_version)` via `addCheckSessionOp`, so once another instance of the same server establishes a session and bumps the version, the older instance's writes are atomically rejected with `ZBADVERSION`. It is a fencing token against a stale process still committing data — the payload is irrelevant, which is why it is empty. **Problem.** When Keeper crosses `max_memory_usage_soft_limit`, clients can no longer re-establish their Keeper session — it fails with `Coordination error: Out of Memory` on a path under the sessions path — so tables that depend on session re-establishment stay read-only for as long as the memory event lasts, rather than as long as the write pressure lasts. In one observed incident this single write accounted for 96.7% of everything Keeper refused (1,330,721 of 1,376,368 requests over roughly two hours), and only 6 refused `Set`s were on any other path — so the limit was almost exclusively blocking recovery rather than holding back data volume. **Root cause.** The soft-limit gate refuses whatever `checkIfRequestIncreaseMem` reports as memory-increasing, and that function decides by op type: it returns true for every `Set`. But a `Set` cannot allocate a znode — the node must already exist, otherwise the request fails with `ZNONODE` and stores nothing — so a `Set` with empty data cannot increase the amount of data Keeper stores. `ZooKeeper::initSession` registers a session by writing empty data to `<sessions_path>/<zookeeper_name>/<server-uuid>`, so that write is refused and the client cannot recover from the condition that caused the refusal. The `Multi` branch has the same defect by a different route: it sums `set_req.bytesSize()`, which includes the path, the version and the xid, so an empty `Set` inside a `Multi` counts as growth proportional to its path length. **Solution.** Classify `Set` by its data, so empty data is not memory-increasing, and make the `Multi` branch sum `set_req.data.size()`. The classifier was duplicated verbatim in both dispatchers, so it moves to `KeeperCommon` where the copies cannot drift — which also makes it directly testable. `Create`, `Remove`, `SetACL` and `Auth` classification are unchanged, and writes that genuinely allocate are still refused: a `Multi` that creates an ephemeral node remains rejected, so this does not by itself return a replicated table to a writable state. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Keeper no longer refuses a `Set` request with empty data when it is over `max_memory_usage_soft_limit`. Such a request cannot increase the amount of stored data, and refusing it prevented clients from re-establishing their session, which could keep tables read-only for the whole duration of a Keeper memory event.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112386",
          "createdAt": "2026-07-29T03:13:12Z",
          "updatedAt": "2026-08-13T13:38:36Z",
          "timestamp": "2026-08-13T13:38:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "al13n321"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:df807241946f8d2d8345",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114608",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114608",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make the 04780 index-analysis allocation oracle the minimum of several runs",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/104948 `04780_json_subcolumn_index_match_not_quadratic` compares `MemoryAllocatedWithoutCheckBytes` of a dotted-constant `EXPLAIN indexes = 1` against a no-dots control with a 150% threshold, measuring each arm with a single query. A single run is not a stable oracle: whichever query happens to be the first to touch a cache or spin up a thread pool absorbs a transient multi-megabyte allocation, and the dotted arm always runs first in the loop, so such a one-off lands on it and inflates the ratio arbitrarily. Locally the effect is easy to see: the first query after a server start reports 8–22 MB in this counter against a ~2.8 MB steady state for the identical query. The test failed this way on at least 6 unrelated PRs since 2026-08-12 (e.g. `Stateless tests (amd_asan_ubsan, distributed plan, parallel)` on https://github.com/ClickHouse/ClickHouse/pull/104948 at commit 81300c20: `longidx index analysis over a constant with 100000 dots allocated 4216780 bytes, more than 150% of the no-dots control (22224 bytes)` — reruns passed; CIDB shows the same failure on #106011, #108522, #114475, #96130, #114476). Measure each arm three times and take the minimum: a genuine quadratic regression is deterministic and shows up in every run, so the oracle keeps discriminating (the minimum can only remove one-sided transient noise), while a one-off transient can no longer fail the test. Verified against the current master binary: the modified test passes repeatedly. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114608",
          "createdAt": "2026-08-13T09:13:17Z",
          "updatedAt": "2026-08-13T13:38:02Z",
          "timestamp": "2026-08-13T13:38:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:be9583af0e8d0472dab8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111394",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111394",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fsync backup files and directories when writing a backup to local disk",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111320 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `BACKUP ... TO File(...)` / `Disk(...)` now fsyncs the backup data files, the `.backup` manifest and the containing directories before reporting `BACKUP_CREATED`, so an acknowledged backup to local storage survives power loss. Controlled by the new backup setting `fsync_backup_files` (default `true`). Object-storage destinations (`S3`/`Azure`) are unaffected. ### Description Fixes #111320. `BACKUP ... TO File()/Disk()` returned `BACKUP_CREATED` without issuing any `fsync`/`fdatasync` at the destination: not the data files, not the `.backup` manifest, and not the destination directories (there was no `fsync` anywhere in `src/Backups/`). On power loss after the acknowledgement the backup could be lost entirely or left torn, even though `BACKUP_CREATED` is exactly what an operator relies on before dropping the source data. Object-storage destinations were already durable (a completed upload is persisted server-side); only local `File()`/`Disk()` were affected. Report URL: https://github.com/ClickHouse/ClickHouse/issues/111320 (reproduced 3/3 with a `dm-flakey` power-loss simulation). Fix, gated on the new backup setting `fsync_backup_files` (default `true`), following the durability audit family (#68958 -> #111346, #111269 -> #111335): - Two writer hooks with a no-op default on `IBackupWriter`, overridden only by the local `File`/`Disk` writers (`S3`/`Azure`/`Memory`/`Null` inherit the no-op): `syncFileToDisk(file_name)` (fdatasync a written file, covering both the buffered and the native `fs::copy`/`IDisk::copyFile` paths) and `syncDirectoriesToDisk()` (fdatasync every directory the backup created, deepest-first, plus the backup root's parent, via `LocalDirectorySyncGuard` / `IDisk::getDirectorySyncGuard`). - Each data file is synced right after it is written in `BackupImpl::writeFile` (safe under the concurrent write path: each call fsyncs its own file). - In `BackupImpl::finalizeWriting` the `.backup` manifest (or, for archives, the archive file) is synced last, after all data files, so a persisted manifest never precedes its payload. Directory syncing runs for every writer, including the internal writers of `BACKUP ON CLUSTER` which write their own data files. Verified locally with ProfileEvents: `fsync_backup_files=1` issues `FileSync`/`DirectorySync` for the whole backup (data files + manifest + every nested directory); `fsync_backup_files=0` issues none (matching the previous behavior); the backup still restores correctly. Regression test `tests/queries/0_stateless/04412_backup_to_file_fsync.sh` asserts, via the `FileSync`/`DirectorySync` ProfileEvents of the `BACKUP` query in `system.query_log`, that the fsyncs are issued when `fsync_backup_files=1` and are absent when `fsync_backup_files=0`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111394",
          "createdAt": "2026-07-22T13:30:01Z",
          "updatedAt": "2026-08-13T13:37:42Z",
          "timestamp": "2026-08-13T13:37:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "jkartseva"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:1d4d2afda4b6733659fd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114619",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114619",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112573 to 26.7: Fix async bounded read buffer readbigat race",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114619",
          "createdAt": "2026-08-13T10:34:20Z",
          "updatedAt": "2026-08-13T13:37:32Z",
          "timestamp": "2026-08-13T13:37:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll",
          "state": "closed",
          "assignees": [
            "kssenii",
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:66897d8699221b71176f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111451",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111451",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Release pull request for branch 26.7",
          "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111451",
          "createdAt": "2026-07-22T17:40:24Z",
          "updatedAt": "2026-08-13T13:37:32Z",
          "timestamp": "2026-08-13T13:37:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "release"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c5998e74681889869594",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108786",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108786",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Switch the default compression to ZSTD(3) for table data and network",
          "text": "Switch the default compression in ClickHouse from `LZ4` to `ZSTD(3)`, for both on-disk table data and network communication. ### Motivation `LZ4` has been the default for a long time and optimizes for speed, but `ZSTD(3)` gives substantially better compression ratios at a still-modest CPU cost, reducing both storage footprint and network traffic out of the box. ClickHouse Cloud already uses a stronger default; this aligns the self-managed defaults. ### What changes - **On-disk table data.** `CompressionCodecFactory::getDefaultCodec` now returns `ZSTD(3)` instead of `LZ4`. For `MergeTree` column data the built-in default is *size-aware*: when a column specifies no codec and no `<compression>` server-config case matches, `CompressionCodecSelector::choose` uses the faster `LZ4` for parts smaller than 100 MB and `ZSTD(3)` for larger parts, so freshly inserted data starts as `LZ4` and the bigger parts produced by background merges switch to `ZSTD(3)`. The direct `getDefaultCodec` users — the `Log` family, part checksums, the HTTP `compress=1` output, the `StripeLog` data stream, and the persistent `Set`/`Join` files — are not size-aware and switch uniformly from `LZ4` to `ZSTD(3)`. External temporary (spill) files are unaffected: they keep their own `temporary_files_codec` default (`LZ4`). - **Network.** The defaults of `network_compression_method` (`LZ4` → `ZSTD`) and `network_zstd_compression_level` (`1` → `3`) are changed, so the native client/server and Distributed (server/server) protocol uses `ZSTD(3)` by default. The changes are recorded in `SettingsChangesHistory.cpp` so the `compatibility` setting restores the previous behavior for these paths. The distributed-query streaming exchange (`StreamingExchangeSink`) does not read `network_compression_method` and always uses the server default codec; this is safe because each frame is self-describing (the receiver auto-detects the codec) and the exchange is a transient, same-version channel — `StreamingExchangeProtocol` rejects peers on a different protocol version, so a stream is never read back by a node expecting a different codec. - **Upgrade safety.** The append-only `Log`/`TinyLog`/`StripeLog` engines resolve the default codec at write time, so a table written before the upgrade (`LZ4`) and appended to after it (`ZSTD(3)`) ends up with mixed-codec blocks in one file. Their readers now pass `allow_different_codecs = true` (as the `MergeTree` reader already does) so such files still read back correctly. The legacy custom-frame Keeper snapshot format (`compress_snapshots_with_zstd_format = false`) is pinned to `LZ4` so it stays the documented format. Columns, tables, and connections that specify a codec or method explicitly are unaffected — with one exception: streams written through the built-in default codec directly, which ignore per-column codecs and the `<compression>` config. The `StripeLog` data stream, the persistent `Set`/`Join` backup files (written by `SetOrJoinSink` and `StorageJoin::mutate` through a `CompressedWriteBuffer` with no codec), and `MergeTree` auxiliary streams such as part checksums (`checksums.txt`, via `MergeTreeDataPartChecksums::write`) always follow the new default; see the changelog entry below for the per-path rollback. ### Validation Built locally and verified against the new binary: - A `MergeTree` table created with no explicit codec follows the size-aware default: small parts report `LZ4` and parts larger than 100 MB report `ZSTD(3)` as their `default_compression_codec` in `system.parts`. Direct `getDefaultCodec` streams (e.g. `StripeLog`) report `ZSTD(3)` regardless of size. - `system.settings` shows `network_compression_method = ZSTD` and `network_zstd_compression_level = 3`. - End-to-end `SELECT`/`INSERT` over the native protocol with default (now `ZSTD`) network compression work; all of `LZ4`/`lz4hc`/`zstd`/`none` still work and an invalid method still errors with `BAD_ARGUMENTS`. - The `02995_new_settings_history` consistency check passes (both setting changes are recorded). - The on-disk byte-dump tests (`02047_log_family_*_data_file_dumps`) were regenerated for the new default: `Log`/`TinyLog` keep `LZ4`-pinned columns, while `StripeLog`'s shared `data.bin`/`index.mrk` reflect the `ZSTD(3)` default (it uses `getDefaultCodec` and ignores the per-column codec). Size-sensitive stateless and integration tests that assert compressed byte sizes were pinned back to `LZ4` (cache segment sizes, distributed-batch corruption offsets, full-disk thresholds, frozen-part checksums). The CI performance comparison does **not** measure this tradeoff: the harness pins the built-in default codec to `LZ4` (`tests/performance/scripts/config/config.d/compression.xml`) and `network_compression_method` to `LZ4` (`tests/performance/scripts/config/users.d/perf-comparison-tweaks-users.xml`) on *both* the reference and the tested server, deliberately, so the report measures query logic rather than drowning in expected codec regressions. It therefore serves only as a regression filter for the non-codec parts of this change; the storage/network-vs-CPU tradeoff of `LZ4` → `ZSTD(3)` itself is a well-known property of the two codecs and is not measured by this pull request's CI. Documentation updated accordingly, including the `Native` format and protocol specifications and the `<compression>` config examples. ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The default compression method is changed from `LZ4` to `ZSTD(3)`. For client/server and server/server network communication it switches uniformly to `ZSTD(3)`. For on-disk table data, `MergeTree` column data now uses a size-aware built-in default — parts smaller than 100 MB use `LZ4` and larger parts use `ZSTD(3)` — while the direct built-in-default streams (the `StripeLog` data stream, the persistent `Set`/`Join` files, the HTTP `compress=1` framed output, and part checksums) switch uniformly from `LZ4` to `ZSTD(3)`. This improves compression ratios and reduces storage and network usage out of the box, at a modest increase in CPU usage. None of the on-disk or HTTP paths are controlled by `compatibility`; the previous behavior can be restored per path as follows. For client/server and Distributed network compression, use the `compatibility` setting (or set `network_compression_method = 'LZ4'`). For `MergeTree` column data, set a column/table `CODEC(LZ4)` or configure the server `<compression>` default to `LZ4`. For the `Log`/`TinyLog` engines, set a column `CODEC(LZ4)` (their default does not consult the `<compression>` config). The `StripeLog` data stream, the persistent `Set`/`Join` backup files, the HTTP `compress=1` framed output, and `MergeTree` auxiliary streams such as part checksums (`checksums.txt`) use the built-in default codec directly — they ignore per-column codecs and the `<compression>` config — so they always use the new `ZSTD(3)` default and have no runtime rollback; reading stays correct because every compressed frame is self-describing.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108786",
          "createdAt": "2026-06-29T02:28:41Z",
          "updatedAt": "2026-08-13T13:36:12Z",
          "timestamp": "2026-08-13T13:36:12Z",
          "metrics": {
            "reactions": 1,
            "comments": 69
          },
          "labels": [
            "pr-performance",
            "pr-backward-incompatible",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6ea0221cd9dedecb2a08",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109531",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109531",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix text index direct read with patch parts (lightweight updates)",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/106460 (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/99543 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/106460 Related: https://github.com/ClickHouse/ClickHouse/pull/99543 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results (and an `UNKNOWN_IDENTIFIER` error for `MATERIALIZED` indexed columns) when a query reads from a `text` index together with a column updated by a lightweight update (patch parts). Direct reading from the text index is now disabled for parts that have patch parts. ### Description Direct reading from a `text` index (`query_plan_direct_read_from_text_index`, enabled by default) produces the search result as a virtual column via a dedicated index-read step prepended to the reader chain. That step reads no physical data, so it cannot anchor patch application, which aligns patches to the block by `_part_offset`. When a queried part has patch parts (from a lightweight `UPDATE`) and a patch-applied column is read together with the direct-index virtual column, the direct-read path and the patch-application path do not line up. The result was either dropped rows (wrong results) or, when the indexed column was `MATERIALIZED`, an `UNKNOWN_IDENTIFIER` error while evaluating the virtual column's default expression. This happens even when the patched column is unrelated to the indexed column, so the existing per-index `canUseIndex` check was not sufficient. The fix disables direct text-index reading for the whole query when any queried part has patch parts, falling back to regular index reading (identical to `query_plan_direct_read_from_text_index = 0`, which always produced correct results). Bisected to #99543. Reproducer from the issue (`json_title String MATERIALIZED data_as_json.title::String` with a `text` index) plus a plain-column variant are added as `03100_lwu_51_text_index_patched_column`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109531",
          "createdAt": "2026-07-06T15:52:24Z",
          "updatedAt": "2026-08-13T13:36:08Z",
          "timestamp": "2026-08-13T13:36:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 20
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.5-must-backport"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "CurtizJ"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:50268256b7c98ac05b9d",
        "signalId": "github:ClickHouse/ClickHouse:issue:113763",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113763",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Propagating `getArgumentsThatCanBeOnlyNull` through the combinators silently changes the result type and value of every pre-existing `<agg>If<Combinator>(x, NULL)` expression",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ ### Describe what's wrong After this PR, `countIfOrNull(number, NULL)` returns `Nullable(UInt64)` `NULL` where it returned `UInt64` `0`; `uniqIfOrNull(number, NULL)` likewise flips `0` -> `NULL`; `sumIfResample(0, 2, 1)(number, NULL, number % 2)` returns `Array(UInt64)` `[0,0]` where it returned the scalar `Nullable(Nothing)` `NULL`; and `sumIfState(number, NULL)` now produces a real `AggregateFunction(sumIf, UInt64, Nullable(Nothing))` state where it produced the constant `Nullable(Nothing)`. The same shift applies to `count`, `uniq`, `any`, `min`, `max`, `avg`, `groupArray`, `quantile`, `varSamp` and `topK` under `-State`, `-OrNull`, `-OrDefault`, `-Resample`, `-ArgMin` and `-ArgMax`. None of this is mentioned in the PR title, description or changelog entry, all of which speak only about the new `gini` function. **Root cause:** src/AggregateFunctions/Combinators/AggregateFunctionState.h:163 (and the six sibling overrides listed in affected_locations) newly forward `getArgumentsThatCanBeOnlyNull` from the nested function. Combined with the existing consumer at src/AggregateFunctions/Combinators/AggregateFunctionNull.cpp:88, this removes the `AggregateFunctionNothing` fold for every `<agg>If<Combinator>(x, NULL)` expression - a user-visible change to functions completely unrelated to `gini`, shipped under the changelog category `New Feature`. **Why we believe this is a bug:** `AggregateFunctionFactory::get` (src/AggregateFunctions/AggregateFunctionFactory.cpp:106-127) applies the `Null` combinator whenever any argument type is `Nullable`. `AggregateFunctionCombinatorNull::transformAggregateFunction` (src/AggregateFunctions/Combinators/AggregateFunctionNull.cpp:88) asks the nested function for `getArgumentsThatCanBeOnlyNull` and, if an only-null (`Nullable(Nothing)`) argument is NOT in that set, folds the whole aggregate into `AggregateFunctionNothing` (line 116-118). Before this PR only `AggregateFunctionIf` overrode the method, so as soon as a second combinator wrapped the `-If` function the information was lost and the fold happened. The PR makes `-State`, `-SimpleState`, `-OrFill`, `-Distinct`, `-Resample`, `-ArgMin`/`-ArgMax` and `AggregateFunctionNullBase` forward the nested set, so the fold no longer happens and the real combinator chain is built instead. Risk if broken: result types of persisted `AggregateFunction` columns and materialized-view schemas change on upgrade, and aggregate values silently flip from `0` to `NULL`. **Affected locations:** - `src/AggregateFunctions/Combinators/AggregateFunctionState.h:163` — new getArgumentsThatCanBeOnlyNull forwarding to nested_func - `src/AggregateFunctions/Combinators/AggregateFunctionOrFill.h:388` — new forwarding override, drives -OrNull/-OrDefault - `src/AggregateFunctions/Combinators/AggregateFunctionResample.h:247` — new override, nested set plus last_col - `src/AggregateFunctions/Combinators/AggregateFunctionDistinct.h:375` — new forwarding override - `src/AggregateFunctions/Combinators/AggregateFunctionCombinatorsArgMinArgMax.cpp:225` — new override, nested set plus key_col - `src/AggregateFunctions/Combinators/AggregateFunctionNull.h:285` — AggregateFunctionNullBase now forwards to nested_function - `src/AggregateFunctions/Combinators/AggregateFunctionSimpleState.h:108` — new forwarding override - `src/AggregateFunctions/Combinators/AggregateFunctionNull.cpp:88` — consumer: decides whether to fold into AggregateFunctionNothing **Impact:** On upgrade, existing queries and materialized views that stack a second combinator on `-If` with a constant-`NULL` (or `Nullable(Nothing)`) filter change their result type, and `countIfOrNull`/`uniqIfOrNull` change their value from `0` to `NULL`. `<agg>IfResample(...)` changes shape from a scalar to an `Array`, which makes downstream expressions that consumed the scalar fail. `<agg>IfState(x, NULL)` becomes a real `AggregateFunction(...)` state, which is schema-affecting for `CREATE MATERIALIZED VIEW ... AS SELECT sumIfState(...)`. ## Assumptions The bot recorded these claims it could not verify directly from source. A maintainer ✅ confirms; ❌ flags a wrong premise (the bot should rework or close the finding). - [ ] **The new results are semantically better than the old ones (a real state beats a folded constant, and `-Resample` must return an Array), so the intent is a fix rather than a regression.** - *Why unverifiable:* The PR description does not mention the combinator change at all, so the author's intent for non-`gini` functions is not stated anywhere. - *Falsifiable test:* Ask the author/maintainer whether the change in `countIfOrNull(x, NULL)` from `0` to `NULL` and in `sumIfResample(...)(x, NULL, k)` from scalar `NULL` to `[0,0]` is intended; if yes, the changelog category must become `Backward Incompatible Change` and the behavior must be covered by a test. ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `31081d9f050140`). ### How to reproduce ```sql -- Test: result of an -If aggregate with a constant NULL filter, combined with a second combinator. SELECT toTypeName(countIfOrNull(number, NULL)), countIfOrNull(number, NULL) FROM numbers(5) ORDER BY 1; SELECT toTypeName(uniqIfOrNull(number, NULL)), uniqIfOrNull(number, NULL) FROM numbers(5) ORDER BY 1; SELECT toTypeName(sumIfResample(0, 2, 1)(number, NULL, number % 2)), sumIfResample(0, 2, 1)(number, NULL, number % 2) FROM numbers(5) ORDER BY 1; SELECT toTypeName(x) FROM (SELECT sumIfState(number, NULL) AS x FROM numbers(5)) ORDER BY 1; ``` [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/f1146cf3-cf28-4eea-bbd7-9ecef5328573) ### Expected behavior ``` UInt64 0 UInt64 0 Nullable(Nothing) \\N Nullable(Nothing) ``` ### Error message and/or stacktrace ``` Nullable(UInt64) \\N Nullable(UInt64) \\N Array(UInt64) [0,0] AggregateFunction(sumIf, UInt64, Nullable(Nothing)) ``` ### Additional context **Open risks:** - `AggregateFunctionArray`, `AggregateFunctionForEach`, `AggregateFunctionMap` and `AggregateFunctionMerge` also expose `getNestedFunction` but were NOT given the override. I audited all four: each rejects a top-level `Nullable(Nothing)` argument before the `Null` combinator can matter (`Illegal type Nullable(Nothing) of argument for aggregate function with Array/ForEach suffix. Must be array.`, `Aggregate function Map requires map as argument.`, and `-Merge` requires an `AggregateFunction(...)` argument), so the omission is not reachable - not filed as a separate finding. - State binary compatibility is unaffected: `CREATE TABLE (s AggregateFunction(sumIf, UInt64, Nullable(Nothing)))` and `CAST(... AS AggregateFunction(sumIf, UInt64, Nullable(Nothing)))` behave identically on 26.7 and on the PR build (both expect 8 bytes), so only the inferred type of `sumIfState(x, NULL)` moved. **Suggested fix:** Keep the combinator propagation (it is the more consistent behavior), but re-categorize the changelog entry as `Backward Incompatible Change`, spell out the affected expression shapes in the entry, and add the non-`gini` cases to a stateless test so the new behavior is pinned. If the change to `countIfOrNull(x, NULL)` (`0` -> `NULL`) is not intended, restrict the propagation to the combinators `gini` actually needs. **Analysis details:** Confidence HIGH | Severity P2 | Testability: `STATELESS_SQL` Found during automated review of [PR #112280](https://github.com/ClickHouse/ClickHouse/pull/112280). --- _ClickGapAI · Confidence: HIGH · Severity: P2 · Finding: `h_pr112280_001`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #113868 - Merged into: `26.8.1.1317` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113763",
          "createdAt": "2026-08-07T04:57:14Z",
          "updatedAt": "2026-08-13T13:36:04Z",
          "timestamp": "2026-08-13T13:36:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "comp-aggregate-functions"
          ],
          "author": "clickgapai",
          "state": "closed",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6dcb1189de5a39d5a065",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110626",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110626",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain data/expected structure mismatches in INSERT parse errors",
          "text": "Improves the error message shown when parsing the data of an `INSERT` fails. Previously, inserting data whose structure does not match the destination produced a confusing low-level parse error (for example `Cannot parse input: expected '\\t' before ...`) with no hint about the real cause. In the linked issue, the destination's inferred schema had integer columns while the inserted `TSV` data had string columns, and the reported error pointed at a tab that was actually present. Now, on a parse failure, ClickHouse infers the structure of the data being inserted (only when the format has a data schema reader, and solely for diagnostics) and, if it does not correspond to the expected structure, appends an explanation listing both the inferred and the expected structure. For example: ``` Code: 27. DB::Exception: Cannot parse input: expected '\\t' before: 'page_view... ... The structure of the data being inserted does not match the structure expected by the query, which is likely the cause of the parsing error. Inferred structure of the input data (in format `TSV`): c1 Nullable(Int64) c2 Nullable(String) c3 Nullable(String) Expected structure: c1 Int64 c2 Int64 c3 Int64 ``` The check is wired through a lazy provider on `IInputFormat` that runs only on a genuine parse error, so there is no cost on the happy path. The comparison ignores the artificial `Nullable` wrapper that schema inference adds by default, so inserting valid data into non-nullable columns is not falsely flagged. It covers the synchronous local/server path, client-side parsing, and the asynchronous insert queue (the default path now that `async_insert` is enabled by default). Closes: https://github.com/ClickHouse/ClickHouse/issues/110622 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When parsing of the data being inserted by an `INSERT` fails, the error message now explains a likely structure mismatch by comparing the structure inferred from the data with the structure expected by the query.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110626",
          "createdAt": "2026-07-15T23:41:14Z",
          "updatedAt": "2026-08-13T13:35:45Z",
          "timestamp": "2026-08-13T13:35:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 25
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:072f5d389fc9235fde60",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114570",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114570",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix presentation URL in `clickhouse-git-import` help",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Documentation entry for user-facing changes The help message of `clickhouse-git-import` referenced the presentation at `https://presentations.clickhouse.com/matemarketing_2020/`. Replace it with `https://presentations.clickhouse.com/2020-matemarketing/`. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1328` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114570",
          "createdAt": "2026-08-13T02:12:26Z",
          "updatedAt": "2026-08-13T13:35:12Z",
          "timestamp": "2026-08-13T13:35:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:034cfd3f01d65ee817b6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114490",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114490",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "ci: remove stale cache status files",
          "text": "Though note that the server was still active, not sure why (since it got SIGTRAP, maybe some deadlock in sanitizers build). Fixes: https://pastila.nl/?004f7797/8607b82ba8c324e4f5081e6c747ad583#ad4mwd7GRSHyUgNa46pJHw==GCM ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1327` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114490",
          "createdAt": "2026-08-12T14:22:41Z",
          "updatedAt": "2026-08-13T13:35:08Z",
          "timestamp": "2026-08-13T13:35:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "azat",
          "state": "closed",
          "assignees": [
            "maxknv"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7e5975a39a418c7abde6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110892",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110892",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain analyze join stats",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/110668 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add information of internal state of joins to `EXPLAIN ANALYZE`",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110892",
          "createdAt": "2026-07-17T16:06:34Z",
          "updatedAt": "2026-08-13T13:34:35Z",
          "timestamp": "2026-08-13T13:34:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "Fgrtue",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:97c5d24dfd01715a9453",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114003",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114003",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix Keeper CORRUPTED_DATA after cross-segment writeAt crash (#112101)",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix ClickHouse Keeper refusing to start with `CORRUPTED_DATA` after a crash during cross-segment Raft log truncation (`writeAt`), which could leave an acknowledged log entry stranded behind stale changelog segments. Closes #112101 ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Detail When `Changelog::writeAt` rewrote an entry in an earlier changelog segment, later segments were removed asynchronously and the rewrite could be fsynced/acked before those unlinks finished. A crash in that window left a duplicated index in the earlier file plus stale higher-index files; startup treated the gap as unrecoverable corruption. This change waits for those `RemoveChangelog` operations **outside** `writer_mutex` before appending the rewrite (durability ordering; no lock cycle with the remove thread). Startup still reports `CORRUPTED_DATA` for changelog gaps so recovery stays manual. ### Test plan - [x] `gtest_coordination_changelog` / `ChangelogTestWriteAtPreviousFile` — after cross-segment `write_at`, superseded changelog files are already gone before the rewrite is acknowledged",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114003",
          "createdAt": "2026-08-09T02:44:31Z",
          "updatedAt": "2026-08-13T13:34:31Z",
          "timestamp": "2026-08-13T13:34:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "nishant-uxs",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:491b83db814d2ea70d8e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108820",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108820",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow Distributed and Remote tables in Replicated databases",
          "text": "`database_replicated_allow_only_replicated_engine` rejects table engines that keep their own unreplicated data on disk in a `Replicated` database. Before this change, `Distributed` was also rejected because its optional background `INSERT` queue makes `storesDataOnDisk` return true, even though the queue is a transient send buffer and the table's actual data belongs to its destination shards. This change distinguishes table-owned on-disk data from auxiliary delivery state. The restriction applies to non-`ATTACH` `CREATE` queries when `database_replicated_allow_only_replicated_engine = 1`. | Engine or group | Result | Reason | |---|---|---| | `ReplicatedMergeTree` family and `SharedMergeTree` | Allow | The engine provides a replication or shared-storage contract. | | Writable non-replicated `MergeTree` with a local or remote storage policy | Reject | It owns unreplicated on-disk table data; a remote disk alone does not establish replication. | | Static read-only non-replicated `MergeTree` | Allow | It cannot create new table-owned data. | | `Log`, `TinyLog`, `StripeLog`, `Set`, `Join`, `EmbeddedRocksDB`, and database `File` | Reject | They own unreplicated on-disk table data. | | `MaterializedPostgreSQL` | Reject | It owns a local nested table. | | `Distributed`, `Remote`, and `RemoteSecure` | Allow | Their optional local queue is auxiliary delivery state rather than data of the table itself. | | `Memory`, `Buffer`, and `Null` | Allow | They do not own on-disk table data. | | `Merge`, `Alias`, and `View` | Allow | They are metadata-only. | | Views with inner tables | Depends | Each generated inner table is created and checked separately. | | External-storage engines, data lakes, external databases, and external queues | Allow | Their data is managed outside ClickHouse-owned table storage. | | Lazy `StorageTableProxy` | Reject conservatively | The nested storage is unknown without loading it. | `ATTACH` remains outside the existing gate. The `Distributed` background `INSERT` queue is local to the node accepting an insert and is not replicated. Users requiring acknowledgement only after data reaches the destination shards should set `distributed_foreground_insert = 1`. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow `Distributed`, `Remote`, and `RemoteSecure` tables in `Replicated` databases when `database_replicated_allow_only_replicated_engine` is enabled, while continuing to reject writable non-replicated `MergeTree` tables using local or remote storage policies.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108820",
          "createdAt": "2026-06-29T15:29:14Z",
          "updatedAt": "2026-08-13T13:34:15Z",
          "timestamp": "2026-08-13T13:34:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "UberDever",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:30745f80d3720718f9f9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113754",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113754",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject STREAM with parallel replicas at plan-build time",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/110144 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `LOGICAL_ERROR` when a `SELECT` used the `STREAM` modifier on a named table together with parallel replicas in the read-tasks mode. With a materialized CTE used as an `IN` set in `PREWHERE`, the server raised `Reading from materialized CTE ... DelayedPortsProcessor gate is missing in the query plan` (a server abort in debug and sanitizer builds). A table carrying `STREAM` is now rejected when parallel replicas are requested: at `enable_parallel_replicas = 2` the query fails with `SUPPORT_IS_DISABLED`, at `1` it runs without them, exactly as already happens for `FINAL`. ### Description `MergeTreeDataSelectExecutor::read` already refuses `STREAM` with parallel replicas, but its `enable_parallel_reading` argument is `canUseParallelReplicasOnFollower` (`StorageMergeTree.cpp:385`, `StorageReplicatedMergeTree.cpp:6305`), so the refusal only fires on a follower: the initiator builds the plan it should have refused and runs it. `STREAM` also suppresses the eager set build: `ReadFromMergeTree::applyFilters` early-returns for a streaming read (`ReadFromMergeTree.cpp:2593`), and that inplace build is what materializes a CTE, so a materialized CTE used as an `IN` set stays unmaterialized and its ungated `StorageMemory` read raises the error at `ReadFromMemoryStorageStep.cpp:87`. The fix extends the existing `FINAL` check in `Planner::buildPlanForQueryNode` (`Planner.cpp:2306` on master) to `STREAM`. It runs before `buildJoinTreeQueryPlan`, so no storage read is reached, and it clears `allow_experimental_parallel_reading_from_replicas` rather than skipping one branch, covering `parallel_replicas_plan_based` too. For `STREAM` it applies on an initiator only. The enclosing `canUseTaskBasedParallelReplicas` is role-blind, and a follower that cleared the setting here would lose its own read-side refusal, so an old initiator plus a new follower silently did a full local `STREAM` read instead of failing; measured on a two-server cluster at `enable_parallel_replicas = 1`, it now fails with `ILLEGAL_STREAM` as it does on two old servers. `FINAL` keeps its unconditional handling, having no read-side refusal to preserve. Scope, measured: this covers the read-tasks mode on a named table, the reported carrier. The custom-key and sampling-key modes never reach the guard's enclosing `canUseTaskBasedParallelReplicas` branch, and `STREAM` with `parallel_replicas_mode = 'custom_key_sampling'` hangs identically before and after this change; that hang stays open. A `TableFunctionNode` carrying `STREAM` is not covered either. An earlier revision covered both and was reduced on review. #110972 rewrites this same loop; a merge has to keep both sides. Found by the AST fuzzer on #110144",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113754",
          "createdAt": "2026-08-07T02:36:15Z",
          "updatedAt": "2026-08-13T13:33:49Z",
          "timestamp": "2026-08-13T13:33:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8fa57ab5570585d89b25",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111152",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111152",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "docs(h3): fix invalid unidirectional-edge examples (#102845)",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ## What The `h3GetOriginIndexFromUnidirectionalEdge` / `h3GetDestinationIndexFromUnidirectionalEdge` examples used directed edge `1248204388774707197`, which fails `h3UnidirectionalEdgeIsValid` and raises `INCORRECT_DATA` instead of returning the documented indexes. ## Fix - Switch the worked examples to the valid edge `1248204388774707199` (already used by the `h3UnidirectionalEdgeIsValid` example on the same page). - Correct origin → `599686042433355775` and destination → `599686043507097599`. - Keep `docs/en/.../h3.md` aligned with the `REGISTER_FUNCTION` examples in the two C++ sources (those strings feed generated docs). ## Why Repro from #102845 — readers should be able to copy-paste the docs without an exception. ## Notes - AI-assisted; human-reviewed. - Docs-only + FunctionDocumentation string updates; no runtime logic change. - Closes #102845 ### Check - [x] Valid edge per H3 (`is_valid_directed_edge`) - [x] Origin/destination pair matches library on that edge <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1324` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111152",
          "createdAt": "2026-07-20T23:19:41Z",
          "updatedAt": "2026-08-13T13:33:44Z",
          "timestamp": "2026-08-13T13:33:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 18
          },
          "labels": [
            "pr-documentation",
            "manual approve",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "Bartok9",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:5d3ed8b2dd1c3332b266",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114320",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114320",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "feat(clickpipes, docs): document clickpipe DDL DEFAULT propagation logic",
          "text": "- https://github.com/PeerDB-io/peerdb/pull/4635 - https://github.com/PeerDB-io/peerdb/pull/4632 ### Changelog category - Documentation (changelog entry is not required) ### Changelog entry Document `DEFAULT ...` DDL propagation for mysql and postgres ClickPipes. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1323` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114320",
          "createdAt": "2026-08-11T13:05:16Z",
          "updatedAt": "2026-08-13T13:33:41Z",
          "timestamp": "2026-08-13T13:33:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "pr-synced-to-cloud"
          ],
          "author": "dtunikov",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e7270bb0d7ecd3ff92c1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111332",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111332",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix RIGHT JOIN with parallel_replicas_min_number_of_rows_per_replica",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111206 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NOT_FOUND_COLUMN_IN_BLOCK` / `THERE_IS_NO_COLUMN` errors (and, on some releases, silently wrong results) for a `RIGHT JOIN` when `parallel_replicas_min_number_of_rows_per_replica` is set and a left-table column is projected. ### Description Closes: #111206 With `parallel_replicas_min_number_of_rows_per_replica` > 0, a `RIGHT JOIN` that projects a left-table column failed: ```sql SELECT r.ver, (l.a + 2) FROM tl AS l RIGHT JOIN tr AS r USING (k); -- Code: 10. Column `a` not found in table default.tr (NOT_FOUND_COLUMN_IN_BLOCK) ``` It also surfaced as `Code: 8 THERE_IS_NO_COLUMN` when aggregating a right-table column, through `RIGHT ANTI JOIN`, and as a silently wrong result on some releases. Leaving the setting at 0, and `LEFT` / `INNER` joins, were unaffected. Root cause: the initiator runs index analysis on the leftmost leaf to estimate the replica count and hands that scan to `createLocalPlanForParallelReplicas`. For a `RIGHT JOIN` the local plan parallelizes the right table (`findReadingSteps` descends the right child), so the leftmost leaf's parts and column list were applied to the right table's read, requesting the left table's columns from the right table. Fix: reuse the pre-analyzed result only when the parallelized scan was not reached through a `RIGHT JOIN` right-branch descent; otherwise let that scan analyze itself, which is what already happens when no analysis is passed. This is also correct for self-joins, where the two sides share the same storage but are distinct occurrences. Note this is unrelated to `automatic_parallel_replicas_mode`, which is a separate feature: a non-zero mode forces `enable_parallel_replicas` to 0, so the two never apply at once. An earlier revision of this PR mislabelled the bug as \"automatic parallel replicas\"; thanks to @nickitat for catching it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111332",
          "createdAt": "2026-07-22T08:41:13Z",
          "updatedAt": "2026-08-13T13:33:36Z",
          "timestamp": "2026-08-13T13:33:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "nickitat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:43160479ea3a23592dcd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114274",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114274",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for replication before asserting a row count in test_rename_distributed",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `test_rename_column/test.py::test_rename_distributed` asserts an exact row count right after inserting into a `Distributed` table, with no barrier - only a 30 second poll in `select()`. Two things must finish first, and neither is awaited: 1. The INSERT is background-spooled: `insert_sync` at `StorageDistributed.cpp:1143` is false here (`distributed_foreground_insert` defaults false, table is persistent), so `DistributedSink` takes `writeAsync` and the INSERT returns while a block is still on the initiator's disk. 2. `internal_replication` is true, so a shard's peer replica must FETCH every part (~100 tiny parts per INSERT per shard given `PARTITION BY num % 100`), and `load_balancing` defaults to `RANDOM`, so a shard leg can be served by a replica still fetching. Found by inspection plus a local repro; public CI has not lost this race yet. Forcing the assert onto its first read (`poll=0`) fails reproducibly with `assert '1473\\n' == '1998\\n'`. There, `system.distribution_queue` held 2 un-flushed files while all four `system.replication_queue`s were empty, so a bare `SYSTEM SYNC REPLICA` returns vacuously. The spool drains first. The fix adds a helper running `SYSTEM FLUSH DISTRIBUTED` then `SYSTEM SYNC REPLICA` on all four nodes, after each insert. The first insert needs it too: the spool replays stored query text naming the columns, so a block outliving the `num2 -> foo2` rename replays `INSERT ... (num, num2)` and fails permanently with `NO_SUCH_COLUMN_IN_TABLE`, reproduced separately. `poll=30` is unchanged; raising it only lengthens the bet. Validated with the same `poll=0` probe in both arms on one binary, so the barrier does the work and not the poll: 2/2 fail without it, 3/3 pass with it. Then 50/50 green with `poll=30` restored, whole file green. The sibling `test_rename_distributed_parallel_insert_and_select` also reads a `Distributed` table, but its `select()` calls pass no `expected_result`, so they assert nothing about counts and are left alone. The other eight tests read plain `ReplicatedMergeTree`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114274",
          "createdAt": "2026-08-11T05:45:22Z",
          "updatedAt": "2026-08-13T13:33:32Z",
          "timestamp": "2026-08-13T13:33:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5a352c8f5c33eaaff4d6",
        "signalId": "github:ClickHouse/ClickHouse:issue:109308",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:109308",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Implement serde for `WindowStep` and support it for `makeDistributed`",
          "text": "We have basic infrastructure for `QueryPlan` serialization and deserialization: https://github.com/ClickHouse/clickhouse/blob/master/src/Processors/QueryPlan/QueryPlan.h#L101-L104. It's implemented for multiple steps, for example `GROUP BY`: https://github.com/ClickHouse/clickhouse/blob/master/src/Processors/QueryPlan/AggregatingStep.cpp#L945-L1105. However implementation for `WindowStep` is missing https://github.com/ClickHouse/clickhouse/blob/master/src/Processors/QueryPlan/WindowStep.h#L11. It makes it impossible to run distributed queries with `make_distributed_plan=1` if they contain any window functions. So the task is to support `WindowStep`. <!-- ch-version-info:start --> ### Version info - Resolved by: #109802 - Merged into: `26.8.1.615` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/109308",
          "createdAt": "2026-07-03T15:00:14Z",
          "updatedAt": "2026-08-13T13:33:28Z",
          "timestamp": "2026-08-13T13:33:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "unfinished code"
          ],
          "author": "alesapin",
          "state": "closed",
          "assignees": [
            "davenger"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:18fd4dca7caaa5838465",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113076",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113076",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Populate query_log views column for inserts through Alias",
          "text": "Share query access info with the forwarded target insert so materialized views triggered on the target are recorded in the `views` column of `system.query_log`. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `views` column in `query_log` for inserts to an `Alias` table whose target triggers materialized views, which was always empty.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113076",
          "createdAt": "2026-08-03T10:16:33Z",
          "updatedAt": "2026-08-13T13:33:17Z",
          "timestamp": "2026-08-13T13:33:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "eclbg",
          "state": "open",
          "assignees": [
            "scanhex12"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:7dc0ef1208c7185d4e36",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113333",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113333",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Skip the aggregation hash-table stats cache key without GROUP BY keys",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/113037 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a performance regression where an aggregation without `GROUP BY` keys serialized its whole input plan subtree, including the full contents of constant-folded literals, on every execution to compute a hash-table-size cache key it can never use. A query such as `SELECT quantileMerge(arrayJoin(arrayMap(x -> state, range(5000000))))` spent about 63 ms per execution in query-plan optimization and now spends 0.15 ms. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/113037 Related: #107643 (introduced the pass) `setAggregationHashTableCacheKeys` admitted every `AggregatingStep` and serialized its input subtree to derive a cache key. `ActionsDAG::serialize` writes the full binary content of constant columns, so a folded `range(5000000)` literal materialized about 40 MB into a buffer on every execution. A keyless aggregation cannot use that key. `AggregatedDataVariants::chooseMethod` starts with `if (keys_size == 0) return Type::without_key;`, `init` discards the `size_hint` for `without_key`, and `without_key` is absent from `APPLY_FOR_VARIANTS_CONVERTIBLE_TO_TWO_LEVEL`. Such aggregations also wrote a meaningless entry (`median_size` = 1) into the shared statistics cache, evicting useful ones. The fix admits an `AggregatingStep` only when `!getParams().keys.empty()`. The early return when no aggregation is present, the per-node `try/catch` and the compute-all-then-stamp-all ordering are unchanged, and the key stays an optional flagged field in `AggregatingStep::serialize`, so no format or setting default changes. Validated on x86_64 (the issue reports aarch64), release, against pristine master with distinct build IDs. `QueryPlanOptimizeMicroseconds` for the reported query, median of 9: 63000 to 154. Keyed aggregations keep preallocating, and the existing keyed tests are the regression guard here; a keyless assertion would need a process-global asynchronous metric or a timing counter, both flaky. <details><summary>Validation matrix</summary> - `04509_hash_table_sizes_stats_table_functions` still reports 650000 and 520000 on both branches. - `HashTableStatsCacheEntries` grows by 12 for 12 distinct keyed aggregations on both branches; for 12 keyless ones it grows by 12 on master and by 0 with this change. - `hash_table_sizes_stats`, `hash_table_sizes_stats_small` and `group_by_consecutive_keys` queries: ratios 0.93 to 1.02. - Full `aggregat`/`group_by` stateless suite, 460 tests: identical per-test verdicts. - 50 runs each of `04506` and `04507`: 100 OK, 0 FAIL on both branches. - Wall clock for the reported query, min of 15: 275 ms to 214 ms. </details> `optimizeJoin` and `considerEnablingParallelReplicas` reuse these subtree hashes and run after this pass, so with a keyless aggregation under a join their key values shift: the flag bit recording whether a key was stamped now differs. `HashTablesStatistics` is process-local, so the cost is one un-preallocated execution. Constant serialization is still costly for keyed aggregations, about 62 ms here, and is left to a separate fix. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1275` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113333",
          "createdAt": "2026-08-04T14:04:48Z",
          "updatedAt": "2026-08-13T13:32:43Z",
          "timestamp": "2026-08-13T13:32:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-performance",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "nickitat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:98003d181f71803ce2e0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107305",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107305",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make ALTER MODIFY COLUMN on named Tuple metadata-only when adding subfields",
          "text": "`ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` that only adds new subfields is now metadata-only (no mutation). Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1` (default `false`). Subfield additions through `Array`/`Map`/nested `Tuple` wrappers are also handled. `Nullable(Tuple(...))` is blocked (null map incompatibility). Removing/renaming subfields or changing types still triggers a mutation. Key/index/projection guards reject the metadata-only path when `primary.idx` or skip-index bytes would become invalid (whole tuple in key, or subcolumn whose type changes). Subcolumn references with unchanged types (e.g. `ORDER BY t.a` when only `t.c` is added) are allowed. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` is now metadata-only when only adding subfields, matching the speed of top-level `ADD COLUMN`. Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1`. ### Documentation entry for user-facing changes - [x] Documentation is not required (behavioral improvement; semantics unchanged)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107305",
          "createdAt": "2026-06-12T07:42:58Z",
          "updatedAt": "2026-08-13T13:31:51Z",
          "timestamp": "2026-08-13T13:31:51Z",
          "metrics": {
            "reactions": 1,
            "comments": 7
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3d2a81e89e08462b2df7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:99495",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:99495",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `GradualResizeProcessor` to limit effective parallelism for GROUP BY on small data volumes",
          "text": "When ClickHouse processes GROUP BY, it often overestimates the number of threads needed. With `max_threads = 64` but only a few thousand rows, all 64 `AggregatingTransform` instances get data, produce 64 partial hash tables, and the merge phase has to combine all of them — most nearly empty. This wastes time on merging overhead, which is especially noticeable for heavy aggregate states such as `uniq`, `uniqExact`, `groupArray`, etc. The new `GradualResizeProcessor` starts by pushing data to a single output port (or one port per split group when `min_outstreams_per_resize_after_split` applies), and activates all aggregation streams at once as soon as the configured row or byte threshold is crossed. For small datasets, only one aggregating thread receives data (or one per split group); for large datasets, all threads are used as before. New settings: - `min_rows_per_stream_for_gradual_resize` (default: `1000`) - `min_bytes_per_stream_for_gradual_resize` (default: `0`) When either threshold is non-zero, the pre-aggregation `StrictResize` is replaced with `GradualResize` in the pipeline. The optimization is enabled by default; set both `min_rows_per_stream_for_gradual_resize = 0` and `min_bytes_per_stream_for_gradual_resize = 0` to opt out. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve performance of GROUP BY on small data volumes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/99495",
          "createdAt": "2026-03-14T07:35:33Z",
          "updatedAt": "2026-08-13T13:31:40Z",
          "timestamp": "2026-08-13T13:31:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 30
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:8e398a54ad3d536572b6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114642",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114642",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: document filtering for `system.query_log` initial queries",
          "text": "Document the recommended filters for analyzing `system.query_log`. The page now explains that `is_initial_query = 1` selects top-level client queries, while `initial_query_id` correlates the full cascade of a distributed query across nodes. The existing basic example also filters for initial queries so child executions are not treated as separate client queries. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114642",
          "createdAt": "2026-08-13T13:10:53Z",
          "updatedAt": "2026-08-13T13:31:35Z",
          "timestamp": "2026-08-13T13:31:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f730a6c60e2af2635b1e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109891",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109891",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reintroduce borrowed threadgroup async uaf fix",
          "text": "Reintroduce #108988 Related: https://github.com/ClickHouse/ClickHouse/pull/107030 Related: https://github.com/ClickHouse/ClickHouse/pull/108577 CI: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105890&sha=3dc0e76362eb18e27f1fffcd3f61ca9f13725fe8&name_0=PR&name_1=Stateless%20tests%20%28amd_tsan%2C%20s3%20storage%2C%20sequential%2C%201%2F2%29 Fixes a use-after-free risk in asynchronous work scheduled while a borrowed `ThreadGroup` is current. Borrowed `ThreadGroup` objects used by materialized view and async insert flush paths point their `performance_counters` and `memory_tracker` to the parent query group, so they are valid only while that parent group is alive. Async callbacks could capture such a borrowed group and later attach it on a pool thread after the parent query group had finished. Instead of keeping the parent `ThreadGroup` alive with a `shared_ptr`, this change keeps borrowed accounting scoped. Borrowed groups are marked explicitly, async callback capture drops borrowed groups, and thread pool callback runners capture the normalized group at task enqueue time rather than when a potentially long-lived runner object is created. This preserves synchronous borrowed accounting, but async work started from a borrowed scope runs under normal thread/global accounting instead of writing into, or prolonging the lifetime of, an already finished query group. Full ASAN reports https://gist.github.com/filimonov/1ec59047c65e3a5367c5c83f6021cc27 Compared to #108988 - added one commit with code comments + fix of the test failure https://github.com/ClickHouse/ClickHouse/issues/109841 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a memory safety issue where asynchronous work scheduled from materialized view processing could keep using query-level accounting after the query had finished.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109891",
          "createdAt": "2026-07-09T12:34:20Z",
          "updatedAt": "2026-08-13T13:30:49Z",
          "timestamp": "2026-08-13T13:30:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 47
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "comp-query-execution"
          ],
          "author": "filimonov",
          "state": "open",
          "assignees": [
            "azat",
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:4cd8b6d37031607fb519",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111427",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111427",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Refactor columns deserialization to avoid all assumeMutable calls, simplify the substreams cache, and remove rows_offset",
          "text": "PR #109105 replaced `assumeMutable` with `IColumn::mutate` in the some pars of the deserialization code. When a column is shared via the substreams cache, `IColumn::mutate` clones the whole accumulated column, giving O(rows × granules) cost and a 3–4× deserialization slowdown. Rework the deserialize path so it never clones: - `deserializeBinaryBulkWithMultipleStreams` and its helpers take `IColumn &` instead of `ColumnPtr &`. - The substreams cache never shares a `ColumnPtr`; consumers copy the current range out via `insertRangeFrom`. - `assumeMutable`/`IColumn::mutate`/`const_cast` are removed from the deserialize paths; mutable children come from non-const accessors. - Cache-shared members are dropped from the `Deserialize*State` structs. - The MergeTree reader stack reads into `MutableColumns`. Additionally, remove `rows_offset` from the deserialization API. `rows_offset` (the number of leading rows to skip while deserializing a range) is provably always `0` for every caller, so it is dropped from `deserializeBinaryBulkWithMultipleStreams`, `deserializeBinaryBulk` and all their helpers. This deletes the read-side skip machinery: - the skip loops / `istr.ignore(size * rows_offset)` seeks; - the `Map` bucketed reorder-with-skipped-rows path (which also removes a latent substreams-cache over-count on bucketed `Map`); - the `Object` shared-data `granules_offsets` / `last_incomplete_granule_offset` cursors (the load-bearing `StructureGranule` continuous-read state is kept); - the `SerializationArrayOffsets` helper (`rows_offset`-only); - assorted dead offset arithmetic. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Refactor columns deserialization to avoid all assumeMutable calls, simplify the substreams cache, and remove the always-zero rows_offset parameter.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111427",
          "createdAt": "2026-07-22T16:39:29Z",
          "updatedAt": "2026-08-13T13:30:46Z",
          "timestamp": "2026-08-13T13:30:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "Avogar",
          "state": "open",
          "assignees": [
            "KochetovNicolai"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:80f7f43305ea8d99f9db",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114596",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114596",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix target access checks for the Alias engine",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> `Alias` table now requires `SHOW COLUMNS` on its target when a new definition is submitted. This prevents users from using `Alias` to reveal a target table's schema. The trivial `count` optimization is disabled when the user has no `SELECT` privilege on the target, allowing the normal read path to enforce access checks. `rows` and `bytes` statistics are also hidden unless the user has `SHOW TABLES` on the target. Closes: https://github.com/ClickHouse/ClickHouse/issues/114511 ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an RBAC bypass in the `Alias` table engine that allowed users without privileges on the target table to reveal its schema, row count, size, and existence.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114596",
          "createdAt": "2026-08-13T05:30:42Z",
          "updatedAt": "2026-08-13T13:30:16Z",
          "timestamp": "2026-08-13T13:30:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-must-backport",
            "can be tested",
            "pr-critical-bugfix"
          ],
          "author": "nauu",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d760a20e253d8cd26230",
        "signalId": "github:ClickHouse/ClickHouse:issue:111211",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:111211",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "ORDER BY is not applied globally when reading a Distributed table through a Merge engine: per-shard-sorted streams are concatenated, LIMIT returns wrong rows",
          "text": "**Describe what's wrong** Reading a `Distributed` table through a `Merge` engine drops the global sort: `SELECT ... ORDER BY x` returns each shard's stream sorted individually and concatenated, not a globally ordered result. With `LIMIT n` this silently returns the wrong rows (the first shard's top-n instead of the global top-n). Querying the same `Distributed` table directly is correct. Aggregations over the `Merge` table (`count`, `max`) are correct — only the sorted path is wrong. **Does it reproduce on the most recent release?** Reproduces on `26.7.1.408` and near-HEAD master `3090a4fc` (`26.7.1.653`). **How to reproduce** Cluster `two_shards`: two shards on one server with `default_database` `sh0` / `sh1`. ```sql CREATE DATABASE sh0; CREATE DATABASE sh1; CREATE TABLE sh0.t (w Int64) ENGINE = MergeTree ORDER BY w; CREATE TABLE sh1.t (w Int64) ENGINE = MergeTree ORDER BY w; INSERT INTO sh0.t SELECT number * 3 FROM numbers(100); -- max 297 INSERT INTO sh1.t SELECT number * 3 + 1000 FROM numbers(100); -- max 1297 CREATE TABLE dist_t AS sh0.t ENGINE = Distributed(two_shards, '', t); CREATE TABLE merge_t AS dist_t ENGINE = Merge(currentDatabase(), '^dist_t$'); -- correct: 1297, 1294, 1291 SELECT w FROM dist_t ORDER BY w DESC LIMIT 3 SETTINGS enable_analyzer = 1; -- WRONG: 297, 294, 291 (shard 0's top-3; global maximum 1297 not returned) SELECT w FROM merge_t ORDER BY w DESC LIMIT 3 SETTINGS enable_analyzer = 1; -- also wrong WITHOUT LIMIT: 297, 294, 291, 288, ... (per-shard sorted streams -- concatenated; the result multiset is complete but the order is not global) SELECT w FROM merge_t ORDER BY w DESC SETTINGS enable_analyzer = 1; ``` **Characterization** - `max_threads = 1` (no thread-interleave explanation); deterministic. - Independent of `optimize_read_in_order` and `optimize_skip_unused_shards`. - `count()` / `max()` over `merge_t` are correct (500 rows / global max) — the row set is complete; only the final sort is missing on the `Merge`-over-`Distributed` read path, so any `ORDER BY ... LIMIT` silently returns wrong rows. Found by the optimizer-tester topology differential (`tests/optimizer_tester`, branch `optimizer-tester-framework`) the first run after a `Merge`-over-`Distributed` arm was added; detected by the order-sensitive comparison for total `ORDER BY` queries. Related: https://github.com/ClickHouse/ClickHouse/issues/111206",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/111211",
          "createdAt": "2026-07-21T10:51:36Z",
          "updatedAt": "2026-08-13T13:29:45Z",
          "timestamp": "2026-08-13T13:29:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "bug",
            "comp-distributed",
            "comp-storage-merge"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:c906fa21c33d27320ae8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86353",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cascades cost-based optimizer for distributed query plans",
          "text": "A Cascades-style cost-based optimizer that chooses distribution strategies for the multi-stage distributed query plans of #106020. It explores alternatives in a memo (a shared store of equivalent plan fragments) with top-down, goal-directed search and picks the cheapest plan satisfying the required distribution and sorting properties, inserting exchange operators (plan steps that move rows between nodes) as needed. Implemented: - **Join strategies**: shuffle hash join, broadcast hash join (with `ReplicatedRead` — every worker repeats the same read of a small table instead of a network broadcast, assuming shared storage where all workers see the same data), replicated join (a small deterministic join is recomputed identically on every node over such reads, so its result never crosses the network; nested joins compose), local join. - **Aggregation strategies**: two-phase (partial + merge), shuffle by group keys, local; `distributed_aggregation_memory_efficient` and `distributed_plan_force_shuffle_aggregation` are honored. - **Top-N**: two-stage distributed top-N (per-node bounded sort, sorted-merge gather, coordinator limit); disabled under `exact_rows_before_limit`, which needs the full row count. - **Read strategies**: parallel N-way read, replicated read, local read. For `FINAL`, #108148 (already in master) taught the rule-based distributed plan to split a `FINAL` read into disjoint primary-key-range buckets where that is safe; Cascades now reuses that machinery, so `FINAL` no longer forces a serial read here either. The coordinator ships each bucket's marks in the `read_bucket` task parameters. - **Properties and enforcers**: distribution (node count, replication, partitioning columns with equivalence classes and the types the keys are cast to before hashing) and sorting; when a plan alternative lacks a required property, an enforcer inserts the step that provides it (`Gather`/`Shuffle`/`Broadcast`/`ScatterExchange`, `Sort`). - **Transformations**: join commutativity (only for semantics-preserving joins: `INNER ALL`, `CROSS`, `SEMI`/`ANY`/`ANTI`; never `ASOF`), two-phase aggregation split, two-stage top-N split. - **Cost model**: `work`, `network`, and `sequential` components, a fixed per-exchange overhead, broadcast costed per receiving node, and statistics clamped to join kind and strictness semantics; exchange costs use per-row byte widths measured from the parts' column sizes (followed through renames, not derived from types); all weights and calibration constants are overridable at query time. What this improves over the rule-based distributed planner, on TPC-H plans. The rule-based planner broadcasts a small table when its read is below `distributed_plan_max_rows_to_broadcast`, but it often cannot size the result of a join, so a small join result (`nation x region`, 5 rows after the region filter) is scattered across nodes, joined there, and shuffled again (repartitioned across nodes) by the next join key. It also often inserts a shuffle at join and aggregation boundaries even when the rows are already divided by the right key. Cascades estimates sizes through joins, knows which partitioning already holds, compares broadcast against shuffle by cost for each join, and recomputes a small deterministic join on every node when that is cheaper than moving its result. On TPC-H SF100 over 8 nodes (same binary, same run window; times are server-side averages of the hot runs, Cascades averaged over two runs that produced identical plans) the join-heavy queries improve: | Query | Rule-based -> Cascades | What changed in the plan | |---|---|---| | Q18 | 12.03 s -> 6.94 s | Fewer shuffles (8 -> 5): the aggregated `lineitem` subquery joins without a re-shuffle (`l_orderkey = o_orderkey`), and the final GROUP BY reuses the existing partitioning. | | Q21 | 8.02 s -> 5.54 s | 34 exchanges -> 14: all four `supplier x nation` joins are recomputed by every node over full local reads, so nothing is gathered or broadcast for them. | | Q09 | 4.19 s -> 2.25 s | 11 exchanges -> 6: `part` and `nation` are read in full by every node, so `lineitem` and `supplier` are not shuffled to meet them. | | Q08 | 2.32 s -> 1.00 s | 15 exchanges -> 7: `nation x region` (5 rows) is recomputed by every node; `part` is read in full per node, so `lineitem` is not shuffled to join it. | | Q02 | 2.24 s -> 0.93 s | 16 exchanges -> 5: the `supplier x nation x region` chain is recomputed by every node, so only `partsupp` and `supplier` are shuffled. The top-100 sort becomes two-stage, sending at most 100 rows per node. | | Q17 | 2.77 s -> 1.87 s | The small per-part average is broadcast to every node, so the outer 600M-row `lineitem` read is not shuffled. | | Q11 | 0.49 s -> 0.23 s | 5 exchanges -> 2: `supplier x nation` is recomputed by every node and `partsupp` joins it in place. | The remaining queries change less. Summed over all 22 queries, hot server time drops from 46.0 s to 32.1 s (about 30% lower). The largest regression is Q15 (0.22 s -> 0.39 s): the extra time is initiator-side planning — the memo search and statistics loading currently repeat for the view subqueries; fixing that is a follow-up. `EXPLAIN pretty = 1, estimates = 1` shows the chosen plan with a row estimate and the accumulated cost for each step. Also in this PR, two improvements to the shared bucketed-read machinery (they benefit the rule-based path too): the `FINAL` layer split no longer depends on the coordinator's core count, and a many-partition `FINAL` split groups its layers into the target task count instead of falling back to a serial read. Design, a worked example on TPC-H data (a simplified 3-table query traced through the memo), and current limitations are documented in `src/Processors/QueryPlan/Optimizations/Cascades/ARCHITECTURE.md`. Plan-shape tests cover the actual TPC-H queries (`03836_tpch_join_order_plans`). Disabled by default. Requires the analyzer; remote execution requires the stateless-worker configuration, while `distributed_plan_execute_locally = 1` runs the stages in-process without it: ```sql SET enable_cascades_optimizer = 1, make_distributed_plan = 1; ``` For tests, `param__internal_cascades_cluster_node_count` overrides the cluster size, `param__internal_cascades_cost_config` overrides the cost model configuration, `param__internal_join_table_stat_hints` injects table statistics, `param__internal_cascades_task_limit` lowers the task budget (it can never raise it). Related: https://github.com/ClickHouse/ClickHouse/pull/106020 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental Cascades cost-based optimizer for distributed query plans, enabled by `enable_cascades_optimizer = 1` together with `make_distributed_plan = 1`. It chooses between shuffle, broadcast, replicated, and local join strategies, two-phase, shuffle, and local aggregation, two-stage distributed top-N, and parallel and replicated reads by estimated cost, inserting exchange operators as needed. ### Documentation entry for user-facing changes - [ ] Documentation written in [/docs](https://github.com/ClickHouse/ClickHouse/tree/master/docs)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86353",
          "createdAt": "2025-08-28T11:27:09Z",
          "updatedAt": "2026-08-13T13:29:43Z",
          "timestamp": "2026-08-13T13:29:43Z",
          "metrics": {
            "reactions": 21,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:993cd701784e8244b308",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110838",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110838",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add introspection TCP port",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ClickHouse server now has an introspection port. This is a native protocol TCP listener that starts before the server begins attaching tables and stops only after the tables' detach completes. During these windows, an operator can connect to it with `clickhouse client` and run queries such as `SHOW PROCESSLIST`, `SELECT * FROM system.stack_trace`, or `SYSTEM INSTRUMENT ADD 'QueryMetricLog::startQuery' SLEEP ENTRY 0.5`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110838",
          "createdAt": "2026-07-17T11:07:23Z",
          "updatedAt": "2026-08-13T13:29:35Z",
          "timestamp": "2026-08-13T13:29:35Z",
          "metrics": {
            "reactions": 1,
            "comments": 16
          },
          "labels": [
            "pr-feature"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "evillique"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ea458ef03826a04606a9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111464",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111464",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix ORDER BY not applied globally when reading Distributed through Merge",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/111211 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results for a global `ORDER BY` (and `DISTINCT`) when reading a `Distributed` table through a `Merge` engine: the per-shard sorted streams were concatenated instead of merge-sorted, so `LIMIT` could return rows local to one shard instead of the global top rows, and `DISTINCT` could return a doubled result multiset. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/111211 `ReadFromMerge::initializePipeline` narrows the united child pipeline with `narrowPipe`, which concatenates streams via `ConcatProcessor` and does not preserve per-stream order. When a child is read at a partial stage that emits already-sorted streams (e.g. a `Distributed` child produces one sorted stream per shard), the outer plan adds a merge-only sorting step that assumes each input stream is individually sorted. With `max_threads = 1` the two sorted shard streams were concatenated into one, the merge-only sort then had a single unsorted input and became a no-op, and `LIMIT` returned one shard's local top rows. The `should_not_narrow` guard already suppresses narrowing for the two other order-sensitive cases (read-in-order and memory-efficient distributed aggregation). This adds the third: a global `ORDER BY` (no aggregation, or any after-aggregation stage) over a child read above `FetchColumns`. The condition mirrors when the planner/interpreter choose a merge-only sort (`Planner.cpp` `isFromAggregationState` / `InterpreterSelectQuery.cpp` `from_aggregation_stage`). Reproducer (before: `297, 294, 291`; after: `1297, 1294, 1291`), see #111211. The same `narrowPipe` concatenation also corrupts the result multiset under `DISTINCT` (with `optimize_distinct_in_order`, the second shard's sorted run survives adjacent-only dedup, doubling the rows). That query shape is already covered by the guard added here, so no extra code change was needed; the regression test now also asserts the `DISTINCT` / `DISTINCT ON` row counts (reported by @ zlareb1 on #111211).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111464",
          "createdAt": "2026-07-22T18:56:16Z",
          "updatedAt": "2026-08-13T13:29:24Z",
          "timestamp": "2026-08-13T13:29:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 15
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:107815bd42d6f0436eea",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114220",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114220",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.6: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114220",
          "createdAt": "2026-08-10T20:07:09Z",
          "updatedAt": "2026-08-13T13:28:22Z",
          "timestamp": "2026-08-13T13:28:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d66121e0eadcd1468a3b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114328",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114328",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enhance MetadataStorageFromMemory",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Pure refactoring change, doesn't affect any working part of code.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114328",
          "createdAt": "2026-08-11T13:44:39Z",
          "updatedAt": "2026-08-13T13:28:21Z",
          "timestamp": "2026-08-13T13:28:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-object-storage-disks"
          ],
          "author": "alesapin",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:33c69ce2109dcc47ffc3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111287",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111287",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix double free when finalizing -State aggregates under looping combinators",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server crash (double free) that could happen when finalizing an aggregate function with the `-State` combinator nested under a looping combinator (`-Resample`, `-ForEach`, `-Map`), for example `groupArrayStateResample`, if a memory limit was reached during finalization. ### Description Reported on https://github.com/ClickHouse/ClickHouse/pull/110975 (unrelated to that PR). Found by the Stress test (amd_debug): a segfault in `Aggregator::prepareChunkAndFillWithoutKey`, reached from `ConvertingAggregatedToChunksTransform::initialize`. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110975&sha=838d0b61235b06939c8acd923ebd396f51cf5b10&name_0=PR&name_1=Stress%20test%20%28amd_debug%29 Root cause: the `-State` combinator transfers its result by aliasing the raw aggregate state pointer into a `ColumnAggregateFunction` (`AggregateFunctionState::insertResultInto` -> `getData().push_back(place)`); ownership passes to the column. `Aggregator::insertAggregatesIntoColumns` relies on this transfer being atomic per place: on an exception it destroys the whole place exactly once. A looping combinator nested over `-State` aliases many sub-states one at a time; the `push_back` into the column's pointer array can reallocate and, being memory-tracked, throw `MEMORY_LIMIT_EXCEEDED` mid-loop. The already-transferred sub-states are then freed once by the aggregator's full `destroy()` and again by `~ColumnAggregateFunction`, i.e. a double free. Reproducer (crashes without the fix, returns a memory-limit error with it): ```sql SELECT arrayMap(x -> finalizeAggregation(x), state) FROM (SELECT groupArrayStateResample(0, 1048576, 1)(number, number % 20) AS state FROM numbers(100000)) SETTINGS max_memory_usage = 150000000, max_rows_to_read = 0; ``` Fix: reserve the destination columns before the transfer loop so the aliasing `push_back`s cannot reallocate (and therefore cannot throw) once a transfer has started. `ColumnAggregateFunction` used the no-op `IColumn::reserve`, so a real `reserve()`/`capacity()` over its state-pointer array is added. For `-Map`, the (possibly variable-width) key inserts are moved into their own loop before the value transfer, keeping the throwing work out of the aliasing loop. Reserving happens before any aliasing, so a throw there is harmless. The transfer loop is now non-throwing at the point of aliasing, restoring the atomic-per-place contract; results are unchanged. The fix covers all three looping transfer combinators (`-Resample`, `-ForEach`, `-Map`), which share the aliasing path; non-looping combinators delegate a single call and are already atomic. The added stateless test reproduces the crash deterministically via `-Resample` (empty buckets keep memory low until the finalization transfer, so a memory limit reliably lands the throw mid-transfer). `-ForEach` and `-Map` build their sub-states eagerly during aggregation, so they are not deterministically reproducible under a memory limit, but are fixed as the same class via the shared transfer path.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111287",
          "createdAt": "2026-07-21T20:34:03Z",
          "updatedAt": "2026-08-13T13:28:04Z",
          "timestamp": "2026-08-13T13:28:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:df82f752024d001d6e7e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114628",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114628",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for DETACH DATABASE to release tables in three integration tests",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/93064 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/issues/93064 Three integration tests intermittently fail with `Code: 219 ... Database <db> cannot be detached, because some tables are still in use. Retry later.` thrown from `DatabaseAtomic::assertCanBeDetached` (`src/Databases/DatabaseAtomic.cpp:521`). Root cause: non-SYNC `DETACH DATABASE` is best-effort by contract, and these tests assume it is atomic. `assertCanBeDetached` throws if any table of the database still has a live refcount. The engine already has the deterministic wait, `waitDetachedTableNotInUse`, but `executeToDatabaseImpl` calls it only under `query.sync` (`src/Interpreters/InterpreterDropQuery.cpp:779-790`). Stateless CI installs `tests/config/users.d/database_atomic_drop_detach_sync.xml` (`database_atomic_wait_for_drop_and_detach_synchronously=1`); the integration helpers install no such profile and run at the server default of 0. Hence the failures are confined to the integration suite, and 16 of the 17 integration occurrences in 180 days are sanitizer builds, where the holder's window is wider. Code 219 is intended behaviour of the asynchronous form and is pinned in-tree: `01107_atomic_db_detach_attach.sh:19` sets the setting to 0 to provoke it and asserts it. So this is a test-side defect, and the change asks for the wait the engine already implements rather than altering it. I added `SYNC` to the five `DETACH DATABASE` statements with observed CI failures: `test_drop_replica` (11 hits/180d), `test_replicated_table_structure_alter` (4), and `test_drop_database_replica:192` (2, both stacks confirm that line is the thrower). All 21 non-SYNC sites under `tests/integration/` were enumerated; the other 16 are excluded because they have zero CIDB hits in 180 days, already pass `database_atomic_wait_for_drop_and_detach_synchronously`, or sit inside the existing `detach_database_with_retry` helper, whose docstring documents a different holder. Validation: with a holder injected the way `01107` does it, the DETACH returns 219 without `SYNC` and succeeds with it, 10/10 on `Atomic` and on `Replicated`. The three tests pass 42/42 runs. Each `DETACH` in `test_drop_replica` measures ~0.115 s over 25 measurements, so the wait costs nothing when no table is held, and it is cancellable and shutdown-aware. CI report for the master failure: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=91b700711adf0a07172e1046c8d8f013f1e0b82c&name_0=MasterCI&name_1=Integration%20tests%20%28amd_tsan%2C%206%2F6%29",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114628",
          "createdAt": "2026-08-13T12:36:13Z",
          "updatedAt": "2026-08-13T13:25:55Z",
          "timestamp": "2026-08-13T13:25:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:e0f6289910a4affb4946",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114247",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114247",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Clear plain LIMIT/OFFSET in the window-view backfill source query",
          "text": "Follow-up to https://github.com/ClickHouse/ClickHouse/pull/113759: the last AI review finding on that PR landed after the PR had been added to the merge queue, so the branch could no longer be updated. This addresses it. `StorageWindowView::getSourceTableSelectQuery` builds the raw-source backfill query for `CREATE WINDOW VIEW ... POPULATE`. The rows it produces are inserted into the window view, where `writeIntoWindowView` executes the mergeable view query over them, so the backfill query must deliver the raw source rows and leave all row transformations of the original query to the view query — otherwise the initialized state diverges from live behavior. The PR originally made the helper clear the leftovers of the original `SELECT` that violated this invariant one by one — `LIMIT`/`OFFSET` (and `WITH TIES`), plain `DISTINCT`, `ARRAY JOIN`, `WHERE`/`PREWHERE`, table-expression `SAMPLE`/`FINAL` — on top of the `JOIN`/`GROUP BY`/`ORDER BY`/`LIMIT BY`/`WINDOW`/`QUALIFY`/`INTERPOLATE` handling it already had. Review then found that this strip-the-leftovers approach misses wrapped sources: the same constructs inside a `FROM (SELECT ...)` subquery or a CTE definition survived the rewrite, and covering them would have required recursing the rewrite into every nested select. So the helper now builds the backfill query from scratch instead of stripping a clone of the view query. The contract makes this valid: `writeIntoWindowView` always receives raw source-table blocks (`getInputHeader` is the source table header no matter how the view query wraps or transforms the table — `PushingToWindowViewSink` is created with exactly that header), so the correct backfill query is exactly `SELECT <source columns> FROM <source table>`, plus `ORDER BY` on the timestamp column so the watermark is initialized from the earliest record. Wrapped sources, joins, and every row-shaping clause are covered by construction because the user query is no longer cloned at all. The helper shrinks by ~100 lines. Note: the divergence is currently unobservable because `CREATE WINDOW VIEW ... POPULATE` fails before writing any rows (https://github.com/ClickHouse/ClickHouse/issues/113493), so no regression test is possible yet; this keeps the rewrite invariant consistent for when `POPULATE` is fixed. Verified with `clickhouse-local` probes that `POPULATE` over wrapped-subquery/CTE/`JOIN`/`WHERE`/`FINAL` sources now analyzes the backfill query cleanly and proceeds to the pre-existing #113493 sink error. Related: https://github.com/ClickHouse/ClickHouse/pull/113759 Related: https://github.com/ClickHouse/ClickHouse/issues/113493 ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114247",
          "createdAt": "2026-08-10T23:40:02Z",
          "updatedAt": "2026-08-13T13:25:40Z",
          "timestamp": "2026-08-13T13:25:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8bb08d5492a771a78b5f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110613",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110613",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add experimental PCO compression codec (linking the pcodec Rust crate)",
          "text": "Adds an experimental `PCO` compression codec that links the [pcodec](https://github.com/pcodec/pcodec) (`pco`) Rust crate — the reference implementation — rather than reimplementing it. `pco` is a lossless codec specialized for sequences of fixed-width numbers; on numeric columns with smooth, multimodal, or high-entropy distributions (measurements, timings, counters, identifiers) it often beats `Gorilla`/`FPC`/`ZSTD`/`ALP` on ratio. This is the Rust-library counterpart to the native C++ port explored in https://github.com/ClickHouse/ClickHouse/pull/106222. Instead of maintaining a C++ reimplementation, it links the upstream crate. ### Runtime CPU dispatch `pco` has no explicit SIMD: it relies on the compiler autovectorizing its per-batch loops, which only happens when the crate is compiled with the relevant instruction sets enabled (its `build.rs` warns to build with `-C target-feature=+avx2,+bmi1,+bmi2`). That does not work for ClickHouse, which compiles all Rust once at a fixed baseline micro-architecture level for a binary that must run on a range of CPUs. So `pco` is taken from a ClickHouse fork patched to add runtime CPU-feature dispatch of its hot loops — the same mechanism ClickHouse's C++ uses via `TargetSpecific.h`. The vectorizable loop bodies (decode `read_offsets`; encode `write_short_uints`/`write_uints`/`set_offsets`) are compiled at the crate baseline and again for `x86-64-v3` (AVX2 + BMI1/2 + FMA) and `x86-64-v4` (AVX-512), and the widest variant the running CPU supports is selected once and cached. On aarch64 the baseline already includes NEON, so the baseline body is used directly. Verified at baseline `-C target-cpu=x86-64`: the `read_offsets` v3 trampoline emits AVX2 (`ymm`) and v4 emits AVX-512 (`zmm`). The patch is merged into `ClickHouse/pcodec` (a fork of pcodec/pcodec in the ClickHouse organization) via its own PR: https://github.com/ClickHouse/pcodec/pull/1. The crate is added as the `contrib/pcodec` submodule and linked through a thin FFI wrapper crate (`rust/workspace/pco`). Its one not-yet-vendored dependency, `rand_xoshiro`, was added to `contrib/rust_vendor` in https://github.com/ClickHouse/rust_vendor/pull/72. ### Codec Supported types are all fixed-width numerics of 1/2/4/8 bytes via their underlying integer/float representation: `Int8`..`Int64`, `UInt8`..`UInt64`, `Float32`/`Float64`, and the types backed by them (`Date`, `DateTime`, `Decimal32`/`Decimal64`, `IPv4`, `Enum`, ...). The on-disk block stores a 2-byte header (element width + partial-tail byte count) followed by a raw partial-value tail (as in `Gorilla`/`FPC`) and the payload. The payload is a standalone `.pco` stream, wire-compatible with the reference pcodec implementation; when compression would not shrink a block the raw bytes are stored instead (a per-block \"stored\" flag), so the output never expands by more than the 2-byte header and `getMaxCompressedDataSize` is tight. The codec is gated behind `allow_experimental_codecs`. Because it needs the column type, it can only be specified per column: it is rejected in `TTL ... RECOMPRESS` and in the untyped compression settings that resolve codecs without a type. `CompressionCodecMultiple` propagates the experimental / column-type-requiring properties so a chain such as `CODEC(Delta, PCO)` is still gated, and a codec-only `ALTER TABLE ... MODIFY COLUMN x CODEC(PCO)` validates against the existing column type. Tests: `04512_pco_codec` round-trips every supported numeric and backed type (per-element verification, edge cases, codec chaining, compression-ratio check) and `04513_pco_codec_gating` covers the experimental gate, the `TTL RECOMPRESS` rejection, and the codec-only `ALTER`. The FFI wrapper has its own Rust unit tests (round-trip of all types, the no-expansion fallback, and fail-closed handling of malformed/mismatched streams). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new experimental compression codec `PCO`, which links the [pcodec](https://github.com/pcodec/pcodec) library (patched for runtime CPU dispatch), specialized for fixed-width numeric columns. It is wire-format compatible with pcodec `.pco` streams and is enabled with `allow_experimental_codecs`. ### Documentation entry for user-facing changes: - [x] Documentation is written (the `PCO` codec is documented in `docs/en/sql-reference/statements/create/table.md`, and the compression-frame method byte in `docs/en/interfaces/specs/NativeFormat.md`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110613",
          "createdAt": "2026-07-15T19:32:58Z",
          "updatedAt": "2026-08-13T13:25:30Z",
          "timestamp": "2026-08-13T13:25:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 15
          },
          "labels": [
            "submodule changed",
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:78e6cab3c4e3768fc00f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106734",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106734",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "More settings to randomize",
          "text": "Extends the stateless-test randomizers in `tests/clickhouse-test` with settings added since they were last swept, and pins the tests that were implicitly relying on the old defaults. `SettingsRandomizer`: * Broadens `query_plan_optimize_join_order_algorithm` to cover `dpsub` and `dphyp` (always with `greedy` kept in the fallback chain, since exhausting the chain throws `EXPERIMENTAL_FEATURE_ERROR`), and randomizes `query_plan_optimize_join_order_max_searched_plans`. * Adds ~35 further query-level settings (`LIMIT BY` / `FINAL` / lazy-`FINAL` / runtime-filter / text-index / statistics / regexp-compilation toggles and their numeric companions). * Adds `use_statistics_for_part_pruning` → `use_statistics` and `use_projection_index_in_read_pools` → `optimize_use_projection_filtering` to `conditional_settings`, so a child setting is not enabled while its parent is off. * Randomizes `function_implementation` over the values the server actually supports, probed at startup: it globally filters every `ImplementationSelector`-backed function and each has a different arch-tag set, so a forced value that some function lacks would raise `NO_SUITABLE_FUNCTION_IMPLEMENTATION`. * A few settings are deliberately left commented out with the reason recorded in place, e.g. `use_constant_folding_in_index_analysis` (unsafe per-part min-max prune, and the trigger for https://github.com/ClickHouse/ClickHouse/issues/109893), `correlated_subqueries_use_in_memory_buffer`, `min_filtered_ratio_for_lazy_final` (https://github.com/ClickHouse/ClickHouse/issues/112332) and `max_streams_for_union_step`. `MergeTreeSettingsRandomizer`: * Randomizes `concurrent_part_removal_threshold_for_remote_disk`. * The implicit min-max index settings (`add_minmax_index_for_block_number_column`, `add_minmax_index_for_block_offset_column`, `part_minmax_index_columns`) are explicitly kept out: they are not transparent to queries. The first two add real `MINMAX` skipping indices named `auto_minmax_index__block_number` / `auto_minmax_index__block_offset`, which surface in `SHOW INDEXES`, `system.data_skipping_indices`, part checksums and the `Skip` section of `EXPLAIN indexes = 1`; `part_minmax_index_columns = with_block_number_offset` gives every table a part-level min-max index even with no partition key, so `EXPLAIN indexes = 1` grows an extra `Min-Max / Condition: true` block everywhere. This is the same reason their siblings `add_minmax_index_for_{numeric,string,temporal}_columns` were never randomized — all five are `isReadonlySetting`, i.e. part of the table schema. The rest of the diff pins the affected settings in the stateless tests whose output depends on them. Related: https://github.com/ClickHouse/ClickHouse/pull/107586 (fixes the `Invalid binary search result in MergeTreeSetIndex` logical error this randomization made much more likely to be hit; merged, so it is pulled in by the branch update) Related: https://github.com/ClickHouse/ClickHouse/issues/109893 Related: https://github.com/ClickHouse/ClickHouse/issues/112332 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Running CI a few times before merging.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106734",
          "createdAt": "2026-06-08T16:44:21Z",
          "updatedAt": "2026-08-13T13:25:06Z",
          "timestamp": "2026-08-13T13:25:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-ci"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e5dad24b2e93ded91d59",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113909",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113909",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check the query cancellation while filling `system.parts` and its siblings",
          "text": "The tables based on `StorageSystemPartsBase` (`system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`) build the whole result eagerly in `initializePipeline`, so a cancelled or timed out query kept building rows over every storage and part until the very end. In a stress test with ThreadFuzzer this took minutes and tripped the hung check: a `SELECT` over `system.parts_columns` (with a filter matching every active part on the server) stayed in the process list for 252 seconds with `is_cancelled = 1` and `max_execution_time = 10`. Check the query status per storage and per part, following the pattern of `system.zookeeper` and `system.remote_data_paths`. The check also covers the storage-discovery prepass in `StoragesInfoStream` (the eager enumeration of all databases and tables), the analogous prepass in `StoragesDroppedInfoStream` (so `system.dropped_tables_parts` is interruptible as well), the skip loop in `StoragesInfoStreamBase::next`, and the per-storage part enumeration itself: the `MergeTreeData` helpers (`getDataPartsVectorForInternalUsage`, `getAllDataPartsVector`, `getProjectionPartsVectorForInternalUsage`, `getAllProjectionPartsVector`) take an optional `need_stop` callback that is checked periodically while the parts snapshot is being built. The two column-oriented tables additionally check the query status inside their column-enumeration loops and their metadata prepass, so a very wide table does not create a long uninterruptible stretch inside a single storage. The return value of `checkTimeLimit` is honored, so with `timeout_overflow_mode = 'break'` the eager build stops at the soft deadline and returns the rows collected so far. The table-lock acquisition in `StoragesInfoStreamBase::tryLockTable` is interruptible as well: instead of a single wait inside `RWLockImpl::getLock` for the whole `lock_acquire_timeout`, the lock is acquired in 100 ms slices with a query-status poll between the attempts (the total timeout and the `DEADLOCK_AVOIDED` semantics are preserved), so a killed or soft-timed-out query does not sit in the lock wait while a concurrent DDL query holds the drop lock. The test `04869_system_parts_lock_wait_cancellation` pins this with a share lock held by a long `SELECT` and a `DROP TABLE` in an `Ordinary` database queued behind it. The test uses the new `slowdown_system_parts_enumeration` failpoint, which only affects the specially named test tables (so concurrently running tests are unaffected). It sleeps 500 ms on every enumerated part, so building the full result for a 20-part table takes at least 10 seconds, and it sleeps 1 second per `COLUMNS_CANCELLATION_CHECK_PERIOD` (128) enumerated columns of a part, so building the full `system.parts_columns` / `system.projection_parts_columns` result over a single part with 1301 columns (and a projection over all of them) also takes at least 10 seconds. The test asserts that queries with a 1 second deadline in the `break` mode finish well under that, which is only possible by stopping at the per-part and per-column cancellation checkpoints. Timed assertions are needed because a plain row-count assertion cannot distinguish a build with the fix from one without: in the `break` mode the executor drops the eagerly built result after the deadline in both cases. The test also asserts partial row counts under a pre-expired deadline for all five tables, including `system.dropped_tables_parts` over a deterministic dropped-table fixture. The pre-expired-deadline checks also run under the failpoint, so the fewer-rows assertion is deterministic even on a machine fast enough to build the whole result in under a millisecond. For tables with the `_snap` name marker, the failpoint additionally slows down the parts-snapshot walks inside `MergeTreeData` (500 ms per enumerated part) and makes them poll the stop callback on every element, and for tables with the `_meta` name marker it slows down the column-metadata prepass of the column-oriented tables (1 second per 128 enumerated metadata columns), so the timed checks also prove that the snapshot materialization and the prepass are interruptible: all six of these checks fail against a binary with the `need_stop` polls and the prepass checkpoints disabled. Caught by `Stress test (arm_asan_ubsan, s3)`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113722&sha=ceedb772f6aa36e221d25ce48f9af71f032a7b08&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Related: https://github.com/ClickHouse/ClickHouse/pull/113722 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Queries over `system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`, and `system.dropped_tables_parts` now react to cancellation and `max_execution_time` while the result is being built.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113909",
          "createdAt": "2026-08-08T01:18:22Z",
          "updatedAt": "2026-08-13T13:24:22Z",
          "timestamp": "2026-08-13T13:24:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7c45b729ff47eccfaa74",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114457",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114457",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix header column order after the no-rescoring vector search rewrite",
          "text": "<!--- A technical description of your changes with a motivation --> The second-pass vector search optimization (`vector_search_with_rescoring = 0`) removed the distance function from the `ExpressionStep` outputs and re-appended the rewritten `_distance` alias at the end, changing the header column order. Steps created above the `Sorting` before that rewrite runs — the local top-N `Limit` and the exchange steps of a distributed plan — kept the original column order, and `makeDistributedPlan` failed to rebuild the plan fragments with a logical error: `Cannot add step Limit to QueryPlan because it has incompatible header with root step Sorting`. Non-distributed plans never re-validate step headers after optimization, which is why this only surfaced with `make_distributed_plan = 1`. The fix reinserts the rewritten output node at the original position of the distance column, so the header column order is preserved and all previously created parent steps stay consistent. Found by AST fuzzer on [#42701](https://github.com/ClickHouse/ClickHouse/pull/42701) ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=019244e0d46b4d9029fc261a549e1c5cd79b1bbb&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)); the same signature also hit the unrelated #98789 on 2026-08-08. The regression test asserts only the row count: the distributed plan currently returns an incorrect top-N for vector search queries regardless of the rescoring mode — a separate, pre-existing bug. Related: https://github.com/ClickHouse/ClickHouse/issues/114456 Related: https://github.com/ClickHouse/ClickHouse/pull/42701 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a logical error `Cannot add step Limit to QueryPlan because it has incompatible header with root step Sorting` when a vector search query with `vector_search_with_rescoring = 0` was executed with the experimental `make_distributed_plan` setting.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114457",
          "createdAt": "2026-08-12T09:43:36Z",
          "updatedAt": "2026-08-13T13:24:05Z",
          "timestamp": "2026-08-13T13:24:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0b61db5d34e9e97c464b",
        "signalId": "github:ClickHouse/ClickHouse:issue:113360",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113360",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Paimon primary-key tables silently return duplicate/stale rows: no merge-on-read, primary keys ignored",
          "text": "Reading a Paimon primary-key table returns duplicate/stale rows silently: the reader collects the raw union of base and delta data files with no merge-on-read, so superseded row versions from upserts are returned alongside current ones. Primary keys are parsed from the table schema (`PaimonSchemaProcessor::getPrimaryKeys`) and the LSM merge metadata is parsed from manifests (`_LEVEL`, `_MIN_SEQUENCE_NUMBER`, `_MAX_SEQUENCE_NUMBER` in `PaimonClient`), but none of it is used by the read path in `PaimonMetadata`. Since primary-key tables are the default choice for upsert/CDC workloads in Paimon, this produces silently wrong results on a very common table type — no exception, no warning. **How to reproduce** Write a primary-key table with an upsert (Paimon 1.1.1, Spark 3.5): ```sql CREATE TABLE paimon.default.pk_t (id INT, val STRING) TBLPROPERTIES ('primary-key'='id', 'bucket'='1', 'file.format'='parquet'); INSERT INTO paimon.default.pk_t VALUES (1, 'old'), (2, 'two'); INSERT INTO paimon.default.pk_t VALUES (1, 'new'); SELECT * FROM paimon.default.pk_t ORDER BY id; -- Spark (correct): 1 new -- 2 two ``` Read it with ClickHouse (master, commit `7c826d816cd5`): ``` $ clickhouse local --query \"SELECT * FROM paimonLocal('/tmp/paimon_pk_wh/default.db/pk_t') ORDER BY id, val\" 1 new 1 old <-- superseded row version returned silently 2 two ``` The same happens via `paimonS3` / `paimonAzure` and the `Paimon*` table engines. **Suggested behavior** Until merge-on-read is implemented, reading a table whose schema declares `primary-key` should throw an exception (fail-close), the same way other unsupported Paimon features already do (`ROW` type, unsupported `scan.mode` values). Silent wrong data is strictly worse than an error. Found while adding Paimon-over-object-storage integration tests (the suites deliberately avoid primary-key tables because of this). Related: https://github.com/ClickHouse/ClickHouse/issues/113337 Caused by: https://github.com/ClickHouse/ClickHouse/pull/102343",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113360",
          "createdAt": "2026-08-04T17:16:17Z",
          "updatedAt": "2026-08-13T13:22:19Z",
          "timestamp": "2026-08-13T13:22:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "bug",
            "unfinished code",
            "unexpected behaviour",
            "experimental feature",
            "comp-datalake"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "scanhex12"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:f447557ca893900662cb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114602",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114602",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: regenerate reference documentation from source",
          "text": "This pull request is opened automatically by the nightly documentation autogeneration workflow. It regenerates settings, functions, table and database engines, data types, formats, table functions, window functions, system tables, and asynchronous metrics from the structured documentation embedded in the ClickHouse source and exposed through the corresponding `system.*` tables. The generator preserves page frontmatter and hand-written content outside the `{/*AUTOGENERATED_START*/}` / `{/*AUTOGENERATED_END*/}` regions. Some pages are fully generated below their frontmatter. Do not edit generated content by hand -- edit the structured documentation in the defining source code instead; the next nightly run regenerates the pages. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114602",
          "createdAt": "2026-08-13T08:02:58Z",
          "updatedAt": "2026-08-13T13:21:42Z",
          "timestamp": "2026-08-13T13:21:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-autogenerated-docs"
          ],
          "author": "clickhouse-gh[bot]",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e5b092a07da54a784b8c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112327",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112327",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use-after-free on a sparse join key in a direct dictionary join",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109225 --> Related: https://github.com/ClickHouse/ClickHouse/pull/109225 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-after-free when a `JOIN` with `join_algorithm = 'direct'` onto a dictionary was given a join key that is stored with sparse serialization. `getColumnVectorData` returned a reference to a temporary column, so the dictionary lookup read freed memory: release builds could return wrong results and debug or sanitizer builds aborted. ### Description `getColumnVectorData` (`src/Dictionaries/DictionaryHelpers.h`) materializes its key column into a function-local `ColumnPtr` and then returns a `PaddedPODArray` reference **into that local**. It copied the data into the caller's `backup_storage` only when the input was `Const`. For a dense column that was still safe, because every conversion is a no-op returning `getPtr()` and the caller's own `ColumnPtr` keeps the buffer alive. It is not safe for a column that has to be materialized: `ColumnSparse::convertToFullColumnIfSparse` and `ColumnReplicated::convertToFullColumnIfReplicated` each allocate a **new** column that nothing else owns, so the returned reference dangles as soon as the function returns. The fix takes the copy whenever a conversion actually produced a different column (`full_column.get() != column.get()`). This is pointer identity rather than a type test, so it covers `Const`, `Sparse` and `ColumnReplicated` with one predicate, and it is fail-closed for any representation added later. The previous `Const` check is strictly subsumed: `ColumnConst::convertToFullColumn` returns either the inner column or a `replicate` result, never the `ColumnConst` itself, so `Const` behaviour is unchanged. The dense path is unaffected and adds no copy, which I verified by instrumenting both live call sites: a dense key reports zero copies, a sparse key reports one at each site. A sparse key is the case that is a use-after-free today, and it already paid for a full materialization inside `removeSpecialRepresentations`, so the extra `memcpy` of that same buffer is negligible. Reaching the bug requires a path that hands a non-materialized key to the dictionary. `dictGet`, `dictHas` and the hierarchy functions cannot: `IFunction::useDefaultImplementationForSparseColumns()` and `...ForReplicatedColumns()` both default to true and no dictionary function overrides them, so a dense, caller-owned column arrives. `IDictionary::getByKeys` (the direct join) is the reaching path, because its own `removeSpecialRepresentations` call sits inside a Nullable-only branch and a non-Nullable sparse key passes through untouched. That is also why `04627_direct_join_dictionary_nullable_key`, which does exercise a sparse key, never caught this: its key is Nullable, so it gets materialized. All 14 `getColumnVectorData` call sites are fixed by this single change. Two of them are reachable today, both in `FlatDictionary` and both on the same `getByKeys` call: `hasKeys` (the site in the reports below) and `getColumn`, reached through `getColumns`. The remaining 12 are hierarchy-only and reachable solely through the pre-converting function path. `Hashed` and `HashedArray` never reach the helper on the `getByKeys` path at all, because `DictionaryKeysExtractor` holds its converted column by value and therefore owns it. I also swept every other `convertToFullColumnIf*` / `recursiveRemove*` / `removeSpecialRepresentations` call site under `src/` for the same \"derived data escapes the owning local\" shape and found no second instance, so no sibling fix is needed. The bug dates to 2021 (`b5b624f3d7e9bf`, which introduced the conversion here) and was widened in 2025 by `2b6cb36d1dc936`, which added `ColumnReplicated` as a second carrier. The same code is present on 26.7, 26.6, 26.5, 26.4 and 26.3. Found while triaging CI on #109225 and reproduced on unmodified master. It is latent in CI only because no existing test combined a sparse-serialized left key with a direct join over a `FLAT()` dictionary; CI randomizes `ratio_of_defaults_for_sparse_serialization`, so any test that does hit this combination fails roughly 40% of the time. Reports on `31b4a2961ef4c5183f7d15dda7f755541a77a98c`: - `AddressSanitizer: heap-use-after-free`, allocated by `ColumnSparse::convertToFullColumnIfSparse`, freed at the end of `getColumnVectorData`, read by `FlatDictionary::hasKeys`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20flaky%20check%29 - `Logical error: '(n >= (static_cast<ssize_t>(pad_left_) ? -1 : 0)) && (n <= static_cast<ssize_t>(this->size()))'` from the same read: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20flaky%20check%29 The new test `04652_direct_join_dictionary_sparse_key` covers both live call sites, the aggregation shape from the report, and a key carried through `ARRAY JOIN` over a sparse base column, which is a third shape where the key has to be materialized. It pins its results against a dense table and against `join_algorithm = 'hash'` instead of hand-written constants, and asserts both that the key really is sparse and that `DirectKeyValueJoin` is still chosen, so it cannot pass vacuously. On master it aborts; with the fix it passes 50/50 with and without randomized settings. Reverting only the new predicate makes it abort again. A follow-up cleanup worth doing separately: the helper carries a `/// TODO: Remove` and would be better returning the owning `ColumnPtr` alongside the data, which removes the need for `backup_storage` entirely. That touches all 14 call sites and four dictionary classes, so it does not belong here.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112327",
          "createdAt": "2026-07-28T17:38:47Z",
          "updatedAt": "2026-08-13T13:21:21Z",
          "timestamp": "2026-08-13T13:21:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c58767c252dbddf60dbc",
        "signalId": "github:ClickHouse/ClickHouse:issue:113993",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113993",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Minmax index incorrectly prunes fractional DateTime64 values for integer bounds",
          "text": "### Company or project name _No response_ ### Describe what's wrong A minmax data-skipping index can change the result of a strict comparison between DateTime64 and an integer. Without skip-index pruning, `time > 0` correctly matches a value of 0.01 seconds. With a minmax index, the same row is pruned. ### Does it reproduce on the most recent release? Yes ### How to reproduce https://fiddle.clickhouse.com/66f7a197-cd87-4248-8e8c-6c44a4b85e4b ```sql CREATE TABLE t ( time DateTime64(9), INDEX idx_time time TYPE minmax GRANULARITY 1 ) ENGINE = MergeTree ORDER BY tuple() SETTINGS index_granularity = 1; INSERT INTO t VALUES (0.01::Decimal(9, 2)::DateTime64(9)); SELECT count() FROM t WHERE time > 0 SETTINGS use_skip_indexes = 0; -- 1 SELECT count() FROM t WHERE time > 0 SETTINGS force_data_skipping_indices = 'idx_time'; -- 0 EXPLAIN indexes = 1 SELECT * FROM t WHERE time > 0; ``` `EXPLAIN` shows the minmax condition as `time in [1, +Inf)`. ### Expected behavior Both queries should return 1. A data-skipping index must not change the query result. ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context `Range::shrinkToIncludedIfPossible` changes the open UInt64 bound `> 0` into the closed bound `>= 1`. That transformation is valid for an integer value domain, but not for DateTime64(9), which can contain values between 0 and 1. KeyCondition treats native integers and DateTime64 as directly comparable, so the integer bound reaches this normalization without first being converted into the DateTime64 domain.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113993",
          "createdAt": "2026-08-08T23:59:27Z",
          "updatedAt": "2026-08-13T13:21:03Z",
          "timestamp": "2026-08-13T13:21:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "clickgap-analyzed",
            "culprit-pr-pinned"
          ],
          "author": "EmeraldShift",
          "state": "open",
          "assignees": [
            "amosbird"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:56de2d3ee4910248b771",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114623",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114623",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "CI: Cache: restrict cross-branch reuse to pull_request workflows only",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112358 Cache reuse in praktika was only branch-gated for `push`; `schedule`, `dispatch`, and `merge_queue` could reuse a record produced on any branch. This tightens that as a deliberate design choice, on two grounds: - **Security.** A record can be produced on any branch, including untrusted ones. Letting a trusted lane (`push`, `merge_queue`) reuse an arbitrary-branch record is a cache-poisoning vector, so trusted lanes should reuse only records from branches they trust. - **Quality gating.** Same-branch (and, for `merge_queue`, main-branch) reuse is a stronger gate. In particular it keeps the merge-queue drift guard (#112358) re-running against the merged-with-`master` state instead of trusting a PR-side green run. New reuse policy per workflow event (`ci/praktika/hook_cache.py`): | Event | May reuse a record produced on | |---|---| | `pull_request` | any branch | | `merge_queue` | its own branch, or the main branch | | `push` / `schedule` / `dispatch` | the same branch only | `merge_queue` reuses main-branch records because the merge group is the PR merged with the current main, so a record whose digest was produced on main has identical inputs and is safe to reuse. A PR's record is on a different branch and is never reused by other lanes, so the drift guard holds. The write side is the symmetric dual: only `push` may overwrite an existing record; every other event writes only into an empty slot (`if_not_exist=True`), so a narrow-branch record never clobbers the main-branch record. The stateless flaky check no longer needs distinct PR/merge-queue cache digests, so the comment justifying that and `test_pr_run_does_not_cache_away_the_merge_queue_drift_guard` are removed. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114623",
          "createdAt": "2026-08-13T11:50:30Z",
          "updatedAt": "2026-08-13T13:20:40Z",
          "timestamp": "2026-08-13T13:20:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [
            "leshikus"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:49fd62f2fd87d197c7f6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114413",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114413",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix loss of primary-key pruning for `DateTime64` columns under a `toUnixTimestamp` sorting key",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114407 Related: https://github.com/ClickHouse/ClickHouse/issues/79977 Related: https://github.com/ClickHouse/ClickHouse/pull/101814 For a `MergeTree` table with `ORDER BY toUnixTimestamp(ts)` where `ts` is `DateTime64`, a plain range filter on `ts` (e.g. `WHERE ts >= '2026-06-15'`) stopped using the primary key in 26.7: `EXPLAIN indexes = 1` shows `PrimaryKey Condition: true` and every granule of the matched parts is read. The regression came from #101814, which correctly changed the monotonicity gate in `KeyCondition::extractMonotonicFunctionsChainFromKey` to ask `getMonotonicityForRange` over the function's argument type instead of its result type. That exposed a gap: `ToNumberMonotonicity` (the monotonicity implementation behind `toUnixTimestamp` and the `toInt*`/`toUInt*` family) did not support `DateTime64` arguments at all and answered \"unknown\", so the monotonic chain was rejected and the pruning was silently lost. Keys like `toDate(ts)` or `toStartOfHour(ts)` were unaffected because their monotonicity classes handle `DateTime64` explicitly, and `toUnixTimestamp` over a plain `DateTime` column was unaffected because `DateTime` is on the whitelist of `ToNumberMonotonicity`. This PR teaches `ToNumberMonotonicity` to answer for `DateTime64` arguments. The conversion takes the whole number of seconds, truncating the fractional part toward zero, and throws a `DECIMAL_OVERFLOW` exception instead of wrapping around (see `DecimalUtils::convertTo`), so it preserves order everywhere it is defined: - For a signed target of at least 64 bits (e.g. `toInt64`), the conversion is total — the whole number of seconds of any `DateTime64` fits — so it is reported as `is_always_monotonic`, for concrete ranges as well. As a bonus, the mirror-image filter `toInt64(ts) >= c` over `ORDER BY ts` now prunes too (the batched application over columns of index values can never throw), and stays consistent with the exact-ranges `count()` optimization. - For narrower or unsigned targets (e.g. `toUnixTimestamp`, whose result is `UInt32`), the conversion throws for a part of the domain, so it must not claim `is_always_monotonic`: `matchesExactContinuousRange` takes that claim as a promise that the per-range analysis can confirm every granule of a found range, while a part may hold out-of-range values, for which the per-range analysis has to answer \"unknown\". The first version of this PR made exactly that inconsistent claim, and the AST fuzzer found the debug assertion \"Inconsistent `KeyCondition` behavior\" (a logical error, so the stateless runs on the debug build also showed \"Server died\"). Instead, the conversion now reports a new flag `Monotonicity::is_always_monotonic_where_defined` — monotonic over the subset of the domain where the evaluation succeeds — which is consumed only by the constant-pushdown gate in `canConstantBeWrappedByMonotonicFunctions`. It is sound there: stored keys cannot correspond to out-of-range values, because computing the sorting key at insert time would have thrown, and an unrepresentable constant is rejected gracefully by the guards in `applyFunctionChainToColumn`. Per the review findings, those guards are also fixed: the pre-execution range checks are extended from `Date`/`DateTime`/`UInt32` to the whole native integer family, so an out-of-range constant pushed through e.g. `ORDER BY toUInt8(ts)` is rejected gracefully instead of throwing a `DECIMAL_OVERFLOW` exception during index analysis, and the negative-value fast reject is now unsigned-specific, so signed integer keys (e.g. `ORDER BY toInt64(ts)`) keep pruning for pre-`1970-01-01` filters. For partial conversions like `toUnixTimestamp`, the mirror-image case #79977 (`WHERE toUnixTimestamp(ts) >= c` over `ORDER BY ts`, a full scan since 23.1) remains out of scope: index analysis applies chain functions to whole columns of index values (see `applyFunction` in `KeyCondition.cpp`), and a part may contain out-of-range values next to the checked range, for which the batched conversion would throw an exception. For total conversions like `toInt64` it is fixed here. The test covers the restored pruning (with `force_primary_key`), sub-second bounds of the relaxed atom (positive and negative), constants outside the `UInt32` range, `toInt64` sorting keys including pre-`1970-01-01` filters, out-of-range constants over a narrow `toUInt8` key, the mixed-part case where index analysis must not throw an exception, and the exact-ranges `count()` optimization over a total conversion. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a regression in 26.7: for a `MergeTree` table ordered by `toUnixTimestamp` (or another integer conversion) of a `DateTime64` column, a plain range filter on that column no longer used the primary key and read all granules of the matched parts. Additionally, a filter like `toInt64(ts) >= c` over a table ordered by the raw `DateTime64` column now uses the primary key.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114413",
          "createdAt": "2026-08-12T02:15:56Z",
          "updatedAt": "2026-08-13T13:20:23Z",
          "timestamp": "2026-08-13T13:20:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b7bd741bab734eefffed",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114492",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114492",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113484 to 26.6: Push down plan level constants from joins",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113484 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114488 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31605886263/job/94144581099)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114492",
          "createdAt": "2026-08-12T14:35:23Z",
          "updatedAt": "2026-08-13T13:20:12Z",
          "timestamp": "2026-08-13T13:20:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [
            "diegomestre2",
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b2b082596b7ebb6b3ed3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110084",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110084",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add UUID2 data type with correct sorting",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/110066 Introduces `UUID2`, a variant of the `UUID` data type with correct (lexicographic) sorting. ## Motivation For historical reasons, the `UUID` data type sorts by the *second half* of the value. This is unexpected and, in particular, hurts the performance of primary indexes built on `UUIDv7` columns, whose most significant bits are a timestamp: with sorting by the second half, primary-key analysis cannot prune granules by the timestamp. ## What this does `UUID2` stores the 128-bit value as a plain big-endian integer of the 16 canonical bytes, so that natural integer comparison of the underlying value matches the textual (lexicographic) order and the canonical byte order used by most other systems. It reuses the `UUID` column and `Field` representation (like `DateTime` reuses `UInt32`), so sorting is correct with no extra comparison code, and its binary/interchange serialization is the canonical big-endian byte order. - `UUID1` is an alias of the current `UUID` type. - A new setting `uuid_type_version` (default `1`) controls whether the bare name `UUID` resolves to `UUID` (`1`) or `UUID2` (`2`) at `CREATE`/`ALTER` time. The resolved concrete type is materialized into the stored table definition (including nested types such as `Array(UUID)`), so reads never depend on the session setting and existing tables are never rewritten. The default will be flipped to `2` in a later, separate change. - Conversions to/from `String`, `UInt128`, `FixedString(16)` and `UUID`, plus `toUUID2` / `toUUID2OrZero` / `toUUID2OrNull`. - Parity across functions (`hex`/`bin`, `reinterpretAs*`, `UUIDv7ToDateTime`, `UUIDToNum`, `empty`/`notEmpty`, `min`/`max`, hashing, `uniq`), formats (`RowBinary`, `Native`, `JSON`, `CSV`, `TSV`, `Arrow`, `Parquet`, `Avro`, `BSON`, `MsgPack`, `Protobuf`, `CapnProto`, `JSONExtract`) and storage (`generateRandom`, `bloom_filter` skip index). The `UUID` type is unchanged (verified: still sorts by second half, identical conversions, default `uuid_type_version = 1`). ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new data type `UUID2`, a variant of `UUID` that sorts by its textual (lexicographic) representation instead of by the second half of the value. The setting `uuid_type_version` (default `1`) selects whether the type name `UUID` resolves to `UUID` or `UUID2`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110084",
          "createdAt": "2026-07-11T11:29:15Z",
          "updatedAt": "2026-08-13T13:19:35Z",
          "timestamp": "2026-08-13T13:19:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e8dd885926c01af12621",
        "signalId": "github:ClickHouse/ClickHouse:issue:72380",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:72380",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "`RENAME DATABASE` query doesn't work for materialized views",
          "text": "**Company or project name** Prefer not to specify **Describe what's wrong** `RENAME DATABASE` query doesn't seem to be working correctly with materialized view statements. After renaming a database with `RENAME DATABASE old_name TO new_name` query, the materialized views would still reference tables with the `old_name` as could be seen with `SHOW CREATE new_name.materialized_view` statement in the repro. That leads to a bunch of issues with database permissions and `SELECT` queries. And while `INSERT` queries into the main table seem to be working, they are not triggering any relative materialized views. Repro: https://fiddle.clickhouse.com/797f730d-d037-4b64-8d17-ce47911fccdd **Does it reproduce on the most recent release?** Yes **Enable crash reporting** Doesn't crash **How to reproduce** I first got this on `24.8.1.10452` in ClickHouse Cloud, but I'm also able to reproduce it on current latest. ``` CREATE DATABASE test; CREATE TABLE test.sample ( id UUID DEFAULT generateUUIDv4(), data TEXT DEFAULT '' ) ENGINE = MergeTree PRIMARY KEY (id) ORDER BY (id); -- Create materialized view without explicitly specifying target table CREATE MATERIALIZED VIEW test.inline_mat_view ( uuid UUID, data TEXT ) ENGINE MergeTree ORDER BY (uuid) AS SELECT id as uuid, data FROM test.sample; -- Create table for materialized view explicitly CREATE TABLE test.explicit_table ( uuid UUID, data TEXT ) ENGINE MergeTree ORDER BY (uuid); -- Create materialized view pointing to explicitly created table CREATE MATERIALIZED VIEW test.explicit_mat_view TO test.explicit_table AS SELECT id as uuid, data FROM test.sample; -- Output original CREATE statements for both materialized views SHOW CREATE test.inline_mat_view; SHOW CREATE test.explicit_mat_view; RENAME DATABASE test TO dev; -- Inserting data into main table doesn't trigger any error, but materialized views are not updated INSERT INTO dev.sample (data) VALUES ('test1'), ('test2'), ('test3'), ('test4'), ('test5'); -- CREATE statements for materialized views after renaming the database still points to the old database name SHOW CREATE dev.inline_mat_view; SHOW CREATE dev.explicit_mat_view; -- Exception due to missing 'test' database SELECT * FROM dev.inline_mat_view; SELECT * FROM dev.explicit_mat_view; ``` **Expected behavior** When renaming the database I would expect materialized view to be updated accordingly and to point to the correct table in the renamed database.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/72380",
          "createdAt": "2024-11-25T10:53:38Z",
          "updatedAt": "2026-08-13T13:18:04Z",
          "timestamp": "2026-08-13T13:18:04Z",
          "metrics": {
            "reactions": 1,
            "comments": 4
          },
          "labels": [
            "potential bug"
          ],
          "author": "biased-badger",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:764d158a16e1e18676e8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96130",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96130",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Randomize tests with DETACH/ATTACH table before query execution",
          "text": "Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that enable randomly detaching and reattaching tables used in a query before its execution. This is a testing-only feature designed to find bugs related to table reattachment. Before executing a query, the system collects all tables referenced in the AST, and for each eligible table (stores data on disk, supports detaching, has no action locks or dependencies), it performs a `DETACH` followed by `ATTACH`. Changes: - Add `supportsDetachingTables` virtual method to `IDatabase` (overridden to `false` for engines that do not support non-permanent `DETACH TABLE`: `DatabaseDictionary`, `DatabaseReplicated`, `DatabaseSQLite`, `DatabaseBackup`, `DatabaseFilesystem`, `DatabaseHDFS`, `DatabaseS3`, `DatabaseURL`, `DatabaseRemote`, `DatabaseDataLake`, `DatabaseMaterializedPostgreSQL`) - Add `has`/`hasAny` methods to `ActionLocksManager` for checking existing locks (skipping expired `weak_ptr` entries) - Add table collection visitor and reattach logic in `executeQuery` (runs after AST validations, process list admission, and external tables initialization; skips `EXPLAIN`, transactions, internal/non-initial queries, and CTE name collisions) - Fix off-by-one in `MergeTreeDeduplicationLog::dropOutdatedLogs` (don't drop the active log) and add `sync` call in shutdown - Add `no-random-detach` tag to tests incompatible with this feature - Add `--no-random-detach` and `--reattach-tables-probability` options to `clickhouse-test` - Add `02461_reattach_tables` test Continuation of #55943. Continuation of #42336 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that randomly `DETACH` and `ATTACH` tables used in a query before its execution. This is a testing-only feature that helps find reattachment-related bugs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Introduces new pre-execution mutations (internal `DETACH`/`ATTACH`) in `executeQuery`, which can affect table availability and concurrency behavior if enabled; guarded by new experimental settings but touches core query execution paths. > > **Overview** > Adds experimental settings `reattach_tables_before_query_execution` and `..._probability` to optionally **DETACH and ATTACH back** eligible tables referenced by a query immediately before execution, including AST table discovery that accounts for CTE scoping, privilege checks, dependency/lock checks, and safety skips (e.g. `system`, non-disk storages, dynamic-structure columns, transactions, `EXPLAIN`, internal/non-initial queries). > > Extends `IDatabase` with `supportsDetachingTables()` and marks multiple database engines as not supporting non-permanent detach; adds `ActionLocksManager::has/hasAny` helpers to avoid detaching tables with active action locks. Updates the test runner and stress tooling to randomize this behavior (with `--no-random-detach` and probability control), adds a new `02461_reattach_tables` test, and tags many existing tests to opt out where DETACH/ATTACH would add flakiness/overhead. Also fixes `MergeTreeDeduplicationLog` cleanup to avoid dropping the active log and ensures writer `sync()` on shutdown. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit a369371ff61cb1934815ea1e8caf7debd83c98c4. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96130",
          "createdAt": "2026-02-05T23:27:47Z",
          "updatedAt": "2026-08-13T13:17:36Z",
          "timestamp": "2026-08-13T13:17:36Z",
          "metrics": {
            "reactions": 2,
            "comments": 91
          },
          "labels": [
            "pr-build"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0d4d7454c04e7e59aa92",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114604",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114604",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix flaky 02999_scalar_subqueries_bug_2",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Failure report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=2485f1496f7dff85c6468cc59d1856cec0f92a9d&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_tsan%2C%20parallel%29 Related: https://github.com/ClickHouse/ClickHouse/pull/111983 Related: https://github.com/ClickHouse/ClickHouse/pull/91744 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02999_scalar_subqueries_bug_2` checked that the scalar subquery in a materialized view definition is not executed at CREATE time by racing a 2 second `max_execution_time` against a 3 second `sleepEachRow`. That makes the test depend on machine load rather than on the property under test, and it failed on master in `Stateless tests (amd_tsan, parallel)`: ``` Code: 159. DB::Exception: Timeout exceeded: maximum: 2000 ms. (TIMEOUT_EXCEEDED) (query 10, line 17) ``` The diagnostics rerun passed 43/43 with the same randomized settings. The scalar was not executed. The message carries no `elapsed ... ms` clause, while the in-sleep deadline check (`sleep.cpp:168`) goes through `checkTimeLimit`, which always reports a non-zero elapsed; the statement was cancelled while still pending, since every query including DDL is registered with the `CancellationChecker` watchdog unconditionally (`ProcessList.cpp:381`). The engine behaved correctly and only the clock failed. `max_execution_time` is not randomized by the runner, so there is no setting to pin, and widening the bound was already tried on the structural twin (#91744) without holding. I removed the timing oracle and assert the property directly with `throwIf(1)`, following @ alexey-milovidov's fix for that twin in #111983. If the scalar is not executed `throwIf` never fires; if it ever is, the statement fails `FUNCTION_THROW_IF_VALUE_IS_NON_ZERO`. `throwIf::isSuitableForConstantFolding` returns `false` (`throwIf.cpp:94`), the same guard `sleepEachRow` relies on, so the mechanism the old oracle depended on is preserved, and the check is now stronger: it detects execution at all, not only execution slower than 2 seconds. I also cover `CREATE TABLE ... EMPTY AS SELECT`, which reaches a different analysis site, and added an executed-position arm that must fail so the check has teeth. The reference file stays empty. Validation: the new test is green 50/50 under `enable_analyzer=1` and 50/50 under `=0`, green with 8 concurrent copies, and two mutations of the new oracle each redden it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114604",
          "createdAt": "2026-08-13T08:49:21Z",
          "updatedAt": "2026-08-13T13:17:22Z",
          "timestamp": "2026-08-13T13:17:22Z",
          "metrics": {
            "reactions": 2,
            "comments": 6
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8d3e787fb376aaf9f76e",
        "signalId": "github:ClickHouse/ClickHouse:issue:109476",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:109476",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Automatically change some settings if `make_distributed_plan=1` is enabled",
          "text": "* New distributed execution doesn't support some kinds of `IN (subquery)` queries. However we have amazing feature which rewrites `IN -> JOIN`: https://github.com/ClickHouse/ClickHouse/pull/83991. So `rewrite_in_to_join=1` should be enabled automatically if `make_distributed_plan=1` is specified. * Parallel replicas-style reads are not supported by `make_distributed_plan=1`. Instead it uses less efficient static distribution of work. So `enable_parallel_replicas=0` should be set automatically. * Recent optimization for correlated subqueries is not supported yet https://github.com/ClickHouse/ClickHouse/pull/91205. Should be disabled automatically with `correlated_subqueries_use_in_memory_buffer=0`. * `use_skip_indexes_on_data_read` should be disabled just in case. In case of huge table we can use distributed index analysis. * `compile_expressions=0` -- doesn't work yet, need to fix <!-- ch-version-info:start --> ### Version info - Resolved by: #112463 - Merged into: `26.8.1.506` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/109476",
          "createdAt": "2026-07-06T10:38:23Z",
          "updatedAt": "2026-08-13T13:17:21Z",
          "timestamp": "2026-08-13T13:17:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "unfinished code"
          ],
          "author": "alesapin",
          "state": "closed",
          "assignees": [
            "alesapin"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:5dae41ab73badcf5dcd2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101791",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101791",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "In case of trivial views, push whole outer query to shards.",
          "text": "### Changelog category: - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): In case of trivial views over distributed table push whole outer query to shards. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Description / Proposed Solution When a VIEW is defined over a Distributed table, ClickHouse traditionally executes it on the shards without enclosing outer query. This means filters and expressions declared in the outer query are evaluated on the coordinator after pulling raw data from shards. For views whose body is a plain SELECT (column references, *, or arbitrary expressions — but no aggregation, grouping, ordering, joins, window functions, or scalar subqueries) over a single Distributed table, we can do better: inline the view body as a subquery and hand the whole thing to StorageDistributed. Each shard then receives the full outer query with the view body inlined, evaluates it against its local table, and only ships the result back. A view qualifies as \"trivial\" if its inner query: - Has a single SELECT (no UNION) - Selects only column references, *, or expressions — but no window functions (require the full dataset) and no scalar subqueries in the SELECT list - Has no WITH, PREWHERE, GROUP BY, HAVING, QUALIFY, ORDER BY, LIMIT, LIMIT BY, DISTINCT, or ARRAY JOIN - Has no subqueries in the WHERE clause - Reads from exactly one table with no joins, no table functions, no FINAL, no SAMPLE - Is not a parameterized view and does not use SQL SECURITY DEFINER **The optimization can be disbaled by setting (enabled by default):** ``` SET optimize_trivial_view_pushdown_to_distributed = 0; ``` ### Example: Env setup: ``` create table x engine = MergeTree ORDER BY tuple() AS SELECT intDiv(number,100000) as a, number as b FROM numbers(1000000000); SET prefer_localhost_replica = 0; CREATE TABLE x_dist AS x ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), x); CREATE VIEW v_computed AS SELECT a + 1 AS x, b AS y FROM x_dist WHERE a != 0; ``` Performance: ``` :) SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: 6baedc55-c8e7-4b2b-9946-b0d828abaf25 ┌─plus(a, 1)─┬──────────sum(b)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └────────────┴─────────────────┘ 1 row in set. Elapsed: 17.199 sec. Processed 2.00 billion rows, 32.00 GB (116.29 million rows/s., 1.86 GB/s.) Peak memory usage: 38.58 MiB. :) SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 7c4f2854-3de2-4d40-b960-efcc44a7b26d ┌─────x─┬──────────sum(y)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └───────┴─────────────────┘ 1 row in set. Elapsed: 16.497 sec. Processed 2.00 billion rows, 32.00 GB (121.24 million rows/s., 1.94 GB/s.) Peak memory usage: 38.82 MiB. ``` Plan: ``` :) explain SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; EXPLAIN SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: e2d32d5f-7024-4e6a-b686-45c7c0d5f2ef ┌─explain──────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression (Before GROUP BY) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression (Before GROUP BY) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └──────────────────────────────────────────────────────────────────────────────────┘ :) explain SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; EXPLAIN SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 88213305-46a5-493e-862c-a80c675c9452 ┌─explain───────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes query planning/execution for a subset of views over `Distributed` tables and touches access checks/row policy enforcement and SQL SECURITY semantics, which can affect correctness and security-sensitive behavior. > > **Overview** > Adds a new default-on setting `optimize_trivial_view_pushdown_to_distributed` to inline *trivial* views over `Distributed` tables and push the full outer query down to shards, reducing coordinator-side filtering/processing and network transfer. > > Implements planner rewrites to swap the view table expression with an analyzed subquery, merge `FINAL`/`SAMPLE` modifiers, and preserve semantics by suppressing pushdown when the outer query contains non-deterministic functions, while also explicitly handling SQL SECURITY modes, row-policy injection/logging, and column-pruned privilege checks. > > Extends integration/stateless tests to cover modifier propagation, non-determinism suppression, row-policy enforcement, SQL SECURITY behavior, and interactions with `max_rows_to_read_leaf`. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 583e1e6c0e8e25081391d7a07af086c6f9888c6f. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101791",
          "createdAt": "2026-04-04T19:32:01Z",
          "updatedAt": "2026-08-13T13:16:57Z",
          "timestamp": "2026-08-13T13:16:57Z",
          "metrics": {
            "reactions": 3,
            "comments": 15
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "simonmichal",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cbc6c5a85901f8458a7a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104965",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104965",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add server setting `additional_memory_tracking_per_thread`",
          "text": "Each thread accumulates up to `max_untracked_memory` (4 MiB by default) of allocations before reporting them to the server-wide `MemoryTracker`. With many threads, this unreported memory can sum to a large amount, causing the server's tracked memory usage to under-count actual consumption and leading to OOM. This PR introduces a server-level setting `additional_memory_tracking_per_thread` (default 4 MiB) and speculatively charges this amount to the server-wide `MemoryTracker` around every job executed in our `ThreadPool` workers. The global tracked memory becomes a safe upper bound on actual consumption. The reservation is charged on the server-wide (total) tracker only — deliberately not through the query's tracker chain. Query-level accounting feeds heuristics that compare memory deltas against byte thresholds (conversion of aggregation hash tables to two-level via `group_by_two_level_threshold_bytes`, spill-to-disk decisions), and phantom reservations of `num_threads * 4 MiB` trip those thresholds immediately: an earlier revision of this PR that charged the query tracker showed consistent slowdowns of GROUP BY queries in performance tests for exactly this reason. With the server-wide-only reservation, query-level and user-level accounting (`max_memory_usage`, `memory_usage` in `system.processes`) are unaffected. The speculative reservation uses the throwing path of the memory tracker, so when it would exceed the server memory limit the corresponding job is treated as failed with `MEMORY_LIMIT_EXCEEDED` — the same behavior as if the job itself had exceeded the limit. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new server setting `additional_memory_tracking_per_thread` (default 4 MiB) which speculatively reserves this amount on the server-wide memory tracker around every `ThreadPool` job. It compensates for the up to `max_untracked_memory` of un-reported allocations per thread, making the server's tracked memory a safe upper bound on actual consumption and reducing the risk of OOM with many concurrent threads. Query-level and user-level memory accounting are not affected. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASan - [ ] <!---ci_include_tsan--> All with TSan - [ ] <!---ci_include_msan--> All with MSan - [ ] <!---ci_include_ubsan--> All with UBSan - [ ] <!---ci_include_coverage--> All with Coverage - [ ] <!---ci_include_aarch64--> All with Aarch64 #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASan - [ ] <!---ci_exclude_tsan--> All with TSan - [ ] <!---ci_exclude_msan--> All with MSan - [ ] <!---ci_exclude_ubsan--> All with UBSan - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 #### Extra options: - [ ] <!---ci_set_arm--> Add tests with aarch64 builds - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104965",
          "createdAt": "2026-05-14T16:39:55Z",
          "updatedAt": "2026-08-13T13:16:10Z",
          "timestamp": "2026-08-13T13:16:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "pr-improvement",
            "memory"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:541f129898ac0a696dd9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110144",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110144",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support per-authentication-method GRANTS clause in CREATE USER and ALTER USER",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109117 Implements per-authentication-method grant limits, the first part of the linked issue: ```sql ALTER USER vasya ADD IDENTIFIED WITH password BY 'WYmdFyas8PftrHbHQQo8' VALID UNTIL '2026-12-31' GRANTS (SELECT ON db.table) ``` When a user logs in with such a method, the access rights of the session are the intersection of the user's access rights (including granted roles) with the listed elements. The clause never adds rights: a listed privilege that is not granted to the user stays unavailable. This provides a way to create tokens for applications: an additional credential with an expiration date and a limited set of grants, which is tied to the user — it is displayed in `query_log` and `processlist` as the user, stops working if the user is deleted, and is narrowed when the user loses grants. Details: - The clause is parsed after the per-method `VALID UNTIL`, works with `CREATE USER`, `ALTER USER [ADD] IDENTIFIED`, and `NOT IDENTIFIED`, is shown by `SHOW CREATE USER`, and persists through the SQL serialization of access entities (and therefore backups). Elements without a database name are bound to the current database when the query is interpreted. - The intersection erases all grant options (the clause cannot express them), so such sessions cannot `GRANT` anything. Role administration is denied entirely (fail-close), including per-role admin option. `EXECUTE AS`, `ALTER USER`, `CREATE USER` and similar escapes require the corresponding rights to be listed explicitly and granted to the user. - Reattaching to a named session (`session_id`) with a different credential re-applies the limit of the credential used by the new connection, so a limited credential cannot pick up the full rights of a session created by an unrestricted one. - The limit is captured at login: `ALTER USER` affects new sessions, not established ones (same as `VALID UNTIL`). - A new `auth_grants` column in `system.users` exposes the limit of each authentication method. - `ContextData`'s copy constructor now preserves `external_roles` and the new field: previously a recalculation of access rights on a copied context silently dropped external roles; for the new field that would mean silently widening the rights. Known limitations (consistent with `VALID UNTIL` and external roles): the limit is not propagated to other nodes of a cluster in distributed queries or `ON CLUSTER` DDL (the query is checked on the initiator), and the clause is not available in `users.xml`. The `CREATE TOKEN` syntactic sugar and the separate grant for self-service `ADD IDENTIFIED` mentioned in the issue are left for a follow-up. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Authentication methods in `CREATE USER` and `ALTER USER ... ADD IDENTIFIED` support a `GRANTS (SELECT ON db.table, ...)` clause which limits the access rights of sessions authenticated with that method to the intersection with the listed grants. This allows using additional credentials as tokens for applications: `ALTER USER vasya ADD IDENTIFIED WITH password BY '...' VALID UNTIL '2026-12-31' GRANTS (SELECT ON db.table)`. Closes [#109117](https://github.com/ClickHouse/ClickHouse/issues/109117). ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110144",
          "createdAt": "2026-07-12T05:16:00Z",
          "updatedAt": "2026-08-13T13:15:45Z",
          "timestamp": "2026-08-13T13:15:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c5ab7dfe4925ad9b42e9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104437",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104437",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add spatial_bbox skip index for MergeTree geometry columns",
          "text": "Part of making ClickHouse fastest spatial analytical engine on Earth https://github.com/bacek/chgeos/blob/main/BENCHMARK.md ;) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Adds `spatial_bbox` skip index for MergeTree geometry columns. The index stores a bounding box per granule and skips granules whose geometry cannot intersect the query geometry, reducing work for spatial predicates. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104437",
          "createdAt": "2026-05-08T22:48:54Z",
          "updatedAt": "2026-08-13T13:14:54Z",
          "timestamp": "2026-08-13T13:14:54Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "bacek",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6050337cf318b5a58256",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114636",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114636",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix logical error in text index lazy apply mode on cancelled queries",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114603 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed logical error `Multi-block postings must be compressed` in queries over tables with a text index in the lazy posting-list apply mode, when the query was canceled during the read.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114636",
          "createdAt": "2026-08-13T12:51:52Z",
          "updatedAt": "2026-08-13T13:14:42Z",
          "timestamp": "2026-08-13T13:14:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:1fd4d097510f5c5356e4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109130",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109130",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Stress test: do not force use_query_cache for non-throw overflow-mode tests",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/107907 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description The stress runner enables `use_query_cache=1` as a global per-runner client option (`ci/jobs/scripts/stress/stress.py`, probability `1/15`). The server refuses that together with a non-throw `*_overflow_mode` and raises `QUERY_CACHE_USED_WITH_NON_THROW_OVERFLOW_MODE` (error 731), because such a mode can truncate the result, which must never be cached (Bug 67476, guarded in `executeQuery.cpp`). Tests that set such a mode then fail. `00107_totals_after_having` sets `group_by_overflow_mode = 'any'`, and three consecutive failures trip `--max-failures-chain`, aborting the runner. 72 stateless tests set a non-throw overflow mode. Reported by @ alexey-milovidov while triaging a `Stress test (amd_msan)` failure on #107907 (report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107907&sha=f7e21726f54a9&name_0=PR). That block now also pins the overflow modes to `'throw'`, but only as a client option, so it does not cover this: a test's own `SET` overrides a client option. Verified against a local server: with `use_query_cache=1` and every mode pinned to `'throw'` on the command line, a query after `SET group_by_overflow_mode = 'any'` still returns 731. Fix: detect tests that set a non-throw overflow mode and turn the forced query cache back off for exactly those, on the command line and on the HTTP carrier used in cloud mode. The trailing value wins in both. The detected settings are the ten the server guards, so `read_overflow_mode_leaf` is covered; comment-only mentions are ignored. Query-level `SETTINGS` still override the override, so the intentional 731 coverage in `02494_query_cache_bugs` is preserved, and the query-cache coverage of every other test is unchanged. Treating 731 as benign in the runner was rejected: it would let a test's queries silently abort and could mask real 731 regressions.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109130",
          "createdAt": "2026-07-02T09:28:03Z",
          "updatedAt": "2026-08-13T13:14:34Z",
          "timestamp": "2026-08-13T13:14:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:09447643bda4bbee241b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114401",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keeper: do not lose a session request when the Raft leader changes",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/78474 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a bug where a client connecting to ClickHouse Keeper during a Raft leader change could be held for the whole `session_timeout_ms` (30 seconds by default) before its connection was rejected, instead of being rejected as soon as the in-flight request was dropped. The connecting client can now reconnect to another replica sooner. ### Description A client's `Connect` makes Keeper submit an internal `SessionID` request. If the Raft append stream breaks while it is in flight, during a leader election say, that request is lost silently. **Root cause.** Such a request carries `session_id = -1` and no xid; its identity lives in `(server_id, internal_id)`. Two places used the wrong key: * `KeeperRequestDispatcher::onCommit` correlated a commit with its in-flight head by `(session_id, xid)`. Every `SessionID` request shares `(-1, 0)`, and `onCommit` runs on every node for every commit, so a `SessionID` committed for another server retired a still-uncommitted local request. * The error path queued the response for lookup by `session_id`, which `-1` has no callback for, so it was discarded and `getSessionID` timed out. The real waiter is a promise keyed by `internal_id`. **The change.** `onCommit` additionally requires `(server_id, internal_id)` to match for `OpNum::SessionID`. Error responses go through one `SessionID`-aware routing helper on `KeeperDispatcher`, wired into **both** dispatchers: `use_new_dispatcher` is a setting and the old one shares the defect. The request now fails at the in-flight drain bound rather than at `session_timeout_ms`. `KeeperTCPHandler` does not branch on the error code, so that earlier rejection is the whole user-visible gain; the accurate `ZCONNECTIONLOSS` only improves the server log. `test_keeper_force_recovery` also gets a retry around the connect after the election, since dropping in-flight appends is deliberate. The correlation fix is scoped to `OpNum::SessionID`. The garbage collectors also use negative session ids, but their `TryRemove` is idempotent and unwaited, so they are unaffected. **Validation.** New gtests cover both the routing decision and the production wiring behind it, each verified to go red under a mutation of the change it covers; the old dispatcher was exercised with `use_new_dispatcher = false`. An injected fault at the connect under test reddens the integration test without the retry and passes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114401",
          "createdAt": "2026-08-12T00:10:11Z",
          "updatedAt": "2026-08-13T13:14:30Z",
          "timestamp": "2026-08-13T13:14:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1e0218931e02c7ffd60f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113651",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113651",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for container boot before installing packages in `Install packages`",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Install packages (amd_release)` intermittently fails its `Install server rpm` substep with one line and nothing else: ``` + yum localinstall '--disablerepo=*' --allowerasing -y /packages/clickhouse-server-...rpm ... [Errno 2] No such file or directory: '/var/cache/dnf/metadata_lock.pid' ``` Example: [PR #109299](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109299&sha=3c7aad076ea307634a39a3089615b6a73e208621&name_0=PR&name_1=Install%20packages%20%28amd_release%29). **Root cause.** The image boots systemd, and the Dockerfile deliberately keeps `systemd-tmpfiles-setup.service`, which runs `systemd-tmpfiles --create --remove --boot`. centos:8 ships `/usr/lib/tmpfiles.d/dnf.conf`, whose entire content is remove directives for the dnf lock files, `metadata_lock.pid` among them. `test_install` starts the container `--detach` and `docker exec`s `install.sh` immediately, so `yum` and that lock wipe run concurrently. dnf takes the metadata lock in `Base.fill_sack` before it even opens the rpm files and does not guard against it vanishing, so the substep aborts before any package work: install, start and the smoke test never run. **Change.** Wait for the boot transaction before the first `docker exec`, in the shared `test_install` helper, so every substep of both images is covered. Only the rpm image is affected in practice: ubuntu:22.04 ships no `dnf.conf`, and the dpkg and apt locks survive the same run. Waiting for D-Bus first is the load-bearing half: in the first milliseconds `systemctl` cannot reach systemd at all, so a gate built on it alone silently does nothing in the window it guards. A bare `systemctl start` returned `Failed to connect to bus` in 5 of 8 tries; with the bus wait in front, `rc=0` in 8 of 8. **Validation.** Built the real image locally and amplified the race with 24 concurrent containers, no fault injection: **6 of 144** runs reproduced the exact CI line without the wait, **0 of 144** with it, gate `rc=0` in 144/144. Cost is ~0.17 s per container, so ~3.1 s over the 18 containers a job starts, against a job whose 30-day median is ~240 s. No related open issue found.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113651",
          "createdAt": "2026-08-06T11:22:40Z",
          "updatedAt": "2026-08-13T13:12:16Z",
          "timestamp": "2026-08-13T13:12:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:7d594d5d0ce69f555a4a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112816",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112816",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow running queries detached from client session",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New setting `run_query_in_background`. The server accepts the query, immediately returns an empty result, and runs it to completion regardless of what happens to the connection. The result is discarded. Track the query by its `query_id` in `system.processes` and `system.query_log`. Intended for long queries like `INSERT ... SELECT`, `CREATE TABLE … AS SELECT`, or `CREATE MATERIALIZED VIEW … POPULATE` that must not die with a dropped connection.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112816",
          "createdAt": "2026-07-31T20:49:40Z",
          "updatedAt": "2026-08-13T13:11:52Z",
          "timestamp": "2026-08-13T13:11:52Z",
          "metrics": {
            "reactions": 3,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:89e9a413f7c2a5fa450c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113707",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113707",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add Erathos connector to data ingestion docs",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds Erathos, an ELT platform, to the data ingestion integrations list.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113707",
          "createdAt": "2026-08-06T18:04:58Z",
          "updatedAt": "2026-08-13T13:10:44Z",
          "timestamp": "2026-08-13T13:10:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "gelsonbagetti",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:856189c95f4833d09e90",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114212",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114212",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain the required column order when a dictionary QUERY returns columns in the wrong order",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a dictionary with a PostgreSQL source and a custom `QUERY` returns its columns in an order that does not match the dictionary structure, the error now names the destination dictionary column and the column order the query must return, instead of only reporting that a value could not be parsed. ### Description A `RANGE_HASHED` dictionary over `SOURCE(POSTGRESQL(... QUERY '...'))` failed with `Cannot parse LocalDate: 97955`, where `97955` was an unrelated `id` column's value: the DDL was valid, only the column order was wrong, and the message pointed at the data. A dictionary's expected source structure is keys-first: key(s), `RANGE MIN`, `RANGE MAX`, then the attributes. ClickHouse aliases every column when it builds the `SELECT` itself, but a user `QUERY` passes through verbatim and the result is read strictly by position, so `id` was deserialized into the `Date` attribute `contract_time`. On a conversion failure `PostgreSQLSource` now names the destination column and its position, and for a custom `QUERY` states the required order, derived from the dictionary's own structure so it suits every layout (the `RANGE` wording appears only for range dictionaries). Successful loads are untouched, and `BAD_ARGUMENTS` is preserved because `MaterializedPostgreSQL` relies on it. <details><summary>The message for the reported case</summary> ``` Cannot parse PostgreSQL value '97955' as Date: Cannot parse LocalDate: 97955: while reading column 2 of the result into `contract_time`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `meter_no`, `contract_time`, `end_date`, `id` (the key column(s) first, then the RANGE MIN and RANGE MAX columns, then the remaining attributes) ``` For a `FLAT` dictionary the same failure omits the `RANGE` clause: ``` Cannot parse PostgreSQL value 'abc' as Int64: Could not convert string to l: 'abc': while reading column 2 of the result into `num`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `id`, `num`, `name` ``` </details> A new integration test covers the range case, the documented order, and a `FLAT` dictionary asserting no `RANGE` wording appears; it fails on master. `04401_composite_key_dict_non_leading_key` gains the range shape over the ClickHouse source, pinning its by-name reordering against regression. Scope, since the mechanism is wider than the fix. The MySQL dictionary source (and likely XDBC and Cassandra) maps positionally too, but its parse errors do not funnel through one place, so folding it in grows the diff; say the word and I will extend it. Two cases stay untouched: an unnamed-column query still matches positionally and can silently return a wrong value, and a short row silently defaults trailing columns. Matching by name is possible, not ruled out: a `LIMIT 0` probe, or projecting the expected names inside the streamed query. I avoided it because it silently changes behaviour for every dictionary relying on positional matching, and arbitrary queries can return duplicate or unnamed columns.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114212",
          "createdAt": "2026-08-10T19:51:47Z",
          "updatedAt": "2026-08-13T13:10:14Z",
          "timestamp": "2026-08-13T13:10:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:de00475c65385f909804",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114597",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114597",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix zero-capacity hash-table statistics cache seeded by lazy FINAL",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/101647 Related: https://github.com/ClickHouse/ClickHouse/pull/113333 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed hash-table statistics — and with them aggregation hash-table preallocation — being silently and permanently disabled for the whole server when a query using the lazy `FINAL` optimization (`query_plan_optimize_lazy_final`) was the first aggregation to run after server startup. ### Description The `FINAL`-replacing dedup aggregation of the lazy `FINAL` optimization (`LazyReadReplacingFinalSource`, #101647) built its `AggregatingStep` with default-constructed `StatsCollectingParams`, i.e. `max_entries_for_hash_table_stats = 0`. The query-plan optimization pass `setAggregationHashTableCacheKeys` stamps a cache key onto every keyed `AggregatingStep`, which enables statistics collection. The process-wide statistics cache is created lazily by the first query that touches it, with `max_entries_for_hash_table_stats * sizeof(Entry)` as its capacity — so when a lazy `FINAL` query was the first aggregation in a freshly started server, the cache was created with zero capacity: every insertion was evicted immediately and every lookup missed, permanently disabling hash-table statistics (and preallocation) for the whole server lifetime. In the affected CI run's server log the aggregation statistics cache had 3515 `Statistics updated` entries and zero `found in cache` hits over the entire lifetime, while the join statistics cache (a separate instance created with proper parameters) worked normally. This stayed unnoticed while keyless aggregations, which carry properly constructed params, also touched the cache and almost always created it first. #113333 (merged 2026-08-12 19:00 UTC) stopped keyless aggregations from touching the cache, after which the lazy `FINAL` tests started to win the first-toucher race, and `04625_hash_table_sizes_stats_table_expression_modifiers` began to fail with `final-prealloc 0` / `no-final-prealloc 0` across unrelated jobs (`amd_msan`, `amd_asan_ubsan, distributed plan`, `Fast test`, `Fast test (arm_darwin)`) starting 2026-08-12 22:39 UTC — 6 failures, no failures of this kind before that date. Because the poisoned cache is server-instance state, the flaky-check diagnosis reproduced the failure 5/5 within the affected run while adjacent master commits passed. CI reports: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=2790ac4a68cf412d05db2d1573685a1a8467ecdb&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_msan%2C%20WasmEdge%2C%20parallel%2C%201%2F3%29 - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=86363e819f9cb704a892063da9998e4c1281eb76&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_msan%2C%20WasmEdge%2C%20parallel%2C%201%2F3%29 The fix: - `LazyReadReplacingFinalSource` now constructs real `StatsCollectingParams` (same pattern as `Planner`), so the dedup aggregation also benefits from size hints on repeated `FINAL` queries. - `setAggregationHashTableCacheKeys` skips steps whose builder left `stats_collecting_params` unconfigured (`max_entries_for_hash_table_stats == 0`), so a stamped key can never enable collection with a zero-capacity configuration; this also makes an admin-configured server-level `max_entries_for_hash_table_stats = 0` a clean disable for the aggregation path instead of a zero-capacity cache. - The new test `04869_lazy_final_hash_table_stats_cache` runs the first-toucher scenario deterministically in `clickhouse-local` (a fresh process owns a fresh statistics cache): the lazy `FINAL` query runs first, then a 650000-group victim aggregation twice; the second run must preallocate, observed through the `AggregationPreallocatedElementsInHashTables` event in `system.events`. On the unfixed binary this deterministically reports nothing; with the fix it reports 650000. Verified locally: the `clickhouse-local` repro flips from empty to 650000 with the fix; `04625_hash_table_sizes_stats_table_expression_modifiers`, the new `04869` test and the lazy `FINAL` family (`03988`, `03990`, `03991`, `04092`, `04093`) all pass against a patched server, including the exact CI scenario (fresh server, `03990_lazy_final_index_analysis` first, then `04625`). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114597",
          "createdAt": "2026-08-13T05:56:02Z",
          "updatedAt": "2026-08-13T13:10:04Z",
          "timestamp": "2026-08-13T13:10:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:03a1342b8db7466fb3c1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112605",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112605",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-apply LIMIT BY on the initiator for custom-key parallel replicas",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/111555 Follow-up to https://github.com/ClickHouse/ClickHouse/pull/111919, which fixed the `WITH FILL` half of #111555. This fixes the remaining `LIMIT BY` half. ### Problem Under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`), `LIMIT n BY` was applied per replica and the initiator never re-applied it, so it returned up to `n * replicas` rows per group: ```sql CREATE TABLE wf_min (g UInt16, k UInt32) ENGINE = MergeTree ORDER BY k; INSERT INTO wf_min SELECT number % 3, number FROM numbers(30); SELECT g FROM wf_min ORDER BY g LIMIT 2 BY g SETTINGS enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_range', parallel_replicas_custom_key = 'k', parallel_replicas_custom_key_range_upper = 30; -- returned 18 rows (6 per group); correct is 6 (0,0,1,1,2,2) ``` ### Root cause Custom-key parallel replicas splits rows across replicas by an arbitrary key and forces the initiator's input `from_stage` to `WithMergeableStateAfterAggregation(AndLimit)` (`PlannerJoinTree`), which tells the planner \"shards already finalized aggregation-stage processing\". The initiator therefore skips `LIMIT BY` (`Planner.cpp`, guarded by `!isFromAggregationState()`). That is correct for genuine sharding-key-aligned pushdown (each group lives on one shard), but the custom key does not align with the `LIMIT BY` key, so the per-replica `LIMIT BY` is not final. This is the same class of \"no initiator-side finalization over custom-key replica streams\" as the `WITH FILL` half fixed in #111919. ### Fix Re-apply `LIMIT BY` on the finalizing initiator (`isFinalizingStage()`) for custom-key parallel replicas, in addition to the normal `!isFromAggregationState()` case. On a replica that still emits a mergeable state, `LIMIT BY` runs as a preliminary that keeps `offset + length` rows and defers `OFFSET` to the initiator, so `LIMIT n OFFSET m BY` stays correct too. Custom-key parallel replicas is detected via a new `JoinTreeQueryPlan::is_parallel_replicas_custom_key` flag set only where the custom-key path is actually built — the replica custom-key filter, and the `Distributed` / MergeTree initiator dispatch in `PlannerJoinTree` — rather than the ambient `canUseParallelReplicasCustomKey()` setting (which is true whenever the profile enables a `custom_key_*` mode, even when the query used genuine sharding-key pushdown and never built a custom-key plan). This keeps `optimize_distributed_group_by_sharding_key` pushdown (e.g. `01244_optimize_distributed_group_by_sharding_key`) unaffected. The change is analyzer-only; the deprecated legacy interpreter's custom-key path is left unchanged (its custom-key handling is inconsistent across modes and a blanket re-application there regressed `custom_key_sampling` + `OFFSET`). ### Verification Built ClickHouse from an earlier revision of this branch and checked against a running 3-replica custom-key cluster: - `LIMIT 2 BY g` returns 6 rows (was 18) for both `custom_key_range` and `custom_key_sampling`. - `LIMIT 1 OFFSET 1 BY g` returns 3 rows (`0,1,2`), matching the non-distributed result — `OFFSET` applied once. - No regression: normal distributed `LIMIT 2 BY g` over `remote(...)` still returns the correct 6 rows. (The later review fix — switching from the ambient setting to the plan-tied flag — was validated by review/CI; the sandbox VM was recycled so it was not re-run locally. CI runs both `01244_optimize_distributed_group_by_sharding_key` and the new `04657_parallel_replicas_custom_key_limit_by`.) ### Test Added `tests/queries/0_stateless/04657_parallel_replicas_custom_key_limit_by.sql` (tagged `no-old-analyzer`), covering both custom-key modes and the `OFFSET` case. Confirmed it returns 18 (buggy) on unpatched and 6 on patched. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `LIMIT n BY` returning up to `n * replicas` rows per group under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`). `LIMIT BY` is now re-applied on the initiator over the merged replica streams. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112605",
          "createdAt": "2026-07-30T14:04:23Z",
          "updatedAt": "2026-08-13T13:10:02Z",
          "timestamp": "2026-08-13T13:10:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:5e59606a55ea8b70c0e2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110479",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110479",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "EXPLAIN SYNTAX: return a single record (default on)",
          "text": "<!-- Linked issues and pull requests. --> Closes: https://github.com/ClickHouse/ClickHouse/issues/80410 Related: https://github.com/ClickHouse/ClickHouse/pull/107925 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `EXPLAIN SYNTAX` now returns the reformatted query as a single `String` record (with embedded newlines) instead of one record per line, so its output is a recoverable single row that is directly usable (for example, `SELECT count() FROM (EXPLAIN SYNTAX ...)` returns `1`). This is controlled by the new `single_record` option, which defaults to `1`; set `single_record = 0` to restore the historical one-record-per-line output. Other `EXPLAIN` kinds (`PLAN`/`PIPELINE`/`AST`) keep their per-line tree output. ### Description Per the discussion on #80410, `EXPLAIN SYNTAX` is a reformatted, copy-pasteable query, so returning the whole query as a single record is far more usable than splitting it across N rows. This is done in two commits: 1. Add a `single_record` option to `EXPLAIN SYNTAX`, off by default (no behavior change). The single-record code path already existed (used by `EXPLAIN PLAN` JSON output); this wires it to the new option for the `SYNTAX` kind and renames the internal `single_line` flag to `single_record`. 2. Flip the default to `single_record = 1`. This is backward-incompatible: existing `EXPLAIN SYNTAX` consumers now see one row instead of N. The `oneline` option still controls physical-line rendering (default `0`, i.e. embedded newlines). The affected stateless `.reference` files (including `.oldanalyzer.reference` variants) were regenerated for the new default. The `optimize_syntax_fuse_functions` docstring example was switched from `FORMAT TSV` to `FORMAT TSVRaw` so its multi-line output stays literal now that the single record carries the embedded newlines. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110479",
          "createdAt": "2026-07-15T00:40:02Z",
          "updatedAt": "2026-08-13T13:09:34Z",
          "timestamp": "2026-08-13T13:09:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "Fgrtue"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:10d00efc3ccc820bcbd8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112945",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112945",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use `pread` when `preadv2` with `RWF_NOWAIT` cannot be used, and recognize `EPERM` from it",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/104634 Related: https://github.com/ClickHouse/ClickHouse/issues/49149 Related: https://github.com/ClickHouse/ClickHouse/issues/39753 The `pread_threadpool` read method hands every read off to a thread pool, unless the data is already in the page cache, which it checks with the `preadv2` system call and the `RWF_NOWAIT` flag. Two things can go wrong with that check, and both were handled badly. **`EPERM` was not recognized.** It is what a `seccomp` profile of a container runtime answers for a system call that is not in its allow list. `ThreadPoolReader::submit` handed the read off to the thread pool for `ENOSYS` and `EOPNOTSUPP`, but let `EPERM` through to the throw, failing the query with `CANNOT_READ_FROM_FILE_DESCRIPTOR`. This is the signature reported in #49149. **The check was never verified in advance.** `hasBugInPreadV2` only compared the kernel version, and nothing else was checked, so on a system where the check cannot work `pread_threadpool` kept paying for a thread pool hand-off on every read - including the reads that only had to copy the data from the page cache. In #104634, on Amazon Linux 2 (kernel 5.10), the profile shows `ThreadPoolReaderPageCacheMiss` 1,971,514 out of `LocalThreadPoolJobs` 1,972,484 - every read went to the pool - while the device only moved ~22 GB of the 258 GB the file descriptor delivered, i.e. ~90% of the data was in the page cache and still paid for the hand-off. The system call is now probed once, before it is used, by `preadNoWaitUnavailableReason`. The probe passes an invalid file descriptor on purpose: `seccomp` filters and the system call table are consulted before the descriptor is looked up, so an available system call answers `EBADF` without reading anything, while a blocked one answers `EPERM` or `ENOSYS`. When it says the check cannot be used, `applySettingsQuirks` switches the default value of `local_filesystem_read_method` from `pread_threadpool` to `pread` at start time, and says why in the server log. Nothing downstream has to know: the reader, the userspace page cache eligibility in `DiskLocal::prepareRead` and the prefetched read pool all see a plain `pread` setting. As with the other settings quirks, a value set explicitly - in the configuration, or with `SET` at runtime - is left alone. Such a session keeps `pread_threadpool` and keeps paying for the hand-off, which is what happens today. The switch is a property of the host, so the setting is left marked as unchanged: only the changed settings are serialized into the query the initiator sends to the remote shards, and a host that cannot use the system call must not impose `pread` on the shards that can. An explicitly requested value stays changed and is still sent. The per-read `errno` handling in `ThreadPoolReader::submit` is kept, with `EPERM` added to it: the probe answers for the system call, but a particular filesystem can still reject the flag (`tmpfs` answers `EOPNOTSUPP`, for example), and such a read is handed off to the thread pool instead of failing the query. ### How it was tested `preadv2` was rejected the way a container runtime does it, with a `seccomp` filter installed by a small wrapper (`SECCOMP_RET_ERRNO`), and an old kernel was simulated with `setarch --uname-2.6`. Reading a 2 million row `MergeTree` table with the default `local_filesystem_read_method`, before (the released 26.7.1 binary) and after: | | before | after | |---|---|---| | no filter | `ThreadPoolReaderPageCacheHit` 39, `LocalThreadPoolJobs` 90 | `ThreadPoolReaderPageCacheHit` 34, `ThreadPoolReaderPageCacheMiss` 3, `LocalThreadPoolJobs` 93 - the read method stays `pread_threadpool` | | `preadv2` → `EPERM` | `Code: 74 ... errno: 1, Operation not permitted (CANNOT_READ_FROM_FILE_DESCRIPTOR)`, already while attaching the table | the read method is `pread`, the query succeeds, no `ThreadPoolReader` events | | `preadv2` → `ENOSYS` | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 - every read to the pool | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 80 | | kernel reported as older than 5.11 | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 81 | In all three rejected cases the reason is in the log, for example: ``` <Warning> SettingsQuirks: The default value of local_filesystem_read_method has been switched from 'pread_threadpool' to 'pread' (you can explicitly set it back still), because the `preadv2` system call is not available (the probe with an invalid file descriptor answered errno: 1, strerror: Operation not permitted instead of `EBADF`); it is typically rejected by a `seccomp` profile of a container runtime, and can be allowed in the runtime configuration. ... ``` An explicitly requested `local_filesystem_read_method = 'pread_threadpool'` is kept, on every one of those systems, and this is where the per-read `EPERM` handling earns its place: under the `EPERM` filter the same query fails on the released binary and succeeds here, with `ThreadPoolReaderPageCacheMiss` 45 and `LocalThreadPoolJobs` 135 - every read handed off to the pool, which is the documented cost of asking for it there. Under the same filters, `system.settings` reports `local_filesystem_read_method = 'pread'` with `changed = 0`, so nothing is forwarded to the remote shards, while an explicitly requested `pread_threadpool` reports `changed = 1` and is still sent. Unit tests cover the `errno` classification, the probe's `EBADF` contract, and the quirk itself (the default is switched exactly when the probe says the system call is unusable, the switched value stays out of `Settings::changes()`, and an explicitly set value is never switched). An automated end-to-end test would need an instance with a restrictive `seccomp` profile, which the integration test framework cannot express today - it starts every instance with `seccomp:unconfined`. ### Documented behavior impact On a system where the page cache cannot be checked without waiting for the disk - Linux older than 5.11, a sandbox that rejects `preadv2` with an error code, and systems other than Linux, where `preadv2` does not exist and the method never checked the page cache in the first place - the default value of `local_filesystem_read_method` becomes `pread` instead of `pread_threadpool`, so local reads are performed in the calling thread. This includes the reads with `O_DIRECT`, which never look at the page cache and do not need the check; the read method is now resolved once, for the whole server, so they follow the same value. Nothing changes on a supported system, and nothing changes for an explicitly configured read method. The setting description in `src/Core/Settings.cpp`, from which the docs are generated, is updated in this PR. A `seccomp` profile that terminates the process instead of rejecting the system call with an error code cannot be detected: the startup probe is itself a `preadv2` call, so such a profile kills the server there. On `master` it kills it at the first read instead, under the same default read method. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `pread_threadpool` read method needs the `preadv2` system call with the `RWF_NOWAIT` flag to read the data that is already in the page cache without handing the read off to a thread pool. It is now checked at start time whether that system call can be used, and if it cannot - the Linux kernel is older than 5.11, or a `seccomp` profile of a container runtime rejects the system call - the default value of `local_filesystem_read_method` is switched to `pread`, and the reason is reported in the server log. Previously, every read paid for a thread pool hand-off on such systems, and a `seccomp` profile that answers `EPERM` made queries fail with `CANNOT_READ_FROM_FILE_DESCRIPTOR`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112945",
          "createdAt": "2026-08-01T20:09:06Z",
          "updatedAt": "2026-08-13T13:09:33Z",
          "timestamp": "2026-08-13T13:09:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 26
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:dbe6cde7b49db085cf7c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:103483",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:103483",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve sanitizer robustness in parser and UTF-8 subsequence paths",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> This follows up on sanitizer failures discovered while running CI for https://github.com/ClickHouse/ClickHouse/pull/99740. The branch keeps the original `getURLScheme` empty-input guard from that work and adds two more hardening fixes that became visible after the first fix: - guard `Lexer::nextToken` `max_query_size` check against null input pointers (UBSan path); - return early in `HasSubsequenceImpl::hasSubsequenceUTF8` for empty haystack before first UTF-8 decode (MSan path). It also adds focused regression coverage: - parser gtest for null lexer input with `max_query_size`; - stateless SQL test for empty-haystack UTF-8 subsequence behavior. Original PR where issues were discovered: https://github.com/ClickHouse/ClickHouse/pull/99740 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved sanitizer robustness in parser and string-function edge cases: `protocol` now handles empty input safely in `getURLScheme`, `Lexer::nextToken` no longer performs null-pointer arithmetic in `max_query_size` checks, and UTF-8 subsequence evaluation returns early on empty haystacks before decoding. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-b9b84d83-fd58-456b-97e2-c3386af36d84\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-b9b84d83-fd58-456b-97e2-c3386af36d84\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/103483",
          "createdAt": "2026-04-23T23:58:04Z",
          "updatedAt": "2026-08-13T13:09:23Z",
          "timestamp": "2026-08-13T13:09:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [],
          "author": "yakov-olkhovskiy",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f413b1f61af69c5b0183",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100371",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100371",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `borrow_from_cache` object storage and `memory` metadata types",
          "text": "Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache using ephemeral `FileSegment`s. Each stored object is backed by a cache segment held alive via `FileSegmentsHolder`; when released, the cache reclaims the space. Add a new metadata type `memory` that keeps all file-to-blob mappings and directory structure entirely in memory with no persistence. On server restart all data is lost, which is acceptable for the intended use case of temporary tables. Configuration example: ``` disk(type=object_storage, object_storage_type='borrow_from_cache', metadata_type='memory', cache_name='some_cache') ``` ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A new object storage type and metadata type suitable for temporary tables. Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache and holds it from eviction. Add a new metadata storage type `memory` that keeps the mapping in memory. For example, this could allow creating temporary tables with custom engines in the Cloud without using S3.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100371",
          "createdAt": "2026-03-22T15:18:44Z",
          "updatedAt": "2026-08-13T13:09:05Z",
          "timestamp": "2026-08-13T13:09:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 32
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:739ede307c50b24ac16a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:91062",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:91062",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add an experimental regex-free glob parser",
          "text": "Introduces `GlobAST`, a regex-free glob parser, and wires it into the file and object-storage listing paths. It is **off by default** (the experimental `use_glob_ast_parser` setting), so `master` behavior is unchanged — the goal is to land the parser and all the wiring now so that enabling it later is a one-setting switch. The legacy engine builds a regex (`makeRegexpPatternFromGlobs` + `re2::RE2`), which has noticeable downsides: - lists whole prefixes for enum globs (#73333) - blows up to megabyte regexes on numeric ranges (#43456) - has no formal grammar (#80950) - diverges from POSIX shell semantics in places (e.g. stripping single-element brace groups like `{a}`/`{-}`). `GlobAST` parses a pattern once into typed expressions (constant, `?`/`*`/`**`, `{M..N}` range, `{a,b,…}` enum) and matches/expands directly — ranges by numeric bounds checks, enum-only globs by expanding to concrete keys. A `GlobMatcher` front-end selects the new or legacy backend per the setting, so call sites are unchanged; the grammar is documented in `parseGlobs.h`. Tested by unit tests, a differential fuzzer against the legacy matcher (`GlobASTLegacyMatchFuzz`), and a stateless parity test. Related: https://github.com/ClickHouse/ClickHouse/issues/73333 Related: https://github.com/ClickHouse/ClickHouse/issues/43456 Related: https://github.com/ClickHouse/ClickHouse/issues/80950 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental setting `use_glob_ast_parser` (default `false`) that switches glob matching for the `file`/`s3`/object-storage listing paths to a new regex-free parser (`GlobAST`), avoiding regex blow-up on large numeric ranges and listing fewer keys for brace-enumeration globs.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/91062",
          "createdAt": "2025-11-27T22:07:46Z",
          "updatedAt": "2026-08-13T13:09:02Z",
          "timestamp": "2026-08-13T13:09:02Z",
          "metrics": {
            "reactions": 1,
            "comments": 14
          },
          "labels": [
            "pr-experimental",
            "hold"
          ],
          "author": "thevar1able",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cddbb58ee4cd8a43c119",
        "signalId": "github:ClickHouse/ClickHouse:issue:111206",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:111206",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Automatic parallel replicas (parallel_replicas_min_number_of_rows_per_replica > 0) fails RIGHT JOIN with NOT_FOUND_COLUMN_IN_BLOCK",
          "text": "**Describe what's wrong** With automatic parallel replicas — `parallel_replicas_min_number_of_rows_per_replica > 0` — a valid `RIGHT JOIN` that selects a left-table column fails with `10 NOT_FOUND_COLUMN_IN_BLOCK`: the left table's column is resolved against the **right** table. The same query works with plain parallel replicas (`parallel_replicas_min_number_of_rows_per_replica = 0`) and without parallel replicas. **Does it reproduce on the most recent release?** Reproduces on `26.7.1.408` and near-HEAD master `3090a4fc` (`26.7.1.653`). **How to reproduce** Any cluster works (below: the stateless-test 3-replica localhost cluster; `parallel_replicas_for_non_replicated_merge_tree` only because the tables are non-replicated). ```sql CREATE TABLE tl (k Int32, a Int32) ENGINE = MergeTree ORDER BY k; CREATE TABLE tr (k Int32, ver Int32) ENGINE = MergeTree ORDER BY k; INSERT INTO tl SELECT number, number FROM numbers(1000); INSERT INTO tr SELECT number, number FROM numbers(500); -- OK: 500 rows SELECT r.ver, (l.a + 2) FROM tl AS l RIGHT JOIN tr AS r USING (k) SETTINGS enable_analyzer = 1; -- Code: 10. DB::Exception: Column `a` not found in table default.tr SELECT r.ver, (l.a + 2) FROM tl AS l RIGHT JOIN tr AS r USING (k) SETTINGS enable_analyzer = 1, enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_min_number_of_rows_per_replica = 100; ``` **Characterization** - Only `RIGHT JOIN` fails; `LEFT JOIN` and `INNER JOIN` are correct under the same settings. - `ON l.k = r.k` form fails the same way as `USING (k)`; an `ON` clause with an added one-sided condition (`ON l.id = r.id AND l.v != 64`) fails identically (`Column ts not found in table ...` for whichever left column the projection references). - `parallel_replicas_min_number_of_rows_per_replica = 0` (non-automatic mode): correct. - Deterministic (3/3 runs). The row-estimation path of automatic mode appears to build/propagate a header for the wrong side of a `RIGHT JOIN`. Found by the optimizer-tester topology differential (`tests/optimizer_tester`, branch `optimizer-tester-framework`) the first time it varied the `parallel_replicas_*` sub-settings — the fixed configuration used before could not reach automatic mode. Related: https://github.com/ClickHouse/ClickHouse/commit/4531517ba0a5",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/111206",
          "createdAt": "2026-07-21T10:31:18Z",
          "updatedAt": "2026-08-13T13:08:22Z",
          "timestamp": "2026-08-13T13:08:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "bug",
            "comp-joins",
            "comp-parallel-replicas"
          ],
          "author": "zlareb1",
          "state": "closed",
          "assignees": [
            "devcrafter"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:6f1ad2cb9eebd4d076b9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:88234",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:88234",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add Rewrite rules",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental support for defining custom query rewrite rules. Closes: https://github.com/ClickHouse/ClickHouse/issues/80084 ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds a new query-rewriting/rejection pipeline step in `executeQuery` (gated by `query_rules`) plus persistent rule storage (local or ZooKeeper/Keeper), so mistakes could affect query correctness when enabled and introduce new background reload behavior. > > **Overview** > Adds **Query Rewrite Rules**: new `CREATE RULE`/`ALTER RULE`/`DROP RULE` statements that can rewrite matched queries or reject them with a message. > > Rules are persisted via a new `RewriteRules` subsystem with configurable storage (`query_rules_storage` on disk or ZooKeeper/Keeper, including reload/watch support), exposed through `system.query_rules` and `system.query_rules_log`, and enforced via new access grants (`CREATE_RULE`, `ALTER_RULE`, `DROP_RULE`). > > Query execution now optionally applies these rules by traversing the parsed AST before normal processing (enabled by the new `query_rules` setting), with new error codes and coverage via added stateless tests and documentation. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit fd232f5e9d1b6946b568a8402bba477ee7728c8a. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/88234",
          "createdAt": "2025-10-08T11:03:24Z",
          "updatedAt": "2026-08-13T13:08:12Z",
          "timestamp": "2026-08-13T13:08:12Z",
          "metrics": {
            "reactions": 1,
            "comments": 61
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-experimental"
          ],
          "author": "hinata34",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3d6e3cf9a371a9be14a6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113482",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113482",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #109932 to 26.6: Fix crash in direct JOIN over MergeTree with PREWHERE (shared PrewhereInfo corrupted by column pruning)",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/109932 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31008588537/job/92314616034)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113482",
          "createdAt": "2026-08-05T13:25:04Z",
          "updatedAt": "2026-08-13T13:07:48Z",
          "timestamp": "2026-08-13T13:07:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "diegomestre2",
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bebe5fb6e232ecafce8e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114188",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114188",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Ignore redundant parentheses in stored table definitions",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/110833 Related: https://github.com/ClickHouse/ClickHouse/pull/92340 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `ATTACH PARTITION FROM`, `REPLACE PARTITION`, `MOVE PARTITION TO TABLE` and adding a `ReplicatedMergeTree` replica failing with `Tables have different ...`, `METADATA_MISMATCH` or `INCOMPATIBLE_COLUMNS` for tables whose definitions were written with redundant parentheses, such as `PARTITION BY (a)`, `ORDER BY (b)`, `INDEX ix (b * c) TYPE minmax`, `PROJECTION p (SELECT (b) ...)`, `CONSTRAINT c CHECK (a > 0)`, `TTL (d + INTERVAL 10 YEAR)` or `DEFAULT (a + 1)`. ### Description This is an alternative to https://github.com/ClickHouse/ClickHouse/pull/110833 that contains only the parentheses normalization, so that it is easy to backport. It does not touch `getTreeHash` and does not change how definitions are compared - they are still compared as text. The definition expressions of a table (keys, `TTL`, indices, projections, constraints, column defaults) are stored in ZooKeeper as text and compared as text. #92340 made the AST remember whether the user wrote redundant parentheses around an expression, so the same definition now has two spellings, and every comparison of a stored definition started to reject one against the other: ```sql CREATE TABLE dst (x UInt64, y UInt64) ENGINE = MergeTree PARTITION BY x ORDER BY y; CREATE TABLE src (x UInt64, y UInt64) ENGINE = MergeTree PARTITION BY (x) ORDER BY y; INSERT INTO src VALUES (1, 1); ALTER TABLE dst ATTACH PARTITION 1 FROM src; -- Code: 36. DB::Exception: Tables have different partition key. (BAD_ARGUMENTS) ``` An upgrade or a mixed-version cluster is not needed: both tables above are created by the same binary. Measured on released binaries, 25.8, 26.3 and 26.4 accept this and 26.5, 26.6 and 26.7 reject it, which matches the branches that contain `b38a892dfbfab6`. The fix is a single formatting mode. `IAST::FormatSettings::ignore_redundant_parentheses` suppresses the parentheses of the `parenthesized` flag - the one place that emits them, `decideParensEmission`, is skipped - and `IAST::formatIgnoringRedundantParentheses` is the one-line formatter built on it. Because the formatter visits everything it prints, this covers nested parentheses (`PARTITION BY (x + (1))`) and every kind of definition, not just the top level of a key. It is used where a stored definition is serialized or compared: - `ReplicatedMergeTreeTableMetadata::formattedAST` / `formattedASTNormalized` (the ZooKeeper `/metadata` payload and the replica-join comparison). This replaces the partial `stripArtificialParens` workaround added by #92340, which only cleared the flag on an expression list and its immediate children, and it no longer has to clone the AST. - `IndicesDescription::explicitToString` / `allToString`, `ProjectionsDescription::toString`, `ConstraintsDescription::toString` and `ColumnDescription::writeText` - the serializers of the `indices`, `projections`, `constraints` and `columns` parts of the replicated metadata. `ColumnsDescription::operator==` compares two sets of columns through the same serializer. - `MergeTreeData::checkStructureAndGetMergeTreeData` - the `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` structure gate (sorting/partition/primary keys, secondary indices, projections). - `StorageReplicatedMergeTree::alter` - the definitions an `ALTER` writes back into Keeper, so an `ALTER` never publishes a parenthesized definition either. - `AlterCommand::isTTLAlter` (whether restating a `TTL` schedules a `MATERIALIZE TTL` mutation) and the `MODIFY ORDER BY` no-op detection in `AlterCommands::prepare`. The text these produce is exactly what every server produced before #92340, so a new replica and an old one agree in both directions. Nothing else changes: the parser and the formatter are untouched, the table metadata keeps what the user wrote, and `SHOW CREATE` still prints `PARTITION BY (x + (1))`. Definitions that differ in more than the parentheses are still rejected (`a` vs `b`, a different index expression, a different column default). The `columns` payload of `StorageKeeperMap` and `ObjectStorageQueue` also contains these expressions, but both storages now compare it structurally, so a table created by any version keeps working. `StorageKeeperMap` already parses both sides and re-serializes them with the current serializer before comparing. `ObjectStorageQueueTableMetadata` compared the stored string verbatim; it now falls back to comparing the parsed `ColumnsDescription` when the strings differ, so a queue table created by 26.5-26.7 with redundant parentheses in a column `DEFAULT`, `CODEC` or `TTL` is accepted in both directions, and this also un-breaks such a table created by 25.8 and earlier (which is broken on 26.5-26.7 today). Retrying the creation of a replica that failed between the ZooKeeper transaction and saving the local metadata is also covered: `createReplicaAttempt` recognizes an already-created empty replica by comparing its `/metadata` and `/columns`, and now falls back to a structural comparison (`ReplicatedMergeTreeTableMetadata::checkEquals`, parsed `ColumnsDescription`) when the raw strings differ, so such a retry works across the old and new spellings instead of failing with `REPLICA_ALREADY_EXISTS`. Tests: `04836_parenthesized_definitions_attach_partition_from`, `04837_parenthesized_definitions_replicated_metadata` and `04850_parenthesized_definitions_replica_recovery`. Each case in them fails on master with the error it is named after. The unit test `ObjectStorageQueueTableMetadata.ColumnsComparisonIgnoresRedundantParentheses` covers the stored-JSON compatibility of the queue metadata.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114188",
          "createdAt": "2026-08-10T16:56:26Z",
          "updatedAt": "2026-08-13T13:06:54Z",
          "timestamp": "2026-08-13T13:06:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2b55eff33042d357514b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110183",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110183",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use-of-uninitialized-value in WITH FILL suffix over a merge",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107074 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-of-uninitialized-value in `ORDER BY ... WITH FILL ... STALENESS` when the fill suffix generates no rows (for example when the staleness window leaves nothing to fill) and the result is read through a merge of sorted streams (such as a two-shard `Distributed` table). ### Description Found by the AST fuzzer under MSan on top of #107074. `FillingTransform`, when all input chunks are processed, may run the suffix path. If the fill constraints are satisfied but no fill rows are produced (e.g. `WITH FILL ... STALENESS` with an exhausted staleness window), `generateSuffixIfNeeded` returns `true` while the result columns are freshly `cloneEmpty()`'d and carry no data. Previously the transform still emitted a 0-row chunk built from those empty columns. A downstream `MergingSortedTransform` (as used when reading from a two-shard `Distributed` table) then built a sort cursor over that empty chunk and compared row 0, reading past the end of the empty column. Fix: do not emit the suffix chunk when it has no rows. Minimal reproducer (needs a two-shard merge on the initiator): ```sql CREATE TABLE m (key Int) ENGINE = Memory; INSERT INTO m VALUES (100); CREATE TABLE d2 AS m ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), m); SELECT _shard_num FROM d2 ORDER BY _shard_num ASC WITH FILL TO 46 STALENESS 1; ``` CI finding: `AST fuzzer (amd_msan)`, report https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107074&sha=3ec54c88d4eb535a5d644fc1ab91af31d717f9a5&name_0=PR&name_1=AST%20fuzzer%20%28amd_msan%29 MSan use-of-uninitialized-value in `ColumnVector<UInt32>::doCompareAt` (`MergingSortedAlgorithm::consume`), origin `FillingTransform::initColumns` (`cloneEmpty`) via the suffix path. Verified on a local `amd_msan` build: reproduces before the fix (server aborts), clean after; the new regression test returns `1\\n2`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110183",
          "createdAt": "2026-07-12T21:10:52Z",
          "updatedAt": "2026-08-13T13:06:45Z",
          "timestamp": "2026-08-13T13:06:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:d96c58dcb2182b0098e0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114609",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114609",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix flaky 04780_json_subcolumn_index_match_not_quadratic",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113289 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Requested in https://github.com/ClickHouse/ClickHouse/pull/112705#issuecomment-5276141544. Related: https://github.com/ClickHouse/ClickHouse/pull/113289, which added the test. Two independent test defects; the engine fix is fine. **1. The allocation oracle measured deferred flushes, not the query.** `ProfileEvents['MemoryAllocatedWithoutCheckBytes']` accumulates whatever `CurrentMemoryTracker` flushes during the query. A flush fires when the thread's deferred balance crosses `max_untracked_memory`, or when the shared per-CPU budget cannot cover it, so neighbouring threads decide the value. All four failures were `parallel` jobs, bytes identical across asan and tsan, seven sibling parallel jobs green on each failing commit, master 283 OK / 0 FAIL. Varying only that setting: | `max_untracked_memory` | increments | mean bytes/record | |---|---|---| | 0 | 29201 | 3766 | | 1Mi (CI value) | 103 | 222362 | | 16Mi | 103 | 222362 | The count saturates at 103 from 1Mi upward, so above that the per-thread threshold is not the trigger and raising the setting cannot help. Pinning it to 0 gives each allocation its own record. The measured statement now follows a trivial one on the same connection, whose detach drains the balance deferred before it. **2. The last `EXPLAIN` asserted a master-only default.** `Parts: 0 | Granules: 0` comes from `describeActions`, gated by `actions`, which only master force-enables (`set_default_pretty_explain_settings`, absent from 26.5 and 26.6), so the two open backports of #113289 fail deterministically. Fixed by requesting `actions = 0` and dropping that line. `Granules: 0/1000`, the pruning under test, comes from `describeIndexes` and is untouched. I could not confirm the 150% bound is too tight under sanitizers: the failing ratios were 54x and 190x, and a bound admitting those would also admit the 2.8x regression, so it stays. Validation: 50/50 runs pass with randomization on and off. On a pre-#113289 binary the new oracle reddens at 1296x-2386x on all four arms, sharper than the 2.8x-3.0x before.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114609",
          "createdAt": "2026-08-13T09:13:35Z",
          "updatedAt": "2026-08-13T13:06:35Z",
          "timestamp": "2026-08-13T13:06:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1bb99b80b5a26017851d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114182",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114182",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix top-K dynamic filtering for empty Tuple columns",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/113419 Related: https://github.com/ClickHouse/ClickHouse/pull/113406 `Tuple()` can be sorted, but the scalar comparison functions used by `__topKFilter` reject zero-sized tuples. This change skips top-K dynamic filtering when the sort column is directly `Tuple()`. A nested empty `Tuple`, including through `Nullable`, remains eligible for the optimization and is routed through the general column-comparison path. Composite types with their own comparison implementation, such as `Array(Tuple())`, are not misclassified and keep vectorized filtering. Local validation used the final SQL test blob against the existing debug binary: all 10 Praktika stateless runs passed, `check_cpp` passed, and the produced output matched the reference with an empty diff. These were local checks, not upstream CI. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `ORDER BY ... LIMIT` queries on `Tuple()` columns raising a `NOT_IMPLEMENTED` exception when top-K dynamic filtering is enabled. Direct empty tuples now skip the optimization, while supported nested composite types continue to use top-K filtering.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114182",
          "createdAt": "2026-08-10T15:50:01Z",
          "updatedAt": "2026-08-13T13:05:26Z",
          "timestamp": "2026-08-13T13:05:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "Boulea7",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:ad67fe29998d69960ae7",
        "signalId": "github:ClickHouse/ClickHouse:issue:109216",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:109216",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Wrong DISTINCT results (and debug abort) with partial_merge JOIN + optimize_distinct_in_order — Equal values are not contiguous",
          "text": "# `Logical error: 'Equal values are not contiguous within the range assumed to be sorted'` — `DISTINCT` in order over `partial_merge` JOIN (wrong results in release) ## Summary With `join_algorithm = 'partial_merge'` (or `'prefer_partial_merge'`), the query plan assumes the join preserves the left stream's sort order and applies the pre-`DISTINCT` stage as `DistinctSortedStreamTransform` (`optimize_distinct_in_order`, default on). But `PartialMergeJoin` matches the left stream against the right side **per right block**: when the right table has more than one block/part, the left key ranges are re-emitted for each right block, so equal left-key values are no longer contiguous in the joined stream. - **Debug builds**: `chassert` in `findEqualRangeEndAssumeSorted.h:88` fires → `Logical error: 'Equal values are not contiguous within the range assumed to be sorted'` → abort. Stack: `DistinctSortedStreamTransform::transform → getEqualRangeEndAssumeSorted → checkEqualRangeEndAssumeSorted (ColumnVector<UInt32>)`. - **Release builds**: the check is compiled out, so `DistinctSortedStreamTransform` computes equal-ranges on a stream that violates its precondition — **`DISTINCT` can silently return duplicate rows** (wrong result). Found by the AST-fuzzer correctness-oracle campaign (PR #99980): two independent fleet hits (2026-06-29 `DISTINCT … FULL OUTER JOIN … QUALIFY`, 2026-06-30 `DISTINCT … INNER JOIN … ORDER BY`), minimized from the second snapshot. ## Minimal repro Flaky by nature (depends on chunk arrival order across the two right-side parts; `max_threads = 4` observed rate ≈ 40–50%, single attempt often enough on a debug build). Run a few times: ```sql CREATE TABLE t1 (x UInt32, y UInt64) ENGINE = MergeTree ORDER BY (x, y); CREATE TABLE t2 (x UInt32, y UInt64) ENGINE = MergeTree ORDER BY (x, y); INSERT INTO t1 VALUES (0,0),(1,10),(2,20),(3,30),(4,40); INSERT INTO t2 VALUES (2,21),(2,22),(4,41); INSERT INTO t2 VALUES (0,0),(4,42),(5,50); -- second part is load-bearing SET join_algorithm = 'prefer_partial_merge'; SELECT DISTINCT t1.*, t2.* FROM t1 INNER JOIN t2 ON intDiv(t2.y, 2147483647) = toUInt64(t1.x); ``` Also reproduces in `clickhouse-local` with the same script (abort, exit 134). The original fuzzed query had a baroque `ON and(if(...), key = key)` condition and an `ORDER BY … DESC NULLS LAST` — neither is needed (verified: pure-equality `ON` hits, no-`ORDER BY` hits). ## Mechanism (EXPLAIN PIPELINE) `prefer_partial_merge` (bug path): ``` DistinctTransform DistinctSortedStreamTransform × 2 <-- assumes sorted-by-left-prefix stream JoiningTransform × 2 2 → 1 <-- PartialMergeJoin MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) × 2 <-- left read in order FillingRightJoinSide MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) ``` `hash` (correct path): plain `DistinctTransform`, no in-order read, no sorted assumption. The left read is in `(x, y)` order and the plan carries that sort property through the join to justify `DistinctSortedStreamTransform`. `PartialMergeJoin` breaks the property: it processes the right side block-by-block and re-emits matching left ranges for each right block, so a left key that matches rows in two right blocks appears in two separate runs. Controls (each verified with a 10-attempt harness against the snapshot server): - `join_algorithm = 'hash'` → never reproduces (plain `DistinctTransform`). - `optimize_distinct_in_order = 0` → does not reproduce (MISS ×10). - `join_algorithm = 'partial_merge'` (strict) → reproduces, same as `prefer_partial_merge`. - `SELECT DISTINCT t1.*` only (no `t2` columns) → does not reproduce (MISS ×10); the `DISTINCT` set must include right-side columns. - Baroque original `ON and(if(...), key = key)` and the `ORDER BY … DESC NULLS LAST` → both unnecessary (pure-equality, no-`ORDER BY` variant reproduces). ## Environment - HEAD `01e08fd49182` (2026-06-28 master merge), debug build, aarch64. - Snapshot: `tmp/fuzz_lab/runs/merged7/crashes/inst-06-20260630-184002` (original), also `inst-12-20260629-115230` (same family, first sighting), `inst-05-20260701-012001` (`Sort order of blocks violated` — likely same root: a sort-order property claimed across `PartialMergeJoin` and violated downstream in a merge-sort transform instead of DISTINCT). ## Suggested fix direction `PartialMergeJoin` (and any join that rescans the left side per right block) must not report the left input's sort description as preserved on its output — the plan should either drop the sort property across it (forcing plain `DistinctTransform` / a re-sort before order-dependent consumers) or the join should be excluded from `optimize_distinct_in_order` / read-in-order propagation.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/109216",
          "createdAt": "2026-07-02T19:25:34Z",
          "updatedAt": "2026-08-13T13:05:26Z",
          "timestamp": "2026-08-13T13:05:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "fuzz",
            "sqlancer"
          ],
          "author": "qoega",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:5e113bb82717a4e89a78",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113059",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113059",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Collect SQL stacktraces on the hung-check and server-died abort paths",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/42701 Related: https://github.com/ClickHouse/ClickHouse/pull/112265 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ### Description Related: #42701, #112265 On an ASan build, a stateless run aborting on the hung check records no stack of the hung server: ``` Hung check failed: server is not responding Cannot collect C stacktraces under ASan: debugger attach is disabled. ``` Two things combine. `print_c_stacktraces` declines to attach lldb on ASan builds, because ptrace disables LeakSanitizer; that refusal is correct and stays. And `print_sql_stacktraces`, needing no debugger, was unreachable: its pre-check `check_server_liveness` probes **HTTP** (`http_port`, default 8123), while it collects over native **TCP** (`args.client --port=`, default 9000). Different listeners, different ports, so this signature (alive, not answering HTTP, TCP still serving) failed the gate. The hung-check abort site did not call it at all. This drops the mismatched pre-check and lets the collector be its own liveness test: it is already bounded (`timeout=30`) and reports failure as one trimmed line, so a dead socket costs at most 30 s and cannot re-emit the `Code: 210` tracebacks that motivated the pre-check. The dump is added to the three abort sites that had only the C path: hung check, server died, and the startup check. The stateless job keeps attaching the dump to its result and additionally clears any left by a previous job in the same workspace, so an aborted run cannot upload a stale dump as its own. #114143 has since added the same attachment upstream; this replaces it with the equivalent helper rather than attaching twice. It fixes no hang and does not restore C++ stacks on ASan. A server alive but not answering HTTP now yields the full `system.stack_trace` view with per-thread `query_id`, identifying the wedged query; one dead on both transports records \"tried, got nothing\" instead of silence. Validated against a live server: with HTTP dead and TCP live, master skips and writes nothing, while this branch writes a `sql_stacktraces.log` carrying `thread_name` and `query_id`. Green runs are unaffected. [Prompting report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=1ce152efafc5b00bf31eb7a0a64757fecbbc6e4a&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20distributed%20plan%2C%20parallel%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113059",
          "createdAt": "2026-08-03T06:03:46Z",
          "updatedAt": "2026-08-13T13:05:22Z",
          "timestamp": "2026-08-13T13:05:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f49c6d2235ece213d7b5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112152",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112152",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the function name of a stack frame attributed to a libc++ `__functional` header",
          "text": "Caused by: https://github.com/ClickHouse/ClickHouse/pull/57201 Related: https://github.com/ClickHouse/ClickHouse/pull/100419 `collapseDemangledNames` replaced a stack frame's function name with `?` whenever the frame's source location was a file in a directory ending in `functional`. The intent was to hide `std::function` plumbing frames, whose demangled names spell out the whole captured lambda type and say nothing the neighbouring frames do not already say. But the file of a frame is the source line the *instruction* maps to, which is not necessarily where the enclosing function is defined. An ordinary function can have individual instructions attributed to a libc++ `__functional` header - an inlined `std::function` operation, or compiler-generated code reported with line 0 - and then its name was dropped too, which loses the only useful part of the frame. This PR requires the symbol to name the `std::function` plumbing as well: the type-erasing wrappers (`std::__function::__func`, `__value_func`, `__policy_func`, ...) and the type erasure of `std::function` itself - its `operator()`, its copy, move and callable-taking constructors and assignment operators, and its destructor. The ordinary members of `std::function` (`swap`, `target_type`, `operator bool`, the `nullptr` reset `operator=(std::nullptr_t)`, the empty-constructing `function()` / `function(std::nullptr_t)`, ...) do work of their own, so they keep their names too. Those are still displayed as `?`, exactly as before; every other frame keeps its name - not only an ordinary `DB` function, but also a meaningful libc++ symbol that happens to live in a `__functional` header, such as `std::hash<String>::operator()` from `__functional/hash.h`, or the generic invocation helpers `std::invoke` / `std::__invoke` / `std::mem_fn`, which are not `std::function`-specific and whose frames can name the callable they dispatch to. This is much more likely in a build with ThinLTO enabled, where `std::function` calls are inlined across translation units, so it went unnoticed for a long time: `amd_cfi` is the only integration-test build with ThinLTO on, and its `test_crash_log` failure is what surfaced it. In that build, the frame that actually terminated the server was displayed as ``` 8. contrib/llvm-project/libcxx/include/__functional/function.h:0:7: ? @ 0x1c01f41f ``` both in the fatal log and in `trace_full` in `system.crash_log`, where `0x1c01f41f` is inside `DB::executeQuery` (the symbol resolves correctly - only the display suppressed it). Official release builds also use ThinLTO, so the same frames were being anonymised for users. `collapseDemangledNames` becomes a static member of `StackTrace` so that the heuristic is covered by a unit test in every build, rather than only by the weekly ThinLTO job. Fixes `test_crash_log/test.py::test_crash_log_extra_fields[terminate_with_exception-trace_full]` and `[terminate_with_std_exception-trace_full]`, seen in https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b05161aa67d75ad84151a392f3e87156efd42842&name_0=WeeklyCFI&name_1=Integration%20tests%20%28amd_cfi%2C%202%2F4%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a stack frame being displayed as `?` instead of its function name in fatal log messages and in the `trace_full` column of `system.crash_log`. It affected frames of ordinary functions that have code attributed to a libc++ `__functional` header, which is common in release builds.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112152",
          "createdAt": "2026-07-27T17:43:33Z",
          "updatedAt": "2026-08-13T13:05:01Z",
          "timestamp": "2026-08-13T13:05:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 16
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0a821010049f12347cfc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109225",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109225",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong results with parallel_hash JOIN and read-in-order-through-join",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/issues/109216 Related: https://github.com/ClickHouse/ClickHouse/pull/110671 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results (silently dropped or mis-grouped rows) when `join_algorithm = 'parallel_hash'` is combined with a sorted consumer such as `optimize_aggregation_in_order`, `optimize_distinct_in_order` or `LIMIT BY`. With several join slots and a single-level hash map, `ConcurrentHashJoin` scatters the left block across slots, so the read-in-order-through-join optimization must no longer advertise the left sort order in that case. ### Description The `partial_merge` / `prefer_partial_merge` half of this PR has since been fixed on master by #110671, which added `IJoin::preservesLeftBlockOrder()` (defaulting to `true`) plus the `MergeJoin` / `JoinSwitcher` overrides and the `findReadingStep` gate. What remains here is a different carrier, and it is still a live wrong result on master. **`parallel_hash` (`ConcurrentHashJoin`) with several slots and a single-level map.** It inherits the `true` default from master, but `chooseMethod` leaves a key that materializes to one or two bytes (`key8` / `key16`) single-level - wider keys, including string and fixed-string ones, get a two-level variant. For a single-level map `joinBlock` scatters the left block across slots and `ConcurrentHashJoinResult` emits slot 0, then slot 1, and so on, so equal left-key values stop being contiguous while `findReadingStep` still installs the ordered read. Measured on master (`54dee101`, which contains #110671) against a debug build of this branch: ```sql CREATE TABLE t3 (a UInt32, j UInt8) ENGINE = MergeTree ORDER BY (a, j); CREATE TABLE t4 (j UInt8, v UInt64) ENGINE = MergeTree ORDER BY j; INSERT INTO t3 SELECT intDiv(number, 8)::UInt32, (number % 8)::UInt8 FROM numbers(64); INSERT INTO t4 SELECT (number % 8)::UInt8, number FROM numbers(8); SET join_algorithm = 'parallel_hash', max_threads = 8, optimize_aggregation_in_order = 1, max_bytes_before_external_join = 0, max_bytes_ratio_before_external_join = 0; SELECT a, count() FROM t3 LEFT ALL JOIN t4 ON t3.j = t4.j GROUP BY a ORDER BY a; ``` Master returns `1, 1, 1, 1, 1, 1, 1, 57`; the correct answer is 8 per group, which this branch returns. Ground truth was confirmed three independent ways (`optimize_aggregation_in_order = 0`, `join_algorithm = 'hash'`, `query_plan_read_in_order_through_join = 0`). The fix flips the `IJoin::preservesLeftBlockOrder()` default from `true` (fail-open) to `false` (fail-closed) and makes each join that really does stream the left side through once opt in: `HashJoin`, `DirectKeyValueJoin`, `ConstantJoin`, `PasteJoin` unconditionally, and `ConcurrentHashJoin` only when it does not scatter (`slots == 1 || twoLevelMapIsUsed()`). Flipping the default is what makes the contract hold by property rather than by accident. On master `FullSortingMergeJoin` has no override, so it inherits `true` - it is safe today only because `JoinStepLogical` inserts a `Sorting (Sort Left before JOIN)` step that `findReadingStep` does not descend through. That is a property of the current plan shape, not of the join, so any future plan change would silently reintroduce a wrong result. Under the fail-closed default it is safe by property. Precision was verified in both directions, so the stricter default does not cost the optimization anywhere it was previously correct: a two-level `UInt64` key still reads in order, a single-slot (`max_threads = 1`) `parallel_hash` join still reads in order, and `hash` / `direct` are unchanged. `GraceHashJoin` and `SpillingHashJoin` remain excluded through `hasDelayedBlocks()` as before. `topKThroughJoin.cpp` keeps its explicit `FullSortingMergeJoin` type check for its own mode 2 (a pre-JOIN `Sort` on the preserved input); its comment is updated to say the `preservesLeftBlockOrder()` read already covers that join and the type check is now belt-and-braces. Tests: `04498_distinct_in_order_partial_merge_join` fails on current master on exactly the `parallel_hash` block and passes here, so it is a live regression test rather than a restatement of #110671. `04500_read_in_order_through_constant_join` covers the `ConstantJoin` and `DirectKeyValueJoin` opt-ins, and `04500_limit_by_in_order_partial_merge_join` guards the `LIMIT BY` consumer. Each assertion was verified by mutation: with the corresponding override removed the assertion flips. All of them pin the whole read-in-order trio (`optimize_read_in_order`, `query_plan_read_in_order`, `query_plan_read_in_order_through_join`), since the stateless runner randomizes all three and a drawn `0` would make the plan assertions blind. #109216 is downgraded to `Related:` because the shape it reports (`prefer_partial_merge` + `optimize_distinct_in_order`) no longer reproduces on master after #110671; its reproducer now returns the correct 6 rows over repeated runs. This PR covers the sibling `parallel_hash` carrier of the same class, so it should not auto-close that issue.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109225",
          "createdAt": "2026-07-02T20:34:33Z",
          "updatedAt": "2026-08-13T13:03:48Z",
          "timestamp": "2026-08-13T13:03:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 41
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:78c55615c2765696c5ca",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114090",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114090",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a regression test for duplicate parallel replicas announcements from a self-matching merge() child",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Add a regression test for the parallel replicas coordinator guard trips (`Initiator received more initial requests than there are replicas: replica_num=1`, `Duplicate announcement received for replica number N`) that fired when a `merge()` table function's regex matched the table the outer query reads. #112849 fixed this in code — the `enable_parallel_replicas` clear in `ReadFromMerge::createChildrenPlans` — but its test `04665_merge_table_parallel_replicas_child_plan` never reaches the announcement guards on the pre-fix code: its fixture aborts earlier with `Cannot serialize FutureSetFromSubquery with no query plan`, and it uses the default index granularity, at which the initiator's replica claims every mark range and the followers are cancelled before they announce. So if the `enable_parallel_replicas` clear were ever narrowed (the way the `make_distributed_plan` clear was narrowed to `queryHasSubquerySets`), the signature would regress with `04665` still green. The fixture is the one @groeneai verified against three binaries (see https://github.com/ClickHouse/ClickHouse/pull/102192#issuecomment-5232296945): the outer read and the `merge()` child read of the same table derive the identical `stream_id`, so without the clear one follower builds several read pools under one `replica_num` and announces more than once into the same coordinator. A small `index_granularity` gives the followers enough marks to survive to the announcement; the committed fixture uses 16384 rows with `index_granularity = 16` (~1024 marks). Heavier shapes (500000/128 and 62500/16, both ~3900 marks) exceeded the 180s per-test limit of the flaky check, which runs 50 copies concurrently on a debug build: `system.query_log` from the failed run shows the `parallel_replicas_local_plan = 0` SELECT stalling up to 333s wall at 12.5s CPU with 834 thread-seconds in `NetworkReceiveElapsedMicroseconds` and only 9 coordinator round-trips, followers idle after finishing their physical read in seconds — the round-trips of that mode degrade sharply on an oversubscribed server, so the mark count is kept as low as the reproduction allows (at ~512 marks the local-plan mode stops reproducing, so ~1024 keeps a 2x margin). Verified locally on a 3-replica server: reddens in both `parallel_replicas_local_plan` modes 3/3 runs with the two `enable_parallel_replicas` clears in `StorageMerge.cpp` disabled, passes with them in place, also under 8 concurrent runs. The test asserts the successful result rather than a guard message, since the pre-fix failure surfaces as either of the two adjacent guard messages depending on follower interleaving. To make sure it cannot silently stop exercising the announcement path, it runs with `allow_experimental_parallel_reading_from_replicas = 2` (an unsupported-shape fallback to a plain local read is an error, not a silent success) and additionally asserts `ProfileEvents['ParallelReplicasHandleRequestMicroseconds'] > 0` on the initiator's `system.query_log` entry, like `04545_parallel_replicas_projection_short_circuit_unknown_stream.sql` — a follower must survive past the announcement into the coordinator's request path for it to fire, so the initiator claiming every range and cancelling the followers early fails the test instead of passing it (a coordinator-creation log line alone could not distinguish that). The regex is anchored to the table name so concurrent tests cannot leak into the `merge()`, and the table lives in `default` because a single-argument `merge()` resolves against the default database on each hop. Because `default` is shared by every concurrently running test, the test is a `.sh` so the table name can carry `$CLICKHOUSE_DATABASE`: the first revision used a fixed literal name, and two runs of the test then raced on `CREATE`/`DROP` of the same table (`UNKNOWN_TABLE` / `TABLE_ALREADY_EXISTS`) - the flaky check runs each new test 50 times with `--jobs nproc-1`, and the ordinary parallel jobs repeat newly modified tests. Related: https://github.com/ClickHouse/ClickHouse/pull/112849 Related: https://github.com/ClickHouse/ClickHouse/pull/110972",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114090",
          "createdAt": "2026-08-10T01:00:49Z",
          "updatedAt": "2026-08-13T13:03:38Z",
          "timestamp": "2026-08-13T13:03:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:34fde94cc67b526e0dae",
        "signalId": "github:ClickHouse/ClickHouse:issue:114641",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114641",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Text index: support arbitrary `LIKE` patterns with `tokenizer = 'array'`",
          "text": "### Company or project name ClickHouse Inc. ### Use case A text index with `tokenizer = 'array'` stores the whole column value as a single token, so its dictionary is the set of distinct values of the column, and a `LIKE` filter over that dictionary is exactly the query predicate. Today only patterns of the shape `%needle%`, with an alphanumeric needle, are evaluated using the index. Any other pattern reads the whole column, although the index already contains everything needed to answer it. ```sql CREATE TABLE t ( name String, INDEX idx name TYPE text(tokenizer = 'array') ) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t VALUES ('alpha-service-prod'), ('beta-service-prod'), ('gamma-svc-4999-dev'); ``` | Query | Uses the index today | Why not | |---|---|---| | `SELECT * FROM t WHERE name LIKE '%4999%'` | yes | | | `SELECT * FROM t WHERE name LIKE '%svc-4999%'` | no | punctuation in the needle | | `SELECT * FROM t WHERE name LIKE 'alpha%'` | no | anchored at the start | | `SELECT * FROM t WHERE name LIKE '%prod'` | no | anchored at the end | | `SELECT * FROM t WHERE name LIKE '%alpha%prod%'` | no | more than one needle | | `SELECT * FROM t WHERE name LIKE 'alpha_service%'` | no | `_` wildcard | ### Describe the solution you'd like For `tokenizer = 'array'`, evaluate arbitrary `LIKE` and `ILIKE` patterns using the text index: anchors, punctuation, `_` wildcards, several `%`-separated needles, escaped metacharacters. The existing cost guards should keep working — the minimum required literal length and the limit on how much of the index may be read, falling back to reading the column when the limit is exceeded. ### Describe alternatives you've considered An additional `ngrams(3)` index on the same column answers these patterns, but doubles index storage and write amplification for data the `array` dictionary already describes exactly. ### Additional context Actually, any predicate over the column with an `array` tokenizer can be supported, but it requires non-trivial development. The `LIKE` improvement comes almost for free, so let's start with this. Related: https://github.com/ClickHouse/ClickHouse/pull/98149 Related: https://github.com/ClickHouse/ClickHouse/issues/97723",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114641",
          "createdAt": "2026-08-13T13:03:00Z",
          "updatedAt": "2026-08-13T13:03:06Z",
          "timestamp": "2026-08-13T13:03:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "feature",
            "comp-text-index"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:bc213a5e763f0ab2b928",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114629",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114629",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix DPsub join reordering silently dropping single-table ON-clause filters",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111898 The `dpsub` join-order algorithm silently dropped single-table filter and constant predicates that live in a `JOIN ... ON` clause (e.g. `t1.value = 'x'`), returning extra rows. `greedy` uses a different placement path and was unaffected. The predicates are placed in `collectJoinEdgesMask`. Two placement conditions there each silently dropped such a predicate: - `two_relations` required the whole join step to be exactly two relations, so the predicate was dropped whenever its relation was introduced against an already-multi-relation subplan (e.g. `t1` at the top of `t1 JOIN (t2 JOIN t3)`). - `fromLeft() || fromRight() || fromNone()` dropped the predicate for any single-table filter on a relation whose id is >= 2, because `fromLeft`/`fromRight` test relation ids 0 and 1 specifically (they describe the two inputs of a binary join step, not \"references a single relation\"). The predicate is now attached at the join that introduces its relation (the split whose one side is exactly that relation), and pure constants at the earliest two-relation join. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fixed `dpsub` join-order optimization (`query_plan_optimize_join_order_algorithm = 'dpsub'`) silently dropping single-table filter conditions from a `JOIN ... ON` clause, which could return extra rows.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114629",
          "createdAt": "2026-08-13T12:38:37Z",
          "updatedAt": "2026-08-13T13:02:56Z",
          "timestamp": "2026-08-13T13:02:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "fkastrati",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:3e4a993503e19de1e231",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109005",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109005",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallel full sorting merge join (parallel_full_sorting_merge)",
          "text": "Resolves https://github.com/ClickHouse/ClickHouse/issues/48165 Adds a new `join_algorithm` value `parallel_full_sorting_merge`. `full_sorting_merge` streams both sides and joins them with a single merge, so it keeps memory bounded but runs the merge on one thread — on many-core machines it is often slower than `parallel_hash` even though it uses far less memory. `parallel_full_sorting_merge` keeps the streaming, low-memory profile of a merge join but shards the join by the hash of the join keys into independent per-shard merge joins that run in parallel (up to `max_threads`). A query-plan optimization (`optimizeParallelFullSortingMergeJoin`) switches each side's pre-join `SortingStep` to scatter the rows by the hash of the join keys into a fixed number of shards (`max_threads`) and sort each shard, then marks the join to run shard-by-shard, reusing the existing sharded pipeline (`joinPipelinesYShapedByShards`). Because the partitioning depends only on the join-key values — and `FullSortingMergeJoin` already requires matching key types — equal keys land in the same shard on both sides, so shards join independently. The result is unordered. Unlike the by-primary-key-ranges sharding (`query_plan_join_shard_by_pk_ranges`), this works on unsorted inputs by scattering each side before sorting. Already-sorted inputs are not scattered by this rewrite; they fall back to a single merge join, while in-order MergeTree reads can still be sharded at the source by primary-key ranges. When a side reads a single stream the scatter still produces the fixed shard count on both sides, so joins between inputs with different parallelism (e.g. different part counts) stay co-partitioned. Benchmark, `20M ⋈ 20M` on `UInt64` keys, `max_threads = 8`: | `join_algorithm` | wall | peak RSS | |---|---|---| | `parallel_hash` | 1.27 s | 3201 MB | | `parallel_full_sorting_merge` | **0.52 s** | **974 MB** | | `full_sorting_merge` (single merge) | 0.76 s | 552 MB | i.e. ~2.4× faster and ~3.3× less memory than `parallel_hash`, and ~1.5× faster than the single-threaded merge join. Correctness is verified against the `hash` algorithm for `INNER`/`LEFT`/`RIGHT`/`FULL` joins, many-to-many keys, and `join_use_nulls` in `04492_parallel_full_sorting_merge_join`. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new `join_algorithm` value `parallel_full_sorting_merge`: for hash-compatible equality joins, it shards a full sorting merge join by the hash of the join keys into independent per-shard merge joins running on all threads. It keeps the low, streaming memory usage of a merge join while parallelizing it (in a benchmark, ~2.4x faster and ~3.3x less memory than `parallel_hash`). `ASOF` joins fall back to a single `full_sorting_merge`; hash-incompatible key types (floating-point, `JSON`, `Object`, `Dynamic`) skip only the hash-scatter rewrite and can still be sharded at the source by primary-key ranges when `query_plan_join_shard_by_pk_ranges` is enabled. The result is not ordered. ### Documentation entry for user-facing changes - [x] Documentation is written (the new value is described in the `join_algorithm` setting). <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.592` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109005",
          "createdAt": "2026-07-01T00:38:58Z",
          "updatedAt": "2026-08-13T13:02:56Z",
          "timestamp": "2026-08-13T13:02:56Z",
          "metrics": {
            "reactions": 2,
            "comments": 37
          },
          "labels": [
            "pr-performance",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "m-selmi"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:53a9918e5a35568a65f0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113505",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113505",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "S3 tables engine",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): S3 tables engine catalog for datalakes. Same as https://github.com/ClickHouse/ClickHouse/pull/103220, but with working INSERT",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113505",
          "createdAt": "2026-08-05T14:34:19Z",
          "updatedAt": "2026-08-13T13:02:55Z",
          "timestamp": "2026-08-13T13:02:55Z",
          "metrics": {
            "reactions": 3,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:83eb043bfe36eb434d38",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107091",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107091",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not terminate the server on retryable errors while loading outdated parts",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/106736 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Do not terminate the server when a transient retryable error (such as `MEMORY_LIMIT_EXCEEDED` or a network error) occurs while loading outdated or unexpected data parts in the background. Such errors are not a sign of an inconsistent on-disk set of parts, so the loading task is now retried instead of aborting the whole server. The existing fail-fast behaviour is preserved for genuinely inconsistent state. ### Description Closes #106736 `MergeTreeData::loadOutdatedDataParts()` and the sibling `loadUnexpectedDataParts()` wrap their whole body in a `catch (...)` that calls `std::terminate()`, to fail fast on a genuinely inconsistent on-disk set of parts. The same `catch` also fires for transient runtime errors such as `MEMORY_LIMIT_EXCEEDED` (including the stress test memory fault injector) or network errors, which are not on-disk corruption. This takes the whole server down in the experimental `serverfuzz` stress jobs. `loadDataPartWithRetries()` already classifies and retries such exceptions via `isRetryableException()`, but only around `loadDataPart()` itself; an exception thrown elsewhere in the loader (part removal, ZooKeeper cleanup, rename, the runner machinery, or newly memory-tracked allocations) escapes straight to the terminating `catch`. The fix reuses `isRetryableException()` in both `catch` blocks: for a retryable error in the asynchronous background loader, it logs a warning and reschedules the loading task instead of terminating, so the server stays up and retries once the transient condition clears. Non-retryable errors and the synchronous drop-table path keep the existing fail-fast behaviour. The unexpected-parts loop is made idempotent so a rescheduled retry skips already-loaded parts. A `REGULAR` failpoint `mergetree_load_outdated_parts_inject_retryable_exception` and an integration test (`test_load_outdated_parts_retryable`) inject a retryable error into the outdated-parts loader and verify the server survives and finishes loading. The retry must not requeue a part whose retryable error happened after it was already published. `loadDataPartWithRetries()` inserts a normal Outdated part into `data_parts_indexes` before its post-load cleanup (`preparePartForRemoval` writes the removal TID and can throw retryably); requeueing it as-is made the retry reload the same directory as a fresh part, hit the duplicate-part path, and `remove()` could delete the directory while the published part still referenced it. The published part is now rolled back out of the index before requeueing so the retry reloads it cleanly. The unexpected-parts loop has the same shape (its optional `broken-on-start` detach runs after the part is set), so it now uses a separate `finished` marker set only after the detach succeeds. Two more failpoints exercise these post-load paths and the integration test gains a third case asserting nothing is wrongly detached. The broken-part detach in both loaders no longer hides its errors. It used to call `renameToDetached(\"broken-on-start\", /*ignore_error=*/ replicated)`, and `renameToDetached` swallows `ErrnoException` and `fs::filesystem_error` when `ignore_error` is set - exactly the exception types a failing rename raises, and exactly the class `isRetryableException` treats as transient. On a replicated table a transient detach failure therefore never reached the new classification: the worker continued as if the cleanup had succeeded, leaving the broken outdated part on its original path with its ZooKeeper entry already removed, or setting the unexpected part's `finished` marker so the retry skipped it. Both call sites now pass `ignore_error = false` and classify the error themselves; the tolerance `ignore_error` gave replicated tables is preserved, but only for a permanent failure and only after the retryable case has been handled. A failpoint that throws the swallowed exception type and a regression case per loader cover it. #### Provenance Surfaced on master as the `'px != 0' failed` family whose actual fatal is the `loadOutdatedDataParts` terminate. Latest master hit: `Stress test (experimental, serverfuzz, arm_release)`, commit `81a8bcef7e8ba2dae3dfdc762dde87a853f40070`, STID 3782-2f00. Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=81a8bcef7e8ba2dae3dfdc762dde87a853f40070&name_0=MasterCI&name_1=Stress%20test%20%28experimental%2C%20serverfuzz%2C%20arm_release%29",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107091",
          "createdAt": "2026-06-10T21:59:10Z",
          "updatedAt": "2026-08-13T13:02:49Z",
          "timestamp": "2026-08-13T13:02:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.6-must-backport"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:2b0d5268a778c51a1a1e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106011",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106011",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "DeltaLake: Use new create table transaction",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Use new create table transaction resolves https://github.com/ClickHouse/ClickHouse/issues/103155 this needs https://github.com/ClickHouse/ClickHouse/pull/105861",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106011",
          "createdAt": "2026-05-27T21:12:24Z",
          "updatedAt": "2026-08-13T13:00:43Z",
          "timestamp": "2026-08-13T13:00:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-feature"
          ],
          "author": "SmitaRKulkarni",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0d7237a4e76aa8b812ec",
        "signalId": "github:ClickHouse/ClickHouse:issue:109326",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:109326",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Enable `rewrite_in_to_join` for makeDistributed by default",
          "text": "Currently pre-built sets for `IN (subquery)` clauses don't work well, so as a workaround it make sense to enable `rewrite_in_to_join=1` by default for distributed queries v2.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/109326",
          "createdAt": "2026-07-03T15:59:04Z",
          "updatedAt": "2026-08-13T13:00:40Z",
          "timestamp": "2026-08-13T13:00:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "unfinished code",
            "make it worse"
          ],
          "author": "alesapin",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:80cebe7588fcee338ff3",
        "signalId": "github:ClickHouse/ClickHouse:issue:106460",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:106460",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "FTS index after Light-Weight Update is not working for MATERIALIZED column",
          "text": "### Company or project name ClickHouse ### Describe what's wrong During implementation of Full-Text Search framework I found the issue when using Light-Weight Updates. Particularly, when patch parts have to be applied, `MATERIALIZED` column can't be utilized in FTS index - `serverError UNKNOWN_IDENTIFIER` is thrown. However, other non-`MATERIALIZED` columns can be searched using FTS index. ### Does it reproduce on the most recent release? Yes ### How to reproduce ```sql SET allow_experimental_full_text_index = 1; SET allow_experimental_json_type = 1; SET enable_lightweight_update = 1; DROP TABLE IF EXISTS tab; CREATE TABLE tab ( id Int64, text String, data_as_json JSON, json_title String MATERIALIZED data_as_json.title::String, INDEX idx_text text TYPE text(tokenizer = 'splitByNonAlpha'), INDEX idx_json_title json_title TYPE text(tokenizer = 'splitByNonAlpha') ) ENGINE = MergeTree ORDER BY id SETTINGS enable_block_number_column = 1, enable_block_offset_column = 1; INSERT INTO tab (id, text, data_as_json) VALUES (1, 'database is fast', '{\"title\": \"performance tuning guide\"}'), (2, 'no match here', '{\"title\": \"unrelated\"}'), (3, 'database stuff', '{\"title\": \"performance test results\"}'); -- Sanity check before the lightweight UPDATE (no patch parts yet) — -- compound WHERE works. SELECT 'compound WHERE, no patch parts yet'; SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance'); -- Lightweight UPDATE on a non-materialised column → creates a patch -- part. The materialised column's source (data_as_json) is NOT -- touched. UPDATE tab SET text = 'database updated' WHERE id = 1; -- Single-column WHEREs still work after the patch part is created. SELECT 'WHERE on text only, post-UPDATE'; SELECT count() FROM tab WHERE hasToken(text, 'database'); SELECT 'WHERE on json_title only, post-UPDATE'; SELECT count() FROM tab WHERE hasToken(json_title, 'performance'); -- The compound WHERE now fails with UNKNOWN_IDENTIFIER. SELECT 'compound WHERE, post-UPDATE (BUG)'; SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance'); -- { serverError UNKNOWN_IDENTIFIER } -- The same compound WHERE with apply_patch_parts = 0 succeeds -- (workaround). SELECT 'compound WHERE, apply_patch_parts = 0 (workaround)'; SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance') SETTINGS apply_patch_parts = 0; DROP TABLE tab; ``` [04305_text_index_bug_json_materialized_patch_part.sql](https://github.com/user-attachments/files/28594469/04305_text_index_bug_json_materialized_patch_part.sql) ### Expected behavior Patch parts should be applied and `MATERIALIZED` column should work with FTS index ### Error message and/or stacktrace ``` {acd9a2a5-2c75-4ce3-a848-babcfcaad691} <Error> executeQuery: Code: 47. DB::Exception: Unknown expression or function identifier `data_as_json.title` in scope _CAST(hasToken(json_title, 'performance'), 'UInt8') AS __text_index_idx_json_title_hasToken_b115b01d1cac96a451f08eb68b7633a0, _CAST(CAST(data_as_json.title, 'String'), 'String') AS json_title: (while reading from part /home/alexey/Documents/ClickHouse/store/f2a/f2a80d0a-0b1f-44ee-ba88-640b07bc0411/all_1_1_0/ located on disk default of type local): While executing MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder). (UNKNOWN_IDENTIFIER) (version 26.6.1.374 (official build)) (from 127.0.0.1:55680) (comment: 04305_text_index_bug_json_materialized_patch_part.sql-test_u6trcoga) (query 15, line 79) (in query: SELECT count() FROM tab WHERE hasToken(text, 'database') AND hasToken(json_title, 'performance');), Stack trace (when copying this message, always include the lines below): 0. /ClickHouse/contrib/llvm-project/libcxx/include/__exception/exception.h:115:14: Poco::Exception::Exception(String&&, int) @ 0x0000000026eb20b3 1. /ClickHouse/src/Common/Exception.cpp:139:7: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x0000000014b73ca9 2. /ClickHouse/src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d08ac96 3. /ClickHouse/src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d08a758 4. /ClickHouse/src/Common/Exception.h:189:77: DB::Exception::Exception<char const*, String&, String, String, String>(int, FormatStringHelperImpl<std::type_identity<char const*>::type, std::type_identity<String&>::type, std::type_identity<String>::type, std::type_identity<String>::type, std::type_identity<String>::type>, char const*&&, String&, String&&, String&&, String&&) @ 0x00000000198e909a 5. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3226:27: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198cb0ce 6. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 7. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 8. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 9. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 10. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 11. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 12. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:1243:9: DB::QueryAnalyzer::tryResolveIdentifierFromAliases(DB::IdentifierLookup const&, DB::IdentifierResolveScope&, DB::IdentifierResolveContext) @ 0x00000000198d6e15 13. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:1548:34: DB::QueryAnalyzer::tryResolveIdentifier(DB::IdentifierLookup const&, DB::IdentifierResolveScope&, DB::IdentifierResolveContext) @ 0x00000000198d7f75 14. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3029:57: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c947b 15. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 16. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 17. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 18. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 19. /ClickHouse/src/Analyzer/Resolve/resolveFunction.cpp:1081:39: DB::QueryAnalyzer::resolveFunction(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool) @ 0x0000000019b50135 20. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3299:46: DB::QueryAnalyzer::resolveExpressionNode(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool, bool) @ 0x00000000198c8c46 21. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:3451:49: DB::QueryAnalyzer::resolveExpressionNodeList(std::shared_ptr<DB::IQueryTreeNode>&, DB::IdentifierResolveScope&, bool, bool, bool) @ 0x00000000198c82fa 22. /ClickHouse/src/Analyzer/Resolve/QueryAnalyzer.cpp:285:17: DB::QueryAnalyzer::resolve(std::shared_ptr<DB::IQueryTreeNode>&, std::shared_ptr<DB::IQueryTreeNode> const&, std::shared_ptr<DB::Context const>) @ 0x00000000198c18e5 23. /ClickHouse/src/Interpreters/inplaceBlockConversions.cpp:238:14: DB::(anonymous namespace)::createExpressionsAnalyzer(DB::Block const&, boost::intrusive_ptr<DB::IAST>, bool, std::shared_ptr<DB::Context const>) @ 0x000000001a941cbe 24. /ClickHouse/src/Interpreters/inplaceBlockConversions.cpp:312:16: DB::evaluateMissingDefaults(DB::Block const&, DB::NamesAndTypesList const&, DB::ColumnsDescription const&, std::shared_ptr<DB::Context const>, bool, bool) @ 0x000000001a943559 25. /ClickHouse/src/Storages/MergeTree/IMergeTreeReader.cpp:269:20: DB::IMergeTreeReader::evaluateMissingDefaults(DB::Block, std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) const @ 0x000000001f25b0ea 26. /ClickHouse/src/Storages/MergeTree/MergeTreeReadersChain.cpp:250:28: DB::MergeTreeReadersChain::executeActionsBeforePrewhere(DB::MergeTreeRangeReader::ReadResult&, std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&, DB::MergeTreeRangeReader&, DB::Block const&, unsigned long) const @ 0x000000001f5589ea 27. /ClickHouse/src/Storages/MergeTree/MergeTreeReadersChain.cpp:120:9: DB::MergeTreeReadersChain::read(unsigned long, DB::MarkRanges&, std::vector<DB::MarkRanges, std::allocator<DB::MarkRanges>>&, std::function<void (std::vector<DB::ColumnWithTypeAndName, AllocatorWithMemoryTracking<DB::ColumnWithTypeAndName>> const&, unsigned long, std::optional<bool>&)> const&) @ 0x000000001f556d4d 28. /ClickHouse/src/Storages/MergeTree/MergeTreeReadTask.cpp:396:38: DB::MergeTreeReadTask::read() @ 0x000000001f553de3 29. /ClickHouse/src/Storages/MergeTree/MergeTreeSelectAlgorithms.h:53:103: DB::MergeTreeThreadSelectAlgorithm::readFromTask(DB::MergeTreeReadTask&) @ 0x000000001f56ce0c 30. /ClickHouse/src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:251:31: DB::MergeTreeSelectProcessor::readCurrentTask(DB::MergeTreeReadTask&, DB::IMergeTreeSelectAlgorithm&) const @ 0x000000001f561f91 31. /ClickHouse/src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:430:23: DB::MergeTreeSelectProcessor::read() @ 0x000000001f5642a5 Received exception from server (version 26.6.1): Code: 47. DB::Exception: Received from localhost:9000. DB::Exception: Unknown expression or function identifier `data_as_json.title` in scope _CAST(hasToken(json_title, 'performance'), 'UInt8') AS __text_index_idx_json_title_hasToken_b115b01d1cac96a451f08eb68b7633a0, _CAST(CAST(data_as_json.title, 'String'), 'String') AS json_title: (while reading from part /home/alexey/Documents/ClickHouse/store/f2a/f2a80d0a-0b1f-44ee-ba88-640b07bc0411/all_1_1_0/ located on disk default of type local): While executing MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder). (UNKNOWN_IDENTIFIER) ``` ### Related issues and pull requests _No response_ ### Additional context Tested on `ClickHouse local version 26.6.1.374 (official build)`",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/106460",
          "createdAt": "2026-06-04T12:05:07Z",
          "updatedAt": "2026-08-13T13:00:32Z",
          "timestamp": "2026-08-13T13:00:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "bug",
            "comp-text-index",
            "clickgap-analyzed",
            "culprit-pr-pinned"
          ],
          "author": "alexbakharew",
          "state": "open",
          "assignees": [
            "CurtizJ"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:9ed1a22c20dd5c063957",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113937",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113937",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: update gui.mdx with CHOPs tool",
          "text": "Added CHOPs UI and Admin tool in the docs ### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113937",
          "createdAt": "2026-08-08T10:13:10Z",
          "updatedAt": "2026-08-13T13:00:20Z",
          "timestamp": "2026-08-13T13:00:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "rva-quantrail",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:779c3da9ad3157ce8430",
        "signalId": "github:ClickHouse/ClickHouse:issue:114640",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114640",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "uniqExact regression between 26.3 and 26.7 at high core counts (up to 3.3x on ClickBench Q10/Q11 on c7a.metal-48xl)",
          "text": "Between the official ClickBench mainline runs of 2026-03-27 (≈26.3) and 2026-08-11 (latest stable, 26.7.x) on `c7a.metal-48xl` (192 vCPU), the `COUNT(DISTINCT ...)` queries regressed while overall hot performance stayed flat (hot geomean 0.97): | query | hot 2026-03-27 | hot 2026-08-11 | ratio | |---|---|---|---| | Q10 `SELECT MobilePhoneModel, COUNT(DISTINCT UserID) AS u FROM hits WHERE MobilePhoneModel <> '' GROUP BY MobilePhoneModel ORDER BY u DESC LIMIT 10` | 0.114 s | 0.196 s | **1.7x** | | Q11 `SELECT MobilePhone, MobilePhoneModel, COUNT(DISTINCT UserID) AS u FROM hits WHERE MobilePhoneModel <> '' GROUP BY MobilePhone, MobilePhoneModel ORDER BY u DESC LIMIT 10` | 0.078 s | 0.257 s | **3.3x** | | Q37 (possibly related) | 0.023 s | 0.042 s | 1.8x | `COUNT(DISTINCT ...)` maps to `uniqExact` by default (`count_distinct_implementation`), and the evidence points at `uniqExact` state merging at high thread counts specifically: - The versions benchmark (`c7a.4xlarge`, 16 vCPU) shows **no** regression on these queries across 26.3 → 26.7 → master — but it substitutes `uniq(UserID)` for `COUNT(DISTINCT UserID)` (oldest-compatible query set), so it does not exercise `uniqExact` at all. - On a 96-core aarch64 box with the same 100M-row dataset, 26.7.1.1315 and current master both run Q10/Q11 in ~0.07 s hot — no regression at 96 threads. - So the degradation manifests only on the 192-vCPU machine, i.e. it scales badly with thread count rather than being a general codegen regression. Q10/Q11 have few groups (mobile phone models) with large per-group `uniqExact` states over ~100M rows — the merge of many per-thread large states is the hot path. A bisect of the 26.3 → 26.7 window on a high-core-count machine is needed; candidates are changes to parallel aggregation state merging or `uniqExact`-specific merge paths. Found while investigating the ClickBench dashboard delta for PR #81944 — this regression is unrelated to that PR but inflates the hot side of its dashboard comparison. Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114640",
          "createdAt": "2026-08-13T12:58:55Z",
          "updatedAt": "2026-08-13T12:58:55Z",
          "timestamp": "2026-08-13T12:58:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:ad552b906e91bd2b17f1",
        "signalId": "github:ClickHouse/ClickHouse:issue:114639",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114639",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Push dynamic TopN thresholds into MergeTree reads for ORDER BY ... LIMIT (2-3x on ClickBench Q24/Q26)",
          "text": "**Use case.** `ORDER BY ... LIMIT n` over a sorted-by-something-else table with a narrow projection, e.g. ClickBench Q24/Q26: ```sql SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10; SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime, SearchPhrase LIMIT 10; ``` The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 implemented a **dynamic TopN threshold pushdown** (`Push TopN threshold to MergeTreeSource`, commit `2e2f594308f4`, plus `Better query condition cache: make TopN dynamic filters deterministic and reusable`, gated by `max_limit_to_push_down_topn_predicate = 100`): while the partial-sort transform maintains the current top-`n`, the running n-th-best value of the `ORDER BY` key is pushed down into the `MergeTree` read as a dynamic threshold, so granules whose key range cannot beat the current top-`n` are skipped instead of read, decompressed, and sorted. This mechanism was dropped during the upstreaming of that PR (the scaffolding was removed from the branch; the surviving `RewriteOrderByLimitPass` is a different, row-offset-based approach — it is off by default and measures performance-neutral on ClickBench when enabled). Master's lazy materialization (`query_plan_optimize_lazy_materialization`) covers the wide-`SELECT *` case (Q23), but does not prune reads for narrow projections: every granule passing the `WHERE` is still fully processed. Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64), original bench-opt build (25.9.1.1) vs master `405e218ff`: Q24 0.009 s vs 0.013 s, Q26 0.008 s vs 0.013 s (1.2–1.3x with times this small). The gap is much larger on the official `c7a.metal-48xl` numbers: Q24 0.014 s vs 0.046 s, Q26 0.014 s vs 0.044 s (**2–3x**). A related consideration from the original design: the dynamic filter interacts with the query condition cache, so the thresholds need to be deterministic/reusable (or excluded from the cache key) — the original branch had a follow-up commit specifically making the TopN dynamic filters deterministic for that reason. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114639",
          "createdAt": "2026-08-13T12:58:53Z",
          "updatedAt": "2026-08-13T12:58:53Z",
          "timestamp": "2026-08-13T12:58:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:6fc16bc2035ccc717a9c",
        "signalId": "github:ClickHouse/ClickHouse:issue:114638",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114638",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Aggregation: precompute hashes and prefetch for preallocated hash-table variants (up to 1.5x on multi-key GROUP BY)",
          "text": "**Use case.** Multi-key `GROUP BY` over large blocks, e.g. ClickBench Q16/Q18: ```sql SELECT UserID, SearchPhrase, COUNT(*) FROM hits GROUP BY UserID, SearchPhrase ORDER BY COUNT(*) DESC LIMIT 10; SELECT UserID, extract(minute FROM EventTime) AS m, SearchPhrase, COUNT(*) FROM hits GROUP BY UserID, m, SearchPhrase ORDER BY COUNT(*) DESC LIMIT 10; ``` The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 carried `Precompute hash and prefetch for prealloc variants` (commit `7663500fdebf` in the `bench-opt` branch): for aggregation methods that go through the preallocated/two-level hash table path, compute the key hashes for the whole block up front and software-prefetch the target buckets before the insert/lookup pass, hiding DRAM latency on the hash-table probes. This part of the PR was never upstreamed (the single-`String`-key improvements landed separately as the packed string keys, #93271). Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64), original bench-opt build (25.9.1.1) vs master `405e218ff`: | query | original | master | ratio | |---|---|---|---| | Q16 | 0.209 s | 0.326 s | 1.53x | | Q18 | 0.387 s | 0.487 s | 1.25x | | Q30 | 0.059 s | 0.078 s | 1.28x | | Q35 | 0.060 s | 0.078 s | 1.26x | | Q31 | 0.103 s | 0.122 s | 1.17x | | Q13/Q14/Q15 | — | — | ~1.1x | The official ClickBench numbers on `c7a.metal-48xl` agree in shape (e.g. Q18: 0.186 s vs 0.264 s). These queries all use multi-key aggregation (`keys128`/`serialized`-family methods), where master currently issues dependent random accesses per row. Precomputed hashes + prefetching is the remaining unmerged aggregation win from that PR. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772 Related: https://github.com/ClickHouse/ClickHouse/pull/93271",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114638",
          "createdAt": "2026-08-13T12:58:50Z",
          "updatedAt": "2026-08-13T12:58:50Z",
          "timestamp": "2026-08-13T12:58:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:4f22ad6b6da1f44bc04d",
        "signalId": "github:ClickHouse/ClickHouse:issue:114637",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114637",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Trivial GROUP BY LIMIT optimization: support aggregate functions in the projection (2x on ClickBench Q17)",
          "text": "**Use case.** ClickBench Q17: ```sql SELECT UserID, SearchPhrase, COUNT(*) FROM hits GROUP BY UserID, SearchPhrase LIMIT 10; ``` The trivial `GROUP BY ... LIMIT` optimization (`optimize_trivial_group_by_limit_query`, default `true`) rewrites such queries to `max_rows_to_group_by = n + offset` with `group_by_overflow_mode = 'any'`, so the aggregation stops accepting new keys once enough distinct keys are produced. Today it is restricted to projections **without aggregate functions**, so it never fires for Q17 (or for any ClickBench query). The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 (commit `adda6e0ce599`, `Trivial group by limit`) applied the same rewrite **with** aggregate functions in the projection. Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64): | build | Q17 hot | |---|---| | original bench-opt (25.9.1.1) | 0.092 s | | master `405e218ff` | 0.201 s | i.e. **2.1x**. The official ClickBench numbers on `c7a.metal-48xl` show the same shape: 0.048 s vs 0.116 s. **Why it is restricted.** With `group_by_overflow_mode = 'any'`, rows for already-kept keys continue to be aggregated, so the aggregate values of the returned keys stay correct in the single-threaded case. But with parallel aggregation, each thread keeps its own first `n + offset` keys: a key kept by thread A and rejected by thread B loses B's rows, and the merged result returns that key with an undercounted aggregate. Returning an *unspecified subset* of keys is fine for `LIMIT` without `ORDER BY`; returning *wrong aggregate values* for them is not. The restriction to aggregate-free projections sidesteps this, at the cost of never firing on realistic queries. **Possible directions.** - Share the cutoff key set across threads (e.g. a concurrent filter of \"kept keys\"): once the global set reaches `n + offset`, all threads keep aggregating rows whose key is in the set and drop the rest — aggregate values for kept keys stay exact. - Alternatively, keep per-thread sets but make the final merge drop keys that were not kept by *every* participating thread (kept-by-all keys have exact values; needs `n + offset` sized so enough survive). - Or fall back to a single aggregation thread when the rewrite fires and the estimated cardinality is small — for `LIMIT 10`-style queries the single-threaded early-exit can still beat the full parallel scan. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114637",
          "createdAt": "2026-08-13T12:58:40Z",
          "updatedAt": "2026-08-13T12:58:40Z",
          "timestamp": "2026-08-13T12:58:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:960a6fe6320877d23e0e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114479",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114479",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enable reading in reverse order with FINAL for ReplacingMergeTree",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/58035 Related: https://github.com/ClickHouse/ClickHouse/pull/58361 Related: https://github.com/ClickHouse/ClickHouse/pull/111609 When a query with `FINAL` sorts in reverse order of the sorting key, for example `ORDER BY key DESC LIMIT n`, the read-in-order optimization was disabled and the query read the whole table. It now applies for `ReplacingMergeTree`. `ReplacingSortedAlgorithm` learns a `read_in_reverse` mode. A row with a strictly higher version always replaces the selected one; among rows with equal (or absent) versions, the previously selected row is kept unless the current row comes from a newer data part. That mirrors the \"last written row wins\" rule of the direct reading order, because in the reverse reading order rows within one part arrive backwards while the parts still arrive from the oldest to the newest one. This is what makes the reverse read select the same row of a duplicate key group as a direct read, which is the correctness concern that stopped https://github.com/ClickHouse/ClickHouse/pull/58361. The other engines keep the previous behavior, since their merging algorithms depend on the direct order of rows: the sequence of sign rows in `CollapsingMergeTree`, the order of rows fed to order-dependent aggregate functions in `AggregatingMergeTree`, and so on. On a 110 million row `ReplacingMergeTree` table with two overlapping parts, `SELECT x FROM t FINAL ORDER BY x DESC LIMIT 1` (measured on a `RelWithDebInfo` build): | | before | after | |---|---|---| | Rows read | 110.1 million | 1.71 million | | Elapsed | 4.5 s | 0.22 s | | Peak memory | 41 MB | 23 MB | Trade-off: as with the already existing direct-order in-order reads with `FINAL`, an in-order plan disables vertical `FINAL` and the splitting of parts ranges into intersecting and non-intersecting ones. A query that reads the full result with `ORDER BY key DESC` and no small `LIMIT` may therefore become slower on a wide or mostly merged table. The new setting `optimize_read_in_reverse_order_final` (default enabled) turns the optimization off, and `compatibility` with an earlier version restores the previous plans. Out of scope, to keep this change reviewable: `Merge` tables over `ReplacingMergeTree`, and the interaction with `topKThroughJoin`, which keeps its own optimization for `... FINAL LEFT JOIN ... ORDER BY key DESC LIMIT n`. Both can follow separately. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable the read-in-order optimization for queries with the `FINAL` modifier that sort in reverse order of the sorting key on `ReplacingMergeTree` tables, so that queries such as `SELECT ... FROM t FINAL ORDER BY key DESC LIMIT n` read only the relevant tail of the data instead of the whole table. Can be disabled with the new setting `optimize_read_in_reverse_order_final`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114479",
          "createdAt": "2026-08-12T12:07:10Z",
          "updatedAt": "2026-08-13T12:58:04Z",
          "timestamp": "2026-08-13T12:58:04Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "cwurm",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c88b098064eb1ffc1906",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113401",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix paimon timestamp precision",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> fix: https://github.com/ClickHouse/ClickHouse/issues/112768 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables partitioned by a `TIMESTAMP` or `TIMESTAMP WITH LOCAL TIME ZONE` column of precision above 3. Such tables previously failed with `scale 6 is not supported, only support scale <= 3` before returning any row, which affected every timestamp-partitioned table written by Spark, since Spark maps both `TIMESTAMP` and `TIMESTAMP_NTZ` to Paimon `TIMESTAMP(6)`. Partition pruning on such a column now uses the full precision as well.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113401",
          "createdAt": "2026-08-05T01:12:47Z",
          "updatedAt": "2026-08-13T12:57:52Z",
          "timestamp": "2026-08-13T12:57:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "JiaQiTang98",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:734c4ca20336006d03e6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:81944",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:81944",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "WIP some perf optimizations",
          "text": "This pull request bundles two query-execution changes: - `count_distinct_optimization` is enabled by default. The rewrite is applied only by the analyzer's `CountDistinctPass`, which skips `Nullable` / `LowCardinality(Nullable)` arguments and refuses to fire for remote storages or with `distributed_group_by_no_merge`. The legacy AST-level `RewriteCountDistinctFunctionMatcher` was removed: it never fired (its `table_expr->size() != 1` guard uses the recursive `IAST::size`, which is always `>= 2` for a real table expression) and it lacked all of those guards. - A new opt-in `query_plan_rewrite_order_by_limit` optimization that rewrites `ORDER BY ... LIMIT` over wide `MergeTree` reads into a row-offset form. It is **off by default**: it is not yet neutral for the `RuntimeDataflowStatisticsInputBytes` read-bytes estimation (the estimate diverges from `ReadCompressedBytes` by ~24x on wide `SELECT * ... ORDER BY ... LIMIT` reads). `RewriteOrderByLimitPass` rejects `FINAL`, `QUALIFY`, window functions, `WITH FILL`, `arrayJoin`, and non-deterministic `ORDER BY` expressions such as `rand`. The `ABStringRef` string-key aggregation method that this branch previously carried has been dropped: master merged https://github.com/ClickHouse/ClickHouse/pull/93271, which upstreams the same idea as `PackedStringRef` and makes `AggregationMethodPackedString` the default `key_string` / `key_string_two_level` method. Keeping the branch's variant would mean maintaining a duplicate hash method plus a `base/base/StringRef.h` compatibility shim that existed only to keep it compiling after master replaced `StringRef` with `std::string_view`. The trivial `GROUP BY ... LIMIT` optimization and the top-N threshold pushdown into `MergeTree` reading are no longer part of this branch either; the leftover, callerless scaffolding for the latter (`TopNFilterParameters`, `SortColumnDescription::column_name_in_storage`, `FilterTransform::updateQueryConditionHash`, the `getPrewhereInfo` accessors, and the `getFilterMask` rework in `PartialSortingTransform`) has been removed as well. Related: https://github.com/ClickHouse/ClickHouse/pull/93271 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Enable `count_distinct_optimization` by default for eligible local-table `countDistinct` / `uniqExact` queries, and add the opt-in `query_plan_rewrite_order_by_limit` optimization for wide `MergeTree` reads. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/81944",
          "createdAt": "2025-06-16T14:32:02Z",
          "updatedAt": "2026-08-13T12:56:58Z",
          "timestamp": "2026-08-13T12:56:58Z",
          "metrics": {
            "reactions": 4,
            "comments": 35
          },
          "labels": [
            "pr-performance"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4b4a3d5bc0cc3d4d1883",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114087",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114087",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check column structure against the declared type in `collectOffsetsColumns`",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113925 Related: https://github.com/ClickHouse/ClickHouse/pull/113225 Related: https://github.com/ClickHouse/ClickHouse/issues/113891 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a column's declared type and its data diverge during a `MergeTree` read (mixed type provenance, e.g. after an `ALTER TABLE ... MODIFY COLUMN` whose mutation has not finished), the server now reports a clear exception naming the column and both structures, instead of an unchecked cast: debug and sanitizer builds previously aborted with a bare `Bad cast from type A to B` naming no column, and release builds walked mismatched memory silently. ### Description The type-directed `enumerateStreams` walk in `collectOffsetsColumns` (`src/Interpreters/inplaceBlockConversions.cpp`) pairs each available column's declared type with its data. When an entry's type was resolved from the table's metadata while the column was read from a data part with an older type - the mixed type provenance behind #113925 - the walk `assert_cast`s the column to the wrong class at the first diverging wrapper. In debug/sanitizer builds that is an abort whose message names two column classes and no column (the crash-classifier issue #113891, STID 4256-3fc1, took a cross-file hunt to attribute); in release builds `assert_cast` does not check at all, so the walk misbehaves on memory-unsafe reinterpretation. The new `columnMatchesTypeStructure` check runs right before the walk, only on the missing-columns path where the walk already happens. It descends exactly the levels whose serializations pair the declared type's structure with the column's: - `Array`, `Nullable`, `Tuple`, `Map` - the wrappers whose serializations `assert_cast` the column; - `Variant` - `SerializationVariant::enumerateStreams` walks every alternative declared by the type and pairs it with the column's variant of the same global discriminator, so the alternative list has to match in length and element-wise; - typed paths of `Object` - `SerializationObject::enumerateStreams` walks every typed path declared by the type and looks it up in the column, so the typed paths have to match by name and structure. Dynamic paths and shared data are taken from the column itself and need no check; - `ColumnReplicated` is unwrapped. `Dynamic` is checked by class only, because its `enumerateStreams` takes both the type and the column of its variant from the column itself (`column_dynamic->getVariantInfo().variant_type`), so the two cannot diverge. Every leaf the check does not know is accepted - leaf divergence, such as a part storing `UInt32` for a column widened to `UInt64`, is legitimate. Because the checked structure mirrors exactly what the `enumerateStreams` implementations themselves assert, the check cannot fire on any pairing that debug CI does not already abort on (or, for `Object`, fail with a raw typed-path lookup) - it only converts that failure into a diagnosable exception and closes the release-build hole. With the #113925 reproducer on a `RelWithDebInfo` build of master before that fix, the witness query fails with: ``` Code: 49. DB::Exception: Column `arr.n` is listed with type Array(Nullable(String)) among available columns, but its data has incompatible structure Array(size = 1, UInt64(size = 1), String(size = 2)). It is likely that a type resolved from the table's metadata was combined with a column read from a data part with an older type: (while reading from part .../all_1_1_0/ ...) ``` instead of silent unchecked-cast behavior. #113925 (which fixes the type selection) has since merged and is included here, so this check now guards the invariant against future regressions of the same family. **Validation.** Witness reproduced as above on a local `RelWithDebInfo` build. False-positive sweep: all 381 stateless tests matching `nested`/`subcolumn` run against that binary - 294 passed, 87 failed for documented bare-server environmental reasons (no Keeper, no clusters, no `protoc`, no `/var/lib/clickhouse`), and the check's message appears zero times in the whole run. After the `Variant`/`Object` extension, all 826 stateless `.sql` tests matching `variant`/`dynamic`/`json`/`nested`/`subcolumn`/`object` were run again - the check's message and `Bad cast` both appear zero times - plus a targeted check that reads old parts through newly added `Variant`, `JSON`, `Dynamic`, `Tuple`, `Map` and `Nested` columns and through unfinished `ALTER TABLE ... MODIFY COLUMN` mutations of `Variant` and `JSON`. The `Memory`-engine caller of `fillMissingColumns` always casts columns to the requested types first (`tryGetColumnFromBlock`), so it cannot trip the check either. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114087",
          "createdAt": "2026-08-10T00:15:02Z",
          "updatedAt": "2026-08-13T12:56:49Z",
          "timestamp": "2026-08-13T12:56:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3a6c593134500536a04d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114626",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114626",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not merge-sort a distributed gather whose sort description is all-constant",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 Related: https://github.com/ClickHouse/ClickHouse/issues/106237 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description With `make_distributed_plan = 1`, a window `PARTITION BY <constant>` builds a `GatherSend` fragment whose sort description is entirely constant. `GatherSendStep::updatePipeline` adds an order-preserving `MergingSortedTransform` for any non-empty description, and that merge waits for *every* input stream to have data (`IMergingTransformBase::prepareInitializeInputs`). Hashing a constant key sends all rows to one bucket, so the other branches are never fed and stay `NeedData` while the loaded branch is `PortFull`. Neither side can move: the query deadlocks at zero CPU until `receive_timeout` and then raises `Pipeline stuck` (a logical error, so the server aborts in debug and sanitizer builds). An all-constant description orders nothing, so this change takes the `pipeline.resize(1)` branch that already exists in that function. A `ResizeProcessor` pairs any waiting output with any ready input and has no all-inputs barrier, which is why the code before [4a1ab1e](https://github.com/ClickHouse/ClickHouse/commit/4a1ab1e3d94fb0d) - which replaced an unconditional `resize(1)` with this conditional merge - could not wedge this way. Every row compares equal under an all-constant description, so any interleaving is validly sorted and the contract `GatherReceiveStep` relies on still holds. A description with at least one real column keeps the merge. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 (where this failure was reported) Related: https://github.com/ClickHouse/ClickHouse/issues/106237 (context only, different mechanism) Found by `AST fuzzer (amd_debug, targeted, old_compatibility)`: [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113558&sha=b9fe1650232abf4fefc16992f3efcce44222bebc&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The `BufferedShardByHashTransform` deadlock tracked in #106237 is a different mechanism: that transform does not appear in the failing pipeline at all. The new test `04888` fails on master with `Pipeline stuck` for a constant key, a `LowCardinality` constant and a `Nullable` constant, and passes with this change. Its controls - a real column key, a mixed constant-plus-column key, and the non-distributed plan - pass both before and after, so the change is narrow. `04837_distributed_plan_window_partition_shuffle`, which pins the full distributed plan for a column-keyed window, still passes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114626",
          "createdAt": "2026-08-13T12:23:35Z",
          "updatedAt": "2026-08-13T12:56:43Z",
          "timestamp": "2026-08-13T12:56:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "davenger"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cf9bbdb5253f11808ff2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100377",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100377",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Disable read-in-order when primary key selectivity is poor",
          "text": "When the WHERE clause cannot use the primary key effectively (e.g., `LIKE '%...'`), the `optimize_read_in_order` optimization kills parallelism: each part is read by a single stream instead of many. This can make `ORDER BY` queries 4x slower than the same query without `ORDER BY`. Example on a 2.5B row table with `ORDER BY path` and `WHERE path LIKE '%stderr.log'`: - Without ORDER BY: ~9.5s - With ORDER BY (read-in-order): ~40s - With `optimize_read_in_order = 0`: ~9.6s The fix adds a runtime check in `ReadFromMergeTree::spreadMarkRanges`: if the primary key index selected more than a configurable fraction of all granules and there is no LIMIT, we fall back to parallel reading with per-stream `PartialSortingTransform` + `MergeSortingTransform`, so the output is still sorted as the `SortingStep` above expects. New setting `read_in_order_max_primary_key_ratio` controls the threshold. The default is 1.0, which preserves the previous behavior (never disable read-in-order based on primary key selectivity): the heuristic is opt-in until the default threshold is tuned against the performance report. Set it below 1.0 (e.g. 0.5) to enable the guard. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add a new setting `read_in_order_max_primary_key_ratio` that can disable the read-in-order optimization when the primary key selectivity is poor (more than the given fraction of granules selected), falling back to parallel reading with sorting. This avoids severe parallelism loss for queries like `SELECT ... WHERE path LIKE '%...' ORDER BY path`. The default 1.0 preserves the previous behavior; set a lower value (e.g. 0.5) to enable the heuristic. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) New setting `read_in_order_max_primary_key_ratio` (Float, default 1.0): maximum ratio of selected to total primary key granules for `optimize_read_in_order` to stay enabled. When the ratio exceeds this value, read-in-order is disabled in favor of parallel reading with per-stream sorting. The default 1.0 preserves the old behavior.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100377",
          "createdAt": "2026-03-22T18:02:39Z",
          "updatedAt": "2026-08-13T12:55:23Z",
          "timestamp": "2026-08-13T12:55:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 40
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ae24cc73b2bd334b3b2c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:64184",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:64184",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Adding storage pulsar",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category: - Experimental Feature ### Changelog entry: Added the experimental `Pulsar` table engine for reading from and writing to Apache Pulsar topics. Creating tables with the engine requires enabling the `allow_experimental_pulsar_storage_engine` setting. Closes: https://github.com/ClickHouse/ClickHouse/issues/4226",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/64184",
          "createdAt": "2024-05-21T10:18:47Z",
          "updatedAt": "2026-08-13T12:55:13Z",
          "timestamp": "2026-08-13T12:55:13Z",
          "metrics": {
            "reactions": 2,
            "comments": 7
          },
          "labels": [
            "submodule changed",
            "manual approve",
            "can be tested",
            "pr-experimental"
          ],
          "author": "SteveBalayanAKAMedian",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:06203dbb2c151c131728",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105499",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105499",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Drop for detached tables",
          "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental `DROP DETACHED TABLE` support, gated by `allow_experimental_drop_detached_table`, to remove metadata and data for detached tables. Continues #62490",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105499",
          "createdAt": "2026-05-21T09:59:53Z",
          "updatedAt": "2026-08-13T12:54:54Z",
          "timestamp": "2026-08-13T12:54:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-experimental"
          ],
          "author": "UberDever",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:022051e78e075c58963d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111923",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111923",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix signed integer overflow in the DateLUTImpl sunday-first week helpers",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/107366 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `toStartOfWeek` and `toLastDayOfWeek` with a Sunday-first week mode returning a wrapped day number for a `Date32` value outside the representable calendar, for example after `toDate32('1970-01-01') + INTERVAL 2147483647 DAY`. The week boundary is now computed at the calendar boundary, consistently with the Monday-first week modes. The wrapping was also a signed integer overflow (undefined behavior). ### Description UBSan reported this in `Stress test (arm_asan_ubsan, s3)` on an unrelated PR, from an AST-fuzzer query: ``` src/Common/DateLUTImpl.h:1316:11: runtime error: signed integer overflow: 2147483647 + 6 cannot be represented in type 'int' ``` Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107366&sha=2f1747ab6b4ba58638193a2da4a7c3f0fc8415e1&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 The two `week_mode` overloads `toFirstDayNumOfWeek(v, week_mode)` and `toLastDayNumOfWeek(v, week_mode)` are the only week helpers in `DateLUTImpl.h` that lack the out-of-LUT-range escape branch all their siblings have. `toDayOfWeek(v)` internally takes the escape and returns the weekday of the clamped day, but the arithmetic that follows runs on the raw unclamped day number, so `v += 6` is undefined behaviour for an `ExtendedDayNum` at `INT32_MAX`. A `Date32` reaches such a value through wrapping `addDays`. This PR fixes two bugs with that one root cause: the reported `INT32_MAX` overflow in `toLastDayNumOfWeek`, and its `INT32_MIN` mirror in `toFirstDayNumOfWeek` (`DateLUTImpl.h:1302`, `-2147483647 - 6`). I added the escape branch to both overloads, mirroring the monday-first siblings, so the day number is clamped into the representable calendar before any arithmetic. `day_of_week % 7` maps Sunday (7) to 0 because these overloads start the week on Sunday. As with the monday-first overloads, the resulting week boundary may lie a few days past the calendar boundary; I kept that regime so the two paths agree for the same input. The in-range code is unchanged: `isOutOfLUTRange` passing bounds the day number to `[-25567, 120273]`, where the arithmetic cannot overflow, so every representable date keeps its current result. The bug is observable without a sanitizer, which is why the changelog category is `Bug Fix` rather than the `CI Fix or Improvement` used for the sibling overflows in this file. Before the fix `toLastDayOfWeek(toDate32('1970-01-01') + INTERVAL 2147483647 DAY, 0)` returned the wrapped day number `-2147483648`, and the first and last day of the same week were `-4294967290` days apart instead of 6. I found no open issue for this. The same UBSan signature fired in 5 unrelated pull requests over the last 45 days and never on master.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111923",
          "createdAt": "2026-07-25T23:38:45Z",
          "updatedAt": "2026-08-13T12:54:32Z",
          "timestamp": "2026-08-13T12:54:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:bcf52a1b89dbfc7fe22a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110104",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110104",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check for cancellation in AggregatingInOrderTransform",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107941 --> Related: https://github.com/ClickHouse/ClickHouse/issues/107941 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `optimize_aggregation_in_order` ignoring query cancellation. `AggregatingInOrderTransform` now checks for cancellation while aggregating a chunk, so a query stopped by `KILL QUERY` or by `max_execution_time` (in the default `timeout_overflow_mode = 'throw'`) stops promptly instead of running the whole chunk to completion. ### Description `AggregatingInOrderTransform::consume()` splits one input chunk into runs of equal keys in a loop. Query time and cancellation limits are only checked between pipeline steps (between `work()` calls), so a chunk with many distinct keys makes a single `consume()` call run for a long time (`O(distinct_keys)` iterations, each an `upper_bound` over the remaining rows) with no cancellation checkpoint. As a result a cancelled query (`KILL QUERY`, `max_execution_time`) using `optimize_aggregation_in_order` kept aggregating until the whole chunk was done; the connection thread then blocked in `PullingAsyncPipelineExecutor::cancel() -> ThreadFromGlobalPool::join()` waiting for that loop. The server-side AST fuzzer repeatedly hit this as `Hung check failed, possible deadlock found` (Stress test, all sanitizers), with `system.processes` showing `is_cancelled = 1` and `elapsed` far past the 90s hung-check window while the worker thread sat in `AggregatingInOrderTransform::consume -> Aggregator::executeImpl`. The loop now checks `isCancelled()` once per key interval (cheap) and returns early; the partial aggregation state is discarded because the pipeline is being torn down. This mirrors the existing per-loop cancellation checks in `WindowTransform` and `FillingTransform`. Scope: this covers cancellation that sets `is_cancelled` on the pipeline, i.e. `KILL QUERY` and `max_execution_time` in the default `timeout_overflow_mode = 'throw'`, which is what the reproduced hung check hit (`system.processes` showed `is_cancelled = 1`). The non-default `break` mode is a soft limit that returns a partial result and never sets `is_cancelled`; honoring it mid-chunk (as `FillingTransform` does via `process_list_element->checkTimeLimit()`) is a separate partial-result change, out of scope here. The regular `AggregatingTransform` behaves the same way. Regression test `04512_aggregation_in_order_cancellation` forces one long `consume()` over 40M distinct-key rows in a single chunk, `KILL QUERY ... SYNC` once every row is read: with the fix the KILL returns in a fraction of a second, without it it blocks for the several seconds the loop needs to finish.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110104",
          "createdAt": "2026-07-11T17:17:10Z",
          "updatedAt": "2026-08-13T12:53:40Z",
          "timestamp": "2026-08-13T12:53:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yakov-olkhovskiy"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c5e63e72be360e1d24c6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114216",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114216",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for the RemovePart part_log row in 02950 and 02491",
          "text": "### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02950_part_log_bytes_uncompressed` and `02491_part_log_has_table_uuid` read `system.part_log` immediately after a DDL that removes a part and assert the `RemovePart` row is already there. That row is written asynchronously, so the assertion is unsound and the tests fail with just that line missing. `RemovePart` has a single emitter, `MergeTreeData::removePartsFinally`, which on both DDL paths here is reached only through `grabOldParts()`. Both DDLs (`DROP PART`, `TRUNCATE`) do run cleanup in the query thread, so usually the row lands first. But that grab selects nothing if it loses `try_lock` on `grab_old_parts_mutex`, or if the part's `shared_ptr` is not unique, e.g. while a concurrent read holds a reference. The part then stays `Outdated` and a later cleanup pass writes the row: late, not lost, so the engine is correct and only the tests need fixing. `01686_event_time_microseconds_part_log.sh` already asserts the same shape behind a bounded poll. Both tests become `.sh`, since a bounded poll is not expressible in `.sql`, and wait for the row under a 60 s bound. Each iteration issues `SYSTEM START CLEANUP <table>` to schedule a pass, because a cleanup thread that found nothing to do backs off up to `max_cleanup_delay_period`. Assertion queries, tags and both `.reference` files are unchanged; no `CREATE TABLE` setting is added here (02491's `old_parts_lifetime` pin and 02950's tags are pre-existing, kept verbatim); no `no-parallel` and no blanket `no-random-*`. Validation: holding a reference across the DDL reproduces the non-unique-ownership path deterministically. With that lever and no poll both tests fail 8/8; with the poll they pass 10/10. Dropping only the `SYSTEM START CLEANUP` line fails at the deadline once the cleanup thread is backed off. Deleting the DDL makes the poll time out rather than pass, so it is not satisfied by a stale row. Then 150/150 green per test across default, `-j 8` and randomized-order batches.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114216",
          "createdAt": "2026-08-10T20:03:37Z",
          "updatedAt": "2026-08-13T12:53:22Z",
          "timestamp": "2026-08-13T12:53:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c938bcd251fd15a0e31d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114622",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114622",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not apply DROP fault injection to a refreshable view's cleanup DROP",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/114420 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a refreshable materialized view leaking its rotated-out target table when `ignore_drop_queries_probability` is enabled. The `DROP` a refresh issues to clean up the previous target is a step of the refresh, not a `DROP` the user asked for, so the fault injection no longer applies to it. ### Description Follow-up to #114420, requested by @ tiandiwonder in https://github.com/ClickHouse/ClickHouse/pull/114420#discussion_r3772879738. `ignore_drop_queries_probability` makes a `DROP TABLE` silently do nothing (or become a `TRUNCATE`) so the stress suite exercises \"the table you dropped is still there\". It must only affect `DROP`s the user issued. **Root cause.** After a refresh swaps in a new target, `StorageMaterializedView::dropTempTable` drops the rotated-out one via `InterpreterDropQuery(drop_query, refresh_context)`. `createRefreshContext` never marks that context internal, so the gate treated the cleanup `DROP` as a user `DROP`. Its existing refreshable-view exemption does not cover this: that test inspects the table *being dropped*, which here is the inner target, not the view. Both refresh exits are affected. When it fires, the old target survives as `.tmp.inner_id.<uuid>` still holding a full copy of the view's data, outside the view's metadata and surviving restart. **The change.** `InterpreterDropQuery` gains an `internal` member mirroring the one `InterpreterCreateQuery` already has, `dropTempTable` sets it, and the shared `refresh_context` is untouched. Marking that context instead would break the refresh outright in a `Replicated` database: the publishing `RENAME` runs on it, and `DatabaseReplicated` rejects a non-initial query unless the interpreter also passes `flags.internal`, which neither the Rename nor the Drop interpreter did. That is why #114420's one-liner was safe there and is not here. `QueryFlags{ .internal = internal }` at the replicated enqueue is behaviour-neutral for every other `DROP`, since all its members default to false. The interpreter's other policy branches (`ON CLUSTER` dispatch, access and dependency checks) deliberately keep applying. **Validation.** New test `04887_refreshable_mv_cleanup_drop_not_ignored`: three positive arms (success path, failure path, `Replicated` database) leak on master and are clean with the fix; a fourth asserts a genuine user `DROP` is still skipped. 50 randomized runs plus 50 with `ignore_drop_queries_probability=0.2` were green, as were `04247`, `04796`, `04218` and `04327`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114622",
          "createdAt": "2026-08-13T11:43:28Z",
          "updatedAt": "2026-08-13T12:51:45Z",
          "timestamp": "2026-08-13T12:51:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a22149a77edf6f03ec9e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114635",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114635",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113742 to 26.7: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114635",
          "createdAt": "2026-08-13T12:50:24Z",
          "updatedAt": "2026-08-13T12:51:12Z",
          "timestamp": "2026-08-13T12:51:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:587d9833e76cf8c3207e",
        "signalId": "github:ClickHouse/ClickHouse:issue:102845",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:102845",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Documentation examples for `h3GetDestinationIndexFromUnidirectionalEdge` and `h3GetOriginIndexFromUnidirectionalEdge` use invalid edge index that throws `INCORRECT_DATA`",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _Retrospective finding from a historical scan of [PR #82286](https://github.com/ClickHouse/ClickHouse/pull/82286) (merged 2025-10-17). Confirmed on current codebase — close with a note if already fixed._ ### Describe what's wrong Both `h3GetDestinationIndexFromUnidirectionalEdge` and `h3GetOriginIndexFromUnidirectionalEdge` documentation examples use edge index 1248204388774707197, which is not a valid H3 directed edge. Running the documented query throws `INCORRECT_DATA` exception instead of returning the documented values. **Root cause:** h3GetDestinationIndexFromUnidirectionalEdge.cpp:117 and h3GetOriginIndexFromUnidirectionalEdge.cpp:117: the example input 1248204388774707197 is not a valid H3 directed edge (h3UnidirectionalEdgeIsValid returns 0), and the documented output values are also wrong even if the correct edge (1248204388774707199) is used **Why we believe this is a bug:** h3GetDestinationIndexFromUnidirectionalEdge.cpp:117 and h3GetOriginIndexFromUnidirectionalEdge.cpp:117 — both examples use index 1248204388774707197 which fails `isValidDirectedEdge` check in h3Common.cpp:38. The valid edge used elsewhere in the PR is 1248204388774707199 (differs by 2). The documented output values (599686043507097597 and 599686042433355773) also differ by 2 from the actual values for the valid edge (599686043507097599 and 599686042433355775). **Affected locations:** - `src/Functions/h3GetDestinationIndexFromUnidirectionalEdge.cpp:117` — example uses invalid edge 1248204388774707197, documented output 599686043507097597 - `src/Functions/h3GetOriginIndexFromUnidirectionalEdge.cpp:117` — example uses invalid edge 1248204388774707197, documented output 599686042433355773 **Impact:** Users running the documented example query get an unexpected `INCORRECT_DATA` exception instead of the promised result. Even with `functions_h3_default_if_invalid=1`, the function returns 0 (not the documented value). ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `19cbcd782ed8`). ### How to reproduce ```sql -- Verify documented edge index is invalid SELECT h3UnidirectionalEdgeIsValid(1248204388774707197) AS invalid_edge; SELECT h3UnidirectionalEdgeIsValid(1248204388774707199) AS valid_edge; -- Actual correct results with valid edge SELECT h3GetDestinationIndexFromUnidirectionalEdge(1248204388774707199) AS destination; SELECT h3GetOriginIndexFromUnidirectionalEdge(1248204388774707199) AS origin; ``` [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/8eabd45c-b67c-477b-939f-3946f68a10a4) ### Expected behavior ``` Documentation claims h3GetDestinationIndexFromUnidirectionalEdge(1248204388774707197) returns 599686043507097597 and h3GetOriginIndexFromUnidirectionalEdge(1248204388774707197) returns 599686042433355773, but index 1248204388774707197 is invalid (throws INCORRECT_DATA) and even the correct index produces different values ``` ### Error message and/or stacktrace ``` 0 1 599686043507097599 599686042433355775 ``` ### Additional context **Open risks:** - The existing test 02292_h3_unidirectional_funcs.sql already expects INCORRECT_DATA for index 1248204388774707197 (line 3), confirming the documentation is wrong **Suggested fix:** Replace edge index 1248204388774707197 with the valid edge 1248204388774707199 in both files, and update output values: destination from 599686043507097597 to 599686043507097599, origin from 599686042433355773 to 599686042433355775 **Analysis details:** Confidence HIGH | Severity P3 | Testability: `STATELESS_SQL` Found during automated review of [PR #82286](https://github.com/ClickHouse/ClickHouse/pull/82286). --- _ClickGapAI · Confidence: HIGH · Severity: P3 · Finding: `h_pr82286_002`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #111152 - Merged into: `26.8.1.1324` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/102845",
          "createdAt": "2026-04-15T18:38:21Z",
          "updatedAt": "2026-08-13T12:50:47Z",
          "timestamp": "2026-08-13T12:50:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "comp-documentation",
            "comp-geo"
          ],
          "author": "clickgapai",
          "state": "closed",
          "assignees": [
            "scanhex12"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:c37a35b60472c039e7e0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114634",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114634",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113742 to 26.6: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114634",
          "createdAt": "2026-08-13T12:49:55Z",
          "updatedAt": "2026-08-13T12:50:30Z",
          "timestamp": "2026-08-13T12:50:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:3ed0bce0baf51b7e8e16",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114508",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114508",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add ProbeDeck SQL client integration",
          "text": "Adds a community-maintained ProbeDeck integration guide for iOS and iPadOS. The guide covers ClickHouse Cloud and self-hosted setup, database authentication, mTLS, SSH bastions, monitoring sources, a reproducible query, limits, and troubleshooting. It also adds ProbeDeck to the SQL client navigation and overview. The screenshots use deterministic synthetic data and contain no production credentials or infrastructure. Related: https://github.com/ClickHouse/clickhouse-docs/pull/6625 ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114508",
          "createdAt": "2026-08-12T15:25:38Z",
          "updatedAt": "2026-08-13T12:50:26Z",
          "timestamp": "2026-08-13T12:50:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "manual approve",
            "can be tested"
          ],
          "author": "chamav",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:578bc6b4abe9d5c811c1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114620",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114620",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix serialization of Map-valued settings in access entities",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114591 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a bug where an access entity carrying a `Map`-valued setting, such as a settings profile with `http_response_headers` or `additional_table_filters`, was stored in a form that ClickHouse could not read back, so the entity became permanently unloadable after a restart. Closes #114591. ### Description Closes #114591. `CREATE SETTINGS PROFILE p SETTINGS http_response_headers = '{...}'` succeeds, but the entity is stored as `http_response_headers = [('k', 'v')]`, which nothing can parse back. No error appears at `CREATE` time, so the entity is **permanently unloadable** after a restart: ``` stored: ATTACH SETTINGS PROFILE `p` SETTINGS http_response_headers = [('a', 'b')] CONST; after restart: Code: 62. Syntax error: failed at position 65 ([) ... Could not parse <path>/access/<uuid>.sql ``` The list rebuild drops it silently, reading it back throws, so `SELECT` from `system.settings_profile_elements` fails while it is present, as does `RESTORE` of a backup holding it. Root cause: the value is cast to the setting's native type, so a Map setting holds a `Map` Field, which `FieldVisitorToString` renders as an array of tuples; but `ParserSettingsProfileElement` reads values with a scalar-only `ParserLiteral` and cannot open a `[`. That spelling is rejected everywhere, `SET http_response_headers = [('a','b')]` included, so the write side is wrong. Fix: when a profile element's value, MIN or MAX is a `Map` and the setting is builtin, emit the setting's canonical text as a quoted string. Write side only, no grammar change. Custom settings are excluded because `castValueUtil` returns their value unchanged, so a string would come back a `String` rather than a `Map`. Covers `CREATE USER`/`ROLE`, `ALTER ... SETTINGS` and MIN/MAX, and transitively BACKUP/RESTORE and both storages. `SHOW CREATE` now prints a quoted string rather than `[('k', 'v')]`, intended since the new form is copy-pasteable. Entities already stored in the broken form are not repaired, as they were never parseable; recreate them. Downgrade is safe: a pre-fix binary reads the new form correctly, so no versioning is needed. New test `04902_access_entity_map_setting_round_trip`: 11 of its 15 arms fail on pristine master and pass here, covering empty, multi-key and hostile maps plus a two-process on-disk reload.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114620",
          "createdAt": "2026-08-13T11:13:29Z",
          "updatedAt": "2026-08-13T12:49:35Z",
          "timestamp": "2026-08-13T12:49:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:025ddf9d76165abb2680",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114633",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114633",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113742 to 26.5: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114633",
          "createdAt": "2026-08-13T12:49:24Z",
          "updatedAt": "2026-08-13T12:49:32Z",
          "timestamp": "2026-08-13T12:49:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:ab70285349e845883afc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114632",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114632",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113742 to 26.3: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114632",
          "createdAt": "2026-08-13T12:48:47Z",
          "updatedAt": "2026-08-13T12:48:55Z",
          "timestamp": "2026-08-13T12:48:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:c809349d61c64145733c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112309",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112309",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add hierarchicalKMeans and assignCentroid",
          "text": "Adds functions for computing cluster centroids and assigning new vectors to clusters. Ref : https://github.com/ClickHouse/ClickHouse/issues/112578 ### Changelog category - Experimental Feature ### Changelog entry - Added` hierarchicalKMeans()` and `assignCentroid()` functions.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112309",
          "createdAt": "2026-07-28T16:11:44Z",
          "updatedAt": "2026-08-13T12:48:40Z",
          "timestamp": "2026-08-13T12:48:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "shankar-iyer",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c0d5e8dbb60fd6f445bf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110886",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110886",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "MaterializedPostgreSQL: coordinated Replicated/Shared nested tables for HA",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/47655 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): `MaterializedPostgreSQL` can now create its nested tables as `ReplicatedReplacingMergeTree`/`SharedReplacingMergeTree` for high availability, using a new Keeper-based single-active-worker coordination of the replication slot. Controlled by the new settings `materialized_postgresql_table_engine`, `materialized_postgresql_keeper_path` and `materialized_postgresql_replica_name`. ### Documentation entry for user-facing changes The three new settings and the coordinated-failover mechanism are documented in the embedded `Documentation` blocks of `DatabaseMaterializedPostgreSQL.cpp` and `StorageMaterializedPostgreSQL.cpp`, from which the published engine pages are generated. --- ## Problem The nested tables that `MaterializedPostgreSQL` auto-creates were hardcoded to plain `ReplacingMergeTree` ([`StorageMaterializedPostgreSQL.cpp`](https://github.com/ClickHouse/ClickHouse/blob/master/src/Storages/PostgreSQL/StorageMaterializedPostgreSQL.cpp)), so the engine could not be made highly available (issue #47655). Simply allowing `ReplicatedReplacingMergeTree` is unsafe on its own: a PostgreSQL logical replication slot permits only **one** active consuming session, so two ClickHouse replicas consuming the same slot would race on `pg_replication_slot_advance` and silently drop WAL. ## Solution Allow the nested engine to be `ReplicatedReplacingMergeTree` / `SharedReplacingMergeTree`, **gated** behind a new Keeper-based coordination that elects exactly one active worker — the same pattern used by `S3Queue`, the Keeper-coordinated `Kafka` engine and refreshable materialized views: - The active worker holds an ephemeral `/leader` node under `materialized_postgresql_keeper_path` and is the only replica that consumes the slot. - Standby replicas create the nested tables as replicas of the same replicated tree and receive data (including the initial snapshot) through ClickHouse replication, without touching the slot. - When the active worker's Keeper session ends, a standby wins `/leader` and resumes consuming from the slot's `confirmed_flush_lsn` — no reload, no duplication. PostgreSQL's own single-active-session rule on the slot is the ultimate backstop against double-advance during a handover. A graceful stop of the active worker (`DETACH`, a non-last `DROP`, server shutdown) does not wait for that: it releases `/leader` with a **confirmed** removal - the node lives under the server's shared Keeper session, which outlives the database, and nothing re-enters the election on that replica after shutdown, so an unconfirmed (lost-response) removal is resolved on the spot with owner- and version-checked re-checks instead of leaving a stale node that would keep every peer on standby for as long as that session lives. - The slot and the publication are **shared state**, and the whole lifecycle honors that: - A durable `snapshot_completed` marker in Keeper records that the initial snapshot loaded every table. A new leader may resume from `confirmed_flush_lsn` only when the marker exists; otherwise the previous worker died mid-snapshot, and the new one **clears the nested tables**, drops the slot and redoes the snapshot from scratch, so pre-slot rows cannot be lost. Clearing first is required for correctness: a row the dead worker had copied and that PostgreSQL then `DELETE`d has no counterpart in the new snapshot, so without a clear the stale copy would survive (a `ReplacingMergeTree` collapses duplicate keys by `_version` but never turns a now-absent row into a tombstone). The marker is fenced on the live leadership session: it is written through the Keeper session that backs `/leader` (in a multi-request with a check on `/leader`), the snapshot load aborts as soon as that session is lost, and a consumer is never started over a dead leadership session - so a worker deposed mid-snapshot can never mask its successor's replacement snapshot with a stale marker. The redo of the snapshot is fenced the same way: a worker whose leadership session is no longer alive aborts before truncating the nested tables (re-checked per table) and before dropping or recreating the shared slot, so a deposed worker cannot wipe the tables its successor has already reloaded or discard the slot the successor just created. A startup attempt that fails before a consumer got running (most importantly, a coordinated single-table engine whose one snapshot load failed) aborts instead of starting a consumer with nothing to apply - which would advance the shared slot's `confirmed_flush_lsn` while applying no rows - and releases the leadership so a healthy peer can take over. That release never leaves a stale leadership claim behind: a `remove` of `/leader` that could not be confirmed does not prove the node survived, so the claim is dropped in either case, and a leader node this replica created under the current Keeper session but no longer tracks is recognized (by its stored replica name and its owning session) and removed on its next election attempt - a replica never keeps acting as the active worker, or touches the shared slot and snapshot state, without provably holding `/leader`. - A second coordinated `CREATE DATABASE` adopts an existing publication instead of dropping it from under the active worker. - Every replica registers itself under `<keeper_path>/replicas`; dropping the database on one replica keeps the shared slot/publication for the others, and only dropping the last replica removes them from PostgreSQL (fail-close when Keeper is unavailable). The last-replica decision is fenced on the shared `replicas` node so concurrent `DROP DATABASE` on different replicas cannot both act as last, it removes the replica's registration only atomically with winning that fence - a replica that is not the last one keeps its registration until its local nested tables are actually gone, so even a server killed mid-drop stays visible to every later last-replica check - and it runs even for a `DROP DATABASE` issued immediately after a restart, before the background startup task has rebuilt the replication handler. - Recreating a coordinated database right after dropping it is safe: the nested tables are dropped without the delayed-drop window (so their shared Keeper subtrees do not outlive the `DROP DATABASE`), the snapshot insert is never silently deduplicated against block hashes surviving from a previous incarnation of the shared table, a publication leaked by an incompletely dropped setup (no surviving coordination state in Keeper) is dropped by a fresh `CREATE` instead of silently adopted with its stale table set, and a refused (failed) drop never leaves the replica silently dead: every refusable Keeper step of the pre-data teardown runs before the replica's consumer is stopped, and if the drop still fails after that point - the last replica's removal of the shared coordination nodes, or the deletion of this replica's own local nested tables (e.g. Keeper disappearing while a nested replicated table removes its own Keeper metadata) - replication is rebuilt in the background - for the database engine through a new generic `IDatabase::onDropDatabaseFailed` hook that discards the stopped handler and re-runs the startup task, for the single-table engine by re-arming the handler's retrying startup path - so the replica rejoins the setup once Keeper is reachable again. - All replicas of one coordinated setup must agree on the naming-affecting settings (`materialized_postgresql_table_engine`, `materialized_postgresql_schema`, `materialized_postgresql_schema_list`, `materialized_postgresql_tables_list_with_schema`) and must replicate the same PostgreSQL source — the same source database and, for the single-table engine, the same source table (so a single-table engine and a database engine can never share one keeper path): these determine how the ClickHouse names of the shared nested tables and the names of the shared slot and publication are derived. The first replica publishes a canonical fingerprint of them at `<keeper_path>/naming`; a disagreeing replica is rejected — synchronously at `CREATE` time when the setup already exists in Keeper, and fail-close at startup before it registers itself — instead of adopting the same publication yet building a disjoint replicated tree that never receives the other replicas' data. - The shared **table set** is fenced the same way, *before* any nested table is built: the first replica publishes its derived set at `<keeper_path>/table_set`, and a replica whose derived set differs is refused fail-close (the shared publication - from which joining replicas later derive their set - is only created by the elected active worker, so without this fence two fresh replicas starting concurrently could silently build diverging nested tables on one keeper path). The refusal converges by itself once the publication exists, because joining replicas then derive their set from the publication, which was created from the fenced set. - The last-replica teardown is generation-safe: winning the last-replica fence atomically (in one Keeper multi-request) creates an ownership token at `<keeper_path>/teardown`, which is removed only after the shared PostgreSQL slot/publication have actually been dropped. While the token exists, a fresh coordinated `CREATE` on the same keeper path is rejected (synchronously by validation and fail-close at startup), so the pending by-name drops can never delete a new setup's freshly created slot/publication. A replica recovering its own refused drop reclaims its own token, and a retried drop resumes its earlier teardown instead of leaking the slot. - Per-table and database-wide destructive changes are refused (`NOT_IMPLEMENTED`) in coordinated mode: `ATTACH TABLE` / `DETACH TABLE PERMANENTLY` / `DROP TABLE` / `RENAME TABLE` / `EXCHANGE TABLES`, and a database-wide `TRUNCATE DATABASE` / `TRUNCATE ALL TABLES FROM`. Each would only change the local replica (or wipe its copy of the shared replicated data) while the shared publication, tables list, slot/snapshot marker and peer replicas keep the old state, silently diverging the replicas. Recreate the database with an updated `materialized_postgresql_tables_list` instead. New settings (both the database engine and the single-table engine): | Setting | Default | Purpose | |---|---|---| | `materialized_postgresql_table_engine` | `ReplacingMergeTree` | `ReplacingMergeTree` / `ReplicatedReplacingMergeTree` / `SharedReplacingMergeTree` | | `materialized_postgresql_keeper_path` | (empty) | opt-in gate that enables coordination; supports the `{shard}` macro; a per-replica/per-server macro (`{replica}`/`{server_uuid}`, including reached through a config macro) is rejected at `CREATE` time, and so is `{uuid}` unless the DDL carries the UUID (`ON CLUSTER`, a table inside a `Replicated` database, or an explicit `UUID '...'` clause) so that it is provably identical on every replica, and so is a misspelled/unsupported macro (in this path or in `materialized_postgresql_replica_name`) — both settings are macro-expanded during validation exactly as the handler expands them later | | `materialized_postgresql_replica_name` | `{replica}` | replica identity for coordination and the nested replicated engine; must resolve to a distinct value on every replica, which is enforced: the `/replicas/<name>` registration node stores the owning replica's identity, and a name already registered by another replica is rejected (synchronously at `CREATE` time when the registration is already visible); it must also resolve to a single Keeper node name (empty or containing `/` is rejected, since a nested path under `/replicas` would break the last-replica bookkeeping); together with `materialized_postgresql_keeper_path` it forms the coordination identity of the replica, which is treated as immutable once the setup exists: a configuration-only change of a macro these settings expand through is refused at startup, while a `DROP` tears down the identity persisted in the nested tables; a name change made in the one window where the registration already exists while no nested table does yet is recovered from the registration itself, which stores an owner identity no macro feeds into, so the stale registration is removed (on startup and on drop) instead of keeping `<keeper_path>/replicas` non-empty forever and stopping every future drop from becoming the last-replica drop | The replicated/shared engines require `materialized_postgresql_keeper_path` and vice versa (coordination with a plain `ReplacingMergeTree` would leave the standbys without data), and coordination is mutually exclusive with `materialized_postgresql_use_unique_replication_consumer_identifier` (which gives each replica its own slot). Coordination also requires Keeper/ZooKeeper to be configured on the server: a coordinated `CREATE DATABASE` on a server with no Keeper is rejected synchronously at `CREATE` time rather than being accepted and left retrying in the background. All of this validation also applies to a user `ATTACH DATABASE` / `ATTACH TABLE` that spells out the full definition - it is fresh user input, exactly like a `CREATE`; only replaying an already-persisted definition (server startup, and the short `ATTACH` syntax, which re-reads the stored definition) is exempt. Behaviour is unchanged when coordination is not configured. ## Testing - Added integration test `test_postgresql_replica_database_engine/test_coordination.py`: convergence + single leader, leader failover with no data loss/duplication, rejoin, takeover before snapshot completion redoes the snapshot without losing pre-slot rows, a mid-snapshot takeover after a row is deleted in PostgreSQL drops the stale copy (`test_takeover_after_partial_snapshot_drops_stale_deleted_rows`), a second `CREATE` adopts the publication, `ATTACH`/`DETACH`/`DROP`/`RENAME`/`EXCHANGE`/`TRUNCATE` rejection, shared slot/publication kept until the last replica is dropped, a `DROP DATABASE` immediately after restart still unregisters the replica (`test_drop_immediately_after_restart_unregisters_replica`), concurrent `DROP DATABASE` on both replicas tears down the shared state exactly once (`test_concurrent_drop_on_both_replicas_removes_shared_state`), a per-server keeper path macro is rejected (`test_keeper_path_rejects_per_server_macro`), a plain coordinated `CREATE` with `{uuid}` in the keeper path is rejected while an explicit `UUID '...'` clause makes it acceptable (`test_keeper_path_rejects_uuid_macro_for_a_plain_create`), a leaked publication is replaced instead of adopted (`test_leaked_publication_is_not_adopted_by_fresh_coordinated_create`), a refused drop in the restart window keeps the startup alive (`test_refused_drop_in_restart_window_does_not_disable_startup`), a drop refused (via a failpoint) after the replication handler was already stopped recovers without a server restart for both the database and the single-table engine (`test_refused_drop_after_handler_shutdown_recovers_database`, `test_refused_drop_after_handler_shutdown_recovers_single_table_engine`), a drop refused because the local nested-table deletion itself fails likewise recovers for both engines (`test_refused_drop_when_nested_table_drop_fails_recovers_database`, `test_refused_drop_when_nested_table_drop_fails_recovers_single_table_engine`), coordinated `CREATE` rejected without Keeper configured (`test_coordination_requires_keeper_configured`), a joining replica with different naming-affecting settings is rejected at `CREATE` while identical settings converge (`test_join_with_different_naming_settings_is_rejected`), a coordinated single-table engine cannot join a database engine's keeper path because the fenced identity includes the PostgreSQL source (`test_single_table_engine_cannot_join_database_engine_keeper_path`), a bad macro in the keeper path or the replica name fails the `CREATE` synchronously (`test_bad_macro_in_coordination_settings_is_rejected_at_create`), a replica name that is not a single Keeper path component (empty, or containing `/`) is rejected while a plain name on the same keeper path is accepted (`test_replica_name_must_be_a_single_keeper_component`), a duplicate `materialized_postgresql_replica_name` is rejected without disturbing the registered replica (`test_duplicate_replica_name_is_rejected`), a joining replica whose derived table set differs from the fenced one is refused before building any nested table (`test_join_with_different_table_set_is_rejected`), a fresh `CREATE` on a keeper path whose teardown is still pending is rejected and succeeds once the teardown token is released (`test_create_is_rejected_while_teardown_token_is_held`), the three coordination settings rejected by `ALTER DATABASE ... MODIFY SETTING` with an actionable CREATE-time-only message (`test_coordination_settings_cannot_be_altered`), public DDL on a plain database in the startup window (before the background task has built the replication handler) does not touch a null handler and a `DROP DATABASE` in that window still removes the PostgreSQL publication and replication slot (`test_plain_database_ddl_and_drop_in_startup_window`), a plain `DROP DATABASE` quiesces the retrying background startup task so a retry waking mid-drop cannot recreate the publication/slot while the drop is in flight (`test_plain_drop_database_quiesces_retrying_startup_task`), a refused plain drop re-arms the startup task instead of leaving the database mounted but dead (`test_plain_refused_drop_rearms_startup_task`), an `ALTER` of a mutable setting on a coordinated standby is accepted before its consumer exists and survives a failover (`test_alter_mutable_setting_on_standby_survives_failover`), the same `ALTER` is accepted on a former leader demoted back to standby and applied on its next takeover (`test_alter_mutable_setting_on_demoted_leader`), a registration racing the very start of a last-replica teardown is refused by the atomic teardown-token fence (`test_registration_is_fenced_against_concurrent_teardown_token`), a joiner that adopted the shared publication's table set over its own mismatching `materialized_postgresql_tables_list` recreates a missing publication from the adopted set rather than its stale local list (`test_adopted_table_set_survives_publication_recreation`), that adopted set also survives a restart of the joiner, so a replica restarted while the publication is missing recreates it from the table set fenced in Keeper instead of its stale local list (`test_adopted_table_set_survives_restart_and_publication_recreation`), a coordination identity changed by a configuration-only macro change is refused at startup - with nothing registered under the new identity - and the setup resumes once the configuration is restored (`test_coordination_identity_must_stay_stable_across_restart`), a `DROP` after such a change (with the macro's value changed, or the macro removed entirely) tears down the original coordination identity persisted in the nested tables for both engines (`test_drop_after_coordination_identity_change_tears_down_original_identity`, `test_single_table_drop_after_coordination_identity_change_tears_down_original_identity`), a worker whose Keeper session expires while it is loading the initial snapshot aborts without publishing the `snapshot_completed` marker or starting a consumer while its successor redoes the snapshot (`test_lost_leadership_during_snapshot_does_not_publish_stale_marker`), a worker deposed right after entering the redo-the-snapshot recovery branch aborts at the leadership fence without truncating the tables its successor reloaded or dropping the successor's slot (`test_deposed_worker_aborts_redo_snapshot_before_touching_shared_state`), a coordinated single-table engine recreates an externally dropped publication idempotently - with the correctly quoted table name - and resumes replicating (`test_single_table_publication_recreated_after_external_drop`), a coordinated worker whose snapshot load keeps failing aborts each attempt before a consumer exists and releases the leadership, so a healthy peer completes the full snapshot and no WAL is discarded, for both the single-table and the database engine (`test_failed_single_table_snapshot_releases_leadership`, `test_failed_database_snapshot_releases_leadership`), a coordinated `DETACH TABLE ... PERMANENTLY` / `DROP TABLE` issued in the attach/restart window (before the background startup has published the table wrappers) is refused as a true no-op that leaves the nested table consuming (`test_coordinated_detach_in_startup_window_is_a_no_op_rejection`), a read in the recovery window of a refused `DROP DATABASE` still hides PostgreSQL-deleted row versions instead of falling back to the raw nested tables (`test_refused_drop_recovery_window_keeps_wrapped_reads`), a stale registration left behind by a replica whose coordination name changed before it owned any nested table is purged so the last-replica teardown still removes the shared slot and publication (`test_stale_registration_of_a_renamed_replica_is_purged`), a server hard-killed in the middle of a non-last `DROP DATABASE` stays registered in Keeper (the last-replica decision is one atomic operation), so a peer's drop keeps the shared state around the killed replica's surviving data and the restarted replica resumes replicating before a retried drop tears everything down (`test_hard_stop_during_non_last_teardown_keeps_replica_registered`), a graceful stop of the active worker whose fenced `/leader` removal fails (via a failpoint) still frees `/leader` through the confirmed-release re-check, so the peer takes over promptly while the stopped replica's server and Keeper session keep running (`test_graceful_stop_releases_leader_even_when_removal_fails`), a user `ATTACH DATABASE` / `ATTACH TABLE` with a full definition goes through the coordination validator like a `CREATE` (`test_full_attach_database_definition_is_validated`, `test_full_attach_table_definition_is_validated`), and the other negative validation cases. - Added `test_plain_single_table_engine_refused_drop_recovers` to `test_postgresql_replica_database_engine/test_3.py`: the plain (non-coordinated) single-table engine now drops its local nested table before the authoritative PostgreSQL teardown, so a refused (thrown) nested-table drop keeps the replication slot/publication and re-arms the handler to resume from the existing slot, instead of leaving the table mounted but dead with the PostgreSQL objects already removed. - Added `test_rename_and_exchange_table_are_rejected`, `test_database_wide_truncate_is_rejected` and `test_drop_of_individual_table_is_rejected` to `test_postgresql_replica_database_engine/test_1.py`: `RENAME TABLE` / `EXCHANGE TABLES`, a database-wide `TRUNCATE`, and a `DROP` / `TRUNCATE` of an individual table are now also rejected for a plain (non-coordinated) `MaterializedPostgreSQL` database, which previously fell through to the generic `Atomic` DDL and silently diverged the local tables from the replication state (a dropped table stayed in `materialized_postgresql_tables_list` and in the publication, so the consumer marked it skipped while the slot kept advancing); `DETACH TABLE ... PERMANENTLY` remains the supported removal path. - Added `test_count_does_not_include_deleted_rows` to `test_postgresql_replica_database_engine/test_1.py`: `SELECT count()` reads the cheapest column of the nested table - the one-byte `_sign` column itself - and the wrapped read used to skip its `_sign = 1` filter whenever the sign column was among the requested columns, so a count included the durable tombstones a PostgreSQL `DELETE` leaves in the nested `ReplacingMergeTree` table and disagreed with `SELECT *`. The filter is now applied unconditionally (an explicit read of `_sign` is filtered the same way). - Added `test_failed_attach_rolls_back_setting_and_table` to `test_postgresql_replica_database_engine/test_2.py`: a failed `ATTACH TABLE` now rolls back the already persisted `materialized_postgresql_tables_list` extension and the published table wrapper together with the nested table, so the database does not keep claiming a table is attached that never joined the publication; a retry starts from a clean state. - Added `test_failed_detach_rolls_back_and_table_keeps_replicating` to `test_postgresql_replica_database_engine/test_1.py`: a `DETACH TABLE ... PERMANENTLY` whose local nested-table drop throws is now rolled back completely - the table is re-added to replication through a fresh snapshot and the persisted `materialized_postgresql_tables_list` is restored - so the table keeps replicating and the `DETACH` can simply be retried, instead of stranding a live nested table outside the logical database. - Added `test_detach_database_and_reattach` to `test_postgresql_replica_database_engine/test_1.py`: `DETACH DATABASE` used to fail half-way (the generic detach path had already stopped replication when the per-table walk hit the unconditional `DETACH TABLE not allowed` guard), leaving the database mounted but no longer replicating. The internal walk is now let through, so `DETACH DATABASE` unmounts the database cleanly and `ATTACH DATABASE` resumes replication, catching up on changes made while it was detached. - Added `test_detach_permanently_of_last_table_is_rejected` to `test_postgresql_replica_database_engine/test_1.py`: detaching the last replicated table would persist an empty `materialized_postgresql_tables_list`, and an empty list does not mean \"replicate no tables\" - the table set would be re-derived from the current PostgreSQL schema on the next startup, so the detach would not stick. It is now refused up front, leaving the table replicating. - Verified end-to-end against a locally built binary with two real `clickhouse-server` nodes + `clickhouse-keeper` + PostgreSQL (`wal_level=logical`): coordinated `CREATE DATABASE` builds the `ReplicatedReplacingMergeTree` nested tables with correctly macro-expanded per-table Keeper paths; snapshot + ongoing INSERT/UPDATE are consumed by the single leader; a standby serves the same data via ClickHouse replication; killing the leader triggers takeover (~8s) that resumes consumption with matching row/key counts (no loss, no duplication); and a restarted node rejoins as a standby without stealing leadership. Note: `SharedReplacingMergeTree` is only available in ClickHouse Cloud; open-source runs are covered with `ReplicatedReplacingMergeTree`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110886",
          "createdAt": "2026-07-17T15:45:03Z",
          "updatedAt": "2026-08-13T12:48:29Z",
          "timestamp": "2026-08-13T12:48:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 46
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:517704da8be7d2794778",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114631",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114631",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113742 to 25.8: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114631",
          "createdAt": "2026-08-13T12:48:05Z",
          "updatedAt": "2026-08-13T12:48:13Z",
          "timestamp": "2026-08-13T12:48:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:d31181f2107bab78460a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114152",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114152",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "[WIP]Add incremental rmv core",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add incremental refreshable materialized views managed via refresh_incremental, which when used each refresh reads only the data committed to the single source table since the previous run and persists the advanced cursor in the RMV's Keeper CoordinationZnode for at-least-once resumption. depends on https://github.com/ClickHouse/ClickHouse/pull/111794 cc @alesapin @Michicosun",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114152",
          "createdAt": "2026-08-10T12:45:59Z",
          "updatedAt": "2026-08-13T12:47:58Z",
          "timestamp": "2026-08-13T12:47:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "SmitaRKulkarni",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c7f8e9acd4719a9b4897",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114461",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114461",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "CI: route documentation review requests",
          "text": "Enable documentation review requests now that the relevant teams have repository access. ClickPipes documentation changes request reviews from both `docs` and `clickpipes`; language-client and connector documentation changes request reviews from both `docs` and `integrations-ecosystem`. Documentation reviews are skipped when a PR also changes files under `src/`. GitHub App installation tokens cannot resolve private teams when requesting reviews, even with the documented repository permission. Use the existing robot credential for internal PRs and for a `pull_request_target` workflow that checks out only the trusted base revision. Fork PRs must have the `can be tested` label before the trusted workflow runs. This internal draft exercises the credential and API path from the PR branch. It includes a one-line Python language-client documentation edit, which should request reviews from both `docs` and `integrations-ecosystem`. Related: https://github.com/ClickHouse/ClickHouse/pull/114458 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable automatic documentation team review requests through a trusted workflow.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114461",
          "createdAt": "2026-08-12T09:57:54Z",
          "updatedAt": "2026-08-13T12:47:29Z",
          "timestamp": "2026-08-13T12:47:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-ci"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:66b87b9baf2cf010e4ce",
        "signalId": "github:ClickHouse/ClickHouse:issue:114630",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114630",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "KeyCondition's pointInPolygon primary-key analysis skips is_valid validation (unlike pointInPolygon/spatial_bbox/GeoParquet pruning)",
          "text": "### Describe the unexpected behaviour `KeyCondition`'s primary-key range analysis for `pointInPolygon` (`analyze_point_in_polygon` in `src/Storages/MergeTree/KeyCondition.cpp:3583-3663`) builds the query polygon directly from the literal argument, calls `boost::geometry::correct` and `boost::geometry::envelope`, but never calls `boost::geometry::is_valid`. This is inconsistent with the other two places in the codebase that parse the same kind of literal: - `FunctionPointInPolygon::parseConstPolygon`/`parseConstMultiPolygon` (`src/Functions/pointInPolygon.cpp:842-861,894-913`), which validate the assembled polygon with `bg::is_valid` and throw `BAD_ARGUMENTS` when `validate_polygons` is enabled (the default). - The `spatial_bbox` `MergeTree` skip index and GeoParquet row-group pruning (`src/Common/GeoBbox.h`, added in #104437), which validate every constant geometry argument and fail closed (decline to prune) rather than derive a bbox from an invalid one. `analyze_point_in_polygon` also only recognizes the plain 2-argument `pointInPolygon(point, ring)` form (see the `Case1 no holes in polygon` comment at line 3715); a polygon-with-holes literal (3+ arguments) isn't analyzed for primary-key pruning at all. That part is safe (it just declines to prune, so `KeyCondition` falls back to scanning), it's the missing validity check on the single-ring case that's the actual gap. ### Practical impact Under the default `validate_polygons = 1`, this gap is effectively unreachable through SQL: ClickHouse evaluates `WHERE`-clause constant expressions once on a zero-row block before any index analysis or pruning runs, so an invalid constant polygon literal passed to `pointInPolygon` always raises `BAD_ARGUMENTS` immediately — `KeyCondition`'s analysis, which runs later during part/granule selection, never gets a chance to act on the unvalidated bbox. However, if a user explicitly sets `validate_polygons = 0` — a documented, supported way to bypass `pointInPolygon`'s own geometry validation — the dry-run exception no longer fires, and `KeyCondition`'s PK-range analysis still unconditionally computes a bbox/envelope from the same, now genuinely unvalidated, polygon and may use it to prune primary-key ranges. `boost::geometry`'s query algorithms (`within`, `intersects`, etc.) have undefined results for invalid geometries, so pruning decisions derived from such a polygon are not guaranteed to be sound. ### How to reproduce Not reproducible with a concrete wrong-result example yet — this is a code-review finding, not an observed bug. The scenario would require: a `MergeTree` table with a `pointInPolygon`-friendly primary key, `SETTINGS validate_polygons = 0`, and a self-intersecting/otherwise invalid constant polygon literal in the `WHERE` clause, compared against a full scan of the same query. ### Expected behavior `analyze_point_in_polygon` should either validate the polygon the same way `FunctionPointInPolygon` and `Common/GeoBbox.h` do (and fail closed / decline to prune when invalid, honoring `validate_polygons` the same way the function itself does), or the inconsistency should be a deliberate, documented decision. ### Additional context Found while auditing ClickHouse's geospatial pruning code for duplicated/drifted logic during #104437 (which unified the equivalent bbox-extraction-and-validation logic between the `spatial_bbox` skip index and GeoParquet row-group pruning in `src/Common/GeoBbox.h`). Filing this as a lower-priority follow-up for future reference rather than addressing it in that PR, since it's a different subsystem (primary-key range analysis) and not currently reachable under default settings. Related: https://github.com/ClickHouse/ClickHouse/pull/104437",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114630",
          "createdAt": "2026-08-13T12:47:16Z",
          "updatedAt": "2026-08-13T12:47:16Z",
          "timestamp": "2026-08-13T12:47:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "bacek",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:0bb5bfd3570512299a0a",
        "signalId": "github:ClickHouse/ClickHouse:issue:111879",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:111879",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "pointInPolygon: throwing CAST(tuple, 'Point') hoisted into PREWHERE ahead of its NULL guard (CANNOT_INSERT_NULL), regression from #107988",
          "text": "**Describe what's wrong** A `WHERE` predicate that casts a tuple to `Point` (a non-`Nullable` type) throws `CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN` at execution, even though a preceding filter guarantees the cast never sees a NULL. The throwing `CAST(tuple(...), 'Point')` is materialized in a PREWHERE read step **before** the guarding filter is applied, so it runs on rows the guard would have removed. The query returns a result on 26.2 and throws on 26.6. ``` Code: 349. DB::Exception: Cannot convert NULL value to non-Nullable type: while executing 'FUNCTION CAST(tuple(__table1.longitude, __table1.latitude) :: 7, 'Point'_String :: 9) -> CAST(tuple(__table1.longitude, __table1.latitude), 'Point'_String) Point : 4': While executing MergeTreeSelect(pool: PrefetchedReadPool, algorithm: Thread). (CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN) ``` **Does it reproduce on the most recent release?** Yes, reproduces on 26.6. Good on 26.2. **How to reproduce** ```sql DROP TABLE IF EXISTS repro_pip SYNC; CREATE TABLE repro_pip ( id Int64, session_id Int64, payload JSON, filler String ) ENGINE = MergeTree ORDER BY id; INSERT INTO repro_pip SELECT number, number, if(number % 3 = 0, '{}', '{\"longitude\":4.35,\"latitude\":52.06}')::JSON, -- 1/3 of rows have no coords repeat('x', 200) FROM numbers(500000); -- Throws Code 349 on 26.6; returns 333333 on 26.2. SELECT count(DISTINCT session_id) FROM ( SELECT session_id, CAST(payload.latitude AS Nullable(Float64)) AS latitude, CAST(payload.longitude AS Nullable(Float64)) AS longitude FROM repro_pip WHERE payload.longitude IS NOT NULL -- guards the cast; latitude itself is never guarded ) a WHERE 1=1 -- constant-true term is the trigger AND pointInPolygon((longitude, latitude)::Point, readWKTPolygon('POLYGON ((4.3 52.0,4.4 52.0,4.4 52.1,4.3 52.1,4.3 52.0))')) = 1; ``` Essential ingredients (established by minimization; each is required): - A `JSON` column read via subcolumns (`payload.longitude` / `payload.latitude`). Plain `Nullable(Float64)` columns do **not** reproduce it — the dynamic-subcolumn read is what makes the plan materialize the cast in a PREWHERE read step. - `latitude` / `longitude` NULL together on some rows. - A subquery/CTE guard `payload.longitude IS NOT NULL` (latitude itself is never guarded). - A wide row (`filler`) so move-to-prewhere engages. - The outer `WHERE 1=1 AND pointInPolygon(...)`. The constant-true `1=1` is the trigger: it leaves a residual `Filter` in addition to the PREWHERE, and the throwing cast ends up materialized in a PREWHERE read step evaluated over the whole granule, before the guard filters the NULL rows. **Expected behavior** The query should not throw: the `payload.longitude IS NOT NULL` guard removes the NULL rows before the cast, as it did on 26.2 (returns `333333`). A potentially-throwing expression must not be evaluated in a read step that precedes the filter guarding it. Notably, adding `AND payload.latitude IS NOT NULL` to the subquery does **not** help — the cast is hoisted ahead of that guard too. So this cannot be worked around with NULL guards. **Bisect** First bad commit: `74a2218e61bf3bee47bad4851817d372bbfe4ee1`, the merge of #107988 (\"Fix performance regression for Map subcolumns with PREWHERE\"), merged to master 2026-06-24. Bisected on `origin/master` with the repro above (good `38ac1a2` 26.2-era → bad `bb9c5e5` post-26.6). The changed files match the mechanism: `MergeTreeSplitPrewhereIntoReadSteps.cpp`, `MergeTreeWhereOptimizer.cpp`, `ReadFromMergeTree.cpp`, `MergeTreeSelectProcessor.*`, and subcolumn serialization (`SerializationSparse` / `SerializationMapKeyValue` / `ISerialization`). The PR fixed a Map-subcolumn PREWHERE *performance* regression; its PREWHERE-split changes introduced this *correctness* regression. Caused by: https://github.com/ClickHouse/ClickHouse/pull/107988 **Workarounds** (both return `333333` on 26.6) - `SETTINGS optimize_move_to_prewhere = 0` (also: `allow_reorder_prewhere_conditions = 0`, `query_plan_merge_filters = 0`). - Make the cast NULL-safe: `pointInPolygon((ifNull(longitude, 0.), ifNull(latitude, 0.))::Point, readWKTPolygon('POLYGON ((...))')) = 1`.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/111879",
          "createdAt": "2026-07-25T07:49:13Z",
          "updatedAt": "2026-08-13T12:47:15Z",
          "timestamp": "2026-08-13T12:47:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug"
          ],
          "author": "fm4v",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:9e96585461099d568dbb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112824",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112824",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cross-compile ClickHouse for Windows",
          "text": "Draft. Cross-compiles `clickhouse.exe` for `x86_64-w64-windows-gnu` (mingw-w64 + clang + lld → a native PE, no emulation layer and no MSVC licence). The goal is `clickhouse-client` and `clickhouse-local` running natively on Windows. **State: it links, it has never been run.** The development host is aarch64 and its Wine has no x86-on-ARM emulator, so everything here is compiled and linked but not executed. That is the next step and needs an x86-64 Windows machine. ``` programs/clickhouse.exe: PE32+ executable for MS Windows 6.00 (console), x86-64 520 MB · imports ADVAPI32, IPHLPAPI, KERNEL32, USER32, WS2_32, dbghelp, msvcrt ``` Every import ships with Windows, so the binary is a single self-contained file. Linux builds and links `clickhouse` unchanged after every commit in the branch. CI builds `clickhouse` for Windows on every pull request, so the port cannot silently regress. ## What is here The runtime (libc++/libc++abi/libunwind with SEH, compiler-rt builtins), 117 contribs behind a self-maintaining CI gate, all of Poco — whose Windows layer had been deleted from this fork and is restored from upstream 1.9.3 — every library under `src`, and `programs`. Roughly 900 of the changed lines are one mechanical class: `std::filesystem::path` used where a `String` is wanted, which compiles on POSIX through an implicit conversion that does not exist on Windows. Implemented natively rather than stubbed, because a client needs them: terminal size and console encoding, raw console mode and keystroke reading, Ctrl+C and Ctrl+Break, socket liveness, `statvfs` via `GetDiskFreeSpaceEx`, file mapping, `pread`, per-thread CPU time, `setThreadName`, `isLocalAddress`, `readpassphrase`, and a `WakeupFd` built on a loopback socket pair (a Windows pipe cannot be waited on alongside sockets). Compiled out with the reason recorded at each site, all server-side or POSIX-only: the sampling profiler and signal handlers (Windows reports faults through SEH), `ThreadFuzzer`, the `fork`-based watchdog, `ShellCommand` and everything built on it, the pseudo-terminal features, and the `su`/`docker-init`/`install` tools. ## Known gaps - Not executed, as above. - `-g0` on Windows: a PE image cannot exceed 4 GiB and the DWARF alone is several times that, and PE has no `.gnu_debuglink` equivalent to carry it separately. A crash symbolizes to module and offset, not file and line. - No `Epoll` backend, so nothing that polls sockets through it works yet. - Local syslog, archives (`libarchive` needs a hand-written Windows `config.h`), conditional writes to the local object storage (they need `flock` on a directory and sub-second modification times), and the web terminal report `NOT_IMPLEMENTED`. `docs/en/development/build-cross-windows.md` has the build instructions and a per-subsystem inventory of what remains. Related: https://github.com/ClickHouse/ClickHouse/pull/112185 Related: https://github.com/ClickHouse/ClickHouse/pull/112767 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a cross-compilation target for Windows (`x86_64-w64-windows-gnu`), producing a native `clickhouse.exe`. The build is not yet tested at runtime. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112824",
          "createdAt": "2026-07-31T21:57:07Z",
          "updatedAt": "2026-08-13T12:46:59Z",
          "timestamp": "2026-08-13T12:46:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 15
          },
          "labels": [
            "pr-build",
            "submodule changed"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7f30047b2da008d64fc7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111946",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111946",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a data race on publication of per-user ProfileEvents counters",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/105056 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a data race on the `ProfileEvents::Counters` parent chain. A freshly constructed per-user counters object was published into a chain that other threads traverse lock-free using a relaxed store, so a thread that observed the pointer was not guaranteed to see the pointee fully constructed. The publication is now release-ordered and the loads that dereference it are acquire-ordered. ### Description `ProcessList::insert` constructs a new `ProcessListForUser` for a user that has no entry yet. Its embedded `user_performance_counters` is initialized with ordinary non-atomic writes, and its address is then published into the shared `ThreadGroup`'s counters chain by `setUserCounters`, which stored `parent` with `std::memory_order_relaxed`. A relaxed store pairs with nothing, so the constructor's writes were not ordered before a consumer's reads. Any thread walking the same chain (`Counters::increment` / `incrementNoTrace` / `incrementSignalSafe`) could therefore dereference a pointer to an object it was not guaranteed to see initialized. `setParent` had the same relaxed publication, and it is used on every thread-group attach and by `attachProfileCountersScope`, which publishes a scope-local `Counters`. The fix is local to `Counters`, which owns both the chain and its traversal: the two publication stores become `memory_order_release`, and every load that dereferences what it loads becomes `memory_order_acquire`. The traversal loads were previously implicit `std::atomic` conversions, that is `seq_cst`, so on the increment path this is a small relaxation rather than a strengthening. On x86-64 all of these compile to a plain `mov`; on ARM the loads go from `ldar` to `ldapr`. The publication stores are cold: once per thread-group attach, and once per `ProcessList` insertion. This is a publication-ordering defect, not a use-after-free and not a rehash invalidation. `user_to_queries` entries are never erased (`ProcessListEntry::~ProcessListEntry` documents this, and `getUserInfo` relies on it), and `UserToQueries` is node-based so element addresses are stable. The write side is the initial construction. The report shape only became possible after #105056, which introduced the `cpus` field and `fetchAdd` that the read side touches. Reproduced locally on a ThreadSanitizer build, where the reported stacks are: ``` Read of size 8 by main thread (mutexes: write M0): ProfileEvents::Counters::fetchAdd(...) src/Common/ProfileEvents.cpp ProfileEvents::Counters::incrementSignalSafe(...) src/Common/ProfileEvents.cpp:1980 DB::(anonymous namespace)::writeTraceInfo(...) src/Common/QueryProfiler.cpp:132 DB::QueryProfilerReal::signalHandler(...) src/Common/QueryProfiler.cpp:595 ... std::condition_variable::wait interrupted in DB::ExternalLoader::LoadingDispatcher::loadImpl(...) DB::registerStorageDictionary(...) src/Storages/StorageDictionary.cpp:385 DB::InterpreterCreateQuery::execute() DB::(anonymous namespace)::loadStartupScripts(...) programs/server/Server.cpp:1134 DB::Server::main(...) programs/server/Server.cpp:3541 Previous write of size 8 by thread T190 (mutexes: write M1): ProfileEvents::Counters::Counters(VariableContext, ProfileEvents::Counters*) src/Common/ProfileEvents.cpp:1678 DB::ProcessListForUser::ProcessListForUser(...) src/Interpreters/ProcessList.h:335 ... operator new of the __hash_node<..., DB::ProcessListForUser> for ProcessList::insert ``` The mutexes on the two sides are disjoint, so nothing synchronized the publication with the traversal. The regression test added to `tests/integration/test_startup_scripts/` reproduces this without any private configuration: a startup script creates a non-lazy dictionary whose `CLICKHOUSE` source authenticates a user other than `default`. `registerStorageDictionary` then blocks the main thread in `ExternalLoader::LoadingDispatcher::loadImpl` while the loader thread, sharing the main thread's `ThreadGroup`, performs the first `ProcessList::insert` for that user. With the profiler sampling the startup thread every 1 ms, the signal handler walks the chain while the new counters are being published. Verified in both directions on a ThreadSanitizer build: without the change the test fails on every attempt with the stacks above, with the change it passes 12 consecutive runs (144 server restarts) with no report.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111946",
          "createdAt": "2026-07-26T09:39:06Z",
          "updatedAt": "2026-08-13T12:45:50Z",
          "timestamp": "2026-08-13T12:45:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:d31dc8bc2d7541066b1d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111494",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111494",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Text index: add trivial count optimization",
          "text": "Currently, the text index direct read optimization deserialize the sparse index, dictionary block and postings when there is token that exists in the index. Once the postings is read from disk, it fills the newly created boolean virtual column with postings data. With this optimization, we aim to reduce reading postings from disk and creating a virtual column. Instead we can answer queries using the token metadata from the dictionary block for specific query patterns as follows:. 1. `SELECT count() FROM table WHERE hasToken(column, 'foo');` 2. `SELECT count() FROM table WHERE hasAnyTokens(column, ['foo', 'bar']);` 3. `SELECT count() FROM table WHERE hasAllTokens(column, ['foo', 'bar']);` For the 1. case, we can avoid reading postings at all and use the cardinality metadata stored in the dictionary block to answer the query. For 2. and 3. cases, we would still read the postings but can avoid creating a virtual column. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Returns `COUNT()` queries directly from the text index cardinality metadata.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111494",
          "createdAt": "2026-07-22T22:33:56Z",
          "updatedAt": "2026-08-13T12:45:34Z",
          "timestamp": "2026-08-13T12:45:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-performance"
          ],
          "author": "ahmadov",
          "state": "open",
          "assignees": [
            "Ergus",
            "CurtizJ",
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:aeec2073ab45acd40763",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104217",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104217",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix SQLite WHERE predicate pushdown for strings with special characters",
          "text": "`StorageSQLite::read` used `LiteralEscapingStyle::Regular`, which escapes single quotes as `\\'`. SQLite does not recognise backslash escapes; its only valid string escape is `''`. A pushed-down predicate like `WHERE col = 'it\\'s'` causes SQLite to parse `'it\\'` as a closed string and `s'` as a stray token — a SQL syntax error or injection vector. Switching to `LiteralEscapingStyle::PostgreSQL` would fix single quotes but still emit `\\n`, `\\r`, `\\t` as backslash sequences (which `writeAnyEscapedString` applies unconditionally). SQLite does not interpret those, so predicates on control-character strings would silently return no rows. This PR adds a dedicated `LiteralEscapingStyle::SQLite` backed by `writeQuotedStringSQLite`: only `'` → `''`; all other bytes (including `\\`, newline, tab) are embedded literally. NUL bytes cannot be embedded — SQLite's tokenizer loop in `sqlite3GetToken` terminates on `c==0` even inside a string literal, returning `TK_ILLEGAL` — so a predicate whose string literal (possibly nested in an `IN` tuple, array or map) contains a NUL byte is not pushed down at all: ClickHouse evaluates it locally, and with `external_table_strict_query = 1` the query is rejected instead of silently returning wrong rows. This is a follow-up to PR #74144 which fixed the DDL/PRAGMA and INSERT paths for SQLite but left the SELECT pushdown path using the wrong escaping style. During review the same class of bug was fixed on the PostgreSQL pushdown path as well: strings nested inside `Array` / `Tuple` / `Map` literals (e.g. the elements of a pushed-down `IN` list) now stay in the selected dialect all the way down instead of falling back to the regular ClickHouse escaping, and PostgreSQL string literals that contain backslashes or control characters are emitted as escape string constants (`E'...'`), so the server reads back exactly the original bytes regardless of `standard_conforming_strings` (a real tab used to be sent as the two characters `\\t`). Predicates whose string literals contain a NUL byte are not pushed down to PostgreSQL either, since a PostgreSQL string value cannot contain NUL. The same row-value restriction is applied on the normal `WHERE` pushdown path: a multi-column tuple is written as the row value `(a, b)`, which SQLite and MySQL accept only next to a comparison or `IN`, so a predicate such as `WHERE (id, val) IS NOT NULL` is no longer pushed down to them (ClickHouse evaluates it, and with `external_table_strict_query = 1` the query is rejected) instead of being sent as SQL the external database cannot parse (SQLite reports `row value misused`). For PostgreSQL, whose row constructors are ordinary value expressions, it is still pushed down. A tuple used as the whole condition is ClickHouse's list-of-predicates form and keeps being pushed down as a conjunction, `WHERE (\"a\" > 0) AND (\"column\" > 10)`, for every dialect - no external database accepts a row value as a condition. The user-provided `(SELECT ...)` table argument of `sqlite` / `postgresql` / `mysql`, which is re-serialized from the parsed AST and sent to the external database as is, no longer leaks ClickHouse-only syntax into that SQL: `Array` / `Map` literals and tuples with fewer than two elements (which could only be written back as `tuple(...)`) now throw `BAD_ARGUMENTS` instead of producing SQL the external database cannot parse, an explicit `tuple(a, b)` call is re-serialized as the parenthesized row value `(a, b)` - for SQLite and MySQL only in positions where those databases accept a row value (an operand of a comparison or `IN`); in any other position, such as the SELECT list, both the `tuple(...)` call and the equivalent tuple literal throw `BAD_ARGUMENTS`, because the parenthesized form is a syntax error there (SQLite reports `row value misused`). PostgreSQL row constructors are ordinary value expressions, valid in any expression position (`SELECT (a, b)`, `WHERE (a, b) IS NOT NULL`), so for PostgreSQL such tuples are sent through as row values everywhere instead of being rejected - everywhere except a boolean position, since no database accepts a record as a condition. A tuple in a boolean position - the `WHERE` / `HAVING` of the passed query, or an operand of `AND` / `OR` / `NOT` - is ClickHouse's list-of-predicates form, and is lowered to a conjunction for every dialect: `(SELECT ... WHERE (a > 0, b > 10))` reaches the external database as `WHERE (a > 0) AND (b > 10)`, the same rewrite the normal pushdown path applies; `PREWHERE`, which is ClickHouse-only syntax no external database can parse, is lowered into `WHERE` on that path as well (merging with an existing `WHERE` via `AND`), and the lowered filter gets the same boolean-position normalization. the equivalent tuple literal of constants is not a list of predicates the external database could evaluate and throws `BAD_ARGUMENTS` there instead. `array` / `map` calls on that path are rejected for all three databases. The internal `_CAST(literal, 'Type')` wrapper that the analyzer's `ConstantNode::toAST` puts around a tuple literal used as a plain expression operand (e.g. `WHERE (id, val) = (2, 'y')`) when it re-serializes the subquery argument from the query tree is unwrapped back to the literal, instead of leaking the ClickHouse-internal `_CAST` function into the SQL sent to the external database. A single-row multi-column `IN` set keeps its outer parentheses for both carriers - the fast-path literal `(a, b) IN ((1, 'x'))` and the explicit call `(a, b) IN (tuple(1, 'x'))` - so it reaches the external database as `IN ((1, 'x'))` instead of collapsing to the scalar list `IN (1, 'x')`. This normalization applies to MySQL as well: it shares the same re-serialization path, and although its `Regular` literal escaping style is correct for MySQL string literals (MySQL interprets backslash escapes like ClickHouse), the `tuple(...)` / `array(...)` / `map(...)` forms and `Array` / `Map` / single-element-tuple literals are not MySQL syntax either. The JDBC/ODBC (`StorageXDBC`) pushdown path is intentionally out of scope: the bridge protocol only reports the identifier quoting style, not the literal escaping dialect of the remote database, so plumbing a dialect-aware escaping style through it needs a bridge protocol extension. That path keeps the historical `Regular` escaping, and the limitation is now documented at the call site. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect SQL literal escaping in `StorageSQLite` and `sqlite()` table function when pushing `WHERE` predicates to SQLite: single quotes and control characters (`\\n`, `\\r`, `\\t`, `\\`) were escaped with backslashes, which SQLite does not interpret, causing syntax errors or wrong query results. Also fixed the escaping of string literals pushed down to PostgreSQL: strings nested inside `IN` lists kept ClickHouse escaping, and control characters were sent as backslash sequences that PostgreSQL reads back as different bytes. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104217",
          "createdAt": "2026-05-06T11:55:23Z",
          "updatedAt": "2026-08-13T12:45:08Z",
          "timestamp": "2026-08-13T12:45:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 18
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:960f6aea0802152b099f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114578",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114578",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not erase the source column of a filter deferred after FINAL",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114512 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `THERE_IS_NO_COLUMN` / `NOT_FOUND_COLUMN_IN_BLOCK` when a filter that runs after `FINAL` is a plain column reference. This affected an explicit `PREWHERE b` under `FINAL` and, with default settings, a row policy whose expression is a bare column. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/114512 `SELECT count() FROM t FINAL PREWHERE b` threw `Code: 8. Cannot find column 'b' in source stream`. No `SETTINGS` clause is needed: `apply_row_policy_after_final` defaults to 1, so a row policy defers `PREWHERE` past `FINAL`. Root cause: a filter deferred after `FINAL` runs in a `FilterTransform` built with `remove_filter_column = true`. When the predicate is a bare column reference, the predicate node *is* the DAG's own input node, so the transform erases the source column the stream carries. The existing `restoreDAGInputs` guard cannot help: it only re-adds an input that is not already an output, and a bare predicate is already one. A wrapped predicate (`b = 1`, `NOT b`) has a distinct result node, so erasing it leaves the input alone. Hence only the bare form failed. The fix clears the local `remove_column` flag in `add_deferred_filter` when the filter column is one of the DAG's own inputs, which is how `optimizePrewhere.cpp` already resolves the same case. That lambda is shared by the deferred row policy and the deferred `PREWHERE`, so one change covers both. The defect is broader than reported: a bare-column row policy with no `PREWHERE` anywhere in the query fails identically under default settings. That shape is in the test. Validated in both directions on a debug build. Every shape in the issue's matrix now returns what its `WHERE` equivalent returns, and the new test fails on an unfixed binary. Covered: all four projections, subcolumns, a stacked policy plus `PREWHERE`, the four FINAL engines, and `Nullable`/`LowCardinality`/`Float` predicates. `SELECT *` stays byte-identical to the `WHERE` equivalent. A versioned fixture whose predicate flips across deduplication pins that the filter runs after `FINAL`, on the source column. Pre-existing and unchanged here: a row policy on a bare subcolumn fails `NOT_FOUND_COLUMN_IN_BLOCK` on the non-deferred read path, with and without `FINAL`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114578",
          "createdAt": "2026-08-13T02:56:02Z",
          "updatedAt": "2026-08-13T12:43:49Z",
          "timestamp": "2026-08-13T12:43:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:273ec602dc1d7fc9c8d3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:63383",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:63383",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve the performance of `MODIFY TTL`",
          "text": "`ALTER TABLE ... MODIFY TTL` currently rewrites every part of the table, which on a large table means reading and writing all of its data just to change when rows expire. Very often the new TTL is the old one shifted in time - the retention period is extended or shortened, e.g. `create_time + INTERVAL 300 DAY` becomes `create_time + INTERVAL 10 DAY`. In that case every row's expiry time moves by the same constant number of seconds, so the parts do not have to be rewritten at all: it is enough to shift the expiry timestamps ClickHouse already stores per part. This pull request adds that fast path, and the result for the user is that such a `MODIFY TTL` completes almost instantly instead of taking minutes or hours: ```sql CREATE TABLE test_fast_ttl (`id` UInt32, `name` String, `create_time` DateTime) ENGINE = MergeTree ORDER BY id TTL create_time + toIntervalDay(300); INSERT INTO test_fast_ttl SELECT number, 'AAA', date_sub(day, 100, now()) from numbers(100000000); -- Before ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 25.564 sec. -- After ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 0.046 sec. ``` There is nothing to enable and no new syntax: the optimization is applied automatically inside the `MATERIALIZE TTL` mutation that `MODIFY TTL` already produces, and a plain `ALTER TABLE ... MATERIALIZE TTL` benefits from it as well. The observable result is exactly the same as before - the same rows expire and the parts end up with the same TTL bounds - only the work is avoided. Per part, the mutation now does one of the following: - the part is fully expired under the new TTL - it is replaced with an empty part; - no row of the part is expired yet - the part is cloned (its data files hardlinked) and only its stored TTL bounds are shifted; - otherwise - the part is rewritten exactly as before. The fast path is only taken when it is provably equivalent to the rewrite. It requires that the unconditional rows TTL (`TTL <expr>`) is the only TTL of the table, and that the old and the new TTL are the same date/time column shifted by constant fixed-length intervals, so that `new_ttl(row) - old_ttl(row)` is one constant for every row. Calendar `MONTH`/`YEAR` intervals, `DAY`/`WEEK` intervals in a time zone with daylight saving time, and row-dependent expressions are all rejected. The proof is redone for each part against the TTL expression (and time zone) that the part's stored timestamps were actually computed under, so a part that lags the table metadata, or was written by an older server, falls back to the regular rewrite rather than being shifted unsoundly. The same applies to the boundary cases of the stored timestamps themselves: a part containing a row whose TTL timestamp is exactly `1970-01-01 00:00:00` UTC (which ClickHouse treats as \"no TTL\"), a shift that would move some timestamp onto that value, and a part whose stored TTL is already fully expired (which the regular rewrite drops wholesale, even when the new TTL is longer) all take the regular rewrite. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `ALTER TABLE ... MODIFY TTL` no longer rewrites the table's data when the new TTL is the old one shifted by a constant amount of time (the same date/time column plus fixed-length intervals), which is the common case of extending or shortening the retention period. Fully expired parts are replaced with empty ones and the rest are cloned with only their stored TTL metadata shifted, which makes such an `ALTER` nearly instant. Cases where the shift is not provably constant - calendar month/year intervals, day/week intervals in a time zone with daylight saving time, or row-dependent TTL expressions - fall back to the regular rewrite. `ALTER TABLE ... MATERIALIZE TTL` benefits from the same optimization. <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> > Information about CI checks: https://clickhouse.com/docs/en/development/continuous-integration/ <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration Tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASAN - [ ] <!---ci_include_tsan--> All with TSAN - [ ] <!---ci_include_analyzer--> All with Analyzer - [ ] <!---ci_include_azure --> All with Azure - [ ] <!---ci_include_KEYWORD--> Add your option here #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration Tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASAN - [ ] <!---ci_exclude_tsan--> All with TSAN - [ ] <!---ci_exclude_msan--> All with MSAN - [ ] <!---ci_exclude_ubsan--> All with UBSAN - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 - [ ] <!---ci_exclude_KEYWORD--> Add your option here #### Extra options: - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 <details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/63383",
          "createdAt": "2024-05-05T15:40:47Z",
          "updatedAt": "2026-08-13T12:43:38Z",
          "timestamp": "2026-08-13T12:43:38Z",
          "metrics": {
            "reactions": 1,
            "comments": 33
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "zhongyuankai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a0daa7c27a22899746ed",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112890",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112890",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix transform_null_in=1 for a non-Nullable key vs a Nullable IN set",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/111340 Closes: https://github.com/ClickHouse/ClickHouse/issues/112905 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes several problems with a non-`Nullable` primary key column compared against a subquery or array whose values are `Nullable`, for example `s IN (SELECT s FROM t UNION ALL SELECT NULL)` under `transform_null_in = 1`. Such a query no longer fails with `Cannot convert NULL value to non-Nullable type`. `NOT IN` and `NOT has` no longer drop rows whose key value happens to equal the key type's default (`''` for `String`, `0` for numbers), which previously happened because the `NULL` was folded into that default and then used to prune partitions. And `IN` / `NOT IN` no longer return wrong results for three specific conversions that map two distinct key values onto one set value: between different text types, across a loss of `Decimal` / `DateTime64` / `Time64` scale, and where a temporal element cannot represent one of the key's components. Other conversions that can collapse are not addressed here and are unchanged. ### Description Three defects in MergeTree `KeyCondition` set-index analysis (`tryPrepareSetColumnsForIndex`). **1. Exception.** `canBeSafelyCast(Nullable(X), T)` returned `true` for a non-`Nullable` `String` target, where a `NULL` has no representation, so the strict `castColumn` branch threw error 349 instead of the NULL-safe fallback. Fixed in the predicate. **2. Rows silently dropped.** A source-`NULL` position carried the key type's DEFAULT, injecting a value the query never wrote into the pruning set. That weakens `IN` and, after negation, STRENGTHENS `NOT IN`: ```sql CREATE TABLE t (s String) ENGINE = MergeTree ORDER BY s PARTITION BY s; INSERT INTO t VALUES ('a'), ('b'), (''); SELECT s FROM t WHERE s NOT IN (SELECT 'a' UNION ALL SELECT NULL) ORDER BY s SETTINGS transform_null_in = 1; -- returned only 'b': the '' row was pruned, because the dropped NULL had been folded to '' ``` The row is now dropped instead, which is result-neutral: such a `NULL` can never match a key reaching this block, whose outer type is never `Nullable` there. **3. Exactness was not gated on the conversion.** Index preparation casts set values INTO the key type while runtime membership casts the KEY into the set's type, and `castColumnAccurateOrNull` only proves nothing overflowed. Both are now checked: a set-to-key cast that is not equality-preserving marks the atom relaxed, and a key-to-set cast that can collapse two distinct keys onto one set value makes it DECLINE, since relaxation only forces `can_be_false`. **Scope.** That predicate detects the three classes above, not every possible one, as its header comment says. Not closed here: a key type with no strict round-trip check at all, because `accurate::convertNumeric(strict = true)` needs an integer or float SOURCE and here the source is the KEY, so a `DateTime` key against a `UInt8` element still collapses as on master. Closing it by inverting the default was measured and rejected: it costs three of this PR's own liveness controls for zero correctness gain, since collapse depends on the conversion DIRECTION rather than the key type. Four residual shapes are tracked separately; none is regressed here. Validated by 42 labelled cases in `04545_transform_null_in_non_nullable_key`, each asserting the result plus, via `EXPLAIN indexes = 1`, the set size and pruned part count; seven are controls. Two `03733` reference lines move as the smaller-but-still-exact set is built.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112890",
          "createdAt": "2026-08-01T09:23:38Z",
          "updatedAt": "2026-08-13T12:41:51Z",
          "timestamp": "2026-08-13T12:41:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.5-must-backport"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:0413d4fd51698596b944",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108653",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108653",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support `GROUPS` frame mode for window functions",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> The query below applies the same `1 PRECEDING AND 1 FOLLOWING` bounds as a `ROWS`, a `RANGE`, and a `GROUPS` (this PR) frame. The `order` column contains duplicate and non-consecutive values, so the three modes cover different rows: ```sql CREATE TABLE wf_frame_groups (`order` UInt64, value UInt64) ENGINE = Memory; INSERT INTO wf_frame_groups FORMAT Values (10, 1), (10, 2), (20, 3), (30, 4), (30, 5); SELECT order, value, groupArray(value) OVER (ORDER BY order ROWS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS rows_frame, groupArray(value) OVER (ORDER BY order RANGE BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS range_frame, groupArray(value) OVER (ORDER BY order GROUPS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS groups_frame FROM wf_frame_groups ORDER BY order, value; ``` ```response ┌─order─┬─value─┬─rows_frame─┬─range_frame─┬─groups_frame─┐ │ 10 │ 1 │ [1,2] │ [1,2] │ [1,2,3] │ │ 10 │ 2 │ [1,2,3] │ [1,2] │ [1,2,3] │ │ 20 │ 3 │ [2,3,4] │ [3] │ [1,2,3,4,5] │ │ 30 │ 4 │ [3,4,5] │ [4,5] │ [3,4,5] │ │ 30 │ 5 │ [4,5] │ [4,5] │ [3,4,5] │ └───────┴───────┴────────────┴─────────────┴──────────────┘ ``` Each mode interprets the bounds differently: - `ROWS` counts physical rows, so the frame is at most three adjacent rows: the current row plus one on each side. - `RANGE` counts `order` values, so `1 PRECEDING` and `1 FOLLOWING` cover rows whose `order` is within 1 of the current row's. With gaps of 10, no neighbouring row qualifies, so the frame holds only the rows that share the current `order`. - `GROUPS` (added by this PR) counts peer groups, so `1 PRECEDING` and `1 FOLLOWING` always include the adjacent groups in full, whatever the gaps between `order` values. cc: @cwurm ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `GROUPS` frame mode for window functions (SQL:2011), e.g. `any(price) OVER (PARTITION BY symbol ORDER BY ts GROUPS BETWEEN CURRENT ROW AND 1 FOLLOWING)`. In a `GROUPS` frame the boundaries count whole peer groups — sets of rows that are equal on the `ORDER BY` key — so `N PRECEDING`/`N FOLLOWING` mean `N` peer groups before/after the current row's peer group, rather than physical rows (`ROWS`) or `ORDER BY` value distances (`RANGE`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108653",
          "createdAt": "2026-06-26T20:45:31Z",
          "updatedAt": "2026-08-13T12:40:57Z",
          "timestamp": "2026-08-13T12:40:57Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "nihalzp",
          "state": "open",
          "assignees": [
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8e2781a587b647b13eeb",
        "signalId": "github:ClickHouse/ClickHouse:issue:111838",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:111838",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "`fsync_part_directory = 1` on an `encrypted` disk fails every `INSERT` with `FILE_DOESNT_EXIST`: the INSERT-path directory sync guard double-wraps the absolute part path",
          "text": "### Describe what's wrong Enabling `fsync_part_directory = 1` on a `MergeTree` table stored on an `encrypted` disk (with a non-empty `path`, as in the documented configuration) makes **every `INSERT` fail** with `FILE_DOESNT_EXIST` (Code 107). The durability hardening setting is therefore unusable on encrypted disks — users who want crash-safe part commits on encrypted storage have to turn it off. Root cause is a path double-wrap on the INSERT path only: - `MergeTreeDataWriter::writeTempPartImpl` requests the directory sync guard with the **absolute** part path (`new_data_part->getDataPartStorage().getFullPath()`, `src/Storages/MergeTree/MergeTreeDataWriter.cpp:855-858`). - On an encrypted disk this reaches `DiskEncryptedTransaction::wrappedPath` (`src/Disks/DiskEncryptedTransaction.h:36-42`), which unconditionally prepends `disk_path` again. With the docs' own layout (`<path>encrypted/</path>` over a local disk at `/disk/`) the guard tries to open `/disk/encrypted//disk/encrypted/store/<uuid>/tmp_insert_all_1_1_0/` → `ENOENT` → `LocalDirectorySyncGuard` constructor throws `FILE_DOESNT_EXIST` (`src/Disks/LocalDirectorySyncGuard.cpp:31-37`) → the INSERT fails. - The merge/mutate/fetch path is unaffected because `DataPartStorageOnDiskBase::getDirectorySyncGuard` passes the **relative** `root_path/part_dir` (`src/Storages/MergeTree/DataPartStorageOnDiskBase.cpp:956-958`) — so `OPTIMIZE` on the same table works while `INSERT` throws. - Plain `DiskLocal` (and an encrypted disk with an empty `path`) only work by accident: `std::filesystem::path::operator/` with an absolute right-hand side discards the left-hand side. ### Does it reproduce on the most recent release? Yes — master (26.7.1.1380) and 26.3.2.3 (longstanding). ### How to reproduce * Which ClickHouse server version to use: master; also reproduced on 26.3.2.3. Storage configuration (as in `docs/en/operations/storing-data.md`, encrypted disk with non-empty `path`): ```xml <storage_configuration> <disks> <local_disk><type>local</type><path>/data/local_disk/</path></local_disk> <enc> <type>encrypted</type> <disk>local_disk</disk> <path>enc/</path> <key>0123456789abcdef</key> </enc> </disks> <policies> <enc_policy><volumes><main><disk>enc</disk></main></volumes></enc_policy> </policies> </storage_configuration> ``` ```sql CREATE TABLE t_bug (x UInt64) ENGINE = MergeTree ORDER BY x SETTINGS storage_policy = 'enc_policy', fsync_part_directory = 1; INSERT INTO t_bug VALUES (1); -- Code: 107, FILE_DOESNT_EXIST — every time ``` Controls (all pass): - same table with `fsync_part_directory = 0` → INSERT OK; - same `fsync_part_directory = 1` on the underlying plain `local` disk → INSERT OK (and issues the directory `fdatasync`); - merge path on the encrypted disk: insert with the setting off, `ALTER TABLE ... MODIFY SETTING fsync_part_directory = 1`, `OPTIMIZE TABLE ... FINAL` → OK (relative-path caller). 3/3 on master, 1/1 on 26.3.2.3. ### Expected behavior `INSERT` with `fsync_part_directory = 1` on an encrypted disk succeeds and fsyncs the part directory of the underlying storage, same as on a plain local disk. The INSERT-path caller should pass the disk-relative part path (as the merge path does), or `wrappedPath` should not double-prepend an already-absolute path. ### Error message and/or stacktrace ``` Received exception from server (version 26.7.1): Code: 107. DB::Exception: Received from localhost:9000. DB::Exception: Cannot open file /data/local_disk/enc//data/local_disk/enc/store/b6a/b6a09e61-c268-4145-a690-ae9ae563557e/tmp_insert_all_1_1_0/: , errno: 2, strerror: No such file or directory: While executing WaitForAsyncInsert. (FILE_DOESNT_EXIST) ``` Note the double-wrapped path: `/data/local_disk/enc/` + `/data/local_disk/enc/store/...`. ### Additional context - File-data fsync on encrypted disks is intact (`WriteBufferFromEncryptedFile::sync` forwards down to `fdatasync`, `src/IO/WriteBufferFromEncryptedFile.cpp:44-50`) — only the directory-sync guard on the INSERT path is broken, and loudly. - No integration test covers fsync settings on encrypted disks (`tests/integration/test_encrypted_disk*` contain no `fsync` mentions). - Found by a crash-durability testing framework while auditing whether `DiskEncrypted` drops fsync (it does not — but the directory-fsync knob turns out to be unusable).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/111838",
          "createdAt": "2026-07-24T17:55:12Z",
          "updatedAt": "2026-08-13T12:40:00Z",
          "timestamp": "2026-08-13T12:40:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "bug",
            "comp-mergetree",
            "comp-disk-abstractions"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "CheSema"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:c2de6028e3ac35512ba4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114155",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114155",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Text index: fix out-of-bounds write in front-coding deserialization",
          "text": "Fixes https://github.com/ClickHouse/clickhouse-private/issues/65753. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reject corrupted/malicious dictionary where `lcp` exceeds the previous token or `lcp + data_size` overflows, before the buffer write. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1320` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114155",
          "createdAt": "2026-08-10T13:09:24Z",
          "updatedAt": "2026-08-13T12:39:10Z",
          "timestamp": "2026-08-13T12:39:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud"
          ],
          "author": "ahmadov",
          "state": "closed",
          "assignees": [
            "CurtizJ"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7acc21c692b424a6175d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114510",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114510",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: Improve supported regions page layout",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Improves the [supported cloud regions page](https://clickhouse.com/docs/products/cloud/reference/supported-regions) by grouping regions into provider tabs and presenting Private Region, HIPAA, and PCI availability as comparison flags. This makes the information easier to scan without duplicating each provider name or maintaining separate compliance inventories. Linear issue: [DOC-967](https://linear.app/clickhouse/issue/DOC-967/improve-supported-regions-page-layout) ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved the layout of the supported cloud regions reference page. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1321` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114510",
          "createdAt": "2026-08-12T15:51:33Z",
          "updatedAt": "2026-08-13T12:39:06Z",
          "timestamp": "2026-08-13T12:39:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation",
            "pr-synced-to-cloud"
          ],
          "author": "dhtclk",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:667ec73d2cb4115afc34",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105045",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105045",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support column matcher expansion for default value expressions and index expressions",
          "text": "This PR closes https://github.com/ClickHouse/ClickHouse/issues/92266 Support column matchers in column `DEFAULT`, `ALIAS`, `MATERIALIZED`, and `EPHEMERAL` expressions, and in data skipping index expressions. This allows expressions such as `*`, `COLUMNS('...')`, `COLUMNS(a, b)`, `EXCEPT`, `APPLY`, and `REPLACE` to be expanded before expression validation and execution. ~~The change also adds `namedTuple` function to make matcher-expanded named tuple expressions ergonomic in tests and user queries.~~ The tests cover these use cases, direct and indirect cyclic default-expression dependency detection, and nested matcher expansion. ### Changelog category: - New Feature ### Changelog entry: - Support column matchers such as `*` and `COLUMNS` in column default value expressions, `DEFAULT`, `ALIAS`, `MATERIALIZED`, and `EPHEMERAL` expressions, and in data skipping index expressions. ### Note ~~About the newly added `namedTuple` function: I read the previous discussions in [1], [2], and [3], and my impression is that the existing `enable_named_columns_in_function_tuple` setting is not very discoverable. A separate function name may make the intention clearer and the feature easier to use, especially in this use case. It also avoids changing the behavior of the existing tuple function, so it should not introduce compatibility issues. Please let me know if this direction is not desirable.~~ [1] https://github.com/ClickHouse/ClickHouse/issues/63524 [2] https://github.com/ClickHouse/ClickHouse/issues/54921 [3] https://github.com/ClickHouse/ClickHouse/pull/54881 <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches core expression/DDL validation paths (defaults, aliases, ALTER, mutations, skip indexes), which can affect query planning and error behavior if matcher expansion or cycle detection is incorrect. > > **Overview** > Enables column matchers (e.g. `*`, `COLUMNS(...)` plus `EXCEPT`/`APPLY`/`REPLACE`) inside column `DEFAULT`/`MATERIALIZED`/`ALIAS`/`EPHEMERAL` expressions by expanding matchers before validation/execution, honoring `asterisk_include_*` settings and rejecting qualified matchers. > > Updates DDL/default validation, alias expansion, read-order optimization, merge/SELECT paths, and mutation/materialize flows to use a shared `cloneAndExpandColumnDefaultExpression` helper and adds cycle detection that accounts for matcher-expanded dependencies. > > Extends skip index parsing/analysis to normalize matcher and alias usage (including cyclic-alias detection) before `TreeRewriter` analysis, and adds docs + comprehensive stateless tests covering matcher expansion, errors, and ALTER/mutation/index scenarios. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 48ab80fd074599c63f967f624435e0a2e7f166ed. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105045",
          "createdAt": "2026-05-15T15:10:34Z",
          "updatedAt": "2026-08-13T12:39:00Z",
          "timestamp": "2026-08-13T12:39:00Z",
          "metrics": {
            "reactions": 3,
            "comments": 49
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested"
          ],
          "author": "niyue",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f27aac21b67775d3f090",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110883",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110883",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Disable TopK dynamic filtering when a sorting projection makes the read in-order",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/110862 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a performance regression where `use_top_k_dynamic_filtering` installed a redundant `__topKFilter` prewhere for `ORDER BY ... LIMIT` queries served in-order by a sorting projection, causing the sort column to be read twice. ### Description Closes: #110862 `optimizeTopK` disables `use_top_k_dynamic_filtering` when the `ORDER BY` column is a prefix of the read's sorted order: there the dynamic prewhere filter is counterproductive, because once the running threshold stabilizes it rejects all remaining rows in sorted order, defeating the early pipeline cancellation that `LIMIT` relies on and forcing a full scan. That guard only checked the base table's sorting key. When a sorting projection whose `ORDER BY` differs from the base table is selected, the read is `ReadType: InOrder` with respect to the projection's sort key, but the guard never matched, so the redundant `__topKFilter` prewhere was installed on top of the projection read, re-reading the sort column that in-order reading already provided. `tryOptimizeTopK` runs in the first plan pass, before projection selection and read-in-order (both second pass), so the guard is predictive. It now also disables dynamic filtering when a normal sorting projection whose `ORDER BY` starts with the sort column and which stores every read column is available and projection optimization is enabled — the exact condition under which such a projection is later chosen to serve the read in-order. Results were correct in all cases; this is a performance-only fix. Verified with `EXPLAIN` that the redundant `__topKFilter` is no longer installed while the projection is still selected and the read stays `InOrder`, and that dynamic filtering is still applied when no projection can serve the order.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110883",
          "createdAt": "2026-07-17T15:32:49Z",
          "updatedAt": "2026-08-13T12:38:11Z",
          "timestamp": "2026-08-13T12:38:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.4-must-backport"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:5084b70d0ae1dae50f08",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114625",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114625",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix the bit-sliced full adder in `groupNumericIndexedVector`",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/106208 Related: https://github.com/ClickHouse/ClickHouse/pull/110072 `addValue` set a bit when the computed sum bit was 1 but never cleared it when the sum bit was 0, so every carry left the lower bit set and the write path was not addition: ```sql SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), val)) FROM (SELECT arrayJoin([toInt64(10), toInt64(10)]) AS val); -- {5:30}, expected {5:20} SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), toInt64(1))) FROM numbers(8); -- {5:255}, expected {5:8} ``` Any repeated index whose addends share a set bit is affected, on every index type and in both the small and the promoted representation. `n` rows of value `1` at one index accumulate to `2^n - 1`. Additions whose bits are disjoint need no carry and were already correct (`10 + 5` gives `15`), which is why the existing tests and the documented examples — all of which use distinct indexes — did not catch it. `merge` and `numericIndexedVectorPointwiseAdd` share `pointwiseAddInplace`, which computes whole-bitmap XORs and assigns the result, so clearing is implicit there and those paths were already correct. Only the per-row path was wrong, which is why `numericIndexedVectorAllValueSum` disagreed with `sum(value)` over the same rows. `RoaringBitmapWithSmallSet` had no way to clear an element, so this adds a `remove`. `SmallSet` has no erase and the small set holds at most `small_set_size` elements, so that path rebuilds it without the removed value. `zero_indexes` is now maintained too. It holds the present indexes whose value is zero, so it has to gain an index when an update drives the value to zero and lose it when the value becomes non-zero — the same invariant the pointwise operations restore when they finish: ```cpp /// For any of the total_indexes, if it is not in the non-zero index of the result, the result is 0. total_indexes->rb_andnot(*getAllNonZeroIndex()); zero_indexes = total_indexes; ``` With that, adding `5` and then `-5` row by row produces `{5:0}`, matching what merging the two values already produced, and `numericIndexedVectorGetValue` and `numericIndexedVectorCardinality` agree with the map. `numericIndexedVectorBuild` also goes through `addValue`, but from a map whose keys are unique, so it never carried and is unaffected. The test asserts `numericIndexedVectorAllValueSum` equals `sum(value)` over repeated indexes with negative and fractional values, that the row-by-row path agrees with the pointwise path, and that an index driven to zero is present with value zero. Eight of its eleven assertions fail without the fix; the three that pass are controls — an addition with disjoint bits, the pointwise reference path, and adding zero to an index that already holds a value. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `groupNumericIndexedVector` producing wrong values when the same index appears in more than one row. The bit-sliced adder never cleared a bit when the computed sum bit was zero, so every carry left the lower bit set: eight rows of value `1` at one index accumulated to `255` instead of `8`, and `10 + 10` produced `30` instead of `20`. Values that share no set bits were unaffected. An index whose value is driven to zero is now reported as present with value zero, consistent with `numericIndexedVectorPointwiseAdd` and with merging aggregate states. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114625",
          "createdAt": "2026-08-13T12:12:29Z",
          "updatedAt": "2026-08-13T12:37:10Z",
          "timestamp": "2026-08-13T12:37:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1a5c5fa26811e242abce",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112805",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112805",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not drop a named collection that a detached table still uses",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/96181 Related: https://github.com/ClickHouse/ClickHouse/issues/77366 Related: https://github.com/ClickHouse/ClickHouse/pull/110529 A table detached with a plain `DETACH TABLE` keeps its metadata file, so the server attaches it again on the next start. It is gone from `DatabaseCatalog` though, so `isTableExist` returns false for it, and the `check_named_collection_dependencies` check (added in #96181) treated its dependency as a stale leftover of a failed `CREATE TABLE`: it removed the dependency and let `DROP NAMED COLLECTION` succeed. The `ATTACH` replayed at startup then threw `NAMED_COLLECTION_DOESNT_EXIST`, which aborts loading the metadata, and the server did not start at all. This is how the `Stress test (arm_tsan)` job fails on master with `Cannot start clickhouse-server`: the AST fuzzer makes `04320_url_engine_dispatch_partition_and_format` leave its `URL(named_collection)` table detached (the test's `ATTACH TABLE` never reaches the server), and the test then drops the named collection. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=8aad759007771032aa94d7f4eee3e18103dd63c3&name_0=MasterCI&name_1=Stress%20test%20%28arm_tsan%29 ``` Application: Caught exception while loading metadata: Code: 722. DB::Exception: Waited job failed: Code: 695. DB::Exception: Load job 'load table test_3.04320_..._n' failed: Code: 669. DB::Exception: There is no named collection `04320_..._nc`: Cannot attach table `test_3`.`04320_..._n` from metadata file store/e5b/.../04320_..._n.sql from query ATTACH TABLE ... ENGINE = URL(`04320_..._nc`, format = 'JSON'). (NAMED_COLLECTION_DOESNT_EXIST) ``` The same signature accounts for 6 of the 8 `Cannot start clickhouse-server` failures with a missing named collection in the last 60 days (per `play.clickhouse.com`); the other two come from `03822_named_collection_drop_dependency_check`, which drops the collection with `check_named_collection_dependencies = 0` on purpose, and are the case #110529 handles by tolerating the missing collection at startup. ### Changes Per review feedback, the implementation is a simple in-memory bookkeeping in `NamedCollectionFactory` (an earlier revision inspected the metadata the detached table would be attached from, which required probing database disks and sweeping metadata directories after renames): - `DETACH TABLE` (and `DETACH DATABASE`, for every table inside) moves the dependencies of the table into a list of (collection, database, table) entries. - `DROP NAMED COLLECTION` is refused with `NAMED_COLLECTION_IS_USED` while an entry for the collection exists. - `ATTACH` does not remove the entry by itself: the dependencies are registered while the engine arguments are resolved, and the attach can still fail after that (an unknown format name, a failure creating the storage), leaving the table detached — the entry must keep protecting it. Instead, the entry is removed by the events that prove the metadata under that name is gone or harmless: `DROP TABLE`, `DETACH TABLE ... PERMANENTLY`, and `RENAME` of the (necessarily re-attached) table; `DROP DATABASE` removes the entries of the database's detached tables, and `RENAME DATABASE` re-keys them. The `DROP NAMED COLLECTION` check itself removes nothing: the table's existence in `DatabaseCatalog` is racy against in-flight attaches and detaches (the table can exist while nothing in the drop query has validated its live dependency), so the drop path is read-only and every recorded entry refuses the drop. - `DETACH TABLE ... PERMANENTLY` does not record an entry: a permanently detached table is not loaded at startup, so dropping a collection it references cannot break the server start. A later explicit `ATTACH` of such a table fails cleanly with `NAMED_COLLECTION_DOESNT_EXIST` and is recoverable by recreating the collection. - The list lives in memory only, which is consistent across a restart: a plainly detached table is attached again at the next start, where regular dependency tracking picks it up, and a permanently detached one records no entry at all. The list is deliberately imprecise in one direction: a stale entry may keep refusing the drop for a while after the detached table itself is gone, or after the table was attached back (until the table is dropped or renamed). In exchange, the `DROP NAMED COLLECTION` path performs no disk access at all. - A `RENAME TABLE` that moves a table between an `Ordinary` and an `Atomic` database changes the identity the dependency is keyed by: the move into `Atomic` assigns a fresh UUID to the table and the move out of it drops the UUID, while a dependency is keyed by the UUID for tables of `Atomic` databases and by the name for tables of `Ordinary` ones. The rename interpreter only knows the names, so the entry used to keep the identity the table had before the move and nothing found it afterwards: the detach recorded no entry and, for the `Atomic -> Ordinary` direction, the drop check even classified the entry of the still attached table as a leftover of a failed `CREATE` and dropped the collection from under it. `DatabaseOnDisk::renameTable` now re-keys the entries of the moved table to its new `StorageID`, where both identities are known. `EXCHANGE` is unaffected: it is only supported between two `Atomic` databases, where the UUIDs do not change. - A table of a database with `lazy_load_tables = 1` is attached as a `StorageTableProxy` and its real storage is built only on the first access, so the engine arguments are not resolved at load time and the dependency on the named collection they name stayed unregistered. `DROP NAMED COLLECTION` was then allowed while such a table still referenced the collection - breaking it at the first access with `NAMED_COLLECTION_DOESNT_EXIST` - and a `DETACH` of it had no dependency to move to the list of the detached ones, so the protection above silently disappeared in that mode. `DatabaseOrdinary::loadTableLazy` now registers the dependency straight from the metadata, via the new `tryGetUsedNamedCollectionName` helper. An identifier first argument counts as a collection reference only for the engines that resolve their arguments through named collections (a new `StorageFactory::StorageFeatures::supports_named_collections` flag) - for other engines an identifier means something else, e.g. a cluster name for `Distributed` - and for them the signal is time-stable: whether the collection currently exists is deliberately not checked, so the dependency of a collection that is missing at load time (say, after a drop with `check_named_collection_dependencies = 0`) protects it when it is recreated later. The exception is `Remote`/`RemoteSecure`, where the same identifier is also a valid positional argument - a cluster name - when the named-collection lookup does not resolve (marked by the new `StorageFeatures::named_collection_argument_is_ambiguous`; every other flagged engine treats an unknown collection as an error, not as a fallback to a positional form). Syntax alone cannot prove that such a table uses a collection, so the helper replicates the decision the engine's own argument parsing would make at the same moment - which is exactly what a non-lazy load of the same metadata does: the named-collection branch is taken only when a collection with that name exists, and only `key = value` overrides may follow the collection name, which a positional argument list never looks like. `MongoDB` and `MaterializedPostgreSQL` were the only flagged engines whose eager argument resolution did not pass the dependent table to `tryGetNamedCollectionWithOverrides` (they registered the dependency at the lazy load but not when the storage is built); they now pass it, so both load modes register the same dependency. `addDependency` ignores an exact duplicate, because the same dependency is registered again when the proxy is materialized. Tables created with `CREATE TABLE ... AS f(...)` need no lazy branch: a database with `lazy_load_tables = 1` deliberately loads them eagerly as a `StorageTableFunctionProxy` (see `DatabaseOrdinary::shouldLazyLoad`), and that load registers the dependency via `ITableFunction::getUsedNamedCollectionName`. - The cleanup of stale *active* dependencies (leftovers of a failed `CREATE TABLE`, pre-existing from #96181) no longer treats the table's absence from `DatabaseCatalog` alone as a proof of staleness: the dependency of an in-flight `CREATE`/`ATTACH` is registered while the engine arguments are resolved, before the table is committed to the catalog, and a concurrent `DROP NAMED COLLECTION` could prune it and drop the collection while the create later succeeds — recreating the broken metadata this PR fixes. The creating query holds the `DDLGuard` of the table name for the whole window between the registration and the commit, so the drop re-checks the table's existence under that guard before pruning: once the guard is acquired, no create is in flight, and the table's absence proves the entry is stale. (Entries with an empty database name come from dictionaries defined in the configuration files, which are not created through DDL; they are pruned as before.) The pruning removes only the exact stale entry (the collection and the recorded database, table and UUID): `CREATE TABLE ... UUID` can reuse the UUID of a failed create under a different table name, which the guard of the recorded name does not synchronize with, and removing everything under the UUID would erase the live dependency of such an in-flight create — the collection it uses could then be dropped from under the committed table. A new `create_table_pause_before_commit` failpoint keeps a create inside the window for the tests. ### Documented behavior impact `DROP NAMED COLLECTION` now rejects a collection that a detached table or a table in a detached database references, where it previously succeeded (and left a server that could not start). This is what `check_named_collection_dependencies` already promises - \"Check that DROP NAMED COLLECTION will not break tables that depend on it\" - so the documented behavior of the setting does not change, and setting it to `0` still allows the drop. No documentation update is needed. ### Verified locally (release build) - Before: `CREATE NAMED COLLECTION` + `CREATE TABLE ... ENGINE = URL(nc)` + `DETACH TABLE` + `DROP NAMED COLLECTION` succeeds, and the server then fails to start with the exact error chain above (exit code 210). Same with `DETACH DATABASE`. - After: the drop is refused, the collection stays, the table attaches back, and a restart of a server with the detached table present succeeds. - `04660_drop_named_collection_detached_table`, `04698_drop_named_collection_detached_after_rename`, and `04823_drop_named_collection_broken_attach` (all new, covering plain `DETACH TABLE` (blocks the drop), `DETACH TABLE ... PERMANENTLY` (does not block; the later `ATTACH` fails cleanly), `DETACH DATABASE`, `Ordinary` databases, renames of the table and of the database before and after the detach, stale dependencies of failed `CREATE TABLE`, and an `ATTACH` that fails after the dependencies were registered — the drop stays refused), and `04836_drop_named_collection_inflight_create` (new, runs `DROP NAMED COLLECTION` against a `CREATE TABLE` and an `ATTACH TABLE` paused between the dependency registration and the commit to the catalog: the drop blocks on the `DDLGuard` and is refused), and `04840_drop_named_collection_cross_engine_rename` (new, moves a table between an `Ordinary` and an `Atomic` database in both directions and then detaches the table and its database), and `04848_drop_named_collection_reused_uuid` (new, prunes the stale entry of a failed `CREATE TABLE ... UUID` while a create of a different table reusing the UUID is paused inside the window: the drop of the old collection succeeds, and the drop of the collection the new table uses stays refused; verified to fail without the fix), plus `03822_named_collection_drop_dependency_check` and `04003_named_collection_drop_dependency_check_dict` pass. - `test_named_collections/test.py::test_drop_while_used_by_lazily_loaded_table` (new integration test: a table using a named collection in an `Atomic` database with `lazy_load_tables = 1`, a server restart so the table comes back as a never-accessed proxy, and the drop refused both while it is attached and after `DETACH TABLE`). It is an integration test because the hole is only reachable once the in-memory list is empty, i.e. after a restart: without one, the `DETACH DATABASE` that precedes `ATTACH DATABASE` leaves its own entry behind and that entry refuses the drop on its own. Verified that it fails without the fix (the drop succeeds) and passes with it. - `test_drop_collection_recreated_under_lazily_loaded_table` (new integration test: the collection is dropped with `check_named_collection_dependencies = 0`, the server restarts while it is missing, and the drop of the recreated collection is refused; verified to fail without the fix), and `test_drop_not_used_by_lazily_loaded_distributed_table` (new integration test: a collection named after the cluster of a lazily loaded `Distributed` table is droppable; passes before and after, pinning the behavior). - `test_drop_not_used_by_lazily_loaded_remote_table` (new integration test: a collection named after the cluster of a lazily loaded `ENGINE = Remote(cluster, system, one)` table, created after the table, is droppable after a restart; verified to fail without the fix - the drop was refused with `NAMED_COLLECTION_IS_USED` by the unrelated table), and `test_drop_while_used_by_lazily_loaded_table_function` (new integration test pinning that a `CREATE TABLE ... AS bigquery(collection)` table in a database with `lazy_load_tables = 1` keeps blocking the drop after a restart, before the first access and after a `DETACH TABLE`: such tables are loaded eagerly as a `StorageTableFunctionProxy`, which re-registers the dependency; passes without any code change, confirming no lazy table-function branch is needed). ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the server failing to start after a named collection was dropped while a detached table (or a table in a detached database) still referenced it. `DROP NAMED COLLECTION` now counts detached tables as dependents and is refused with `NAMED_COLLECTION_IS_USED`, as `check_named_collection_dependencies` already does for attached tables.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112805",
          "createdAt": "2026-07-31T19:56:49Z",
          "updatedAt": "2026-08-13T12:35:37Z",
          "timestamp": "2026-08-13T12:35:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:84a85690dfdf89364ff7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104591",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104591",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add optimize_row_order_if_no_order_by (reopen #103919)",
          "text": "Reopen of https://github.com/ClickHouse/ClickHouse/pull/103919. This closes #103839. Adds a new `MergeTree` setting `optimize_row_order_if_no_order_by` (default `1`) that enables `optimize_row_order` automatically for tables without an explicit `ORDER BY`. **Behavior change and migration path.** With an empty sorting key (`ORDER BY ()` / `ORDER BY tuple()`) no query can rely on the physical row order, so the rows of every inserted block are reordered to improve compressibility. This makes such inserts slower (a 5M-row insert benchmark shows roughly `+180%`..`+230%` on the insert itself) in exchange for a smaller on-disk size and faster filters on low-cardinality columns; the CI `clickbench` and `tpch_adapted` runs report no significant change. Existing tables are affected on upgrade. To keep the old behavior: - per table: `SETTINGS optimize_row_order_if_no_order_by = 0` (or an explicit `optimize_row_order = 0`, which also opts out); - server-wide: set it in the `merge_tree` config section; - by version: it is recorded in the `MergeTree` settings changes history, so `compatibility` set to a version before `26.8` keeps it off. Note that `MergeTree` setting defaults are resolved once, when the server materializes its global `MergeTreeSettings`, so `compatibility` has to come from the default profile (`users.xml`) - a `SET compatibility` in an already-running session does not change them. The `MergeTree` documentation is updated accordingly. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `optimize_row_order_if_no_order_by` `MergeTree` setting. When enabled (default), row order optimization is applied to inserts into tables without an explicit `ORDER BY` clause. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes default insert-time behavior for `MergeTree` tables with `ORDER BY ()`, which can affect CPU cost and on-disk row ordering/compression. Also alters row-order optimization internals to swallow some `NOT_IMPLEMENTED` errors, which could mask type-specific issues if incorrect. > > **Overview** > Introduces a new `MergeTree` setting `optimize_row_order_if_no_order_by` (default **on**) and wires it into `MergeTreeDataWriter` so row-order optimization is automatically applied on insert for tables that *lack a sorting key* (`ORDER BY ()`), while preserving existing behavior for tables with an explicit `ORDER BY`. > > Hardens `RowOrderOptimizer` by catching `NOT_IMPLEMENTED` during cardinality estimation and falling back to an all-distinct upper bound instead of failing inserts, and records the setting in settings-change history. > > Updates many stateless/perf tests to explicitly disable the new default (`SETTINGS optimize_row_order_if_no_order_by = 0`) to keep deterministic baselines, adds new coverage for the setting’s default/override behavior and the regression on unsupported nested types, and extends `tests/performance/scripts/perf.py` to strip this setting from `CREATE TABLE` when running against older servers that don’t recognize it. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 02d68c628d9a3625824d33b91f87bd12d6fd3938. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104591",
          "createdAt": "2026-05-11T14:01:19Z",
          "updatedAt": "2026-08-13T12:33:15Z",
          "timestamp": "2026-08-13T12:33:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 44
          },
          "labels": [
            "pr-improvement",
            "pr-performance",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:37691d8d7048a25ef06b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:42701",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:42701",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add table function `obfuscate`",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/39067 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add table function `obfuscate` which applies the same transformation as the `clickhouse-obfuscator` tool to the result of an arbitrary query, e.g. `SELECT * FROM obfuscate(SELECT * FROM table)`. The seed can be controlled with the `obfuscate_seed` setting; with an empty seed a fresh random one is derived per execution.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/42701",
          "createdAt": "2022-10-26T13:41:41Z",
          "updatedAt": "2026-08-13T12:32:43Z",
          "timestamp": "2026-08-13T12:32:43Z",
          "metrics": {
            "reactions": 2,
            "comments": 44
          },
          "labels": [
            "pr-feature"
          ],
          "author": "evillique",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:349a63f30a69392cf252",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113742",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113742",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `CANNOT_CONVERT_TYPE` and an exception in `GroupingAggregatedTransform` when reading a `Merge` table with one child being a `Distributed` table under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_sampling'` or `'custom_key_range'`). Closes [#113741](https://github.com/ClickHouse/ClickHouse/issues/113741). ### Documentation entry for user-facing changes The custom-key parallel replicas branch of the planner replaces the plan of a table expression with a remote read at the fixed stage `WithMergeableStateAfterAggregationAndLimit`, ignoring the stage the plan was requested up to. A `Merge` table over a `Distributed` child plans all of its children up to `WithMergeableState` through an interpreter, so a `MergeTree` child's plan produced finalized (post-aggregation, post-`LIMIT`) data where the parent `ReadFromMerge` expected partial aggregation states: `CANNOT_CONVERT_TYPE` for `count`, and the exception `Chunk should have AggregatedChunkInfo/ChunkInfoWithAllocatedBytes in GroupingAggregatedTransform` when the finalized type structurally coincides with the state type. Only the analyzer path is affected. The fix allows the custom-key read only when the requested stage is `Complete` or `WithMergeableStateAfterAggregationAndLimit` itself; a child planned to a partial stage now runs as a plain local read, as it does when parallel replicas are off. Found by the targeted AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/110972, where it is unrelated: the failure reproduces on master without `parallel_replicas_allow_merge_tables` (verified on a binary with that PR's changes swapped out to the merge base). Fuzzer report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110972&sha=ac0a584ce316ace31f4dbc16b38e1262a2344751&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29 (STID `3970-479a`). Closes: https://github.com/ClickHouse/ClickHouse/issues/113741 Related: https://github.com/ClickHouse/ClickHouse/pull/110972",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113742",
          "createdAt": "2026-08-06T23:22:57Z",
          "updatedAt": "2026-08-13T12:31:42Z",
          "timestamp": "2026-08-13T12:31:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:77e4b183424c08e15a8d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112973",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112973",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `sorted_merge` and `parallel_sorted_merge` join algorithms",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109005 Implements the two algorithms planned in [this discussion](https://github.com/ClickHouse/ClickHouse/pull/109005#discussion_r3694159098): `full_sorting_merge` and `parallel_full_sorting_merge` are always supported, so anything listed after them in `join_algorithm` is unreachable — listing them is an unconditional choice, not a preference. The new `sorted_merge` and `parallel_sorted_merge` algorithms execute the same merge join, but are **available only when both join inputs can be efficiently read in the order of the join keys** (e.g. MergeTree tables whose primary key starts with the join keys), so the pre-join sorts become cheap `FinishSorting` or disappear. When the tables' order cannot be exploited, the selection falls through to the next algorithm in the list. That makes them meaningful as a high-priority preference: `join_algorithm = 'sorted_merge,parallel_hash'` uses the streaming, low-memory merge join exactly when it is certainly beneficial, and a hash join otherwise. `sorted_merge` runs a single in-order merge join. `parallel_sorted_merge` additionally shards the join by ranges of the tables' common primary-key prefix into independent per-shard merge joins running in parallel — the same source-side sharding `query_plan_join_shard_by_pk_ranges` applies, enabled for this join by the algorithm itself; the in-order reads stay intact (no scatter, no re-sort). When the sharding cannot apply (e.g. an `ASOF` join), it degrades to a single `sorted_merge`. Implementation notes: - Eligibility is decided during plan physicalization, where the input subplans are visible: `JoinStepLogical::inputsCanBeReadInJoinKeyOrder` finds the `ReadFromMergeTree` below each input (mirroring `findReadingStep` of `optimizeReadInOrder`, without descending through nested joins) and probes the actual read-in-order matcher (`wouldReadInOrderBeUseful`, the same side-effect-free probe `topKThroughJoin` uses) with the join-key sort description. The predicted decision degrades gracefully in both directions: a false positive runs like `full_sorting_merge` (with a full sort), a false negative falls through to the next algorithm. - The same memoized predicate makes `tryAddJoinRuntimeFilter` keep its hands off an eligible join listed before the first hash-family algorithm. Without this, planting a runtime filter erases the merge algorithms from the list (a merge join reads both sides concurrently and cannot use a runtime filter), silently overriding the priority order — the defining feature of these algorithms. For non-eligible joins the runtime filter (and the erasure) stays, because those algorithms are not selectable anyway. When `applyParallelReplicas` later breaks the eligibility (a join input becomes a distributed read), the filter pass is re-run for exactly the joins it had skipped, so the `hash` fall-through gets its runtime filter back. - `FullSortingMergeJoin` now carries the selected algorithm (`getSelectedAlgorithm`) instead of an `is_parallel` flag; the hash-scatter rewrite (`optimizeParallelFullSortingMergeJoin`) stays exclusive to `parallel_full_sorting_merge`, and `optimizeJoinByShards` runs in a restricted mode (only `parallel_sorted_merge`-selected joins, with a cheap pre-scan bail-out) when `query_plan_join_shard_by_pk_ranges` is off. When the sharded stream counts diverge at pipeline-building time (e.g. a data-dependent `PREWHERE` prunes one side to a single empty stream), `JoinStep` merges each side's per-shard sorted streams back into one sorted stream and runs the single-stream merge join instead of failing (the same approach as #109393, applied at the sharding fallback). - The `CreateSetAndFilterOnTheFlyStep` pair (`max_rows_in_set_to_optimize_join`) is not added for sorted-merge joins: it sits between the read and the sort and would defeat the in-order read the algorithm was selected for. - The old analyzer has no query plan at selection time, so there the algorithms are never selected and the list falls through (documented). With only `sorted_merge` listed and no exploitable order, the query fails with `NOT_IMPLEMENTED`, like other unsupported single-algorithm configurations. - The known lower-priority-fallback side effects of listing merge algorithms (stricter `USING` key-type inference, `topKThroughJoin` deferral) extend to the new values and are documented in the `join_algorithm` setting description. Tests: `04669_sorted_merge_join_selection` pins the selection gating via `EXPLAIN PIPELINE` (selected on a primary-key join with no re-sort, falls through on non-key joins / disabled read-in-order / old analyzer, priority order respected, error when listed alone without exploitable order) and correctness against `hash` for `INNER`/`LEFT`/`RIGHT`/`FULL`/`ANY`/`join_use_nulls`. `04670_parallel_sorted_merge_join` pins the primary-key-range sharding (`Sharding:` marker with `query_plan_join_shard_by_pk_ranges = 0`, no `ScatterByPartitionTransform`, no `MergeSortingTransform`), the `ASOF` degradation, and correctness. `04824_sorted_merge_join_parallel_replicas_fallthrough` and `04894_sorted_merge_join_parallel_replicas_runtime_filter` pin the parallel-replicas edge (fall-through to `hash` with the runtime filter restored), `04893_parallel_sorted_merge_join_shard_stream_divergence` pins the diverged-shard degradation, and `04760`/`04850` pin the `join_use_nulls` and `query_plan_join_shard_by_pk_ranges` contracts. The PR #109005 regression tests and the join runtime filter tests pass unchanged. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added new `join_algorithm` values `sorted_merge` and `parallel_sorted_merge`: merge-join algorithms that are available only when both join inputs can be efficiently read in the order of the join keys (so the join benefits from the tables' order instead of sorting), and otherwise fall through to the next algorithm in the list. `parallel_sorted_merge` additionally shards the join by primary-key ranges into independent per-shard merge joins running in parallel. Listing them first, e.g. `join_algorithm = 'sorted_merge,parallel_hash'`, uses the streaming low-memory merge join exactly when it is certainly beneficial. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112973",
          "createdAt": "2026-08-02T03:03:03Z",
          "updatedAt": "2026-08-13T12:31:38Z",
          "timestamp": "2026-08-13T12:31:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2e75f24e71a128d475a3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112847",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112847",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make SQL SECURITY views an optimization barrier",
          "text": "### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A view with `SQL SECURITY DEFINER` or `SQL SECURITY NONE` is now an optimization barrier, so an expression in the query reading the view is never evaluated on rows that the view itself filters out. Previously such an expression could observe the hidden rows through an exception, and a view used to restrict which rows a user may see did not actually restrict them. --- ## The problem `SQL SECURITY DEFINER` is widely used to build a view that restricts which rows a user may see: ```sql CREATE VIEW user_query_log DEFINER = default SQL SECURITY DEFINER AS SELECT * FROM system.query_log WHERE user = currentUser(); GRANT SELECT ON user_query_log TO alice; ``` `alice` has no grant on the source table, only on the view. But the outer `WHERE` and the view's own `WHERE` are merged into a single filter over the source table, and nothing guarantees which of the two decides first. Any function that can signal through a side channel therefore observes the rows the view is supposed to hide: ```sql -- as alice, whose only grant is SELECT ON user_query_log SELECT event_time FROM user_query_log WHERE throwIf(query LIKE '%some probe%', 'DISCLOSED'); -- DB::Exception: DISCLOSED ... While executing MergeTreeSelect ``` That is a one-bit oracle per query. Exception messages carry the offending value, which turns it into a plain read of the hidden rows: ```sql SELECT * FROM v_secrets WHERE if(owner = 'alice', 1, toUInt8(secret)) = 1; -- DB::Exception: Cannot parse string 'TOP-SECRET-BOB-42' as UInt8 ``` Both work on `SQL SECURITY NONE` as well, and both work with `enable_analyzer = 0`. Row policies are *not* affected: a row policy is a separate `row_level_filter` that the reading step always applies before PREWHERE and before any pushed-down filter. This PR gives a view's own filtering the same standing. ## The fix `IQueryPlanStep` gets a `security_barrier` flag. After `StorageView::readImpl` builds the view's subplan, every step that can drop rows is marked, and optimizations that move a step down the plan refuse to cross a marked step: - `tryMergeExpressions`, `tryMergeFilters`, `tryPushDownFilter` and `tryMergeFilterIntoJoinCondition` refuse when the child is a barrier; - `optimizePrewhere` refuses to pull an outer filter into a barrier source — conditions are combined into the prewhere DAG with `and`, which gives no ordering guarantee — and transfers the barrier onto the source when it absorbs the view's own filter; - `trySplitFilter` moves the flag onto the new lower `FilterStep`, which is the one that still drops rows. - `tryPushDownLimit` refuses to move the invoker's `LimitStep` below a barrier step: once across the seal it would seed `DistinctStep::limit_hint` or a sorting limit inside the view's subplan, and `optimizeLimitForAggregationInOrder` walked through the seal to seed `AggregatingStep::limit_hint` the same way — hints that stop reading the source once enough visible rows are produced, so `read_rows`, progress and timing depended on the rows the view drops or collapses. Both walks, and `pushLimitByIntoSort`, fail closed on a barrier step. Blocking the steps that evaluate expressions is not enough on its own, because index analysis reaches the source by a different route and skips granules by the values of the rows the view hides — `read_rows` then tells the invoker whether such a row exists, with no exception needed. Eight walks are fenced as well: - `optimizePrimaryKeyConditionAndLimit` walks up from the reading step and hands every `FilterStep` it meets to the source's key condition; it now stops once it has consumed a barrier. The barrier's own condition still reaches the source — it is the definer's, and it is what decides which rows exist for the view — but nothing above it does. - `StorageView::readImpl` forwarded `query_info.filter_actions_dag` into the view's inner analyzer, where `Planner::collectFiltersForAnalysis` injects it into the inner plan and the filters it collects reach the inner tables' index analysis. It is no longer forwarded for a barrier view, which costs such a view over `Distributed` its shard skipping on the outer predicate. - `buildSortingDAG` in the read-in-order analysis descended through the view subplan and pulled outer `FilterStep` predicates into the fixed columns and the merged DAG, so an outer `ORDER BY` / `GROUP BY` / `DISTINCT` / `LIMIT BY` could still shape how the source below the barrier reads. It now reports a barrier in the chain and every consumer (sorting, aggregation, `DISTINCT`, `LIMIT BY`, the normal-projection choice, top-K, and the `Merge` child-plan walk) skips the analysis for that chain. Unlike the primary-key walk, the barrier step cannot be consumed here: the sort description belongs to the top of the chain, and a DAG missing the steps above the barrier could resolve a renamed column by name to the wrong source column and change the result order — so the analysis is skipped entirely, which is fail-closed and keeps correctness because the sorting step stays in the plan. - `tryOptimizeTopK` rewrites an `ORDER BY ... LIMIT` into a dynamic `__topKFilter` PREWHERE and minmax-skip-index granule pruning on the source, walking `LimitStep` → `SortingStep` → `ExpressionStep` → `FilterStep` → `ReadFromMergeTree` — and both rewrites are on by default. The walk now fails closed on the first barrier step it meets, including a reading step that carries the barrier after `optimizePrewhere` absorbed the view's own filter. - `tryTopKThroughJoin` peels the expression chain between the invoker's `Sorting` and a `Join` and grafts the invoker's `Sort + Limit` onto the join's preserved input, re-running the optimization passes on that subtree. For a barrier view whose inner query is a join, the graft landed below the seal — verified live pre-fix on both analyzers. The pass now fails closed on a marked `Limit`/`Sorting` (the pattern lies inside a view), a marked peeled expression (the seal), and a marked join (a join of a barrier view is always marked, not being row-preserving). Grafting above the seal of the invoker's own join input stays allowed: the inserted `Sort` consumes its whole input and the re-run passes are individually fenced. - `registerLeftSideIndexAnalysisSecondPass` of the join runtime filters walked from the `__applyFilter` step (which the fenced filter pushdown correctly keeps above the seal) down through every single-child expression or filter step — sealed or not — to the `ReadFromMergeTree` inside the view, and registered the invoker's build-side keys for granule pruning there. The walk now fails closed on the first barrier step. A live disclosure could not be constructed — the descriptor's key name has to match the read's namespace and the consumption path declined in every configuration tried — so this fence makes the contract structural rather than incidental. - The vector search rewrites — the vector-similarity-index pass and the quantized-codes shortlist — walk the same chain and prune the source to the top-N candidates of the invoker's `ORDER BY`. Both now fail closed on a barrier step the same way. - Projection planning: `QueryDAG::build` in `projectionsCommon.cpp` collects every filter of the chain below the aggregation — the invoker's predicates together with the view's own filtering — and both `optimizeUseAggregateProjections` (including `minmax_count_projection`) and `optimizeUseNormalProjections` prune parts and marks with it, so a projection-enabled table under a filtering barrier view still let the invoker's predicate shape the read. The DAG build now fails closed on the first barrier step, which makes the projection optimizations decline the read entirely. One more family reaches the source without going through its index analysis at all. `optimizeDistinctPerPartition`, `optimizeLimitByPerPartition` and `optimizeAggregationPerPartition` walk down through the sealing step and ask the reading to output each partition through a separate port, and `applyStreamDisjointness` carries the resulting partition disjointness back up across the seal, so the invoker's `DISTINCT` / `LIMIT BY` / `GROUP BY` skips its stream merging as well. Read scheduling, progress and resource consumption below the view then depend on how the rows the view drops are spread over the partitions, and all three `allow_*_partitions_independently` settings default to `1`. Both directions now fail closed on a barrier step. Four passes endanger the barrier by rebuilding steps rather than by walking past them, and each now fails closed on a barrier step: - lazy materialization splits every `Expression` / `Filter` step of the chain into a main and a lazy half, and the rebuilt steps do not carry the barrier flag, so the post-lazy `tryMergeExpressions` / `tryMergeFilters` passes saw an unmarked chain and could merge an invoker predicate into the view's own filtering — reopening the exception oracle itself, not just the read-shaping one; - `tryLiftUpUnion` rebuilds the `UnionStep` and clones the parent step into the branches as fresh unmarked steps, so a barrier view over `UNION ALL` lost its seal and `tryPushDownFilter` could then duplicate an invoker predicate into the branches; - `tryExecuteFunctionsAfterSorting` replaces the expression under a `SortingStep` with two new unmarked steps, which would strip the seal of a wrapper view and let the read-in-order and top-K walks descend through it again. - `tryLiftUpArrayJoin` splits the expression or filter above an `ArrayJoinStep`, moves one half below the `ARRAY JOIN`, and rebuilds both halves as fresh unmarked steps. When the parent is the seal of a view whose plan contains `ARRAY JOIN` (the seal is non-trivial whenever the view declares explicit column names or types), the invoker's predicate descended below the `ArrayJoinStep` and was evaluated on rows hidden by empty arrays — a live disclosure through the exception oracle, on both analyzers. Measured on a `DEFINER` view that exposes no row, over 100000 rows sorted by `key`, reading `WHERE key = <a key only a hidden row has>` against `WHERE key = <a key nothing has>`: 576 rows read against 0 without this, and 1000000 against 1000000 with it, on both `enable_analyzer = 1` and `enable_analyzer = 0`. `EXPLAIN SYNTAX` is left alone. It builds `InterpreterSelectQuery` with `only_analyze`, whose plan reads from `ReadNothingStep`, so no expression of the outer query is ever evaluated on a row and the view is still inlined for the diagnostic. Two paths substitute the view into the outer query before a plan exists, so a plan-level barrier cannot see them, and both are closed the same way — such a view is not inlined and is read through `StorageView::read`, which keeps the outer predicate in a step above it: - with `enable_analyzer = 0`, `InterpreterSelectQuery` replaces the view with a subquery and `TreeRewriter` merges the predicates; - with `analyzer_inline_views = 1`, `QueryAnalyzer::inlineViewSubqueryIfNeeded` does the same in the query tree. Without this, `SET enable_analyzer = 0` or `SET analyzer_inline_views = 1` would bypass the fix entirely. The flag also travels with a serialized query plan. A distributed worker deserializes a fragment and optimizes it again, so a barrier it does not know about is a barrier it will optimize across. `QueryPlan::serialize` writes the flag per step and fails closed when the negotiated query plan serialization version predates it (the version is bumped to 6), rather than sending a plan that silently loses its protection. ## Views that hide nothing are left alone Only a view that can actually drop rows becomes a barrier. A step is row-preserving if it is an `ExpressionStep`, a `SortingStep` without a limit, or a source step without PREWHERE; if none of the view's steps is anything else, nothing is marked and the plan is what it is today. The pre-plan paths make the same distinction. `StorageView::canHideRows` proves over the view's definition that the inner query preserves every row of a plainly readable source, and only a view for which that proof fails loses inlining and the forwarded outer filter. The proof fails closed: filters, limits, aggregation, `DISTINCT`, joins, `ARRAY JOIN`, `SAMPLE`/`FINAL`, multi-select unions, table functions, and a `FROM` that is itself a view or a view-wrapping engine (`Merge`, `Buffer`, anything remote) all count as able to hide rows. The proof classifies the storage that actually serves the read, not the object the name resolves to: proxy layers (a lazily loaded table of a database with `lazy_load_tables = 1`, or a table created from a table function) and `Alias` tables are unwrapped first, failing closed on a chain that cannot be resolved, and a storage that rewrites its own reads with `FINAL` and a `_sign` filter (`MaterializedPostgreSQL`) counts as able to hide rows even without another wrapper. So a projection-only `DEFINER` view produces exactly the plan of the same view declared `SQL SECURITY INVOKER` on every path, which the test pins byte-for-byte. Measured on 20M rows with `SELECT sum(length(payload)) FROM v WHERE tag = 'RARE'`, from `system.query_log`: | view | barrier | read_rows | read_bytes | ms | |---|---|---|---|---| | `DEFINER`, filters rows | on | 20 000 000 | 1.43 GiB | 35 | | `DEFINER`, filters rows | off | 1 638 400 | 82.84 MiB | 14 | | `DEFINER`, projection only | on | 1 638 400 | 82.84 MiB | 9 | | `DEFINER`, projection only | off | 1 638 400 | 82.84 MiB | 11 | A projection-only view is unaffected. A filtering view does pay: PREWHERE then holds only the view's own condition, so it no longer skips granules on the outer predicate. That is the inherent price of the guarantee — PostgreSQL's `security_barrier` views behave the same way — and it applies only to views that restrict rows, which are exactly the ones where it matters. The new server setting `sql_security_views_are_optimization_barriers` (default `1`) turns it off. It is deliberately a **server** setting and not a user setting: a user setting would be turned off by the very query that is trying to read the hidden rows. ## Testing `04758_sql_security_view_barrier_read_rows` covers the `read_rows` oracle through index analysis, on both analyzers, and prints `DISCLOSED` on both with the setting off. `04670_sql_security_view_barrier` covers the leak on `DEFINER` and on `NONE`, with `enable_analyzer = 1`, with `enable_analyzer = 0` and with `analyzer_inline_views = 1`, the value leak through a cast error message, the same oracle through a shard with `serialize_query_plan = 1`, that a projection-only `DEFINER` view and an `INVOKER` view still have the outer predicate merged into the view's own filter, and that a projection-only view keeps PREWHERE. `04813_sql_security_view_barrier_read_in_order` pins the read-in-order fence: the `INVOKER` twin and a projection-only `DEFINER` view read `InOrder`, a filtering `DEFINER` view does not, under both analyzers, with unchanged results. `04817_sql_security_view_barrier_top_k` pins the top-K fence: the `INVOKER` twin gets the `__topKFilter`, the filtering `DEFINER` view does not, and `read_rows` of an `ORDER BY ... LIMIT 1` over twin views is identical whether or not the hidden row holds the extreme minimum of the sort column that minmax pruning would rank first. `04818_sql_security_view_barrier_lazy_materialization` pins the lazy-materialization fence the same way and checks that an invoker predicate over the view is never evaluated on the hidden row. `04821_sql_security_view_barrier_projections` pins the projection fence: the `INVOKER` twin uses both a normal and an aggregate projection, the filtering `DEFINER` view uses neither, and `read_rows` of a predicate probe over twin views is identical whether or not the hidden row matches it. `04825_sql_security_view_barrier_union` pins that an outer predicate over a filtering `DEFINER` view on `UNION ALL` stays in a single filter above the union — before the `tryLiftUpUnion` fix it was duplicated into the branches — while the `INVOKER` twin keeps the pushdown. `04826_sql_security_view_barrier_functions_after_sorting` pins that an `ORDER BY ... LIMIT` over a wrapper `DEFINER` view (a `Merge` table over a nested filtering view) produces no in-order reading and no `__topKFilter` with `query_plan_execute_functions_after_sorting` on, while the `INVOKER` twin exploits the source order. `04827_sql_security_view_barrier_masked_wrappers` pins that the classification survives engine masking: a `DEFINER` view over a `Merge` wrapper behind a lazy `TableProxy` (re-masked before every round, since planning materializes the proxy) or behind an `Alias` table plans differently from its `INVOKER` twin on both analyzers and with `analyzer_inline_views = 1`. `04832_sql_security_view_barrier_limit_pushdown` pins the LIMIT fence: over a `DISTINCT` `DEFINER` view the invoker's `LimitStep` stays above the sealing step on both analyzers — before the fix it crossed the seal and sat directly on the `DistinctStep`, where it seeds the hint — and `read_rows` of an `ORDER BY ... LIMIT 1` over twin in-order `GROUP BY` views is identical whether the first group holds one raw row or almost all of them. `04837_sql_security_view_barrier_per_partition` pins the per-partition fence, with the `allow_*_partitions_independently` settings left at their defaults: an outer `DISTINCT` / `LIMIT BY` over a filtering `DEFINER` view produces none of the `Skip stream merging` / `Read each partition through separate port` markers its `INVOKER` twin gets, and the disjointness of a view whose own inner `DISTINCT` legitimately requests per-partition reading does not propagate across the seal into the invoker's `GROUP BY` / `LIMIT BY` — before the fix every `DEFINER` case was identical to its twin. `04840_sql_security_view_barrier_array_join` pins the `ARRAY JOIN` lift-up fence with an exception oracle on both analyzers: the `INVOKER` twin's predicate legitimately descends below the `ARRAY JOIN` and throws on the row an empty array hides, while the `DEFINER` view counts without throwing and its plan keeps every `throwIf` line above the `ArrayJoin` step — before the fix the `DEFINER` view threw as well. `04891_sql_security_view_barrier_top_k_through_join` pins the top-K-through-join fence: the `INVOKER` twin of a view over a `LEFT JOIN` gets the preserved-side `Sort + Limit` graft below the join, the `DEFINER` twin keeps its join input untouched — before the fix the `DEFINER` plan got the graft below the seal. `04892_sql_security_view_barrier_join_runtime_filter` pins the join-runtime-filter contract: with `enable_join_runtime_filters_index_analysis = 1`, twin filtering `DEFINER` views over tables identical except for the hidden row's primary-key value read exactly the same number of rows, while the `INVOKER` control is pruned by the build-side key. Every setting the plan shape depends on is pinned, because the test also runs with randomized settings. With `sql_security_views_are_optimization_barriers = 0` every one of those lines changes, so none of them passes vacuously. Ran 542 existing tests matching `view`, `prewhere`, `push_down`, `pushdown`, `row_policy`, `sql_security` and `definer`. 60 fail in my local environment, and the identical 60 fail with the barrier disabled on the same binary, so this introduces no regressions among them. ## Not covered here - The serialized-plan fence is defensive. With the serialization change reverted I could not make the barrier loss observable — neither the `throwIf` nor the failing-cast oracle leaks through `serialize_query_plan = 1`, `make_distributed_plan = 1`, or a `Distributed` table, because the initiator optimizes the fragment before shipping it. The flag is serialized so that the guarantee does not depend on that. - A separate hole remains: `additional_table_filters` and `additional_result_filter` from the invoker are copied verbatim into the definer's context by `StorageInMemoryMetadata::getSQLSecurityOverriddenContext`. Keyed on the view's *inner* table, the filter lands inside the view and its expression — which may contain scalar subqueries and table functions — is evaluated with the definer's privileges. That is a privilege escalation rather than a row disclosure, it is unaffected by this PR, and it needs its own fix. Until then it can be mitigated with a constraint on the definer's profile (`CREATE SETTINGS PROFILE p SETTINGS additional_table_filters = '' CONST TO <definer>`), which does not help for `SQL SECURITY NONE`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112847",
          "createdAt": "2026-08-01T02:18:25Z",
          "updatedAt": "2026-08-13T12:30:44Z",
          "timestamp": "2026-08-13T12:30:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-must-backport",
            "pr-critical-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:eb68e33ac2952e068470",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114627",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114627",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112217 to 26.7: Measure cancellation server-side in test_cancel_backup.py",
          "text": "Backport of https://github.com/ClickHouse/ClickHouse/pull/112217 to `26.7`. Related: https://github.com/ClickHouse/ClickHouse/pull/112217 Related: https://github.com/ClickHouse/ClickHouse/pull/114478 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Description `test_backup_restore_on_cluster/test_cancel_backup.py::test_cancel_backup` is flaky on the `26.7` branch with the same signature that #112217 fixed on `master`: the client-side `time_to_cancel` stopwatch measures the integration harness (several `docker exec` clickhouse-client launches plus `wait_status` poll quantum) rather than the cancellation itself, and the resulting body exception is masked in reports as `NoTrashChecker.__exit__` asserting `'QUERY_WAS_CANCELLED' in []`. It just failed this way on the backport PR https://github.com/ClickHouse/ClickHouse/pull/114478 (report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114478&sha=f5364950f96d60206363d2b2c1bff278a6976ee8&name_0=BackportPR&name_1=Integration%20tests%20%28amd_asan_ubsan%2C%20db%20disk%2C%20old%20analyzer%2C%201%2F6%29), and #112217 itself documents an occurrence on the `26.6` release branch, so release branches keep hitting it. This is a clean cherry-pick of the test-only fix (measure cancellation server-side via `system.backups` timestamps).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114627",
          "createdAt": "2026-08-13T12:29:22Z",
          "updatedAt": "2026-08-13T12:30:08Z",
          "timestamp": "2026-08-13T12:30:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:7497bc96d8d79e3ecfac",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114478",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114478",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113289 to 26.7: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114478",
          "createdAt": "2026-08-12T12:07:07Z",
          "updatedAt": "2026-08-13T12:29:31Z",
          "timestamp": "2026-08-13T12:29:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll3",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:f60bc1f35cf87378c569",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109602",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109602",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix the streaming-insert block wait not expiring under the query profiler",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109592 Related: https://jira.mariadb.org/browse/CONC-834 **Problem.** The file-descriptor poll timeout can silently never expire on a query thread. `ReadBufferFromFileDescriptor::poll` restarts the interrupted `poll` with the full original timeout after `EINTR`, resetting the deadline on every signal. Query threads receive periodic sampling-profiler timer signals (`query_profiler_real_time_period_ns`, default 1 s; the handler's `SA_RESTART` does not apply — per `signal(7)`, `poll` is never auto-restarted), so whenever the signal period does not exceed the timeout, the wait becomes unbounded while the fd stays silent. The user-visible path is the streaming-insert block wait: `IRowInputFormat::read` calls `poll` with the remaining `input_format_max_block_wait_ms` budget over a `StorageFile` fd/file or stdin buffer. With the profiler active, a stalled input source delays the partial-block flush indefinitely instead of flushing when the wait limit is reached. (All socket paths use `ReadBufferFromPocoSocket*::poll`, which is already deadline-aware, and are not affected.) **Root cause.** Same defect class as the MySQL `connect_timeout` fix in #109592 (mariadb-connector-c, upstream [CONC-834](https://jira.mariadb.org/browse/CONC-834)): an `EINTR` retry loop that passes the original timeout instead of the remainder. This was the last such site in `src/` — a sweep of all timed waits (`poll`/`epoll_wait`/`select`/`nanosleep`/`sigtimedwait`/io_uring/timerfd) found every other one deadline-aware (Poco sockets, `Epoll`, `KeeperTCPHandler`, `ShellCommandSource`, `base/sleep`). **Fix.** Re-poll with the remaining time computed **in microseconds from a single monotonic anchor**, reporting a timeout once the budget is exhausted. Per-retry whole-millisecond accounting (as in some existing call sites) would truncate a sub-millisecond retry interval to zero and make no progress under a sub-millisecond signal cadence, so the remainder is derived from the untouched anchor instead. The unit test (`gtest_fd_read_buffer_poll_under_signals.cpp`) waits on a pipe with no writer while a per-thread kernel timer (`timer_create` + `SIGEV_THREAD_ID`, the query profiler's own mechanism, so delivery deterministically targets the polling thread) interrupts the poll at two cadences: 10 ms (the classic deadline reset) and 0.5 ms (pins the microsecond-precision accounting). A watchdog thread disarms the timer after 3 s so a regressed build fails the elapsed assertion at ~3.2 s instead of hanging; the fixed poll returns in ~200 ms. Verified in both directions: the naive loop fails both cadences, a millisecond-accounting variant fails exactly the sub-millisecond one. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the block-wait timeout of streaming inserts (`input_format_max_block_wait_ms`) potentially never expiring while the query profiler is active: the file-descriptor poll restarted with the full timeout after every profiler signal, so a stalled input source could delay the partial-block flush indefinitely instead of flushing when the wait limit is reached.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109602",
          "createdAt": "2026-07-07T07:53:58Z",
          "updatedAt": "2026-08-13T12:28:07Z",
          "timestamp": "2026-08-13T12:28:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5245b33d164072575b92",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114601",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114601",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a renamed and dropped column being read from the wrong file while the mutation is pending",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114562 ## Problem A column that is renamed, dropped and added again in one `ALTER` silently returns the **dropped** column's data instead of its own default, while the mutation is still pending: ```sql CREATE TABLE t (a UInt64, h UInt8 DEFAULT 0) ENGINE = MergeTree ORDER BY tuple() SETTINGS min_bytes_for_wide_part = 0, min_rows_for_wide_part = 0; INSERT INTO t SELECT number + 100, 0 FROM numbers(1000); SET alter_sync = 0; -- mutation left pending ALTER TABLE t RENAME COLUMN a TO b, DROP COLUMN b, ADD COLUMN b UInt64 DEFAULT 7; SELECT count(), min(b), max(b) FROM t; -- 1000 100 1099 <- the old `a` values -- expected: 1000 7 7 ``` No error and nothing in the log; it affects `Wide` and `Compact` parts alike. The added `b` is a different column from the one that was dropped, so its rows have to come from its `DEFAULT`. ## Root cause `AlterConversions` records a pending drop under the name the command used: ```cpp else if (command.type == DROP_COLUMN) { dropped_columns.emplace(command.column_name); } ``` while every consumer of `isColumnDropped` asks about a column under the name it has **in the part**: - `IMergeTreeReader::isColumnDroppedByPendingMutation` takes the name out of `columns_to_read`, which is built from `getColumnInPart` and therefore already has the rename mapping applied (`IMergeTreeReader.cpp:81`); - `injectRequiredColumnsRecursively` passes `column_name_in_part` (`MergeTreeBlockReadUtils.cpp:119`). Without a rename the two names are identical, so the hazard those checks exist for — a column dropped and re-added under the same name, whose data in the part is stale — behaves correctly. That is why this was never noticed. With a rename they diverge: the part holds `a`, the drop is recorded for `b`, the check asks about `a`, finds nothing, and the reader streams the `a` file into the new `b`. ## Solution Resolve the drop to the part-side name while the commands are still in order. Commands arrive as issued, so the renames recorded so far are exactly those preceding the drop: ```cpp auto name_in_part = command.column_name; for (const auto & entry : rename_map) { if (entry.rename_to == name_in_part) { name_in_part = entry.rename_from; break; } } dropped_columns.emplace(std::move(name_in_part)); ``` Every existing caller already passes a part-side name, so all of them become correct at once and none of them changes. Chained renames are collapsed into one entry by `addMutationCommand`, so `a -> b -> c` resolves `c` straight back to `a`. The same ordering has to retire a rename mapping whose source was dropped, once that name is taken over. `RENAME a TO b, DROP b, RENAME c TO b` left **two** entries pointing at `b` — from the dropped `a` and from the live `c` — and the lookup returns whichever comes first, so `b` read the wrong column: ```cpp std::erase_if(rename_map, [&](const RenamePair & entry) { return entry.rename_to == command.rename_to && dropped_columns.contains(entry.rename_from); }); ``` This one is pre-existing too, and both halves are needed to get it right: | variant | `min(b), max(b)`, want `5000 5999` | |---|---| | `master` | `100 1099` — the **dropped** column's data | | part-side drop only | `0 0` — the default | | both | `5000 5999` ✔ | ### Why not fix it in the reader Checking the renamed-to name in the reader as well looks equivalent and is wrong. The opposite command order produces exactly the same `AlterConversions` state, and `01281_alter_rename_and_other_renames` already pins the required behaviour for it: ```sql ALTER TABLE t DROP COLUMN value2_old, RENAME COLUMN value2 TO value2_old; ``` Here the drop hit a *different* column that merely carried that name, and `value2`'s data is live and must be readable as `value2_old`. The only thing separating the two situations is where the drop sits among the renames, which the readers cannot see. Measured with the reader-side variant applied, live data is replaced by defaults: ``` dropped, then the name reused by a rename -1000 100 1099 +1000 0 0 ``` Hence the ordering has to be consumed at construction time. ### Tests `tests/queries/0_stateless/04872_read_renamed_then_dropped_and_readded_column.sh` — six cases: the defect on `Wide` and on `Compact`, the `RENAME/DROP/RENAME` case above, plus three controls that must not move — drop-and-re-add without a rename (what the check was written for), a pending rename with no drop, and the mirror case. Because the two possible wrong answers fail in opposite directions, the result is checked three ways: | variant | the defect | the mirror case | |---|---|---| | `master` | `1000 100 1099` ✗ | `1000 100 1099` ✔ | | reader-side variant | `1000 7 7` ✔ | `1000 0 0` ✗ | | this change | `1000 7 7` ✔ | `1000 100 1099` ✔ | `01281_alter_rename_and_other_renames`, `01278_alter_rename_combination`, `03905_chained_rename_column_mutation`, `04053_alter_add_rename_column_in_single_query` and `04011_detach_rename_attach_column` all pass. ### Notes for reviewers - `isColumnDropped` now consistently answers about part-side names. That matches all four existing call sites, so it is a fix rather than a contract change for them, but a future caller holding a metadata name would have to resolve it first. - **Overlaps with #114562, in a way that is safe in either merge order.** That PR adds a second lookup in the empty-candidate accounting loop of `injectRequiredColumns`, mapping a part name forward to its renamed-to name and testing the drop under that, so that a single-statement `RENAME a TO b, DROP b` is recognised as accounted for. This change makes the *first* lookup answer that correctly, because the drop is then already recorded as `a`, so the second lookup becomes dead code once both are in. It is deliberately left in place here rather than removed pre-emptively: while only #114562 is merged the second lookup is what makes that case work. A small cleanup to delete it belongs after this merges. Checked on a scratch branch carrying both changes with that lookup deleted — `04870`, `04871`, `04872`, `03830`, `04011` and `01281_alter_rename_and_other_renames` all still pass, so the cleanup is safe. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a column that was renamed, dropped and added again in one `ALTER` returning the dropped column's values instead of its own default while the mutation was still pending.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114601",
          "createdAt": "2026-08-13T07:45:51Z",
          "updatedAt": "2026-08-13T12:25:54Z",
          "timestamp": "2026-08-13T12:25:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f06ca25bb6070f04c26d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114414",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114414",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add SHARED REGEXP path placement policy to JSON",
          "text": "Adds `SHARED REGEXP 'pattern'` to `JSON` type declarations so matching root-relative flattened paths are always stored in shared data instead of competing for dedicated dynamic-path subcolumns. This is useful for high-cardinality path families whose promotion would displace more useful paths. Matching is partial by default, like `SKIP REGEXP`. Set the persisted `JSON` type parameter `shared_regexp_use_partial_match=0` to require full-string matching. Typed paths take precedence over `SHARED REGEXP`; `SKIP` and `SKIP REGEXP` continue to discard matching data. Rules are evaluated against the complete path from the original `JSON` root, including derived sub-objects. The rules are compiled once into an immutable `RE2` matcher, with bounded rule count and pattern sizes. A single rule uses direct `RE2` matching and multiple rules use `RE2::Set`. `JSON` columns without rules keep a null matcher, their existing binary type encoding, and their existing row-placement path. Metadata snapshots are copied lazily only when retained placement provenance actually changes a type. Placement provenance is stored in part column metadata and retained by default through horizontal and vertical merges, wide and compact mutations, column renames, lightweight-update patch materialization, and `Array`/`Nullable`/`Tuple`/`Map` wrappers. Removing a rule therefore does not unexpectedly promote already-shared paths during a later rewrite. Set the `MergeTree` table setting `allow_json_shared_data_paths_repromotion=1` to opt into reconsidering those paths. Policy-only metadata changes inside `Variant` are documented as unsupported. The `Native` binary type encoding uses `JSON` encoding version 1 only when a policy is present; ordinary `JSON` types remain byte-identical version 0. The canonical `JSON` documentation and `Native` format specification are updated accordingly. Testing includes 18 focused unit tests and 9 stateless scenarios covering syntax, partial/full matching, root-relative sub-objects, all supported wrappers, flattened `Native` and `RowBinary` paths, horizontal/vertical merges, compact/wide mutations, renames, statistics, and both lightweight-patch application paths. An 88.6 MB real Elasticsearch `indices_stats` document produced 143,555 flattened paths; all 143,290 paths targeted by `^indices[.]` remained in shared data and none became dynamic. A one-run end-to-end smoke comparison was 1.66 s without the policy and 1.69 s with it, with a 0.055% max-RSS difference; these small differences are noise-level, not a statistical benchmark. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `SHARED REGEXP` rules to the `JSON` data type to keep matching paths in shared data, with configurable full-string matching and explicit control over re-promoting paths after rules are removed.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114414",
          "createdAt": "2026-08-12T02:35:41Z",
          "updatedAt": "2026-08-13T12:24:44Z",
          "timestamp": "2026-08-13T12:24:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "valerypetrov",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6daf5118aa6404fa559d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113558",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113558",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix missing materialized-CTE gate when a `Merge` table has several children",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/113184 Related: https://github.com/ClickHouse/ClickHouse/pull/113489 Related: https://github.com/ClickHouse/ClickHouse/pull/111194 Related: https://github.com/ClickHouse/ClickHouse/pull/113043 Related: https://github.com/ClickHouse/ClickHouse/pull/108924 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the `LOGICAL_ERROR` `Reading from materialized CTE '...' before its materialization completed - DelayedPortsProcessor gate is missing in the query plan` raised when a `Merge` table with more than one child reads a materialized CTE that the outer query references. ### Description `ReadFromMerge::createChildrenPlans` optimizes every child plan on its own, and `resolveMaterializingCTEs` claims a materialized CTE globally through `MaterializedCTE::is_materialization_planned`. The first child plan to be optimized therefore moved the CTE's plan into *its own* tree, and every other `DelayedMaterializingCTEsStep` for that CTE - in the sibling children and in the outer plan - degenerated into a gate-less `MaterializingCTEsStep`. The CTE's writer then sat in one child's pipeline while readers sat in another, with no `DelayedPortsProcessor` between them, so a sibling's in-place `IN`-set build read the storage while it was still empty and `ReadFromMemoryStorageStep` raised the exception (in debug and sanitizer builds it aborts the server). Reproducer on master, deterministic (20/20), also with `max_threads = 1`, and a plain `EXPLAIN` is enough because the set is built during plan optimization: ```sql SET enable_analyzer = 1, enable_materialized_cte = 1; CREATE TABLE t (x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t SELECT number FROM numbers(10); CREATE TABLE tdist AS t ENGINE = Distributed(test_shard_localhost, currentDatabase(), t); CREATE VIEW tconst AS SELECT toUInt64(1) AS x; WITH t AS MATERIALIZED (SELECT number AS c FROM numbers(2)) SELECT count() FROM merge(currentDatabase(), '^(tconst|tdist)$') WHERE (x IN (t)) AND (x NOT IN (t)); ``` Children are visited in table-name order, so `tconst` is planned first and claims the CTE, and the `Distributed` child that follows builds the set in place while the CTE is unbuilt. Renaming so the `Distributed` child sorts first makes the same query pass, which is what pins the mechanism. **Fix.** A child plan no longer claims a CTE that the outer query references as well. `removeDelayedMaterializingCTEsStepFor` strips those steps from the child plan before it is optimized, leaving the outer plan - whose `MaterializingCTEsStep` sits above the whole merge - as the single owner that gates every child. This is the same reasoning `DelayedCreatingSetsStep::makePlansForSets` already applies to pre-built `IN`-subquery plans. The set of CTEs to strip comes from walking the outer `query_info.query_tree`, so a CTE defined *inside* one child (a `View` with its own `WITH ... AS MATERIALIZED`) is left owned by that child - it is the only reader, and stripping it unconditionally would leave it with no materialization at all. **Validation.** New `04811_materialized_cte_merge_child_gate` covers the failing child order, the explicit-subquery form, a satisfiable predicate that pins the data rather than only the absence of the exception, the `EXPLAIN` route, the reverse child order that always worked, and the view-owned-CTE case that must keep materializing inside the child. Every failing arm reproduces 5/5 on a master binary and passes 5/5 after the change. The `materialized_cte` suite is green. This is one shape of a recurring family - the same assertion is also reported in #113184 and addressed for other shapes by #113489, #111194 and #113043 - so the underlying `is_materialization_planned` claim being global while the gate is per-plan is worth revisiting separately. It surfaces constantly in the AST fuzzer; found via https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=9d0b1a25ba7aa4579c95a65baca002d1dd7a1e47&name_0=MasterCI&name_1=AST%20fuzzer%20%28amd_debug%29",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113558",
          "createdAt": "2026-08-05T19:32:08Z",
          "updatedAt": "2026-08-13T12:24:43Z",
          "timestamp": "2026-08-13T12:24:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:365bd66e70fc68e1bc1e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112601",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112601",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Iterate ColumnObject subcolumns in sorted path order",
          "text": "Make `ColumnObject::forEachSubcolumn`, `forEachMutableSubcolumn` and their recursive variants iterate over the already-maintained `sorted_typed_paths`/`sorted_dynamic_paths` lists instead of the raw `typed_paths`/`dynamic_paths` `unordered_map`s. Motivation: the iteration order of an `unordered_map` is not guaranteed to be preserved by its copy constructor, and `IColumn::mutate` clones the column (copy-constructing those maps) when it is shared. `IColumn::convertToFullIfWrapped` collects the unwrapped subcolumns while iterating the source column and then reassigns them positionally while iterating the mutated (possibly cloned) column, so both passes must visit subcolumns in the same order. If the cloned maps iterated in a different order, converted subcolumns would be assigned to the wrong paths. Iterating the sorted path lists makes the visiting order deterministic and stable across cloning, removing the reliance on unspecified `unordered_map` copy-order behavior. This is currently latent — the bundled `libc++` happens to preserve copy order — so it is a safety change with no user-visible effect. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1131` (included in `26.8` and later) - Backported to: `26.7.4.17`, `26.6.3.32`, `26.5.7.40`, `26.3.18.22`, `25.8.30.14` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112601",
          "createdAt": "2026-07-30T13:55:18Z",
          "updatedAt": "2026-08-13T12:22:37Z",
          "timestamp": "2026-08-13T12:22:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-must-backport",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "Avogar",
          "state": "closed",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:88cfdc15774fa538b422",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112479",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112479",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "RabbitMQ related fix",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Security related bugfix. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1209` (included in `26.8` and later) - Backported to: `26.7.4.26`, `26.6.3.31` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112479",
          "createdAt": "2026-07-29T18:56:18Z",
          "updatedAt": "2026-08-13T12:22:33Z",
          "timestamp": "2026-08-13T12:22:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "submodule changed",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "kssenii",
          "state": "closed",
          "assignees": [
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ca342bd41e7526ccc4aa",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111794",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111794",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add unordered stream modifier",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM UNORDERED modifier: skip the per-snapshot commit-order sort depends on https://github.com/ClickHouse/ClickHouse/pull/110653 (not for functional reason, only test) cc @alesapin @Michicosun",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111794",
          "createdAt": "2026-07-24T13:31:48Z",
          "updatedAt": "2026-08-13T12:22:31Z",
          "timestamp": "2026-08-13T12:22:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "SmitaRKulkarni",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a59a8dbb4f056b715b7a",
        "signalId": "github:ClickHouse/ClickHouse:issue:114404",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114404",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Inlined ALIAS body adds constants to the shipped WITH FILL header; parallel replicas throw",
          "text": "_Found via ClickGap automated review — close or comment if this is wrong._ ### Describe what's wrong With parallel replicas enabled, `SELECT k, a_v FROM t ORDER BY k WITH FILL ... INTERPOLATE (...)` where `a_v` is an `ALIAS` column whose body is an expression fails with `Code: 20. Number of columns doesn't match (source: 5 and result: 4). (NUMBER_OF_COLUMNS_DOESNT_MATCH)`. The same query returns rows correctly without parallel replicas, and returns rows correctly under parallel replicas if the alias body is written inline in the `SELECT` list instead of declared as an `ALIAS` column. **Root cause:** `inlineAliasColumns` at src/Interpreters/ClusterProxy/executeQuery.cpp:1048 changes the structure of the shipped query tree, but the initiator's `expected_header` is still computed from the un-inlined tree; when the shipped tree's `WithMergeableState` header gains columns (any plan step that keeps ActionsDAG intermediates, e.g. `Filling`), no reconciliation path can absorb it. <details> <summary>Analysis details (evidence, affected locations, impact)</summary> **Why we believe this is a bug:** `PlannerJoinTree::buildQueryPlanForTableExpression` (src/Planner/PlannerJoinTree.cpp:2048) calls `ClusterProxy::executeQueryWithParallelReplicas`, which since this PR runs `inlineAliasColumns` on the shipped tree (src/Interpreters/ClusterProxy/executeQuery.cpp:1048) and derives the replica `header` from that inlined tree (executeQuery.cpp:1052). Back in `buildQueryPlanForTableExpression`, `expected_header` is re-planned from the ORIGINAL, un-inlined `select_query_info.query_tree` (PlannerJoinTree.cpp:2262-2268). For `ORDER BY ... WITH FILL ... INTERPOLATE` the `WithMergeableState` header is the `Filling` step header, which keeps every ActionsDAG output including the constants of the inlined body — so the inlined tree yields one more column (`2_UInt8` for a `v * 2` body) than the un-inlined tree. `buildShardCollapseFanOut` bails out because it only handles a SMALLER shard header (src/Storages/buildQueryTreeForShard.cpp:1313), and the positional `ActionsDAG::makeConvertingActions` at PlannerJoinTree.cpp:2293 then throws. **Affected locations:** - `src/Interpreters/ClusterProxy/executeQuery.cpp:1048` — new `inlineAliasColumns` call on the shipped tree; `header` at line 1052 is derived from it - `src/Planner/findParallelReplicasQuery.cpp:622` — sibling new `inlineAliasColumns` call; `initial_header` at line 615 is taken from the un-inlined tree and reconciled positionally at line 650 - `src/Planner/PlannerJoinTree.cpp:2293` — positional `makeConvertingActions` that throws NUMBER_OF_COLUMNS_DOESNT_MATCH - `src/Storages/buildQueryTreeForShard.cpp:1313` — `buildShardCollapseFanOut` returns {} when the shard header is not strictly smaller, so a LARGER shard header is unhandled **Impact:** Any query combining an `ALIAS` column whose body is an expression with `ORDER BY ... WITH FILL ... INTERPOLATE` is rejected once `enable_parallel_replicas = 1`. Deterministic, reproduces with both `parallel_replicas_local_plan = 0` and `= 1`. It also masks the correct user error: `INTERPOLATE (k AS k)` (an ORDER BY column as interpolate target) reports `NUMBER_OF_COLUMNS_DOESNT_MATCH` instead of `INVALID_WITH_FILL_EXPRESSION`. </details> ## Assumptions _Unverified assumptions — tick to confirm, comment to refute:_ - [ ] **Before this PR the same query succeeded on the parallel-replicas path** - *Why unverifiable:* no pre-PR binary is available in this environment to run the query against - *Falsifiable test:* Build master (without this PR) and run the repro; expect `0 10 / 2 10 / 4 14 / 6 14 / 8 18`. Pre-PR `executeQueryWithParallelReplicas` shipped the un-inlined tree, so `header` and `expected_header` came from the same query node and matched structurally. ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `48b91073fbabef`). ### How to reproduce ```sql CREATE TABLE t (k UInt32, v Int64, a_v Int64 ALIAS v * 2) ENGINE = MergeTree ORDER BY k; INSERT INTO t VALUES (0,5),(4,7),(8,9); then run SELECT k, a_v FROM t ORDER BY k WITH FILL FROM 0 TO 10 STEP 2 INTERPOLATE (a_v AS a_v) SETTINGS enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, automatic_parallel_replicas_mode = 0, serialize_query_plan = 0; ``` ### Expected behavior ``` 0 10 2 10 4 14 6 14 8 18 0 10 2 10 4 14 6 14 8 18 ``` ### Error message and/or stacktrace ``` 0 10 2 10 4 14 6 14 8 18 Received exception from server (version 26.8.1): Code: 20. DB::Exception: Received from 127.0.0.1:19020. DB::Exception: Number of columns doesn't match (source: 5 and result: 4). (NUMBER_OF_COLUMNS_DOESNT_MATCH) ``` ### Additional context **Open risks:** - The same repro over a 2-shard `Distributed` table fails identically. That path has always inlined, so it is likely broken on master too and is out of scope for this PR — but a fix should cover both, since they share `buildQueryTreeForShard`. - Only `WITH FILL`/`INTERPOLATE` was found to leak ActionsDAG intermediates into the `WithMergeableState` header. Other steps with the same property would fail the same way; not audited exhaustively. **Suggested fix:** Compute the initiator-side `expected_header` from the SAME inlined tree that is shipped (or run `inlineAliasColumns` on the tree used for `expected_header` in `PlannerJoinTree::buildQueryPlanForTableExpression`), instead of relying on the shipped and expected headers happening to agree. Alternatively extend `buildShardCollapseFanOut` to drop shard columns the initiator does not expect, not only to fan out missing ones. Found during automated review of [PR #107700](https://github.com/ClickHouse/ClickHouse/pull/107700). --- _ClickGapAI · Severity: P2 · Finding: `h_pr107700_001`_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114404",
          "createdAt": "2026-08-12T00:43:39Z",
          "updatedAt": "2026-08-13T12:21:58Z",
          "timestamp": "2026-08-13T12:21:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "bug",
            "comp-query-analyzer",
            "comp-parallel-replicas"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "yakov-olkhovskiy"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:38973037a8631234f03b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114226",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114226",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112601 to 26.6: Iterate ColumnObject subcolumns in sorted path order",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112601 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.32` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114226",
          "createdAt": "2026-08-10T20:10:17Z",
          "updatedAt": "2026-08-13T12:21:42Z",
          "timestamp": "2026-08-13T12:21:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "closed",
          "assignees": [
            "Avogar",
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7da887574e79459c3be4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110997",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110997",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix data race on DataTypeAggregateFunction version during Native serialization",
          "text": "Related: found by the `arm_tsan` and `azure, amd_tsan` Stress tests (STID 3977-4818, ThreadSanitizer data race). No existing issue. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results reading `AggregateFunction` states after one client requested them at an older protocol revision. The serialization version chosen for that one response was written onto the data type object shared by the whole table, so it stayed there: every later query read the states at that version, and for `sumMap` over `Decimal32` that returns wrong values, while the column also lost the version in `system.columns` and on the wire. The same in-place write was a data race between concurrent queries serializing such a column in the `Native` format. ### Description A single `DataTypeAggregateFunction` instance is shared across query result blocks: it lives once in the table's column description and is aliased by shallow column copies. `NativeWriter`/`NativeReader` called `setVersionToAggregateFunctions`, which walked to the leaf type and wrote its `mutable version` field in place. Two concurrent `Native` serializations of the same aggregate-function-typed column then raced on that field. Reports: * https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109496&sha=62eeb400eafe86cedff9b5a4e9a36aa722633cda&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29 * https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b59441bd06c2fcb6b103a30874528cc398afc723&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 Both racing stacks are `setVersionToAggregateFunctions` -> `DataTypeAggregateFunction` version setter, via `NativeWriter::write` -> `TCPHandler::processOrdinaryQuery`/`sendData`. Fix: instead of mutating the shared type object, replace the versioned leaf with a copy that carries the version via the constructor (the same way the binary-encoding decode path builds versioned types). The in-place `setVersion`/`updateVersionFromRevision` mutators and the `mutable` qualifier are removed so the object is immutable after construction. This also removes a latent issue that was worse than the race itself. `NativeWriter` passes `if_empty = false` for a client older than `DBMS_MIN_REVISION_WITH_AGGREGATE_FUNCTIONS_VERSIONING`, which unconditionally forced version 0 onto the shared type. Every later query then kept that 0 (`if_empty = true` sees a version already set), so it advertised a type name without a version while serializing version-0 states, and the receiving client - deriving the version from the server revision - deserialized them as version 1. #### Preserving custom type names Because the leaf is now replaced rather than mutated, the rebuilt tree is what the caller ends up with, so the rebuild must not lose anything. Rebuilding the wrappers through `transformTypesRecursively` recreates `Array`/`Tuple`/`Map` via `make_shared` and drops custom type names: | type | expected | with a naive rebuild | |---|---|---| | `Nested(x AggregateFunction(sumMap, ...))` | preserved | `Array(Tuple(...))` | | `SimpleAggregateFunction(anyLast, AggregateFunction(sumMap, ...))` | preserved | `AggregateFunction(...)` | Both are user-visible: the type is sent to the client over `Native`, and on `ATTACH` it becomes the column type in the table metadata. Losing the `SimpleAggregateFunction` name is worse than cosmetic - `AggregatingSortedAlgorithm` and `SummingSortedAlgorithm` recognise such a column by `dynamic_cast` on that very name object, so the column would silently start merging as a plain aggregate function state. So `setVersionToAggregateFunctions` walks the type itself, over exactly the wrappers `transformTypesRecursively` descended into, and returns the original pointer when no leaf changes. `Nullable` is among them: a state cannot be directly inside `Nullable`, but a `Tuple` can, and `Nullable(Tuple(AggregateFunction(...)))` is reachable with `enable_nullable_tuple_type`. A custom name can also sit on the wrapper rather than on the leaf, as in `SimpleAggregateFunction(anyLast, Array(AggregateFunction(...)))`, so a rebuilt wrapper carries the customization of the original too. `DataTypeCustomNamePtr` becomes a `shared_ptr` so a copy of a type can carry the very same custom name object, via the new `IDataType::cloneCustomization`. `Nested` is rebuilt with its custom name kept in sync with the new element types, directly rather than through `createNested`: the latter derives the type from the printed name, and version 0 is deliberately not printed, so a name round trip would turn a leaf explicitly pinned to version 0 back into an unversioned one using the latest version. `callOnNestedSimpleTypes` had no other caller and is removed, so `transformTypesRecursively` (shared with schema inference) is left untouched. ### Testing * `gtest_aggregate_function_version_race` covers the shared-object mutation, nested types, both custom-name cases above, and stress-tests concurrent version assignment over one shared type object. * `04612_aggregate_function_version_custom_type_names` round-trips both types through `Native` (which assigns the version in the writer and again in the reader) and through `DETACH`/`ATTACH`. * `04613_aggregate_function_version_not_sticky` is the one that fails on `master` HEAD. The two tests above assert output that is byte-identical to `master` by design, so neither can. It asks for a `Native` response at a revision below the one that introduced versioning, then checks the column again: on `master` the version 0 forced for that one response stays on the shared type, so a later plain `SELECT finalizeAggregation(s)` reads `([1,2],[10.5,20.25])` back as `([1],[10.5])` and `system.columns` loses the version. The revision is pinned explicitly and the type name the response carries is asserted, so the test cannot pass without that path having run - checked against both ways of it not running, a request that fails and a version assignment that does nothing. * The serialized version bytes are unchanged; the `Native` wire type names are byte-identical to `master` for both types. * 472 related stateless tests (`simple_aggregate`, `nested`, `native`, `geo`, `point`, `polygon`, `aggregate_function`) were run against this build and against a `master` build on the same server config: the failure sets are identical, i.e. no test fails only with this change.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110997",
          "createdAt": "2026-07-19T14:41:39Z",
          "updatedAt": "2026-08-13T12:20:14Z",
          "timestamp": "2026-08-13T12:20:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 23
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b63c5757e5e4b59a510b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108336",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108336",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Release pull request for branch 26.6",
          "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108336",
          "createdAt": "2026-06-23T22:33:53Z",
          "updatedAt": "2026-08-13T12:18:57Z",
          "timestamp": "2026-08-13T12:18:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "release"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e63246a60e797ca643e6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:102192",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:102192",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix \"Not-ready Set\" exception when buildOrderedSetInplace fails",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/107924 `FutureSetFromSubquery::buildOrderedSetInplace` (the speculative set build run during primary key analysis for `IN` subqueries) used to consume the subquery `source` plan up front. If the in-place build then failed silently (e.g. due to subquery timeout with `timeout_overflow_mode = 'break'`, where the executor stops without throwing and without setting `is_created`), the set remained permanently unbuilt and `FunctionIn` threw \"Not-ready Set is passed as the second argument\". The core fix — running the in-place pipeline against a clone of the source plan so that `DelayedCreatingSetsStep::makePlansForSets` can still build the set after a silent failure — has meanwhile been merged in #107924, which falls back to the original destructive build when the source plan contains a step that does not implement `IQueryPlanStep::clone`. After merging master, this PR carries the remaining hardening on top of #107924: - `clone` implementations for eleven more query-plan steps: `JoinStep`, `FillingStep`, `ExtractColumnsStep`, `FractionalLimitStep`, `FractionalOffsetStep`, `MergingAggregatedStep`, `NegativeLimitStep`, `NegativeLimitByStep`, `NegativeOffsetStep`, `StreamInQueryResultCacheStep`, `ReadFromQueryResultCacheStep`. `IN` subqueries whose plans contain these steps (e.g. `ORDER BY ... WITH FILL`, joins planned as `JoinStep`, subqueries going through the query result cache) previously took the destructive fallback, where a silent in-place failure still reproduces the \"Not-ready Set\" exception; with clone support they take the non-destructive path and recover. - `JoinStep::clone` covers only the join algorithms whose `IJoin` implementation supports cloning: `HashJoin`, `ConcurrentHashJoin`, `ConstantJoin` and `FullSortingMergeJoin`. `JoinSwitcher` (`join_algorithm = 'auto'`), `SpillingHashJoin`, `MergeJoin` / `PartialMergeJoin` and `GraceHashJoin` keep the pre-existing destructive fallback: `IJoin::isCloneSupported` also gates `optimizeJoin`, `convertOuterJoinToInnerJoin` and `QueryPipelineBuilder`, so widening it would silently enable join swapping and outer-to-inner conversion for never-validated stateful spilling algorithms. This limitation is documented at the throw site, and `04649_not_ready_set_with_non_clonable_join_algorithms` pins that those algorithms still produce correct results (the `NOT_IMPLEMENTED` stays internal). - `FilledJoinStep` (`StorageJoin` and dictionary joins) and `JoinStepLogicalLookup` (direct key-value joins) intentionally remain non-clonable as well: they wrap live storage-backed join state with no safe copy semantics (`IJoin::clone` constructs an empty join rather than copying the filled state). The constraint is documented at both classes. - Copying a `QueryResultCacheWriter` (which is what `StreamInQueryResultCacheStep::clone` does) now carries over both the original's `query_start_time` and its `skip_insert` decision. On the successful in-place build path the copy is the only writer that is ever finalized, so otherwise `query_cache_min_query_duration` would be measured from the clone point and an already cache-resident key would be buffered into a throwaway buffer instead of being skipped. Pinned by `04652_query_result_cache_subquery_clone_min_query_duration` and `04653_query_result_cache_subquery_clone_skip_insert`. - `QueryPlan::clone` now preserves `max_threads` and `concurrency_control`, so a pipeline built from a cloned plan runs under the same resource contract as the original instead of defaulting to `max_threads == 0` with concurrency control disabled. - The Planner no longer plants query result cache steps into a *logical* plan. A logical plan is not executed where it is built: it is serialized and shipped to another node (parallel replicas with `serialize_query_plan = 1`, see `createRemotePlanForParallelReplicas`), while `StreamInQueryResultCacheStep` and `ReadFromQueryResultCacheStep` hold node-local state (a `QueryResultCacheWriter`, or the cached chunks) that has no serialized representation. Before this, `query_cache_for_subqueries = 1` together with parallel replicas and `serialize_query_plan = 1` failed the whole query with `Method serialize is not implemented for StreamInQueryResultCache`. The cache is still populated and read by the plan the initiator executes itself. - Regression tests using the `prepared_sets_build_ordered_set_inplace_fail` failpoint: `04095_global_not_in_parallel_replicas` (the original parallel-replicas reproducer), `04492_not_ready_set_with_fill_subquery` (a `WITH FILL` subquery source plan), `04550_not_ready_set_with_join_subquery`, `04648_not_ready_set_with_query_result_cache_subquery` (both the cache write and the cache read path) and `04649_not_ready_set_with_non_clonable_join_algorithms`. The original failure was observed in the [Stress test (amd_tsan)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=d0432097aed783bd35054dce2edcefe0c4e5122c&name_0=MasterCI&name_1=Stress%20test%20%28amd_tsan%29) on master, where `considerEnablingParallelReplicas` triggers `selectRangesToRead`, which calls `buildOrderedSetInplace` for primary key analysis. Note on the changelog category: the \"Not-ready Set\" exception is a `LOGICAL_ERROR`, so reproducing it on the unfixed build aborts the server on the sanitizer builds that the per-arch Bugfix validation jobs run on, and the validation harness deliberately treats a mid-run server death as inconclusive rather than as a reproduced bug. Automated bugfix validation is therefore structurally impossible for this fix, and the user-visible bug fix for the common plan shapes already shipped with #107924; the remaining hardening here is classified as an Improvement. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Extend the non-destructive speculative in-place set build for `IN` subqueries (introduced in [#107924](https://github.com/ClickHouse/ClickHouse/pull/107924)) to more query-plan shapes: `clone` is implemented for eleven more query-plan steps, including clone-supported joins via `JoinStep`, `ORDER BY ... WITH FILL` via `FillingStep`, and the query result cache steps. Such subqueries now preserve the original source plan and can recover from a silent in-place build failure instead of throwing \"Not-ready Set is passed as the second argument\". A cloned query plan now preserves `max_threads` and `concurrency_control` of the original plan. Also fixed `Method serialize is not implemented for StreamInQueryResultCache` when `query_cache_for_subqueries = 1` is used together with parallel replicas and `serialize_query_plan = 1`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/102192",
          "createdAt": "2026-04-09T08:31:34Z",
          "updatedAt": "2026-08-13T12:18:14Z",
          "timestamp": "2026-08-13T12:18:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 53
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3df9e40a7b984e59a1f4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114300",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114300",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "TimeSeries: store all tags in the `tags` column",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): TimeSeries: store all tags in the `tags` column The `tags` column of the tags target table now contains all the tags, including the `__name__` tag with the metric name and the tags with dedicated columns from the `tags_to_columns` setting, so the map alone fully identifies a time series. The default id generator now hashes just `tags`, i.e. now it looks like `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(tags)))` (instead of `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(metric_name, all_tags)))` )",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114300",
          "createdAt": "2026-08-11T10:29:30Z",
          "updatedAt": "2026-08-13T12:17:59Z",
          "timestamp": "2026-08-13T12:17:59Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-promql"
          ],
          "author": "vitlibar",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:febd6df52970a5774927",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114417",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114417",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add the `Cluster` database engine",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114411 Related: https://github.com/ClickHouse/ClickHouse/issues/59304 Related: https://github.com/ClickHouse/ClickHouse/pull/110975 Implement the `Cluster` database engine, a follow-up to the `Remote` database engine (#110975). It provides real-time access to the tables of a database on a cluster from the server configuration — the named-cluster counterpart of `Remote`, exactly as the `cluster` table function relates to the `remote` table function: ```sql CREATE DATABASE db ENGINE = Cluster('cluster_name', 'database'); ``` The engine shares the whole metadata and query machinery with the `Remote` database engine (the list of tables and their structure are fetched from the cluster on demand, each table is a `Distributed` proxy forwarding `SELECT` and `INSERT`, the same local-shard visibility rules and remote-replica fallback). The differences: - The cluster is resolved from the server configuration by name on every access, so the database follows configuration reloads and cluster auto-discovery, like a `Distributed` table does. Macros such as `{cluster}` are supported. The cluster must exist at `CREATE`, but a database whose cluster later disappears from the configuration does not prevent the server from starting — it reports the missing cluster until the configuration brings it back. - Connections use the per-replica settings of the cluster configuration (credentials, secure connections, compression, the inter-server secret), so the engine takes no credential arguments and stores no secrets. - `SHOW CREATE TABLE` prints a re-executable `Distributed('cluster_name', 'database', 'table')` definition (including the implicit `rand()` sharding key of a multi-shard database). The only exception is a table currently served through the remote-replica fallback: no equivalent re-executable definition exists for that transient state (a `Distributed` table over the whole cluster performs no such fallback, and the per-replica settings of the configuration cannot be carried by an explicit address list), so `SHOW CREATE TABLE` reports `THERE_IS_NO_QUERY` until the local replica has the objects again. Notes: - A chain of `Remote`/`Cluster` databases on the same server that refers back to itself is rejected at `CREATE DATABASE` of the database that would complete it (`INFINITE_LOOP`), because a lazily-reported cycle would fail every whole-server scan (`system.tables` and the like) for every user. A cycle that comes into existence later (configuration reload, explicit `ATTACH`) is skipped in listings and reported only by table resolution, and does not prevent the server from starting. - The remote-only metadata-lookup fallback (the cluster with the local replicas stripped from their shards) is derived by the new `Cluster::tryGetClusterWithoutLocalReplicas`, which preserves the per-replica settings and the inter-server secret — the string-based rebuild used by the `Remote` engine could not do that for configured clusters, and the `Remote` engine now uses the new method as well. - The follow-up `ddl_mode` setting (forwarding DDL to the cluster) is tracked separately in https://github.com/ClickHouse/ClickHouse/issues/114412. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added the `Cluster` database engine that provides real-time access to the tables of a database on a cluster from the server configuration, forwarding `SELECT` and `INSERT` queries to it. It is the named-cluster counterpart of the `Remote` database engine, as the `cluster` table function relates to the `remote` table function. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114417",
          "createdAt": "2026-08-12T03:48:20Z",
          "updatedAt": "2026-08-13T12:13:40Z",
          "timestamp": "2026-08-13T12:13:40Z",
          "metrics": {
            "reactions": 1,
            "comments": 4
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:857a231864b3a5062154",
        "signalId": "github:ClickHouse/ClickHouse:issue:111271",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:111271",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Distributed ANY INNER JOIN returns one row per key per shard: deduplication is per-shard, not global",
          "text": "**Describe what's wrong** `ANY INNER JOIN` deduplicates per join key — but over a `Distributed` table it deduplicates on each shard independently, so a key whose rows span shards is returned once **per shard**. When the sharding key differs from the join key (the common case), the distributed result is a multiple of the correct local result. **Does it reproduce on the most recent release?** Reproduces on `26.7.1.408` and near-HEAD master `3090a4fc` (`26.7.1.653`). **How to reproduce** Cluster `two_shards`: two shards with `default_database` `sh0`/`sh1`. Sharding is by `k`; the join is on `a`, whose values appear on both shards. ```sql CREATE TABLE sh0.sha (k UInt32, a UInt32) ENGINE = MergeTree ORDER BY k; CREATE TABLE sh1.sha (k UInt32, a UInt32) ENGINE = MergeTree ORDER BY k; INSERT INTO sh0.sha SELECT number, number % 20 FROM numbers(0, 50); INSERT INTO sh1.sha SELECT number, number % 20 FROM numbers(50, 50); CREATE TABLE sha_d AS sh0.sha ENGINE = Distributed(two_shards, '', sha); CREATE TABLE shr (a UInt32) ENGINE = MergeTree ORDER BY a; INSERT INTO shr SELECT number FROM numbers(20); -- 20 (correct: one row per key) SELECT count() FROM (SELECT * FROM sh0.sha UNION ALL SELECT * FROM sh1.sha) AS l ANY INNER JOIN shr AS r ON l.a = r.a SETTINGS enable_analyzer = 1; -- 40 (WRONG: one row per key PER SHARD) SELECT count() FROM sha_d AS l ANY INNER JOIN shr AS r ON l.a = r.a SETTINGS enable_analyzer = 1, distributed_product_mode = 'global'; ``` **Characterization** - Deterministic. The count is exactly `#shards × correct` when every key spans all shards. - The single-node semantics (\"completely disables the cartesian product\" — one row per key) require the deduplication to happen globally (initiator-side or with `ANY`-awareness in the shard-merge), or the distributed rewrite of `ANY` strictness should be rejected/documented. Found by the optimizer-tester topology differential's genuinely-sharded arm (`tests/optimizer_tester`, branch `optimizer-tester-framework`) in a 10,000-query hunt. Related: https://github.com/ClickHouse/ClickHouse/issues/111195",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/111271",
          "createdAt": "2026-07-21T17:52:52Z",
          "updatedAt": "2026-08-13T12:13:18Z",
          "timestamp": "2026-08-13T12:13:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "bug",
            "comp-joins",
            "comp-distributed"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3ad0ec532aea25489c10",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114441",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114441",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112479 to 26.6: RabbitMQ related fix",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112479 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31577531496/job/94052923751) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.31` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114441",
          "createdAt": "2026-08-12T08:36:35Z",
          "updatedAt": "2026-08-13T12:21:44Z",
          "timestamp": "2026-08-13T12:21:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "closed",
          "assignees": [
            "kssenii",
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:051ea17b4d2dbd081b53",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114442",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114442",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112479 to 26.7: RabbitMQ related fix",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112479 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31577531496/job/94052923751) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.26` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114442",
          "createdAt": "2026-08-12T08:37:03Z",
          "updatedAt": "2026-08-13T12:21:46Z",
          "timestamp": "2026-08-13T12:21:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "closed",
          "assignees": [
            "kssenii",
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0fc6b329638c5fed1f78",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114323",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114323",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: require canonical internal links",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114230 This is a one-off cleanup of repository-authored documentation links that use legacy redirect aliases. It updates the current English documentation and source-embedded reference documentation to use routes relative to the docs root, so the automated translation PR can parse and localize them without producing missing locale routes. This PR intentionally adds no permanent CI checks or ongoing enforcement. Its scope is limited to the current link corrections needed to get the automated translation PR parsing successfully. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114230&sha=e599855a281a4dd841be20794036908a58acbf63&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114323&sha=86d2d7e194d7fbf75cc84616a8fa0da1ed84802e&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114323",
          "createdAt": "2026-08-11T13:26:47Z",
          "updatedAt": "2026-08-13T12:12:22Z",
          "timestamp": "2026-08-13T12:12:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci",
            "pr-autogenerated-docs"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ed0579ae8289134ff9d6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114472",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114472",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the patch release version bump increasing across recoveries",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113834 Related: https://github.com/ClickHouse/ClickHouse/pull/113528 --> Scheduled patch releases stopped incrementing the patch number — successive releases on a branch reused the same `vX.Y.P.*` line (e.g. `26.6.2.81`, `26.6.2.158`, `26.6.2.160`), because the post-release version bump was lost whenever a release was interrupted after the tag push, and every later recovery skipped it too. Prepare now classifies a run from the ref and the branch-tip version file into two flags: `is_recovery` (this run re-publishes an existing release rather than creating one) and `is_late_recovery` (the branch has already advanced to a newer release). The deferred bump is gated on `not is_late_recovery`, so a normal run and a current-release recovery complete the interrupted bump, while a superseded recovery never rewrites the version backwards. `stage_bump` refuses to write a version older than the branch tip, and asserts `is_recovery` on an empty bump — a fresh release must advance the version, a recovery may legitimately find it already done. Split out of #113834 — this is the version-bump half with a single push. Handling a non-fast-forward push (a backport moving the branch mid-release) is left to #113834. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114472",
          "createdAt": "2026-08-12T12:04:19Z",
          "updatedAt": "2026-08-13T12:10:07Z",
          "timestamp": "2026-08-13T12:10:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "leshikus",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e3a5f5eb0591214c26c5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114463",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114463",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Report mutation cancellation instead of stopping nested pipelines silently",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107619 Related: https://github.com/ClickHouse/ClickHouse/pull/112152 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `Not-ready Set is passed as the second argument for function 'in'` error when a mutation whose predicate contains `IN (subquery)` is cancelled, for example by `KILL MUTATION` or `DETACH DATABASE`, while the subquery's set is still being built. Related: #107619. ### Description Reported by @ alexey-milovidov on https://github.com/ClickHouse/ClickHouse/pull/112152#issuecomment-5180281674 after a `Stress test (amd_debug)` run aborted this way. `buildSetInplace` materializes the right side of `IN (subquery)` through a nested `CompletedPipelineExecutor` and polls the mutation's interactive-cancel callback. That callback returned a bare `bool`, so on cancellation the executor called `cancel()` and the poll loop returned normally: `Set::finishInsert` never ran, while `build()` had already moved the source plan out. `buildSetsForDAG` returns `void`, so `getMinMaxCountProjectionBlock` evaluated the filter in the same call and `FunctionIn` threw. A constant left-hand side (`1 IN (...)`) makes this the first materialization attempt: it maps to no key column, so primary-key analysis returns first. The callback now throws `ABORTED`, mirroring `MergeTask::checkOperationIsNotCanceled`, which reports the merge path the same way and is likewise used as a nested-pipeline callback. `MergeTreeBackgroundExecutor` already treats `ABORTED` as a normal outcome and logs it at DEBUG. Only the mutation installs this callback shape, so the installers in `TCPHandler` and `LocalConnection` are untouched and client cancellation still returns `QUERY_WAS_CANCELLED` quietly. Validation: `KILL MUTATION` and `DETACH DATABASE` mid-build abort the server before the change and are clean after it, and the table stays mutatable. New test `04865_cancel_mutation_in_subquery_minmax_projection`, 50 of 50 runs. The other cancellation routes and overflow modes checked are listed in the validation gate comment below. Out of scope: a caller that executes a filter DAG synchronously still does not check set readiness, and `buildSetInplace` still leaves the set unbuildable if some future path stops it silently. I found no input reaching either independently of this cancellation. Note for review: #113939 rewrites the same lambda for the S3 read path but keeps the interactive callback non-throwing, so it does not fix this. The two conflict textually; a resolution must keep both the throw and that PR's cancellation persistence.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114463",
          "createdAt": "2026-08-12T10:03:39Z",
          "updatedAt": "2026-08-13T12:08:02Z",
          "timestamp": "2026-08-13T12:08:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6951969e162455c46556",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112313",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112313",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep an S3Queue file retryable after losing the race for its `processing` node",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `S3Queue`/`AzureQueue` skipping a file forever after losing the race for its `processing` node in Keeper: if the other processor released the file without committing it, the file was not retried by this table until the server restart. ### Description Found during the review of https://github.com/ClickHouse/ClickHouse/pull/108977 (the regression test there has to accept losing one file for exactly this reason). When `ObjectStorageQueueIFileMetadata::trySetProcessing` fails because the `processing` node in Keeper already exists, that node belongs to another processor - another server, or another table on the same server. `afterSetProcessing` then updated the local `FileStatus` to `Processing`, and since the non-processable checks in `trySetProcessing` and `prepareSetProcessingRequests` treat `Processing` as terminal, every later attempt short-circuited on that cached state without ever rechecking Keeper. Unlike `Processed` and `Failed`, `Processing` is not backed by a persistent Keeper node: the foreign processor can release the file without committing it - it can die, or fail the file and reset the node. Stale `processing` node cleanup does not evict `local_file_statuses` either; only the processed/failed node cleanup (`ObjectStorageQueueMetadata.cpp`) does. So the file stayed skipped by this table until its file status was evicted from the cache or the server was restarted. The cached `Processing` state is now remembered as a timestamped observation of the foreign `processing` node (`FileStatus::onProcessingByAnotherProcessor`): it is still reported in `system.s3queue`, and it is respected - the file is skipped without touching Keeper - while the observation is fresh (the new table setting `foreign_processing_node_cache_ttl_seconds`, 5 minutes by default; zero means to always check Keeper). After that, the next attempt probes Keeper again and refreshes the observation if the node is still there, so a file released without a commit is retried within the timeout, while a file being processed by another server costs at most one Keeper probe per timeout instead of one per listing pass. As soon as the foreign processor commits the file, the next probe fails on the `processed` (or `failed`) node instead and the state becomes terminal again. `afterSetProcessing` also keeps the cached state untouched when it is already `Processing` and not marked as foreign: in that case the node belongs to a concurrent local processor sharing this `FileStatus` (tables on the same server using the same Keeper path, insert threads), and its owner updates the state on commit. The cached observation participates in the listing pre-filter (`FileIterator::filterProcessableFiles`) as well: a file with a fresh foreign-`Processing` observation is dropped from the batch before the `processed`/`failed` multi-read is built, so the fresh observation avoids Keeper requests on the listing path too. The terminal states which that pre-filter does discover in Keeper are written back into the cached `FileStatus` (unless the state is owned by an active local processor, whose owner updates it on commit), so `system.s3queue_metadata_cache` follows Keeper instead of keeping a stale `Processing` after another processor has committed the file. The write-back replaces the whole cached record, not only the status: the per-attempt data of an abandoned local attempt is cleared, and a file failed by another processor carries the exception and retries from the `failed` node. The write-back is skipped when the cached state already equals the discovered terminal state: such a record describes a finished local attempt (its `rows_processed` and timings must survive relistings). A cached `Failed`, however, also describes retriable local attempts (`retries < loading_retries`), so it is kept only when its retry count matches the `failed` node payload: when another processor exhausts the retries after a retriable local failure, the cached record follows the terminal node instead of keeping the stale local exception. The set-processing probe (`trySetProcessing`/`prepareSetProcessingRequests`) follows the same contract: when it discovers a `processed`/`failed` node (which can appear after the pre-filter ran), it returns the metadata of that node and `afterSetProcessing` refreshes the whole cached record with the same guards, instead of flipping only the status. A skipped file is not simply dropped from the current listing pass: the file iterator keeps it in a recheck list (also filled when `trySetProcessing`/`prepareSetProcessingRequests` observe a foreign `processing` node), and every batch boundary of the pass takes the files whose observation has expired and runs them through the regular filtering. Within a long listing pass the TTL is therefore honored with batch granularity; files whose observation is still fresh when the listing is exhausted are dropped with the iterator, because the observation timestamps live in the shared file status cache and the next pass re-queues them with the original deadlines. The TTL bounds the retry latency even on an otherwise idle queue, where the polling backoff after zero-row cycles can far exceed it: the streaming task schedules its next cycle no later than the earliest pending recheck deadline. In `Ordered` mode a foreign-held file also blocks the later files of its ordering domain (the scope of one `processed` pointer: a bucket, and a partition within it when partitioning is used) for the current listing pass. Without this, committing a later file advances the `processed` pointer past the held file, and the next listing pass drops it as already processed - losing it forever if the foreign processor never commits it. The file iterator records foreign-held files per ordering domain and drops the later files of the domain both at the listing pre-filter and before handing a file out (a file returned for retry releases its `processing` node); they are re-listed by the next pass. A held file stops blocking its domain as soon as this server wins its `processing` node or a terminal state for it is discovered in Keeper. With several processing threads sharing one ordering domain (`buckets = 1`), the block alone is not enough: it is recorded only when the set-processing attempt of the held file fails, and a later file handed out to another thread before that could still be committed first. The file iterator therefore registers a file whose set-processing outcome is not yet known at hand-out time (following the hand-out order), and a later file of the domain waits until the outcomes of the smaller files are known before starting its own attempt - so the set-processing attempts of one domain serialize for the duration of one Keeper round trip, and a foreign-held discovery blocks the later files before any of them starts processing. `foreign_processing_node_cache_ttl_seconds` is a per-table setting: `ObjectStorageQueueMetadataFactory` shares a single `ObjectStorageQueueMetadata` between all tables with the same `keeper_path`, so the value is not kept there - it travels from `StorageObjectStorageQueue` through `FileIterator` to `ObjectStorageQueueMetadata::getFileMetadata`, and each table uses (and reports in `system.s3_queue_settings`) the value from its own DDL. The setting can be changed on a live table with `ALTER TABLE ... MODIFY SETTING`: the storage keeps the value in an atomic member which the file iterators read through a reference, so the new value (for example, zero, to get a stuck file retried immediately) applies to the running streaming task without recreating the table. Tests: `tests/integration/test_storage_s3_queue/test_foreign_processing_node.py` emulates the foreign processor with a real `processing` node in Keeper, checks that the file is not committed while that node exists, removes it, and asserts that the file is then processed - without the fix the count stays one short. A second test in the same file creates two tables sharing one `keeper_path` with different values of the setting and checks both the reported value and the retry window. A third test emulates the foreign processor committing one file and failing another, and asserts that `system.s3queue_metadata_cache` reports `Processed` (respectively `Failed`, with the exception of the processor which failed the file) instead of a stale `Processing`, and that neither file is ingested by this table. A fourth test checks that `ALTER TABLE ... MODIFY SETTING` shortens the retry window of an already-running table. A fifth test repeats the retry scenario in `ordered` mode, with the foreign processor holding the lexicographically greatest file, asserting that the max processed path does not swallow the skipped file. A sixth test holds the only file of the queue with a two-minute polling backoff configured and asserts that it is retried within the TTL after its holder releases it, not after the backoff. A seventh test parks the `ordered` set-processing attempt at a failpoint between the initial state read and the multi request, fails the file from a fake foreign processor inside that window, and asserts that the cached record carries the exception of the `failed` node instead of an empty one. An eighth test feeds the table a file it cannot parse (a retriable local failure), writes the terminal `failed` node from a fake foreign processor, and asserts that the cached record follows its payload. A ninth test holds a file in the middle of the ordering domain in `ordered` mode and asserts that the later files wait for it: only the files before it are processed until the foreign `processing` node disappears, then the held file and the files after it follow. A tenth test runs two processing threads over one ordering domain, parks the set-processing attempt of the smallest file at a failpoint (the file is held by a fake foreign processor), and asserts that the other thread does not ingest the later files while the outcome of the smallest file is unknown, and that no file is lost after the holder releases it. `src/Storages/ObjectStorageQueue/tests/gtest_file_status_foreign_processing.cpp` covers the shared `FileStatus` state machine, including the same-server contention case and the reset of the per-attempt data (exception, processed rows, processing end time) of a previous local attempt when the file becomes observed as processed by another processor, and the whole-record refresh on a terminal node discovered by the set-processing probe. The members added by the fix are referenced from `if constexpr (requires ...)` branches, so the gtest also compiles at the merge base (where it fails at runtime), which is what the `Bugfix validation (unit tests)` job checks. Related: https://github.com/ClickHouse/ClickHouse/pull/108977",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112313",
          "createdAt": "2026-07-28T16:32:50Z",
          "updatedAt": "2026-08-13T12:05:35Z",
          "timestamp": "2026-08-13T12:05:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6e562133f1e327bce764",
        "signalId": "github:ClickHouse/ClickHouse:issue:113741",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113741",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Custom-key parallel replicas over a Merge table with a Distributed child: children are offloaded to a finalized stage (CANNOT_CONVERT_TYPE, logical error in GroupingAggregatedTransform)",
          "text": "🕵 Reading a `Merge` table (or the `merge` table function) with custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_sampling'` or `'custom_key_range'`) fails with an exception when the common processing stage of the children is `WithMergeableState` — for example, when one of the underlying tables is a `Distributed` table. ```sql CREATE TABLE t_mrg_ck_1 (k UInt64) ENGINE = MergeTree ORDER BY k AS SELECT number FROM numbers(100000); CREATE TABLE t_mrg_ck_2 (k UInt64) ENGINE = MergeTree ORDER BY k AS SELECT number FROM numbers(100000); CREATE TABLE t_mrg_ck_3 (k UInt64) ENGINE = Distributed('test_shard_localhost', currentDatabase(), 't_mrg_ck_1'); SET enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_sampling', parallel_replicas_custom_key = 'k'; SELECT count() FROM merge(currentDatabase(), '^t_mrg_ck_'); ``` ``` Code: 70. DB::Exception: Conversion from UInt64 to AggregateFunction(count) is not supported: while converting source column `count()` to destination column `count()`: Child table: default.t_mrg_ck_1. (CANNOT_CONVERT_TYPE) ``` Depending on the aggregate function, it instead trips an assertion in the pipeline — an exception in debug and sanitizer builds (found by the AST fuzzer, STID `3970-479a`): ```sql SELECT 47, quantileExactInclusive(visibleWidth(['1', '2'])) IGNORE NULLS FROM merge(currentDatabase(), '^t_mrg_ck_') GROUP BY ALL LIMIT 973; ``` ``` Logical error: 'Chunk should have AggregatedChunkInfo/ChunkInfoWithAllocatedBytes in GroupingAggregatedTransform.'. ``` **Root cause.** The `Distributed` child reports `WithMergeableState` from its `getQueryProcessingStage`, so `StorageMerge::getQueryProcessingStage` sets the common stage of all children to `WithMergeableState`, and `ReadFromMerge::createPlanForTable` plans each `MergeTree` child through an interpreter with `SelectQueryOptions(WithMergeableState)`. Inside that child interpreter, the custom-key branch of `PlannerJoinTree` (`src/Planner/PlannerJoinTree.cpp`, the `canUseParallelReplicasCustomKey` block) offloads the child query to the replicas at the hard-coded stage `WithMergeableStateAfterAggregationAndLimit`, ignoring the requested `to_stage`. The child plan therefore produces finalized rows (post-aggregation, post-`LIMIT`) where the parent `ReadFromMerge` pipeline expects partial aggregation states: `convertAndFilterSourceStream` throws `CANNOT_CONVERT_TYPE` when the finalized type differs from the state type, and when the types coincide structurally, the chunks without `AggregatedChunkInfo` reach `GroupingAggregatedTransform` and trip the assertion. Only the analyzer path is affected (`enable_analyzer = 0` returns correct results). Reproduced on current master (verified on a binary with no unrelated changes); found by the targeted AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/110972 (which is unrelated: the failure reproduces without `parallel_replicas_allow_merge_tables`, a setting that does not exist on master), report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110972&sha=ac0a584ce316ace31f4dbc16b38e1262a2344751&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29 The custom-key offload should be skipped when the requested `to_stage` is below `WithMergeableStateAfterAggregationAndLimit`: a plan that must stop at a partial stage cannot accept a finalized remote read. Related: https://github.com/ClickHouse/ClickHouse/pull/110972",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113741",
          "createdAt": "2026-08-06T23:14:04Z",
          "updatedAt": "2026-08-13T12:04:37Z",
          "timestamp": "2026-08-13T12:04:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "comp-distributed",
            "potential bug",
            "comp-parallel-replicas"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e45d8f19d4c3eb111787",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108329",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108329",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Bump `libssh` to 0.12.0",
          "text": "Update the `contrib/libssh` submodule to the upstream release `libssh-0.12.0` (from an upstream master snapshot, `libssh-0.11.0-368-g47305a2f`), tracked via the `ClickHouse/libssh-0.12.0` branch of the fork mirror. Both the old and new pins are clean upstream commits, so no ClickHouse patches needed to be carried forward. Build integration changes in `contrib/libssh-cmake`: - Bump the `libssh_VERSION_*` variables to 0.12.0. - Add the new ML-KEM sources. In 0.12.0, `kex.c` references `ssh_client_hybrid_mlkem_remove_callbacks` from an unguarded switch case, so `hybrid_mlkem.c` and `mlkem.c` are now mandatory. ClickHouse bundles OpenSSL 3.5, which provides the EVP ML-KEM API, so the OpenSSL backend `mlkem_crypto.c` is used (matching upstream's `OPENSSL_VERSION >= 3.5.0` logic). - In every per-platform `config.h`: define `GLOBAL_CONF_DIR` (now used unconditionally via string concatenation in `options.c`), and enable `HAVE_OPENSSL_MLKEM` / `HAVE_MLKEM1024` to select the OpenSSL ML-KEM backend and the ML-KEM-1024 variants. ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Update `libssh` to 0.12.0. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.1.24` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108329",
          "createdAt": "2026-06-23T20:48:38Z",
          "updatedAt": "2026-08-13T12:00:33Z",
          "timestamp": "2026-08-13T12:00:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-build",
            "submodule changed",
            "pr-synced-to-cloud"
          ],
          "author": "thevar1able",
          "state": "closed",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:20bd48c12050789bb731",
        "signalId": "github:ClickHouse/ClickHouse:issue:113711",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113711",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "MATERIALIZED CTE is not materialized when queried through a VIEW",
          "text": "### Company or project name _No response_ ### Describe the unexpected behaviour ClickHouse runs CTE twice if we select from a view based on a query with MATERIALIZED CTE. ### Which ClickHouse versions are affected? Latest (26.7) ### How to reproduce https://fiddle.clickhouse.com/eb4a0c4a-7efa-4554-a3a6-67c964162b15 ```sql -- 1. Create a sample table CREATE TABLE test_cte ( customer_id UInt32, amount Decimal(10, 2) ) ENGINE = MergeTree ORDER BY customer_id; -- 2. Insert sample data INSERT INTO test_cte (customer_id, amount) VALUES (1, 500.00), (1, 750.00), (2, 200.00), (2, 300.00), (3, 1500.00); --3. CREATE VIEW with MATERIALIZED CTE CREATE VIEW view_cte AS WITH summary AS MATERIALIZED ( SELECT customer_id, sum(amount) AS total_amount FROM test_cte GROUP BY customer_id ) SELECT * FROM summary AS s1 INNER JOIN summary AS s2 ON s1.customer_id = s2.customer_id SETTINGS enable_materialized_cte = 1, final = 1; -- 4. Check usage of MATERIALIZED CTE -- raw query use MATERIALIZED ! EXPLAIN indexes=1 WITH summary AS MATERIALIZED ( SELECT customer_id, sum(amount) AS total_amount FROM test_cte GROUP BY customer_id ) SELECT * FROM summary AS s1 INNER JOIN summary AS s2 ON s1.customer_id = s2.customer_id SETTINGS enable_materialized_cte = 1, final= 1; MaterializingCTEs (Materialize CTEs before main query execution) ... └──MaterializingCTE (Materializing CTE: summary) └──Aggregating │ Keys: customer_id │ Aggregates: sum(amount) │ Skip merging: 0 └──ReadFromMergeTree (default.test_cte) ... -- select from the view => NO Materialized CTE EXPLAIN indexes=1 SELECT * FROM view_cte; Join (JOIN FillRightFirst) ... ├──Aggregating │ │ Keys: customer_id │ │ Aggregates: sum(amount) │ │ Skip merging: 0 │ └──ReadFromMergeTree (default.test_cte) ... └──BuildRuntimeFilter (Build runtime join filter on customer_id) │ Filter id: RF1 │ Source table: default.test_cte └──Aggregating │ Keys: customer_id │ Aggregates: sum(amount) │ Skip merging: 0 └──ReadFromMergeTree (default.test_cte) ... ``` ### Expected behavior MATERIALIZED CTE is materialized when queried through a VIEW ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113711",
          "createdAt": "2026-08-06T18:31:10Z",
          "updatedAt": "2026-08-13T11:59:13Z",
          "timestamp": "2026-08-13T11:59:13Z",
          "metrics": {
            "reactions": 2,
            "comments": 2
          },
          "labels": [
            "bug",
            "experimental feature"
          ],
          "author": "SaltTan",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f2311de7fccb2d25d915",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112650",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112650",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject WITH FILL bounds that do not fit the ORDER BY column type",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/109454 Related: https://github.com/ClickHouse/ClickHouse/issues/109216 --> `WITH FILL FROM`/`TO` values are converted to a type wide enough for the arithmetic - `Int64` for every integer column type - while the generated values are written into a column of the column's own type, which truncates whatever does not fit its range. A truncated value wraps around, so the filled stream stops being monotonic while the query plan keeps claiming that it is still sorted by the fill columns. `DISTINCT` in order relies on that claim and reads the stream as a sequence of sorted runs, so it deduplicates within wrong ranges: ```sql SELECT count() FROM (SELECT DISTINCT x, s FROM (SELECT toUInt8(5) AS x, 'Hello' AS s ORDER BY x ASC WITH FILL FROM 1 TO 1025)); ``` returns `1024` instead of `257` in a release build, because `WITH FILL FROM 1 TO 1025` over a `UInt8` column generates `1..255, 0, 1..255, 0, ...`. In a debug or sanitizer build the same query aborts in `DistinctSortedStreamTransform` with `Equal values are not contiguous within the range assumed to be sorted`, which is what the AST fuzzer hit: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109454&sha=b0030d87c1f4e4b31d2766c17286e42492a2da94&name_0=PR&name_1=AST%20fuzzer%20%28amd_msan%29 https://github.com/ClickHouse/ClickHouse/pull/109454 The fuzzer query is unrelated to that pull request; the abort reproduces on `master`: ```sql SELECT DISTINCT x, isZeroOrNull(materialize(true)), s FROM ( SELECT 5 AS x, 'Hello' AS s ORDER BY x ASC NULLS LAST WITH FILL FROM 1 TO 10 INTERPOLATE (`s` AS concat(s, 'A')) LIMIT 1048576 UNION ALL SELECT 5 AS x, 'Hello' AS s ORDER BY x ASC NULLS LAST WITH FILL FROM 1 TO 1025 INTERPOLATE (`s` AS concatAssumeInjective(s, 'A')) LIMIT 1048576 ) ORDER BY s ASC; ``` ### Fix Reject `FROM`/`TO` values that cannot be represented in the type of the `ORDER BY` column with `INVALID_WITH_FILL_EXPRESSION`, next to the existing check that rejects negative bounds for an unsigned column type. This matches how the equivalent `Date`/`DateTime` bound mismatch is already rejected (https://github.com/ClickHouse/ClickHouse/issues/30421). `TO` is an exclusive bound, so it may still be outside of the range as long as the values that filling actually generates fit: `WITH FILL FROM 0 TO 256` over a `UInt8` column keeps generating `0..255`, and so does `WITH FILL FROM 0 TO 257 STEP 3`, which stops at `255`. The last generated value is known up front only when `FROM` and a plain numeric `STEP` are given. Without `FROM` the sequence is anchored at a data value, so which values are generated is known only at execution time: `WITH FILL TO 257 STEP 3` over a `UInt8` column stops at `254` when the data ends at `11` but reaches `256` when it ends at `13`. Still, whenever filling generates anything at all, its last value lands within one step before `TO`, so such a bound is rejected only when even the value one whole step away from `TO` (clamped towards zero, so that anchors close enough to `TO` to generate nothing keep being accepted) does not fit the column type - that is, only when filling provably wraps for every possible anchor. An `INTERVAL` step performs its calendar arithmetic in the column's own native type, so unlike a plain numeric step it wraps around within the column domain and can never reach a `TO` outside of it: filling would generate wrapped-around values forever (it does on current master, e.g. `SELECT toDate(0) AS d ORDER BY d ASC WITH FILL FROM toDate(0) TO 70000 STEP INTERVAL 100 YEAR` never terminates). Such a `TO` is therefore rejected regardless of `FROM` - but only when it is out of range in the fill direction. A `TO` out of range against the fill direction (e.g. `ORDER BY d DESC WITH FILL TO 70000 STEP INTERVAL -1 YEAR` over a `Date` column) is a guaranteed no-op instead: every possible anchor is already past it, filling never takes a single step, and the bound is accepted - the same clamp towards zero as in the numeric case. `STALENESS`, which is allowed only without `FROM`, replaces `TO` as the effective bound whenever it comes first, so the sequence can stop arbitrarily far below `TO` and such a `TO` is not checked either. Float and Decimal bounds are not checked: they saturate instead of wrapping, and a `Float64` bound is generally inexact in a `Float32` column, so requiring exact representability there would reject ordinary queries. On top of the storage range, the calendar arithmetic of an `INTERVAL` step clamps at the boundaries of the representable calendar - the `[0000-01-01, 9999-12-31]` window of `DateLUTImpl`, taken in the local civil calendar of the column's time zone, so for `DateTime64` the boundary in raw ticks shifts by the UTC offset (the last reachable second of a `DateTime64(0, 'Etc/GMT-14')` column is `253402250399`, not the UTC `253402300799`) - and for `Date32` and `DateTime64` that window is strictly narrower than the storage type. A `TO` beyond the calendar boundary in the fill direction fits the storage type but can never be reached: the filling keeps generating the clamped boundary value forever (it does on current master, e.g. `SELECT toDate32('9999-12-31') AS d ORDER BY d ASC WITH FILL TO 3000000 STEP INTERVAL 1 YEAR` never terminates). Such a `TO` is rejected against the calendar limits, scale-aware for `DateTime64`. For `Date` the calendar clamp coincides with the `UInt16` storage boundary, and `DateTime` wraps within `UInt32`, where any in-range `TO` stays reachable from some anchor, so the storage check covers those two. Beyond reachability, for `Date32` and `DateTime64` the values between the calendar boundary and the boundary of the storage type are invalid in themselves: no conversion produces them (they all clamp at the calendar boundary), yet a `FROM` bound in that gap is materialized into the column as is and serialized as the clamped boundary date - a spurious duplicate of the genuine boundary value next to it (e.g. `SELECT d FROM (SELECT toDate32('2000-01-01') AS d ORDER BY d ASC WITH FILL FROM -719529 TO -719528 STEP INTERVAL 1 YEAR)` on current master returns a `0000-01-01` row holding the day number `-719529`, which is not `0000-01-01`). The representability check therefore tests bounds of these two types against the calendar window (in the local civil calendar of the column's time zone for `DateTime64`), not just the storage range: an out-of-calendar `FROM` is rejected for any kind of step - like a `FROM` out of the storage range already is - and a numeric-step `TO` is rejected when the last value generated under it provably lands in the gap, under the same rules as the storage range. The last-generated-value computation covers the `Decimal64`-carried `DateTime64` bounds as well as the `Int64`-carried types: a numeric step over `DateTime64` advances raw ticks of the column's scale, and ticks beyond the calendar do not wrap but are equally invalid (they all serialize as the clamped boundary date). Finally, the calendar clamp makes an `INTERVAL` step able to stagnate: adding the interval to a value whose result would leave the calendar returns the value unchanged, so the sequence can stop advancing strictly below a perfectly representable `TO` and never terminate (e.g. `WITH FILL FROM toDateTime64('9999-06-01 00:00:00', 0, 'UTC') TO toDateTime64('9999-12-31 00:00:00', 0, 'UTC') STEP INTERVAL 1 YEAR` hangs on current master). With an explicit `FROM` the whole sequence is known up front - `FillingRow::next` advances it by one application of the step function at a time - so it is walked at construction time with the same step function, under a bounded budget (65536 steps), and rejected when it provably stagnates before reaching `TO`. The walk mirrors the runtime step exactly, so it can never misjudge a terminating sequence; a fill whose stagnation lies beyond the budget (a fine-grained step over a huge span) is accepted as before. Note that a query that previously returned wrapped-around values now gets an error instead. There is no in-tree test with such bounds. ### Not fixed here `WITH FILL` has other routes to the same wraparound, all pre-existing and all producing garbage values rather than tripping the sortedness assertion in the shapes I could build. They are data-dependent, so they need a different fix: ```sql -- the INTERVAL step function itself wraps in the column type SELECT groupArray(d) FROM (SELECT toDate('2149-06-01') AS d ORDER BY d ASC WITH FILL FROM toDate('2149-06-01') TO toDate('2149-06-06') STEP INTERVAL 1 YEAR); -- ['2149-06-01','1970-12-26','1971-12-26',...] -- the STALENESS border is computed from a data value and leaves the range SELECT groupArray(x) FROM (SELECT toUInt8(250) AS x ORDER BY x ASC WITH FILL STALENESS 20); -- [250,251,252,253,254,255,0,1,...,13] -- an out-of-range TO without FROM is rejected only when it wraps for every possible anchor; -- when only some anchors wrap, the wrapping ones still do so at execution time SELECT groupArray(x) FROM (SELECT toUInt8(13) AS x ORDER BY x ASC WITH FILL TO 257 STEP 3); -- [13,16,...,253,0] -- the INTERVAL step function returns its input unchanged when the result would leave the representable -- calendar, so an anchor near the boundary stagnates below an in-range TO and the filling never terminates; -- without FROM, which anchors stagnate depends on the data and the step (with an explicit FROM this shape -- is now rejected up front, unless the stagnation lies beyond the bounded walk budget of 65536 steps) SELECT * FROM (SELECT toDateTime64('9999-06-01 00:00:00', 0, 'UTC') AS t ORDER BY t ASC WITH FILL TO toDateTime64('9999-12-31 00:00:00', 0, 'UTC') STEP INTERVAL 1 YEAR); -- hangs: 9999-06-01 + 1 year would be out of range, so the step keeps returning 9999-06-01 -- the same shape over Date32, reachable since https://github.com/ClickHouse/ClickHouse/pull/111534 extended -- the parsed range of Date32 to the whole calendar (before that, such an anchor clamped to 2299-12-31) SELECT * FROM (SELECT toDate32('9999-06-01') AS d ORDER BY d ASC WITH FILL TO 2932896 STEP INTERVAL 1 YEAR); -- hangs the same way; with `FROM toDate32('9999-06-01')` added it is now rejected up front -- the DateTime INTERVAL step arithmetic wraps within UInt32, so a TO near the top of the storage range can be -- unreachable from a given anchor while staying reachable from others; the wrapped sequence cycles instead of -- stagnating, which the construction-time walk does not detect even with an explicit FROM SELECT * FROM (SELECT toDateTime('2106-01-01 00:00:00', 'UTC') AS t ORDER BY t ASC WITH FILL TO 4294967295 STEP INTERVAL 100 YEAR); -- hangs: 2106 + 100 years wraps to 2069 and cycles below TO forever ``` ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong `DISTINCT` results (and a logical error `Equal values are not contiguous within the range assumed to be sorted` in debug builds) for `ORDER BY ... WITH FILL FROM/TO` when a bound does not fit the type of the `ORDER BY` column: the generated values were silently truncated and wrapped around, so the filled stream was no longer sorted. An out-of-range `FROM` is now rejected with `INVALID_WITH_FILL_EXPRESSION`, and an out-of-range `TO` is rejected whenever the values that filling generates provably wrap the column type for every possible starting point - including always under an `INTERVAL` step for a `TO` out of range in the fill direction, whose calendar arithmetic stays within the column domain and can never reach such a `TO`, making the filling non-terminating (a `TO` out of range against the fill direction is a guaranteed no-op and stays accepted). For `Date32` and `DateTime64`, the bounds are additionally checked against the representable calendar (`[0000-01-01, 9999-12-31]` in the column's time zone), which is narrower than the storage type: values in between are invalid - everything else clamps at the calendar boundary, and filling materialized them as is, serialized as a spurious duplicate of the boundary date - so an out-of-calendar `FROM` is rejected for any kind of step, an `INTERVAL`-step `TO` is rejected when it is beyond the calendar in the fill direction (the clamping calendar arithmetic can never reach it), and a numeric-step `TO` over `Date32` or `DateTime64` is rejected when the last generated value provably lands out of the calendar. An `INTERVAL`-step fill with an explicit `FROM` is additionally rejected when the sequence provably stagnates before reaching `TO` (the calendar clamp makes the step return its input unchanged near the boundary, so the filling would never terminate). Data-dependent wraparound or stagnation (e.g. via `STALENESS`, numeric-step starting points that wrap only at execution time, or calendar-boundary anchors without `FROM` that make the `INTERVAL` step stagnate below an in-range `TO`) is not covered by this check.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112650",
          "createdAt": "2026-07-30T19:13:54Z",
          "updatedAt": "2026-08-13T11:58:37Z",
          "timestamp": "2026-08-13T11:58:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7902f725fcb392209107",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114043",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114043",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a full-featured AI agent to the client (the `?` command)",
          "text": "Turns the primitive `??` \"generate one SQL query\" helper in the client into a full interactive AI agent, available with a single `?` (`??` is kept as an alias). The agent works over a live connection: it is given the context of the recent queries in the session (with results truncated to the first and last lines to save tokens, plus error messages) and can use tools to read the query history from `system.user_query_log`, inspect the schema (`SHOW`/`DESCRIBE`/`SHOW CREATE`), consult the embedded documentation (`system.documentation`, the same source as the `help` command), run read-only queries without confirmation under a sandbox (`readonly = 1`, a 30-second and a 10 GiB limit, no table functions reaching outside of the server's tables; if the session forbids applying these limits, the query fails instead of running without them), and run any other query after asking the user for confirmation. Queries it runs are echoed and executed on the user's connection and displayed exactly as if the user had typed them; it prints its thoughts and tool calls as it works, and the conversation keeps its context across `?` invocations within a session. When no client-side AI provider is configured (neither the `ai` section of the client configuration nor the `OPENAI_API_KEY`/`ANTHROPIC_API_KEY` environment variables), the agent falls back to the server-side `aiGenerate` function of the connected server (or of `clickhouse-local`) when it has default credentials configured for it (`ai_function_text_default_credentials`), so it works with no client-side setup when the administrator has already enabled AI functions. Implementation lives in `src/Client/AI/`: the agent loop (`AIAgent`), two model backends (`AIAgentTransport`: native tool-calling via `ai-sdk-cpp`, and a text tool-call protocol on top of `aiGenerate`), the tools (`AIAgentTools`), the read-only statement allowlist (`AIQueryValidation`), and the recent-query context buffer (`QueryContextBuffer`). Unit tests cover the tool-call protocol parsing, the read-only validation, and the context buffer; the end-to-end loop (both backends and the confirmation gate) was validated by driving `clickhouse-local` against a mock OpenAI endpoint. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The embedded AI assistant of `clickhouse-client` and `clickhouse-local` is now a full agent, invoked with a single `?`. It sees the recent queries and their results, explores the schema and the documentation, runs read-only queries on its own (and other queries with confirmation) displayed as if you typed them, and can use the server-side `aiGenerate` function when no client-side AI provider is configured. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114043",
          "createdAt": "2026-08-09T15:33:43Z",
          "updatedAt": "2026-08-13T11:57:26Z",
          "timestamp": "2026-08-13T11:57:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fd5cc46ae6b065343356",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114316",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114316",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use the vector similarity index for integer reference vectors",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112233 Related: https://github.com/ClickHouse/ClickHouse/issues/114291 The reference vector of an ANN query is extracted only when its array type is `Float64`, `Float32` or `BFloat16` and every element is a `Float64` field. An integer literal such as `[1, 2]` is typed `Array(UInt8)`, so `tryUseVectorSearch` bails out and the query silently falls back to a brute-force scan over the whole table, although `[1, 2]` denotes the same point as `[1.0, 2.0]` and `L2Distance` accepts it. `EXPLAIN indexes = 1` shows no `vector_similarity` entry and gives no hint why, so a one-character difference in a literal becomes a sharp performance cliff on large tables. Native integer arrays are now accepted as reference vectors and their elements are converted to `Float64`, which is the type the reference vector is stored in anyway. Added `02354_vector_search_bug112233`, covering unsigned, signed, mixed integer/float, and not-exactly-representable reference vectors, plus an equality check between the integer and float spellings of the same query. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Vector search queries now use the `vector_similarity` index when the reference vector is written as an integer array literal, e.g. `ORDER BY L2Distance(vec, [1, 2])`. Previously such queries silently fell back to a brute-force scan.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114316",
          "createdAt": "2026-08-11T12:52:06Z",
          "updatedAt": "2026-08-13T11:56:34Z",
          "timestamp": "2026-08-13T11:56:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "hamidr",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c6cbfadbcfdd82b100a0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114422",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114422",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Convert only SEMI JOIN to IN in the convertJoinToIn optimization",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/101698 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results with `query_plan_convert_join_to_in = 1`. The optimization replaced a JOIN with `key IN (subquery)` for `ALL` and `ANY` strictness, but `IN` only tests membership: `ALL INNER JOIN` lost rows on a duplicated right key, and `ANY INNER JOIN` returned extra rows on a duplicated left key. Only `SEMI JOIN` is converted now. Closes #101698. ### Description `tryConvertJoinToIn` rewrites a JOIN into `key IN (subquery)` over the left side, emitting each matching left row exactly once: semi-join semantics. The gate excluded strictnesses by name rather than requiring that property, so it admitted two that break it in opposite directions. - **`ALL`**, the default and the reported bug, emits `left_count(k) * right_count(k)` rows, so a duplicated right key multiplies left rows and `IN` loses them: the reproducer gives 3 off, 2 on. - **`ANY`** deduplicates the *left* side at the default `any_join_distinct_right_table_keys = 0`, so `IN` instead adds rows: 2 off, 5 on. The gate now admits only `Semi`; `Left` joins the kind gate since `SEMI` needs `LEFT`/`RIGHT`. Five more declines, each a divergence I measured off vs on: - a `Join` engine right side, whose declared kind and strictness the rewrite stops validating: on `Join(ANY, LEFT, id)`, `SEMI LEFT JOIN` threw `INCOMPATIBLE_TYPE_OF_JOIN` off, returned rows on; - an active `max_rows_in_join`, `max_bytes_in_join`, `max_rows_to_transfer` or `max_bytes_to_transfer`: the join bounds its stored right side, the set only its hash table, so a limit could stop being enforced. The setting is therefore inert in a profile bounding joins, as its description now says; - a key whose type has dynamic structure, which `IN` rejects; - a key-value prepared right side (dictionary, `EmbeddedRocksDB`, any `IKeyValueEntity`), which the planner probes by key: `DirectKeyValueJoin` off, `id in 2000000-element set` on. The `Join` engine check above now tests the whole `PreparedJoinStorage`, covering both; - a key expression consuming a left column the projection needs, e.g. `ON arrayJoin(l.tags) = r.tag` selecting `l.tags`, throwing `NOT_FOUND_COLUMN_IN_BLOCK`. #104809 fixes that hole from the `ALL` side, so I reuse its predicate; it also asserts an `ALL` join converts, which this PR disables. `ALL` being the default, the pass now fires for far fewer queries; it stays opt-in. All five existing tests drove this pass through `ALL` joins, so each would have gone silent; each now gets a converting shape and an oracle. Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/101698#issuecomment-5254250810",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114422",
          "createdAt": "2026-08-12T05:30:34Z",
          "updatedAt": "2026-08-13T11:55:53Z",
          "timestamp": "2026-08-13T11:55:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:425d506e8ef3e4024891",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114580",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114580",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: Duplicate TLS argument rejection and positional-arity stripping untested",
          "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #110615](https://github.com/ClickHouse/ClickHouse/pull/110615). That PR: (1) Adds TLS/SSL to every PostgreSQL integration: `sslmode` plus certificate/key either as server-local paths (`sslrootcert`/`sslcert`/`sslkey`, config-only) or as literal contents (`*_pem`, accepted from SQL, materialized into `TemporarySecretFile` and masked as secrets). New code: … **1. Duplicate TLS argument rejection and positional-arity stripping untested** `src/Storages/StoragePostgreSQL.cpp:726`, `src/Databases/PostgreSQL/DatabasePostgreSQL.cpp:573` **Risk:** `StoragePostgreSQL::extractSSLParamsFromArguments` strips trailing TLS `key = value` pairs and rejects repeats at `StoragePostgreSQL.cpp:726-727`; the stripped list then feeds the arity check at `DatabasePostgreSQL.cpp:573`. Risk if broken: a repeated `sslmode = 'require', sslmode = 'disable'` … **Unique vs PR tests:** 04820 formats queries with distinct TLS keys and checks masking plus path rejection; 04846 checks query-tree masking; test_postgresql_ssl exercises real handshakes through named collections. None repeats a TLS key (the `specified more than once` branch) and none combines the maximum positional … **Tags:** `-- Tags: no-fasttest` — `no-fasttest`: the PostgreSQL integration is not built in the fast test build; the PR's own 04820 uses the same tag cc @alexey-milovidov (author of #110615) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114580",
          "createdAt": "2026-08-13T03:32:59Z",
          "updatedAt": "2026-08-13T11:53:46Z",
          "timestamp": "2026-08-13T11:53:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ea3b09bfd3b8b4087a79",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113383",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113383",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Push tuple element predicates into Parquet and ORC subcolumn reads",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/112575 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Filter pushdown now works for `Tuple` subcolumns in Parquet and ORC files. A predicate such as `WHERE tup.1 = 555555` over `file()`, `s3()` or `url()` now prunes row groups and row index strides using the tuple element's own statistics instead of reading the whole file. ### Description Closes: #112575 Three independent defects, all needed for the reported query to prune. Parquet's reader itself was fine. **The analyzer never produced a subcolumn.** `StorageFile`/`StorageURL`/`StorageObjectStorage` return `false` from `supportsOptimizationToSubcolumns`, so `tupleElement(tup, 1)` was never rewritten to `tup.1`; the pass also accepted only `TableNode`, skipping table functions. That blanket `false` keeps #106147 fixed (`NOT_FOUND_COLUMN_IN_BLOCK` on `.null` in PREWHERE), so rather than flipping it this adds a narrow `supportsOptimizationToTupleElementSubcolumns` virtual defaulting to the existing one, with a `{Tuple, tupleElement}`-only allow-list. `04303_object_storage_prewhere_isnotnull_subcolumn` passes unmodified. Source identity also moves to the table-expression node: every `file()` resolves to the same `_table_function.file` ID, so two such sources shared a key. Accepting table functions generally also makes `format()`, `values()` and `view()` eligible for the other transformers; those storages already answer `supportsSubcolumns()`, so the default covers them. **ORC's search argument builder resolved top-level names only,** so any dotted name emitted `YES_NO_NULL` while the read path in the same file resolved them recursively. Resolving recursively also reaches the flattened-`Nested` descent, which rewrites the type it is given, so the builder keeps the key's own type: an array-typed predicate over a flattened `Nested` leaf is not pushed, since scalar element statistics cannot decide it. **ORC built its KeyCondition from the reader header,** which carries only the parent column, so the predicate degraded to `unknown`. ORC now passes `initKeyConditionOnce` a local copy extended with the tuple element paths the filter references; `FormatFilterInfo`, the Parquet call site and the reader header are untouched. Admission requires a named tuple at every level (unwrapping `Nullable`/`LowCardinality`/`Array`), which refuses Map `.keys`/`.values`: they use `SubstreamType::TupleElement` but have no per-element statistics. <details> <summary>Measurements (100k rows, one row group / stride per 10k)</summary> | Arm | master | this PR | |---|---|---| | ORC `WHERE tup.1 = 55555` | 200000 rows read | **20000** | | ORC `WHERE id = 55555` (control) | 20000 | 20000 | | Parquet `WHERE tup.1 = 55555` | 7 row groups / 0 pruned | **1 / 6** | | Parquet `WHERE id = 55555` (control) | 1 / 6 | 1 / 6 | Results identical in every arm. Refusal arms (Map `.keys`/`.values`, `.null`, `.size0`, unnamed tuple, `Array(Tuple)`, type-mismatch structure hint) return correct results with pruning off. Multi-level `tup.2.1` stays unpruned: correct, and a separate optimization. Regression sweep over `*functions_to_subcolumns*`, `*tuple_element*`, `*_parquet_*`, `*_orc_*` and the named pushdown tests, run on this build and on an unmodified master build for attribution: no regression attributable to this change. New tests are 50/50 green under randomized settings. </details> Also noticed, not touched here: reading a standalone dotted ORC tuple element with its inferred type returns column defaults, because `Nested::flatten` does not descend a `Nullable(Tuple)`. #109741 (open) rewrites that helper for the `Arrow` spelling.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113383",
          "createdAt": "2026-08-04T20:17:22Z",
          "updatedAt": "2026-08-13T11:51:25Z",
          "timestamp": "2026-08-13T11:51:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:384a43c87c3ff16de704",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114219",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114219",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.5: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114219",
          "createdAt": "2026-08-10T20:06:41Z",
          "updatedAt": "2026-08-13T11:49:07Z",
          "timestamp": "2026-08-13T11:49:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:83b0cbcea6f5d45f46b6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105848",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105848",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use text index for LIKE/ILIKE with ESCAPE",
          "text": "Follow-up to #99774 (qoega's LIKE ESCAPE clause). Before this change, a query like `... WHERE col LIKE pattern ESCAPE 'c'` silently bypassed the primary key (`KeyCondition`), the text index (`TYPE text`), and the bloom-filter text indexes (`ngrambf_v1`, `tokenbf_v1`, `sparse_grams`), falling back to a full scan even on tables with such an index defined or with `col` in the sorting key. Root cause: `LIKE pattern ESCAPE 'c'` is parsed into a 3-argument function call `like(col, pattern, escape_char)`. `MergeTreeIndexConditionText::traverseAtomNode`, `MergeTreeConditionBloomFilterText::extractAtomFromTree`, and `KeyCondition::extractAtomFromTree` only accepted 2-argument forms; any 3-arg `like`/`ilike` was rejected without ever reaching the LIKE handler. The execution layer in `FunctionsStringSearch::executeImpl` already calls `likePatternWithCustomEscapeToLikePattern` to fold the custom escape into standard backslash escapes before matching, so result correctness was never affected, only performance. Apply the same fold at index-condition time: - Text index (`MergeTreeIndexConditionText`): handles 3-arg `like` and `ilike` (the same arities `isSupportedFunction` allows for the index). - Bloom-filter text indexes (`MergeTreeConditionBloomFilterText`, for `ngrambf_v1` / `tokenbf_v1` / `sparse_grams`): handles 3-arg `like` and `notLike` (the like-family forms this index supports; `ilike` is not supported by this index type). - Primary key (`KeyCondition`): handles 3-arg `like` and `notLike` (the like-family forms present in `atom_map`). Case-insensitive `ilike`/`notILike` are not in `atom_map` and continue to fall back to row-level evaluation, same as the existing 2-arg behavior. In all paths the pattern must be a constant String and the escape argument a constant String of length 1; otherwise we bail out. Invalid escape sequences cause analysis to skip the index/key so row-level evaluation throws the same error the user would have seen before. The `nextInStringLike` tokenizer used by both text-index variants drops the backslash for an unknown escape `\\c` while row-level matching keeps it, so an unknown/trailing backslash is declined (the `likePatternHasUnknownBackslashEscape` guard) and falls back to row-level matching to avoid wrongly pruning a matching granule. In the bloom-filter path this guard is placed in the shared `traverseTreeEquals`, so it protects both the folded 3-argument form and the pre-existing 2-argument `like` / `notLike` / `mapContainsKeyLike` / `mapContainsValueLike` paths. Regression tests use `EXPLAIN indexes = 1` to assert that pruning now happens: - `04277_text_index_function_like_escape.sql` covers the `TYPE text` index for `LIKE`/`ILIKE`, both `splitByNonAlpha` and `array` tokenizers, and both the operator form and the functional form `like(a, p, 'c')`. - `04292_105885_primary_key_like_escape.sql` covers the primary key for `LIKE`/`NOT LIKE` (with a remaining wildcard so the perfect-prefix extraction succeeds) and the functional form. - `04351_bloom_filter_text_index_like_escape.sql` covers the `ngrambf_v1` / `tokenbf_v1` / `sparse_grams` text indexes for `LIKE`/`NOT LIKE` ESCAPE and the functional form, including a consumed escape character, `force_data_skipping_indices`, and the unknown/trailing-backslash decline for both the 2-argument and 3-argument forms. The tests FAIL on master without this change and PASS with it. Closes #105885 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `LIKE ... ESCAPE 'c'` and `ILIKE ... ESCAPE 'c'` predicates (added in #99774) now use `TYPE text` skip indexes to prune granules; `LIKE ... ESCAPE 'c'` and `NOT LIKE ... ESCAPE 'c'` additionally use the `ngrambf_v1`, `tokenbf_v1`, and `sparse_grams` skip indexes, and the primary key when the column is in the sorting key, instead of falling back to a full scan. ### Documentation entry for user-facing changes - [x] Documentation is unchanged (the existing text-index docs already describe LIKE-based pruning; this change only fills a coverage gap for the recently added ESCAPE clause).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105848",
          "createdAt": "2026-05-26T11:56:24Z",
          "updatedAt": "2026-08-13T11:49:00Z",
          "timestamp": "2026-08-13T11:49:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 38
          },
          "labels": [
            "pr-improvement",
            "can be tested",
            "hold"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "ahmadov",
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:637b71c328933e1d7c44",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113609",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113609",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a logical error comparing arrays whose element type is Nothing",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Comparing two arrays whose element types have no least supertype aborted the server with `Bad cast from type DB::ColumnNothing to DB::ColumnVector<char8_t>` when an aligned element position held a bare `Nothing`, for example `SELECT CAST([], 'Array(Nullable(Nothing))') > [[1]]`, or over real columns of `Array(Tuple(Nothing, UInt64))` and `Array(Tuple(Array(UInt8), Int64))`. The abort happened during constant folding as well as at execution, so analysis alone could kill the server. Root cause: `getReturnTypeImpl`'s array-no-supertype branch declared a plain `UInt8`, while the element comparator it later builds declares `Nothing` for a value-less position. Execution then `assert_cast`-ed that `ColumnNothing` to `ColumnUInt8`. A bare `Nothing` position is genuinely undecidable: it holds no values, a `Tuple` member has no null map covering it, and there is no length to tie-break on. The fix makes the declared type honest. A new per-side `containsUndecidableNothing` predicate rejects such a pair during analysis with `ILLEGAL_TYPE_OF_ARGUMENT`, and `compareGatheredElements` keeps a fail-closed guard for defence in depth. Each side is classified independently, because one side's null map must never decide a position on the other side. The predicate deliberately mirrors `FunctionsNullSafeCmp::containsNothing` and does not descend into `Array`/`Map`; deeper positions stay covered because the caller recurses once per array level. Two `Nothing` shapes remain decidable and keep answering, now correctly rather than crashing: `Nullable(Nothing)` (decided by its own null map) and `Array(Nothing)` (shares a supertype with `Array(T)`). Those answers were checked against the same comparison on a pair that does have a supertype, and match exactly in both operand orders. Introduced by #110245, which is on `master` only, so no released version is affected and no backport is needed. Tracked in #113640 (CI signature STID 1499-2747), which also collects the `Equality` sibling at `FunctionsComparison.h:1485`. <details> <summary>Validation</summary> Both directions, one command per binary (pre-fix and fixed builds of the same branch): | query | pre-fix | fixed | |---|---|---| | `SELECT CAST([],'Array(Nullable(Nothing))') = [[1]]` | abort, `Bad cast ... ColumnNothing ...` | `0` | | `SELECT CAST([],'Array(Nullable(Nothing))') > [[1]]` | abort, same message | `0` | | `SELECT [NULL] = [[1]]` | `Code: 44` | `0` | The values the fixed build returns were compared against the supertype path (`[1]` in place of `[[1]]`) and match exactly: `0 1 0 0 1 1` for the six operators, `0 1 1 1 0 0` for an empty aligned prefix, `0 1` for `isNotDistinctFrom`/`isDistinctFrom`, and `0 0` in the reversed operand order. Coverage: all eight operators the introducing PR added; `Nothing` direct, under `Nullable`, nested in `Tuple`, and under 1-, 2- and 3-deep `Array` wrappers; `Map(k, Nothing)` and `Array(Nothing)` as must-still-answer controls; empty and non-empty aligned prefixes. An empty range is rejected too, so validity never depends on the data. No existing assertion was weakened: all 47 pre-existing reference lines of `04549_array_comparison_bigger_types_nullable` are unchanged. 200/200 runs pass across `--test-runs 50 --order random` and `--test-runs 50 --no-random-settings`, and the feature's own eight-test suite has no failures. Nine mutation arms were run against the new tests; eight are caught, and the ninth only makes the secondary fail-closed guard unreachable, which no input can pin because the analysis-time check rejects every such pair first. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1319` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113609",
          "createdAt": "2026-08-06T03:27:50Z",
          "updatedAt": "2026-08-13T11:48:45Z",
          "timestamp": "2026-08-13T11:48:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "diegomestre2"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:54691da6e3b14d062bfe",
        "signalId": "github:ClickHouse/ClickHouse:issue:112569",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112569",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Request for office hours for new contributors.",
          "text": "### Company or project name N/A ### Question -- Is there any office hours for new contributors such that we can contribute to the project?",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112569",
          "createdAt": "2026-07-30T09:59:20Z",
          "updatedAt": "2026-08-13T11:46:42Z",
          "timestamp": "2026-08-13T11:46:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "question"
          ],
          "author": "harshil15999",
          "state": "closed",
          "assignees": [
            "antonkovalenko"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f397ccd8ac0443993dcc",
        "signalId": "github:ClickHouse/ClickHouse:issue:112241",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112241",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "A `WHERE` predicate on an `arrayJoin`'d column can be rewritten into `arrayFilter` so the array is filtered before expansion",
          "text": "### Company or project name ClickHouse customer ### Use case Queries that unnest one or more `Array` columns with `arrayJoin` (or `ARRAY JOIN`) and then select a few specific elements in `WHERE` are a very common pattern (tag/label arrays, event attribute arrays, key-value arrays stored as parallel arrays). Today ClickHouse first materializes the full expansion and only then applies the filter, so a query that ultimately returns one row per source row can transiently produce `length(A) * length(B) * length(C)` rows per source row. The work and the memory spent on those rows is unnecessary, and with several `arrayJoin` calls in one query it's multiplicative. Example: ```sql SELECT arrayJoin(A) AS a, arrayJoin(B) AS b, arrayJoin(C) AS c FROM t WHERE a = 'X-A' AND b = 'X-B' AND c = 'X-C'; ``` is equivalent to, but much slower and more memory hungry than, the manual rewrite: ```sql SELECT arrayJoin(arrayFilter(x -> x = 'X-A', A)) AS a, arrayJoin(arrayFilter(x -> x = 'X-B', B)) AS b, arrayJoin(arrayFilter(x -> x = 'X-C', C)) AS c FROM t; ``` ### Describe the solution you'd like Add an optimization that pushes a WHERE conjunct that constrains an arrayJoin'd column into an arrayFilter over the source array, i.e. rewrite `arrayJoin(expr) AS a ... WHERE f(a)` into `arrayJoin(arrayFilter(x -> f(x), expr)) AS a` and drop the conjunct from WHERE when it becomes redundant. Conditions under which the rewrite is applicable: 1. The predicate is a top-level conjunct of WHERE (an AND operand). OR across different array-join columns must not be rewritten. 2. The conjunct depends on exactly one arrayJoin'd column. Other columns it references must be row-constant (not produced by an arrayJoin); those are captured by the lambda, which arrayFilter already supports. A conjunct relating two different arrayJoin'd columns (e.g. a = b) cannot be rewritten. 3. The predicate must be deterministic and stateless (exclude rand, runningAccumulate, neighbor, and anything else whose result depends on evaluation order or position in the stream). 4. Multi-array `ARRAY JOIN a, b` is not supported ### Describe alternatives you've considered Writing `arrayJoin(arrayFilter(...))` by hand. This works and is the current workaround. ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112241",
          "createdAt": "2026-07-28T08:40:27Z",
          "updatedAt": "2026-08-13T11:45:58Z",
          "timestamp": "2026-08-13T11:45:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "feature"
          ],
          "author": "ayakovlev-clickhouse",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ee642e5c2ff2fc26d70d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114210",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114210",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use higher quality hash for nullable fixed-width keys in external aggregation",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): 64-bit hash function is used now for external nullable fixed-width aggregation methods (avoids collision disaster for very high cardinality aggregation). --- External aggregation merges partitions whose combined key count can far exceed 4 billion, which is why `mergeBlocks` re-aggregates the spilled stream with a method hashing over more than 32 bits rather than `HashCRC32`. `key64`, `keys128`, `keys256` and the serialized methods all have such a counterpart and are remapped to it; the nullable fixed-width methods never got one, so a `GROUP BY` on a nullable key that spills is still re-aggregated with a 32-bit hash and takes the collision blowup the remap exists to avoid. Add `nullable_key64_hash64`, `nullable_keys128_hash64` and `nullable_keys256_hash64`, and remap to them. The packed forms need no new data type: a nullable packed key carries its null map inside the key, so they reuse `AggregatedDataWithKeys128Hash64` / `…256Hash64` with `has_nullable_keys`. Only the single nullable key needs one, mirroring how `nullable_key64` wraps the `HashCRC32` map in `AggregationDataWithNullKey`. `nullable_key32` is deliberately left out, matching the existing choice not to remap `key32`: a 32-bit key space cannot exceed 4 billion distinct values.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114210",
          "createdAt": "2026-08-10T19:33:14Z",
          "updatedAt": "2026-08-13T11:45:24Z",
          "timestamp": "2026-08-13T11:45:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "nickitat",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f4d9aaec211ea503f5fe",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109594",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109594",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support `arrayExists` predicates for text-like indexes",
          "text": "Previously, `text`, `tokenbf_v1`, and `ngrambf_v1` indexes on `Array(String)` columns could not be used for predicates such as `arrayExists(x -> x LIKE '%needle%', arr)`. These predicates had to read every granule even though the index already stores tokens from the array elements. This PR lets index analysis use `arrayExists` lambdas where the lambda tests the element against a constant, for example `arrayExists(x -> x LIKE '%needle%', arr)` or `arrayExists(x -> x IN ('a', 'b'), arr)`. The index can then check the tokens for `arr` before reading rows. This is safe because any matching element must have added the required tokens to the index granule. The original `arrayExists` expression is still evaluated for rows that pass the index filter. The supported functions are listed explicitly for each index type. They include positive string predicates such as `equals`, `LIKE`, `ILIKE`, `startsWith`, `endsWith`, `match`, `multiSearchAny`, `hasToken`, and `IN`; the `text` index also supports `hasAnyTokens`, `hasAllTokens`, `hasPhrase`, `multiSearchAnyUTF8`, and `multiMatchAny`. Negative predicates such as `notEquals`, `NOT LIKE`, and `NOT IN` are not supported because they are not valid filters for empty arrays. For `text` indexes, functions whose result depends on tokenization, such as `hasToken`, are supported only when the index uses `splitByNonAlpha` without preprocessors or postprocessors. Other `arrayExists` lambdas are unchanged: they simply do not use these indexes. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `text`, `tokenbf_v1`, and `ngrambf_v1` indexes can now prune granules for predicates such as `arrayExists(x -> x LIKE '%needle%', arr)` on `Array(String)` columns.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109594",
          "createdAt": "2026-07-07T02:19:49Z",
          "updatedAt": "2026-08-13T11:45:00Z",
          "timestamp": "2026-08-13T11:45:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "EmeraldShift",
          "state": "open",
          "assignees": [
            "ahmadov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1c2f8a361380a954f442",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114420",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114420",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not apply DROP fault injection to CREATE OR REPLACE internal DROPs",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/110893 Related: https://github.com/ClickHouse/ClickHouse/pull/110971 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `CREATE OR REPLACE` leaking an internal `_tmp_replace_*` table, and `CREATE OR REPLACE VIEW` failing with `NOT_IMPLEMENTED`, when `ignore_drop_queries_probability` is set. The DROPs it issues internally are steps of one user statement, so DROP fault injection no longer applies to them. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/110893 Related: https://github.com/ClickHouse/ClickHouse/pull/110971 `CREATE OR REPLACE` builds the replacement under a temporary `_tmp_replace_*` name and publishes it by rename. It issues two DROPs internally: cleaning up that temporary table if the statement fails, and dropping the replaced table after the swap. Both took their context from `make_drop_context`, which did not mark it DDL-internal, so `InterpreterDropQuery` treated them as *user* DROPs and injected faults. Two things follow. If the storage keeps data on disk the DROP is skipped outright and a populated `_tmp_replace_*` table is stranded: listed by `SHOW TABLES`, holding its data, surviving a restart, one per statement. Otherwise (a materialized view, say) the DROP is rewritten to `TRUNCATE`, whose branch takes an exclusive lock under the outer statement's query id while that id already holds read locks, raising `RWLockImpl::getLock(): Cannot acquire exclusive lock while RWLock is already locked`. That fired twice on master, under [`Stress test (amd_tsan)`](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=d134c6da364233695c456f9c51fbbc521a4fbd5b&name_0=MasterCI&name_1=Stress%20test%20%28amd_tsan%29) and [`Stress test (arm_debug)`](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=1033651ae6f23b6c10f0023dd1247c6feb8bde4f&name_0=MasterCI&name_1=Stress%20test%20%28arm_debug%29). A third internal DROP is misclassified the same way. `CREATE OR REPLACE VIEW` uses that machinery only on `Atomic` or `Replicated`; other engines drop the view in place through `doCreateTable`. `StorageView` implements no `TRUNCATE`, so there the rewrite fails the statement with `NOT_IMPLEMENTED` and the stale view survives. The fix marks both contexts DDL-internal, as the sibling helper in the same function and `InterpreterDropQuery::executeDropQuery` already do; that asymmetry is why plain `CREATE ... AS SELECT` was immune and only REPLACE was exposed. User DROPs are still skipped, which the new test pins. The setting defaults to 0, so a default configuration is unaffected. The read lock the assertion reports is not identified; the fix does not depend on it, since without the rewrite nothing requests an exclusive lock. <details> <summary>Validation</summary> - New test `04796_create_or_replace_internal_drop_not_ignored` fails on master and passes with the fix, on every shape it covers. - All three sites covered by shape: `CREATE OR REPLACE TABLE ... AS SELECT` over MergeTree takes the skip path (stray `_tmp_replace_*` count accumulates 1, 2, 5, 6 on master); `CREATE OR REPLACE MATERIALIZED VIEW ... POPULATE` takes the rewrite-to-`TRUNCATE` path; `CREATE OR REPLACE VIEW` on a non-Atomic database fails on master with `Truncate is not supported by storage View`, leaving the stale definition readable. Failure path, success path (which leaks the *replaced* table), and repeated replaces all covered. - 50 randomized runs each, on the final binary, of the new test, `03013_ignore_drop_queries_probability`, `00916_create_or_replace_view` and `01866_view_persist_settings`: 200/200. Same for `04492`, `04524`, `04326`, `04328` on the first binary: 200/200. - A/B sweep of the 41 tests matching `create_or_replace`/`replace_table`/`tmp_replace`/`create_as_select`, plus a focused sweep of the 14 that reach the view site: each removes exactly one failure (the new test) and introduces none. - All 8 `Common.RWLock*` unit tests pass. - Replicated databases measured on both binaries across every reachable `CREATE OR REPLACE` shape: no change, no `INCORRECT_QUERY`. That path already marks the context internal (`DatabaseReplicatedTask::makeQueryContext`), so the fix is a no-op there. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1308` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114420",
          "createdAt": "2026-08-12T04:39:37Z",
          "updatedAt": "2026-08-13T11:44:46Z",
          "timestamp": "2026-08-13T11:44:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:819241366f6ef7b46c72",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108862",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108862",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use `HashSet` for aggregations without aggregates",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Aggregation queries without aggregate functions now use `HashSet`-based methods instead of `HashMap` (for supported key types). Speedups up to 1.8x times were observed. --- <img width=\"1015\" height=\"430\" alt=\"Screenshot 2026-07-03 at 00 34 25\" src=\"https://github.com/user-attachments/assets/15cc6a20-7568-49a4-94a5-854c7751ae02\" /> Further steps are `distinct` -> `group by` rewrite and key-columns-only (and perhaps semi) joins.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108862",
          "createdAt": "2026-06-29T22:01:12Z",
          "updatedAt": "2026-08-13T11:42:45Z",
          "timestamp": "2026-08-13T11:42:45Z",
          "metrics": {
            "reactions": 1,
            "comments": 7
          },
          "labels": [
            "pr-performance"
          ],
          "author": "nickitat",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a0c30f34abf00440817c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113248",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113248",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not run the server-side AST fuzzer on the stress harness's own queries",
          "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) The stress test aborts with `Test script failed` before running a single test: ``` Error on processing query: Timeout exceeded while receiving data from server. Waited for 15 seconds, timeout is 15 seconds. (query: SELECT value FROM system.server_settings WHERE name = 'cannot_allocate_thread_fault_injection_probability') ... subprocess.CalledProcessError: ... returned non-zero exit status 159. ``` The fail-close verification in `install_thread_pool_fault_injection` (added in #104782) ran a single `clickhouse client` query with `--receive_timeout=15` and no retry, so one slow answer killed the whole job. The first commit retries it, mirroring `call_with_retry`, keeping the fail-close semantics: persistent failure or a zero probability still aborts the run. Retrying alone is not enough, because the server is not slow. The stress profile (`stress_tests.lib`) enables the server-side AST fuzzer for the `default` user with `ast_fuzzer_runs=5` and `ast_fuzzer_any_query=true`, and that also applies to the maintenance queries of the harness itself. The fuzzer runs as a query-finish callback, so the connection thread executes the five mutated follow-up queries before the response completes. In [Stress test (arm_release) on #113224](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113224&sha=297082adba654b8bafac191c6c61b5023bcc76ef&name_0=PR&name_1=Stress%20test%20%28arm%5Frelease%29) (https://github.com/ClickHouse/ClickHouse/pull/113224) the verification query read its 439 rows in 67 ms and the handler returned after 16.3 s, past the client's 15-second `receive_timeout`: ``` 21:12:45.393 executeQuery: Read 439 rows, 131.72 KiB in 0.067 sec. 21:12:45.395 ASTFuzzer: Fuzzed query: SELECT value FROM system.server_settings WHERE name = ... ... 5 fuzzed follow-up queries ... 21:13:01.597 TCPHandler: Processed in 16.274 sec. ``` Every one of the five queries the fuzzer fired on in that run was the harness's own: the `SYSTEM RELOAD CONFIG` and the verification query of `stress.py` (11.3 s and 16.3 s in the connection thread), the two `SELECT 1` readiness probes and the `SYSTEM STOP DISTRIBUTED SENDS` of `stress_tests.lib`. No test query was fuzzed, because `clickhouse-test` runs are already started with `ast_fuzzer_runs=0`. The second commit pins `ast_fuzzer_runs=0` on the harness's own queries, in `stress.py` and in `stress_tests.lib`. That is what `clickhouse-test` already does for its infrastructure queries (`clickhouse_execute_http`) and what `stress.py` already did for the smoke check and the hung check. An explicit value on the command line is marked `changed` even though it equals the default, so the client sends it and it overrides the profile. Besides the timeouts this also stops two hazards that `ast_fuzzer_any_query` made possible: * A fuzzed `DETACH DATABASE` / `KILL QUERY` from `prepare_for_hung_check`, or a fuzzed `SYSTEM STOP DISTRIBUTED SENDS` during shutdown, running some other statement. * Fuzzed copies of a `system.processes` query showing up in the very processlist the hung check is about to inspect. A fuzzed readiness probe can also burn most of the 30-second `receive_timeout` of `start_server`, which reports `Cannot start clickhouse-server` for a server that is up. Related: https://github.com/ClickHouse/ClickHouse/pull/109496 Related: https://github.com/ClickHouse/ClickHouse/pull/113224 Related: https://github.com/ClickHouse/ClickHouse/pull/104782",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113248",
          "createdAt": "2026-08-04T07:30:34Z",
          "updatedAt": "2026-08-13T11:42:36Z",
          "timestamp": "2026-08-13T11:42:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:93bb5dafe44f638dbd85",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114262",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114262",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Lazy materialization for reading local Parquet files (`file` / `File`)",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110970 Follow-up to the lazy materialization for object storage (#110970), as discussed in https://github.com/ClickHouse/ClickHouse/pull/110970#issuecomment-5247858385: implement it for plain local Parquet files read through `StorageFile` — the `file` table function and the `File` table engine. For `ORDER BY ... LIMIT n` queries, the columns that are not needed for sorting and filtering are read only for the `n` rows that survive the `LIMIT`. The format side (row-selective Parquet reads via `FormatFilterInfo::rows_to_read`) and the plan split (`JoinLazyColumnsStep`, `LazyMaterializingTransform`) from #110970 are reused as-is; this PR adds the `StorageFile` counterpart of the two branches: - The main pass appends a `__global_row_index` column (file index in a per-query `LazyFileRegistry` + physical row numbers from `ChunkInfoRowNumbers`) in `StorageFileSource::generate`. - The lazy branch (`LazilyReadFromFile` → `LazyReadFromFileSource` → `StorageFileLazyRowsSource`) reopens only the surviving files (at most `LIMIT n` of them) with a per-file set of rows to read and `parquet.preserve_order`. - The storage-agnostic column-split logic of `ReadFromObjectStorageStep::keepOnlyRequiredColumnsAndCreateLazyReadStep` (which columns can be deferred: `DEFAULT` expression dependencies, PREWHERE inputs, hive partition and virtual columns) is extracted into the shared `splitLazilyReadColumnsFromFormatInfo` and reused by both storages. **Generation safety.** POSIX has no conditional read, so the reread cannot be pinned the way `If-Match` pins it on S3. Instead the reread fails close with the new `FILE_CHANGED_DURING_READ` error when the file's generation token — sub-second mtime + inode + size, reusing `computeFileCacheVersionToken` from the query condition cache integration — no longer matches the one captured when the main pass opened the file. The token is validated at file registration time in the main pass, and both before and right after the reopen in the lazy pass. Replace-by-rename (the common atomic-update pattern) is always caught since it changes the inode; an in-place rewrite is caught up to the filesystem timestamp tick (and already tears a single-pass read today). **Gates** (`ReadFromFile::canUseLazyMaterialization`): Parquet format only, no file descriptor reads (stdin cannot be reopened), no archive entries, no `distributed_processing`, no `--rename_files_after_processing`, uncompressed files only. Controlled by the new setting `query_plan_optimize_lazy_materialization_for_file` (enabled by default), on top of `query_plan_optimize_lazy_materialization`. **Bug fix for #110970 shared via the helper:** deferring a requested subcolumn (e.g. of a `JSON` column) left the lazy branch's format header empty and the query failed with `Not found column or subcolumn ... in block`, because the format header contains the parent column of a requested subcolumn while the split filtered it by subcolumn names. The split now maps requested columns to their storage-level names — the parent of a deferred subcolumn moves to the lazy branch, and stays in the main branch as well when a sort key still needs another subcolumn of it. Covered for both storages by the new tests. On a 1 GB local Parquet file (2M rows × 105 columns equivalent shape: two 300-byte strings), `SELECT * ... ORDER BY k DESC LIMIT 5` runs 3× faster (0.15 s vs 0.45 s, page-cache warm), with results verified identical with the optimization on and off. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Lazy materialization for `ORDER BY ... LIMIT n` queries (#110970) now also applies to local Parquet files read with the `file` table function and the `File` table engine: the columns that are not needed for sorting and filtering are read only for the `n` rows that survive the `LIMIT`. The second read of a surviving file fails close with the new `FILE_CHANGED_DURING_READ` error if the file was modified between the two passes. Controlled by the new setting `query_plan_optimize_lazy_materialization_for_file` (enabled by default). Also fixes lazy materialization for object storage failing with `Not found column or subcolumn ... in block` when a requested subcolumn (e.g. of a `JSON` column) is deferred.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114262",
          "createdAt": "2026-08-11T03:20:36Z",
          "updatedAt": "2026-08-13T11:39:56Z",
          "timestamp": "2026-08-13T11:39:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5c6027149ce88a82942f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114618",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114618",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112573 to 26.6: Fix async bounded read buffer readbigat race",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114618",
          "createdAt": "2026-08-13T10:33:52Z",
          "updatedAt": "2026-08-13T11:39:38Z",
          "timestamp": "2026-08-13T11:39:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "kssenii",
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:dbd65e3778907e871d27",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114437",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114437",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Refuse Iceberg data compaction until it can publish its result",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/114194 Related: https://github.com/ClickHouse/ClickHouse/pull/114324 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `OPTIMIZE TABLE` on an Iceberg table (data compaction, `allow_experimental_iceberg_compaction = 1`) deleted files that retained snapshots still referenced, including `version-hint.text` and every `vN.metadata.json`, which left the table unreadable and could permanently lose an acknowledged snapshot's rows. In the open-source build it now reports `NOT_IMPLEMENTED` instead of running. `OPTIMIZE TABLE ... MANIFEST`, `expire_snapshots` and `remove_orphan_files` are unaffected. ### Description Related: #114194 (the same stale-hint root reaching `remove_orphan_files`). Measured on master (`26.8.1.1`), v2 table with position deletes: `OPTIMIZE TABLE` returns success silently, drops the object count 18 to 10, and deletes `metadata/version-hint.text` with `v1..v5.metadata.json`; every later read fails `Code: 107 FILE_DOESNT_EXIST`. No stale hint is needed. When the hint is behind the newest metadata, the rewrite is rooted at the older version, so an acknowledged snapshot's data files are not carried forward yet are still deleted: `groupArray(x)` returns `[1]` where `[1,99]` was committed. Root cause: `getOldFiles` takes a raw listing of `metadata/` and `data/` with no reachability filter and no age gate, and `clearOldFiles` deletes all of it. Those files are reachable by the engine's own definition - `collectSnapshotReferencedFiles` walks every entry of the `snapshots` array - so deleting them is wrong at any age, under any guard. The rewrite is also never published atomically: `getPlan` never calls `setVersion`, so it always writes `v0.metadata.json`, and `writeMetadataFiles` commits with a bare `WriteMode::Rewrite`, no hint and no catalog commit. Since the old files are listed before it runs, cleanup deletes the previous `vN` and leaves `v0`, so a reader resolves whichever generation survived - in the measured run, the truncated one. Publishing correctly needs the current-snapshot replacement model and catalog plumbing that do not exist here, so this refuses the path rather than repairing it, following the existing format-version 3 refusal for `OPTIMIZE ... MANIFEST`. The data-compaction success contracts in the tests were deleted rather than inverted into \"the feature is absent\" assertions, which would have to be deleted again once publication lands. The new integration test asserts the property that holds either way: no pre-existing file is removed and the committed rows survive. A stateless test pins the refusal. `OPTIMIZE ... MANIFEST` coverage is untouched.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114437",
          "createdAt": "2026-08-12T08:24:40Z",
          "updatedAt": "2026-08-13T11:38:19Z",
          "timestamp": "2026-08-13T11:38:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:bcfea36d905048493a03",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113024",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113024",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Run the documentation examples in CI",
          "text": "#108556 made the SQL examples embedded in `system.documentation` runnable, but nothing runs them, so they go stale again as soon as behaviour changes: since then a steady stream of one-off fixes has been needed (#109421, #109965, #110459, #112287, ...). This adds a runner and a CI job that execute every one of them, and brings the examples and their documented responses back in line with what the server actually does. ### The runner `tests/docs_examples/runner.py` reads every example out of `system.documentation` on a running server and executes it. The examples of one entity run **in order, in a single session, in a database of their own**: a documentation page is written to be followed from top to bottom, so one example commonly creates the table a later one queries. It is plain Python and can be pointed at any server: ``` python3 tests/docs_examples/runner.py --port 8123 python3 tests/docs_examples/runner.py --port 8123 --filter '^argM' --verbose ``` Each example gets one of three outcomes: * `ok` — it ran, and its output matches the documented response (or it documents no response, or it documents an exception and indeed threw it); * `error` — it failed to run (or documents an exception and did not throw it); * `output` — it ran, but its output differs from the documented response. Everything that is not `ok` is listed in `tests/docs_examples/known_failures.txt` with the reason for it. The run fails if an example that is not on the list fails, and also if a listed one starts passing, so the list can only shrink. A few entries are marked `unstable`, for the examples whose output is random enough to sometimes match the documented one. An example that documents an exception is expected to throw it from its **last** statement — what comes before is setup that has to succeed — and when the documented response names an error code, the exception has to carry the same one. The message text is not compared: it holds a version number and a query pipeline description that are not part of what the example teaches. The server the job starts runs in `Etc/UTC`. Many examples convert between a date with time and a number, or hash a `DateTime`, without naming a time zone, and the documented responses are the ones of a `UTC` server; without pinning it, the run would only reproduce on a machine whose time zone matches. The comparison pins the Pretty rendering to the plain form the documented responses are written in — no row numbers, no colour, long column names spelled out in full, no readable-number tip, a named tuple printed as a tuple — so that it is about the data and the shape of the result rather than about a rendering default that changed after a page was written. ### The job `ci/jobs/docs_examples_job.py` starts a server from the shipped configuration plus the fragments in `programs/server/config.d` (macros, the legacy geobase, the natural language processing data, Keeper, the test clusters) and `programs/server/users.d` (the localhost-only network of the `default` user, access management, query logging), and two fragments of its own in `tests/docs_examples/config.d` and `tests/docs_examples/users.d`, so the features the examples demonstrate are actually configured, and configured the way the shipped server configures them. The one thing the examples add to the shipped user configuration is stated explicitly: the `queryID`, `initialQueryID` and `initialQueryStartTime` examples read from three shards at `127.0.0.{1..3}`, so `default` is allowed in from the whole loopback network rather than from `127.0.0.1` alone. It publishes an HTML report naming every failing example with its source file, its query and both responses. The job runs on pull requests and on master, next to the other jobs that run a corpus of queries against a server. ### What it found **Examples that did not run.** All of these are fixed here: * examples that were never a query: a bare expression (`factorial(10)`), a syntax template with placeholders sitting in an `Examples` block (the `iceberg*`, `paimon*`, `deltaLake*`, `oss`, `cosn` and `mergeTree*` table functions — moved to the `syntax` field, where they belong), a leaked C++ string literal, a MySQL session transcript, an unbalanced parenthesis, a stray `\\G` left over from a `clickhouse-client` session; * an example calling the wrong function: `YYYYMMDDhhmmssToDateTime` demonstrated `YYYYMMDDToDateTime`; * examples reading a table nothing creates: `salary`, `Employees`, `t`, `key_val`, `encryption_test`, `student_ttest`, `points`, `example_table`, ... — each page now creates its own; * pages that cannot be followed top to bottom, because every example re-creates the same table and the second one hits `TABLE_ALREADY_EXISTS`; * arguments the function rejects: an H3 index that is not a valid directed edge, S2 cell ids that do not form a valid rectangle, a signed weight for a weighted quantile, an `INSERT ... FORMAT JSONEachRow` without the semicolon that ends its data, a cipher the build does not provide; * examples of an error whose response was prose (\"Raises a `NO_COMMON_TYPE` exception\") or a paste from version 19.14, now written as the exception the server prints — the runner treats a documented exception as an expectation to fail. **Responses that no longer match.** 621 of them are regenerated from what the server prints, after checking that the output is identical across two runs on a fresh server. Most were stale column names (`avg(x)` for a column named `t`, `argMax(a, tuple(b, a))` for what is now printed as `argMax(a, (b, a))`) or hand-typed values that never came from a server (`[4, 3, 2, 1]` where the server prints `[4,3,2,1]`, `'dcba'` where it prints `dcba`), but some were genuinely wrong results. **Examples that needed a dataset nobody has.** `anyHeavy`, `categoricalInformationValue`, `topK`, `IPv4NumToStringClassC`, `IPv6NumToString`, `arrayEnumerateUniq`, `bar`, `indexHint`, `transform` and `evalMLMethod` demonstrated themselves on `ontime`, `metrica.hits`, `test.hits`, `hits_all`, `test.visits` or `trips`. Each of them now builds a small table of its own, so the example is one a reader can run. `naiveBayesClassifier` and its two variants train a dictionary on an inline set of token counts instead of naming one that does not exist. The four `flameGraph` \"examples\" were not examples at all — their documented response was a `clickhouse client ... | flamegraph.pl` command line — so they are recipes in the description now, and the function has one example that runs. **What is left in the known-failures file** is what a test server cannot produce: examples that call an external model provider, need a trained CatBoost model, a TLS certificate or a geobase hierarchy the test configuration does not carry; responses that describe the machine or the build (a host name, a source path, a version, a disk size); and outputs that are random or depend on the time. ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113024",
          "createdAt": "2026-08-02T19:33:37Z",
          "updatedAt": "2026-08-13T11:34:16Z",
          "timestamp": "2026-08-13T11:34:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:aabc94272a83d28adbb4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114465",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114465",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix an error reading an Iceberg table whose `current-snapshot-id` is JSON null",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109739 ## Problem The Iceberg spec lets `current-snapshot-id` be absent, `-1`, or JSON null, and all three mean the same thing: the table has no current snapshot. External writers do emit `\"current-snapshot-id\": null`. Reading such a table, several commands fail with an unrelated Poco conversion error instead of taking the no-snapshot path: - `SELECT ... FROM system.iceberg_history` reports the table as broken and skips it, so its history is silently missing from the result: `Ignoring broken table <db>.<t>: Poco::Exception. Code: 1000, Invalid access: Can not convert empty value.` - `ALTER TABLE ... EXECUTE expire_snapshots(...)` fails. - `DELETE` / `UPDATE` fail. ## Root cause `Poco::JSON::Object::has` returns true for a key whose value is JSON null — the key is present, the value is an empty `Poco::Dynamic::Var`. A reader guarded only by `has` therefore proceeds to `getValue<Int64>`, and `Var::convert<Int64>` throws `Poco::InvalidAccessException`: ``` Poco::Dynamic::Var::convert<long>() DB::IcebergMetadata::getHistory(std::shared_ptr<DB::Context const>) const DB::StorageSystemIcebergHistory::fillData(...) ``` Each affected site is immediately followed by a `< 0` / `>= 0` test, so treating null like a negative id is what the surrounding code already intends; it simply never handled that spelling. ## Fix Three reads now also check `isNull`, matching the idiom already used by their neighbours in the same files (`IcebergMetadata.cpp:581`, `Mutations.cpp:601`, `IcebergWrites.cpp:1120`): | Site | Reached by | |---|---| | `IcebergMetadata::getHistory` | `system.iceberg_history` | | `expireSnapshots` | `ALTER TABLE ... EXECUTE expire_snapshots(...)` | | `mutate` | `DELETE` / `UPDATE` | Regression test: `tests/queries/0_stateless/04846_iceberg_null_current_snapshot_id.sh`, which rewrites `current-snapshot-id` to JSON null and exercises all three. Verified closed-loop — with the three guards reverted and rebuilt, the test fails with the error above; with them restored it passes. Each guard was attributed individually rather than only as a group. Two related reads are deliberately **not** changed here: - `writeMetadataFiles` (`Mutations.cpp`) has the same pattern, but no repro could be constructed: it runs only after a mutation has matched rows, and a null `current-snapshot-id` makes the table read as empty, so the two conditions are mutually exclusive. Reverting only that line leaves the new test green. - The two reads in `Compaction.cpp` are fixed by #109739, which also adds the manifest-compaction validation and null coverage there. Both pull requests target `master` independently; their changes are complementary and can merge in either order. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an error when reading an Iceberg table whose `current-snapshot-id` is JSON null, which some writers emit for a table with no current snapshot.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114465",
          "createdAt": "2026-08-12T10:37:46Z",
          "updatedAt": "2026-08-13T11:32:44Z",
          "timestamp": "2026-08-13T11:32:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:17e738077ff06d1df85c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114476",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114476",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113289 to 26.6: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114476",
          "createdAt": "2026-08-12T12:06:37Z",
          "updatedAt": "2026-08-13T11:32:09Z",
          "timestamp": "2026-08-13T11:32:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll3",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4c0eeaa1193b5785cea2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114621",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114621",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #94748 to 25.8: Fix invalid result of joining two `-Cluster` table functions",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/94748 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31694431101/job/94428840299)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114621",
          "createdAt": "2026-08-13T11:31:25Z",
          "updatedAt": "2026-08-13T11:31:52Z",
          "timestamp": "2026-08-13T11:31:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "thevar1able",
            "KochetovNicolai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7346605d82befbba955e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:94748",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:94748",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix invalid result of joining two `-Cluster` table functions",
          "text": "Assisted-by: Claude Sonnet 4.5 via GitHub Copilot ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix invalid result on joining multiple table expressions, when leftmost table expression is a `-Cluster` table function. Resolves https://github.com/ClickHouse/ClickHouse/issues/89996 <!-- ch-version-info:start --> ### Version info - Merged into: `26.1.1.899` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/94748",
          "createdAt": "2026-01-21T18:10:36Z",
          "updatedAt": "2026-08-13T11:31:33Z",
          "timestamp": "2026-08-13T11:31:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v25.8-must-backport"
          ],
          "author": "thevar1able",
          "state": "closed",
          "assignees": [
            "KochetovNicolai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a9d9c85ec210b7f038a1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:99981",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:99981",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Iceberg: propagate table UUID from REST catalog to avoid metadata cac…",
          "text": "REST catalog inline responses already contain the table UUID and metadata location. Propagate the UUID through DataLakeSpecificProperties -> StorageObjectStorageConfiguration.catalog_uuid_hint -> initializePersistentTableComponents so the metadata cache is checked before fetching metadata.json from network storage. ### Performance mechanism **Before:** On first table initialization, `getMetadataJSONObject` is called without a known UUID. The cache probe is skipped (`table_uuid == nullopt`), so metadata.json is always fetched from remote storage as a cold read. Even if the same table is initialized again later, the cache was never populated on the first pass. **After:** The REST catalog provides the table UUID upfront via `catalog_uuid_hint`. `getMetadataJSONObject` probes the cache with `uuid:path` key before doing any remote I/O. If the metadata was previously cached (e.g., by a prior query that retroactively populated it), the cold read is eliminated entirely. If not, the cold read still happens but the result is cached with the correct UUID for future lookups. **Impact:** Deterministic reduction in remote `metadata.json` fetches for REST catalog tables. Each avoided fetch saves one round-trip to S3/GCS/ABFS plus deserialization overhead. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve Iceberg catalog metadata caching by using table UUID from catalog responses to warm the metadata cache, avoiding a redundant remote metadata.json read on table initialization. ### Documentation entry for user-facing changes Improve Iceberg catalog metadata caching. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Medium risk because it changes Iceberg metadata cache key composition and the metadata-loading path to use a new UUID hint, which could affect cache hit rates or correctness if UUIDs/paths are inconsistent. > > **Overview** > **Improves Iceberg REST catalog metadata caching** by propagating the table UUID from REST inline responses through `DataLakeSpecificProperties` into `StorageObjectStorageConfiguration` as `catalog_uuid_hint`, so the first metadata fetch can probe the metadata cache before doing extra remote IO. > > Updates metadata loading to accept an optional known UUID, capture raw metadata JSON for reuse, and retroactively populate the cache once the real UUID is discovered. Also fixes potential cache-key collisions by changing `IcebergMetadataFilesCache::getKey` to use a `uuid:path` delimiter, and adds gtest coverage for key uniqueness and cache hit/miss behavior. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 2e7400b8dc007641c702d41b3fd6cafdc10eb823. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/99981",
          "createdAt": "2026-03-18T21:44:06Z",
          "updatedAt": "2026-08-13T11:30:37Z",
          "timestamp": "2026-08-13T11:30:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 29
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "bacek",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cd3b03c6b460a66be8ec",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114409",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114409",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Revert the PromQL topk/limitk streaming plan and its shared-subquery materialization\"",
          "text": "Reverts ClickHouse/ClickHouse#114326 Depends on https://github.com/ClickHouse/ClickHouse/pull/113397",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114409",
          "createdAt": "2026-08-12T01:15:49Z",
          "updatedAt": "2026-08-13T11:29:41Z",
          "timestamp": "2026-08-13T11:29:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:001f59b1d2e49eb980e8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111895",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111895",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add exclude_data_from_backup MergeTree setting",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111827 ### Changelog category (leave one): - New Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added exclude_from_backup and exclude_data_from_backup MergeTree table settings so BACKUP can skip a table entirely or skip only its data while still restoring its DDL. ### Description Implements #111827: a table-level setting so that `BACKUP` can skip a table's data while still including its DDL, so the table is restorable (empty) later. Useful for tables whose data can be regenerated from a source table (e.g. materialized-view targets), to reduce backup size. - New `Bool` MergeTree setting `exclude_data_from_backup` (default `false`). - Hooked into `BackupEntriesCollector::shouldBackupTableData()`: when the setting is enabled on a `MergeTreeData`-derived table, data collection is skipped for that table; DDL collection is unaffected (existing code path already handles DDL/data independently). - Added `tests/integration/test_exclude_data_from_backup/test.py` covering both the default (`false`, data backed up) and enabled (`true`, data skipped) cases; both pass. **Not yet tested** (feedback welcome): behavior on `ReplicatedMergeTree` specifically, and `BACKUP DATABASE`/`BACKUP ... ALL TABLES` paths (the hook is in the shared per-table code path used by all backup forms, so it should behave the same, but I haven't added explicit coverage for these yet). The setting name/interface is intentionally not MergeTree-specific in wording, per @alexey-milovidov's suggestion on the issue, so it could be adopted by other engines later.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111895",
          "createdAt": "2026-07-25T12:25:30Z",
          "updatedAt": "2026-08-13T11:27:24Z",
          "timestamp": "2026-08-13T11:27:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "adityaksolves",
          "state": "open",
          "assignees": [
            "jkartseva"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b4801fa509450fe109f4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114617",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114617",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112573 to 26.5: Fix async bounded read buffer readbigat race",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114617",
          "createdAt": "2026-08-13T10:33:24Z",
          "updatedAt": "2026-08-13T11:26:27Z",
          "timestamp": "2026-08-13T11:26:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "kssenii",
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:49276f5f6e8a0cdf046b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104993",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104993",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Handle non-constant RHS for `IN`",
          "text": "Fix `IN` and `NOT IN` expressions with non-constant right-hand side operands that reference columns from the current row. # old analyzer Previously, the old analyzer tried to build a standalone `Set` for expressions such as `number % 2 IN (number % 3, number % 5)`, which made the right-hand side unable to resolve `number` and produced an `UNKNOWN_IDENTIFIER` exception. This change rewrites such expressions to row-wise `has` expressions instead. ```sql --- error was reported for the query below for old analyzer as mentioned in #58242 SET enable_analyzer = 0; SELECT number FROM numbers(10) WHERE number % 2 IN (number % 3, number % 5) ORDER BY number; ``` ### Out of scope: bare source-column RHS under the old analyzer Under the old analyzer (`enable_analyzer = 0`), a bare column of the `FROM` source as the right-hand side, such as `x IN (arr)` where `arr` is a column of the current row, still fails with `UNKNOWN_TABLE`: `MarkTableIdentifiersVisitor` rewrites `x IN ident` into `x IN (SELECT * FROM ident)` before source columns are collected, so the expression never reaches the row-wise rewrite. The new analyzer resolves the same query as a column and succeeds; tests `04234` and `04812` pin this divergence explicitly. Closing it needs either reordering that visitor after source columns are known, or falling back from a table to a column when the table does not exist, plus a compatibility decision for `x IN t` when a column shadows an existing table name - that is tracked in the review discussion and left out of this PR on purpose. # new analyzer The new analyzer already handled the basic non-constant right-hand side case, but some tuple and NULL cases still failed. This change fixes: * tuple-typed right-hand side expressions produced by functions other than tuple * tuple left-hand side membership checks that previously tried to create `Nullable(Tuple(...))` * `NULL` operands in non-constant tuple right-hand side operands, where the old cast target could become `Nullable(Nothing)` Examples: ```sql --- Before this fix, the new analyzer treated the tuple-typed if RHS as a single tuple value and failed with a type error; after this fix, it expands the tuple value one level for scalar IN, so the query returns 1 SELECT number IN (if(number >= 0, tuple(number, number + 1), tuple(0, 0))) FROM numbers(1); ``` ```sql --- tuple in tuple, user exepects some rows to match, however, error like `Cannot create column with type 'Nullable(Tuple(UInt8, UInt8))' because Nullable Tuple type is not allowed` will be reported before this fix SELECT number, (1, 1) IN ((number % 3, number % 2), (2, 2)) FROM numbers(6) ORDER BY number; ``` ```sql --- user expects `NULL` but error like `Conversion from UInt8 to Nothing is not supported` will be reported before this fix SELECT x IN (y, 1) FROM ( SELECT materialize(NULL) AS x, materialize(2) AS y ); ``` Issue: https://github.com/ClickHouse/ClickHouse/issues/58242 ### Changelog category (leave one): - Bug Fix ### Changelog entry: - Fix IN and NOT IN expressions with non-constant right-hand side operands referencing columns from the current row, and align new analyzer tuple right-hand side handling with existing ClickHouse IN semantics. Under the old analyzer, a bare source column as the right-hand side (`x IN (arr)`) still resolves as a table name and stays out of scope. This closes [#58242](https://github.com/ClickHouse/ClickHouse/issues/58242)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104993",
          "createdAt": "2026-05-15T04:11:24Z",
          "updatedAt": "2026-08-13T11:24:10Z",
          "timestamp": "2026-08-13T11:24:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 27
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "niyue",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:17ad21df19a764b953d7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114074",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114074",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix vector search with `arrayJoin` below the sort and with a row policy",
          "text": "Fixes for the vector search `ORDER BY <distance> LIMIT` optimization. **`arrayJoin` below the sort.** `arrayJoin` changes the number of rows - an empty array drops its base row - so the rewrite must not shortlist base rows before it runs. `optimizeTopK` rejects `arrayJoin` for this reason (https://github.com/ClickHouse/ClickHouse/issues/82279), but the vector search optimization did not, so the rows that a later base row would have contributed were never produced: ```sql CREATE TABLE t (id UInt32, tags Array(UInt32), vec Array(Float32), INDEX idx vec TYPE vector_similarity('hnsw', 'cosineDistance', 2)) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 4; INSERT INTO t SELECT number, if(number < 16, [], [number]), [toFloat32(number), toFloat32(number + 1)] FROM numbers(64); SELECT arrayJoin(tags) FROM t ORDER BY cosineDistance(vec, [0., 1.]) LIMIT 1; -- returned nothing; with `use_skip_indexes = 0` it returns 16 ``` **A row policy is an additional filter.** A row policy restricts rows inside the reader just like a `WHERE` or a `PREWHERE`, but it did not participate in `additional_filters_present`, so a query filtered only by a policy was treated as unfiltered: the `vector_search_filter_strategy = 'prefilter'` bailout was skipped, so an explicit request for exact search was ignored, and the index fetched only `LIMIT` neighbours without the `vector_search_index_fetch_multiplier` compensation, which the policy can then discard. **Filters that read the vector column broke the non-rescoring rewrite** (found during review). The rewrite drops the physical vector column from the read list and replaces it with the virtual `_distance` column, so any filter that still needs the column was left without its input and the query failed with `NOT_FOUND_COLUMN_IN_BLOCK` (an exception, pre-existing on master). The rewrite is now skipped - the same treatment as the vector column in `SELECT` - when the column is read by: - a row policy (including a policy carried in the deferred filter under `FINAL` with `apply_row_policy_after_final`), - a plain `WHERE` filter below the sort, e.g. `WHERE length(vec) > 0` (reachable with default settings, because the implicit `PREWHERE` optimization is disabled for vector-search candidates). The quantized-codes rewrite (`useVectorSearchWithQuantizedCodes`) is not affected by these problems: it splices the shortlist above the whole `Expression`/`Filter` chain, so the expansion and the filters run below it, and its reader-side filters prefilter the approximate ranking. This was checked with the same queries. `FINAL` with an explicit `PREWHERE` is covered by a separate regression test. `FINAL` can defer a `PREWHERE` until after the merge, but the deferral copies the filter and clears `query_info.prewhere_info` only when the pipeline is built, i.e. past every query plan optimization, so the existing `getPrewhereInfo` bailout still fires and the physical vector column survives for the deferred filter. Related: https://github.com/ClickHouse/ClickHouse/issues/82279 Related: https://github.com/ClickHouse/ClickHouse/pull/110188 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed vector search returning too few rows for a query with `arrayJoin` below `ORDER BY <distance> LIMIT`. A row policy now counts as an additional filter for vector search, so `vector_search_filter_strategy = 'prefilter'` is honored and the index fetch multiplier is applied for a query filtered by a row policy. Fixed `NOT_FOUND_COLUMN_IN_BLOCK` for a non-rescoring vector search query when a `WHERE` filter or a row policy reads the vector column. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114074",
          "createdAt": "2026-08-09T21:37:59Z",
          "updatedAt": "2026-08-13T11:23:41Z",
          "timestamp": "2026-08-13T11:23:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f8323966a6bc867f7f0b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106199",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106199",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Push down volume-reducing functions in query plan",
          "text": "Adds a query plan optimization that pushes volume-reducing functions (`length`, `lengthUTF8`, `empty`, `notEmpty`) below the `Sorting` and `Filter` steps, so those steps carry the fixed-size result instead of the original `String` / `FixedString` / `Array` / `Map` argument. The rewrite is done with `ActionsDAG::split`: the function nodes form the first part, which becomes a new step below the child, and the second part becomes the parent step's `ActionsDAG` and reads the results as inputs. It only fires when the wide argument column really stops flowing through the child step — the column is removed from the child's output, and unless the child reads it itself it is not even produced below the child. `tryExecuteFunctionsAfterSorting` and `trySplitFilter` keep the pushed functions where they are, so the optimizations do not move the same nodes in opposite directions. For the shape from the issue below, `SELECT avg(length(s)) FROM test WHERE notEmpty(s)`, the filter still needs `s` to evaluate its condition, but `length(s)` is now computed before the filter, so the wide column is not copied for the surviving rows. Enabled by default; can be turned off with `query_plan_push_down_volume_reducing_functions = 0`. An earlier attempt at the same optimization was https://github.com/ClickHouse/ClickHouse/pull/86139 by @talmawash, credited as a co-author. Closes: https://github.com/ClickHouse/ClickHouse/issues/82378 Related: https://github.com/ClickHouse/ClickHouse/pull/86139 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a query plan optimization that pushes volume-reducing functions (`length`, `lengthUTF8`, `empty`, `notEmpty`) below the `Sorting` and `Filter` steps, replacing the wide `String` / `FixedString` / `Array` / `Map` argument with the fixed-size result, so it is neither buffered by a sort nor copied by a filter. Controlled by the new setting `query_plan_push_down_volume_reducing_functions`, enabled by default. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106199",
          "createdAt": "2026-05-31T13:26:13Z",
          "updatedAt": "2026-08-13T11:21:58Z",
          "timestamp": "2026-08-13T11:21:58Z",
          "metrics": {
            "reactions": 2,
            "comments": 12
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "fastio",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ced81c6300a3a7737797",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108371",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108371",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Array subscript operator supports array of integers as index.",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/108095 The array subscript operator now accepts an array of integers as the index, so `arr[indexes]` gathers the elements at all of those positions at once. It is equivalent to `arrayMap(i -> arr[i], indexes)`, including the result type, the handling of negative indexes and the value produced for an out-of-range position, but it has its own implementation: a constant source array is not materialized per row, and numeric element types are gathered through the `PODArray` directly. ```sql SELECT [10, 20, 30, 40][[2, 4, 1]]; -- [20,40,10] SELECT [10, 20, 30][[1, 5, -1]]; -- [10,0,30] SELECT arrayElementOrNull([10, 20, 30], [1, 5]); -- [10,NULL] SELECT [10, 20, 30][[1, NULL, 2]]; -- [10,NULL,20] ``` The positions may be nullable, and a `NULL` position produces `NULL`, just as for a scalar index. The main use case is a lookup table: a constant dictionary array indexed by a per-row array of positions. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The array subscript operator supports an array of integers as the index: `arr[indexes]` returns the elements at all of the given positions, equivalently to `arrayMap(i -> arr[i], indexes)`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108371",
          "createdAt": "2026-06-24T12:15:02Z",
          "updatedAt": "2026-08-13T11:21:36Z",
          "timestamp": "2026-08-13T11:21:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "ucasfl",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:338d600b2c81b79fab8b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110171",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110171",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add VALID FOR clause for users and credentials",
          "text": "We support `VALID UNTIL` for users and credentials. This adds `VALID FOR <interval>` as a convenience shorthand. Instead of an absolute date and time, `VALID FOR` accepts an interval, and the expiration deadline is computed as the current time plus that interval at the moment the query is executed. The result is stored in the `VALID UNTIL` form, so `SHOW CREATE USER` always displays the resolved absolute deadline. It can be used everywhere `VALID UNTIL` can - at the user level and per authentication method, in both `CREATE USER` and `ALTER USER`. Examples: ```sql CREATE USER u1 VALID FOR INTERVAL 1 DAY; CREATE USER u2 IDENTIFIED WITH plaintext_password BY 'x' VALID FOR INTERVAL 3 MONTH; ALTER USER u1 VALID FOR INTERVAL 1 DAY + INTERVAL 12 HOUR; ``` Implementation notes: the current time is injected as a literal (rather than using `now`, which is non-deterministic and would not fold to a constant expression), and the deadline is computed with `toDateTime64` so that large intervals saturate at the `DateTime64` upper bound instead of overflowing the year-2106 boundary of `DateTime`. Note on the system table schema: to represent deadlines beyond the year 2106 exactly, the `valid_until` column of `system.users` changes from `Array(DateTime)` to `Array(DateTime64(0))`. This is a backward-incompatible change of a documented system table for tooling that introspects the column type; the values themselves keep second precision. ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added the `VALID FOR <interval>` clause to `CREATE USER` and `ALTER USER` as a shorthand for `VALID UNTIL`. The expiration deadline is computed as the current time plus the given interval at query execution time and stored in the `VALID UNTIL` form. The `valid_until` column of the `system.users` table now has the type `Array(DateTime64(0))` instead of `Array(DateTime)`, so that deadlines beyond the year 2106 are represented exactly; tooling that reads this column should handle the new type. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110171",
          "createdAt": "2026-07-12T18:12:03Z",
          "updatedAt": "2026-08-13T11:20:58Z",
          "timestamp": "2026-08-13T11:20:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-backward-incompatible",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b6ee9d84e7823bb40b6d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114340",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114340",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Populate the submodule working trees in parallel in the build job",
          "text": "The build job populated the submodule working trees with a bare `git submodule update` on a submodule cache hit, which walks the submodules one at a time on a single core. In `Build (arm_tidy)` on master that step takes 39 s at ~3% CPU of a 32-core `m8g.8xlarge` — an otherwise idle machine waiting on one `git` process, right before the equally single-threaded cmake configuration and long before `ninja` can saturate the box. The first 253 s of that job produce no compilation at all; this is one of the three serial phases in it. Report the numbers come from: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=03ad5002ff1501ceb5d584b6d0bb028674b276e7&name_0=MasterCI&name_1=Build%20%28arm_tidy%29 Fan the checkouts out over the idle cores instead, the same way the cache-miss branch right below it (`contrib/update-submodules.sh --max-procs 10`) and `ci/jobs/fast_test.py` already do: feed the submodule paths from `.gitmodules` to `xargs --max-procs`. The submodule list is read with `get_output_or_raise` and asserted non-empty, so failing to enumerate `.gitmodules` fails the step rather than silently checking out nothing — `xargs --no-run-if-empty` would otherwise exit 0 on an empty pipeline. Measured on a synthetic superproject with 40 submodules and a populated `.git/modules` (the CI cache-hit state): **2.13 s serial vs 0.16 s** with `--max-procs=20`, both producing an identical working tree and a clean `git status`. In CI the win is bounded by the largest submodule rather than by the sum of all of them, so expect the step to shrink to roughly the cost of `contrib/llvm-project` alone; the `Checkout Submodules` duration in this PR's build reports against 39 s on master is the real measurement. ### Changelog category (leave one): - CI Fix or Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114340",
          "createdAt": "2026-08-11T14:18:28Z",
          "updatedAt": "2026-08-13T11:17:16Z",
          "timestamp": "2026-08-13T11:17:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "maxknv"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bba65ad94f8b43d592f1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107567",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107567",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallelize listing of globbed `s3` table function paths",
          "text": "Speeds up listing files for the `s3` table function over globbed paths, in two complementary ways. The listing was a single serial stream of `ListObjectsV2` requests paginated by a continuation token, bound by the per-request latency of S3 (~75 ms measured against `clickhouse-public-datasets`); this dominates queries over buckets with many objects. An earlier attempt (#66504) used speculative keyspace splitting with fixed split points assuming a uniform key distribution and a fixed character set, and was closed as fragile. Controlled by the new setting `s3_list_object_parallelism` (default `1` = the previous serial behavior). **1. Hierarchical layouts** — walk the \"directory\" tree formed by the `/` delimiter with several worker threads in parallel, pruning subtrees that cannot contain a matching key (so for partitioned layouts like `year=*/month=*/` fewer objects are fetched, not merely listed in parallel). Recursive globs (`**`) keep serial listing. **2. Big flat directories** (e.g. files named by hash/UUID — `musicbrainz/mlhdplus-complete/<uuid>.txt.zst`, ~600k files in one prefix) — when a listed prefix is truncated but has no sub-directories, split it by keyspace: tile the remaining range into contiguous `(start_after, end]` sub-ranges listed concurrently, with boundaries derived from the byte alphabet observed in the listed page. Because the sub-ranges are contiguous they tile the interval **with no gaps** — a key whose byte is absent from the sampled alphabet still falls into the range that brackets it — so the split is complete regardless of the key distribution or character set. The split is a single level and is **probe-gated**: before splitting, one cheap `ListObjectsV2` checks whether any key exists beyond the current bucket; if not (keys share a common prefix, e.g. `pageviews-YYYYMMDD-HH`), the directory is paginated serially instead of issuing a fan of empty boundary probes. This is the robust counterpart to #66504 — it never regresses for clustered keys. On S3 Express / directory buckets — which reject `StartAfter` and only allow the `/` delimiter — keyspace splitting is disabled automatically (detected via `isS3ExpressBucket`), so flat ranges are paginated serially while hierarchical layouts are still listed in parallel. Correctness of glob matching is guaranteed by the existing per-file `RE2::FullMatch`, so the directory pruning only needs to be conservative. **Measured against real S3** (`clickhouse-public-datasets`, ~75 ms/list-request) across ~16 datasets — all return identical results (same count, no duplicates, identical path checksum), serial vs parallel: - Big flat, uniform keys — `musicbrainz/mlhdplus-complete` (594415 files): **37s → 3.7s (~10x)**. - Big flat, clustered keys — `wikistat` (97529) / `gharchive` (100091): probe sends them serial; parallel matches serial (no regression). - Hierarchical — `web/store`: ~2.5x (bounded by its small directory count; scales toward the parallelism factor for larger trees). - Small/medium directories (`nyc-taxi`, `ontime`, `hits/native`, `sensors`, `tranco`, `youtube`, `adsblol`, `bluesky`, …): correct, no regression. A unit test (`gtest_object_storage_parallel_listing`) asserts every key is produced exactly once across parallelism 1–64, including adversarial cases (bytes outside the sampled alphabet, shared long prefixes, boundary keys, mixed hierarchical+flat). Stateless tests (`04339_s3_parallel_glob_listing`, `04340_s3_parallel_flat_listing`) verify parallel and serial listing return identical results. Closes: https://github.com/ClickHouse/ClickHouse/issues/65572 Related: https://github.com/ClickHouse/ClickHouse/pull/66504 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a setting `s3_list_object_parallelism` to list the files of the `s3` table function over a globbed path in parallel. Hierarchically partitioned layouts are listed by walking the tree of common prefixes concurrently; a single big flat directory is listed by splitting its keyspace into contiguous sub-ranges. This speeds up queries over buckets with many objects (e.g. ~10x for a directory of 600k files).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107567",
          "createdAt": "2026-06-15T21:43:46Z",
          "updatedAt": "2026-08-13T11:15:29Z",
          "timestamp": "2026-08-13T11:15:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 32
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b3b11c87c3212f0e298a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110958",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110958",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix toTime key-expression type mismatch under use_legacy_to_time",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107951 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server abort (`Bad cast from type ColumnVector<UInt32> to ColumnVector<Int32>`, a `LOGICAL_ERROR`) when inserting into a `MergeTree` table whose `PRIMARY KEY` or `ORDER BY` uses `toTime(...)` from a session whose `use_legacy_to_time` value differs from the one under which the key metadata was built. The `toTime` legacy/new resolution now follows the context that builds the expression, so a persisted key expression's type no longer depends on the writing session. ### Description `toTime()` resolves to two different functions depending on the `use_legacy_to_time` setting: the new `toTime` returns `Time` (Int32-backed), the legacy variant returns `DateTime` (UInt32-backed). The selection in `FunctionFactory::tryGetImpl` read the thread-local query context and ignored the `context` argument the caller passed. A `MergeTree` table with `PRIMARY KEY (toTime(c1))` / `ORDER BY toTime(c1)` persists only the expression AST. Its primary-index on-disk type comes from `metadata_snapshot->getPrimaryKey().data_types`, derived by rebuilding the key expression with the storage's global context (server-default `use_legacy_to_time`). The part-writer serializes the index with that persisted type. When an `INSERT` runs in a session with a different `use_legacy_to_time`, the write-path key expression re-resolved `toTime` to the other variant, producing a column whose physical type mismatched the persisted serialization, so `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow` hit `assert_cast<ColumnVector<Int32>>(ColumnVector<UInt32>)` and aborted the server (a handled exception in release builds, an abort under debug/sanitizers). Fix: resolve the `toTime` legacy swap from the caller-provided `context` (falling back to the thread-local query context only when no context is supplied). Stored key/sorting expressions are rebuilt with the storage global context, so they now resolve `toTime` consistently with the type persisted in the table metadata, while normal query resolution still honors the session setting. Found by the BuzzHouse fuzzer. - CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109351&sha=b0cca3209ae52b188fa092a31eb1e754db40608d&name_0=PR&name_1=BuzzHouse%20%28amd_msan%29 - Check: `BuzzHouse (amd_msan)`; assertion `Bad cast from type DB::ColumnVector<unsigned int> to DB::ColumnVector<int>` at `SerializationNumber<int>::serializeBinary` <- `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow`. Reproducer: ```sql CREATE TABLE t (c0 Int32, c1 DateTime64 MATERIALIZED nowInBlock64()) ENGINE = MergeTree() PRIMARY KEY (toTime(c1)); INSERT INTO t (c0) SETTINGS use_legacy_to_time = 1 SELECT number FROM numbers(1000); ``` ### DDL behaviour change carried by the fix Previously, `CREATE TABLE` in a session whose `use_legacy_to_time` differed from the server-wide default stamped the stored key type with the session's resolution (e.g. `DateTime` when the session set `use_legacy_to_time = 1` on a server defaulting to `0`). With this fix, the stored key type always resolves under the server-wide default, so the session setting at `CREATE` time no longer affects the persisted key type. This is observable via `DESCRIBE mergeTreeIndex(...)` and is pinned by the test. Upgrade note for that narrow window (table created while the session setting differed from the server global, on a pre-fix binary): after the upgrade the same table resolves its key as `Time`, so parts written before and after store different raw key values for the same timestamp (e.g. `90000` vs `3600` for `01:00:00`); reads, inserts and merges succeed, but a key-range predicate may miss rows from old parts. Where session and global agreed (the overwhelmingly common case), nothing changes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110958",
          "createdAt": "2026-07-18T23:24:10Z",
          "updatedAt": "2026-08-13T11:15:28Z",
          "timestamp": "2026-08-13T11:15:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 21
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2e5e31b2ddfb37491d61",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114507",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114507",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Column IDs for MergeTree: the data structure",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ## Background First part of [#99754](https://github.com/ClickHouse/ClickHouse/pull/99754) This one adds the data structure — `ColumnId`, `ColumnIdMapping`, its durable store, and the schema fields that carry a mapping: - `ColumnId`: a strong typed struct to replace the `std::string column_name` in most places - plug in `NameAndTypePair` & `ColumnDescription`, which are used in table schema - `ColumnIdMapping`: in-memory structure to store the `name <-> id` mapping - plug in `StorageInMemoryMetadata`: so that it is captured atomically with the schema - `ColumnIdMappingStore`: persistence store for the mapping ## Not in this PR Stamping ids on columns, id-keyed reads and writes, the experimental setting that turns the feature on, metadata-only `RENAME` / `DROP COLUMN`, and `ReplicatedMergeTree` / `SharedMergeTree` support. Each is planned as a follow-up on top of this one. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114507",
          "createdAt": "2026-08-12T15:18:12Z",
          "updatedAt": "2026-08-13T11:14:51Z",
          "timestamp": "2026-08-13T11:14:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "murphy-4o",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d677cd3b09279be597da",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109368",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109368",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Require a join subquery alias only when it removes a real ambiguity",
          "text": "`joined_subquery_requires_alias = 1` (the default) rejected every unaliased subquery, table function or union used in a multi-table join, even when the missing alias could not cause any ambiguity. That is stricter than necessary: an alias only serves to qualify a column, so it is only needed when the unaliased table expression exposes a column whose name also occurs in another table expression of the same join. `validateJoinTableExpressionWithoutAlias` now throws `ALIAS_REQUIRED` only on such a name collision (computed with the existing `getColumnsFromTableExpression` helper over the sibling table expressions), and otherwise allows the missing alias. Genuine ambiguities involving non-sibling table expressions are still caught later by the normal `AMBIGUOUS_IDENTIFIER` resolution, exactly as they are for ordinary tables. This lets standard queries such as TPC-DS q14 (whose `cross_items` derived table has no correlation name) run without setting `joined_subquery_requires_alias = 0`: ```sql SELECT i_item_sk FROM item, (SELECT iss.i_brand_id AS brand_id FROM store_sales, item AS iss ...) WHERE i_brand_id = brand_id; -- no shared column name -> no alias needed ``` Notes: - Validation in `resolveJoin`/`resolveCrossJoin` is moved to run after all table expressions of the join are resolved, so sibling columns are known when the collision is checked. - The change is purely permissive: it never turns a previously-succeeding query into an error. When the columns of any side cannot be determined it falls back to the old strict behavior. - Only the analyzer is affected; the deprecated non-analyzer path keeps the stricter behavior. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): With `joined_subquery_requires_alias = 1` (the default), the analyzer now rejects an unaliased subquery or table function in a join only when the missing alias would make one of its columns unreachable, for example because the name collides with another joined table expression or is shadowed by an in-scope alias; otherwise the query is allowed. If the analyzer cannot determine the exposed names up front, it keeps the old strict behavior. Unambiguous queries, such as some standard TPC-DS queries, no longer require adding an alias. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109368",
          "createdAt": "2026-07-03T21:53:19Z",
          "updatedAt": "2026-08-13T11:14:24Z",
          "timestamp": "2026-08-13T11:14:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 26
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "m-selmi"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d0a4b3fd6a047c1a3e66",
        "signalId": "github:ClickHouse/ClickHouse:issue:114591",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114591",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Settings profile with a Map setting (`http_response_headers`) does not survive a serialization round trip",
          "text": "### Company or project name ClickHouse ### Describe the unexpected behaviour An access entity with a Map-valued setting — e.g. a settings profile with `http_response_headers` — is serialized into a form that ClickHouse's own parser cannot read back. The stored entity becomes permanently unloadable. ### How to reproduce ```sql -- A string literal is the only accepted way to write a Map setting in a profile -- (map literals {...}, array literals [...], and map(...) are all rejected here): CREATE SETTINGS PROFILE test_profile SETTINGS http_response_headers = '{''Content-Type'':''application/json'', ''Access-Control-Allow-Origin'':''*''}' CONST; SHOW CREATE SETTINGS PROFILE test_profile; -- CREATE SETTINGS PROFILE `test_profile` SETTINGS http_response_headers = [('Content-Type', 'application/json'), ('Access-Control-Allow-Origin', '*')] CONST -- Feeding the emitted statement back — which is exactly what DiskAccessStorage and -- ReplicatedAccessStorage store and re-parse via deserializeAccessEntity: CREATE SETTINGS PROFILE test_profile2 SETTINGS http_response_headers = [('Content-Type', 'application/json'), ('Access-Control-Allow-Origin', '*')] CONST; -- Code: 62. DB::Exception: Syntax error: failed at position 72 ([): -- Expected one of: literal, NULL, NULL, number, Bool, TRUE, FALSE, string literal, end of query ``` No value of the setting round-trips — even an empty map serializes as `[] CONST`, which is rejected the same way. ### Root cause The write path and the read path disagree: - Write: `SettingsProfileElement` (`src/Access/SettingsProfileElement.cpp`) casts the value to the setting's native type (`Map`), and the AST formats the Field as a literal, producing an array-of-tuples literal `[('k', 'v'), ...]`. - Read: `ParserSettingsProfileElement` (`src/Parsers/Access/ParserSettingsProfileElement.cpp`) parses the value with a scalar-only literal parser. Unlike the query-level `SETTINGS` clause parser, it was never extended to accept collection literals (cf. #75065, where the query-level parser expects `literal or map, ..., OpeningCurlyBrace`). So `serializeAccessEntity` → `deserializeAccessEntity` fails for any user/role/profile whose profile elements contain a collection-valued setting (`http_response_headers`, `additional_table_filters`, ...). > *Generated by [Nerve](https://github.com/ClickHouse/nerve)*",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114591",
          "createdAt": "2026-08-13T04:17:50Z",
          "updatedAt": "2026-08-13T11:14:20Z",
          "timestamp": "2026-08-13T11:14:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "potential bug"
          ],
          "author": "pufit",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:561717a75baccead0ec7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100391",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100391",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallelize read-in-order from a single part with PrefetchingConcatProcessor",
          "text": "When a single MergeTree part is split into multiple streams for read-in-order, use the new `PrefetchingConcatProcessor` instead of relying on the downstream `MergingSortedTransform`. The key difference: - `MergingSortedTransform` marks all inputs as needed but only consumes from one stream at a time (for non-overlapping ranges), and each port can only buffer 1 block — so after initial prefetch, reading becomes sequential. - `PrefetchingConcatProcessor` marks all inputs as needed AND pulls data from non-current inputs into internal buffers, keeping upstream sources busy. It outputs data in strict input order (concatenation), which is correct because ranges from a single part are non-overlapping and pre-sorted. This avoids both the merge comparison overhead and the sort overhead, while enabling true parallel IO and filtering across range groups. Benchmark on 100M rows (4 threads, single part, poor PK selectivity): - Single-thread read-in-order: 1.57s - Parallel + sort (no read-in-order): 1.27s - PrefetchingConcat: 0.61s (2.6x faster) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve the performance of queries that are reading in the order of the primary key. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes read-in-order execution and introduces a new buffering processor in the query pipeline; incorrect gating or backpressure/buffering behavior could impact result ordering or memory usage in some MergeTree reads. > > **Overview** > **Adds `PrefetchingConcatProcessor`**: a new concat-like processor that keeps *all* inputs marked needed and buffers per-input chunks (bounded by `max_buffered_chunks`) to prefetch non-current streams while still outputting streams strictly in order. > > **Wires it into MergeTree read-in-order**: `ReadFromMergeTree` now conditionally replaces `MergingSortedTransform` with `PrefetchingConcatProcessor` when reading in ascending order from a *single part* split into multiple streams and only when there is filtering work (`PREWHERE` or row-level filter); it also adds `prefer_multiple_streams`/`setPreferMultipleStreams()` so aggregation-/distinct-in-order paths can opt out to avoid collapsing parallel streams. > > **Tests/bench coverage updated**: adds a new stateless test asserting `PrefetchingConcat` appears only for single-part reads and preserves sort order, adjusts an existing buffering test to force multi-part behavior, and adds a performance scenario guarding against regressions for filter-less `ORDER BY` reads. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit c0ba40d2f630976681024d2ac941261c4abdbe57. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100391",
          "createdAt": "2026-03-22T20:21:02Z",
          "updatedAt": "2026-08-13T11:12:54Z",
          "timestamp": "2026-08-13T11:12:54Z",
          "metrics": {
            "reactions": 1,
            "comments": 41
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cf384023d4788e53d084",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112304",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112304",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Supporting lazily replicated arrays on arrayElement and arraySlice (reduces memory usage and improves performance)",
          "text": "arrayElement and arraySlice now consume lazily replicated arrays (ColumnReplicated, produced by lazy ARRAY JOIN and JOIN under enable_lazy_columns_replication) without materializing them. Follow up of : https://github.com/ClickHouse/ClickHouse/pull/111581 and https://github.com/ClickHouse/ClickHouse/pull/111749 Closes: https://github.com/ClickHouse/ClickHouse/issues/54967 ### Implementation: - A new ReplicatedSource<Base> adapter in GatherUtils wraps any array source (numeric, generic, or nullable) over the compact nested column and remaps each logical row to its nested row through the replication indexes. - arrayElement gets a dedicated path: a per-row gather from the nested data for non-constant indexes (supporting negative indexes, out-of-range defaults, Nullable elements, and arrayElementOrNull), and for constant indexes it executes on the nested rows only and re-wraps the result lazily with the same indexes. - Per-row index arguments (arraySlice offset/length, arrayElement index) are materialized if they arrive replicated - Unsupported shapes (Map, LowCardinality elements) Measured on the test workload (1000-element String arrays, 50× replication by ARRAY JOIN, 100 rows): peak memory drops from 1.12 GB to 29 MB (38×). ### First Query ``` sql WITH materialize(range(1000)) AS large_array SELECT count() FROM system.numbers WHERE NOT ignore( arrayMap(idx -> arraySlice(large_array, idx, 5), arrayEnumerate(large_array))) SETTINGS max_rows_to_read = 262144, read_overflow_mode = 'break', max_memory_usage = 20000000000, enable_lazy_columns_replication = 1 ``` ### Second Query ``` sql WITH materialize(range(1000)) AS large_array SELECT count() FROM system.numbers WHERE NOT ignore( arrayMap(idx -> arraySlice(large_array, idx, 5), arrayEnumerate(large_array))) SETTINGS max_rows_to_read = 16384, read_overflow_mode = 'break', max_block_size = 512, max_threads = 1, max_memory_usage = 20000000000, enable_lazy_columns_replication = {1|0} ``` | Configuration | Rows | Time | Peak RSS | Result | |---|---|---|---|---| | Lazy ON, default block size, 262k rows | 327,045 | ~4.5 s (~72M lambda calls/s) | 3.15 GB | ✅ completes | | Lazy OFF, same settings | — | fails in 0.5 s | — | ❌ `MEMORY_LIMIT_EXCEEDED`: tries to allocate 121.83 GiB for one 65536-row block | | Lazy ON, `max_block_size=512`, 1 thread, 16384 rows | 16,384 | 0.28 s | 272 MB | ✅ | | Lazy OFF, same | 16,384 | 7.26 s | 2.28 GB | ✅ | Tests: a stateless test covering numeric/String/Nullable/Tuple elements, dynamic and constant offsets, negative and Nullable indexes, empty arrays, ARRAY JOIN and JOIN producers, UInt16 replication indexes, plus a performance test comparing both settings. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): arrayElement and arraySlice now read lazily replicated arrays (produced by ARRAY JOIN and JOIN when enable_lazy_columns_replication is enabled) directly, without materializing them. This greatly reduces memory usage and improves performance of queries that index or slice a large array.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112304",
          "createdAt": "2026-07-28T15:36:42Z",
          "updatedAt": "2026-08-13T11:11:30Z",
          "timestamp": "2026-08-13T11:11:30Z",
          "metrics": {
            "reactions": 2,
            "comments": 4
          },
          "labels": [
            "pr-performance"
          ],
          "author": "diegomestre2",
          "state": "open",
          "assignees": [
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1406c463ffd90971ba05",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114533",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114533",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong results for non-boolean conditions taken out of `and`",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112236 `and` implicitly converts its arguments to booleans, so any non-zero value is true. When a query plan optimization takes a part of a conjunction away and a single conjunct is left as the new predicate, that conjunct was converted with a cast to the type of the original predicate. A cast is not a boolean conversion: it maps values like 256 or 0.1 to 0, so rows whose condition value is a non-zero multiple of 256 silently disappeared. ```sql CREATE TABLE t (id UInt32) ENGINE = MergeTree ORDER BY id; INSERT INTO t SELECT number FROM numbers(600); SELECT count() FROM t AS l LEFT JOIN t AS r ON l.id = r.id WHERE r.id AND l.id = r.id; ``` returned 597 instead of 599, the rows with `id = 256` and `id = 512` were dropped. `mergeFilterIntoJoinCondition` moves `l.id = r.id` into the JOIN and leaves `CAST(r.id, 'UInt8')` as the filter: ``` Filter column: CAST(id AS UInt8) ``` The same happens in `ActionsDAG::removeUnusedConjunctions` when a conjunct is pushed down and the filter column is still needed in the result. That one is reachable without a JOIN, and the value of the condition was wrong there as well (the raw value instead of a boolean): ```sql SELECT count(), sum(f) FROM ( SELECT id, (id != 1000 AND s) AS f FROM (SELECT id, sum(id) AS s FROM t GROUP BY id) WHERE id != 1000 AND s ); ``` returned `597 597` instead of `599 599`. It only converted floating point types, now every non-boolean type is converted. Both places now wrap the remaining conjunct into `and(x, true)`, the same way `toBoolIfNeeded` does it in `JoinStepLogical.cpp`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results when a non-boolean condition, such as a bare integer column in `WHERE`, is left alone after the other conditions are merged into the JOIN condition or pushed down. Rows whose condition value was a non-zero multiple of 256 were skipped. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114533",
          "createdAt": "2026-08-12T18:40:57Z",
          "updatedAt": "2026-08-13T11:09:02Z",
          "timestamp": "2026-08-13T11:09:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ac254375dcf3ecaa118c",
        "signalId": "github:ClickHouse/ClickHouse:issue:114004",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114004",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Inconsistent AST formatting: `view((SELECT ...))` in table-function arguments loses subquery parentheses and cannot be parsed back (STID: 1941-1bfa)",
          "text": "🕵 Found by `AST fuzzer (amd_debug, targeted, old_compatibility)` on an unrelated PR (https://github.com/ClickHouse/ClickHouse/pull/91993, which only touches hex encoding): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=91993&sha=eae3a3cbbfb5d81bed061634d53c701167dce04a&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The fuzzer produced a query containing `view((SELECT ...))` as an argument of a comparison function nested inside `file(...)` table-function arguments. Formatting the AST drops the parentheses around the `view` subquery argument (`view((SELECT ...))` → `view(SELECT ...)`), and the formatted text cannot be parsed back, so the format-consistency check fails with a logical error (`abortOnFailedAssertion` in the debug build): ``` Logical error: 'Inconsistent AST formatting: the query: DESCRIBE TABLE file(concat(xor(if(indexHint(and(greaterOrEquals(alias5764._time, toLowCardinality(toNullable(1048575)) + number), lessOrEquals(alias5764._time, view((SELECT toUInt16OrDefault(65536, equals(minus(materialize(NULL, arrayElement(b)), greater(lowCardinalityIndices('\\\\\\\\', toLowCardinality(NULL)), -1)), toInt8OrZero(2147483646)))))))), a <= toInt256OrDefault(-2147483647), materialize(isNullable(NULL))), and(b >= 257, 1.1754943508222875e-38 > b)), currentDatabase(), toFixedString(toNullable('_04513.orc'), assumeNotNull(256)))) SAMPLE 100 / 5 SETTINGS input_format_orc_use_fast_decoder = 0 cannot parse query back from DESCRIBE TABLE file(concat(xor(if(indexHint(and(greaterOrEquals(alias5764._time, toLowCardinality(toNullable(1048575)) + number), lessOrEquals(alias5764._time, view(SELECT toUInt16OrDefault(65536, equals(minus(materialize(NULL, arrayElement(b)), greater(lowCardinalityIndices('\\\\\\\\', toLowCardinality(NULL)), -1)), toInt8OrZero(2147483646))))))), a <= toInt256OrDefault(-2147483647), materialize(isNullable(NULL))), and(b >= 257, 1.1754943508222875e-38 > b)), currentDatabase(), toFixedString(toNullable('_04513.orc'), assumeNotNull(256)))) SAMPLE 100 / 5 SETTINGS input_format_orc_use_fast_decoder = 0'. ``` The two texts differ only in the parentheses around the `view` subquery: the original has `view((SELECT ...))`, the re-formatted text has `view(SELECT ...)`. On a recent master-based **release** build (`clickhouse local`) the asymmetry is observable in the opposite direction — the parenthesized form is the one that does not parse inside table-function arguments: ```sql -- parses fine (special `view` handling in expression context): SELECT lessOrEquals(t, view((SELECT 1))); -- parses fine (unparenthesized subquery): SELECT * FROM file(if(indexHint(lessOrEquals(t, view(SELECT 1))), 'a', 'b')); -- SYNTAX_ERROR (parenthesized subquery inside table-function arguments): SELECT * FROM file(if(indexHint(lessOrEquals(t, view((SELECT 1)))), 'a', 'b')); ``` So whether `view((SELECT ...))` / `view(SELECT ...)` parses depends on context (plain expression vs. table-function argument) and on build/compatibility configuration, while the formatter always emits the unparenthesized form. Either the parser should accept both forms in all contexts where `view` is accepted at all, or the formatter should print the form that is guaranteed to parse back in the surrounding context. Previous distinct manifestations of this dedup bucket (STID: 1941-1bfa) were fixed individually: #109501 (back-quoted numeric type name, fixed by #109720), #106850, #106358, #100131. This `view` parenthesization case is a new one. CIDB shows this check failing on 5 unrelated PRs in the last 30 days (112921, 110886, 113533, 91993, 111061) and never on master — expected, since the AST fuzzer only runs per-PR with random queries.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114004",
          "createdAt": "2026-08-09T02:45:03Z",
          "updatedAt": "2026-08-13T11:07:08Z",
          "timestamp": "2026-08-13T11:07:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "fuzz",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:59dd96f9d1dbc86a3100",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104948",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104948",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Declarative function signatures, continuation of #3775",
          "text": "This is an experimental continuation of [#3775 (2018)](https://github.com/ClickHouse/ClickHouse/pull/3775) by @alexey-milovidov, which proposed a declarative way to describe function signatures so that argument validation and return-type inference can be driven by a small DSL instead of hand-written `getReturnTypeImpl` logic per function. ### What's in this branch A `getSignatureString()` method on `IFunction` and `IFunctionOverloadResolver`. When a function exposes a non-empty signature, the base `getReturnTypeImpl(ColumnsWithTypeAndName)` parses and applies it via a small grammar of type **matchers** (`UInt`, `Number`, `Array(T)`, `MaybeNullable(T)`, `Function((args), R)`, `T : Any` capture, …) and type **functions** (`leastSupertype`, `nativeNumber`, `aggregateFunctionReturnType(AggregateFunction(name, …))`, `subcolumnTypeOf`, `typeFromString`, `DateTime64(scale, tz)`, …). The DSL supports variadic positions (`…`), ellipsis grouping for repeated argument units (`T1, V1, …` repeats the pair), `OR` between alternatives, optional positions `[T]`, const-value capture (`const name String`), lambdas, etc. The grammar / parser / type matchers / type functions live under `src/DataTypes/FunctionSignature.h`, `FunctionSignature.cpp`, `TypeMatchers.cpp`, `TypeFunctions.cpp`. ### Coverage ~191 commits on top of master, each is a small per-family round titled `Function signatures: round N — …`. Current state on this branch: ``` SELECT count(*) AS total, countIf(signature != '') AS with_sig, round(countIf(signature != '') * 100.0 / count(*), 1) AS pct FROM system.functions WHERE NOT is_aggregate AND alias_to = '' AND origin = 'System'; 1416 1395 98.5 ``` (With \\`allow_experimental_nlp_functions = 1\\`; the few remaining unset ones are setting / config-gated functions like \\`aiClassify\\`, \\`region*\\`, \\`synonyms\\`, whose \\`create\\` throws unless the relevant config block / setting is present, so \\`system.functions.tryGet\\` returns null even though the signature is in source. With the right config + setting they all surface — verified in tmp/ test config.) Roughly: - **Authoritative** (DSL drives type-check and return type, the legacy `getReturnTypeImpl` is either gone or bypassed): higher-order array functions (`arrayMap`, `arrayFilter`, `arrayFirst*`, …), `toIntervalX`, comparison, `multiSearch*` / `multiMatch*`, vector L-norms / distances / dot product, `UUIDv7ToDateTime`, `arrayReduce` / `arrayReduceInRanges`, `reverse`, `mapKeys` / `mapValues`, `getSubcolumn`, the `least` / `greatest` resolver, paired-variadic `timeSeriesTagsToGroup` / `timeSeriesStoreTags`, … - **Documentation-only** (signature surfaced via `system.functions` but the legacy `getReturnTypeImpl` still runs because the result type uses promotion / widening / setting-dependent dispatch the current DSL can't express): arithmetic (`plus`, `minus`, `multiply`, `divide`, modulo / intDiv family), array widening (`arraySum` / `arrayCumSum*` / `arrayDifference`), `mapContains*Like`, `dateTrunc`, `toStartOfWeek`, `parseDateTime*`, `transform`, `range`, `toX` conversions, etc. There is a `signature_documentation` opt-in alongside `signature` in the binary-arithmetic, unary-arithmetic, FunctionArrayMapped, and FunctionMapToArrayAdapter families specifically so a function can advertise a signature without the DSL accidentally hijacking the legacy widening logic. ### Status Draft / experimental. - This is an experiment — I'm not asking for it to be merged. There's significant disruption (191 commits touching every function family), and the gain is mostly documentation: today only ~30% of the converted functions are actually DSL-authoritative; the rest are decorative. The arithmetic promotion matrix and the per-Op widening rules in particular would need a richer type-function vocabulary (or new matchers) before they can be expressed declaratively. - The branch has been kept rebased on master throughout and builds cleanly. I've run the stateless test suite on each round. The pre-existing-on-master test failures I hit are listed in commit messages of the rounds where I encountered them (none introduced by this work). - The original PR has been open since 2018 — this branch is intended as a concrete data point on \\\"what fraction of ClickHouse functions can be reasonably described by a declarative signature, and what would the DSL need to grow to cover the rest.\\\" That's the question I'd love feedback on. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104948",
          "createdAt": "2026-05-14T14:35:16Z",
          "updatedAt": "2026-08-13T11:07:00Z",
          "timestamp": "2026-08-13T11:07:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 38
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:20231e639dfc3fdcfb27",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114468",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114468",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix: error message for ALTER DROP COLUMN of a key column",
          "text": "<!-- PR title: Fix error message for ALTER DROP COLUMN of a column used in a key expression --> Closes: https://github.com/ClickHouse/ClickHouse/issues/71776 Related follow-up issues: #114481, #114181 `ALTER DROP COLUMN` and `ALTER CLEAR COLUMN` behaved inconsistently on a column that belongs to `ORDER BY`, `PRIMARY KEY` or `PARTITION BY` An `ALTER` is validated by building the new table schema first and inspecting it afterwards. Once the column is gone the sorting key cannot be resolved, so building the schema fails before the check that knows the real reason. That check is now done first, on the current schema, where the key columns are already known. It understands all the ways a key can refer to a column: the column itself, a column inside a key expression, a subcolumn (`ORDER BY a.x` for `a Tuple(...)`) and a whole nested group dropped by its common prefix Examples: 1. Basic ```sql CREATE TABLE t (a UInt64, b UInt64) ENGINE = MergeTree ORDER BY a; ALTER TABLE t DROP COLUMN a; -- before: Code: 47. Missing columns: 'a' while processing: 'a' ... (UNKNOWN_IDENTIFIER) -- after: Code: 524. Trying to ALTER DROP key a column which is a part of key expression. (ALTER_OF_COLUMN_IS_FORBIDDEN) ``` 2. Dropping a column that is used only in a TTL expression now also reports which TTL expression it breaks Note: `ALTER_OF_COLUMN_IS_FORBIDDEN` would be wrong here. Dropping a column used in a `TTL` is not always forbidden - `ALTER TABLE t DROP COLUMN d, MODIFY TTL a + INTERVAL 1 DAY` is valid and passes. And this is a wrapper not a check: the same step returns `ILLEGAL_TYPE_OF_ARGUMENT` for `MODIFY COLUMN d Array(UInt8)`, which `ALTER_OF_COLUMN_IS_FORBIDDEN` would mislabel ```sql CREATE TABLE t (d Date, a UInt64) ENGINE = MergeTree ORDER BY a TTL d + INTERVAL 1 DAY; ALTER TABLE t DROP COLUMN d; -- before: Code: 47. Missing columns: 'd' ... (UNKNOWN_IDENTIFIER) -- after: Code: 47. Cannot apply ALTER because it breaks the TTL of the table: Missing columns: 'd' ... (UNKNOWN_IDENTIFIER) ``` 3. For a key column the command is rejected whatever the partition is, so the old error only sent the user to fix something that would not help ```sql CREATE TABLE t (a UInt64, b UInt64, c UInt64) ENGINE = MergeTree PARTITION BY b ORDER BY a; ALTER TABLE t CLEAR COLUMN a IN PARTITION 'nonsense'; -- before: Code: 53. Cannot convert string 'nonsense' to type UInt64. (TYPE_MISMATCH) -- after: Code: 524. Trying to ALTER CLEAR key a column ... (ALTER_OF_COLUMN_IS_FORBIDDEN) ``` 4. With `share_nested_offsets` enabled, a name that is not a column of the table denotes the whole nested group `<name>.*`, and `DROP`/`CLEAR` of the group is now rejected the same way when any column of the group is used in a key. Before, this was worse than a confusing message: the checks compare names exactly and did not see the group, so `DROP COLUMN IF EXISTS n` destroyed the group's data with a mutation that then failed halfway and `CLEAR COLUMN n` on a non-empty table left a mutation that can never finish (a key column cannot be rewritten) in `system.mutations` until `KILL MUTATION` ```sql CREATE TABLE t (`n.a` UInt64, `n.b` UInt64, x UInt64) ENGINE = MergeTree ORDER BY `n.a`; ALTER TABLE t DROP COLUMN n; -- before: Code: 47. Missing columns: 'n.a' while processing: '`n.a`' ... (UNKNOWN_IDENTIFIER) -- after: Code: 524. Trying to ALTER DROP whole Nested group n whose columns (`n.a`) are part of key expression. (ALTER_OF_COLUMN_IS_FORBIDDEN) ALTER TABLE t CLEAR COLUMN n; -- before: accepted, data of `n.a` and `n.b` is zeroed and the mutation stays in system.mutations forever -- after: Code: 524. Trying to ALTER CLEAR whole Nested group n whose columns (`n.a`) are part of key expression. (ALTER_OF_COLUMN_IS_FORBIDDEN) ``` Groups with no key columns inside are dropped as before, and with `share_nested_offsets = 0` the name does not denote the group and the check does not apply. The remaining group/exact-name mismatches (silent data destruction for groups without key columns, dependent views, races with background mutations) predate this PR and are reported separately with reproductions ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... DROP COLUMN` of a column that is used in the sorting, primary or partition key now fails with `ALTER_OF_COLUMN_IS_FORBIDDEN` and an explanation, the same as `ALTER TABLE ... CLEAR COLUMN` does. Previously it failed with a confusing `UNKNOWN_IDENTIFIER: Missing columns` error coming from the recalculation of the key expressions. The same applies to dropping or clearing a whole `Nested` group by its common prefix when a column of the group is used in a key; previously that could silently destroy the group's data or leave a mutation that never finishes. Dropping a column that is used only in a `TTL` expression now also reports which TTL expression it breaks",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114468",
          "createdAt": "2026-08-12T11:43:38Z",
          "updatedAt": "2026-08-13T11:04:29Z",
          "timestamp": "2026-08-13T11:04:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "m7kss1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7686a6d02052629c079b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114613",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114613",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `jsonPathValues` tokenizer for JSON text indexes",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 Related: https://github.com/ClickHouse/ClickHouse/pull/110757 I first tried to fix the edge cases between `JSONAllValues` and the text index. I kept finding more cases where the index could return an incorrect result. I therefore needed to disable more useful index paths to keep queries correct. The main problem is that `JSONAllValues` stores plain text. It does not retain the path or type of values from `Dynamic` JSON columns. PR #113376 documents several examples. A tokenizer made specifically for `JSON` seemed like a better fit. The new `jsonPathValues` tokenizer stores the JSON path, type, and value in each token. This gives the index enough information for safe typed comparisons. ```text +-------------------+-------+---------------------+-------+--------+------------------+ | escaped JSON path | 00 00 | binary-encoded type | 00 00 | 1-byte | payload | | | | | | kind | | +-------------------+-------+---------------------+-------+--------+------------------+ Payload: complete value : | full value | truncated value : | value prefix | SipHash-2-4 (8 bytes, LE) | map entry : | escaped key | 00 00 | complete/truncated value | validation : | empty | Kinds: 1/2 = scalar, 3/4 = array element, 5/6 = map entry (complete/truncated), 7 = dynamic validation Escaping: 00 -> 00 01 Component terminator: 00 00 ``` The path-value format also enables direct reads. This part was inspired by [ClickStack's use of text-index direct reads for dynamic map attributes](https://clickhouse.com/blog/making-clickstack-5x-faster-clickhouse-observability). `jsonPathValues` brings this model directly to `JSON` columns. It does not require user-defined alias columns or query rewrites. Long values use a bounded prefix and a hash. ClickHouse validates candidate rows when a token is truncated or a runtime type is unsafe. This prevents false-negative results. The tokenizer supports declared and dynamic paths, arrays, scalar leaves in `Array(JSON)`, and declared `Map(String, String)` paths. It supports equality, `IN`, `has`, map lookups, prefix and suffix searches, `LIKE`, `ILIKE`, regular expressions, and multi-search functions. `startsWith` uses an ordered dictionary lookup, so it does not require a full dictionary scan. Substring searches still scan the dictionary. ## Benchmark We tested `jsonPathValues(1024)` with 99,999,984 distinct events from 100 JSONBench files. Both target tables had the same 192-part layout. Each indexed query returned the same result as the query without an index. | Metric | No index | `jsonPathValues(1024)` | |---|---:|---:| | Build time | 214 s | 863 s | | Total storage | 9.29 GiB | 23.48 GiB | | Exact `cid` lookup | 1,926 ms | 13 ms | | Exact long-value lookup | 637 ms | 8 ms | | `startsWith` | 29 ms | 16 ms | | Substring `LIKE` | 145 ms | 93 ms | The index used 14.19 GiB, or approximately 152 bytes per JSON object. It made the table 2.53 times larger and made index construction 4.03 times slower. The exact lookups were 80 to 148 times faster. `startsWith` was 1.81 times faster. Query times are medians from ten warm runs. This PR also adds a checked-in performance test with 500,000 generated JSON rows. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `jsonPathValues` tokenizer for bounded, type-safe text indexes on `JSON` paths, arrays, and maps.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114613",
          "createdAt": "2026-08-13T09:35:07Z",
          "updatedAt": "2026-08-13T11:00:57Z",
          "timestamp": "2026-08-13T11:00:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5784d28c8b14c7ab2fc9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114590",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114590",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: JSON shared data at the max 256 buckets is untested end-to-end",
          "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #112172](https://github.com/ClickHouse/ClickHouse/pull/112172). That PR: Rewrites the write path of bucketed `JSON` shared data. `splitSharedDataPathsToBuckets` and `flattenAndBucketSharedDataPaths` are replaced by a single `SharedDataBucketsSplitter` that computes each path's bucket once (cached in `PODArray<UInt8>`) plus per-bucket byte sizes, then hands out one … **1. JSON shared data at the max 256 buckets is untested end-to-end** `src/DataTypes/Serializations/SerializationObjectHelpers.cpp:88`, `src/DataTypes/Serializations/SerializationObjectHelpers.cpp:119` **Risk:** This PR replaces `splitSharedDataPathsToBuckets` (which kept the bucket index in a `size_t`) with `SharedDataBucketsSplitter`, which caches the bucket of every path in a `PODArray<UInt8>` via `static_cast<UInt8>(bucket)` (`SerializationObjectHelpers.cpp:88`). The value range now exactly saturates … **Unique vs PR tests:** The PR's own `04617_array_json_shared_data_buckets_insert_memory.sh` is a memory-limit test at the default bucket counts (8/32) that only asserts `count()` and is skipped on every sanitizer build. This test pins the maximum legal bucket count (256), which no test in either tree uses, and asserts … [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/4753df1e-4460-4e86-9262-9fdbe47e2cdb) cc @Avogar (author of #112172), @thevar1able (merged/approved #112172) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114590",
          "createdAt": "2026-08-13T04:14:22Z",
          "updatedAt": "2026-08-13T11:00:48Z",
          "timestamp": "2026-08-13T11:00:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d3e61270a94c109a9440",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112667",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112667",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Silk integration",
          "text": "Splits the silk runtime integration out of https://github.com/ClickHouse/ClickHouse/pull/111275, so that it can be reviewed on its own. This adds the plumbing that lets ClickHouse run work on [silk](https://github.com/ClickHouse/silk) fibers, without yet putting any subsystem on them. - `Silk::initializeFiberScheduler` / `Silk::destroyFiberScheduler`, called by the server when the `enable_silk_runtime` server setting is enabled. The fiber stack size is configurable through the `silk.fiber_stack_size` configuration key (320 KiB by default, which leaves enough room for OpenSSL handshakes). - `FiberLocal` - fiber-local storage. A fiber can migrate between operating-system threads, so it must not observe another fiber's `thread_local` state; the values of the registered slots are swapped in and out on every fiber switch instead. `current_thread` (`ThreadStatus`), the OpenTelemetry tracing context, and the memory-tracker and exception blockers are moved to it. - The silk thread-local-storage sanitizer: an LLVM pass in `utils/silk-thread-local-storage-sanitizer` that instruments every `thread_local` access and aborts when a fiber touches raw thread-local storage. Without it, a variable that was not migrated to `FiberLocal` produces silent corruption rather than a diagnostic. It is enabled in the debug and ASan CI builds. - `Silk::ConnectionPool` and `Silk::streamSocketFactory` - a `Connection` pool and a socket factory that suspend the calling fiber instead of blocking the operating-system thread. `PoolBase` and `ConnectionPool` are templated on the lock and the condition variable to make that possible, and `ConnectionPool` stays an alias of the `std::mutex` instantiation, so the existing call sites are unchanged. - Memory that the runtime maps outside the C++ heap - fiber stacks and `io_uring` rings - is charged to `total_memory_tracker` through silk's mmap accounting hooks. - The low-level silk runtime counters are exported to `system.asynchronous_metrics` under a `Silk` prefix. - `Common/Fiber.h` and `Common/FiberStack.h` are renamed to `Common/StackfulCoroutine.h` and `Common/CoroutineStack.h`. They implement the boost-context coroutines used by `AsyncTaskExecutor`, which are unrelated to silk fibers, and having two different things called \"fiber\" in the same codebase is confusing. Related: https://github.com/ClickHouse/ClickHouse/pull/111275 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental support for the [silk](https://github.com/ClickHouse/silk) fiber runtime, enabled with the `enable_silk_runtime` server setting. When it is enabled, the server initializes the silk fiber scheduler at startup, so that subsystems supporting it can run their jobs on fibers instead of occupying an operating-system thread while waiting for I/O.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112667",
          "createdAt": "2026-07-30T21:32:29Z",
          "updatedAt": "2026-08-13T11:00:33Z",
          "timestamp": "2026-08-13T11:00:33Z",
          "metrics": {
            "reactions": 1,
            "comments": 1
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "CheSema"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ec9c71b3850d24462ab4",
        "signalId": "github:ClickHouse/ClickHouse:issue:111898",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:111898",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "DPsub join-order reordering with `query_plan_enable_optimizations = 0` silently drops a non-equi `JOIN ON` conjunct on chained joins (post-#109638)",
          "text": "## Describe what's wrong With `query_plan_optimize_join_order_algorithm = 'dpsub'` and `query_plan_enable_optimizations = 0`, a chained join whose first `ON` clause carries a non-equi conjunct returns wrong results: the DPsub reordering still runs (it restructures the join tree even though plan optimizations are disabled), but the residual, non-equi part of the `ON` condition is silently dropped — rows that fail it come back as matched. This is the surviving corner of #109617: the fix (#109638, merged 2026-07-07) repairs the default path, but with `query_plan_enable_optimizations = 0` the same query still returns the pre-fix wrong result on current master. Settings must not change results — either the reordering should not run at all when plan optimizations are disabled, or it must preserve the residual predicate. ## How to reproduce Masters `26.8.1.102` (`eb577ff401a3`, 2026-07-24) and `26.7.1.1169` (`38334977be1e`, 2026-07-18), official builds (reduced from `tests/queries/0_stateless/02372_analyzer_join`, same data as #109617): ```sql CREATE TABLE t1 (id UInt64, value String) ENGINE = MergeTree ORDER BY tuple(); CREATE TABLE t2 (id UInt64, value String) ENGINE = MergeTree ORDER BY tuple(); CREATE TABLE t3 (id UInt64, value String) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t1 VALUES (0, 'Join_1_Value_0'), (1, 'Join_1_Value_1'), (2, 'Join_1_Value_2'); INSERT INTO t2 VALUES (0, 'Join_2_Value_0'), (1, 'Join_2_Value_1'), (3, 'Join_2_Value_3'); INSERT INTO t3 VALUES (0, 'Join_3_Value_0'), (1, 'Join_3_Value_1'), (4, 'Join_3_Value_4'); SELECT t1.id, t1.value, t2.id, t2.value, t3.id, t3.value FROM t1 INNER JOIN t2 ON t1.id = t2.id AND t1.value = 'Join_1_Value_0' LEFT JOIN t3 ON t2.id = t3.id ORDER BY ALL SETTINGS query_plan_optimize_join_order_algorithm = 'dpsub', query_plan_enable_optimizations = 0; ``` Observed (both versions, deterministic 5/5): ``` 0 Join_1_Value_0 0 Join_2_Value_0 0 Join_3_Value_0 1 Join_1_Value_1 1 Join_2_Value_1 1 Join_3_Value_1 ``` Expected (and returned at default settings): only the first row — for `id = 1`, `t1.value = 'Join_1_Value_0'` is false, so the `INNER JOIN` must not match. The second row is the `ON` conjunct being ignored. The trigger surface: - only `dpsub` misbehaves; `greedy` and `dpsize` return the correct result with `query_plan_enable_optimizations = 0`; - either co-factor alone is fine: `dpsub` with optimizations on is correct (the #109638 fix), and `query_plan_enable_optimizations = 0` with the default algorithm is correct; - the chain is required — the two-table `t1 JOIN t2 ON t1.id = t2.id AND t1.value = '...'` alone is correct; - a `RIGHT JOIN` first hop misbehaves the same way (matched rows appear where the unmatched-side defaults are expected); - `EXPLAIN` shows the reordering did run despite optimizations being disabled: the plan becomes `t1 ⋈ (t2 ⋈ t3)` and no filter step carries `t1.value = 'Join_1_Value_0'` anywhere. ## Expected behavior The same rows as with default settings. `query_plan_optimize_join_order_algorithm` and `query_plan_enable_optimizations` are plan-shape settings and must never change query results. Workaround: `query_plan_enable_optimizations = 1` (default), or a different join-order algorithm. Related: https://github.com/ClickHouse/ClickHouse/issues/109617 Related: https://github.com/ClickHouse/ClickHouse/pull/109638 Found by an automated optimizer-correctness differential tester (unoptimized-vs-optimized `opt_toggle` oracle on a stateless-suite replay).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/111898",
          "createdAt": "2026-07-25T12:58:39Z",
          "updatedAt": "2026-08-13T10:58:43Z",
          "timestamp": "2026-08-13T10:58:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "comp-joins",
            "comp-query-optimizer",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "fkastrati"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ed8082639dbc3277d041",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114475",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114475",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113289 to 26.5: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114475",
          "createdAt": "2026-08-12T12:06:09Z",
          "updatedAt": "2026-08-13T10:58:02Z",
          "timestamp": "2026-08-13T10:58:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll3",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:96d0ffca511a1d3ff51f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113868",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113868",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Add aggregate function `gini`\"",
          "text": "Reverts ClickHouse/ClickHouse#112280 Closes: https://github.com/ClickHouse/ClickHouse/issues/113763 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Besides adding `gini`, #112280 added nine `getArgumentsThatCanBeOnlyNull` overrides across eight combinator files and made `-If` union the nested set, so those combinators forward that set from the function they wrap. The `Null` combinator uses that set to decide whether an argument that can only be `NULL` folds the aggregate into `AggregateFunctionNothing`. Forwarding it removes the fold one level up, so expressions chaining one of them on top of `-If` changed type and value. Measured on debug builds of this branch and its parent: | expression | before #112280 | after | |---|---|---| | `countIfOrNull(number, NULL)` | `UInt64` `0` | `Nullable(UInt64)` `NULL` | | `sumIfResample(0, 2, 1)(number, NULL, number % 2)` | `Nullable(Nothing)` `NULL` | `Array(UInt64)` `[0, 0]` | | `sumIfState(number, NULL)` | folded `Nullable(Nothing)` | `AggregateFunction(sumIf, UInt64, Nullable(Nothing))` | Two of the new overrides sit outside the `-If` family, reaching expressions with no `-If` at all. A partial revert is not available: `gini` declares argument 0 in `getArgumentsThatCanBeOnlyNull` itself, and one forwarding path carries both that declaration and the `-If` filter index, so keeping `gini` without the propagation contradicts its own test. `gini` is unreleased (the merge is not an ancestor of 26.7, 26.6, 26.5, 26.3 or 25.8, and never reached `CHANGELOG.md`), so this withdraws no released behaviour. A behaviour-neutral re-land following the `sum` family, as @Manerone specified, comes separately. The docs page is removed here too. It did not come from the reverted merge: `2d4afadac21e7b4` moved it into the live Mintlify tree, so it arrived via the master merge. Its navigation entry, legacy redirect and slug-map row go with it, or those would point at a missing target. Regenerating is not an option: the aggregate family in `autogenerate_docs.py` lists that page directory and carries `skip_if_empty`, so an unregistered function's page is never revisited. Verified by byte identity rather than a new test: of the 14 source and test paths the merge touched, 13 match their pre-merge blobs, and `registerAggregateFunctions.cpp` differs only by unrelated `MergedJSONPatch` lines master added in `e3698631b023165`. cc @Manerone <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1317` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113868",
          "createdAt": "2026-08-07T18:25:03Z",
          "updatedAt": "2026-08-13T10:57:09Z",
          "timestamp": "2026-08-13T10:57:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-not-for-changelog",
            "manual approve",
            "can be tested",
            "pr-synced-to-cloud",
            "pr-autogenerated-docs"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "Manerone"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6a5f3162a3061b3d7e54",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114589",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114589",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "sync ErrorCodes.cpp from private",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1318` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114589",
          "createdAt": "2026-08-13T04:01:46Z",
          "updatedAt": "2026-08-13T10:57:05Z",
          "timestamp": "2026-08-13T10:57:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "chhetripradeep",
          "state": "closed",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a972495dbb1334f8d4c9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114606",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114606",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "PostgreSQL: allow an empty TLS contents override when the collection stores no credential",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113947 Related: https://github.com/ClickHouse/ClickHouse/pull/110615 Port of the #113947 narrowing (MySQL) to `StoragePostgreSQL::getSSLParams`. An empty `ssl*_pem` query override is rejected only when it would actually drop a TLS credential the named collection carries — a path (a query cannot override path keys, so the value read is the collection's own) or contents, read in the pre-override form via `NamedCollection::getValueBeforeQueryOverride`. On a collection with no TLS keys at all, the empty override stays the no-op it is for the direct arguments, instead of throwing `BAD_ARGUMENTS`. The rejection cases are unchanged and remain covered by `test_path_overrides_are_rejected` and `test_tls_credentials_in_sql_named_collection`; the new `test_empty_override_without_stored_credential_is_noop` covers the no-op case (it fails without the code change). ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix of the unreleased #110615: an empty `ssl*_pem` override on a PostgreSQL named collection without TLS credentials is a no-op again instead of an error.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114606",
          "createdAt": "2026-08-13T09:11:09Z",
          "updatedAt": "2026-08-13T10:56:42Z",
          "timestamp": "2026-08-13T10:56:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:de1fe8f4e4d9a00bd9a7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113450",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113450",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113337 Related: https://github.com/ClickHouse/ClickHouse/pull/113425 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables that contain a nullable `ARRAY` or `MAP` column. Such a table could not be read at all, because the schema mapper wrapped the composite type in `Nullable`, which ClickHouse forbids, so both `DESC` and `SELECT` failed with `Nested type Array(Nullable(Int32)) cannot be inside Nullable type`. A nullable composite column is now mapped to a non-`Nullable` composite type and a `NULL` value is read as an empty one. Closes #113337. ### Description This takes over #113425 by @zlareb1, who closed it and asked me to carry it forward. The diagnosis and the fixture are his; the source change uses the in-tree capability gate rather than deleting the wrap. **What breaks.** Paimon columns are nullable by default, so a plain `CREATE TABLE paimon.default.t (f ARRAY<INT>)` from Spark produces an unreadable table. `Paimon::DataType::parse` applied its `if (nullable)` wrap in the composite branch as well as the scalar one, and `DataTypeArray`/`DataTypeMap` return `canBeInsideNullable() == false`, so `DataTypeNullable`'s constructor threw. This happens while parsing the schema, so it takes out the whole table rather than one column. Affects `paimonS3`/`paimonLocal`/`paimonAzure`, their `*Cluster` variants, the `Paimon*` engines and the REST catalog. The engines need `allow_experimental_paimon_storage_engine`; the table functions do not. **The change.** The two inner wraps become a single `makeNullableSafe`, which wraps only when the type permits it. Neither the Iceberg nor the DeltaLake schema processor wraps a composite in `Nullable` (Iceberg gates on `canBeInsideNullable()`; DeltaLake keeps the wrap in its scalar branches only), so this aligns Paimon with them. The scalar wrap is untouched, so inner nullability survives: the fixture reads as `Array(Nullable(Int32))` and `Map(String, Nullable(Int32))`. The gate, rather than deletion, keeps the wrap available for a future `ROW`, whose `DataTypeTuple` does permit it. A `NULL` composite reads as an empty one, the Parquet reader's documented behaviour, so the two become indistinguishable. That is forced by the type system and matches Iceberg, DeltaLake, Arrow, ORC and Avro. **Validation.** New test `04757_paimon_nullable_composite_types` over @zlareb1's fixture fails on master with the error above and passes with the fix. The ten existing Paimon tests are identical on both binaries. Restoring either wrap individually reddens the new test on its own message. 50/50 randomized runs pass. `Types.h` is absent on 25.8, so 26.3 through 26.7 are affected; `must-backport` labels look appropriate.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113450",
          "createdAt": "2026-08-05T10:05:17Z",
          "updatedAt": "2026-08-13T10:55:45Z",
          "timestamp": "2026-08-13T10:55:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.5-must-backport"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c1c622dd723665a7be3d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113334",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113334",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Introduce zk leader metrics to Keeper mntr",
          "text": "Adds leader-only Keeper `mntr` metrics: - `zk_leader_uptime` - `zk_sum_election_time` - `zk_cnt_election_time` - `zk_sum_leader_unavailable_time` - `zk_cnt_leader_unavailable_time` `zk_leader_uptime` starts at NuRaft `BecomeLeader`. The cumulative election and leader-unavailability metrics sample `isLeaderAlive` once per `heart_beat_interval_ms`; election completion is recorded at `BecomeLeader`, while leader-unavailability completes when polling observes a live local leader. `srst` resets all four cumulative values. The intended alternative was exact NuRaft lifecycle callbacks: an election-start callback before pre-vote or vote, plus a leader-ready callback for leader-unavailability completion. That requires changing vendored NuRaft and an upstream contribution, whose merge timeline could delay these Keeper monitoring metrics. This PR therefore delivers the metrics now through existing NuRaft state, while documenting the sampling limitation: boundaries can differ by up to one heartbeat interval and short no-leader windows can be missed. Also added keeper-only metrics - `KeeperLastLeaderElectionTime` - `KeeperLastLeaderUnavailableTime` ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Introduce emulated ZooKeeper metrics in Keeper `mntr` command: `zk_leader_uptime`, `zk_sum_election_time`, `zk_cnt_election_time`, `zk_sum_leader_unavailable_time`, `zk_cnt_leader_unavailable_time`; Also introduce related leader-oriented metrics for Keeper only: `KeeperLastLeaderElectionTime`, `KeeperLastLeaderUnavailableTime`",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113334",
          "createdAt": "2026-08-04T14:08:25Z",
          "updatedAt": "2026-08-13T10:50:45Z",
          "timestamp": "2026-08-13T10:50:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested",
            "pr-autogenerated-docs"
          ],
          "author": "UberDever",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:deb12f49fe5f9babf46f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114615",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114615",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "diff-review skill second edition",
          "text": "Reworks the `diff-review` skill in `.claude/skills/diff-review` — the local in-browser diff review that Claude Code serves before a commit or a PR. Heads-up: I tailored this to how *I* like to review, so please treat it as a suggestion rather than a standard. Check the branch out, run it on one of your own diffs, and keep it only if you find it helpful. What changed: - **Multi-round review.** Sending a round no longer ends the review: the server stays up, the agent works on the comments it was handed while you keep reading, and each comment turns green in the page as it gets addressed. - **Comments are durable state.** They are written to the `--out` file as you type, so they survive a server restart or a page reload, and after the agent's fixes move the code they are relocated by their anchor line instead of pointing at the wrong place. Resolutions, replies and dismissals are part of that state. - **Explicit review range.** `--staged`, `--head <sha>` and `--committed` alongside `--base`, so a review of recorded work never picks up local edits; the header always states which range is on screen. - **Navigation.** Directory tree with per-file status, `+a −d` counts and open-comment badges, a path filter, an all-files page, whole-file mode with expandable folded regions, a Comments pane listing open and already-addressed comments, draggable sidebar and split, keyboard shortcuts with a `?` overlay, and light / dark / system themes. - **Tests.** `ui.html` is split into ES modules under `ui/`, covered by `ui_test.mjs`, and `persist_test.mjs` exercises the whole persistence round-trip against real servers on a throwaway repository. - Vendored `@pierre/diffs` bumped from 1.2.12 to 1.3.5. Local agent tooling only — nothing in the server, the build or the tests. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114615",
          "createdAt": "2026-08-13T09:54:07Z",
          "updatedAt": "2026-08-13T10:50:32Z",
          "timestamp": "2026-08-13T10:50:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e83915b4c2563b931b28",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114187",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114187",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Rework the comments in the Web UI",
          "text": "Make the comments in `programs/server/play.html` concise and on point: state the constraint and the failure it prevents in simple terms, and drop meta-narrative and repeated details (the file shrinks from 13416 to 11810 lines). Stale comments are fixed along the way: the syntax highlighting no longer uses `mix-blend-mode: difference` (the textarea text is transparent over the backdrop), the active tab no longer erases the strip's border with a box-shadow (its `::after` does), and the roadmap listed already-implemented items. A few code simplifications found during the pass (verified with a comment-stripped diff to be the only code changes): - Define `@keyframes hourglass-animation` once and use it everywhere: the databases-panel hourglasses referenced it while only `tab-hourglass-animation` was defined, so they never animated. - Remove a redundant `position: relative` in `.tab.active` and a pointless local alias of `setQueryOrRun`. - Put visible characters directly into string literals (📌, 🌈, №, Σ, ↻, ⧗) instead of `\\u` escapes. The `test_play_reconcile_startup` harness (which runs the real script extracted from `play.html`) passes all scenarios. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Web UI: the hourglass loading indicator in the databases panel is animated again.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114187",
          "createdAt": "2026-08-10T16:50:15Z",
          "updatedAt": "2026-08-13T10:47:52Z",
          "timestamp": "2026-08-13T10:47:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:34c056cf7f5490a4f7b4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:97032",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:97032",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Implement Prometheus /api/v1/series, /labels, /label/values endpoints",
          "text": "Implement the three remaining Prometheus HTTP API metadata endpoints for the `TimeSeries` engine, which previously threw `NOT_IMPLEMENTED`. These endpoints are required by Grafana's Prometheus data source for label autocomplete and the metric browser. - `/api/v1/series`: returns all series with their full label sets by querying the tags table, filtered by the `match[]` series selectors. Each row is normalized to its logical Prometheus label set (a `tags_to_columns` tag can live in the dedicated column or in the residual `tags` `Map` of a supported external table) with the same rules as `timeSeriesStoreTags`: both carriers are merged, exact duplicates collapse, empty values are dropped, and conflicting carrier values are rejected with `bad_data`, so deduplication is by series identity rather than by the raw table layout. - `/api/v1/labels`: returns all distinct label names from the tags `Map` column and from the `tags_to_columns` columns, always including `__name__` as a virtual label. An empty label value means the label is absent, so a map entry with an empty value (possible in a supported external `tags` table, e.g. `tags = {'env': ''}`) does not surface its key as a label name and does not count toward `limit`, consistently with the other endpoints. - `/api/v1/label/<name>/values`: returns the distinct values for a specific label, reading the `metric_name` column for `__name__`, the dedicated `tags_to_columns` column (falling back to the residual `tags` `Map`) for a configured tag, or the tags `Map` otherwise. Both of these endpoints validate the carriers of the labels they materialize the same way `/api/v1/series` and the query endpoints do, so a row holding different non-empty values for one tag in its dedicated column and in the residual `tags` `Map` is rejected with `bad_data` instead of being silently reported. The optional `limit` parameter is supported on all the endpoints, and a present-but-empty `limit=`, `start=`, or `end=` is rejected with `bad_data` rather than being treated as an omitted parameter. The limit is enforced in the emission layer rather than as a SQL `LIMIT`, so the carrier validation above sees every matched row and a corrupted row cannot be hidden by a small `limit` and a favorable row order. The optional `start`/`end` parameters filter by the `min_time`/`max_time` columns of the `tags` table and are accepted only when both `store_min_time_and_max_time` and `filter_by_min_time_and_max_time` are enabled - the same gate as the query path's tags-table prefilter. When either setting is disabled (including an external `tags` table that still physically has the bound columns while `store_min_time_and_max_time = 0`), a ranged metadata request is rejected instead of silently diverging from `/api/v1/query` and `/api/v1/query_range`. The `match[]` parameter accepts full Prometheus series selectors (a bare metric name or an instant selector with `=`, `!=`, `=~`, `!~` label matchers, e.g. `cpu_usage{host=\"server1\"}`), parsed with the same PromQL parser as `/api/v1/query` and translated into the same `tags`-table filter the query endpoints use, so all three metadata endpoints select exactly the series the query endpoints would read. A repeated `match[]` is the union of the selectors, as in Prometheus. A matcher on a `tags_to_columns` tag resolves the tag with the same carrier normalization as the endpoints themselves: the dedicated column wins when non-empty (NULL is normalized to the empty label value), the residual `tags` `Map` is used otherwise, and a row carrying different non-empty values in the two carriers is rejected with `bad_data` instead of silently preferring one of them. Regexp matchers are fully anchored (`^(?:...)$`), and non-legacy (UTF-8) label names are written as quoted string literals, e.g. `{\"http.status_code\"=\"200\"}`, both as in Prometheus. A selector whose matchers all match the empty label value (e.g. `{host=~\".*\"}`) is rejected with `bad_data`, as in Prometheus - at least one matcher must not match the empty string, so a `match[]` cannot degenerate into an unfiltered scan - and an invalid regexp anywhere in a selector is likewise rejected with `bad_data` at parse time. Label names and values are JSON-escaped on output and string literals interpolated into the generated SQL are quoted, so values containing quotes, backslashes, or control characters are handled correctly. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Implemented the Prometheus `/api/v1/series`, `/api/v1/labels`, and `/api/v1/label/<name>/values` metadata endpoints for the `TimeSeries` engine. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/97032",
          "createdAt": "2026-02-15T21:08:25Z",
          "updatedAt": "2026-08-13T10:47:46Z",
          "timestamp": "2026-08-13T10:47:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 49
          },
          "labels": [
            "pr-feature",
            "submodule changed",
            "manual approve",
            "can be tested",
            "comp-promql",
            "pr-autogenerated-docs"
          ],
          "author": "ajonkisz",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "nikitamikhaylov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ab51fb2c54692c75bfe3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114177",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114177",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a crash when a primary-key range layer produces an empty pipe",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed a server crash when reading a `MergeTree` table by primary-key range layers and one of the layers produced an empty pipe. `readByLayers` in `PartsSplitter` applies the per-layer border filter to every pipe returned by the reading step getter, and such a pipe can be empty. `applyRangeFilterFromAST` then calls `Pipe::getHeader`, which dereferences the pipe's null `header`. The ASan+UBSan build reports `reference binding to null pointer of type 'element_type' (aka 'const DB::Block')`, the TSan builds a segmentation fault. An empty pipe carries nothing to filter and is dropped by the consumers anyway (`Pipe::unitePipes` starts with `removeEmptyPipes`), so it is skipped now, exactly like a null filter AST. Found by the AST fuzzer in CI on two unrelated pull requests, so it is not caused by either of them: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=99639&sha=e3675c072cc6c7cb1184a060d9c872ac519c78a2&name_0=PR&name_1=Stress%20test%20%28amd_asan_ubsan%29 (see https://github.com/ClickHouse/ClickHouse/pull/99639) and https://github.com/ClickHouse/ClickHouse/pull/103182. There is no deterministic reproducer - the fuzzed query is not recoverable from the logs, because the server dies before it is logged - so no test is added. Closes: https://github.com/ClickHouse/ClickHouse/issues/114176",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114177",
          "createdAt": "2026-08-10T15:31:49Z",
          "updatedAt": "2026-08-13T10:47:24Z",
          "timestamp": "2026-08-13T10:47:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:eb7df5ce334adc018013",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114562",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114562",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix NO_SUCH_COLUMN_IN_TABLE naming a column absent from the table when a part's columns were all renamed or all dropped",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/103654 Closes: https://github.com/ClickHouse/ClickHouse/issues/103116 Related: https://github.com/ClickHouse/ClickHouse/pull/96675 Related: https://github.com/ClickHouse/ClickHouse/issues/79110 ## Problem `NO_SUCH_COLUMN_IN_TABLE` naming a column that is **not** in the table, when reading or merging a part whose columns the metadata no longer knows under those names. Two reported shapes, both with a mutation queued but not yet applied: - every column renamed, and a merge has to read a column missing from the part (#103654) — the merge then fails and is retried forever, so the table stops compacting: ``` Code: 16. DB::Exception: There is no column h in table. (NO_SUCH_COLUMN_IN_TABLE) (in query: OPTIMIZE TABLE t_rename_vm FINAL;) ``` - every column dropped (#103116) — a plain `SELECT` fails, no merge needed: ``` Code: 16. DB::Exception: There is no column a in table. (NO_SUCH_COLUMN_IN_TABLE) (in query: SELECT c FROM t_all_dropped ORDER BY c;) ``` ## Root cause `injectRequiredColumns` injects the smallest physical column of a part when none of the requested columns is physically present, only to learn the number of rows. That column must be **both** readable from the part **and** resolvable by the caller, which looks names up in the metadata. The candidate set kept being taken from one side or the other — which guarantees only one of the two properties — and each fix patched the other side: | commit | candidate set | fixed | broke | |---|---|---|---| | `5c1db5fc661` | metadata columns | unresolvable names | part may have no file for any of them → `LOGICAL_ERROR` in `getColumnNameWithMinimumCompressedSize` | | `6d0b4dc9885` | part columns | that `LOGICAL_ERROR` | part name may be absent from the metadata → `NO_SUCH_COLUMN_IN_TABLE` (#96675) | | `1d2e2d1c6b2` (#96675) | intersection, else part columns | both, when non-empty | the empty case reverts to part-only names → **#103116**; the intersection compares *part* names against *metadata* names, so renamed columns drop out → **#103654** | Two things kept this going: - **Wrong name space.** The rest of the function works in metadata names and maps to part names only for the part lookup (`injectRequiredColumnsRecursively`, lines 105-107). This block alone enumerated part names and filtered them against the snapshot, which is exactly why renames fell through it. - **Wrong problem.** The block's own comment says the goal is to \"know number of rows\", implemented as a column read, while `getReadTaskColumns` already documents that a column is needed *only* for non-adaptive granularity. And `if (available_columns.empty()) available_columns = part_columns;` substituted an unusable value rather than answering whether that set can legitimately be empty — it can, and the answer differs by case. ## Solution Build the candidate set as the **pairing**, so both properties hold by construction: walk the metadata columns, map each to its name in the part through `AlterConversions` exactly as `injectRequiredColumnsRecursively` does, skip the ones a pending mutation drops, and keep those the part has files for. Sizes are looked up under the part's name; the metadata name is what gets injected. When the pairing is empty, decide on the merits instead of falling back, because the two ways of getting there are opposites: - **Everything the part holds is accounted for** by the structure or by a pending drop. Then each requested column is legitimately absent and the part reads as rows of defaults. Nothing needs to be injected under adaptive granularity; if the granularity is not adaptive, throw, since a column really must be read. - **The part holds a column that is neither in the structure nor being dropped** — a part attached after the schema moved on. Its rows exist on disk, so reporting defaults would hide data. Throw, naming the offending column. That second rule is not hypothetical: `04011_detach_rename_attach_column` (issue #79110) requires exactly this error for the renamed variant, and an earlier version of this patch that always injected nothing broke it. #79110's reporter asked whether such an attach should fail and the answer was yes. A drop is recorded under the name the command used, so when a column was also renamed the drop and the part-side name differ and both are consulted. A single `ALTER` can rename and then drop the same column, so this composition is reachable — two statements cannot produce it, because an `ALTER` issued while a `RENAME` is still pending blocks. ### Tests - `04870_vertical_merge_inject_column_after_rename.sh` — #103654. - `04871_inject_column_no_common_column_with_metadata.sh` — four cases: all columns dropped (#103116's repro); a stale re-attached part, which must **error**; renamed columns with `index_granularity_bytes = 0`; and a single-statement pending rename + drop. Both are shell tests so a `trap` always releases the server-global failpoint — their failure mode is a query throwing, which in a `.sql` test would leave mutations disabled for every later test on that server. Every clause was checked by reverting it alone: without the rename mapping the non-adaptive case fails; without the adaptive-granularity branch the dropped-columns case fails; without the second namespace in the drop check the rename+drop case fails. `03830_vertical_merge_inject_column_after_drop` (the test of #96675, which introduced this fallback) and `04011` both keep passing. ### Notes for reviewers - The candidate walk builds `getAllPhysical()` per part rather than taking a reference to the part's column list. The branch only runs when no requested column is physically present in the part, so it is not a hot path, but it is more work than before on a wide table read across many such parts. - Behaviour change: a part whose contents are fully accounted for is now read as rows of defaults where it previously threw. That is the point of #103116. - Two related defects found while verifying this are **not** fixed here and are independent of this function: - A part produced by a merge while a `RENAME COLUMN` is pending reads the renamed columns as their type default, because `StorageMergeTree::MutationsSnapshot::getOnFlyMutationCommandsForPart` decides applicability from the part's data version alone while the merge writes current metadata names. Values reappear once the mutation applies. Mirroring `ReplicatedMergeTreeQueue`'s metadata-version gate does *not* work — I tried it, and the source parts satisfy it too, so a plain `SELECT` regresses immediately. - `RENAME a TO b, DROP b, ADD b` in one statement makes the re-added `b` return the old `a` data (`min(b), max(b)` gave `100, 1099` instead of `7, 7`). The stale mapping is applied by `IMergeTreeReader`'s own name resolution, so it cannot be fixed in `injectRequiredColumns`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NO_SUCH_COLUMN_IN_TABLE` naming a column that is not in the table, when reading or merging a part whose columns had all been renamed or all dropped by a mutation that was not applied yet. The merge kept being retried, so the table stopped compacting.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114562",
          "createdAt": "2026-08-13T00:42:13Z",
          "updatedAt": "2026-08-13T10:46:30Z",
          "timestamp": "2026-08-13T10:46:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:aaf579e31a9721382abd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114067",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114067",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix(Analyzer): skip rerunFunctionResolve for 'exists' nodes created by rewrite_in_to_join",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114026 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed `Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION)` thrown when a `PREWHERE` clause contains an `IN (subquery)` predicate and `rewrite_in_to_join = 1` (or `make_distributed_plan = 1`, which force-enables it) is set. The same query spelled with `WHERE` worked correctly. --- ### Problem With `rewrite_in_to_join = 1`, the analyzer rewrites `x IN (subquery)` into an `exists(...)` `FunctionNode` resolved via `FunctionExists`, a special function that is not registered in `FunctionFactory`. When `PREWHERE` is resolved, `ReplaceColumnsVisitor` calls `rerunFunctionResolve` on every `FunctionNode` in the predicate. `rerunFunctionResolve` (`src/Analyzer/Utils.cpp`) already special-cases `grouping` (also resolved outside the factory), but not `exists`, so it called `FunctionFactory::instance().get(\"exists\", context)` and threw `UNKNOWN_FUNCTION`. ### Fix Two parts: 1. `PREWHERE` is evaluated by the reading step and cannot execute a correlated subquery — the planner rejects one with `ILLEGAL_PREWHERE`. So the `rewrite_in_to_join` rewrite is now skipped while resolving a `PREWHERE` expression, and the plain `IN` is kept there. `PREWHERE x IN (subquery)` then returns the same result as its `WHERE` spelling, which is what the issue asks for. Subqueries nested inside `PREWHERE` still rewrite their own `IN` predicates. 2. `exists` is added to the special-case early return in `rerunFunctionResolve`, next to `grouping`. This matters for an explicitly written `PREWHERE EXISTS (correlated subquery)`, which is genuinely unsupported: it is now reported honestly as `ILLEGAL_PREWHERE` instead of `Unknown function exists`. ### Test `tests/queries/0_stateless/04820_rewrite_in_to_join_prewhere_exists.sql` covers `PREWHERE ... IN (subquery)` against the `WHERE` control arm, `NOT IN`, tuple `IN`, an `IN` nested in a subquery inside `PREWHERE`, and the explicit `PREWHERE EXISTS (...)` case asserting `ILLEGAL_PREWHERE`. ### Reproduction ```sql CREATE TABLE t (k UInt64, s String) ENGINE = MergeTree ORDER BY k; INSERT INTO t SELECT number, toString(number % 2) FROM numbers(1000); -- Threw: Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION) SELECT count() FROM t PREWHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; -- Worked (control) SELECT count() FROM t WHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; ```",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114067",
          "createdAt": "2026-08-09T20:05:33Z",
          "updatedAt": "2026-08-13T10:42:18Z",
          "timestamp": "2026-08-13T10:42:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "close in a month if not active"
          ],
          "author": "RohithPariki",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ee1c2760bd78b7e801fd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114571",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114571",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Interpolate quantiles in the integer domain",
          "text": "Interpolate quantiles in the integer domain <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results from the interpolating quantile functions (`quantile`, `median`, `quantileExactWeightedInterpolated`, `quantileInterpolatedWeighted` and their `quantiles*` forms) over integer-backed types such as `DateTime64` and `Decimal`. The interpolated value was computed as a `Float64` and then narrowed, which loses ticks above magnitude 2^53 and is undefined near the top of the range: `quantile` over a `DateTime64(9)` column at 2262-04-11 returned a date in 1677. `quantileInterpolatedWeighted` also subtracted the endpoints in the column's own type, which overflows once they span more than it, so it could be wrong at any magnitude: over a full-range `Int8` column it returned `-128`. Related: https://github.com/ClickHouse/ClickHouse/pull/44067 ### Description Interpolating quantiles computed the result as a `Float64` and narrowed it to the native integer type. Three defects follow. `Int64::max` has no exact `Float64` representation (the nearest is 2^63, one past it), so the cast is undefined. On x86 it wrapped to `Int64::min`: ```sql SELECT quantile(x) FROM (SELECT fromUnixTimestamp64Nano(9223372036854775807, 'UTC') AS x FROM numbers(2)); -- 1677-09-21 00:12:43.145224192, was 2262-04-11 23:47:16.854775807 ``` On AArch64 it saturates, so only UBSan complained. Above magnitude 2^53, which a nanosecond timestamp passes by two orders of magnitude, the `Float64` spacing exceeds one, so the middle tick between two ordinary values was unrepresentable. Third, `quantileInterpolatedWeighted` subtracted the endpoints in the column's own type, so it was wrong far below 2^53 too: `-128` over a full-range `Int8` column. Each family loses the value inside its own callee, so a fix at the narrowing site cannot work. Each now keeps its own `Float64` expression while both endpoints are below magnitude 2^53 and, for the subtracting family, their difference fits, so results there are unchanged otherwise. Two equal endpoints, and a level equal to a sample's own percentile, return that endpoint directly rather than through the expression, so a result there can move by one, to the true value. Outside it the offset from the lower endpoint is exact in the unsigned domain and never exceeds the endpoint distance, so the result stays inside. Coincident endpoints on the `Float64` result path are also returned directly, so an infinite one no longer forms `inf * 0.0`. `NO_SANITIZE_UNDEFINED` is removed from `QuantileInterpolatedWeighted::interpolate`, whose suppression hid this. `QuantileExactWeighted` keeps its own, now confined to the endpoint selection whose position cast is defective near `UInt64::max` already.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114571",
          "createdAt": "2026-08-13T02:18:38Z",
          "updatedAt": "2026-08-13T10:36:19Z",
          "timestamp": "2026-08-13T10:36:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ebb67ddfc209c4c6f589",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114565",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114565",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Scale the lldb stacktrace budget by build flavor, keep timed-out dumps",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109455 --> Related: https://github.com/ClickHouse/ClickHouse/pull/109455 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `clickhouse-test` collects C-side backtraces with `lldb -o 'thread backtrace all'` under a hard 30s timeout. On a debug server the attach needs longer, the timeout fires, and the partial dump was discarded too, leaving no C stacks for the wedged process. CIDB, 90 days: 181 `suspiciously small stacktraces` rows over 118 PRs, 80 `timed out after 30 seconds` over 59 PRs, 78 carrying both. Two causes: 1. The 30s budget assumed the backtrace \"should finish in seconds even on a debug build\", so anything longer meant a wedged lldb. Measured on a 174-thread debug server in the CI stress image, a healthy attach takes 44-118s and returns a complete dump: the cost is symbolization. 2. `shell_get_output` honoured `keep_output_on_error` only for `CalledProcessError`; `TimeoutExpired` is not a subclass, so a timeout returned `\"\"`, discarding what lldb wrote. The per-PID budget now scales with the build flavor, reusing the `slow_build` predicate (debug, TSan/ASan/UBSan/MSan, coverage) that `MergeTreeSettingsRandomizer` uses for the same reason: 120s on a slow build, 30s unchanged on release (`Stress test (arm_debug)` is 51 of those 80 rows, so the driver is debug). The flavor comes from `args.build_flags`, not the module globals a spawned worker resets; the startup-failure caller runs before those flags exist, so there it is read from the binary via `clickhouse local`, as `is_asan_build` already does. A separate aggregate ceiling bounds the pid loop and reports skipped pids. The per-test timeout handler keeps 30s, since it runs inside its own fired alarm. A timed-out dump is now kept, decoded and marked truncated. Validated against a live debug server with real lldb: before, 30.8s and a 0-byte dump called suspiciously small; after, 63.6s and 834KB with 348 headers. Forcing the budget below the attach cost keeps 48.7KB where master keeps 0. Twelve new tests, each change reverted reddening one.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114565",
          "createdAt": "2026-08-13T01:06:44Z",
          "updatedAt": "2026-08-13T10:36:18Z",
          "timestamp": "2026-08-13T10:36:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:273d81163ef885a756f8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112950",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112950",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support the Vortex file format",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/87327 Related: https://github.com/ClickHouse/rust_vendor/pull/74 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added support for reading and writing the [Vortex](https://github.com/vortex-data/vortex) columnar file format (the `Vortex` input and output format). This closes [#87327](https://github.com/ClickHouse/ClickHouse/issues/87327). ### Documentation entry for user-facing changes The implementation uses the Rust `vortex` crate (v0.83.0) through a new C FFI crate `rust/workspace/vortex` (`_ch_rust_vortex`), following the same pattern as `prql` and `polyglot`. Data crosses the FFI boundary through the Arrow C Data Interface and is converted with the same `ArrowColumnToCHColumn`/`CHColumnToArrowColumn` code as the `Arrow` format. IO is delegated back to ClickHouse through callbacks: reads go through ClickHouse's own read buffers (range reads for seekable inputs, whole-file buffering otherwise), and the produced file is streamed into the output buffer. All work is driven by a single-threaded runtime on the calling thread — the library spawns no threads, and Rust panics are caught at the FFI boundary and turned into exceptions. Features: - Reading with projection pushdown: only the columns used by the query are read from the file. - Schema inference and `count()`-only queries answered from file metadata without reading data. - Writing with the library's default adaptive compression (BtrBlocks-style cascading encodings + zstd), including a valid empty file for empty results. - Graceful errors on malformed and truncated files (fuzzer-friendly: no aborts, Rust panics become exceptions). Limitations (documented in `docs/reference/formats/Vortex.mdx`): - `Map`, `Int128`/`UInt128`/`Int256`/`UInt256`, `IPv6`, and `Interval` columns cannot be written (no corresponding Vortex type). - `String` and `FixedString` are written as Vortex `Binary` (ClickHouse strings are arbitrary bytes, while Vortex requires `Utf8` to be valid UTF-8). - The format is disabled in MSan builds: the MSan-instrumented library (with origin tracking) is so large that linking `unit_tests_dbms` overflows the 2 GiB `R_X86_64_PC32` relocation range (same approach as `wasmtime` and `delta-kernel-rs`). - Reading and writing are single-threaded in this first version: the whole scan (I/O, decompression, and decoding) runs on one thread, so on ClickBench reads are significantly slower than `Parquet`, which ClickHouse decodes with multiple threads (see the [benchmark results](https://github.com/ClickHouse/ClickHouse/pull/112950#issuecomment-5274283519) and [the explanation](https://github.com/ClickHouse/ClickHouse/pull/112950#issuecomment-5274483899)). Filter pushdown (added in https://github.com/ClickHouse/ClickHouse/pull/114373, `input_format_vortex_filter_push_down`, on by default) reduces the amount of data decoded by selective queries, but does not parallelize the scan. Writes are also slower than `Parquet` (the adaptive compressor samples many encodings per column) — parallelism can be added later. The 127 new vendored Rust crates are added in https://github.com/ClickHouse/rust_vendor/pull/74 (the `contrib/rust_vendor` submodule is bumped to that branch).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112950",
          "createdAt": "2026-08-01T21:39:00Z",
          "updatedAt": "2026-08-13T10:35:47Z",
          "timestamp": "2026-08-13T10:35:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "pr-feature",
            "submodule changed",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fce620df8ef3e81605e9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112573",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix async bounded read buffer readbigat race",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix AsynchronousBoundedReadBuffer's readBigAt data race. Closes https://github.com/ClickHouse/ClickHouse/issues/109678. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1240` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112573",
          "createdAt": "2026-07-30T10:54:24Z",
          "updatedAt": "2026-08-13T10:34:38Z",
          "timestamp": "2026-08-13T10:34:38Z",
          "metrics": {
            "reactions": 1,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v26.4-must-backport"
          ],
          "author": "kssenii",
          "state": "closed",
          "assignees": [
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d87c83b9eed2dc39260b",
        "signalId": "github:ClickHouse/ClickHouse:issue:114500",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114500",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "[RFC] Add a `histogram(N)` column statistic for range predicate selectivity",
          "text": "### Company or project name ClickHouse ### Use case Selectivity estimation for range predicates (`<`, `<=`, `>`, `>=`, `BETWEEN`, and range decompositions from `PlainRanges`) on columns with non-uniform value distributions. Today these are estimated with `tdigest` (if declared), linear interpolation over `[min, max]` from `basic`/`minmax`, or a magic default factor (`default_cond_range_factor = 0.33`). The interpolation fallback assumes a uniform distribution and can be off by an order of magnitude on common data shapes. Example: a `ts` column spanning three years where 80% of rows are in the last three months — for `WHERE ts >= '2024-01-01' AND ts < '2024-02-01'` interpolation predicts ~1/36 of rows while the true answer is ~27%. This misleads PREWHERE ordering, join-order decisions, and `RelationProfile` row estimates. `tdigest` helps, but as an optimizer statistic it gives point estimates with no error bounds, is comparatively expensive to build inside background merges, and cannot compose with a most-common-values statistic (no way to subtract heavy-hitter mass from a region of the digest). ### Describe the solution you'd like An opt-in, parameterized bucket-frequency histogram statistic: ```sql CREATE TABLE t ( k UInt64, ts DateTime STATISTICS(basic, uniq_v2, histogram(128)) ) ENGINE = MergeTree ORDER BY k; ALTER TABLE t ADD STATISTICS ts TYPE histogram(128); ALTER TABLE t MATERIALIZE STATISTICS ts; ``` Key properties: - **Equi-width buckets on a dyadic grid**: bucket width is a power of two, boundaries anchored at zero. This is what makes the statistic fit ClickHouse's per-part architecture: it builds streaming, block-by-block, in bounded memory (a fixed counter array, no sort, no second pass), and grids built independently on different parts are hierarchically nested, so query-scope merging across thousands of selected parts is exact re-binning — merging loses resolution, never accuracy. - **Exact counts, deterministic bounds**: counts cover all rows of the part (not a sample), so the estimator gets hard `lower(v) <= true_count < upper(v)` bounds with a policy-driven interpolated point estimate in between. - **Framework reuse**: implemented as a new `StatisticsType` in the existing `IStatistics` / `ColumnStatistics` framework; row-preserving merges combine summaries exactly, row-changing merges (TTL, deletes, collapsing) rebuild from the output stream — matching the current rebuild-vs-merge split in `MergeTask`. - **Supported types (v1)**: integers, floats (with dedicated NaN/±Inf counters), `Decimal`, `Date`/`Date32`/`DateTime`/`DateTime64`, `Enum`, `IPv4`, plus `Nullable`/`LowCardinality` wrappers. `String` and 128-bit ID types deferred. - **Estimation use (v1)**: range/comparison selectivity in `estimateLess`/`estimateRange` first; equality/`IN` refinement via bucket density and composition with the proposed `mcv` statistic (subtracting the heavy-hitter head at estimation time) as follow-ups. - **Guardrails**: `N` required and capped by server-level settings; payload size bounded (~one `UInt64` per bucket); not included in `auto_statistics_types` initially. ### Describe alternatives you've considered - **Equi-height (equi-depth) histograms** (PostgreSQL, MySQL, Oracle, StarRocks): more accurate per bucket under skew, but require sorted data or precomputed quantiles to build and cannot be merged exactly across parts — a non-starter for per-part streaming statistics built inside inserts and merges. An approximate equi-depth _view_ can still be derived at query scope from merged fine-grained buckets. - **Keeping `tdigest` as the only range estimator**: retained, not replaced — but it carries no formal error bounds, and its centroids shift under merging; the histogram is expected to be the better range estimator, to be validated by benchmarks. - **A mergeable quantile sketch (KLL)**: provably mergeable, but duplicates `tdigest`'s rank-space role, gives probabilistic rather than deterministic bounds, and offers no stable bucket identity for future per-bucket metadata. - **The `histogram()` aggregate function's adaptive bins** (Ben-Haim/Tom-Tov): streaming and bounded, but data-dependent boundaries never align across parts, so merging is heuristic — fine for visualization, wrong for an optimizer statistic. ### Additional context Suggestion to implement histograms raised by @fkastrati.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114500",
          "createdAt": "2026-08-12T14:46:49Z",
          "updatedAt": "2026-08-13T10:25:15Z",
          "timestamp": "2026-08-13T10:25:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "feature"
          ],
          "author": "cv4g",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b544a5c9b24d33125e9c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112930",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112930",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve canceling queries with the `url` function",
          "text": "Resubmission of https://github.com/ClickHouse/ClickHouse/pull/104089 by Roman Vasin (@rvasin) into a branch of the main repository, with the review feedback from @Algunenano addressed. The original pull request is closed as superseded. Related: https://github.com/ClickHouse/ClickHouse/pull/104089 Related: https://github.com/ClickHouse/ClickHouse/pull/102801 Related: https://github.com/ClickHouse/ClickHouse/pull/104014 A query which reads over HTTP does not react to cancellation while `ReadWriteBufferFromHTTP` is retrying a request: `KILL QUERY`, `max_execution_time` or a disconnected client take effect only after all `http_max_tries` attempts and the backoffs between them are over, which is minutes with the default settings. ### What the change does {#what-the-change-does} - `doWithRetries` waits for the backoff on a cancellation flag instead of sleeping, so a cancellation interrupts the wait instead of being noticed after it has expired. The flag (`ReadWriteBufferFromHTTP::Cancellation`, a one-shot flag with a condition variable) is owned by `StorageURLSource` and set from its `cancel`. - After the backoff, and at the terminal exit of the loop - the attempt after which there is nothing left to retry, or one which failed with a non-retriable error - the loop asks the query status whether the query is still alive, via `CurrentThread::checkIfNotCancelled`, exactly as `Client::HeadObject` already does it for S3. A killed or timed out query is therefore reported with its own proper error (`QUERY_WAS_CANCELLED`, `TIMEOUT_EXCEEDED`, or the exception recorded for a disconnected client) instead of the network error that happens to be at hand. This part needs no plumbing, so it works for every user of the buffer, including those that pass no cancellation flag (the `web` disk, HTTP dictionaries, the data lake catalogs). - When the read is cancelled for another reason - the pipeline is being torn down because something else in the query has already failed, or the client has disconnected - the loop stops retrying and rethrows the error of its last attempt, which is the very exception the caller would have got once the attempts were exhausted. - `StorageURLSource::cancel` wakes the backoff for every cancellation reason - no one is left to wait for the remaining attempts. A cancellation after which the query must still succeed with what it has already read - a soft `max_execution_time` with `timeout_overflow_mode = 'break'` (`CancelledByTimeout`, which only ever comes from `PipelineExecutor::checkTimeLimitSoft`; a timeout with the `throw` overflow mode kills the query and arrives as `CancelledByUser`), or a consumer that has enough data (`PartialResult`) - is remembered, and `StorageURLSource::generate` then discards the error of the interrupted read and ends the stream, so the query returns its partial result instead of failing with the HTTP error. Only the interruption of the read is discarded: the sites which report an error *because* of the cancellation - the retry loop when it stops retrying, the failover loop when it stops probing the options - mark it (`Cancellation::markReadInterrupted`), and an error the cancellation has nothing to do with, for example a parse error of the data that had already been downloaded when it arrived, fails the query the same way it would with no cancellation at all. What is remembered is the *effective* kind of the cancellation: `ExecutingGraph::cancel` upgrades a `PartialResult` cancellation to the reason of a later hard one - for example, the first cancel of a client with `partial_result_on_first_cancel` followed by a `KILL QUERY` - and after the upgrade the read is not soft anymore, so nothing is discarded or synthesized - and nothing built under the soft state escapes either: an error of the interrupted read that was already in flight when the upgrade landed is suppressed rather than rethrown, because under a hard cancellation the query fails for a reason of its own - the error of the failed peer, the kill reported by the process list, or a disconnected client with no one left to report to - and the stale error must not mask it. - A cancellation stops `StorageURLSource::initialize` even where its helpers swallow the errors of the requests they make: `getFirstAvailableURIAndReadBuffer` rethrows the error of a cancelled read instead of probing the next failover option (and no longer probes them for a killed query at all, including a query killed right after its last option has failed, which is reported as cancelled instead of with the aggregate network error, and a cancellation which lands where no request is in flight for it to interrupt - between the options, or after the last one has failed - stops the choosing of the URI with the same two outcomes, a discarded cancellation error for a soft one and no buffer at all for a hard teardown, so the aggregate error can never stand in for the failure that really happened), and `ReadWriteBufferFromHTTP::tryGetFileSize` / `tryGetLastModificationTime` rethrow it instead of treating the interrupted `HEAD` request as a file without metadata (the rethrow is keyed off the mark the interrupted request leaves - `Cancellation::markReadInterrupted` - not off the cancellation flag at the moment the fallback runs, so an error that had already happened when the cancellation arrived stays swallowed, and a soft cancellation cannot turn a metadata failure the query would have survived into a failure of the query). `generate` re-checks `isCancelled` after `initialize`, so it never pulls a chunk no one needs. `initialize` itself ends the stream when the cancellation arrives after the URI has been chosen but before the metadata of the file - the modification time and the size - is requested, so a cancelled source does not start a fresh metadata `HEAD` request. The check is repeated between the two metadata probes: when the shared `HEAD` request has failed with a network error before the cancellation arrived, the probe of the modification time has nothing to remember, and the probe of the file size would otherwise send a fresh `HEAD` of its own. ### How the review feedback is addressed {#how-the-review-feedback-is-addressed} - *\"it is really easy to misuse, because you don't get any feedback on the thread that was cancelled, and exit normally. A simple look at `ReadWriteBufferFromHTTP::readBigAt` shows how trivial is now to hit asserts due to a pipeline cancellation.\"* - there is no silent `break` anymore: `doWithRetries` either does its work or throws, so a caller cannot mistake a cancellation for a success. The `chassert` in `readBigAt` stays untouched, `nextImpl` cannot report a cancellation as the end of the stream, and the row count cache cannot be filled from an interrupted read. Consequently `StorageURLSource` needs none of the `isCancelled()` checks the original pull request added after each buffer operation - which also answers *\"I don't understand the pattern of doing the HTTP request, and then checking if the pipeline is cancelled. Shouldn't it be the other way around?\"* and *\"Why not check at the start of the loop?\"*. - *\"part of the confusion comes from dealing with query cancellation and pipeline cancellation as the same thing, when they aren't, and how to report this properly both in logs and to the end user.\"* - the two are now separate. Query cancellation is taken from the query status and reported as a cancellation. Pipeline cancellation only stops the retrying and reports the HTTP error that really happened, so it cannot race the exception of a peer source with a wrong message and URI - which is what broke the two reverted attempts. Both cases are logged with the reason at the point where the loop gives up. - *\"We don't propagate any errors here, so I don't understand why we skip the first attempt.\"* - the `attempt > 1` condition is gone. The check now sits where the loop decides to wait and try again, which by construction can only be reached after an attempt has failed, so the error of a genuine first failure is never masked. - *\"we are sleeping without checks, that is, a cancel does not wake up the thread, which means we wait just to exit. We should use a condition variable and wait on it instead.\"* - done, see above. - *\"slow request won't be dealt with, we'll wait until the request either finishes or times out\"* - still not addressed, as agreed in that review. A request that neither answers nor times out is not interruptible; that needs closing the socket from the outside and is left for a separate change. ### Test {#test} `04615_kill_query_url_function` starts a server that always answers `503`, so that the query really is inside the retry loop - the test in the original pull request always answered `200` and never reached a second attempt. With 30 attempts and a backoff of 1 to 2 seconds, for both the `HEAD` and the `GET` request, the query would retry for minutes; the test kills it and checks that the client stops within seconds and reports an error. It also checks that the same query, when it is not killed, still reports the `503` of the server as before. Verified locally against a build of this branch: the killed query ends 3 ms after the `KILL` and is reported as `Code: 394. DB::Exception: Query was cancelled`. On an unpatched server the same query stays in `system.processes` with `is_cancelled = 1` and keeps retrying for more than five minutes. `04691_url_function_partial_result_on_break_timeout` covers the soft timeout: a glob over a file that the server serves completely and a URL that it always answers with `503`, with a backoff of 1 to 2 seconds over 30 attempts. A query with `max_execution_time = 1` and `timeout_overflow_mode = 'break'` succeeds with the rows of the first file in about a second, while the same query without the timeout still reports the `503` of the server. The partial result of the `break` mode consists of the rows already streamed to the consumer - a cancelled aggregation returns nothing even for regular tables - so the test asserts the streamed rows. `04759_url_function_cancel_during_initialize` covers a cancellation during initialization, whose helpers swallow the errors of the requests they make. Its server counts the requests to each path: after a soft `break` timeout interrupts the retries of the metadata `HEAD` request the data is never downloaded, and after it interrupts the retries of the first failover option of an `a|b` URL the second option is never probed - while an uninterrupted query still treats the failing metadata request as non-fatal and reads the data. Both assertions fail on the code before the fix. `04811_url_function_no_count_cache_poisoning_on_break_timeout` covers the row count cache (`use_cache_for_count_from_files`): a read of a slowly streaming URL interrupted by a soft `break` timeout leaves no entry in `system.schema_inference_cache`, while a complete read still caches the correct row count. `generate` records the count only when the read genuinely reached the end of the file - it checks the final status of its `PullingPipelineExecutor` and `isCancelled` on the source, as `StorageMemory` mutations do - so a cancelled read can never record the rows it happened to read as the row count of the file. `04812_url_function_no_fallback_probe_after_soft_cancel` covers a cancellation which lands *between* the failover options, where there is no request in flight for it to interrupt: an `empty|data` URL with `engine_url_skip_empty_files = 1`, where the empty file takes longer to be served than the soft `break` timeout of the query. Once the empty file is skipped, the loop checks the cancellation flag before constructing and probing the next option, so the query succeeds with no rows and the data URL receives no requests - the assertion fails on the code before the fix. An uninterrupted query still skips the empty file and reads the next option. This between-options check is reason-aware: it reports the interruption with a cancellation error - which `generate` then discards - only for the soft cancellations, after which the query must still succeed (`CancelledByTimeout` of the `break` overflow mode, or `PartialResult`). A pipeline torn down because something else in the query has already failed, or a disconnected client, must not be reported with a fabricated cancellation that could reach the user in place of the failure that really happened - and needs no synthetic error at all: `getFirstAvailableURIAndReadBuffer` returns no buffer and `initialize` ends the stream, since nobody is left who needs the data. `04824_url_function_kill_after_partial_result_cancel` covers a hard cancellation arriving after a soft one: the first cancel of a client with `partial_result_on_first_cancel` - after which the query must still succeed with its partial result - is followed by a `KILL QUERY`, both landing while the source is blocked in a request that a cancellation cannot interrupt (an empty file ahead of a failover option, whose response the test server withholds until the test releases it after the `KILL` has returned - the kill delivers the cancellation to the processors synchronously, so no timing can release the source early). The killed query must fail with the cancellation error instead of discarding it as if its result were partial, and must not probe the next failover option - the discard assertion fails on the code before the fix. The reason the executor delivers may understate a kill: its `checkTimeLimitSoft` poll observes a killed query as a soft timeout, and when that poll wins the race for the one-shot cancel reason of `ExecutingGraph`, the kill's own hard broadcast never reaches the source - so before discarding, the source additionally asks the process list, and a killed query fails with the proper cancellation error under either delivery order. `04825_url_function_no_next_option_after_disconnect` covers a hard teardown which does not kill the query, landing between the failover options: the client of a query blocked in the request for a held empty first option disconnects (`kill -9`), the test waits until the delivery of the resulting cancellation to the source is visible in the log (`StorageURLSource::cancel` leaves a debug trace of every delivered cancellation and its reason), and only then lets the server answer the held request. The source finds the empty file, and must end the stream instead of probing the next option, which would succeed - the assertion fails on the code before the fix. The tests which kill their query or assert on its log pin `parallel_replicas_for_cluster_engines` off: the rewrite of `url` to `urlCluster` would move the source into remote queries with their own query ids. `04829_url_function_no_metadata_requests_after_cancelled_head` covers the fallback for servers which do not support `HEAD`: `ReadWriteBufferFromHTTP::getFileInfo` treats a non-retriable 4xx response as \"the server cannot answer this\" and reports no metadata - and used to do so even when the request had been interrupted by a cancellation, so the initialization completed as if the file simply had no metadata instead of failing with the error of the interrupted read, which `generate` discards or fails with depending on the kind of the cancellation. The test answers the metadata `HEAD` with `400` only after a soft `break` timeout has been delivered to the source (the request-counting server withholds the response until the delivery is visible in the log): the query succeeds with its empty partial result, no request follows the cancelled `HEAD`, and the log records that the interrupted read was reported and its error discarded - the last assertion fails on the code before the fix. `04837_url_function_killed_query_after_last_attempt` covers the exit of the retry loop after the attempt which has nothing left to retry. Its reader is the schema inference of the `url` table function, which passes no cancellation flag, so the query status is the only thing that can tell the read that the query is gone. The server withholds its `503` response until the test has killed the query, and `http_max_tries` is 1, so the read is interrupted exactly where the retrying ends: the query must fail with `QUERY_WAS_CANCELLED`, while the same query which is not killed must still report the `503` of the server. On the code before the fix the killed query is reported with the `503` instead - which is the assertion the test makes. `04843_url_function_no_metadata_probe_after_soft_cancel` covers the window after the URI has been chosen and before the metadata of the file is requested: the server withholds the response of the first failover option until a soft `break` timeout has been delivered to the source (visible in the log), and then serves the file without a `Content-Length`, so the modification time and the size could only come from a `HEAD` request. The cancelled source must end the stream instead of probing the metadata no one is left to read: the query succeeds with its empty partial result and no `HEAD` request follows - the assertion fails on the code before the fix. `04844_url_function_parse_error_after_soft_cancel` covers an error which the cancellation has nothing to do with: the server streams a few good rows, withholds the rest until a soft `break` timeout has been delivered to the source, and only then sends a malformed row. The parse error must fail the query even though the soft cancellation - after which the query would otherwise succeed with its partial result - has already been latched: the file is malformed no matter when the query stopped wanting more of it. On the code before the fix the query succeeds, and the malformed row is discarded together with the interruption of the read. `04846_url_function_no_second_metadata_probe_after_cancel` covers a cancellation which lands between the two metadata probes of `initialize`: the server tears down the metadata `HEAD` request without a response - a network error which leaves the probe of the modification time with nothing to remember - and the source is held in the window between the probes with a failpoint (`storage_url_pause_between_metadata_probes`, added for the test: the window is a few instructions wide) until a soft cancellation has been delivered to it - the client is cancelled with SIGINT under `partial_result_on_first_cancel = 1`, so the delivery is under the test's control and cannot race the query the way a short `max_execution_time` did (which made the first version of the test flaky under the sanitizer builds) - and the delivery is visible in the log. The released source must end the stream instead of letting the probe of the file size send a second `HEAD` request: the query succeeds with its empty partial result and the server sees exactly one `HEAD` - the assertion fails on the code before the fix. `04869_url_function_stale_metadata_error_after_soft_cancel` covers the opposite side of the same fallbacks: a metadata `HEAD` request which fails on its own *before* any cancellation arrives must stay non-fatal even when a soft cancellation lands while its failure is being unwound. The server tears down the `HEAD` request without a response, the source is held inside the fallback with a failpoint (`http_read_buffer_pause_before_metadata_fallback`, added for the test: the window between the failed request and the fallback is a few instructions wide), the test delivers the soft cancellation (SIGINT with `partial_result_on_first_cancel = 1`, visible in the log) and only then releases the source. The fallback must swallow the error of the request the cancellation did not interrupt: the query succeeds with its empty partial result and the data is never requested - on the code before the fix the query fails with the stale network error of the `HEAD` request. `04871_url_function_hard_cancel_upgrade_after_interrupted_read` covers the upgrade of a soft cancellation to a hard one landing in the last window: after the error of the interrupted read has been thrown under the soft state, but before `generate` has handled it. One source retries an always-failing URL and its backoff is woken by a soft cancellation (SIGINT with `partial_result_on_first_cancel = 1`); the thrown error is held in the window with a failpoint (`storage_url_pause_before_handling_interrupted_read_error`, added for the test: the window is a few instructions wide). A second source, held by the test server until now, is then released into a parse error - a real failure the cancellation has nothing to do with - which cancels the pipeline hard and upgrades the paused source's reason to `Exception`; the test waits until both deliveries are visible in the log, so the upgrade deterministically lands inside the window, and only then releases the source. The stale error of the interrupted read must be suppressed and the query fails with the parse error of the peer - on the code before the fix it fails with the stale HTTP error of the read its own cancellation interrupted. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `KILL QUERY`, query timeouts and client disconnects now stop a query that reads over HTTP (for example over the `url` table function) while it is retrying a request, instead of waiting until all the retry attempts are exhausted.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112930",
          "createdAt": "2026-08-01T16:30:32Z",
          "updatedAt": "2026-08-13T10:22:51Z",
          "timestamp": "2026-08-13T10:22:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2e589b54a9037a7b8ec2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107125",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107125",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add query plan cache for complex queries (views, joins, subqueries)",
          "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add an experimental query plan cache supporting complex `SELECT` queries: views (including nested views), joins, `UNION`, `IN` subqueries and multiple tables over supported local table engines. When `allow_experimental_query_plan_cache = 1` and `enable_query_plan_cache = 1`, repeated identical queries skip query analysis and logical planning (the query is still parsed) while still executing against current data. `SYSTEM DROP QUERY PLAN CACHE` clears the cache; events `QueryPlanCacheHits`/`QueryPlanCacheMisses`/`QueryPlanCacheValidationMisses`/`QueryPlanCacheStaleMisses` and metrics `QueryPlanCacheBytes`/`QueryPlanCacheEntries` provide observability. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) --- ### Description An experimental **query plan cache** for `SELECT` queries that supports **complex queries** — views (including nested views), joins, `UNION`, `IN` subqueries and multiple tables over supported local table engines. It complements https://github.com/ClickHouse/ClickHouse/pull/99309, which covers single-table `MergeTree` queries with a universalize/materialize design; this PR takes a different approach to reach query shapes where planning cost is actually concentrated (deep view chains, many joins). Unlike the query result cache, a plan cache hit still executes the query: only query analysis and logical planning are skipped (the query is still parsed, since the cache key is built from the parsed AST), so hits are transactionally consistent and always see current data. #### Design On a miss, the plan is built in the planner's existing `build_logical_plan` mode (the machinery used by distributed plan shipping), with a new `SelectQueryOptions::cacheable_logical_plan` mode that makes the logical plan self-contained: - Leaf reads are storage-agnostic `ReadFromTable` placeholders for any eligible local engine. - **Views are expanded at plan time** (only `StorageView`; a materialized view stays a leaf and is read from its target table like an ordinary table), recursively in logical mode (propagated through `StorageView::readImpl` via `SelectQueryInfo::build_logical_plan`), so the cached plan embeds the analyzed view bodies instead of deferring the expensive expansion to execution. - Key-value direct-join lookup steps (`JoinStepLogicalLookup`), which bind live storages into the plan, are not used; a regular join is planned instead and the physical join algorithm is chosen at materialization. Distributed and parallel-replicas logical plans do not set this mode and are unaffected. - The plan is serialized with the standard query plan serialization and built with `compile_expressions = 0` (JIT-fused function nodes cannot be serialized; the JIT still applies when pipelines are built from the plan). Both hits and misses execute through `QueryPlan::resolveStorages`, which re-binds every leaf to a fresh storage snapshot. #### Cache key and invalidation The key is the normalized AST hash, a hash of plan-affecting settings, the current database, the user and the sorted role set. Each entry stores a **dependency fingerprint over every referenced storage**, discovered from the plan's `ReadFromTable` leaves plus an AST closure through view definitions (which also catches tables referenced only from scalar subqueries): UUID, metadata version (or a schema content hash including the view body), and the row policy hash. Every hit revalidates all dependencies and re-checks `SELECT` access, so `DROP`/`CREATE`, `ALTER`, view redefinition (including **nested** views), row policy changes and permission revocation are handled transparently. #### Eligibility - Non-deterministic functions anywhere — including inside expanded view bodies — make a query uncacheable (`arrayJoin` is exempt: it is multi-valued but pure). - Scalar subqueries are evaluated during analysis and baked into the plan as constants, so they are gated behind the opt-in setting `query_plan_cache_allow_scalar_subqueries`. - Table functions, temporary tables, system tables (except `system.one`), remote/`Merge` storages and views whose SQL security is not `INVOKER` (`DEFINER` and `NONE`) are rejected. - Plans containing steps without serialization support (e.g. window functions) execute normally and are simply not stored. #### Performance, proven on DOOMbench [DOOMbench](https://github.com/cedardb/doombench) is a DOOM-like raycasting engine implemented entirely in SQL; its `screen` query renders a frame through a deep chain of views over joins, unions, array joins and scalar subqueries. Ported to ClickHouse SQL, the workload is **query-analysis-bound**: profiling showed `IQueryTreeNode::getTreeHash` and analyzer resolution dominating, and `SELECT ... LIMIT 0` was as slow as the full query. With the plan cache (local server, single connection): | DOOMbench frame render | median | speedup | | --- | --- | --- | | cache off | 551-577 ms | — | | cache on (hit) | 227-246 ms | **2.4x** | | cache on, `max_threads = 4` | 148-176 ms | **3.3x** | On hits the analyzer cost disappears from profiles (`getTreeHash`: 25.8k samples → 174); the remaining time is genuine execution. Correctness validated on the same workload: frames are byte-identical between cache off/miss/hit, game-state mutations are visible on every cached render (HTAP freshness), nested view redefinition is picked up immediately, and `rand()` inside a view body refuses to cache. #### Notes - This PR includes the `ActionsDAG` input-constant serialization fix https://github.com/ClickHouse/ClickHouse/pull/107124 as its first commit; it will be rebased once that PR is merged. - The miss path executes through the same logical-plan machinery as hits, so hit and miss behavior is identical by construction; storage-specific planning shortcuts (e.g. trivial `count()`) are not applied to cache-enabled queries. Related: https://github.com/ClickHouse/ClickHouse/pull/99309 Related: https://github.com/ClickHouse/ClickHouse/pull/107124",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107125",
          "createdAt": "2026-06-11T01:46:36Z",
          "updatedAt": "2026-08-13T10:14:20Z",
          "timestamp": "2026-08-13T10:14:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 28
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:896e10898235c9ddef29",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114614",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114614",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix join NDV propagation for distributed aggregation",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Preserve NDV statistics across optimized inner joins so distributed planning can choose partial aggregation for low-cardinality GROUP BY queries instead of shuffling all joined rows. ### Description An optimized `JoinStepLogical` has two children, so `estimateReadRowsCount` returned before reading its saved result-column statistics. Move the optimized-join handling before the generic single-child guard so row estimates and column statistics propagate to downstream planning. The regression test verifies that a low-NDV GROUP BY after an inner join selects partial aggregation. It pins the join-order limit and distributed aggregation mode to isolate the planner behavior from injected settings and exchange execution details. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114614",
          "createdAt": "2026-08-13T09:48:39Z",
          "updatedAt": "2026-08-13T10:10:05Z",
          "timestamp": "2026-08-13T10:10:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-performance"
          ],
          "author": "XuJia0210",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9ede9d74cc919f8410fa",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109472",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109472",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Log tolerated connection failures to remote MySQL/PostgreSQL databases as warnings instead of errors",
          "text": "Fix CI upgrade test failure: Example failing report — Upgrade check (amd_release) on the unrelated PR https://github.com/ClickHouse/ClickHouse/pull/108255: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=108255&sha=28ca539bd4bf66de37e388238de45d840bc30c52&name_0=PR&name_1=Upgrade%20check%20%28amd_release%29 (`Error message in clickhouse-server.log (see upgrade_error_messages.txt)` with the `getTablesIterator` line above; the report shows green now because the check passed on a later rerun — which is exactly the flaky-victim pattern). **Problem.** A `DatabasePostgreSQL` or `DatabaseMySQL` pointed at an unreachable server writes the connection failure to the server log at `<Error>` level on code paths that explicitly tolerate and retry the failure. A temporarily unavailable remote server is a normal operational state on these paths, yet it floods the log with error-level messages. Concretely: - every `system.tables` / `system.columns` scan that touches the database logs an error (`DatabasePostgreSQL::getTablesIterator`: `Code: 614. DB::Exception: ... Connection to ... failed`); - the background cleaner logs one more error every rescheduling cycle; - `ATTACH DATABASE ... ENGINE = MySQL(...)` logs three error lines for a single tolerated probe failure. This also trips the Upgrade check (which fails on any unexpected error message in `clickhouse-server.log`) on unrelated PRs: the stateless test `04210_show_remote_databases_in_system_tables` (added in #104416) leaves `PostgreSQL`/`MySQL` databases pointed at the unroutable `192.0.2.1` during the check's restart window, and since #109082 made remote databases visible to `system.tables` scans by default, the check failed on 137 distinct PRs in 14 days, and 0 times on master. **Root cause.** Seven log sites report a caught-and-tolerated (or caught-and-propagated) connection failure at error level: the catches in `DatabasePostgreSQL::getTablesIterator` (kept non-throwing so `system.tables` scans do not fail) and `DatabasePostgreSQL::removeOutdatedTables` (the cleaner reschedules and continues) use `tryLogCurrentException` with the default `LogsLevel::error`; the same for the `DatabaseMySQL` constructor catch on `ATTACH` and `DatabaseMySQL::getCreateTableQueryImpl` with `throw_on_error = false`; and the connection-pool layers (`mysqlxx::Pool::allocConnection`, `mysqlxx::PoolWithFailover::get`, `postgres::PoolWithFailover::get`) log at error level before propagating the exception to the caller — who is the one deciding how severe the failure actually is. **Fix.** Downgrade those seven sites to warning. Exception propagation is unchanged everywhere: a non-tolerated failure (e.g. `CREATE DATABASE` with an unreachable host) still reaches the client as an error through the normal query-error path. The new test `04506_remote_database_unreachable_no_error_log` attaches `PostgreSQL` and `MySQL` databases pointing at an unreachable host, exercises the tolerated paths (attach probe, `system.tables` scan, background cleaner), and asserts through `system.text_log` that each path produced a warning (proving the failure path fired) and no error-level lines at all for the corresponding queries. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109472",
          "createdAt": "2026-07-06T09:15:05Z",
          "updatedAt": "2026-08-13T10:09:29Z",
          "timestamp": "2026-08-13T10:09:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-ci"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:720e0337ae486932e89a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114484",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114484",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Document that PREWHERE filters one join input before the JOIN",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required for a documentation change. ### Description `prewhere.mdx` did not mention `JOIN` at all. This is the documentation KochetovNicolai asked for when he closed issue 89097 as not-a-bug: \"We need to document this.\" One `<Note>`, next to the existing note that documents the same class of fact for `FINAL`. It states the rule and gives the equivalent explicit spelling as a filtered subquery. The `SELECT` clause list is annotated to agree with it. Measured on the example as it appears on the page: `PREWHERE b.y > 50` returns `[1,2,3,4]`, the filtered-subquery spelling the same, the `WHERE` spelling `[1,4]`. The `WHERE` result is unchanged whether `query_plan_filter_push_down` is on or off, which is why the note credits `WHERE` to the join result rather than to a fixed position, per PedroTadim's review. The note claims nothing about individual join kinds or strictnesses, since `FULL JOIN` and `ASOF INNER JOIN` both differ too, and it attributes the fill value to [join_use_nulls](https://clickhouse.com/docs/reference/settings/session-settings/join#join_use_nulls) rather than naming one. Related: https://github.com/ClickHouse/ClickHouse/issues/89097 Related: https://github.com/ClickHouse/ClickHouse/issues/114206",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114484",
          "createdAt": "2026-08-12T13:01:26Z",
          "updatedAt": "2026-08-13T10:08:46Z",
          "timestamp": "2026-08-13T10:08:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:eaebc466735c13296026",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113998",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113998",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not merge patch parts across a pending mutation version",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/98898 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a permanently stalled mutation on `ReplicatedMergeTree` tables with lightweight updates enabled. A merge of patch parts could produce a patch whose source data versions span the version of a mutation that is still queued, which no mutation can then apply, so the `MUTATE_PART` entry failed and retried forever (`Found patch part ... that intersects mutation with version ...`). Such merges are now postponed until the mutation completes. ### Description Related: https://github.com/ClickHouse/ClickHouse/issues/98898 A patch part records the data-version range of what it patches, and applies whole or not at all. Merging patch parts unions those ranges, so merging patches for versions 1 and 6 yields one spanning 1..6. A mutation cutting at version 2 cannot apply it, `patchHasHigherDataVersion` throws, and the `MUTATE_PART` entry retries forever. Debug and sanitizer builds abort, hence the stress failures. Root cause: the merge predicate already refuses parts whose current mutation versions differ, but derives them from `mutations_by_partition`, where `getCurrentMutationVersion` returns 0 for an absent partition. The finished-mutation cleaner removes the `/mutations` znodes once every replica's `mutation_pointer` passed them, while the `MUTATE_PART` entries survive. Every patch then maps to 0 and every pair looks mergeable. The logs show `There are no mutations for partition ID all` just before the abort. The fix takes the pending versions from the replication queue instead, which is durable: a queued `MUTATE_PART` entry's `new_part_name` encodes its target version. A patch `MERGE_PARTS` whose sources would span one is postponed, so it proceeds once the mutation completes. No assertion is relaxed, and the cost is one walk of a queue the neighbouring guards already walk. The repro is deterministic: pristine master aborts with the exact message, the fix passes, and deleting only the new guard call brings the abort back. The test asserts the mutation completes, the queue drains and the data is fully patched, so no deadlock or dropped patch passes. Two limits. Like the guards beside it, this is an execution-side per-replica check: it stops a replica merging across a mutation queued there, not every route by which a spanning patch could reach one. And it does not repair a patch already spanning a version on disk, so such a table stays stuck; that needs per-row `_data_version` filtering when applying patches.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113998",
          "createdAt": "2026-08-09T01:04:11Z",
          "updatedAt": "2026-08-13T10:02:32Z",
          "timestamp": "2026-08-13T10:02:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "CurtizJ"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e021dc81cbc5a28fc0fc",
        "signalId": "github:ClickHouse/ClickHouse:issue:113184",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113184",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Materialized CTE over Distributed: 49 LOGICAL_ERROR \"Reading from materialized CTE before its materialization completed - DelayedPortsProcessor gate is missing\" survives the #108924 fix",
          "text": "## Describe the problem With `enable_materialized_cte = 1`, a twice-referenced materialized CTE reading a `Distributed` table, filtered by `IN (SELECT ... FROM <another materialized CTE>)`, fails with exception 49 (`LOGICAL_ERROR`): ``` Code: 49. DB::Exception: Reading from materialized CTE 'ct' before its materialization completed - DelayedPortsProcessor gate is missing in the query plan: While executing Memory. (LOGICAL_ERROR) ``` This is the fail-fast invariant introduced by the materialized-CTE scheduling fix itself (PR https://github.com/ClickHouse/ClickHouse/pull/108924, merged) — verified on a master build that CONTAINS that merge: the PR's own `04036_materialized_cte_distributed_race` shapes pass, but this neighbouring shape on the same `Distributed` path still trips the retained gate check. The identical shapes over a plain `MergeTree` table pass (the `04227_materialized_cte_reused_with_in_subquery` cases). Release builds throw; debug builds abort on the logical-error check. ## How to reproduce Version: `26.8.1.653` (public master build, includes the #108924 merge). Single server; the single-shard `test_shard_localhost` cluster is enough: ```sql SET enable_materialized_cte = 1; CREATE TABLE t (c Int32) ENGINE = MergeTree ORDER BY c; INSERT INTO t VALUES (1), (2), (3); CREATE TABLE dist_t AS t ENGINE = Distributed(test_shard_localhost, currentDatabase(), t); WITH ct AS MATERIALIZED (SELECT 1 AS c), rs AS MATERIALIZED (SELECT * FROM dist_t WHERE c IN (SELECT c FROM ct)) SELECT count() FROM rs AS a, rs AS b; ``` Deterministic: 20/20 exceptions with `max_threads = 1`, 20/20 with `max_threads = 8`, and the same with `serialize_query_plan` 0 and 1 (80/80 total). Expected: `1` (the count the same query returns over the plain `MergeTree` table). Required ingredients (each verified by removing it): - `enable_materialized_cte = 1`. - The twice-referenced materialized CTE (`rs`) reads a `Distributed` table. Only `rs` needs it — `ct` can be a constant `SELECT 1 AS c`, and moving the `Distributed` read into `ct` while `rs` reads the local table makes the query pass. - `rs` filters by `IN (SELECT ... FROM ct)` where `ct` is another *materialized* CTE. With a plain (non-materialized) CTE, a bare `IN (SELECT 1)`, or a literal `IN (1, 2, 3)` the query passes — this looks like the `forceMaterializeCTE` path for IN-set subqueries, which the fix intentionally kept \"collected and gated exactly as before\". - `rs` referenced twice in the join tree. `rs AS a, rs AS b`, `ANY LEFT JOIN`, and `SELECT * FROM rs UNION ALL SELECT * FROM rs` all fail alike. Incidental: the join kind, `max_threads`, `serialize_query_plan`, the number of shards. ## Additional context A *single* reference to `rs` over `Distributed` fails differently — `Code: 60 UNKNOWN_TABLE: Unknown table expression identifier '_materialized_cte_ct_...'` on the shard-local query — the temp-table shipping defect family (#112642 is the parallel-replicas variant), so the double reference selects this gate-missing channel rather than merely amplifying it. Exception site: `src/Processors/QueryPlan/ReadFromMemoryStorageStep.cpp` (`MemorySource::generate`), reached from the materialization pipeline. Distinct from the other open \"gate is missing\" fixes: #111194 (`WITH TOTALS`, per #110176) and #113043 (extremes + `UNION`) — neither involves the `Distributed` + IN-set path. Found by query fuzzing on master (`tests/optimizer_tester`, branch `optimizer-tester-framework`). Related: https://github.com/ClickHouse/ClickHouse/pull/108924 Related: https://github.com/ClickHouse/ClickHouse/issues/112642 Related: https://github.com/ClickHouse/ClickHouse/issues/111194 Related: https://github.com/ClickHouse/ClickHouse/issues/113043",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113184",
          "createdAt": "2026-08-03T20:24:05Z",
          "updatedAt": "2026-08-13T10:01:33Z",
          "timestamp": "2026-08-13T10:01:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "bug",
            "comp-query-optimizer",
            "comp-distributed",
            "comp-query-execution",
            "common table expressions"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:acda43d37cbbece37101",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110653",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110653",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add STREAM BOUNDED modifier",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM BOUNDED modifier, which read only the first snapshot of a streaming query, then finish instead of subscribing for updates. cc @alesapin @Michicosun <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1314` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110653",
          "createdAt": "2026-07-16T08:35:24Z",
          "updatedAt": "2026-08-13T10:01:10Z",
          "timestamp": "2026-08-13T10:01:10Z",
          "metrics": {
            "reactions": 1,
            "comments": 6
          },
          "labels": [
            "pr-improvement",
            "pr-synced-to-cloud"
          ],
          "author": "SmitaRKulkarni",
          "state": "closed",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2c8a7298ea2df77404c3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114544",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114544",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Text index: fix read mode of LIKE/ILIKE with an index built on a container",
          "text": "Evaluating `LIKE`/`ILIKE` by scanning the text index dictionary always used an exact direct read, which removes the original condition from the query plan. That is only valid when a matching dictionary token proves the predicate, which is not the case when the index is built on a container while the predicate reads a single element out of it: with an index on `mapValues(m)`, `m['a'] LIKE '%foobar%'` also returned rows whose match is under another key. However, it may produce correct result due to the data, but it's not guaranteed. To ensure for producing correct results, we need to use `HINT` mode when a text index built on a container (Map or JSON). It is better to be safe than sorry. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix text index evaluation of `LIKE/ILIKE` operator built on Map or JSON containers. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1313` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114544",
          "createdAt": "2026-08-12T21:22:24Z",
          "updatedAt": "2026-08-13T10:01:06Z",
          "timestamp": "2026-08-13T10:01:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud"
          ],
          "author": "ahmadov",
          "state": "closed",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e8c8970a3dee79930c7b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114569",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114569",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: Truncated `Array` elements stream: `CANNOT_READ_ALL_DATA` branch has no test",
          "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #109212](https://github.com/ClickHouse/ClickHouse/pull/109212). That PR: (1) Moves the `Array` offsets monotonicity check in `SerializationArray::deserializeOffsetsBinaryBulk` from `settings.native_format` to `!settings.position_independent_encoding`, and rescopes the scan to values appended by the current call (starting one element early); adds a second consistency … **1. Truncated `Array` elements stream: `CANNOT_READ_ALL_DATA` branch has no test** `src/DataTypes/Serializations/SerializationArray.cpp:537`, `src/DataTypes/Serializations/SerializationArray.cpp:545` **Risk:** `SerializationArray::deserializeBinaryBulkWithMultipleStreams` restructured the offsets/elements consistency check into a nested `if`: the `!nested_column->empty()` arm (`SerializationArray.cpp:536-538`) throws `CANNOT_READ_ALL_DATA`, the new `!settings.position_independent_encoding` arm throws … **Unique vs PR tests:** The PR's gtest covers decreasing offsets and the empty-elements case by calling the serialization directly; nothing covers a short-but-non-empty elements column, and nothing exercises either new/restructured branch through `NativeReader` from SQL. cc @groeneai (author of #109212), @alexey-milovidov (merged/approved #109212) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1315` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114569",
          "createdAt": "2026-08-13T02:09:17Z",
          "updatedAt": "2026-08-13T10:01:03Z",
          "timestamp": "2026-08-13T10:01:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "clickgapai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:905ce94f066f8d1fcf12",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114600",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114600",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add is_nullable to system.columns, use it in information_schema",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/44930 Related: https://github.com/ClickHouse/ClickHouse/pull/48560 `information_schema.columns` decided nullability by matching the printed type name (`type LIKE 'Nullable(%)'`), so a `LowCardinality(Nullable(T))` column — which accepts `NULL` — was reported as `is_nullable = 0`, and MySQL-compatible clients inferred it as `NOT NULL`. On #48560 @alexey-milovidov said about that approach: > I don't like either the modification in this PR or the original code. Because it is doing ad-hoc, > rough text parsing, like what you typically do with Perl. The right way to do it is to introduce > `is_nullable` column in the `system.columns` and reuse it here. So this does that: `system.columns` gains `is_nullable UInt8`, computed from the data type via the existing `isNullableOrLowCardinalityNullable()`, and the view selects it. The text parsing is removed rather than extended, matching how `numeric_precision`, `numeric_scale`, `datetime_precision` and `character_octet_length` are already computed in C++ and passed through. `IS_NULLABLE` is an alias of `is_nullable`, so it is fixed too. A column reports `is_nullable = 1` exactly when it accepts `NULL`: | type | accepts NULL | is_nullable | |---|---|---| | `Nullable(T)` | yes | 1 | | `LowCardinality(Nullable(T))` | yes | 1 | | `LowCardinality(T)` | no | 0 | | `Array(Nullable(T))`, `Tuple(Nullable(T))`, `Map(K, Nullable(V))` | no | 0 | Composite types stay `0` because the column itself cannot hold `NULL`, only its elements can. New test `04669_information_schema_is_nullable_lowcardinality` covers both `system.columns` and the view. Regenerated `02117_show_create_table_system`, `02206_information_schema_show_database` and `01602_temporary_table_in_system_tables` for the added column. `SimpleAggregateFunction(any, Nullable(T))`, `Variant(...)` and `Dynamic` also accept `NULL` while reporting `0`. Left for a follow-up — now a change in one place. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `information_schema.columns` reporting `is_nullable = 0` for `LowCardinality(Nullable(T))` columns. Add an `is_nullable` column to `system.columns`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114600",
          "createdAt": "2026-08-13T07:31:59Z",
          "updatedAt": "2026-08-13T09:58:54Z",
          "timestamp": "2026-08-13T09:58:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "zainulabidin302",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:38e75f532f009e73197a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112250",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112250",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Materialize column statistics on INSERT by default",
          "text": "This PR contains https://github.com/ClickHouse/ClickHouse/pull/109454 minus `materialize_statistics_on_insert_max_table_size` (which I'm happy to introduce in a second step). Made a separate PR to speed up the integration of the feature (the new behavior has high demand and the original PR is stuck since three weeks). If the original PR gets merged first, we can close this one. <!-- Related: https://github.com/ClickHouse/ClickHouse/pull/109454 --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Column statistics are now materialized on `INSERT` by default. This improves estimations for the cost-based join optimizer.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112250",
          "createdAt": "2026-07-28T09:46:11Z",
          "updatedAt": "2026-08-13T09:57:06Z",
          "timestamp": "2026-08-13T09:57:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "rschu1ze",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fcb1b96c15dc1c57d48d",
        "signalId": "github:ClickHouse/ClickHouse:issue:114616",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114616",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Row policy over a file-backed table breaks `DEFAULT` columns missing from the data file (`UNKNOWN_IDENTIFIER`; silently wrong results on 26.7)",
          "text": "🕵 Found while working on https://github.com/ClickHouse/ClickHouse/pull/114262 (lazy materialization for local Parquet files); the behavior is identical with that PR's optimization on or off, and reproduces on current master without it. **Describe what's wrong** When a table over a data file (`File(Parquet)`, and by code inspection the same applies to the object-storage path) has a `DEFAULT` column that is missing from the file, and a row policy references either that column or interacts with reading it, the query fails with `UNKNOWN_IDENTIFIER` — the source prunes the `DEFAULT` expression's input columns before `AddingDefaultsTransform` can compute the column. On the official 26.7 build, the first case below returned **silently wrong results** (0 rows instead of 900): the policy was evaluated against type defaults instead of real row values. Current master turned that into a fail-safe exception, which is safer, but the queries are legitimate and should work. **How to reproduce** ```bash clickhouse local ``` ```sql INSERT INTO FUNCTION file('rp.parquet', Parquet) SELECT number AS k, number % 10 AS a, concat('val_', toString(number)) AS s FROM numbers(1000) SETTINGS engine_file_truncate_on_insert = 1; -- Case 1: the row policy references a DEFAULT column missing from the file. CREATE TABLE t_rp_j (k UInt64, a UInt64, s String, j JSON DEFAULT toJSONString(map('user', map('name', concat('u', toString(a)))))) ENGINE = File(Parquet, 'rp.parquet'); CREATE ROW POLICY pol_j ON t_rp_j USING j.user.name != 'u0' TO ALL; SELECT k, s FROM t_rp_j ORDER BY k LIMIT 3; -- Code: 47. DB::Exception: Unknown expression or function identifier `a` in scope -- _CAST(toJSONString(map('user', map('name', concat('u', toString(a))))), 'JSON') AS j ... While executing File. -- Official 26.7 instead returns 0 rows (silently wrong: expected 900 rows admitted by the policy). -- Case 2: the policy is on a plain file column, and the query selects a DEFAULT -- column depending on another file column, together with a PREWHERE. CREATE TABLE t_rp_d (k UInt64, a UInt64, s String, d UInt64 DEFAULT a * 2) ENGINE = File(Parquet, 'rp.parquet'); CREATE ROW POLICY pol_d ON t_rp_d USING a != 0 TO ALL; SELECT k, d FROM t_rp_d PREWHERE s != 'val_2' ORDER BY k LIMIT 3; -- Code: 47. DB::Exception: Unknown expression or function identifier `a` in scope _CAST(a * 2, 'UInt64') AS d ... -- (the same statement without PREWHERE works and returns 1→2, 2→4, 3→6) ``` **Expected behavior** Both queries return the same result as with the row policy applied above the read (e.g. as `MergeTree` does): case 1 → 900 rows admitted by the policy computed from real `a` values; case 2 → `d` computed from `a` after prewhere filtering. **Where it comes from** `updateFormatPrewhereInfo` / `SourceStepWithFilter::applyPrewhereActions` shrink the format header to the filter DAG's outputs, dropping input columns that a `DEFAULT` expression of another (missing) column needs; `AddingDefaultsTransform` inside the source then cannot compute the default. For case 1 the policy's own input is a missing defaulted column, so its dependency `a` is never part of the read set to begin with. **Additional context** Related: https://github.com/ClickHouse/ClickHouse/pull/114262 Related: https://github.com/ClickHouse/ClickHouse/pull/110970",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114616",
          "createdAt": "2026-08-13T09:56:28Z",
          "updatedAt": "2026-08-13T09:56:28Z",
          "timestamp": "2026-08-13T09:56:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "potential bug"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3e86c383b0c0de652b3b",
        "signalId": "github:ClickHouse/ClickHouse:issue:112018",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112018",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "LOGICAL_ERROR on re-run INSERT … SELECT — sparse Tuple subcolumn breaks the insert-deduplication retry pipeline",
          "text": "When a Tuple column has an element stored with sparse serialization, re-running an INSERT … SELECT that is caught by block deduplication throws LOGICAL_ERROR (code 49) instead of being a silent no-op. Repro (tested in 26.4.1): ``` CREATE TABLE sparse_tuple_src ( id UInt64, body Tuple(key UInt64, flag Bool) ) ENGINE = MergeTree ORDER BY id; CREATE TABLE sparse_tuple_dst ( id UInt64, body Tuple(key UInt64, flag Bool) ) ENGINE = MergeTree ORDER BY id SETTINGS non_replicated_deduplication_window = 100; -- body.flag is the default (false) for all but the first 1000 of 2M rows, i.e. ~99.95% -- defaults. That is above ratio_of_defaults_for_sparse_serialization (default 0.9375), -- so the tuple ELEMENT gets sparse serialization. INSERT INTO sparse_tuple_src SELECT number, (number, number < 1000) FROM numbers(2000000); -- First insert: succeeds. INSERT INTO sparse_tuple_dst SELECT * FROM sparse_tuple_src SETTINGS insert_deduplication_token = 'repro'; -- Second, identical insert: blocks are duplicates, so the deduplication -- retry pipeline is built -> LOGICAL_ERROR. INSERT INTO sparse_tuple_dst SELECT * FROM sparse_tuple_src SETTINGS insert_deduplication_token = 'repro'; Code: 49. DB::Exception: Block structure mismatch in function connect between SourceFromSingleChunk and ConvertingTransform stream: different columns: body Tuple(key UInt64, flag Bool) Tuple(size = 0, UInt64(size = 0), Sparse(size = 0, UInt8(size = 1), UInt64(size = 0))) body Tuple(key UInt64, flag Bool) Tuple(size = 0, UInt64(size = 0), UInt8(size = 0)). (LOGICAL_ERROR) ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #111191 - Backported to: `26.7.3.19`, `26.6.3.7`, `26.5.7.21` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112018",
          "createdAt": "2026-07-27T06:17:23Z",
          "updatedAt": "2026-08-13T09:54:35Z",
          "timestamp": "2026-08-13T09:54:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "bug"
          ],
          "author": "romanbukarev-clh",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:dab6bf4f8fd32c6bc573",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108522",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108522",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add system.s3(azure)_queue_metadata",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add system.s3(azure)_queue_metadata. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1316` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108522",
          "createdAt": "2026-06-25T16:42:44Z",
          "updatedAt": "2026-08-13T09:54:17Z",
          "timestamp": "2026-08-13T09:54:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-improvement",
            "pr-synced-to-cloud"
          ],
          "author": "kssenii",
          "state": "closed",
          "assignees": [
            "bharatnc"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0df3c8febc5cb88ca9d7",
        "signalId": "github:ClickHouse/ClickHouse:issue:114499",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114499",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "[RFC] Add most-common-values (`mcv`) column statistics for better equality/IN selectivity on skewed columns",
          "text": "### Company or project name ClickHouse ### Use case Cardinality estimation for equality and `IN` predicates on skewed columns is currently weak. When only `uniq`/`uniq_v2` statistics are available, `ColumnStatistics::estimateEqual()` falls back to a uniform assumption of `rows / ndv`, which can be off by orders of magnitude for skewed data: ```sql WHERE country = 'US' -- 'US' may be 45% of rows, not 1/ndv WHERE status = 500 WHERE event_type IN ('purchase', 'signup') ``` Example: a scan of 100M rows with 250 distinct countries yields a uniform estimate of 400K rows for `country = 'US'`. If `US` is actually 45% of the column, the true answer is ~45M rows — a 100x error that propagates into join order, PREWHERE, and other plan choices. Status-, enum-, and country-like `LowCardinality` columns are the primary target, where skew is common and the statistic is cheap to build. ### Describe the solution you'd like Add an opt-in, parameterized per-column statistic `mcv(N)` (most common values), reusing the existing `IStatistics` / `ColumnStatistics` framework: ```sql CREATE TABLE events ( user_id UInt64, country LowCardinality(String) STATISTICS(basic, uniq_v2, mcv(64)), status UInt8 STATISTICS(basic, uniq_v2, mcv(16)) ) ENGINE = MergeTree ORDER BY user_id; ALTER TABLE events ADD STATISTICS country TYPE mcv(64); ALTER TABLE events MATERIALIZE STATISTICS country; ``` Key properties: - **Bounded, mergeable heavy-hitter summary**, not an unbounded exact map and not a plain local top-N list. Canonical representation: a Misra-Gries-family summary (per-value lower-bound counts plus one global error term), with an exact mode for low-cardinality data. Local exact top-N lists are insufficient because they don't merge reliably: a value can be a global heavy hitter without appearing in any single part's top-N. - **Query-scope merging**: queries scan many parts, so serialized summaries must be mergeable at planning time (in `ConditionSelectivityEstimatorBuilder`), with well-defined error accumulation. Physical merges rebuild the statistic from the merge output stream when the row set changes. - **Estimation**: for tracked values, use the MCV count (with explicit lower/upper bounds in sketch mode); for untracked values, subtract the MCV head and spread the residual over the remaining NDV — the classic MCV treatment in relational optimizers (cf. PostgreSQL `most_common_vals`). Add a batch estimate for `IN` lists so shared tail mass isn't double-counted. - **Opt-in and guarded**: require the `N` parameter, keep `mcv` out of default `auto_statistics_types`, and bound build memory and serialized payload size via server-level settings. - **Scope of v1**: equality/`IN` selectivity only, per-shard planning. MCV is complementary to `uniq_v2` (head vs. tail) and to a separately planned histogram statistic (head vs. body/tail distribution). A prerequisite is fixing parameter handling for statistics: today `STATISTICS(tdigest(200))` parses but silently drops the argument, and factory/deserialization paths recreate descriptions from the bare type enum. `mcv(N)` needs parameters to be validated, persisted, compared in `structureEquals()`, and survivable through serialization. ### Describe alternatives you've considered - **Exact local top-N per part**: simple, but merging only top-N lists across parts misses true global heavy hitters; slack candidates and an explicit error term are needed for sound query-scope merging. - **`countmin` (existing, requires `USE_DATASKETCHES`)**: estimates the frequency of a known value but cannot enumerate heavy hitters, so it can't drive skew detection or tail-subtraction estimates; it remains useful as a fallback for untracked point values. - **Unbounded exact `value -> count` map**: unbounded memory during background merges; unacceptable. - **Histograms**: serve range predicates and the distribution body/tail; planned separately and complementary — heavy hitters distort buckets, so an MCV head makes a future histogram better, not redundant. - **Naming (`frequent_items`, `topk`, `heavy_hitters`)**: `mcv` matches the established optimizer-statistics term (PostgreSQL `most_common_vals`) and is algorithm-agnostic. ### Additional context - The Misra-Gries family gives deterministic retention (with `k` counters, any value with frequency `> rows/(k+1)` is retained) and provably mergeable summaries (Agarwal et al., \"Mergeable Summaries\", TODS 2013). - A production-proven implementation (Apache DataSketches frequent-items sketch) is already vendored in `contrib/datasketches-cpp`; the in-tree `SpaceSaving.h` (used by `topK`) is a related candidate engine. The persisted format should be engine-agnostic, and `mcv` should not hard-depend on `USE_DATASKETCHES`. - Follow-ups explicitly out of scope for v1: join cardinality from MCV intersections, skew-aware join/aggregation execution, cross-shard statistics transport, multi-column statistics",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114499",
          "createdAt": "2026-08-12T14:46:45Z",
          "updatedAt": "2026-08-13T09:53:59Z",
          "timestamp": "2026-08-13T09:53:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "feature",
            "st-need-info"
          ],
          "author": "cv4g",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8961ed65c30609ae4306",
        "signalId": "github:ClickHouse/ClickHouse:issue:114612",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114612",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Heap-use-after-free: Parquet v3 prefetcher reads and writes through a `ReadBuffer` freed by `IInputFormat::onFinish`",
          "text": "🕵️ ## Describe what's wrong `ParquetV3BlockInputFormat` does not override `resetReadBuffer()`, so `IInputFormat::onFinish()` frees the format's owned `ReadBuffer` while the Parquet `Prefetcher`'s IO tasks are still reading through it on the prefetch thread pool. The tasks then dereference freed memory. ASan on a plain `clickhouse local` run of the reproducer below (master, 26.8.1.1310): ``` ==ERROR: AddressSanitizer: heap-use-after-free on address 0x7112bd001800 WRITE of size 1048576 at 0x7112bd001800 thread T9 (ParquetPrefetch) #3 in DB::ReadBuffer::next() src/IO/ReadBuffer.cpp:113:15 #5 in DB::ReadBuffer::read(char*, unsigned long) src/IO/ReadBuffer.h:169:37 #6 in DB::Parquet::Prefetcher::readSync(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:113:29 #7 in DB::Parquet::Prefetcher::runTask(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:540:13 #8 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:435 #15 in DB::ThreadPoolCallbackRunnerFast::threadFunction() src/Common/threadPoolCallbackRunner.cpp:225:13 ``` It is a **1 MiB write into freed heap**, so this is memory corruption, not only a bad read. The same defect was hit independently by the `La Casa Del Dolor (arm_asan_ubsan)` CI job, where the report is a read and carries the full free/allocation stacks: ``` ERROR: AddressSanitizer: heap-use-after-free on address 0xfcdd2aec39c0 READ of size 8 at 0xfcdd2aec39c0 thread T423 (ThreadPool) #0 in DB::Parquet::Prefetcher::readSync(...) Prefetcher.cpp:110:21 <- reader->setReadUntilEnd() #1 in DB::Parquet::Prefetcher::runTask(...) Prefetcher.cpp:523:13 #2 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() Prefetcher.cpp:424:17 #9 in DB::ThreadPoolCallbackRunnerFast::threadFunction() threadPoolCallbackRunner.cpp:225:13 0xfcdd2aec39c0 is located 0 bytes inside of 272-byte region freed by thread T468 (ThreadPool) here: #1 in std::default_delete<DB::ReadBuffer>::operator()(DB::ReadBuffer*) #7 in std::vector<std::unique_ptr<DB::ReadBuffer>>::~vector() #8 in DB::ISource::work() src/Processors/ISource.cpp:140:13 ... #18 in DB::IPolygonDictionary::loadData() src/Dictionaries/PolygonDictionary.cpp:337:8 previously allocated by thread T468 (ThreadPool) here: #2 in DB::FileDictionarySource::loadAll() src/Dictionaries/FileDictionarySource.cpp:60:19 ``` `ISource.cpp:140` is the `onFinish()` call inside the `catch (...)` handler, and the freed 272-byte region is the `ReadBufferFromFile` that `FileDictionarySource::loadAll` handed to the format via `addBuffer`. ## How to reproduce Reproduces on master (26.8.1.1310) with the official `build_amd_asan_ubsan` binary, **8 out of 10 runs, with no ThreadFuzzer and no unusual settings**. Under ThreadFuzzer it hit on the first run. Run Fiddle: https://fiddle.clickhouse.com/06f26a34-d286-40a1-85f5-7d3e952d47b8 On a release build this prints only the load error and exits 53: `Code: 53. CAST AS Array can only be performed between same-dimensional Array ... While executing ParquetV3BlockInputFormat. (TYPE_MISMATCH)` — the 1 MiB write into freed memory is silent. On an ASan build the process dies with the report above instead (5/5 runs of exactly the commands above; 8/10 in an earlier variant, so it is a race, but a very wide one). The polygon dictionary is only a convenient way to make the pipeline throw mid-read while owning its `ReadBuffer` — the defect is in the format, not in the dictionary. ## Root cause `Prefetcher` protects only *its own* lifetime. `scheduleTask` captures the shutdown handle and each task takes `std::shared_lock(*_shutdown, std::try_to_lock)`; `~Prefetcher` calls `shutdown->shutdown()` (`ShutdownHelper`, `src/Common/threadPoolCallbackRunner.h:507`), which blocks until every in-flight task has released. As its own comment says, \"`this` is safe to access as long as `shutdown_lock` is held\". But `Prefetcher::reader` points at a `ReadBuffer` owned by `IInputFormat::owned_buffers`, an unrelated lifetime that the handshake does not cover: - `IInputFormat::onFinish()` -> `resetReadBuffer()` -> `resetOwnedBuffers()` -> `owned_buffers.clear()` - `ParquetV3BlockInputFormat` overrides `resetParser()` and `onCancel()`, but **not `resetReadBuffer()`**. `resetParser()` happens to be safe only because it destroys the reader *before* delegating to the base. `resetReadBuffer()` has no such ordering, so the buffer dies while the `ReadManager` -> `Reader` -> `Prefetcher` chain is still alive with tasks running. Three routes reach it: `onFinish()` from the `catch (...)` in `ISource::work()` (the one above), `onFinish()` on the normal completion path (`ISource.cpp:133`) with speculative prefetches still outstanding, and `IInputFormat::setReadBuffer` when a format is reused across files. ## Suggested fix Mirror what `resetParser()` already does, so `~Prefetcher` drains the IO tasks before the base frees the buffer: ```cpp void ParquetV3BlockInputFormat::resetReadBuffer() { /// ~Prefetcher waits for in-flight IO tasks, which read through the buffer that /// IInputFormat::resetReadBuffer() is about to free. { std::lock_guard lock(reader_mutex); reader.reset(); } IInputFormat::resetReadBuffer(); } ``` ## Additional context Distinct from #109678 / #112573. That was a data race on `AsynchronousBoundedReadBuffer::prefetch_future` consumed by concurrent `readBigAt` (the `RandomRead` branch, `Prefetcher.cpp:102`). This one is a lifetime bug in the `SeekAndRead` branch, which already holds `read_mutex` — locking cannot help once the object is freed — and the freed buffer here is a plain `ReadBufferFromFile`, so `AsynchronousBoundedReadBuffer` is not involved at all. The binary used above already contains #112573 (merged into 26.8.1.1240). Found by this Dolor run: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=94148&sha=a2ef409981fc77a168eaf8cc68d7ddf67c9fadec&name_0=PR&name_1=La+Casa+Del+Dolor+%28arm_asan_ubsan%29",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114612",
          "createdAt": "2026-08-13T09:21:28Z",
          "updatedAt": "2026-08-13T09:53:04Z",
          "timestamp": "2026-08-13T09:53:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "comp-parquet-reader-v3"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c2dbf31f31ddaff520c8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109433",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109433",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reserve memory for merges up front",
          "text": "Add proactive memory reservation for background merges, enhancing the existing `merges_mutations_memory_usage_soft_limit` mechanism. `background_memory_tracker` measures memory that background tasks have *already* allocated, so `canEnqueueBackgroundTask` is purely reactive: when many merges are scheduled at the same time — for example right after a mutation produces many parts — their IO buffers are not allocated yet, all of them pass the gate, and only then grow their memory usage and collide. Every merge now estimates the memory of its input and output IO buffers (`CompactionStatistics::estimateNeededMemoryForMerge`) as the number of input column streams (over all source parts) times the read IO buffer size, plus the number of output column streams (of the result part) times the write IO buffer size. Object storage (S3) write buffers are large and double-buffered, so they are accounted separately. Since IO buffers only ever hold data that flows through them, the estimate is capped by the data volume of the merge (beyond the eagerly allocated per-stream compressor and file buffers, which honor adaptive write buffers): without this cap, a merge of tiny parts in a many-column table on object storage would reserve gigabytes it can never touch, and concurrent merges would saturate the soft limit and starve each other. The merge reserves this amount at start (`MergeMemoryReservation`) and releases it when it finishes. A merge that will certainly run the vertical algorithm is priced by the streams that are alive at once (the merging columns, one gathering column at a time, and up to `max_merge_delayed_streams_for_parallel_write` delayed streams) rather than all output streams concurrently, and a stream whose data volume is not derivable from the source parts (a rebuilt projection, a `DEFAULT`-filled column of variable size) is priced at the buffers its writer allocates before any data flows through it (its compressor block and file buffer) plus a projected-volume bound, rather than at any multipart upload size - a multipart writer starts from the buffer size its caller passes and grows it only with the data written into it, leaving further growth to the reactive tracker — over-reservation starves all merges, while under-reservation merely degrades to the reactive behavior of the current code for that stream. `canEnqueueBackgroundTask` now consults both the actual usage and the reservation: - a non-replicated background merge reserves at selection and is not scheduled when the reservation would exceed the limit (it retries later); once `CurrentlyMergingPartsTagger` chooses the actual destination disk, the reservation is corrected if the pre-selection guess made from the source parts was wrong, and a merge that waited in the background queue re-prices its reservation at task start against the destination disk's live multipart upload settings (a config reload could have raised them), as the replicated path does. Whether a background merge runs as a `ReplacingMergeTree` cleanup merge is decided at selection and carried to the scheduler, so a row-reducing cleanup merge is priced as one; - a replicated merge, already committed to run locally, reserves unconditionally at execution so the reservation still throttles selection of further merges; - a user-initiated merge (`OPTIMIZE`) also reserves unconditionally, so it is throttled by the gate for other merges but can never be silently skipped by it. A single merge whose estimate exceeds the whole limit is always allowed to proceed alone, so progress is never blocked. Adds the `MergesMutationsMemoryReservation` metric for observability, a `gtest` for the reservation accounting, and stateless tests: a smoke test and a regression test that `OPTIMIZE TABLE ... FINAL` merges everything down to a single part under a pathologically small soft limit instead of silently doing nothing. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Background merges now reserve the memory of their input/output IO buffers up front against `merges_mutations_memory_usage_soft_limit`, so that scheduling many merges at once (for example right after a mutation) can no longer oversubscribe memory as they all start. A user-initiated `OPTIMIZE` is never silently skipped by this mechanism.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109433",
          "createdAt": "2026-07-05T07:51:55Z",
          "updatedAt": "2026-08-13T09:51:04Z",
          "timestamp": "2026-08-13T09:51:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 69
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:018d9fe5446f5b3c0196",
        "signalId": "github:ClickHouse/ClickHouse:issue:114611",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114611",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Logical error: Invalid number of columns in chunk pushed to OutputPort. Expected A, found B (STID: 2270-3258)",
          "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Invalid number of columns in chunk pushed to OutputPort. Expected A, found B (STID: 2270-3258) CI report: [AST fuzzer (amd_debug)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=108522&sha=da345caa0c75914b7749446668b09bb6ca9355e5&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBJbnZhbGlkIG51bWJlciBvZiBjb2x1bW5zIGluIGNodW5rIHB1c2hlZCB0byBPdXRwdXRQb3J0LiBFeHBlY3RlZCBBLCBmb3VuZCBCIChTVElEOiAyMjcwLTMyNTgpJwogICAgLS0gQU5EIGNoZWNrX25hbWUgPSAnQVNUIGZ1enplciAoYW1kX2RlYnVnKScKICAgIEFORCB0ZXN0X3N0YXR1cyBJTiAoJ0ZBSUwnLCAnRVJST1InKQogICAgQU5EIChwdWxsX3JlcXVlc3RfbnVtYmVyID0gMCBPUiBiYXNlX3JlZiBJTiAoJ21hc3RlcicpKQogICAgQU5EIHRlc3RfY29udGV4dF9yYXcgTElLRSAnJUV4Y2VwdGlvbjolJwpHUk9VUCBCWSBkYXkKT1JERVIgQlkgZGF5IERFU0MK) Test output: ``` Error: Logical error: 'Invalid number of columns in chunk pushed to OutputPort. Expected 0, found 1 Header: Chunk: String(size = 1) '. --- Failed query: SELECT '-92233\\02036854775808', rank() OVER (RANGE BETWEEN UNBOUNDED PRECEDING AND UNBOUNDED FOLLOWING) FROM remote('127.0.0.{2..3}:9000', view(SELECT DISTINCT '1' AS c0 LIMIT 318)) AS v0 QUALIFY rank() OVER (ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW) > 0 LIMIT 511 --- Stack trace: pthread_kill @ 0x00000000000969bd gsignal @ 0x0000000000042476 __ieee754_lgamma_r @ 0x00000000000287f3 src/Common/Exception.cpp:66:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x00000000143ef46e src/Common/Exception.cpp:115:13: DB::Exception::handleErrorCode(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x00000000143f04ac src/Common/Exception.cpp:171:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x00000000143f08ec src/Common/Exception.h:201:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d64c0d6 src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d64bb9e src/Common/Exception.h:219:77: DB::Exception::Exception<unsigned long, unsigned long, String, String>(int, FormatStringHelperImpl<std::type_identity<unsigned long>::type, std::type_identity<unsigned long>::type, std::type_identity<String>::type, std::type_identity<String>::type>, unsigned long&&, unsigned long&&, String&&, String&&) @ 0x000000001a484b4e src/Processors/Port.h:431: DB::OutputPort::pushData(DB::Port::State::Data) src/Processors/ISource.cpp:49:19: DB::ISource::prepare() @ 0x000000001ea44754 src/Processors/Sources/RemoteSource.cpp:102:30: DB::RemoteSource::prepare() @ 0x000000001eed2ef9 src/Processors/Executors/ExecutingGraph.cpp:384:59: DB::ExecutingGraph::updateNode(DB::ExecutingGraph::Node*, std::queue<DB::ExecutingGraph::Node*, boost::container::devector<DB::ExecutingGraph::Node*, AllocatorWithMemoryTracking<DB::ExecutingGraph::Node*>, void>>&, std::queue<DB::ExecutingGraph::Node*, boost::container::devector<DB::ExecutingGraph::Node*, AllocatorWithMemoryTracking<DB::ExecutingGraph::Node*>, void>>&) @ 0x000000001ea5eb4d src/Processors/Executors/PipelineExecutor.cpp:407:38: DB::PipelineExecutor::executeStepImpl(unsigned long, DB::WorkloadResources&&, std::atomic<bool>*) @ 0x000000001ea5888c src/Processors/Executors/PipelineExecutor.cpp:355:5: DB::PipelineExecutor::executeSingleThread(unsigned long, DB::WorkloadResources&&) @ 0x000000001ea58eac src/Processors/Executors/PipelineExecutor.cpp:692: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>::type std::__invoke[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:17: ? @ 0x000000001ea5bde7 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000014603329 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1293: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x000000001460c392 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x00000000146004de contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x000000001460974e start_thread @ 0x0000000000094a83 __clone3 @ 0x0000000000126890 ```",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114611",
          "createdAt": "2026-08-13T09:20:40Z",
          "updatedAt": "2026-08-13T09:48:04Z",
          "timestamp": "2026-08-13T09:48:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "testing",
            "fuzz"
          ],
          "author": "kssenii",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8dee6aa0cb4a7495c483",
        "signalId": "github:ClickHouse/ClickHouse:issue:114603",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114603",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Logical error: Unexpected token for lazy mode: A. Multi-block postings must be compressed (STID: 4250-5377)",
          "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Unexpected token for lazy mode: A. Multi-block postings must be compressed (STID: 4250-5377) CI report: [AST fuzzer (amd_debug)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=3b9bc037faeea25fb3dc2841ec3f87fbf3f2d8f8&name_0=MasterCI&name_1=AST%20fuzzer%20%28amd_debug%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBVbmV4cGVjdGVkIHRva2VuIGZvciBsYXp5IG1vZGU6IEEuIE11bHRpLWJsb2NrIHBvc3RpbmdzIG11c3QgYmUgY29tcHJlc3NlZCAoU1RJRDogNDI1MC01Mzc3KScKICAgIC0tIEFORCBjaGVja19uYW1lID0gJ0FTVCBmdXp6ZXIgKGFtZF9kZWJ1ZyknCiAgICBBTkQgdGVzdF9zdGF0dXMgSU4gKCdGQUlMJywgJ0VSUk9SJykKICAgIEFORCAocHVsbF9yZXF1ZXN0X251bWJlciA9IDAgT1IgYmFzZV9yZWYgSU4gKCcnLCAnbWFzdGVyJykpCiAgICBBTkQgdGVzdF9jb250ZXh0X3JhdyBMSUtFICclRXhjZXB0aW9uOiUnCkdST1VQIEJZIGRheQpPUkRFUiBCWSBkYXkgREVTQwo=) Test output: ``` Error: Logical error: 'Unexpected token for lazy mode: zrare. Multi-block postings must be compressed'. --- Failed query: SELECT DISTINCT count() FROM cluster('test_unavailable_shard', currentDatabase(), 'tab_postings_order__fuzz_12') PREWHERE hasAllTokens(s, ['filler', 'zrare']) WHERE hasAllTokens(s, ['zra\\0e', 'filler']) --- Reproduce commands (auto-generated; may require manual adjustment): SELECT DISTINCT count() FROM cluster('test_unavailable_shard', currentDatabase(), 'tab_postings_order__fuzz_12') PREWHERE hasAllTokens(s, ['filler', 'zrare']) WHERE hasAllTokens(s, ['zra\\0e', 'filler']); DROP TABLE IF EXISTS cluster('test_unavailable_shard', currentDatabase(), 'tab_postings_order__fuzz_12'); --- Stack trace: pthread_kill @ 0x00000000000969bd gsignal @ 0x0000000000042476 __lgamma_r_finite@GLIBC_2.15 @ 0x00000000000287f3 src/Common/Exception.cpp:66:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x00000000143ff02e src/Common/Exception.cpp:115:13: DB::Exception::handleErrorCode(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x000000001440006c src/Common/Exception.cpp:171:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x00000000144004ac src/Common/Exception.h:201:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d6533d6 src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d652e9e src/Common/Exception.h:219:77: DB::Exception::Exception<std::basic_string_view<char, std::char_traits<char>>&>(int, FormatStringHelperImpl<std::type_identity<std::basic_string_view<char, std::char_traits<char>>&>::type>, std::basic_string_view<char, std::char_traits<char>>&) @ 0x000000000df7df83 src/Storages/MergeTree/MergeTreeReaderTextIndex.cpp:370:15: DB::MergeTreeReaderTextIndex::makeLazyCursor(std::basic_string_view<char, std::char_traits<char>>, DB::TokenPostingsInfo const&) @ 0x000000001e281f96 src/Storages/MergeTree/MergeTreeReaderTextIndex.cpp:793:30: DB::MergeTreeReaderTextIndex::fillColumnLazy(DB::IColumn&, unsigned long, unsigned long, unsigned long, roaring::Roaring&) @ 0x000000001e285955 src/Storages/MergeTree/MergeTreeReaderTextIndex.cpp:524:17: DB::MergeTreeReaderTextIndex::readRows(unsigned long, bool, unsigned long, unsigned long, std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) @ 0x000000001e28325a src/Storages/MergeTree/MergeTreeRangeReader.cpp:187: DB::MergeTreeRangeReader::DelayedStream::readRows(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&, unsigned long) src/Storages/MergeTree/MergeTreeRangeReader.cpp:259:47: DB::MergeTreeRangeReader::DelayedStream::finalize(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) @ 0x000000001e25f0d7 src/Storages/MergeTree/MergeTreeRangeReader.cpp:375: DB::MergeTreeRangeReader::Stream::finalize(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>>&) src/Storages/MergeTree/MergeTreeRangeReader.cpp:1201:31: DB::MergeTreeRangeReader::startReadingChain(unsigned long, DB::MarkRanges&) @ 0x000000001e268190 src/Storages/MergeTree/MergeTreeReadersChain.cpp:296:36: DB::MergeTreeReadersChain::read(unsigned long, DB::MarkRanges&, std::vector<DB::MarkRanges, std::allocator<DB::MarkRanges>>&, std::function<void (std::vector<DB::ColumnWithTypeAndName, AllocatorWithMemoryTracking<DB::ColumnWithTypeAndName>> const&, std::unordered_set<String, std::hash<String>, std::equal_to<String>, std::allocator<String>> const&, unsigned long, std::optional<bool>&)> const&) @ 0x000000001e2bb9ba src/Storages/MergeTree/MergeTreeReadTask.cpp:442:38: DB::MergeTreeReadTask::read() @ 0x000000001e2b82de src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:270:31: DB::MergeTreeSelectProcessor::readCurrentTask(DB::MergeTreeReadTask&, DB::IMergeTreeSelectAlgorithm&) const @ 0x000000001e2c6e27 src/Storages/MergeTree/MergeTreeSelectProcessor.cpp:477:23: DB::MergeTreeSelectProcessor::read() @ 0x000000001e2c906d src/Storages/MergeTree/MergeTreeSource.cpp:233:41: DB::MergeTreeSource::tryGenerate() @ 0x000000001f16ae23 src/Processors/ISource.cpp:119:26: DB::ISource::work() @ 0x000000001ea558c1 src/Processors/Executors/ExecutionThreadContext.cpp:57: DB::executeJob(DB::ExecutingGraph::Node*, DB::ReadProgressCallback*) src/Processors/Executors/ExecutionThreadContext.cpp:127:28: DB::ExecutionThreadContext::executeTask() @ 0x000000001ea7715e src/Processors/Executors/PipelineExecutor.cpp:387:26: DB::PipelineExecutor::executeStepImpl(unsigned long, DB::WorkloadResources&&, std::atomic<bool>*) @ 0x000000001ea694e8 src/Processors/Executors/PipelineExecutor.cpp:355:5: DB::PipelineExecutor::executeSingleThread(unsigned long, DB::WorkloadResources&&) @ 0x000000001ea69bac src/Processors/Executors/PipelineExecutor.cpp:692: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>::type std::__invoke[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&>(DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:17: ? @ 0x000000001ea6cae7 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000014613869 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1293: operator() contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x000000001461c8d2 contrib/llvm-project/libcxx/include/__functional/function.h:502: ? contrib/llvm-project/libcxx/include/__functional/function.h:754: ? src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x0000000014610a1e contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x0000000014619c8e start_thread @ 0x0000000000094a83 __GI___clone3 @ 0x0000000000126890 ```",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114603",
          "createdAt": "2026-08-13T08:37:23Z",
          "updatedAt": "2026-08-13T09:48:01Z",
          "timestamp": "2026-08-13T09:48:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "testing",
            "fuzz"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [
            "CurtizJ"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fb9195d88d865bda5f32",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111985",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111985",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Measure the compressed size of aggregate states in `estimateSizeOfCompressedState`",
          "text": "`Aggregator::estimateSizeOfCompressedState` estimates how many bytes the aggregate states would take when a replica sends them to the initiator, which is how automatic parallel replicas decides whether distributing a query pays off. It builds a `CompressedWriteBuffer` over a `NullWriteBuffer`, but serializes the sampled states into the `NullWriteBuffer` directly and then reads that buffer's counter: ```cpp NullWriteBuffer wb; CompressedWriteBuffer wbuf(wb); // never written to ... aggregate_functions[j]->serialize(place + offsets_of_aggregate_states[j], wb); ... wbuf.finalize(); res += it ? static_cast<size_t>(table.size() * wb.count() / ((it + period - 1) / period)) : 0; ``` So nothing is ever compressed and, despite the function's name and its own `We only interested in the size of compressed state` comment, it returns the plain serialized size. `recordAggregationStateSizes` then stored that single number as all three of `bytes`, `sample_bytes` and `compressed_bytes`, which pins the compression ratio of the aggregation-state statistics to 1. `recordAggregationStateColumnSizes`, documented as `Mirrors the logic of Aggregator::estimateSizeOfCompressedState` for in-order aggregation, feeds the same `AggregationState` counters via `estimateCompressedColumnSize` and does measure the compressed size - so two producers of one statistic disagreed on what it means. This change serializes the sample through the compressing buffer and returns the sample's uncompressed and compressed sizes alongside the extrapolated total, letting the existing compression-ratio machinery in `RuntimeDataflowStatisticsCacheUpdater` apply the states' real ratio. Taking the ratio from the sample rather than compressing the extrapolated size also keeps the per-compressed-block framing overhead out of the extrapolation, where multiplying it by `table.size() / num_samples` would have inflated it. **Effect on the estimate.** Only queries whose output is dominated by compressible aggregate states move. That is exactly the shape that showed up as a CI failure of `03634_autopr_output_bytes_estimation` on master: `query_28` (`MIN(Referer)` states over a two-level hash table with many groups) estimated 59335657 bytes against a recorded 23722663, a ratio of 2.4996 that left the test's 2.5x bound no margin. Reading the estimates out of the failing job's server log, the ten other queries in that test match their recorded values within 1.00..1.41x, because their output is dominated by aggregation keys or by output columns, both of which were already measured compressed. So the recorded values expect the compressed size and this defect is why that one entry looked wrong. **Two small-sample corrections found while validating this in CI.** The compressed format writes a checksum and a block header in front of every block, so a sample of a few bytes comes out of `CompressedWriteBuffer` larger than it went in and the derived ratio drops below one, *inflating* the estimate. With an early conversion to a two-level hash table (`group_by_two_level_threshold=1`, which the test randomization sets) every one of the 256 buckets holds a handful of states and every per-bucket sample is dominated by that framing, which is what made `04034_autopr_dataflow_cache_reuse_between_different_queries` fail on the first run of this pull request: the two cache-reusing queries lost parallel replicas. When the states are really sent, the framing is amortized over `min_compress_block_size` of data, so such a sample is now reported as incompressible rather than as expanding, keeping the ratio at 1 - exactly what the caller assumed before the ratio was measured at all. The new test `04653_autopr_state_size_estimate_small_buckets` pins `group_by_two_level_threshold` to 1 and fails without that. Reviewing the same statistic end to end also turned up the mirror-image defect on the in-order-aggregation side, reported in review: `recordAggregationStateColumnSizes` totalled `IColumn::byteSize`, which for a `ColumnAggregateFunction` whose states live in a foreign arena (as `AggregatingInOrderTransform` hands them over) counts one pointer per row, so arena-backed states - `min(String)`, `groupArray`, `uniqExact` - were under-counted. That column is now sized from its serialized states too, sampling at most as many of them as the hash-table producer samples per bucket. **Recorded values.** The values in `03634_autopr_output_bytes_estimation` are empirical and need `test.hits`, so they are re-measured from this PR's own CI run rather than guessed - the test reports every query that lands outside the 2.5x band, and I will update whatever it reports. `query_1` and `query_20` are single `count()` states whose estimate becomes the ~29-byte compressed-block framing instead of a 2-3 byte varint; the test's `NOT (res.2 < 100 AND res.3 < 100)` guard already covers those. Verified so far: both changed translation units pass a full `-fsyntax-only` typecheck against master. The estimate itself is validated by CI, since reproducing it needs the stateful dataset. Related: https://github.com/ClickHouse/ClickHouse/pull/111981 That PR is the interim, test-only fix for the master CI failure linked below: it re-records `query_28` as the ~58 MB the estimator currently reports. **The two changes touch the same line and are mutually exclusive.** If #111981 merges first, this PR must set `query_28` back to the post-fix measurement (expected to land near the original 23722663); if this PR merges first, #111981 should be closed as unnecessary. https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=5baed0f5a10c333cd220b9646d6ef4647425079d&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20distributed%20plan%2C%20parallel%29 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the estimation of the size of aggregate states used by automatic parallel replicas: it measured the serialized size of the states instead of their compressed size, overestimating how much data a replica would send for queries that aggregate into large, compressible states.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111985",
          "createdAt": "2026-07-26T20:39:07Z",
          "updatedAt": "2026-08-13T09:45:18Z",
          "timestamp": "2026-08-13T09:45:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 25
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8300a68b6716c89faf59",
        "signalId": "github:ClickHouse/ClickHouse:issue:114286",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114286",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "date_time_overflow_behavior is ignored when reading out-of-range Date32 values via IcebergS3 table engine (works with icebergS3() table function)",
          "text": "### Company or project name _No response_ ### Describe the unexpected behaviour Reading an Iceberg table that contains Parquet DATE values outside the ClickHouse Date32 range ([-25567, 120530] / 1900-01-01..2299-12-31) throws VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE when querying through the IcebergS3 table engine, even with: SETTINGS date_time_overflow_behavior = 'saturate' The same data, same settings, and same named collection work when reading via the icebergS3() table function. Server default is already date_time_overflow_behavior = ignore, but the table-engine path still throws. ### Which ClickHouse versions are affected? Reproduced on ClickHouse 26.7.2.59 (official build). ### How to reproduce Iceberg table on S3 with at least one Parquet DATE column containing an out-of-range day number (example: -645804, -639653). Mount as table engine: ``` CREATE TABLE iceberg.example ENGINE = IcebergS3( s3_iceberg_bucket, url = 'https://s3.example/lakehouse/db/table', format = 'Parquet' ) SETTINGS allow_dynamic_metadata_for_data_lakes = 1; ``` Fails: ``` SELECT * FROM iceberg.example LIMIT 1 SETTINGS iceberg_timestamp_ms = <snapshot_ts>, date_time_overflow_behavior = 'saturate'; ``` Error (example): Code: 321. DB::Exception: Input value -639653 is out of allowed Date32 range, which is [-25567, 120530]: read stage: ColumnData: column: value_date: (in file/uri .../data.parquet): While executing ParquetV3BlockInputFormat: While executing ReadFromObjectStorage. (VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE) Works: ``` SELECT min(start_date) FROM icebergS3( s3_iceberg_bucket, url = 'https://s3.example/lakehouse/db/table', format = 'Parquet' ) SETTINGS iceberg_timestamp_ms = <snapshot_ts>, date_time_overflow_behavior = 'saturate'; ``` Also observed: input_format_parquet_use_native_reader_v3 = 1 by default; stack mentions ParquetV3BlockInputFormat. Disabling V3 + saturate in dbt query_settings on SELECT from the table engine still failed in our runs. ### Expected behavior Out-of-range dates should be saturated to 1900-01-01 / 2299-12-31 (or ignored, depending on the setting) for both: ``` ENGINE = IcebergS3 icebergS3() table function ``` Behavior should be consistent. ### Error message and/or stacktrace ``` Code: 321. DB::Exception: Input value -639653 is out of allowed Date32 range, which is [-25567, 120530]: read stage: ColumnData: column: value_date: While executing ParquetV3BlockInputFormat: While executing ReadFromObjectStorage. (VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE) version 26.7.2.59 (official build) ``` ### Related issues and pull requests _No response_ ### Additional context Iceberg type date is mapped to ClickHouse Date32. Source data legitimately contains sentinel / dirty dates (e.g. year 0005, 0202, 9999, 5005) that already exist in a MergeTree copy of the same dataset; MergeTree SELECT works, IcebergS3 table engine SELECT throws. Related history: [#51402](https://github.com/ClickHouse/ClickHouse/issues/51402) / [#55696](https://github.com/ClickHouse/ClickHouse/pull/55696) introduced date_time_overflow_behavior for Parquet Date32 overflow; this looks like a remaining gap on the Iceberg table engine read path (or settings not applied there), while the table function path respects the setting. Cluster settings of interest: date_time_overflow_behavior = ignore (default) input_format_parquet_use_native_reader_v3 = 1 Workaround: read via icebergS3(...) + SETTINGS date_time_overflow_behavior = 'saturate' instead of SELECT from ENGINE = IcebergS3.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114286",
          "createdAt": "2026-08-11T08:41:13Z",
          "updatedAt": "2026-08-13T09:38:02Z",
          "timestamp": "2026-08-13T09:38:02Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "unexpected behaviour"
          ],
          "author": "alexsubota",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cbe5fee7e97f357a0eee",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113376",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113376",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `JSONAllValues` text index probe coercion",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113022 Fix incorrect result filtering when a `JSONAllValues` text index serializes a comparison constant using a representation that differs from the JSON subcolumn. This includes value-changing coercions such as `IPv4` to `UInt32`, types such as `Bool` whose semantic equality does not imply identical text, and `DateTime` representations that depend on time zones. Equality and `has` predicates now use the index only when the probe representation is compatible with the statically typed JSON subcolumn or array element type. Equality and `IN` predicates on runtime-typed paths, casts from `Dynamic` paths, and values with session-dependent serialization are evaluated without this index. Safe direct access and identity casts remain accelerated. Casts to `String` remain accelerated for statically typed values with stable serialization. Other value-changing casts decline index use because their stored representation cannot be inferred safely. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix incorrect results from `JSONAllValues` text indexes when comparison values require type coercion or have session-dependent or runtime types.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113376",
          "createdAt": "2026-08-04T19:05:43Z",
          "updatedAt": "2026-08-13T09:36:54Z",
          "timestamp": "2026-08-13T09:36:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [
            "CurtizJ"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d5b9cdff9d13566265a9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111219",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111219",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add basic implementation of  `DROP PARTITION` for Iceberg",
          "text": "This the first which introduces support for ALTER DROP PARTITION on iceberg tables - supports transformations in partition expression e.g - removes only manifests which were present on the moment of execution of a query - does not support catalogs - does not support schema evolution - does not support partition evolution - does not support \"mixed manifests\", when one manifest file has data files from different partitions (AI always mentions this case, but it is only spec-related case, all majors engines do not produces such manifests, anyway we detect and reject such tables) Related: https://github.com/ClickHouse/ClickHouse/pull/105198 Related: https://github.com/ClickHouse/ClickHouse/pull/109288 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `ALTER TABLE ... DROP PARTITION` support for Iceberg tables. It supports only simple tables, no catalogs, without schema and partition evolution.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111219",
          "createdAt": "2026-07-21T11:55:43Z",
          "updatedAt": "2026-08-13T09:36:16Z",
          "timestamp": "2026-08-13T09:36:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-feature",
            "hold"
          ],
          "author": "Diskein",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d3ef15c482e79495b609",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96978",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96978",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix JSON/XML format statistics race condition with parallel replicas",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `rows_read` and `rows_before_limit_at_least` reported as 0 or stale in JSON/XML format output when a query with `LIMIT` reads from remote connections (parallel replicas or the `remote` table function). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) No documentation changes needed — this is a bug fix with no user-facing API changes. --- ## Summary - When parallel replicas are used with `LIMIT`, `rows_read` in JSON/XML output is reported as 0 - Root cause: race condition where `finalizeImpl` writes statistics BEFORE `PipelineExecutor::finalizeExecution` collects remaining progress from connection draining - Fix: two-phase output format finalization — statistics are written AFTER all progress has been collected - The same drain also delivers late `ProfileInfo` packets that update the `rows_before_limit_at_least` / `rows_before_aggregation` counters, so the deferral covers the whole trailer, and the native-protocol `ProfileInfo` sent to the client is refreshed from the post-drain counters ## Approach Split the output format's finalization so statistics are written AFTER `finalizeExecution` collects all remaining progress: - **Phase 1** (during pipeline, in `finalizeImpl`): Write everything EXCEPT the trailer after the `rows` field (`rows_before_limit_at_least`, `rows_before_aggregation`, the `\"statistics\"` section) and closing delimiter - **Phase 2** (after `finalizeExecution`): Re-read the rows-before-* counters, write the trailer and close the document via a finalize callback ### Key changes: - `IOutputFormat`: Added `hasDeferredStatistics`, `writeDeferredStatisticsAndFinalize`, `completeDeferredStatistics` mechanism; `snapshotRowsBeforeCounters` re-reads the shared counters before the deferred trailer is written - `PipelineExecutor`: Added `setFinalizeCallback` called at end of `finalizeExecution` after progress collection - `CompletedPipelineExecutor`: Sets the finalize callback to invoke `completeDeferredStatistics` - `JSONRowOutputFormat`, `XMLRowOutputFormat`, `JSONColumnsWithMetadataBlockOutputFormat`: Override deferred statistics methods; the whole trailer (`rows_before_limit_at_least`, `rows_before_aggregation`, `statistics`) is deferred to phase 2 (`JSONUtils::writeRowsBeforeAndStatistics` factors out the shared writer) - `ParallelFormattingOutputFormat`: Propagates deferred statistics through the parallel formatting pipeline - `LazyOutputFormat` / `PullingOutputFormat`: `getProfileInfo` re-snapshots the rows-before-* counters, so the `ProfileInfo` packet sent to native-protocol clients (TCP `clickhouse-client`, gRPC, `LocalConnection`) carries the post-drain values — by the time it is sent, the executor thread has been joined and the counters are final - New failpoint `tcp_handler_sleep_before_secondary_query_trailing_packets` delays a secondary query's trailing `Totals` / `Extremes` / `ProfileInfo` / `Progress` / `EndOfStream`, making the race deterministic for tests Closes https://github.com/ClickHouse/ClickHouse/issues/85785 ## Test plan - [x] Build succeeds (Release) - [x] Test `00365_statistics_in_formats` passes - [x] JSON format tests pass (`00159_parallel_formatting_json_and_friends_1/2`, `00378_json_quote_64bit_integers`, `00685_output_format_json_escape_forward_slashes`, `01447_json_strings`, `01449_json_compact_strings`, `01486_json_array_output`, `02554_format_json_columns_for_empty`) - [x] XML format tests pass (`00307_format_xml`, `02122_parallel_formatting_XML`) - [x] `03918_statistics_in_formats_parallel_replicas` — forward guard: parallel-replicas `rows_read` matches the single-node value - [x] New deterministic reproducer `04603_rows_before_limit_parallel_replicas_late_packets` — fails without the fix (failpoint forces the trailing packets into the drain window); checks `rows_read` under parallel replicas and `rows_read` + `rows_before_limit_at_least` through the `remote` table function, over HTTP (server-side formatting) and native TCP (client-side formatting), across `JSON`, `JSONColumnsWithMetadata`, and `XML` - [ ] CI: Run with parallel replicas to verify `rows_read` is no longer 0 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches query pipeline execution, cancellation, and output formatting/finalization paths; while scoped to statistics/progress reporting, it changes timing and thread-synchronization and could affect query completion/cancellation behavior under parallel replicas. > > **Overview** > Fixes a race where JSON/XML output could report `rows_read=0` with parallel replicas + `LIMIT` by ensuring final `Progress` packets (drained after cancellation/early completion) are incorporated before statistics are emitted. > > This introduces **two-phase output finalization**: `IOutputFormat` can defer writing statistics/closing delimiters (`hasDeferredStatistics`/`writeDeferredStatisticsAndFinalize`), and `PipelineExecutor` now supports a `setFinalizeCallback` invoked at the end of `finalizeExecution()` after collecting remaining progress; `CompletedPipelineExecutor` wires this to `IOutputFormat::completeDeferredStatistics()`. > > Cancellation/draining logic is tightened to preserve trailing progress and avoid blocking hard cancels: `PipelineExecutor` triggers a `PartialResult` cancel on normal completion, `ISource::cancel` skips `onCancel` for `PartialResult` but allows later escalation, `RemoteSource` adds explicit cancel-reason upgrades and aborts drain via `RemoteQueryExecutor::abortDrain`, and `RemoteQueryExecutor::finish()` drains until all replica connections complete while swallowing expected cancel exceptions. Adds a stateless regression test `03918_statistics_in_formats_parallel_replicas` to assert `rows_read > 0` across JSON/XML formats over TCP/HTTP. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 5818266286a7645dcf1b57212bf9b3214cfd2ec8. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96978",
          "createdAt": "2026-02-15T07:04:09Z",
          "updatedAt": "2026-08-13T09:34:49Z",
          "timestamp": "2026-08-13T09:34:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 53
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:eca09cb3b26c79ccb930",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107943",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107943",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix NOT_IMPLEMENTED exception in system.detached_tables for DatabaseDictionary and similar engines",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/104868 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry: Fixed `SELECT * FROM system.detached_tables` throwing `NOT_IMPLEMENTED` exception when databases like `DatabaseDictionary`, `DatabaseOverlay`, `DatabaseFilesystem`, or `DatabaseS3` exist. Such databases are now silently skipped during iteration. ### Problem `SELECT * FROM system.detached_tables` fails with `Cannot get detached tables for DatabaseDictionary. (NOT_IMPLEMENTED)` when any Dictionary or similar database exists on the server. Root cause: `getDetachedTablesIterator` was called without error handling in both `StorageSystemDetachedTables.cpp` and `StorageSystemTables.cpp`. Databases like `DatabaseDictionary`, `DatabaseOverlay`, `DatabaseFilesystem`, `DatabaseHDFS`, and `DatabaseS3` do not override this method and throw `NOT_IMPLEMENTED`. ### Fix Wrapped `getDetachedTablesIterator` calls in `try-catch` blocks in both affected files. Databases that throw `NOT_IMPLEMENTED` are now silently skipped, allowing the query to return results from supported databases. ### Test Added stateless test `04357_system_detached_tables_not_implemented` covering: - `DatabaseDictionary` coexisting with a normal `Atomic` database - Detached table correctly visible from normal database - No exception thrown when unsupported database engines exist",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107943",
          "createdAt": "2026-06-19T09:03:36Z",
          "updatedAt": "2026-08-13T09:34:35Z",
          "timestamp": "2026-08-13T09:34:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 16
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "adityaksolves",
          "state": "open",
          "assignees": [
            "tuanpach"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4d93ebe8229440f778dd",
        "signalId": "github:ClickHouse/ClickHouse:issue:112909",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112909",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "FULL JOIN USING over a Distributed left table: qualified column t1.a returns the coalesced USING value for right-only rows (or exception 8 at pure defaults)",
          "text": "`FULL JOIN ... USING` where the **left** table is read through `Distributed`: selecting the qualified join column `t1.a` is wrong for right-only rows — and at pure defaults the same query throws an exception. **How to reproduce** (26.8.1.561, any single-shard cluster whose replica is the server itself, e.g. `test_shard_localhost` from the standard test configs): ```sql CREATE TABLE t1 (a UInt16, b UInt16) ENGINE = MergeTree ORDER BY tuple(); CREATE TABLE t2 (a Int16, b Nullable(Int64)) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t1 SELECT number + 1, number FROM numbers(10); INSERT INTO t2 SELECT number - 4, number FROM numbers(10); CREATE TABLE dist_t1 AS t1 ENGINE = Distributed(test_shard_localhost, currentDatabase(), t1); CREATE TABLE dist_t2 AS t2 ENGINE = Distributed(test_shard_localhost, currentDatabase(), t2); -- LOCAL, correct: right-only rows show t1.a = 0 (the UInt16 default) SELECT a, t1.a, t2.a FROM t1 FULL JOIN t2 USING (a) ORDER BY (t1.a, t2.a); -- DISTRIBUTED, wrong values: right-only rows show t1.a = -4..-1 -- (the COALESCED USING value leaks into the qualified left column) SELECT a, t1.a, t2.a FROM dist_t1 AS t1 FULL JOIN dist_t2 AS t2 USING (a) ORDER BY (t1.a, t2.a) SETTINGS prefer_localhost_replica = 0; -- DISTRIBUTED, pure defaults (prefer_localhost_replica = 1): exception -- Code: 8. DB::Exception: Cannot find column `a` in source stream, -- there are only columns: [a, __table1.a, t2.a]. (THERE_IS_NO_COLUMN) SELECT a, t1.a, t2.a FROM dist_t1 AS t1 FULL JOIN dist_t2 AS t2 USING (a) ORDER BY (t1.a, t2.a); ``` Deterministic 20/20 for both manifestations. Characterization: - Wrapping ONLY the left table in `Distributed` is sufficient (right side local: same wrong values / exception). - `RIGHT JOIN` is correct; the defect is FULL-specific (left-only default rows vs right-only coalesced rows). - The bare `a` (USING projection) is correct in all variants — only the QUALIFIED `t1.a` is corrupted/lost. - `prefer_localhost_replica` selects which manifestation appears (1, the default → exception 8; 0 → silent wrong values); everything else is at defaults. Related: https://github.com/ClickHouse/ClickHouse/issues/66739",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112909",
          "createdAt": "2026-08-01T14:44:45Z",
          "updatedAt": "2026-08-13T09:34:32Z",
          "timestamp": "2026-08-13T09:34:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "bug",
            "comp-joins",
            "comp-distributed",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "KochetovNicolai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:92211be0c54ae0885c09",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114610",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114610",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Lazy-load part-level column statistics",
          "text": "When loading a part's statistics, the part currently loads statistics for every column via `getEstimates`, even when only a few columns are needed. On wide tables this performs a lot of pointless deserialization. Related: https://github.com/ClickHouse/ClickHouse/pull/104691 This change makes per-part statistics load lazily, only for the requested columns: - `IMergeTreeDataPart::getEstimates` now takes the set of requested columns, loads only the missing ones, and caches the result per column. A deterministic miss (no statistics declared, file absent, or corrupted) is recorded as a `nullopt` entry so the column is not re-probed on every query. - Part pruning asks for only the filter columns: `StatisticsPartPruner` exposes `getCandidateColumns` and `filterPartsByStatistics` passes them to `getEstimates`. - `system.parts_columns` discovers which columns have statistics via the new `getColumnsWithStatistics` and loads only those. A new `LoadedStatisticsColumns` profile event counts the per-column statistics that are actually deserialized. The performance test measures a 100-column wide table with 50 parts and a query filtering on 3 columns; part pruning loads 150 columns instead of 5000. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Part-level column statistics are now loaded lazily, only for the columns that are actually needed, reducing statistics deserialization on wide tables.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114610",
          "createdAt": "2026-08-13T09:17:26Z",
          "updatedAt": "2026-08-13T09:33:44Z",
          "timestamp": "2026-08-13T09:33:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "hold"
          ],
          "author": "zoomxi",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:12d854820342675ff38c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110072",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110072",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix bitmap subset functions for small-set and promoted signed bitmaps",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109974 Related: https://github.com/ClickHouse/ClickHouse/issues/106208 `subBitmap` and `bitmapSubsetOffsetLimit` apply offset/limit in ascending value order. The small-set bitmap path iterated keys in insertion order instead, producing wrong subsets when values were not inserted sorted (e.g. `bitmapBuild([5, 4, 1, 2, 3])`). The small paths of `rb_range` and `rb_limit` also compared element values as `UInt32`, so values above `2^32` in `UInt64` / `Int64` bitmaps were truncated and failed to match their own thresholds — for example `bitmapSubsetLimit(bitmapBuild([4294967297]::Array(UInt64)), 4294967297, 1)` returned an empty bitmap. Beyond that, this aligns `bitmapSubsetInRange` / `bitmapSubsetLimit` / `bitmapMin` / `bitmapMax` / `bitmapContains` / `bitmapTransform` so that small and promoted bitmaps compare elements in the same unsigned element-type domain. That is the domain signed element types were introduced with in https://github.com/ClickHouse/ClickHouse/pull/20171: `01702_bitmap_native_integers` has asserted `bitmapMin` = `251` and `bitmapMax` = `255` for `Int8` `[-1, -2, -3, -4, -5]` since 2021, while `rb_range` and `rb_limit` were left comparing sign-extended `UInt32` values and were annotated at the time as \"currently only support UInt32\". Master therefore contradicted itself: `bitmapSubsetInRange(bm, bitmapMin(bm), bitmapMin(bm) + 1)` returned nothing for `Int8`, `Int16` and `Int64` bitmaps, even though `bitmapMin` reported an element that is present. `BSINumericIndexedVector` looks indexes up in the same bitmaps, so it now uses that domain as well. Without it, `groupNumericIndexedVector` returned different results for `Int8` / `Int16` index columns than for wider index types once a bit-slice bitmap was promoted past the small set, and `numericIndexedVectorGetValue` never found a negative index at all. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect results from `subBitmap`, `bitmapSubsetInRange` and `bitmapSubsetLimit` for bitmaps still held in the small representation, including `UInt64` and `Int64` element values above `2^32`, which were truncated to 32 bits. Comparisons in bitmap functions over signed element types now consistently use the unsigned value of the element type, so in an `Int8` bitmap the element `-1` is compared as `255` instead of as the sign-extended `4294967295`; this also fixes `bitmapMin`, `bitmapMax`, `bitmapContains` and `bitmapTransform` on bitmaps that have grown past the small representation. Queries that passed sign-extended thresholds have to be adjusted: over an `Int8` bitmap, `bitmapSubsetInRange(bm, 4294967168, 4294967296)` becomes `bitmapSubsetInRange(bm, 128, 256)`. Also fixed `groupNumericIndexedVector` returning different results for `Int8` and `Int16` index columns than for wider index types, and `numericIndexedVectorGetValue` returning `0` for negative indexes. Corrected the bitmap function documentation, including the subset functions that were described as using 1-based indexing and the signed `bitmapBuild` / `bitmapToArray` support. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1309` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110072",
          "createdAt": "2026-07-11T01:18:01Z",
          "updatedAt": "2026-08-13T09:33:43Z",
          "timestamp": "2026-08-13T09:33:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "RamiDarwiche",
          "state": "closed",
          "assignees": [
            "yakov-olkhovskiy"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:80429f771d78a301dbd2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104809",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104809",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `NOT_FOUND_COLUMN_IN_BLOCK` in `query_plan_convert_join_to_in` with `arrayJoin` JOIN key",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `NOT_FOUND_COLUMN_IN_BLOCK` exception in `INNER JOIN ... ON arrayJoin(...) = ...` queries when `query_plan_convert_join_to_in` is enabled and the SELECT references the source array column. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- Not applicable -- bug fix, no new feature or API change. --> --- ### Problem With `SET query_plan_convert_join_to_in = 1` (off by default), the optimizer rewrites a hash `INNER JOIN ON arrayJoin(L.col) = R.col` into `[Expression(\"Calculate join left keys\"), Filter(\"IN\"), DelayedCreatingSets, Expression(\"Join output actions\")]`. If the SELECT projection (or any post-JOIN expression) references the source array column `L.col` itself, or `arrayJoin(L.col)` again, the rewritten plan throws ``` Code: 10. DB::Exception: Not found column __table1.tags in block ... (NOT_FOUND_COLUMN_IN_BLOCK) ``` at execution (and at `EXPLAIN actions = 1` planning time). Minimal repro: ```sql CREATE TABLE lt (id UInt64, tags Array(String)) ENGINE = MergeTree ORDER BY id; CREATE TABLE rt (tag_id String) ENGINE = MergeTree ORDER BY tag_id; INSERT INTO lt VALUES (1, ['a','b','c']), (2, ['d','e']); INSERT INTO rt VALUES ('a'), ('d'); -- Throws NOT_FOUND_COLUMN_IN_BLOCK: SELECT lt.id, lt.tags FROM lt INNER JOIN rt ON arrayJoin(lt.tags) = rt.tag_id SETTINGS query_plan_convert_join_to_in = 1; -- Same with arrayJoin in the projection -- also throws: SELECT lt.id, arrayJoin(lt.tags) FROM lt INNER JOIN rt ON arrayJoin(lt.tags) = rt.tag_id SETTINGS query_plan_convert_join_to_in = 1; ``` ### Root cause `tryConvertJoinToIn` builds `left_pre_join_actions = JoinExpressionActions::getSubDAG(<left join keys>)`. The output set is only the JOIN-key expressions — `arrayJoin(__table1.tags)` and `__table1.id` in the repro — **not** the source `__table1.tags` array column. Later, `cloneSubDAGWithHeader(output_header, JoinExpressionActions::getSubDAG(join_output_actions))` clones the post-JOIN expressions onto a new DAG seeded from `output_header`. `cloneSubDAGWithHeader` calls `mergeInplace(..., remove_dangling_inputs=true)`, which remaps `INPUT` nodes by name to the corresponding inputs in `output_header`. When `join_output_actions` references `__table1.tags` (because SELECT did), the source `__table1.tags` INPUT in `second_dag` has no match in `output_header` — `mergeInplace`'s `remove_dangling_inputs` only removes input nodes that collide with `first`'s inputs, not input nodes that have no counterpart at all. The non-INPUT nodes that depend on it (the cloned ARRAY_JOIN node when SELECT had `arrayJoin(lt.tags)`, or the column reference itself when SELECT had `lt.tags`) survive, and the resulting `ExpressionStep` cannot find the column at execution. ### Fix Decline the conversion in `tryConvertJoinToIn` whenever any input of `join_output_actions` is not among the outputs that `left_pre_join_actions` forwards. The check runs immediately after `left_pre_join_actions` / `right_pre_join_actions` are computed and **before** any plan mutation (`makeExpressionNodeOnTopOf` etc.), so the bail-out is safe — the query then runs through the normal JOIN path and produces the correct result. ```cpp { auto join_output_actions_subdag = JoinExpressionActions::getSubDAG(join_output_actions); std::unordered_set<std::string_view> forwarded_columns; for (const auto * out : left_pre_join_actions.getOutputs()) forwarded_columns.insert(out->result_name); for (const auto * input : join_output_actions_subdag.getInputs()) if (!forwarded_columns.contains(input->result_name)) return 0; } ``` The check mirrors the established `appendInputsForUnusedColumns` style (`ActionsDAG.cpp:1453-1461`), which is the idiomatic \"verify every input is present in a sample block\" pattern in this codebase. ### Related fixes for the same root cause This is the third bug in 12 months rooted in `mergeInplace` / `clone` of an ARRAY_JOIN-bearing DAG being merged into a context whose stream header does not provide a required INPUT column. The prior two are: - #96989 (prevention strategy: blacklist `ARRAY_JOIN` in the pushdown eligibility check). - #97239 (bail-out strategy: detect ghost `ARRAY_JOIN` after the merge and abandon the optimization). - #104785 (post-hoc cleanup strategy: snapshot the merged-in `ARRAY_JOIN` pointers and drop them after the merge via a new `ActionsDAG::removeNodes` helper). This PR follows the bail-out pattern of #97239 — `tryConvertJoinToIn` is an optional optimization, so declining the rewrite is safe and the query falls back to the regular JOIN path. A more invasive follow-up could forward the missing left-side columns through `left_pre_join_actions` so the optimization stays on even in these cases; left for a separate PR once we have evidence that real workloads hit the bail-out frequently. The setting `query_plan_convert_join_to_in` is off by default, so the bug is reachable only when explicitly enabled. ### Regression test `tests/queries/0_stateless/03918_convert_join_to_in_arrayjoin_dangling.{sql,reference}` — co-located with #96989's `03918_arrayjoin_function_with_join_and_where`. Four queries: 1. `SELECT lt.id ... INNER JOIN ... ON arrayJoin(lt.tags) = rt.tag_id` — the optimization stays on, returns the correct rows (validates the fix's \"no-regression\" guarantee for the cases it does not bail out on). 2. `SELECT lt.id, lt.tags ...` — before fix: NOT_FOUND_COLUMN_IN_BLOCK; after fix: correct result via the fallback JOIN path. 3. `SELECT lt.id, arrayJoin(lt.tags) ...` — same shape as (2). 4. Control: same as (3) with `query_plan_convert_join_to_in = 0`. Confirms the fallback path returns the same answer the fix produces in (3). Repro-first verified: queries (2) and (3) were observed to fail with NOT_FOUND_COLUMN_IN_BLOCK on un-fixed upstream master (`7bd0fa28146`) before the reference was locked.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104809",
          "createdAt": "2026-05-13T09:09:21Z",
          "updatedAt": "2026-08-13T09:33:28Z",
          "timestamp": "2026-08-13T09:33:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6b4f078bea9faf50e24e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104691",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104691",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Lazy-load column statistics during query planning",
          "text": "When the query planner needs column statistics — for join reordering, for prewhere selectivity estimation, or for part pruning — it currently loads statistics for every column of the table from disk on the first access, even when the query only filters or joins on a handful of columns. This PR reduces statistics-file I/O during query planning on wide tables. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Column statistics are now loaded on demand for only the columns the query planner needs instead of for every column of the table on first access.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104691",
          "createdAt": "2026-05-12T11:13:36Z",
          "updatedAt": "2026-08-13T09:29:44Z",
          "timestamp": "2026-08-13T09:29:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 15
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "zoomxi",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:42c0670c205e7cf80b83",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114530",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114530",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject a Parquet offset index whose first page does not start at row 0",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/114464 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114464 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Parquet files whose offset index does not start at row 0. Such a file could make the native Parquet reader return wrong rows, or report `LOGICAL_ERROR` instead of `INCORRECT_DATA`. The offset index is now validated when it is read. ### Description Reported by @ PedroTadim in #114464 ([his request to take it](https://github.com/ClickHouse/ClickHouse/issues/114464#issuecomment-5265506353)): a single flipped bit in a Parquet offset index made the v3 reader raise `Row passes filters but its page was not selected for reading. This is a bug.` That `LOGICAL_ERROR` aborts on debug and sanitizer builds, and in release blames ClickHouse for a defect in the file. Reachable from untrusted file content on defaults. Root cause: `decodeOffsetIndex`, the only place the offset index is deserialized, validated byte ranges, monotonicity and `first_row_index < num_rows`, but never that the *first* page starts at row 0, as the spec requires since a chunk covers every row of its row group. Page ends come from the next page's `first_row_index` while the row sweep covers `[0, num_rows)`, so a nonzero anchor leaves rows `[0, anchor)` described by no page. The reported abort is the mildest of three symptoms. For a V1 data page in an array column `num_rows_in_page` stays unset, so the row-count cross-check is skipped, the corrupt anchor seeds the row cursor and the read returns wrong rows with no error: on a ClickHouse-written file, `WHERE id = 10` returned the row-6 payload. The anchor is now validated in `decodeOffsetIndex`, which runs before every consumer, so all three become a clear `INCORRECT_DATA`. `Page doesn't contain requested row` is reclassified too (a page the offset index lists can turn out to be an index or dictionary page, which the reader skips). `Row passes filters ...` stays `LOGICAL_ERROR`: with a validated anchor it is unreachable from file content, so it remains the page-selection tripwire @ PedroTadim asked to keep. Such a file is now rejected rather than read, but it already aborts or returns wrong rows. ClickHouse's writer always anchors page 0 at row 0, and all 124 offset-index fixtures in `tests/` conform.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114530",
          "createdAt": "2026-08-12T18:14:39Z",
          "updatedAt": "2026-08-13T09:24:07Z",
          "timestamp": "2026-08-13T09:24:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8116743378a3c42de439",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114599",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114599",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "AST fuzzer: do not create a view that duplicates a non-parallel sink",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110166 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/110166 Reported by @ alexey-milovidov there: a `Stress test (arm_release)` hung check on an `INSERT` stuck in `StorageLog::write`. https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110166&sha=c890bb8e35d0d9138c1fa770e1fd4a1db05f39bf&name_0=PR&name_1=Stress%20test%20%28arm_release%29 It is not lock contention under load. Fuzzing a `CREATE MATERIALIZED VIEW` renames the view but leaves its external `TO` untouched, so the clone shares source *and* target with the original and one `INSERT` builds two sinks for that target. `LogSink` holds the table's exclusive lock from pipeline build until `onFinish`, so the second sink waits out `lock_acquire_timeout` and throws `Code: 159`; the `INSERT` re-arms, the processlist never empties, and the hung check fails. Reproducible in 4 statements with no concurrency and no fuzzer, on three build flavours. This skips such a fuzzed query instead of creating the view. Whether it is hazardous is a sink-graph question, so it is answered in `InsertDependenciesBuilder`, which owns that graph, and it errs only towards skipping: safe shared-target fuzzing keeps running, and a sole writer of a non-parallel target is unaffected. `Buffer` and `Distributed` end a branch, since the write they forward runs as an `INSERT` of their own whose sinks this walk does not model. Anything the walk cannot decide executes normally and is counted, so an undecided answer never suppresses fuzzing. Two `ProfileEvents` make both outcomes attributable; the path is confined to `ast_fuzzer_runs > 0`. A proven duplicate survives a later branch that throws while materializing a lazy storage. And since the answer describes the catalog, where a view appears only once its `CREATE` has executed, the fuzzer reserves the (source, target) pair across the decision and the create; the contract states that boundary. Refusing two writers on one `Log`-family target at INSERT planning time is out of scope: that is user-visible, and on #74080 the same topology was answered as unsupported. <details><summary>Validation</summary> Ground truth measured on a pre-fix binary by building the duplication by hand, against the predicate's verdict on the same 9 topologies (agreement is 9/9): | topology | wedges pre-fix | skipped | |---|---|---| | `src -> mv -> tgt(TinyLog)` | yes | yes | | `src -> mvA -> mt(MergeTree) -> mvB -> lg(TinyLog)` | yes | yes | | `Alias`-hidden: `mv TO Alias(mt)`, `mt -> mv2 -> lg` | yes | yes | | deep cascade `mt1..mt8 -> lg` | yes | yes | | `Alias(MergeTree) -> view -> Memory` | no | no | | directly shared `Buffer` | no | no | | directly shared `Distributed` | no | no | | stale edge (`MODIFY QUERY` away from `mt`) | no | no | | shared `MergeTree` | no | no | Plus the hazard behind a lazily loaded table (`lazy_load_tables`, wedges pre-fix, skipped) and its `MergeTree` counterpart (skipped in neither). Through the real carrier (`--ast_fuzzer_runs=40 --ast_fuzzer_any_query=1`): pre-fix the `INSERT` fails `Code: 159`; with the fix the hazardous arms are skipped, the insert completes, and the safe arms are untouched. Removing the skip restores the failure; removing the name withdrawal makes 345 of 480 later fuzzed queries name a view that was never created (0 with it). `04876_ast_fuzzer_skips_duplicate_non_parallel_sink` states each claim against its own fixture's query_log rows rather than server-global counters, so a concurrent fuzzing process cannot move them: 30/30 green under a deliberate adversary, 50/50 randomized green, 32 consecutive runs in one database with nothing left behind, and eight mutations each flipping exactly their own assertion. It fails with the skip removed (`Code: 159`) and with the proxy hop reverted. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114599",
          "createdAt": "2026-08-13T07:00:18Z",
          "updatedAt": "2026-08-13T09:16:52Z",
          "timestamp": "2026-08-13T09:16:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a60fcbd86fc7b7253a03",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110695",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110695",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Combine I/O cost with selectivity in PREWHERE condition ordering",
          "text": "When `use_statistics=1` (the default since `auto_statistics_types` was introduced), the PREWHERE optimizer sorted conditions by `estimated_row_count` alone, with `columns_size` only as a tiebreaker. This caused expensive conditions (e.g. Map column, ~500KB) to be placed before cheap ones (e.g. scalar column, ~1KB) whenever the expensive condition appeared more selective — ignoring the I/O cost difference. Apply the classic conjunctive filter ordering rule: sort by `cost / (1 - selectivity)`, i.e. the I/O cost per rejected row. This is computed as `columns_size / max(1, total_rows - estimated_row_count)` and replaces the separate `estimated_row_count, columns_size` pair in the condition comparison tuple. When statistics are unavailable (`estimated_row_count=0`, `total_rows=0`), the formula degrades to `columns_size`, preserving the existing behavior. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Combine I/O cost with selectivity in PREWHERE condition ordering. It fixes performance regression in PREWHERE execution in some cases introduced after https://github.com/ClickHouse/ClickHouse/pull/101275. Part of https://github.com/ClickHouse/ClickHouse/issues/110462 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1177` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110695",
          "createdAt": "2026-07-16T13:16:19Z",
          "updatedAt": "2026-08-13T09:11:02Z",
          "timestamp": "2026-08-13T09:11:02Z",
          "metrics": {
            "reactions": 1,
            "comments": 6
          },
          "labels": [
            "pr-performance",
            "pr-synced-to-cloud"
          ],
          "author": "Avogar",
          "state": "closed",
          "assignees": [
            "hanfei1991"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bdca78235241320384ee",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114394",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114394",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject a data lake schema whose column name is empty",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/114350 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114350 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a crash when reading a data lake table whose schema declares a column with an empty name. Malformed Iceberg metadata is rejected with `ICEBERG_SPECIFICATION_VIOLATION`; a schema supplied by a catalog, Delta Lake or Paimon is rejected with `AMBIGUOUS_COLUMN_NAME`. The check is unconditional, so an Iceberg table that merely retains an unused historical schema with an empty field name also becomes unreadable instead of aborting once that schema is read. ### Description Every lake reader copies the provider's field name into a `NamesAndTypesList` unvalidated. That becomes the table's column list, so an empty name reached `ASTIdentifier`, whose constructor asserts no identifier part is empty. `SELECT *` converts the query tree back to an AST during planning, so it aborted the server on debug and sanitizer builds, while `DESCRIBE` and `SELECT count()` succeeded and made the table look readable. In release the assert compiles out and the name surfaces one layer down. An empty column name is unrepresentable in ClickHouse, but nothing validated the point where an external lake schema becomes the table structure. The check goes there rather than into each provider's parser, so it also covers readers added later, and reuses the code `Block::insert` already returns. Three call sites. `tryGetTableStructureFromMetadata` and `buildStorageMetadataFromState` are members of one template class instantiated over every `IDataLakeMetadata` subclass, covering Iceberg, both Delta readers, Paimon and Hudi; the second is needed because the schema-reload path skips the first. `DatabaseDataLake` builds columns straight from the catalog and reaches neither, so `TableMetadata::setSchema` covers the REST, Glue, Unity, Hive and PaimonRest catalogs. Iceberg keeps the earlier check in its schema processor, which fires first when the name comes from metadata.json; a catalog that builds the column list itself, such as Glue, does not reach it, so an Iceberg table read that way reports `AMBIGUOUS_COLUMN_NAME`. Validation: both Delta readers and Paimon aborted before the change and now report the error, measured under `allow_experimental_delta_kernel_rs` 0 and 1. Well-named lakes still read on every arm, and the test fails on a build without this change. The catalog site is covered by a Glue integration test asserting on `SHOW CREATE TABLE`: a read is not a usable oracle there, because the empty name also reaches `Block::insert` and yields the same code without the check.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114394",
          "createdAt": "2026-08-11T22:59:21Z",
          "updatedAt": "2026-08-13T09:09:41Z",
          "timestamp": "2026-08-13T09:09:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:827fb52d9e72a1130a87",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114053",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114053",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not run the LeakSanitizer check on the forced exit path",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/112846 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description On ASan builds the server's forced-shutdown branch calls `safeExit(0)`, which runs `__lsan_do_leak_check()`. That branch is taken precisely because handlers or refresh tasks missed the drain timeout, so the check stops the world while those threads are mid-query and classifies chunks they still own, producing unattributable reports. Two were reported on 2026-08-08 as a `uniqExact` hash-table \"direct leak\" (2 MiB on #112846, [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112846&sha=2e6a38112ad4dacef62f5cc4cdb396cc297282eb&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29); 4 MiB on #101264): each a single allocation with zero indirect leaks, while two green `asan_ubsan` jobs on the same commit also force-shutdown with connections live and report nothing. Since #113608 made such reports nameable, they compete with real ones. `safeExit` takes a second parameter, a `LeakCheck` enum defaulted to `Run`, so existing callers are unchanged. The three callers reached while other threads still run pass a skip: `Server.cpp` writes one stderr line recording that coverage was skipped rather than passed; `clickhouse-local` and the client's SIGINT/SIGQUIT handler skip **silently**, since neither redirects fd 2, so a notice there would be program output read as a test failure. Keeper's identical-looking branch is unchanged: it calls `server_pool.joinAll()` first. Clean shutdowns never reach `safeExit`. That notice is allow-listed in the one scanner that can see it, `sanitizer_hits` in `ci/jobs/scripts/clickhouse_proc.py`, matched as a whole line so a report sharing a line with it is still blamed; a new `ci/tests` test pins that. `tests/clickhouse-test` needs none: its `IGNORED_SANITIZER_ERRORS` filters only the sanitizer runtime's `log_path` files, and the carriers that runner sees skip quietly. Validated on an ASan build: the notice appears on the forced branch and no leak report does, while the same tree with the guard removed runs the check. Side effect: `_exit()` drops the async logger queue, so `Will shutdown forcefully.` is now missing on ASan builds, as it already is where no check delays the exit. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1312` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114053",
          "createdAt": "2026-08-09T17:34:15Z",
          "updatedAt": "2026-08-13T09:31:29Z",
          "timestamp": "2026-08-13T09:31:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:29d19f1806f3db5ceb9a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114518",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114518",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Raise a catchable error instead of LOGICAL_ERROR on spec-violating Iceberg metadata",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114487 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reading an Iceberg table whose metadata describes a schema evolution that the Iceberg specification forbids no longer aborts the server. Four spec violations in `IcebergSchemaProcessor` were reported as `LOGICAL_ERROR`, which is treated as a failed assertion, or reached a fatal assertion unchecked; they now raise `ICEBERG_SPECIFICATION_VIOLATION`. ### Description Validators in `IcebergSchemaProcessor` reject genuinely invalid Iceberg metadata, but did so with `ErrorCodes::LOGICAL_ERROR`, which `Exception::handleErrorCode` treats as a failed assertion (`src/Common/Exception.cpp:112-116`). Metadata content alone took the server down: no `ALTER`, no write path, a corrupt or crafted `.metadata.json` is enough. The three existing detections are correct, so only their error code changes. A fourth shape had no check. Sites fixed, all reachable from metadata content: - `getSchemaTransformationDag`: an old struct/list/map field becomes a primitive. Its message named the opposite direction and is corrected. - `getSchemaTransformationDag`: a new field with no old counterpart is `required` with no default. The frame in the reported stack. - `registerSnapshotWithSchemaId`: one `snapshot-id` bound to two `schema-id`s. Not named in the issue, but `IcebergMetadata.cpp:384-389` registers every entry of the `snapshots` array, both values read straight out of the JSON. - `getSchemaTransformationDag`: the mirror shape, an old primitive becoming a struct, list or map. The complex branch keyed off the new field's type alone, so it built `EvolutionFunctionStruct` and aborted in its `lazyInitialize` `chassert` instead of reporting anything. Now rejected before the transform is built. This file uses both conventions. `Utils.cpp:1196-1199` picks 743 over `BAD_ARGUMENTS` for a versioned specification rule broken by a file being read, which is what all four sites are, matching `SchemaProcessor.cpp:419`; `MetadataGenerator.cpp:452` keeps `BAD_ARGUMENTS` for the write path. The exposure is wider than the issue states: `abort_on_logical_error` (`ServerSettings.cpp:1433`) is linked by `tests/config/install.sh:244-246` for every non-fast-test stateless run, so release-build CI aborts here too. Not touched: `SchemaProcessor.cpp:189/194` fire when stringifying an already-parsed object, a genuine internal invariant. Manifest-file sites need their own proof. A new stateless test covers all four sites. On master it kills the server (`[ FAIL ] Reason: server died`); with the fix each returns `Code: 743`, and 50/50 randomized runs pass.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114518",
          "createdAt": "2026-08-12T16:14:30Z",
          "updatedAt": "2026-08-13T09:05:31Z",
          "timestamp": "2026-08-13T09:05:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5cf630eb8830631569ba",
        "signalId": "github:ClickHouse/ClickHouse:issue:112905",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112905",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Unexpected results with a minmax index and a NULL-containing `SOME` comparison",
          "text": "### Company or project name _No response_ ### Describe what's wrong `c1 = SOME([1, NULL])` is lowered to `has([1, NULL], c1)`. For `c1 = 0`, that is a definite, non-NULL **false** — `0` is not `1`, and \"not found, but the array also contains NULL\" evaluates to false, not unknown, per the row's own projection. So `NOT (c1 = SOME([1, NULL]))` must be true. With a `minmax` skip index on `c1`, the query instead returns **zero rows**, and so do the direct predicate and the `IS NULL` check — the row vanishes from all three. ```sql CREATE TABLE t (c1 INT) ENGINE = MergeTree ORDER BY tuple(); CREATE INDEX idx ON t(c1) TYPE minmax GRANULARITY 1; INSERT INTO t VALUES (0); SELECT c1, NOT (c1 = SOME([1, NULL])) AS p FROM t; -- c1=0, p=1 (true) SELECT count() FROM t WHERE NOT (c1 = SOME([1, NULL])); -- Expected: 1 -- Actual: 0 ``` *Found by an automated fuzzing & triage agent; the repro is verified but the analysis may be wrong.* ### Does it reproduce on the most recent release? Yes ### How to reproduce Can be reproduced on 26.7.1.1315 ```sql CREATE TABLE t (c1 INT) ENGINE = MergeTree ORDER BY tuple(); CREATE INDEX idx ON t(c1) TYPE minmax GRANULARITY 1; INSERT INTO t VALUES (0); SELECT count() FROM t WHERE (c1 = SOME([1, NULL])); -- 0 (correct) SELECT count() FROM t WHERE NOT (c1 = SOME([1, NULL])); -- 0 (WRONG, expect 1) SELECT count() FROM t WHERE ((c1 = SOME([1, NULL]))) IS NULL; -- 0 (correct) -- 0 + 0 + 0 = 0, but the table has 1 row. -- Disabling data-skipping indexes for this query alone (no schema/data -- change) restores the correct result: SELECT count() FROM t WHERE NOT (c1 = SOME([1, NULL])) SETTINGS use_skip_indexes = 0; -- 1 (correct) -- Removing the NULL array element also restores correct indexed execution, -- isolating NULL inside the membership set as part of the trigger: SELECT count() FROM t WHERE NOT (c1 = SOME([1])); -- 1 (correct) ``` ### Expected behavior As mentioned above. ### Error message and/or stacktrace _No response_ ### Related issues and pull requests [#106948](https://github.com/ClickHouse/ClickHouse/issues/106948) / [#110266](https://github.com/ClickHouse/ClickHouse/issues/110266) — `minmax`/`KeyCondition` mis-negating a float range in the presence of NaN. ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112905",
          "createdAt": "2026-08-01T14:07:36Z",
          "updatedAt": "2026-08-13T09:04:34Z",
          "timestamp": "2026-08-13T09:04:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "bug",
            "comp-skip-index",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "suyZhong",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e009c6bf908d4f869b25",
        "signalId": "github:ClickHouse/ClickHouse:issue:114605",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114605",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Infinite uncancellable loop in function `hop` at analysis time: a window interval whose span wraps to 0 modulo 2^32 dodges the time-overflow guard",
          "text": "🕵 Found by the AST fuzzer in the stress test of a CI run ([Stress test (arm_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112930&sha=1c5bf0b613146bd6666b6d2bd2be55f539f6c0d8&name_0=PR&name_1=Stress%20test%20%28arm_debug%29), `Hung check failed, possible deadlock found`): the fuzzed query hung for 1338 s with `is_cancelled: 1` and was still spinning when the hung check gave up. Reproducer (hangs forever, also in `clickhouse local` on current master): ```sql SELECT hop(toDateTime32('1969-12-31'), toIntervalDay(1), toIntervalDay(2147483648), 'US/Samoa') ``` Because the arguments are constants, the loop runs during constant folding in `QueryAnalyzer::resolveFunction`, i.e. at analysis time, where nothing checks the cancellation - the query cannot be killed (`KILL QUERY` marks it cancelled, but the thread never looks). The stack of the hung thread is `TimeWindowImpl<HOP>::dispatchForColumns` <- `IExecutableFunction::execute` <- `QueryAnalyzer::resolveFunction` <- ... <- `InterpreterExplainQuery::execute` (the fuzzed query was an `EXPLAIN`). The mechanism is in `executeHop` (`src/Functions/FunctionsTimeWindow.cpp`). `AddTime<Day>` is a wrapping `UInt32` computation, `static_cast<UInt32>(t + delta * 86400)`. With `window_num_units = 2147483648 = 2^31`, the subtraction of the window is `2^31 * 86400 = 43200 * 2^32 ≡ 0 (mod 2^32)`, so ```cpp wstart = AddTime<kind>::execute(wend, -window_num_units, time_zone); // wstart == wend: subtracting the window is a no-op modulo 2^32 if (wstart > wend) throw Exception(ErrorCodes::BAD_ARGUMENTS, \"Time overflow in function {}\", name); // does not fire: equal, not greater ``` the time-overflow guard added by https://github.com/ClickHouse/ClickHouse/pull/61523 (for the same class of hang, https://github.com/ClickHouse/ClickHouse/issues/61521) does not fire. The following loop ```cpp do { wend_latest = wend; wend = static_cast<ToType>(AddTime<kind>::execute(wend, -hop_num_units, time_zone)); } while (wend > time_data[i]); ``` then decrements `wend` by one day at a time past zero, where it wraps back to ~2^32. Every value it takes stays congruent to the starting point modulo `gcd(86400, 2^32) = 128`, so it never hits a value `<= time_data[i]` exactly when the start is not congruent to a small value, and the loop never terminates. The same loop shape exists in `executeHopSlice` (`windowID`). The guard needs to catch a wrapped subtraction that lands exactly on (or above) `wend`, and the loop needs to detect the wrap of `wend -= hop` (the new `wend` coming out greater than the old one) instead of relying on the comparison with the time.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114605",
          "createdAt": "2026-08-13T09:03:36Z",
          "updatedAt": "2026-08-13T09:03:36Z",
          "timestamp": "2026-08-13T09:03:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug",
            "fuzz"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b6400a2c4009ac82cd76",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110833",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110833",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Compare stored table definition expressions by AST instead of formatted text",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/92340 Related: https://github.com/ClickHouse/ClickHouse/pull/110840 Related: https://github.com/ClickHouse/ClickHouse/pull/108590 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed cross-version metadata compatibility for tables whose keys and other definition expressions were written with redundant parentheses (`PARTITION BY (a)`, `ORDER BY (b, c)`, `DEFAULT (a + 1)`, `TTL (d + INTERVAL 10 YEAR)`, `INDEX ix (b * c)`, `CONSTRAINT c CHECK (a > 0)`, `PROJECTION p (SELECT (b) ...)`) or with an explicit single-element `tuple(a)`. #92340 started preserving these parentheses in stored metadata, so `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` between two such tables failed with `Tables have different ...` even on a single server version, and a replica reading metadata written by another version failed with `METADATA_MISMATCH`. Stored table definition expressions are now compared by their ASTs (`getTreeHash`), not by their formatted text. ### Description #92340 started to preserve redundant parentheses in the stored definition ASTs. Older versions stored the canonical form without them, so the metadata serialized by the two versions differs, and every comparison (`ReplicatedMergeTree` replica join, `ATTACH`/`REPLACE`/`MOVE PARTITION FROM`) rejected definitions that are actually equal. Two sub-cases with different affected ranges: - Redundant parentheses (`PARTITION BY (a)` vs `PARTITION BY a`). This is the #92340 regression: measured on released binaries, 25.8, 26.3 and 26.4 accept `ATTACH PARTITION FROM` between the two forms while 26.5, 26.6 and 26.7 reject it, matching the branches that contain `b38a892dfbfab6`. It does **not** require an upgrade or a mixed-version cluster: both tables created by the same 26.7 binary already fail, because the comparison is between two in-memory ASTs of which only one carries the `parenthesized` flag. - An explicit single-element `tuple(a)` vs `a`. This one fails on every version tried, 25.8 included, so it is a long-standing bug rather than a #92340 regression. It is fixed here as well because `extractKeyExpressionList` unwraps the optional `tuple(...)`. Following https://github.com/ClickHouse/ClickHouse/pull/110833#issuecomment-5040213065, this replaces the earlier text-canonicalization approach entirely: - Expressions are compared by their ASTs using `getTreeHash`, not as text or reformatted text. Comparing formatted text is wrong because the text depends on the formatting logic, which can change for unrelated (e.g. aesthetic) reasons, while the stored form may have been written by any past server version. - The comparison method is reusable: `sameAST` in `src/Parsers/IAST.h` (aliases significant, `nullptr`-tolerant overload for optional expressions). - It is applied at every place where serialized, stored expressions are compared to check whether the table was altered or changed unexpectedly: - `ReplicatedMergeTreeTableMetadata::checkImmutableFieldsEquals` / `checkEquals` / `checkAndFindDiff` (replica join, `ALTER` application). The stored strings are parsed purely syntactically (without resolving against any column set), so the fields of an `ALTER` log entry that adds a column and changes a key in one go compare safely; the `columns`/`context` parameters are gone. Keys additionally go through `extractKeyExpressionList`, so `a`, `(a)` and `tuple(a)` are the same single-column key while `a DESC` stays different. - `MergeTreeData::checkStructureAndGetMergeTreeData` (the `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` structure gate): sorting/partition/primary keys and the secondary-index/projection definition sets. - `StorageReplicatedMergeTree::checkTableStructureAttempt`: the columns comparison (`ColumnsDescription`/`ColumnDescription`/`ColumnDefault` equality now compares the default/codec/TTL expressions as ASTs). The column-by-column comparison also has to ignore in-memory-only state that is never serialized (implicit statistics and the auxiliary `data_type` of `ColumnStatisticsDescription`, hence the new `hasSameExplicitStatistics`), otherwise every replicated `CREATE` would report `INCOMPATIBLE_COLUMNS` against the columns it had just written to ZooKeeper. - `StorageReplicatedMergeTree::alter`: detection of which metadata fields an `ALTER` actually changed. The changed fields are also written back to Keeper through the same backward-compatible serializers the `ReplicatedMergeTreeTableMetadata` constructor uses (`formatDefinition` / `formatDefinitionList`), so an `ALTER` never publishes a parenthesized definition that an older replica would reject. - `AlterCommand::isTTLAlter` (whether restating a TTL schedules a `MATERIALIZE TTL` mutation) and the `MODIFY ORDER BY` no-op detection in `AlterCommands::prepare`. - `ProjectionDescription::operator==`. - For `getTreeHash` to be a faithful identity of a definition, AST nodes that keep semantic state outside of `children` now hash that state (`updateTreeHashImpl` overrides): `ASTTTLElement` (mode, destination, `GROUP BY` keys/assignments, recompression codec), `ASTIndexDeclaration` (name, granularity), `ASTConstraintDeclaration` (name, `CHECK`/`ASSUME`), `ASTProjectionDeclaration` (name), `ASTProjectionSelectQuery` (which clause each child belongs to, previously `SELECT a GROUP BY b` and `SELECT a ORDER BY b` hashed equally), `ASTWindowDefinition` (frame type and boundary kinds). - Since a member that is not a child and is not hashed silently makes two different definitions compare equal, `getTreeHash` documents the requirement, and the classes whose member set is the whole point of the override (`ASTTTLElement`, `ASTIndexDeclaration`, `ASTConstraintDeclaration`, `ASTProjectionDeclaration`, `ASTWindowDefinition`, `ASTSetQuery`, `ASTWithElement`, `ASTWindowListElement`) carry a `sizeof` `static_assert`, so adding a member there fails to compile until it is considered. `ASTSelectQuery`, `ASTProjectionSelectQuery` and `ASTOrderByElement` instead iterate the clause/child-role enumerators, so a newly added one is hashed with no code change. The remaining overrides are not asserted: `ASTColumnsApplyTransformer` (which reaches the `parameters` and `lambda` subtrees) and `ASTSelectWithUnionQuery` hash more than one member, `ASTWithAlias` and `ASTSelectIntersectExceptQuery` one each. Measured: the `static_assert` fires for a new `String`, `UInt64` or pointer member; a lone `bool` can still fit in tail padding, which the comment covers. The negative tests found five ways the comparison was too permissive, each of which master rejects. They are fixed here and each is covered by a test that fails if the fix is reverted: - `ColumnDescription::operator==` used `IDataType::equals`, which ignores the `SimpleAggregateFunction` wrapper (as `MergeTreeData::sortingKeyChanged` already documents), so a replica declaring a plain `UInt64` joined a table whose column is `SimpleAggregateFunction(sum, UInt64)`. It compares `getName` now. - `hasSameExplicitStatistics` kept only the `StatisticsType`, so `STATISTICS(tdigest(1))` and `STATISTICS(tdigest(2))` compared equal although the parameters are a part of the stored definition and survive `SHOW CREATE`. The declaration ASTs are compared as well. - `parameters` and `lambda` of `ASTColumnsApplyTransformer` are not children and were hashed with `updateTreeHashImpl`, which stops at the node itself, so a projection using `APPLY quantile(0.5)` and one using `APPLY quantile(0.9)` hashed equally. `stripArtificialParens` could not reach them either. Both now descend. - The alias was hashed without a length prefix, so `fooIdentifier_bar` and `bar AS Identifier_foo` produced the same byte stream and two projections with different output columns compared equal. - `ASTWindowDefinition` keeps its frame type and boundary kinds outside `children` and did not hash them, so a projection aggregating over `ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW` compared equal to one over `ROWS BETWEEN CURRENT ROW AND CURRENT ROW`. The frame offsets are children and were already covered. One latent bug became reachable and is fixed too: `ASTTTLElement::clone` left `recompression_codec` shared with the source, and `formatDefinition` clones the metadata AST specifically in order to canonicalize it, so `stripArtificialParens` would have written into the live metadata snapshot of a `RECOMPRESS` TTL. No parser or formatter changes and no on-disk/`SHOW CREATE` changes: what the user wrote is preserved; only comparisons became insensitive to formatting. Genuinely different definitions still differ (`a` vs `b`, `a` vs `a DESC`, different index granularity, `CHECK` vs `ASSUME`, different TTL destination). The comparison of whole `CREATE` queries during `RESTORE` (`compareRestoredTableDef`) still compares text: it is a whole-query comparison rather than an expression comparison, and its failure direction is safe (refuses the restore). Tested with a real previous-version server (`clickhouse/clickhouse-server:26.4`): `tests/integration/test_backward_compatibility/test_parenthesized_keys.py` exercises a new replica joining an old table and vice versa, an old replica applying parenthesized `ALTER` log entries written by the new version plus restarts of both replicas, and upgrade + `ATTACH PARTITION FROM`. Unit test `src/Storages/MergeTree/tests/gtest_replicated_metadata_compare.cpp` locks the comparison semantics including the negative cases. Stateless tests: `03471_replace_partition_tuple_key_normalization`, `04604_parenthesized_partition_key_attach_from`, `04612_parenthesized_index_projection_attach_from`, `04622_modify_ttl_parenthesized_no_mutation`, `04646_projection_ttl_parenthesized_zk_metadata`, `04648_alter_parenthesized_zk_metadata`, `04650_replica_column_definition_mismatch_rejected`, `04693_projection_apply_function_name_case` (the function name of a projection's `COLUMNS(...) APPLY` transformer is canonicalized for the comparison only, so `APPLY SUM` and `APPLY sum` compare equal while the stored definition keeps the as-written spelling that older replicas compare byte-for-byte).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110833",
          "createdAt": "2026-07-17T10:35:00Z",
          "updatedAt": "2026-08-13T09:01:25Z",
          "timestamp": "2026-08-13T09:01:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 71
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.6-must-backport"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alesapin",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:668c8c4e7f08d7cfed1a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108468",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108468",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `merge_use_batch_sorting_queue` `MergeTree` setting for ordinary merges",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/38022 Related: https://github.com/ClickHouse/ClickHouse/pull/38859 Adds `merge_use_batch_sorting_queue`, a Boolean `MergeTree` setting for ordinary `MergeTree` merges. It has no effect on merges that change rows, such as those performed by `ReplacingMergeTree` or `AggregatingMergeTree`. The setting defaults to `false`, preserving the existing merge behavior. When set to `true`, it uses the batch sorting queue in the ordinary merge sorting stage. The batch sorting queue is already used in several query sorting paths. This change makes it available for background merges as an opt-in table setting, so users can enable it for merge-heavy workloads where `MergingSortedTransform` is a significant part of merge time. The PR also adds targeted stateless coverage for: - `merge_use_batch_sorting_queue = false` vs `true` correctness across horizontal and vertical merges with a broad set of data types; - vertical merges with small granules and larger merge blocks; - equal primary-key values spread across multiple source parts, checked by merged-part order. Benchmark results from local runs with `merge_use_batch_sorting_queue` disabled and enabled: | Dataset | Wall time speedup | Wall time disabled | Wall time enabled | User CPU change | Peak RSS disabled | Peak RSS enabled | `MergeTotalMilliseconds` change | `MergingSortedMilliseconds` change | |---|---:|---:|---:|---:|---:|---:|---:|---:| | ClickBench hits | 23.9% | 10.270s | 7.813s | -28.3% | 1.38 GiB | 1.36 GiB | -27.1% | -69.2% | | OnTime 2019 | 35.7% | 6.855s | 4.410s | -41.4% | 1.79 GiB | 1.83 GiB | -31.5% | -91.4% | | StackOverflow votes | 13.0% | 0.730s | 0.635s | -17.4% | 374 MiB | 372 MiB | -14.4% | -82.2% | | HackerNews | 2.8% | 3.565s | 3.465s | -3.1% | 798 MiB | 803 MiB | -0.9% | -74.8% | | NYC Taxi | 2.4% | 3.505s | 3.420s | -1.9% | 998 MiB | 1000 MiB | -2.2% | -5.2% | | StackOverflow posts | 2.7% | 4.275s | 4.160s | -2.0% | 1.07 GiB | 1.08 GiB | -2.3% | -32.8% | The benefit is workload dependent. It is largest when sorted merging is a significant fraction of the total merge time, and smaller when writing, compression, or other merge work dominates. This PR was developed with AI assistance from Codex and Claude. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `merge_use_batch_sorting_queue` `MergeTree` setting to optionally use the batch sorting queue for ordinary `MergeTree` merges, reducing CPU and wall-clock time for merge workloads where sorted merging is a significant cost.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108468",
          "createdAt": "2026-06-25T12:18:51Z",
          "updatedAt": "2026-08-13T09:00:02Z",
          "timestamp": "2026-08-13T09:00:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-performance",
            "manual approve",
            "can be tested"
          ],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2c78753844a55a26b455",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114089",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114089",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `bloom_filter` index skipping granules for a `FixedString` constant",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112693 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `hasAny`, `hasAll`, `has`, `indexOf`, `mapContainsKey`, `mapContainsValue` and `mapContains` returning too few rows, or failing with `TOO_LARGE_STRING_SIZE`, when a `bloom_filter` index is queried with a `FixedString` constant. The index hashed the padded form of the constant while the function compares the unpadded one, so a matching granule was skipped. ### Description A `bloom_filter` skip index silently returns too few rows when the constant is a `FixedString`, and in four cases raises `TOO_LARGE_STRING_SIZE` on a query that succeeds without the index: ```sql CREATE TABLE k (id UInt64, v Array(String), INDEX idx v TYPE bloom_filter GRANULARITY 1) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 1; INSERT INTO k VALUES (0,['V0']),(1,['V0\\0']),(2,['X']); SELECT count() FROM k WHERE hasAny(v, [toFixedString('V0',3)]); -- 0, wrong SELECT count() FROM k WHERE hasAny(v, [toFixedString('V0',3)]) SETTINGS use_skip_indexes=0; -- 1 ``` Root cause: the array-search functions coerce the constant with CAST before comparing it to the elements (`hasAny`/`hasAll`, and `has`/`indexOf` over a `FixedString` element, cast both sides to the least supertype; `has`/`indexOf` over a `LowCardinality` element cast the constant straight to the dictionary type), and a CAST of `FixedString` to `String` strips the trailing zero padding. The index instead converted the constant with `convertFieldToType`, which keeps the padding, so it hashed a value the function never compares and pruned granules that match. The error cases are the same cause: an oversized `Field` passed through unchanged and reached a column `insert` during index analysis. The fix replicates the coercion at the `Field` level, which is exact because the string casts involved are fully predictable: strip the padding of a `FixedString` constant, re-pad it to the width of the element type (the stored form of every element the function can match), and decline the index when the constant has no stored representation, so a runtime `TOO_LARGE_STRING_SIZE` stays reachable instead of becoming a silent empty result. The direct cast to a dictionary type rejects an over-wide `FixedString` constant before stripping, while the supertype cast strips first - that one ordering difference is the only divergence between the two coercions, so it is a `bool` rather than a second code path. `has`/`indexOf` over a plain `String` element compare the constant's raw padded bytes (`executeString`) and keep the old conversion, as does `has(<constant array>, <indexed scalar>)`, whose runtime compares `Field`s directly. #### `Map` predicates `mapContainsKey`, `mapContainsValue`, `mapContains` and `has` over a `Map` are adapters of the same `arrayIndex.h` machinery, so a `bloom_filter` index on `mapKeys`/`mapValues` needs the same coercion - and it needs to pick the mode from the `Map` type rather than from the index header, because the `mapKeys`/`mapValues` index expression strips `LowCardinality` from the key/value type while the `mapContains*` adapters run over the keys/values subcolumn, which keeps it. Reading the mode from the stripped header alone would still prune a matching granule on `Map(LowCardinality(String), ...)`, which is broken on `master` today independently of the constant's width. `has` over a `Map` is the one exception, and it goes the other way: `FunctionArrayIndex::executeMap` rewrites the map to an array of its keys and calls `recursiveRemoveLowCardinality` on both arguments before `executeArrayImpl`, so it compares the raw padded bytes exactly like `has` over an `Array(String)`, unlike `mapContainsKey` on the same column. Its element type is therefore stripped of `LowCardinality` before the mode is read. The direct array spellings `has(mapKeys(m), ...)` / `indexOf(mapValues(m), ...)` need no special handling: `mapKeys`/`mapValues` return a full `Array` with `LowCardinality` stripped from the type, so the index header type is already the type the runtime compares against. This supersedes https://github.com/ClickHouse/ClickHouse/pull/112693 with a smaller implementation of the same semantics (one `Field`-level helper instead of a two-mode column-cast round-trip through `getLeastSupertype` plus a batched clone of `createColumnFromConstantArray`). The array test - 115 assertions over the full element-type/function/constant-width matrix, including granule-pruning and `TOO_LARGE_STRING_SIZE` reachability assertions - is taken unchanged from that pull request and passes byte-identically, so the two implementations are behaviorally equivalent on the covered surface. Three further tests cover the `Map` predicates, the direct `mapKeys`/`mapValues` array spellings, and `has` over a `Map` with `LowCardinality` keys; each compares the indexed answer against an unindexed oracle.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114089",
          "createdAt": "2026-08-10T00:59:11Z",
          "updatedAt": "2026-08-13T08:57:11Z",
          "timestamp": "2026-08-13T08:57:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:129d217130c17cb443ae",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108017",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108017",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add explicit DB::NsyncSharedMutex and use it for KeeperLogStore changelog locking",
          "text": "### Description We would like to try [google/nsync](https://github.com/google/nsync) mutex for ClickHouse. We used the existing microbenchmarks and did a comparison: - nsync: https://pastila.nl/?01726aec/f309fa62466e02f4514ab51b0166be86#jxPxN6DPD2sau29I1XHTmw==GCM - current: https://pastila.nl/?000ffb86/366fc7712e4de4898668d9c7ebee6d9b#nM+GdsaPsVmmynDDP8vVIQ==GCM **Readers-only: current SharedMutex wins** | Threads | `Self` current | `Nsync` | Winner | |---:|---:|---:|---| | 32 | 66.7755 ns | 160.620 ns | `Self` | | 64 | 69.9920 ns | 151.107 ns | `Self` | | 128 | 67.0653 ns | 150.980 ns | `Self` | So `NsyncSharedMutex` is worse for pure read-heavy workloads. That confirms we should not replace DB::SharedMutex globally. **Writers-only: NsyncSharedMutex wins strongly** | Threads | `Self` current | `Nsync` | Winner | |---:|---:|---:|---| | 32 | 466.074 ns | 56.4393 ns | `Nsync` | | 64 | 461.444 ns | 56.1529 ns | `Nsync` | | 128 | 460.664 ns | 57.4841 ns | `Nsync` | `Nsync` is about **8x faster** for contended writers. **Mixed read/write: NsyncSharedMutex wins strongly** | Threads | `Self` current | `Nsync` | Winner | |---:|---:|---:|---| | 32 | 268.006 ns | 97.0517 ns | `Nsync` | | 64 | 291.993 ns | 92.9774 ns | `Nsync` | | 128 | 325.001 ns | 94.2072 ns | `Nsync` | `Nsync` is about **3.45x faster** for mixed read/write contention. Since `KeeperLogStore::changelog_lock` has a mixed shared/exclusive access pattern, we selected it as the candidate to try. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Use google/nsync for KeeperLogStore changelog locking",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108017",
          "createdAt": "2026-06-20T05:22:29Z",
          "updatedAt": "2026-08-13T08:42:57Z",
          "timestamp": "2026-08-13T08:42:57Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "submodule changed",
            "can be tested"
          ],
          "author": "chhetripradeep",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b0a1843642a6527ca03d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104435",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104435",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Geoparquet rowgroup pruning",
          "text": "Part of making ClickHouse fastest spatial analytical engine on Earth https://github.com/bacek/chgeos/blob/main/BENCHMARK.md ;) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Adds GeoParquet spatial pruning at row group and page levels to the Parquet reader. Row group pruning skips entire row groups whose bounding box doesn't overlap the query geometry. Page-level pruning uses the `covering.bbox` column index to skip irrelevant pages within row groups. Also generalizes spatial predicate pushdown through `IFunctionBase::isSpatialPredicate()` and adds a `GeoFilter` that evaluates spatial predicates during Parquet row reading. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes Parquet read-time pruning logic by adding spatial predicate extraction and bbox-based skipping at row-group and page granularity, which can impact query correctness if predicate/bbox handling is wrong. Also adjusts Iceberg manifest min/max serialization/pruning behavior and adds new settings/events that need validation across varied Parquet/GeoParquet metadata. > > **Overview** > Adds **GeoParquet spatial filter pushdown** to the Parquet reader, gated by new `input_format_parquet_spatial_filter_push_down`, to skip row groups (and optionally pages) when WHERE contains *conjunctive-only* spatial predicates whose constant-geometry bbox is disjoint from bbox statistics (via `covering.bbox` columns or `geospatial_statistics.bbox`). This introduces a new `Parquet::GeoFilter` extractor/bbox utilities, injects covering bbox columns for stats-only evaluation, wires spatial `KeyCondition`s into row-group/page pruning, and tracks pruned pages via new `ProfileEvents::ParquetPrunedPages`. > > Generalizes spatial predicate identification by adding `IFunction*::isSpatialPredicate()` (propagated through adaptors), marking built-in spatial predicates and enabling WebAssembly UDFs to opt in via `is_spatial_predicate` setting; also improves GeoParquet metadata parsing to capture `covering.bbox` column paths and exposes this flag in `FunctionNode` dumps. > > Extends Iceberg min/max handling to support Float32/Float64 stats correctly and to *filter out only non-serializable columns* instead of disabling bounds entirely, and adds Iceberg manifest pruning based on bbox column bounds for spatial predicates. New integration/stateless tests cover spatial row-group/page pruning, OR-safety, and Float32 stats round-tripping. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 135fb227c7e9a0b9b671b483ae561d0c83e31365. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104435",
          "createdAt": "2026-05-08T22:45:37Z",
          "updatedAt": "2026-08-13T11:24:21Z",
          "timestamp": "2026-08-13T11:24:21Z",
          "metrics": {
            "reactions": 1,
            "comments": 11
          },
          "labels": [
            "pr-performance",
            "can be tested",
            "pr-synced-to-cloud",
            "comp-datalake"
          ],
          "author": "bacek",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:876c8e5fa7b87f44decc",
        "signalId": "github:ClickHouse/ClickHouse:issue:114487",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114487",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Malformed Iceberg schema-evolution metadata aborts the server via LOGICAL_ERROR instead of a normal exception",
          "text": "🕵️ ## Describe what's wrong Reading an Iceberg table whose metadata describes an invalid schema evolution (per the Iceberg spec) crashes the server instead of raising a normal, catchable exception: ``` Thread ... received signal SIGABRT, Aborted. ... #11 DB::Iceberg::IcebergSchemaProcessor::getSchemaTransformationDag (..., old_id=0, new_id=1) at SchemaProcessor.cpp:729 ``` `getSchemaTransformationDag` correctly detects the spec violation (\"a required column can't be added during schema evolution without a default\") but reports it via `ErrorCodes::LOGICAL_ERROR`, which `Exception::handleErrorCode` treats as an assertion failure and **aborts the process** in debug and sanitizer builds (`Common/Exception.cpp`, comment: \"In debug builds and builds with sanitizers, treat LOGICAL_ERROR as an assertion failure\"). In a plain release build without assertions it would merely surface as an oddly-labeled `Code: 49` exception - still wrong, but not fatal. This is reachable purely through metadata content - no ALTER, no live write path - so any corrupted or adversarially-crafted `.metadata.json` (or a bug in a third-party writer) that describes an invalid evolution takes the server down on the next read, in any debug/sanitizer build (which includes ClickHouse's own CI debug/sanitizer stateless-test configurations). Found via a lake-corruption fuzzer that feeds semantically-mutated Iceberg metadata to `clickhouse local` - this was, by a wide margin, the single most frequent abort signature across two independent 15-minute fuzzing sessions (22/635 and 46/1017 iterations respectively landed on this exact frame), yet distinct from the already-filed #114350 (empty-column-name `chassert`) - different function, different throw site, different trigger condition. ## Root cause `SchemaProcessor.cpp`, inside `IcebergSchemaProcessor::getSchemaTransformationDag`, has (at least) two throw sites that validate externally-supplied schema content but use `ErrorCodes::LOGICAL_ERROR` - a code reserved for \"this should never happen, it's our own bug\", not \"the input data is invalid\": ```cpp // line ~697 - old field was a struct/list/map, new field is a primitive if (old_json->isObject(f_type) && !field->isObject(f_type)) { throw Exception( ErrorCodes::LOGICAL_ERROR, \"Can't cast primitive type to the complex type, field id is {}, old schema id is {}, new schema id is {}\", id, old_id, new_id); } ``` ```cpp // line ~727 - a new column with no counterpart in the old schema is `required` (no default) if (!type->isNullable() && !field->isObject(f_type)) { throw Exception( ErrorCodes::LOGICAL_ERROR, \"Cannot add a column with id {} with required values to the table during schema evolution. \" \"This is forbidden by Iceberg format specification. Old schema id is {}, new schema id is {}\", id, old_id, new_id); } ``` Both are genuine, correct spec-violation detections - the bug is only the error code choice. Notably, the *sibling* function in the same file, `getSchemaTransformationDagByIds` (a few lines down), validates its own external input (an unknown schema-id) with `ErrorCodes::BAD_ARGUMENTS` instead - so these two throw sites are also inconsistent with the file's own established convention for \"the metadata says something invalid.\" ## Does it reproduce on the most recent release? Reproduced on `master`, `clickhouse-master/BUILD/bin/clickhouse` built 2026-08-12 (debug build, assertions enabled). ## How to reproduce No lake infrastructure needed beyond a normal `IcebergLocal` table plus one JSON edit: ```bash BIN=/path/to/clickhouse rm -rf lake \"$BIN\" local -q \" SET allow_experimental_insert_into_iceberg = 1; SET allow_insert_into_iceberg = 1; CREATE TABLE t (id Int64, s String) ENGINE = IcebergLocal('lake/', 'Parquet'); INSERT INTO t SELECT number, toString(number) FROM numbers(5); \" # Add an evolved schema (id 1) with a new *required* field that has no counterpart in schema 0 # (the schema the existing data file was written under), and make it current - a real writer # would never emit this (it violates the Iceberg spec), simulating corrupted/adversarial metadata. python3 - lake <<'PY' import json, copy, sys, glob path = sorted(glob.glob(f\"{sys.argv[1]}/metadata/v*.metadata.json\"))[-1] with open(path) as f: doc = json.load(f) evolved = copy.deepcopy(doc[\"schemas\"][0]) evolved[\"schema-id\"] = 1 evolved[\"fields\"].append({\"id\": 3, \"name\": \"extra\", \"required\": True, \"type\": \"long\"}) doc[\"schemas\"].append(evolved) doc[\"current-schema-id\"] = 1 with open(path, \"w\") as f: json.dump(doc, f) PY \"$BIN\" local -q \"SELECT * FROM icebergLocal('lake/') ORDER BY id\" # Aborted (core dumped) ``` Full self-contained script: `tmp/repro_schema_le/repro.sh`. ## Expected behaviour A normal, catchable exception, matching how the sibling `getSchemaTransformationDagByIds` already handles \"the metadata references an unknown schema-id\" a few lines below. Should not abort in any build type. ## Suggested fix Trivial - swap the error code at both throw sites in `getSchemaTransformationDag` (`SchemaProcessor.cpp:~697` and `~727`) from `ErrorCodes::LOGICAL_ERROR` to `ErrorCodes::ICEBERG_SPECIFICATION_VIOLATION`. That code (743) is *already declared* as `extern const int ICEBERG_SPECIFICATION_VIOLATION;` at the top of this exact file (`SchemaProcessor.cpp:50`) - it's simply never used by these two sites. No new include, no new declaration, no signature change. This is also already the established convention for this exact class of check elsewhere in the same directory: - `MetadataGenerator.cpp:452`: `throw Exception(ErrorCodes::BAD_ARGUMENTS, \"Iceberg spec doesn't allow to add non-nullable columns\")` - near-identical wording, on the write path. - `Utils.cpp:1188`: picks `BAD_ARGUMENTS` vs `ICEBERG_SPECIFICATION_VIOLATION` depending on `format_version` for the same category of spec check. Only remaining work is a regression test (no existing test references either throw message) - same shape as `04846_iceberg_null_current_snapshot_id.sh` from #114465: rewrite metadata to trigger it, assert a normal exception instead of a crash. Not yet compiled/tested. ## Additional context Distinct from #114350 (`ASTIdentifier` `chassert` on an empty field `name`, a different function entirely) despite both being \"malformed Iceberg metadata aborts the server via an internal-invariant-style check.\" Given how much more frequently this one surfaced in fuzzing than the already-filed issue, worth treating as its own report rather than folding in. Also checked PR #114465 (\"Fix reads of a null `current-snapshot-id`\") on the chance it was a fix for this - it isn't: different function (`getHistory`/`expireSnapshots`/`mutate`), different root cause (`Poco::JSON::Object::has` returning true for a JSON-null value), no overlap with `getSchemaTransformationDag`. Searched GitHub (issues and PRs, open and closed, plus a code search for `getSchemaTransformationDag`) and checked the file's recent commit history - no existing issue or fix for this bug as of 2026-08-12.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114487",
          "createdAt": "2026-08-12T13:30:48Z",
          "updatedAt": "2026-08-13T08:40:40Z",
          "timestamp": "2026-08-13T08:40:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "comp-datalake"
          ],
          "author": "PedroTadim",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d4ab37275c03e497837c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114546",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114546",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrapped `Time64` values from an overflowing scale conversion in `convertFieldToType`",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/94537 Related: https://github.com/ClickHouse/ClickHouse/pull/111534 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes incorrect, often sign-flipped, `Time64` literals and `IN`-list constants produced when rescaling a lower-scale `Decimal64` overflowed `Int64`. Such a conversion now reports `DECIMAL_OVERFLOW`, matching the `DateTime64` branch and explicit `CAST`. ### Description `convertFieldToTypeImpl` rescales a `Decimal64`-backed field into a `Time64` column at `src/Interpreters/convertFieldToType.cpp:473-476`. The scale-increasing arm multiplied without a range check, so `value * scale_multiplier_diff` could exceed `Int64` and wrap. The wrapped product was then handed to `decimalFromComponentsWithMultiplier<Time64>(value, 0, 1)`, whose own `mulOverflow` check is a no-op at multiplier `1`, so the corrupted value became the field. Observed on master (`bb1bd307`): `Values 'x Time64(6)' (253402207200000::Decimal64(0))` returned `-999:59:59.722624`, and `INSERT` persisted it. The wrapped `Int64` of `253402207200000 * 1000000` is `-4852209831933722624`, which renders as exactly that; the negated input wraps to `+4852209831933722624`, so a negative input returned a positive time. Explicit `CAST(... AS Time64(6))` already reported `DECIMAL_OVERFLOW` here, so the literal path disagreed with `CAST`. The file carried this same statement twice, for `DateTime64` and `Time64`, both unguarded. The related PR above guarded the `DateTime64` one and added test `03797`; its `Time64` twin was left as it was. This change mirrors that guard onto the twin. The operand also becomes `Int64`, which the guard requires: `mulOverflow` on an unsigned operand reports overflow for every negative value, which would reject in-range negative times. Reporting rather than returning Null matches the `DateTime64` twin, which `03797` asserts, and explicit `CAST`. The neighbouring `Date32` branches keep their Null contract and are untouched. Found by a UBSan signed-overflow report on this line; there is no open issue for it. It keeps reproducing on `master`, in both `asan_ubsan` stress jobs, as `signed integer overflow: 253402239600000 * 1000000 cannot be represented in type 'long'` at `src/Interpreters/convertFieldToType.cpp:475`: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=86363e819f9cb704a892063da9998e4c1281eb76&name_0=MasterCI&name_1=Stress%20test%20%28amd_asan_ubsan%29 - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=895af217db902c1458fecc608c6f6b3a78cc985e&name_0=MasterCI&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Only inputs that were producing wrong values change: `9223372036854` at scale 6, the largest whose rescale still fits, is still accepted and returns `999:59:59.000000` identically. Verified by building both arms and diffing: every overflow arm goes from a wrong value to `DECIMAL_OVERFLOW`, while 30 in-range controls across scales 0/3/6/9, both signs and both scale directions are byte-identical. `03797` and `04837` still pass. New test `04883` fails on master and is green here over 100 randomized runs.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114546",
          "createdAt": "2026-08-12T21:31:34Z",
          "updatedAt": "2026-08-13T08:38:05Z",
          "timestamp": "2026-08-13T08:38:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:53a8e6c109fc5131af4d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114171",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114171",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Hash the AST members that `getTreeHash` did not see",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110833 This is the `getTreeHash` part of #110833, split off because it stands on its own: it fixes wrong results that have nothing to do with that pull request's motivation (comparing stored table definitions). Only the `src/Parsers` changes were taken; the `sameAST` helper and the metadata comparison stay in #110833. `IAST::getTreeHash` identifies an AST subtree, and several callers use it as the identity of an expression or of a query: `Context::executeTableFunction` caches the result of a table function under it, `ExecuteScalarSubqueriesVisitor` caches the value of a scalar subquery, `ActionsVisitor` keys prepared sets, `ComparisonGraph` and `WhereConstraintsOptimizer` decide whether two expressions are the same one. The default implementation hashes only `getID` and `children`, so anything a node keeps outside `children` is invisible to it, and a number of nodes keep meaningful state there. Two ASTs that mean different things then get the same hash, and a caller that treats the hash as an identity silently substitutes one for the other. For example, on `master`: ```sql SELECT (SELECT count() FROM view(SELECT 1 AS x UNION ALL SELECT 1)) AS union_all, (SELECT count() FROM view(SELECT 1 AS x UNION DISTINCT SELECT 1)) AS union_distinct ``` ``` ┌─union_all─┬─union_distinct─┐ 1. │ 2 │ 2 │ └───────────┴────────────────┘ ``` The second `view` returns `2` instead of `1`: `union_mode` is not a child of `ASTSelectWithUnionQuery` and was not hashed, so the two calls got the same cache key. The same happens for `INTERSECT` against `EXCEPT`, for `LIMIT 2` against `OFFSET 2` (the same literal in a different role), for `WITH FILL ... TO 5` against `WITH FILL ... STEP 5`, for a window frame bounded from below against one bounded from above, and for `APPLY(quantile(0.1))` against `APPLY(quantile(0.9))`. This change hashes the members that are part of an element's meaning and are not children: * `ASTSelectQuery`, `ASTProjectionSelectQuery`, `ASTOrderByElement` - the role each child plays (`WHERE` against `HAVING`, the limit against the offset, the `WITH FILL` upper bound against the step), which is recorded only in `positions`. `ASTSelectQuery::group_by_with_grouping_sets` was missing as well. * `ASTSelectWithUnionQuery` - `union_mode` and `list_of_modes`; `ASTSelectIntersectExceptQuery` - `final_operator`. * `ASTWindowDefinition` - the frame and the parent window name; `ASTWindowListElement` - the name the window is bound to. * `ASTWithElement` - the name the CTE is bound to, `MATERIALIZED`, and the column aliases. * `ASTQueryWithOutput` - the `INTO OUTFILE` modifiers; `ASTQueryWithTableAndOutput` - `TEMPORARY` and an explicit `UUID`. * `ASTSetQuery` - `is_standalone`, the settings reset with `= DEFAULT`, the query parameters, and the node's own identity, which this override skipped altogether. * `ASTTTLElement` - the mode, the destination, the `GROUP BY` key and assignments, and the recompression codec. * `ASTIndexDeclaration`, `ASTConstraintDeclaration`, `ASTProjectionDeclaration` - the declared name (and the index granularity, and `CHECK` against `ASSUME`). * `ASTCollation` - the collation name; `ASTStreamSettings` - the cursor tree and the watermark column and idle timeout. Follow-up fixes in the same area, found by the debug build's format+parse round-trip check and by review: * A per-column `PRIMARY KEY` is normalized by `ParserCreateQuery` into the storage definition, but `primary_key_specifier` stayed `true` on the column declarations while formatting never printed it, so once the flag was hashed, such a `CREATE` no longer round-tripped format+parse to the same tree hash and the debug build reported the `Inconsistent AST formatting` logical error. The parser now clears the flag when its meaning is transferred; where it does survive - `ALTER TABLE ... ADD/MODIFY COLUMN` - formatting now prints `PRIMARY KEY` instead of silently dropping it. * `ParserQueryWithOutput` canonicalizes the output-option children to the end of `children`, and most parsers add the database/table children first, but several `clone` implementations rebuilt them in a different order, so a clone of `CHECK TABLE t FORMAT JSONEachRow` or `SHOW CREATE TABLE t INTO OUTFILE 'x'` hashed differently than the original. The clones now rebuild `children` in the parser's order. Along the way: `ASTDropQuery::clone` and `ASTUndropQuery::clone` did not clear the copied `children` (the clone kept the source's children and appended the cloned ones on top), `ASTOptimizeQuery::clone` pushed `deduplicate_by_columns` into `children` while the parser keeps it member-only, and `ASTCheckTableQuery::partition`, `ASTWatchQuery::limit_length`, and the `where_expression` / `limit_length` of `SHOW COLUMNS` / `SHOW INDEXES` were left shared with the source instead of being cloned. * The `static_assert`s pinning the size of the AST nodes are checked only on 64-bit targets: the wasm32 parser build has a different layout. Two more fixes in the same area: * `ASTColumnsApplyTransformer` reached its non-child `parameters` and `lambda` through `updateTreeHashImpl` rather than `updateTreeHash`, which stops at the node itself and never descends into it, so the `0.5` of `APPLY(quantile(0.5))` was not hashed. * `ASTWithAlias` hashed the alias without its length, so the alias ran into whatever the node writes next and `foo` followed by `Identifier_bar` produced the same byte stream as the alias `bar` on `Identifier_foo`. `ASTCollation::readJSON` and `ASTWithElement::readJSON` put a member into `children` that neither the parser nor `clone` puts there, so an AST restored from JSON had a shape - and a hash - that no parsed AST has, and `clone` of it dropped the child again. They now reproduce the parser's shape. `ASTTTLElement::clone` left the recompression codec shared with the source, which the new unit test surfaced. New tests: `04836_tree_hash_ast_identity` covers the wrong results above, and `gtest_tree_hash_completeness` covers the members that no pair of queries can differ in on their own, by editing the JSON serialization of the AST, plus the JSON round-trip shape. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results for a query that calls the same table function twice with arguments that differ only in a part of the query that was not taken into account by the AST hash, such as `UNION ALL` against `UNION DISTINCT`, `INTERSECT` against `EXCEPT`, `LIMIT` against `OFFSET`, or the bounds of a window frame. The second call reused the result of the first one.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114171",
          "createdAt": "2026-08-10T15:12:26Z",
          "updatedAt": "2026-08-13T08:31:05Z",
          "timestamp": "2026-08-13T08:31:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0ef4daa65af998e80a95",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113895",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113895",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Strip the cosmetic parenthesized flag before comparing stored definitions",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/92340 Related: https://github.com/ClickHouse/ClickHouse/pull/110833 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed comparison of stored table definitions that were written with redundant parentheses (`PARTITION BY (a)`, `PRIMARY KEY (key)`). Since #92340 the formatter preserves those parentheses, so `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` rejected two otherwise identical tables with `Tables have different partition key`, and a `KeeperMap` table created by 26.5 or 26.6 could not be opened by a server of another version. ### Description #92340 started preserving the parentheses a user writes around a definition expression. They are cosmetic, but stored table metadata is compared as text against a form that may have been written by another server version, so two identical definitions began to differ as strings. Two user-visible consequences: - **`ATTACH PARTITION FROM`.** A table declared `PARTITION BY (a)` no longer matches one declared `PARTITION BY a`. This requires neither an upgrade nor a mixed-version cluster: the comparison is between two in-memory ASTs of which only one carries the flag, so both tables created by the same binary already fail. Measured on released binaries, 26.3 and 26.4 accept the pair and 26.5 and later reject it. ```sql CREATE TABLE src (a UInt32, b UInt32) ENGINE=MergeTree PARTITION BY (a) ORDER BY b; CREATE TABLE dst (a UInt32, b UInt32) ENGINE=MergeTree PARTITION BY a ORDER BY b; INSERT INTO src VALUES (1, 1); ALTER TABLE dst ATTACH PARTITION 1 FROM src; -- 26.4: ok -- 26.6: Code: 36. DB::Exception: Tables have different partition key. (BAD_ARGUMENTS) ``` - **`KeeperMap`.** The primary key is serialized into Keeper and compared there against the text written by whichever version created the table. 26.5 and 26.6 write `primary key: (key)`, every other version writes `primary key: key`, so a server of the other version refuses to open the table: ``` Path ... is already used but the stored primary key definition doesn't match. Stored metadata: ... primary key: (key) local metadata: ... primary key: key ``` On a multi-replica setup the replicas that cannot apply the definition never finish startup. ### Implementation `ReplicatedMergeTreeTableMetadata` already stripped the flag, but only on the top level of an expression list, and the helper was private to that file, so the other two comparison sites never got it. This promotes it to `Parsers/stripArtificialParens.h`, makes it walk the whole tree, and applies it in the two places that were missed. The walk also reaches members that are not in `children` and would otherwise be skipped: the `GROUP BY` keys, the `GROUP BY` assignments and the recompression codec of a TTL element, and the `parameters` and `lambda` of a projection's `APPLY` transformer. `StorageKeeperMap` additionally accepts a stored primary key that still carries the parentheses, so tables already created by 26.5 or 26.6 keep working after this change. If that stored text cannot be parsed the comparison stays strict and the mismatch is reported, rather than being silently accepted. Only the comparison changes. There are no parser or formatter changes: what the user wrote is still what is stored and what `SHOW CREATE` and `system.tables` report, and genuinely different definitions still differ (covered by negative cases in both tests). This also fixes two transposed format arguments in the `KeeperMap` mismatch message, which rendered the path and the field name in the wrong order (`Path columns is already used but the stored /keeper_map_tables/... definition doesn't match`). ### Relationship to #110833 This is the compatibility part of #110833, extracted so it can be reviewed and backported on its own, as requested in https://github.com/ClickHouse/ClickHouse/pull/110833#issuecomment-5058455151. The `getTreeHash` work from that PR is a separate, larger change and is not included here. Credit for the original diagnosis and the wider fix goes to @groeneai. ### Tests - `04821_parenthesized_key_attach_partition_from` covers `ATTACH PARTITION FROM` across the two spellings, asserts that `system.tables` still reports the parentheses the user wrote, and that a genuinely different partition key is still rejected. - `04822_keeper_map_parenthesized_primary_key` asserts that the primary key stored in Keeper does not depend on the spelling, that a second table on the same path with the other spelling opens, and that a different primary key is still rejected.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113895",
          "createdAt": "2026-08-07T22:24:44Z",
          "updatedAt": "2026-08-13T08:28:40Z",
          "timestamp": "2026-08-13T08:28:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "v26.6-must-backport"
          ],
          "author": "fm4v",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:192c41f0b6c62a5440eb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114010",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114010",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not create unused aggregate states in aggregation in order with a partial GROUP BY key",
          "text": "Do not create unused aggregate states in aggregation in order with a partial GROUP BY key <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114000 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114000 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a memory leak with `optimize_aggregation_in_order = 1` when the table sorting key is a strict prefix of the `GROUP BY` key and an aggregate function whose state owns heap memory is used, such as `quantileDD`. Server memory grows with every such query until restart. ### Description `AggregatingInOrderTransform` has two output modes. When the sorting prefix covers the whole `GROUP BY` key, each state created by `createStatesAndFillKeyColumnsWithSingleKey` (stored in `variants.without_key`) is handed to a `ColumnAggregateFunction` by `addSingleKeyToAggregateColumns`, which also clears the pointer. In the *partial* key mode (`group_by_key == true`, e.g. table `ORDER BY parent_key` with `GROUP BY parent_key, child_key`) the result comes from the hash table via `prepareChunkAndFillSingleLevel`, which never reads `without_key`, and both `addSingleKeyToAggregateColumns` call sites are guarded by `if (!group_by_key)`. So that state is write-only: the next call overwrites the pointer and orphans the previous state, and the following `variants.invalidate()` sets the type to `EMPTY`, making `destroyAllAggregateStates` return early. Arena bytes are freed; the state's owned allocations are not. The producer was left unconditional when the mode and all its `!group_by_key` guards were added in `3931dbd848786e8` (2022-03-06), so the guard pair has been asymmetric since then. It is invisible for arena-only states, which is why the in-tree test of this plan shape is green with `count()`. This completes the guard pair: the key-column fill is split into `fillKeyColumnsWithSingleKey`, used by the two partial-key call sites, so no state is created where none is consumed. The `!group_by_key` path is unchanged. `overflow_row` cannot be set in this mode, so the key-only path needs no overflow-row state. The test runs its queries through `clickhouse local`, whose at-exit LeakSanitizer check aborts on a leaked state, so the leak is an ordinary test failure on any ASan build instead of something only a stress job sees. It also pins the plan shape and the results; on a non-ASan build only those are checked. Verified with `clickhouse-test`: FAIL on pristine master (1600 bytes leaked in 30 allocations), OK here, 20/20 randomized. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1276` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114010",
          "createdAt": "2026-08-09T05:37:58Z",
          "updatedAt": "2026-08-13T09:29:26Z",
          "timestamp": "2026-08-13T09:29:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "nickitat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5f9561a253d5cdf9d64f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111459",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111459",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Adaptive Aggregator",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> <img width=\"2938\" height=\"1140\" alt=\"image\" src=\"https://github.com/user-attachments/assets/fb27b9b0-958d-42db-88aa-7a1c0dbe3b35\" /> <img width=\"2928\" height=\"1098\" alt=\"image\" src=\"https://github.com/user-attachments/assets/f115fecc-fa83-44fd-ad37-119c8d831b32\" /> <img width=\"2982\" height=\"898\" alt=\"image\" src=\"https://github.com/user-attachments/assets/9f38d0d8-3019-41ce-8e4d-ecfcc5ac0d67\" /> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New adaptive algorithm for parallel `GROUP BY` (controlled via setting `enable_adaptive_aggregator`, enabled by default): each thread aggregates into its own hash table until it holds `adaptive_aggregator_freeze_threshold` keys and then freezes it, so frequent keys keep updating the small cache-resident tables with no coordination, while rare keys are routed by their hash into per-bucket backlogs and aggregated exactly once, inside the bucket-parallel merge.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111459",
          "createdAt": "2026-07-22T18:45:09Z",
          "updatedAt": "2026-08-13T08:25:02Z",
          "timestamp": "2026-08-13T08:25:02Z",
          "metrics": {
            "reactions": 4,
            "comments": 11
          },
          "labels": [
            "pr-performance"
          ],
          "author": "nihalzp",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:213c7169b2d934c03954",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113443",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113443",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a failpoint inside the Paimon incremental-read at-most-once window",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/102343 Paimon incremental reads have documented at-most-once semantics: the Keeper watermark advances at file-collection time, before the collected batch is delivered, so a crash inside that window loses the batch. This window has been untestable — nothing can deterministically crash a server between two statements inside one query's execution. This adds a pauseable failpoint, `paimon_incremental_read_pause_after_watermark_commit`, exactly between the watermark commit and delivery. Integration tests can enable it, observe the committed watermark in Keeper while the read is paused, restart the server, and assert the batch is lost — pinning the at-most-once contract deterministically. The test flips the day delivery becomes at-least-once. The failpoint is inert unless enabled via `SYSTEM ENABLE FAILPOINT`, like every other failpoint. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113443",
          "createdAt": "2026-08-05T08:43:23Z",
          "updatedAt": "2026-08-13T08:24:23Z",
          "timestamp": "2026-08-13T08:24:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-ci"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:12c035466b2601a82a2a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107292",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107292",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add setting to omit CSV quotes for date and time types",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/34668 Adds `output_format_csv_quote_date_time_types`, a backward-compatible CSV output setting for users that need date and time values emitted without surrounding double quotes while keeping existing `CSV` behavior by default. The setting applies to `Date`, `Date32`, `DateTime`, `DateTime64`, `Time`, and `Time64` values in `CSV` output. Strings remain quoted according to the existing `CSV` serializer and numbers are unchanged. CSV input parsing is unchanged and already accepts both quoted and unquoted date and time values. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds the setting `output_format_csv_quote_date_time_types`. When disabled, `Date`, `Date32`, `DateTime`, `DateTime64`, `Time`, and `Time64` values are written without surrounding double quotes in `CSV` output, while preserving the existing quoted behavior by default. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107292",
          "createdAt": "2026-06-11T22:23:38Z",
          "updatedAt": "2026-08-13T08:23:12Z",
          "timestamp": "2026-08-13T08:23:12Z",
          "metrics": {
            "reactions": 1,
            "comments": 10
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "cristhiank",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:07bb8b0ac6c50f80a9e4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114547",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114547",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Horizon Support in ClickHouse",
          "text": "Add support for Snowflake Horizon. It support read and write path. I tested it against my own catalog. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description] Add support for Snowflake Horizon. You can now query Iceberg table in Iceberg behind the horizon catalog. You can also write to the Iceberg table via the catalog.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114547",
          "createdAt": "2026-08-12T21:33:29Z",
          "updatedAt": "2026-08-13T08:19:48Z",
          "timestamp": "2026-08-13T08:19:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-feature"
          ],
          "author": "melvynator",
          "state": "open",
          "assignees": [
            "asya-ch"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:dc91841ec712a8ba12ac",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111992",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111992",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a mutated part losing files to the temporary directory cleanup",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/107150 Related: https://github.com/ClickHouse/ClickHouse/pull/96376 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a race between a `ReplicatedMergeTree` mutation and the removal of old temporary directories that could publish a mutated part with some of its files missing, making reads of that part and all its later mutations fail. ### Description `MutateFromLogEntryTask::finalize` released `mutate_task` - and with it the RAII guard that registers `tmp_mut_<part>` in `MergeTreeData::TemporaryParts` - before calling `MergeTreeData::Transaction::renameParts`, which is what actually renames the directory on disk. Between those two points the temporary directory of an already precommitted part was not protected from `clearOldTemporaryDirectories`. When the cleanup thread hit that window it started removing the directory while `renameParts` was moving it to the persistent name, so the part became active with part of its files already deleted while `checksums.txt` still listed all of them. From the failing job's `node1` log, the three events are 850 µs apart: ``` 19:34:46.048315 Renaming temporary part tmp_mut_97_0_0_0_7 to 97_0_0_0_7 19:34:46.049102 Removing temporary directory .../tmp_mut_97_0_0_0_7/ <- cleanup thread 19:34:46.049168 Renaming part to 97_0_0_0_7 <- the actual rename ``` Reads of such a part fail, because `MergeTreeMarksLoader` derives the marks path from the part's own checksums and therefore expects the file to exist: ``` Code: 1001. DB::Exception: std::filesystem::filesystem_error: filesystem error: in file_size: No such file or directory [\".../97_0_0_0_7/foo2.cmrk2\"] ``` and every following mutation of the part fails permanently, because it hardlinks the source part's files by the names listed in its checksums: ``` Code: 424. DB::ErrnoException: Cannot link .../97_0_0_0_7/foo2.bin to .../tmp_mut_97_0_0_0_8/num2.bin ... (CANNOT_LINK) ``` This is the flaky `test_rename_column/test.py::test_rename_distributed_parallel_insert_and_select`. That test sets `temporary_directories_lifetime = 1`, which makes the window easy to hit; it has failed this way 12 times in the last two months, including twice on `master`. #107150 reported the same symptom from the reader's side and improved the error message without fixing the race. Now the part is renamed while `mutate_task` is still alive, the way `MergeTreeDataMergerMutator::renameMergedTemporaryPart` already does for merges (\"Explicitly rename part while still holding the lock for tmp folder to avoid cleanup\"). `mutate_task` is still reset before `checkPartChecksumsAndCommit`, so the fallback fetch on a checksum mismatch (#96376) still runs with the guards released. Every other caller of `renameTempPartAndReplace` with `rename_in_transaction=true` renames immediately afterwards; `MutateFromLogEntryTask` was the only one that did not. ### Testing `04512_mutation_temp_dir_cleanup_race` pauses a mutation right before the rename with a new pauseable failpoint and checks that the cleanup thread reports the directory as in use instead of removing it. It fails on the unfixed code (`0 1` - the directory is removed, and the mutation then dies with `Cannot set modification time to file`) and passes with the fix (`1 0`), verified locally on a debug build in both directions. Reported by: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=423ec1ba3ac4288ce5b934235be0c6700e0f7c14&name_0=MasterCI&name_1=Integration%20tests%20%28amd_msan%2C%201%2F8%29 ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111992",
          "createdAt": "2026-07-26T23:31:56Z",
          "updatedAt": "2026-08-13T08:14:44Z",
          "timestamp": "2026-08-13T08:14:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:15eb3c44e481e0632ac9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114504",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114504",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Imply make_distributed_plan from distributed_plan_workers_num",
          "text": "Leasing Stateless Workers only makes sense for a distributed query plan, so `distributed_plan_workers_num` no longer has to be paired with `make_distributed_plan`: a non-zero Worker count enables the plan on its own. An explicit `make_distributed_plan` still wins if provided. The implication is applied before the adjustments added in https://github.com/ClickHouse/ClickHouse/pull/112463, so an implied plan still turns off the features it does not support yet. Closes: https://github.com/ClickHouse/ClickHouse/issues/114501 Related: https://github.com/ClickHouse/ClickHouse/pull/112463 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Setting `distributed_plan_workers_num` to a non-zero value now enables `make_distributed_plan` automatically, unless that setting is set explicitly.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114504",
          "createdAt": "2026-08-12T15:12:42Z",
          "updatedAt": "2026-08-13T08:14:16Z",
          "timestamp": "2026-08-13T08:14:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "andreev-io",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:caabc22fca2a2d7c3a40",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114563",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114563",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Order the /play startup test's load-window keystroke by construction",
          "text": "Order the /play startup test's load-window keystroke by construction <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `test_play_reconcile_startup` can fail one assertion of `dirty-startup-merge-entry-reowned` (\"the URL hash carries the live query, not the dropped blank tab stale hash\", actual `?tab=Scratch#U0VMRUNUIDExMQ==`) while the scenario's other five checks pass. No related open issue found. The Web UI is not at fault. The four `dirty-startup-*` scenarios must deliver their \"user typed while IndexedDB was still opening\" keystroke inside the load window, but the harness expressed that ordering as a duration: `await sleep(config.duringLoadDelayMs || 5)`. The fake `IndexedDB.open` registers its `openDelayMs: 30` timer *during* page-script evaluation, while the keystroke timer is registered only *after* evaluation returns, so the keystroke wins only while `eval_end - open_call` stays under the measured 26-28 ms slack. Evaluating the 647 KB page script takes 1-2 ms idle, up to 17 ms under CPU contention. When the keystroke loses, `markBootstrapDirty` short-circuits on `bootstrap_settled`, `bootstrap_dirty` stays false, and reconciliation runs the non-dirty named-URL path, for which the stale URL is the *correct* result. The test was asserting the dirty contract against the non-dirty path. Yielding only to microtasks before the interaction makes the ordering structural: the bootstrap is synchronous, the open can resolve only through a `setTimeout`. Three premise assertions, one of which is that the interaction precedes the fake open's completion, make a future ordering regression fail loudly instead of being misattributed to `play.html`. The now-unreachable `duringLoadDelayMs` key is removed. So the suite can catch a reintroduction of the timing dependency, `dirty-startup-merge-entry-reowned` opens IndexedDB with zero delay; the other three keep the 30 ms open so the timer path stays covered. `test.py` also pins all five `dirty-startup-*` names, as it already did for `run-marker-*`: dropping any of them used to still report \"All scenarios passed\". No new test file, and no assertion weakened. All 84 checks pass, and the fix repairs all four `duringLoad` scenarios. <details> <summary>Validation</summary> Simulating a slow host (burn N ms of synchronous time after `indexedDB.open()` registers its timer, still inside evaluation). The injection is **external**: each arm runs its own tree's scenarios exactly as committed, no knob overridden. | stall | before | after | |---|---|---| | 26 ms | 84 PASS | - | | 30 ms | **3 FAIL** | 84 PASS | | 40 ms | **3 FAIL** | 84 PASS | | 400 ms | - | 84 PASS | The boundary matches the independently measured 26-28 ms slack; after the change the ordering holds at 16x that margin. On the runtime the test actually uses (`clickhouse/mysql-js-client`, node v22.22.0, page over HTTP), stall 40 ms was 3/3 red with the reported signature verbatim and is 3/3 green after. The suite now detects its own regression: reverting just the microtask yield to the old `sleep(5)` fails with `HARNESS ERROR: duringLoad ran after the IndexedDB open completed`, on both runtimes, and at every post-registration stall from 0 to 400 ms (a stall makes both fake timers overdue, the one case the workspace-state premises alone let through). Deleting any one `dirty-startup-*` scenario now fails `test.py`, where each deletion previously reported \"All scenarios passed\". Other arms: 20/20 consecutive unperturbed runs green with an invariant 84 PASS; keystroke forced to 45 ms reproduces before and is green after; under 8 busy CPU loops 3/3 red before, 3/3 green after. Counting evaluations shows the ordering assertion runs in all four scenarios and fires in none of them unperturbed, so it is neither dead nor trigger-happy. All four `duringLoad` carriers hold their premise after the change, and the adjacent timing knobs (`stale-reload-run-race`'s `openDelayMs`, the `wasmInstantiateDelayMs`/`disableWasm` scenarios) are deliberately untouched and stay green. Also checked against the in-flight `play.html` rewrite in PR 114187: both globals the new assertions read survive there and all 17 dirty-startup checks pass against its `play.html`. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1310` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114563",
          "createdAt": "2026-08-13T00:43:38Z",
          "updatedAt": "2026-08-13T08:36:42Z",
          "timestamp": "2026-08-13T08:36:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:314bb57c95f9b4257179",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:97540",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:97540",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Validate IN tuple/subquery column count mismatch in analyzer",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/74442 The analyzer (`resolveFunction`) now validates that the number of elements on the left side of `IN` matches the number of columns in the right-side subquery. Previously, constant folding could optimize away the `IN` expression and silently hide the mismatch: for example, `(1, 1) IN (SELECT 1)` wrapped in a subquery with a constant-false outer `WHERE` succeeded silently. Now it correctly throws `NUMBER_OF_COLUMNS_DOESNT_MATCH` during analysis. A single `Tuple` column on the right side is still accepted, since `FunctionIn` compares the whole left-side value against it; `LowCardinality` and `Nullable` wrappers are unwrapped before the check. Covered by the new test `03380_in_tuple_subquery_column_count`, including the original reproducer from the issue. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Detect column count mismatch between a tuple and a subquery in the `IN` clause during analysis, preventing constant folding from silently hiding the error. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/97540",
          "createdAt": "2026-02-21T00:39:53Z",
          "updatedAt": "2026-08-13T08:11:20Z",
          "timestamp": "2026-08-13T08:11:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 34
          },
          "labels": [
            "pr-bugfix",
            "submodule changed"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:40c4ea4047f47e2f7522",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113289",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113289",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/113003 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a query plan optimization stall when a `WHERE` clause contains a large string constant with many dots and the table carries a `bloom_filter`, `tokenbf_v1`, `ngrambf_v1` or `text` skip index. Matching a filter column name against `JSONAllPaths(...)` index columns enumerated every dot split of the name, which made skip-index condition building quadratic in the constant's length. Closes #113003. ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/113003 **What breaks.** A `SELECT` whose `WHERE` contains a large dotted string constant, over a table with a skip index, spends unbounded time in query plan optimization. The reporter measured 9.2 hours at 100% of one core with `max_execution_time = 300` set, on 26.4.3.37 and 26.7.1.1315. Nothing on that path observes cancellation, so `max_execution_time` fires and `KILL QUERY` is inert, no `QueryStart` row is written, and the handler thread is leaked until restart. Workaround was `SETTINGS use_skip_indexes = 0`. **Root cause.** `tryMatchJSONSubcolumnToIndex` reads a filter column name as `<json_column>.<path>`. Not knowing where the split is, it enumerated every dot split of the name and per split formatted a lookup key and scanned the index columns. The name embeds a folded constant verbatim (`position('a.a.a...', s)`), so its length is user-controlled, giving O(length^2). **The change.** Scan the index columns instead: keep entries shaped `JSONAllPaths(X)` and test each `X` against the name by prefix-plus-dot. Cost no longer depends on the name. Selection is deliberately unchanged (shortest matching `X`, ties to the first entry), which reproduces what walking dot positions left to right did. This is not `bloom_filter`-specific: `tokenbf_v1`, `ngrambf_v1` and `text` reach the same helper, so all four are fixed at once. **Validation.** The reporter's repro goes from 42.7s to 0.16s on a debug build. All 80 existing `EXPLAIN indexes = 1` cases in `04024_json_skip_index_*` pass unchanged; 3 new cases pin the previously untested ambiguous case where two index columns match one name. The new cost test asserts allocated bytes rather than wall clock, and was verified to fail on the pre-fix binary. The cancellation gap is real and separate; it ships as a follow-up.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113289",
          "createdAt": "2026-08-04T09:31:46Z",
          "updatedAt": "2026-08-13T08:03:04Z",
          "timestamp": "2026-08-13T08:03:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3b90fd68c03947af094e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114119",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114119",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113654 to 26.6: Fix non-atomic Keeper Raft state persistence",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113654 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31365209776/job/93382009349)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114119",
          "createdAt": "2026-08-10T07:37:01Z",
          "updatedAt": "2026-08-13T08:02:50Z",
          "timestamp": "2026-08-13T08:02:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "do not test",
            "pr-cherrypick",
            "pr-critical-bugfix"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "antonio2368",
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:dced37271f83831c25c2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114118",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114118",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113654 to 26.5: Fix non-atomic Keeper Raft state persistence",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113654 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31365209776/job/93382009349)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114118",
          "createdAt": "2026-08-10T07:36:27Z",
          "updatedAt": "2026-08-13T08:02:49Z",
          "timestamp": "2026-08-13T08:02:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "do not test",
            "pr-cherrypick",
            "pr-critical-bugfix"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "antonio2368",
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ae302214d7b33b24c533",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114117",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114117",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113654 to 26.3: Fix non-atomic Keeper Raft state persistence",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113654 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31365209776/job/93382009349)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114117",
          "createdAt": "2026-08-10T07:35:52Z",
          "updatedAt": "2026-08-13T08:02:47Z",
          "timestamp": "2026-08-13T08:02:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "do not test",
            "pr-cherrypick",
            "pr-critical-bugfix"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "antonio2368",
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:83fe22844760ac1e51d8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:76867",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:76867",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Minmax indices by default",
          "text": "<!--- Disable AI PR formatting assistant: true --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): MergeTree tables will have `add_minmax_index_for_numeric_columns` by default. Resolves #70605 ### Details Resolves #70605 <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > Changes the default MergeTree storage behavior by enabling per-numeric-column min-max skipping indexes, which can affect ingest throughput, disk usage, and query plans across all newly created tables unless explicitly disabled. > > **Overview** > **Enables implicit min-max skipping indices on numeric columns by default** by flipping the MergeTree setting `add_minmax_index_for_numeric_columns` to `true` and recording the default change in settings history/compatibility notes. > > To avoid regressions in write-heavy/system tables and test baselines, the PR **forces `add_minmax_index_for_numeric_columns = 0`** when auto-building system log table engines, updates CI/stateful dataset setup and many integration/stateless tests to opt out explicitly, and adds new stateless coverage to verify both the new default and the opt-out behavior. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit e7a97c8b8c10055b2026c1269a674b2e2f73decf. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/76867",
          "createdAt": "2025-02-27T10:47:59Z",
          "updatedAt": "2026-08-13T07:58:42Z",
          "timestamp": "2026-08-13T07:58:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 79
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "devcrafter"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5180abf6cedc453a4cff",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114540",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114540",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject a Variant whose ORC branches read back as one type",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/114169 Related: https://github.com/ClickHouse/ClickHouse/pull/110085 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Closes: https://github.com/ClickHouse/ClickHouse/issues/114169 The native ORC writer accepted a `Variant` whose branches map to *different* ORC types but which the reader parses back to the *same* type. The file was written, and reads fine with an explicit structure, but inference on it failed, breaking the round-trip: ``` $ clickhouse local -q \"SELECT NULL::Variant(String, Int128) FORMAT ORC\" | clickhouse local --input-format ORC -q \"SELECT * FROM table\" Code: 50. ORC union type 'uniontype<binary,string>' has branches with identical types (UNKNOWN_TYPE) ``` Root cause: the writer's duplicate-branch guard keyed its dedup map on the ORC type (`child_type->toString()`), while the reader keys on the ClickHouse type each branch parses back to, mapping both ORC `binary` and ORC `string` to `String`. Two different equivalence relations, so the writer's guard passed where the reader's fired. This keys the guard on a new `orcTypeDedupKey` helper, a recursive rendering of the ORC type that folds `BINARY` into `STRING` and sorts nested `UNION` children (a nested union reads back as a `Variant`, which sorts its branches, whereas ORC keeps them positional). `LIST`, `MAP` and `STRUCT` stay positional, matching `Array`, `Map` and `Tuple`. The writer never emits `CHAR`/`VARCHAR`, so that is the only collapse it can produce. The reader was left alone: remapping ORC `binary` would change the inferred type of every existing ORC `binary` column, and `uniontype<binary,string>` has no type to infer *to* anyway. Validated on a debug build: 16 measured carriers (the 6 reported scalar pairings, `FixedString`, and the same collapse through `Array`/`Tuple`/`Map` and nested unions) wrote and then failed inference before, and are rejected at write time after. 7 positive controls still round-trip. Unreleased feature, so nothing to break and no backport. The `.cpp` adds 26 lines of code, within the 50 @ PedroTadim authorized. Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/114169#issuecomment-5265686354",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114540",
          "createdAt": "2026-08-12T20:34:35Z",
          "updatedAt": "2026-08-13T07:57:35Z",
          "timestamp": "2026-08-13T07:57:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6b4bbb44351325101651",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114408",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114408",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not throw a logical error when an ephemeral node is held by someone else",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/113913 Related: https://github.com/ClickHouse/ClickHouse/issues/86434 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a logical error reported as `Ephemeral node ... still exists after ...s, probably it's owned by someone else` when activating a replica whose `is_active`/`active` node in Keeper is still held by another session. This is expected runtime state, so it is now reported as `ABORTED` instead of `LOGICAL_ERROR`, and no longer aborts the server in debug and sanitizer builds. ### Description Related: #113913, #86434 `deleteEphemeralNodeIfContentMatches` finds the znode held by a foreign owner, logs `isn't owned by us. Will wait until it disappears`, waits `3 * session_timeout_ms`, and on timeout threw `LOGICAL_ERROR`. That is expected runtime state. Another server may legitimately hold the node, and the wait can be shorter than the holder's session: Keeper drops a dead session's ephemerals only at expiry, and the timeout negotiated in the handshake is stored in `Coordination::ZooKeeper`'s own `args` copy, never propagated to the `zkutil::ZooKeeper::args` this wait reads. The message names that case and then called it a logical error anyway. Fixed at the shared site: `LOGICAL_ERROR` -> `ABORTED`. A caller-side `catch` cannot fix it, because `LOGICAL_ERROR` aborts from the `Exception` constructor, before unwinding. The wait and control flow are unchanged, so the operation still fails and activation still refuses to proceed. The message now names the foreign owner as the primary explanation instead of a config mismatch or a bug. Follows #113913, which chose `ABORTED` for the sibling condition in `StorageKafka2::assertActive` (direct-read path; this is the activation path). All six call sites already treat a failed activation as retryable and none branches on the code. Both directions verified with a `TestKeeper` gtest, and end-to-end on a live server via `ReplicatedMergeTreeRestartingThread`: before the change the server aborts, after it the error is retried with backoff and the replica stays read-only. Two latent issues nearby are left alone: the negotiated `session_timeout_ms` is never synced back into `zkutil::ZooKeeper::args`, and `StorageKafka2` writes `active_node_identifier` as a bare UUID where the MergeTree restarting thread uses the quoted `generateActiveNodeIdentifier()` form.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114408",
          "createdAt": "2026-08-12T01:15:24Z",
          "updatedAt": "2026-08-13T08:18:50Z",
          "timestamp": "2026-08-13T08:18:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e4fa45684c1e3a2de60e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110552",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110552",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add support for S3Queue mode='exclusive'",
          "text": "This patch adds a third mode to the S3Queue engine: `exclusive` mode turns off all file tracking and synchronization in ZooKeeper. S3 file processing will be tracked only locally via in-memory structures of this ClickHouse server process. This mode is meant to support high-throughput, high-volume ingestion scenarios. In our particular case, we have a minio instance on each of our ClickHouse nodes, with S3Queue pointing to 127.0.0.1. Exclusive mode allows ingesting many terabytes of data per day while avoid file framing boundary and buffering issues. (S3 multipart uploads solve the buffering issue for ingestion clients.) We have been running our ClickHouse cluster with these patches for many years with no issues. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds support for mode='exclusive' in S3Queue engine, for high-throughput and self-hosted scenarios.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110552",
          "createdAt": "2026-07-15T13:21:15Z",
          "updatedAt": "2026-08-13T07:52:49Z",
          "timestamp": "2026-08-13T07:52:49Z",
          "metrics": {
            "reactions": 1,
            "comments": 17
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "ivan-tkatchev",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5cd223b82633c2444cd2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114538",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114538",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject Delta Lake partition column absent from the table schema",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114462 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `LOGICAL_ERROR` when reading a Delta Lake table whose `metaData.partitionColumns` names a column that `metaData.schemaString` does not declare. Such metadata is now rejected with `BAD_ARGUMENTS` on the delta-kernel reader and `INCORRECT_DATA` on the legacy reader, whenever the table or its schema is read from a snapshot, not only when the query carries a predicate. ### Description Requested by @ PedroTadim in #114462. `partitionColumns` and `schemaString` are independent fields of externally supplied metadata and nothing cross-checked them, so a mismatch was reported as `LOGICAL_ERROR`, which aborts on debug and sanitizer builds. Paimon already threw `BAD_ARGUMENTS` here. - delta-kernel reader: validate in `TableSnapshot::initOrUpdateSchemaIfChanged`, where both values first exist, beside the existing empty-schema rejection. Previously an unfiltered `SELECT *` succeeded and only a predicate threw, since `PartitionPruner` is built only when a filter exists. That site keeps `LOGICAL_ERROR`, now a genuine internal invariant. - legacy reader: `LOGICAL_ERROR` -> `INCORRECT_DATA` where an `add` action's `partitionValues` names an undeclared key, in the JSON-log and checkpoint branches. It also validates `partitionColumns` where `metaData` is loaded, in both branches, so a snapshot with no `add` action to resolve is rejected too. That check compares logical names: a parsed legacy schema is keyed by column-mapping physical names, so comparing against it would reject well formed column-mapped tables. Iceberg is unaffected: it resolves partitions by numeric `source_id` field ids, not by name. Change-data-feed reads go through `TableChanges`, which resolves no partition name, and are out of scope. Behaviour change: such a table was readable without a predicate, and on the legacy reader also with no `add` action; it is now rejected for every query that reads it. It is malformed per the Delta protocol. One case is left alone, unchanged from master: on a table with an explicit column list a bare `count()` is answered from transaction-log row counts, which resolve no column name, so it is not rejected, while reading any column from it is. Validated with the new stateless test `04877`: every failure arm fails on a targeted revert of the line it covers, including a well formed column-mapped control that must keep reading. Also 50/50 randomized runs and the `delta_lake` suite.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114538",
          "createdAt": "2026-08-12T19:55:56Z",
          "updatedAt": "2026-08-13T07:50:11Z",
          "timestamp": "2026-08-13T07:50:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:eb259e1c62816f3a037d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112828",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112828",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Recover a NATS JetStream subscription closed by the broker",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/96651 Related: https://github.com/ClickHouse/ClickHouse/pull/103557 Related: https://github.com/ClickHouse/ClickHouse/pull/112464 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a `NATS` table with `nats_stream` set silently consuming nothing after the NATS server is restarted. The `JetStream` subscription is now re-established automatically instead of requiring `DETACH TABLE` and `ATTACH TABLE`. ### Description A `NATS` engine table reading from a `JetStream` stream stops consuming permanently once the NATS server is restarted. Nothing is logged, the connection reports healthy, and only `DETACH TABLE` plus `ATTACH TABLE` or a server restart recovers it. It is also how the flaky `test_nats_restore_failed_connection_without_losses_on_write` fails on master. Root cause: an asynchronous pull subscription renews its pull request only when a message is delivered, and a reconnect resends the `SUB` line without the outstanding pull request, so with nothing in flight when the server goes away the chain never restarts. The server does report this, answering the outstanding request with `409 Server Shutdown`, and the client then closes the subscription. ClickHouse missed it because `isSubscribed` only tests whether the subscription vector is non-empty, so the existing re-subscribe path was gated on a predicate that cannot see a dead subscription. This adds a per-subscription liveness check and consults it in the streaming task, which drops the subscriptions so the existing re-subscribe runs in the same iteration. Only `JetStream` consumers opt in: core NATS subscriptions are already restored by the client, and recovery drops buffered messages core NATS never redelivers. Validated with five new integration tests. Three restart the broker and fail on master, 9 of 9 repeats, before passing after the change; a fourth asserts a healthy consumer never re-subscribes, so it passes either way and exists to bound the cost. The fifth covers a restart of a table reading two subjects, which nothing covered before. Not covered: a broker loss leaving the subscription with no status at all, such as a hard kill, a partition, or a loss coinciding with a re-subscribe. That needs a local fetch timeout, which would also periodically tear down healthy subscriptions. #103557 targets the same defect from a connection-level reconnect counter, but no longer applies to this code and has no integration test. Close whichever you prefer.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112828",
          "createdAt": "2026-07-31T22:25:07Z",
          "updatedAt": "2026-08-13T07:48:35Z",
          "timestamp": "2026-08-13T07:48:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "antaljanosbenjamin",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6de139e7dab6d9cff88d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112484",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112484",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not deserialize a skip index whose on-disk type is stale",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/112213 Related: https://github.com/ClickHouse/ClickHouse/pull/106988 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes reading a stale secondary (skip) index after an `ALTER TABLE ... MODIFY COLUMN` type change whose mutation never ran, for example because `KILL MUTATION` removed it: the granules on disk were written with the old type but decoded with the new one, raising `LOGICAL_ERROR`, requesting multi-exabyte allocations, or silently returning a wrong result. Also fixes a wrong result from an index over an expression whose meaning changes while every stored byte stays identical, so no mutation is created at all: a `MODIFY COLUMN` altering only a `DateTime` timezone, or only a custom type name such as `UInt8` to `Bool`. Such an index is now skipped for the affected part. Closes #112213. ### Description `MODIFY COLUMN` updates table metadata at once and schedules a mutation to rewrite the parts. Until it runs, a part's granules hold bytes written with the OLD type while index analysis decodes them with the NEW one. The read gate `canUseIndex` infers that mismatch from a *pending mutation entry* and fails open when there is none. Six measured ways reach the deserializer anyway, including `KILL MUTATION` having removed the entry (the report) and conversions for which no mutation is ever created, being metadata-only for the column but not for the index. The top-k minmax read had no gate at all. **The fix** asks the durable question instead: do the part's bytes match the type about to decode them? It lands in `IMergeTreeIndex::getDeserializedFormat`, which already receives the part, so one predicate covers every read path. Physical discovery splits off into a virtual `getPhysicalFormat`, leaving `getDeserializedFormat` **non-virtual** so no override bypasses it. Each required column's part-side type, taken from the part's own list rather than the type-erasing interned cache, is compared against the metadata type: only representation-preserving differences pass, plus a `getName()` same-meaning check on the equals-equal path. The walk recurses pairwise through `Array`, `Nullable` and `LowCardinality`; adding or dropping a wrapper is refused. A non-trivial expression index is refused on any difference. Over-firing costs pruning, not correctness. Two `MergeTask` text-index sites share the predicate, so a stale text index is rebuilt during a merge instead of hardlinked forward. **Out of scope:** index *identity* staleness (a changed expression, a name reused after a killed `DROP INDEX`), which no type check detects. <details><summary>Measured symptoms on master, and validation</summary> | target type | observed on master | |---|---| | `Nullable(UInt64)` | `LOGICAL_ERROR: Sizes of nested column and null map ... not equal after deserialization` | | `Nullable(UInt64)` release / `UInt64` / absent-column part | `Code: 241`, 4 / 2 / 4 EiB | | `Array(UInt64)` | `Code: 33`, \"read just 38 of 18005230136\" | | `Int8` to `Enum8` | wrong: prunes a granule the unindexed read rejects (`Code: 691`) | | `DateTime('UTC')` to `DateTime('Asia/Tokyo')`, `INDEX toHour(dt)` | wrong: 0 vs 3 | | `UInt8` to `Bool`, `INDEX toString(v)` | wrong: 0 vs 32 | | `Tuple(x UInt8)` to `Tuple(x Bool)`, `INDEX toString(p.x)` | wrong: 0 vs 32 | `04165_skip_index_stale_type_after_alter.sql`: 29 cases covering the above, whose 21 control fixtures carry 25 `explain ILIKE` granule assertions pinning the over-fire direction (a simple single-column index across the timezone and `Bool` ALTERs, an unchanged subcolumn index, an unchanged `JSON(a DateTime)` column). On the base commit the test **kills the server** with the reported stack (`SerializationNullable.cpp:178` <- `MergeTreeIndexGranuleSet::deserializeBinary` <- `MergeTreeIndexReader::read` <- `filterMarksUsingIndex`); with the fix it passes 50/50. 36 mutations each confirm one line is load-bearing. A 462-test A/B sweep against pure HEAD leaves the failure set unchanged (27 in both arms, all needing infra this sandbox lacks). Perf over 250 parts and 4 indexes is inside noise. No setting or format changes. #110050 overlaps these files and its helper asks the *physical* question, so it should forward to `getPhysicalFormat` on rebase. My #109616 has since merged and independently reached the same physical-vs-usability split, so it needs no forwarding. Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=108096&sha=25f51f1fddd05350d9c5aa1231aa0e7ee6fc676f&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29 </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112484",
          "createdAt": "2026-07-29T20:03:26Z",
          "updatedAt": "2026-08-13T07:48:24Z",
          "timestamp": "2026-08-13T07:48:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "blocker"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1d6c2ac172266787cec9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111597",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111597",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Recover an intact part from an empty columns.txt instead of losing it",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/111373 Related: https://github.com/ClickHouse/ClickHouse/pull/111414 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed data loss where a MergeTree part with an empty (zero-byte) `columns.txt` was detached as broken on load, discarding all of its rows. An empty `columns.txt` is now treated like a missing one: for wide parts the column list is rebuilt from the part's own metadata, matching the existing behavior for an absent file. ### Description `IMergeTreeDataPart::writeMetadata` rewrites `columns.txt` in place (no atomic rename, no fsync). An interrupted rewrite (crash, power loss, `ENOSPC` between truncate and write) can leave a zero-byte `columns.txt` in an already-committed part directory. On load, an absent `columns.txt` already self-heals: for a wide part the else-branch rebuilds the column list from the table metadata and rewrites the file. But an empty `columns.txt` was read anyway, and `NamesAndTypesList::readText` starts with `assertString(\"columns format version: 1\\n\", buf)`, which throws `CANNOT_PARSE_INPUT_ASSERTION_FAILED` on the empty stream. The otherwise-intact part is then detached as broken and every row is lost. An empty file is strictly less broken than a missing one, so bricking the part for it is illogical. This routes an empty `columns.txt` through the same rebuild path as a missing one. Emptiness is decided via `IMergeTreeDataPart::readFile`, which forces `pread`, so a zero-byte file reports `eof` cleanly instead of faulting under a randomized mmap read method. Compact and patch parts still require `columns.txt` (they cannot rebuild it), so their behavior is unchanged. The rebuild reconstructs the persistent virtual columns the part physically carries (`_row_exists`, `_block_number`, `_block_offset`), not only `getAllPhysical()`. Dropping `_row_exists` would silently discard a lightweight-delete mask (deleted rows would reappear); dropping any of them would fail `columns_substreams.txt` validation and detach the part. Column presence during the rebuild is taken from `columns_substreams.txt`, which records exactly the columns physically written, in order, independent of the on-disk serialization. A column's default serialization can enumerate different streams than were stored (a bucketed Map writes `m.buckets_info`, `m.0.keys`, ... instead of `m.keys`, ...), so probing the default serialization would misjudge such a column absent and drop it, turning recovery back into silent data loss and tripping `columns_substreams.txt` validation. Parts predating that file fall back to enumerating each column's own non-ephemeral streams (present only when every such stream exists), matching `MergeTreeDataPartWide::hasColumnFiles`. These gaps also affected the pre-existing missing-`columns.txt` path. Test `04545_empty_columns_txt_not_fatal` covers wide-part shapes: plain, `_block_number`/`_block_offset`, a lightweight-delete `_row_exists` mask, a `Tuple` and a `Map` (including bucketed serialization), a shared-offset Nested sibling added by `ALTER`, and recovery via the legacy stream-enumeration path when `columns_substreams.txt` is also absent. In each, truncating `columns.txt` to zero and reloading keeps the correct rows/values and persists a complete rebuilt file; a missing `columns.txt` still self-heals. Fails on `master`, passes with the fix.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111597",
          "createdAt": "2026-07-23T12:22:17Z",
          "updatedAt": "2026-08-13T07:48:09Z",
          "timestamp": "2026-08-13T07:48:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e7cec0e07a3383e56c13",
        "signalId": "github:ClickHouse/ClickHouse:issue:110893",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:110893",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "RWLockImpl::getLock re-entrant deadlock in CREATE OR REPLACE internal DROP (STID 2043-3c5c)",
          "text": "## Summary `Logical error: RWLockImpl::getLock(): Cannot acquire exclusive lock while RWLock is already locked` fires from the internal cleanup DROP inside `CREATE OR REPLACE TABLE`. This is a fresh manifestation of the re-entrant DDL-lock LOGICAL_ERROR class previously seen in #79413 / #47023. Surfaced by CI: `Stress test (arm_debug)`, STID 2043-3c5c. Report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110620&sha=ee4b5b869d5d3656964166ce79cdc044545bf273&name_0=PR&name_1=Stress%20test%20%28arm_debug%29 CIDB (30d): 1 occurrence, 0 on master. Rare, timing-dependent (needs the stress thread-fuzzer scheduling). ## Stack ``` RWLock.cpp:145 (LOGICAL_ERROR) <- IStorage::lockExclusively (IStorage.cpp:117) <- InterpreterDropQuery::executeToTableImpl (InterpreterDropQuery.cpp:326/354) <- InterpreterCreateQuery::doCreateOrReplaceTable (InterpreterCreateQuery.cpp:2502 / 2525) ``` ## Root cause `doCreateOrReplaceTable` runs its internal cleanup DROP (the post-EXCHANGE drop at InterpreterCreateQuery.cpp:2502 and the catch-block temp-cleanup drop at :2525) on a context built by `make_drop_context` = `Context::createCopy(current_context)`, which inherits the OUTER statement's `query_id`. `RWLockImpl::getLock` keys its fast-path on `query_id` (RWLock.cpp:130-146). When the same `query_id` already owns `IStorage::drop_lock` and a `Write` acquisition is requested, it throws `Cannot acquire exclusive lock while RWLock is already locked` (read->write upgrade / double-write is unsupported). Note that both `lockForShare` (Read) and `lockExclusively` (Write) operate on the SAME `drop_lock` (IStorage.cpp:79 and :115). So the inner DROP re-enters a lock that the outer `CREATE OR REPLACE` statement still holds under the shared `query_id` (e.g. a share-lock on the source/old-target storage held by an in-flight `AS SELECT` fill pipeline that has not yet released when the internal DROP requests `Write`). ## Relation to prior fixes Same LOGICAL_ERROR class as #79413 / #47023. #86751 (merged, Closes #79413) fixed the sibling `PARALLEL WITH` surface by giving each parallel branch a distinct `query_id`. The `CREATE OR REPLACE` internal-DROP path was not covered and still inherits the outer `query_id` via `createCopy`. ## Candidate fix (direction, mirrors #86751) Give the internal DROP context a distinct `query_id` (or an empty one) in the `make_drop_context` lambda (InterpreterCreateQuery.cpp:2352). An empty `query_id` sets `request_has_query_id = false`, which bypasses the RWLock fast-path entirely, so the internal DROP's exclusive acquisition can never re-enter a lock held under the outer statement's `query_id`. The pre-swap `force_drop`/size-guard semantics are unaffected (the drop context still sets `max_table_size_to_drop`/`max_partition_size_to_drop` when bypassing). ## Reproduction status Not yet deterministic locally. On a debug build with the stress thread-fuzzer env I ran: concurrent same-table `CREATE OR REPLACE` (mixed engines), forced catch-path drops via `max_table_size_to_drop=1`, forced fill-throw via `max_memory_usage`, and self-referential `CREATE OR REPLACE t AS SELECT FROM t` (all under 8-12 concurrent workers). All exercised the internal-DROP path (confirmed via error codes 359/241) but none hit the re-entrant window. The holders destruct synchronously on unwind before the internal DROP, so a foreground load does not land in the exact scheduling window; the CI stress thread-fuzzer does. Per our no-speculative-fix policy the fix will be validated against a failing reproducer (fails without, passes with) before a PR; this issue tracks the analysis and candidate fix in the meantime. <!-- ch-version-info:start --> ### Version info - Resolved by: #114420 - Merged into: `26.8.1.1308` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/110893",
          "createdAt": "2026-07-17T16:14:04Z",
          "updatedAt": "2026-08-13T07:35:51Z",
          "timestamp": "2026-08-13T07:35:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [],
          "author": "groeneai",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6bb60cd8aa65271c1987",
        "signalId": "github:ClickHouse/ClickHouse:issue:109974",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:109974",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "[Bitmap] Function subBitmap execution result error",
          "text": "### Company or project name _No response_ ### Describe what's wrong ### Bug description `subBitmap` / `bitmapSubsetOffsetLimit` can return an incorrect subset when the input bitmap is stored in the small representation. In `src/AggregateFunctions/AggregateFunctionGroupBitmapData.h`, `rb_offset_limit()` : ```cpp if (isSmall()) { UInt64 count = 0; UInt64 offset_count = 0; auto it = small.begin(); for (; it != small.end() && offset_count < offset; ++it) ++offset_count; for (; it != small.end() && count < limit; ++it, ++count) r1.add(it->getValue()); return count; } ``` This assumes the `SmallSet ` is **sorted**, but that is not guaranteed. ### Reproduction ```text VM-20-3-centos :) SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res; SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res Query id: 678e3377-baeb-41a7-88c4-5bf601bd7220 ┌─res───┐ 1. │ [1,2] │ └───────┘ ``` excepted: `[3, 4]` ### Does it reproduce on the most recent release? Yes ### How to reproduce ### Reproduction ```text VM-20-3-centos :) SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res; SELECT bitmapToArray(subBitmap(bitmapBuild([5, 4, 1, 2, 3]), 2, 2)) AS res Query id: 678e3377-baeb-41a7-88c4-5bf601bd7220 ┌─res───┐ 1. │ [1,2] │ └───────┘ ``` excepted: `[3, 4]` ### Expected behavior _No response_ ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context _No response_ <!-- ch-version-info:start --> ### Version info - Resolved by: #110072 - Merged into: `26.8.1.1309` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/109974",
          "createdAt": "2026-07-10T10:16:47Z",
          "updatedAt": "2026-08-13T07:34:21Z",
          "timestamp": "2026-08-13T07:34:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "linrrzqqq",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:07c7083cfd1bca1f4c2d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114586",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114586",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #107028 to 26.6: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114586",
          "createdAt": "2026-08-13T03:41:46Z",
          "updatedAt": "2026-08-13T07:33:42Z",
          "timestamp": "2026-08-13T07:33:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "kssenii",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2dc5749b959a54567b03",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110130",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110130",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Automatically choose between the plain and the secure port in clickhouse-client",
          "text": "When neither `port` nor `secure`/`no-secure` is specified, `clickhouse-client` now probes both the default port 9000 and the secure port 9440 concurrently and uses the protocol of the port that answers first. A server that answers on one port only is connected to without waiting out the connect timeout of the other one, and a server that answers on both is connected to over either of them, since both work (TLS wins a tie, which costs no waiting). This makes the following work out of the box: ``` clickhouse-client --host play.clickhouse.com --user play ``` On `play.clickhouse.com` (and many similarly firewalled servers) the plain port is silently dropped rather than refused, so probing the ports sequentially would stall for the whole connect timeout before TLS could even be attempted — hence the concurrent probe. The addresses of a port, in contrast, are attempted one at a time, the next one only after 250 milliseconds without an answer - the \"Connection Attempt Delay\" of RFC 8305 (Happy Eyeballs) - so a hostname that resolves to several reachable backends is not connected to on all of them at once. The connection the probe establishes to the port it chooses is then handed over to the client instead of being discarded. Together, these keep the automatic choice from leaving sessions that never send anything on the server, which it logs as `Client has not sent any data.` and counts against `max_connections`. When the secure port is the one that answered, the client connects with TLS and the interactive banner shows `Connecting to play.clickhouse.com:9440 (secure) ...`. Because the protocol here is chosen rather than requested, a port that turns out to be unusable is not an error: the client falls back to the other one. It matters the most for the secure port, whose common failure is a self-signed or otherwise untrusted certificate that every client not passing `--accept-invalid-certificate` rejects: the plain port is what the client would have connected to if there were no automatic choice at all, so nothing is taken away from the user. Pass `--secure` to require TLS. The fallback works in the other direction too: when the plain port is the one that answered but the connection to it then fails at the protocol level (e.g. a TLS-terminating proxy in front of the plain port), the secure port is tried. Explicit `--port`, `--secure`, `--no-secure` (on the command line, in the configuration file, or in connection credentials), ClickHouse Cloud hostnames (which already default to TLS), and builds without TLS support bypass the detection entirely. Covered by unit tests for the prober and a new integration test `test_client_auto_secure_port` that reproduces the firewalled-server scenario with `iptables` REJECT/DROP rules: it asserts that the DROP case (the `play.clickhouse.com` one) connects quickly instead of waiting out the connect timeout, that an untrusted certificate falls back to the plain port, and that the probe leaves no extra connection behind on the server. The functional tests now say which protocol the run uses (`--no-secure` for a plain run, symmetrically to the `--secure` a secure run already passes), because the test server listens on both ports. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When neither `port` nor `secure` is specified, `clickhouse-client` tries both the default port 9000 and the secure port 9440 concurrently and uses the one that answers first, so `clickhouse-client --host play.clickhouse.com --user play` connects over TLS without `--secure`, even though the plain port of that server is silently dropped rather than refused. If the port that answered turns out to be unusable — a secure port with an untrusted certificate, for example — the client falls back to the other one, because the protocol was not requested explicitly. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110130",
          "createdAt": "2026-07-12T01:17:48Z",
          "updatedAt": "2026-08-13T07:28:40Z",
          "timestamp": "2026-08-13T07:28:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f05acb14c93af1a16040",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110102",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110102",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallelize aggregation-in-order via key-hash reshuffle (aggregation_in_order_shuffle)",
          "text": "## Description `optimize_aggregation_in_order` is not enabled by default because it is often slower than the default hash aggregation. The reason is structural: for a multi-stream read the whole aggregation is funneled through a single `FinishAggregatingInOrderTransform`, so it runs on ~2–6 cores regardless of `max_threads` and is 8–16× slower than the default parallel hash path for high-cardinality `GROUP BY` (it does *less* total CPU work, but cannot parallelize it). This PR adds an experimental setting **`aggregation_in_order_shuffle`** (default `0`) that removes that funnel by repartitioning, reusing the shuffle primitive introduced for the sharded aggregator (`BufferedShardByHashTransform`): ``` read N sorted streams -> scatter each stream by hash(GROUP BY keys) into num_shards (BufferedShardByHashTransform) -> per shard: MergingSortedTransform(N -> 1) -- groups are disjoint by key -> per shard: streaming AggregatingInOrderTransform + FinalizeAggregatedTransform ``` Each shard aggregates a disjoint set of keys, so there is no cross-shard coordination and no single-threaded merge, while each shard still streams out completed key groups (bounded, O(1) memory). It is only used when the `GROUP BY` order is not relied upon downstream (`!memoryBoundMergingWillBeUsed()`, no bucket-order requirement, no `LIMIT` push-down), and the output is therefore not ordered by the keys. ### Making the shuffle deadlock-free and bounded-memory The M per-shard merges share the N scatters, so a naive scatter deadlocks: a slow/exhausted lane blocks the shared scatter from feeding the others, and a per-shard *sorted* merge is a selective consumer. Two fixes in `BufferedShardByHashTransform` (unbounded mode, used only by this path): - **Demand-driven scheduling**: push to every ready lane first, and pull a new input chunk only to feed an output that is ready *and* starving (empty queue). The shared scatter never stalls on a slow lane (its data is just buffered), so there is no cross-lane cycle, and read-ahead — hence memory — is bounded by how far the fastest consumer runs ahead of the slowest. - **Per-output drain on EOF**: when the input is exhausted, finish each output whose queue is already empty, so a sorted merge gets EOF on exhausted inputs instead of waiting forever (this hung deterministically for non-overlapping parts). Verified: 0 deadlocks over 120+ runs across all cardinalities, overlapping and non-overlapping parts, and `max_threads` 8..128; results are byte-identical to the default for order-independent aggregates. ### Results (200M rows, 50M keys, 96 threads; median of 5) | GROUP BY | default | in-order (funnel) | **shuffle** | |---|---|---|---| | `sum`, 50M groups | 870 ms / 23 GB / 42 cores | 5779 ms / 0.8 GB / 4 cores | **1341 ms / 0.83 GB / 20 cores** | | `uniqExact`, 50M groups | 12040 ms / 56 GB | 5744 ms / 1.5 GB | **1049 ms / 2.0 GB / 22 cores** | For high-cardinality `GROUP BY` the shuffle is **4–11× faster than the current aggregation-in-order at the same O(1) memory**, and for `uniqExact` it is faster than the default while using ~28× less memory. ### Known limitation (why it stays off by default) It scatters raw rows, so it regresses for low/medium cardinality (where in-order is already fast and the shuffle gives no benefit, and long key-runs make the scatter buffer more). Gating it to high cardinality (e.g. from the primary-key granule estimate) is a natural follow-up. `EXPLAIN PIPELINE` is also fixed to render the scatter/merge stage of the in-order pipeline (it was previously omitted). ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added an experimental setting `aggregation_in_order_shuffle` that parallelizes `optimize_aggregation_in_order` by repartitioning the sorted input by the hash of the `GROUP BY` keys into independent shards, removing the single-threaded merge bottleneck while keeping the bounded memory of aggregation-in-order. For high-cardinality `GROUP BY` it is several times faster than the ordinary aggregation-in-order. Disabled by default. ### Documentation entry for user-facing changes - [x] Documentation is written (the new setting is documented in `src/Core/Settings.cpp`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110102",
          "createdAt": "2026-07-11T16:31:59Z",
          "updatedAt": "2026-08-13T07:25:50Z",
          "timestamp": "2026-08-13T07:25:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0ca3ed51273460c14068",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113691",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113691",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `theilsU` window state returning noise when the frame's first argument is constant",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/80373 Related: https://github.com/ClickHouse/ClickHouse/pull/93384 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `theilsU` over a window frame returning an arbitrary value instead of 0 when the first argument is constant within the frame. ### Description `TheilsUWindowData::getResult` (the window-optimized state introduced in https://github.com/ClickHouse/ClickHouse/pull/93384) computes the entropy `H(A)` from cached incremental `Σ n·log n` sums. When the first argument is constant within the frame, the true `H(A)` is zero, and the computed value is pure rounding noise from the incremental updates. The code compared it against exact zero, so a tiny positive noise value passed the check, and `1 - H(A|B) / H(A)` then divided noise by noise: in debug builds this tripped the sanity check as the exception `Logical error: 'res < 1.0 + 1e-4'`, and in release builds the function could return an arbitrary value in $[0, 1]$ instead of 0. The exact (non-window) code path recomputes the entropies from the count maps, where a constant column gives `log(1) = 0` exactly, so it is not affected. The fix compares `H(A)` against an error bound proportional to `N · ε · log N` instead of exact zero, and widens the sanity-check tolerance by the same relative amount so that near-threshold frames do not trip it either. Found by the AST fuzzer on an unrelated PR (it hit https://github.com/ClickHouse/ClickHouse/pull/80373 and https://github.com/ClickHouse/ClickHouse/pull/107667): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=80373&sha=6c271049214aa5a94bd9a5f12fb27a9ffa75648f&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113691",
          "createdAt": "2026-08-06T15:21:02Z",
          "updatedAt": "2026-08-13T07:22:12Z",
          "timestamp": "2026-08-13T07:22:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cb3b5dd104865d166b00",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114474",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114474",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113289 to 26.3: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114474",
          "createdAt": "2026-08-12T12:05:35Z",
          "updatedAt": "2026-08-13T07:18:32Z",
          "timestamp": "2026-08-13T07:18:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll3",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:150a21a2c9f3105f9b8f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114473",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114473",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113289 to 25.8: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114473",
          "createdAt": "2026-08-12T12:04:52Z",
          "updatedAt": "2026-08-13T07:17:25Z",
          "timestamp": "2026-08-13T07:17:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll3",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:743e3f41ebb98d9c2c6e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110968",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110968",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do automatic partition pruning for mutations when it's possible",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ReplicatedMergeTree` mutations now automatically prune the affected partitions based on the `WHERE` part of the query when possible. Also, mutation queries now accept multiple values in the `IN PARTITION` clause. Closes: https://github.com/ClickHouse/ClickHouse/issues/98739 Related: https://github.com/ClickHouse/ClickHouse/pull/99933 This supersedes https://github.com/ClickHouse/ClickHouse/pull/99933 by @alesapin, whose commits are preserved in this branch. The original PR became `CONFLICTING` and could not be updated because maintainer edits are disabled on the fork; the branch had also fallen ~4.5 months behind `master` (in particular, the `MutationCommand` refactor to a text-backed lazy AST required reworking how commands carry `partition`/`partitions`/`predicate`). On top of the original feature, this PR: - Merges current `master` and adapts the code to the `MutationCommand` refactor (commands are re-parsed from `ast_text` via `command.ast()`), the two-argument `getInMemoryMetadataPtr` and the three-argument `ActionsDAGWithInversionPushDown`. - Moves the `optimize_mutations_with_partition_pruning` entry to the current settings-history bucket (review thread on `SettingsChangesHistory.cpp`). - Makes the pruning analysis accept every predicate the mutation itself accepts, instead of hiding failures: the analysis follows the same analyzer selection as the mutation execution - the analysis context matches how the commands will actually be interpreted - for an `ALTER` mutation it is derived from the background context exactly like the background mutation workers derive theirs, so session-only analyzer settings cannot make the submit-time analysis and the asynchronous execution diverge (analyzer-only predicate shapes such as qualified column names work, and a predicate the background execution would reject fails fast at submit time; test `04759_mutation_pruning_background_analyzer_mode`), while a lightweight `UPDATE`, which interprets its commands in the foreground, is analyzed with the submitting context, where the session settings and the current database apply (its predicate is not qualified with the default database, unlike `ALTER` commands; covered by the existing test `03100_lwu_43_subquery_from_rmt`), the column list includes the virtual columns (e.g. `_part`, `_partition_id`) and the `ALIAS` / `EPHEMERAL` columns, and the re-parsed predicate's set operations (`UNION` / `INTERSECT` / `EXCEPT`) are normalized exactly as the mutation execution path does. There is no fallback: an analysis error propagates and fails the mutation rather than silently mutating every partition (review). Without the above, `ALTER TABLE t DELETE WHERE _part = '...'` and similar queries would fail since the setting is enabled by default. Covered by the new test `04612_mutation_pruning_exotic_predicates`. - Requires the `block_numbers` version check only for predicate-pruned commands (review thread on `StorageReplicatedMergeTree.cpp`): for explicit `IN PARTITION` the target set is exact and does not depend on the observed partition list, so the `ZBADVERSION` retry loop is not needed. - In the `ZNONODE` recovery of `EphemeralLocksInPartitions`, reports `ZBADVERSION` to the caller after creating missing partition znodes instead of silently refreshing the version, so a concurrently created partition cannot be missed. - `StorageMergeTree::mutate` validates only explicit `IN PARTITION` ids instead of running the full pruning analysis and discarding the result (on the non-replicated path, parts of unaffected partitions are skipped per part by `canSkipMutationCommandForPart`). - Documents the multi-partition `IN PARTITION` syntax and the automatic pruning behavior. - Recomputes the pruned partition set on every `ZBADVERSION` retry, so a new matching partition created by a concurrent insert on the initiating replica cannot escape the mutation (AI review blocker). Covered by the new failpoint-based test `04613_mutation_pruning_new_partition_race`. - Widens the pruned set with ZooKeeper-only partitions by comparing against the partition set the pruning analysis itself iterated (returned by the pruner from its own parts snapshot), not against a separately read local partition list: a partition the pruner analyzed and ruled out is not re-added, while a partition it could not have seen is still widened in, so a same-replica insert racing with the analysis can neither escape the mutation nor drag an unaffected partition back into it (AI review). Covered by the failpoint-based tests `04614_mutation_pruning_local_partition_race` and `04820_mutation_pruning_analyzed_partition_not_widened`. - Scopes the setting description, the documentation and the changelog entry to the `ReplicatedMergeTree` family: on plain `MergeTree`, predicate-based pruning is not wired in, and an explicit `IN PARTITION` clause should be used (AI review). - Rejects the new `alter->partitions` carrier in `StorageSystemWasmModules` (exactly like the single-partition form) and preserves it in `AlterConversions::createLightweightDeleteCommand`, so the multi-partition clause cannot be silently ignored (AI review). One design note kept as in the original: in the `ZNONODE` recovery path, missing partition znodes are created one per multi-op (create + parent version bump), matching the established pattern of the insert path in `allocateBlockNumber`; batching them into a single parent bump was considered but rejected because a partial `ZNODEEXISTS` would fail the whole multi-op and require a more complex retry. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110968",
          "createdAt": "2026-07-19T04:35:37Z",
          "updatedAt": "2026-08-13T07:14:36Z",
          "timestamp": "2026-08-13T07:14:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 36
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c9df75400e33c29f4a4b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114567",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114567",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check the table name length in `DatabaseOverlay`",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113019 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `clickhouse-local` accepting a `CREATE TABLE` whose name is too long for the configured `--default_database`, which produced a table that could not be dropped. `DatabaseOverlay` now checks the table name length like the on-disk database it writes to. ### Description In `clickhouse-local` with a long `--default_database`, `CREATE TABLE` was accepted and the resulting table could then never be dropped: ``` $ clickhouse local --path=$D -q \"CREATE TABLE tc (a UInt8) ENGINE=MergeTree ORDER BY a; DROP TABLE tc\" \\ -- --default_database=<214 'd's> Code: 1001. std::filesystem::filesystem_error: filesystem error: in rename: File name too long [\".../store/<uuid>/tc.sql\"] [\".../metadata_dropped/<214 d's>.tc.<uuid>.sql\"] ``` The end state survives restarts and has no SQL escape hatch: `DROP`, `DROP ... SYNC` and `DETACH` all fail the same way and the table stays in `system.tables`. The threshold is 212 characters of `--default_database`. `IDatabase::checkTableNameLength` is a no-op default and only `DatabaseOnDisk` overrides it. `DatabaseOverlay` did not, so nothing was checked, while the overlay forwarded the create to a real `DatabaseAtomic` member carrying the same long name. `DROP` then builds `metadata_dropped/{db}.{table}.{uuid}.sql`, whose prefix alone exceeds `NAME_MAX`, so the per-table budget saturates to 0. The fix overrides `checkTableNameLength` in `DatabaseOverlay`, delegating to the first non-read-only member, which is the same member `createTable` writes to. It copies the loop of the adjacent `checkMetadataFilenameAvailability`. No new setting, constant or arithmetic. A real `DatabaseAtomic` with the same name already rejects this up front with `ARGUMENT_OUT_OF_BOUND`, so this makes the overlay consistent with the database that owns the file. The change only affects rejection at DDL time; a table already on disk still attaches and reads, because both `CREATE` call sites are gated by `mode <= LoadingStrictnessLevel::CREATE`. It does not rescue tables already stranded on disk, which would mean changing the `metadata_dropped` naming convention. Related: #113019, where this was found and measured. <details> <summary>Validation: zero regression window, and the object kinds covered</summary> The delegated limit rejects exactly the names that already fail. Longest table name that survives CREATE+DROP today, versus the limit, measured at four database sizes: | escaped db length | limit | longest working today | first failing today | | --- | --- | --- | --- | | 100 | 113 | 113 | 114 | | 150 | 63 | 63 | 64 | | 190 | 23 | 23 | 24 | | 200 | 13 | 13 | 14 | At every size, a name at the limit still works after the change and a name one over is now refused at CREATE instead of being accepted and stranded. `escapeForFileName` emits 3 bytes per non-word byte and the check measures the escaped name: 37 `-` characters (111 escaped bytes) work, 38 (114) do not, against the limit of 113. All six object kinds routed through `InterpreterCreateQuery` reproduced the accepted-then-undroppable behaviour and are covered by the single override: `MergeTree`, `Log` and `Memory` tables, `VIEW`, `MATERIALIZED VIEW` and `DICTIONARY`. `CREATE TEMPORARY TABLE` is not affected, since temporary tables never reach `metadata_dropped`. In the only construction site today the overlay and its writable member share a name, so the delegated limit is the intended one. If a future overlay had a differently named writable member, delegating still gives the correct answer, because the limit follows the database that owns the file. `DatabaseOverlay` is currently built only by `clickhouse-local`; #86768 would add server-side consumers. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114567",
          "createdAt": "2026-08-13T01:34:05Z",
          "updatedAt": "2026-08-13T07:10:12Z",
          "timestamp": "2026-08-13T07:10:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:29db7afeed9646e9b150",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109453",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109453",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add ALTER TABLE ... RECOMPRESS COLUMN",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109432 Introduce a new `ALTER TABLE ... RECOMPRESS COLUMN col` statement that re-compresses the existing data of a column with the column's current compression codec. Changing a column's codec with `MODIFY COLUMN col CODEC(...)` is metadata-only: the new codec applies to newly written data, while data already stored in existing parts keeps its old codec until the parts happen to be merged. `RECOMPRESS COLUMN` rewrites the data of `col` in existing parts so that it is compressed with the codec currently set in the table metadata. Because a compression codec does not change the serialized representation of a column, for `Wide` parts the recompression is done **without deserializing the values**: each compressed block of every substream `.bin` is decompressed and re-compressed one-to-one with the new codec. This keeps the decompressed content and granule boundaries byte-identical, so the marks file only needs its compressed offsets remapped (the decompressed offsets, per-granule row counts, the primary index and skip indexes are preserved and hardlinked). The decompressed content is unchanged, so the `uncompressed_size`/`uncompressed_hash` checksums are carried over from the source part and only the on-disk `file_size`/`file_hash` are recomputed. `Compact` parts (and dynamic-subcolumn types) cannot recompress a single column in isolation, so they fall back to a normal whole-part re-serialization that writes every column with its current codec. Implemented as a mutation. A new `ALTER RECOMPRESS COLUMN` grant is added. The issue also asks to check `RECOMPRESS` TTL mutations, which currently go through a full deserialize/re-serialize merge; that is left for a follow-up. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `ALTER TABLE ... RECOMPRESS COLUMN col`, which re-compresses a column's existing data with its current codec. For wide parts the data is recompressed without deserializing the column values.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109453",
          "createdAt": "2026-07-05T21:19:23Z",
          "updatedAt": "2026-08-13T07:01:14Z",
          "timestamp": "2026-08-13T07:01:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 24
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ec789705e7319f82c0ed",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105249",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105249",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Accessing tables as files, query construction and out-of-band modification in HTTP interface",
          "text": "Implements [#46925](https://github.com/ClickHouse/ClickHouse/issues/46925) according to the [updated spec](https://github.com/ClickHouse/ClickHouse/issues/46925#issuecomment-4475417259). Closes: https://github.com/ClickHouse/ClickHouse/issues/46925 Related: https://github.com/ClickHouse/clickhouse-docs/pull/6398 ## What's in the box - **Path → database/table/format/compression.** New settings `http_allow_database_as_path`, `http_allow_table_as_file`, `http_allow_filters_as_path`, `http_allow_filters_as_unrecognized_url_parameters` let the HTTP interface interpret `/database/table.format.compression` (and hive-style `/name=value/` partitions) in the URL path. When a path identifies a table, the request is processed as `SELECT * FROM database.table`. - **Query construction settings.** `select`, `order`, `sort`, `filter`, `page` wrap the base query as `SELECT [select] FROM (...) [WHERE filter] [ORDER BY order]`. `page=N` translates to `offset = limit * (N - 1)` (errors if `limit` is unset or `offset` is also set). Multiple `filter` URL parameters are combined with `AND`. - **Format/compression overrides.** `default_format`, `format`, `input_format`, `output_format` are now first-class settings; the explicit `format` / `output_format` overrides win over the FORMAT clause in the query and the file extension in the path, while `default_format` is only the fallback used when nothing else selects a format (so the path extension wins over it). A generic `compression` setting wraps the response body (independent of HTTP `Content-Encoding`). The URL path file extension is equivalent; specifying both with conflicting values throws. - **Implicit table.** When a path table coexists with a user `query`, the path table is exposed via `implicit_table_at_top_level`, so `?query=SELECT a, b` against `/hits.csv` reads from `hits` with no `FROM`. If the user's query already has `FROM`, the path component only seeds the download filename. - **Content-Disposition.** Binary or compressed responses get `Content-Disposition: attachment; filename=…`, where the filename comes from the URL path component (or `result.<format>.<compression>` as fallback). - **`database` and the format settings are now real settings.** The `database` URL parameter and the `X-ClickHouse-Database` / `X-ClickHouse-Format` headers all flow through the regular settings pipeline (the header still overrides the URL parameter to preserve historical precedence). `X-ClickHouse-Database` is an alias for `database`; `X-ClickHouse-Format` is an alias for `output_format` (not `default_format`), i.e. an explicit override of the response format that also wins over a `FORMAT` clause in the query. It maps to `output_format` rather than to the bidirectional `format` because the header has always described the response only, so it must not reinterpret the body of an `INSERT`. For the same reason, `--format` in `clickhouse-client` maps to `output_format` (it has always been output-only there), while in `clickhouse-local` it keeps its historical bidirectional meaning and maps to `format`. They are marked `changeable_in_readonly` in the shipped default profile so they remain settable on HTTP GET (which forces `readonly = 2`) and for `readonly = 1` users. Because `database` is a real setting, a profile that constrains it (e.g. marks it readonly or restricts its values) is enforced consistently on every way of choosing the current database: `USE`, `SET database = ...`, the HTTP `database` parameter/header, and the connect-time database of the native TCP, MySQL, and PostgreSQL protocols. In `clickhouse-local`, an explicitly configured database (`--database` or a config-file `database` key) is mirrored into the `database` setting so a profile-inherited value cannot override it. `HTTPHandler::processQuery` is restructured so that authentication and the user profile are applied first, then the auth-related parameter (`role`), then the read-only enforcement, and only then the general settings. ## Changelog category - New Feature ## Changelog entry Added a way to access tables and databases via URL paths in the HTTP interface (e.g. \\`/database/table.format.gz?filter=a>0\\`), plus new settings (\\`http_allow_database_as_path\\`, \\`http_allow_table_as_file\\`, \\`http_allow_filters_as_path\\`, \\`http_allow_filters_as_unrecognized_url_parameters\\`, \\`select\\`, \\`order\\`, \\`sort\\`, \\`filter\\`, \\`page\\`, \\`compression\\`, \\`format\\`, \\`input_format\\`, \\`output_format\\`, \\`default_format\\`, \\`database\\`) that compose with one another and with the existing \\`query\\` URL parameter. ## Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) — [ClickHouse/clickhouse-docs#6398](https://github.com/ClickHouse/clickhouse-docs/pull/6398) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > High risk because it significantly changes HTTP request routing/parsing and introduces many new per-query settings that affect query construction, output formatting, compression, and readonly constraints, which could impact security and compatibility. > > **Overview** > **Enables “table-as-file” and query shaping via the HTTP interface.** The server can now interpret URL paths like `/db/table.format[.compression]` (plus optional hive-style `name=value` path filters) to auto-generate `SELECT * FROM ...`, apply `select`/`filter`/`order`/`sort` wrappers, and support paging via new `page` setting (translated into `limit`/`offset`). > > **Promotes HTTP-specific knobs into first-class settings and aligns client behavior.** Adds new settings for `database`, `default_format`, `format`, `input_format`, `output_format`, `compression`, and HTTP path feature gates; updates HTTP handler to route headers/URL params through the normal settings/constraints pipeline (including readonly carve-outs), adds generic response-body compression and `Content-Disposition` for binary/compressed outputs, and updates client/local tooling to mirror config/CLI format/database options into per-query settings. > > **Adjusts core settings/exec behavior for compatibility.** Widens `limit`/`offset` to `Double` (supporting negative/fractional values), updates analyzers/interpreters/cluster proxy to handle/strip initiator-only settings in distributed queries, refines handler factory ordering/prefix mounting, and adds/updates stateless tests and configs for the new HTTP path semantics and hints. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 8efcce89edbde35344024d1f9306be66470f3e1a. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105249",
          "createdAt": "2026-05-18T16:22:57Z",
          "updatedAt": "2026-08-13T06:53:49Z",
          "timestamp": "2026-08-13T06:53:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 83
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bb25c81ded09959f2ef3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112874",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112874",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "MergeTree: rethrow retryable errors from checkDataPart instead of returning empty checksums",
          "text": "> **Series**: #112871 -> **#112874** (this), #112872, #112873, #112875, #112876 **Problem.** `checkDataPart` swallows a retryable error and returns empty checksums, so `CHECK TABLE` and the fetch path treat a transient failure as \"verified\" and can persist an empty integrity baseline. Failure scenario: ``` checkDataPart(): catch retryable -> return {} CHECK TABLE : empty -> reports OK downloadPartToDisk: empty -> accepts an unverified packed part ``` **Fix.** Rethrow retryable errors so callers' retry/skip guards fire; remove the now-dead empty-checksums guard on the fetch path. **Changes.** - `Storages/MergeTree/checkDataPart`: `return {}` -> `throw` on a retryable error. - `Storages/MergeTree/DataPartsExchange`: delete the now-dead empty-checksums fetch guard. - `tests/integration/test_azure_403_handling`: CHECK-TABLE-surfaces-transient test. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `checkDataPart` swallowing a retryable error and returning empty checksums, so `CHECK TABLE` could report a part as OK and a fetch could accept an unverified part after a transient failure. Retryable errors are now rethrown so the check is retried later instead of masking a transient failure as verified.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112874",
          "createdAt": "2026-08-01T06:53:40Z",
          "updatedAt": "2026-08-13T06:49:01Z",
          "timestamp": "2026-08-13T06:49:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "arsenmuk",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cbc03bdaa842781a94d7",
        "signalId": "github:ClickHouse/ClickHouse:issue:85293",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:85293",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Dictionary credentials rotation",
          "text": "### Company or project name Cabify ### Use case We are using Clickhouse Cloud. We are using dictionaries to populate data from other ClickHouse tables. The source looks sth like this: ``` SOURCE(CLICKHOUSE( USER '${CH_DICT_USER}' PASSWORD '${CH_DICT_PASSWORD}' DATABASE '${CH_DB}' TABLE '${CH_TABLE}' ) ``` The issue we are encountering is that we cannot rotate the credentials; doing so would make the dictionary stop working, as the credentials used during its creation would become invalid. We use short-lived dynamic credentials (managed by [Vault](https://github.com/hashicorp/vault)) as a general approach, so not being able to rotate provided dictionary credentials seems not ideal for us. Requiring users with the `default_role` to provide static credentials for creating a dictionary appears to be quite restrictive and may contradict certain security best practices implemented by different organizations. We are aware that no password is required if the `DEFAULT` user is used for creating the dictionary, sth like: ``` SOURCE(CLICKHOUSE( TABLE '${CH_TABLE}' )) ``` That approach could work, as rotating `DEFAULT` user credentials won't affect the already created dictionaries. However, again this would force to use `DEFAULT` user for creating a dictionary, which goes against PoLP (Principle of Least Privilege). We have followed the following approach for now: > Use a user with `default_role` privileges, static credentials and HOST LOCAL. Although it could met some of our security requirements (as that user can NOT access from outside the instance), the solution is not ideal. ### Goal Allow the provision of user credentials for creating the dictionary, ensuring that the dictionary remains functional even if the credentials are rotated. Additionally, I’m unclear on why the user needs to have the `default_role` just to create a dictionary, as this seems to grant excessive permissions. ### Describe the solution you'd like I expect that any user with the `default_role` will have similar behaviour as the `DEFAULT` user when creating a dictionary. Once dictionaries are created by the `DEFAULT` user, the user’s credentials can be changed, and the dictionaries continue working. So ideally, we provide a user with HOST LOCAL with credentials that can be rotated without affected related dictionaries. For example, instead of hardcoding user credentials at the dictionary level, verify the authentication and authorization of the user during the dictionary creation process. If the user is valid, link their username to the dictionary. Corresponding validations should also be performed during dictionary reloads, similar to the current approach with the `DEFAULT` user. ### Describe alternatives you've considered _No response_ ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/85293",
          "createdAt": "2025-08-08T12:23:39Z",
          "updatedAt": "2026-08-13T06:48:59Z",
          "timestamp": "2026-08-13T06:48:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "feature"
          ],
          "author": "x-martinez",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6082bdae8eb9ecaac740",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105780",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105780",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add SCANN Vector Index Support",
          "text": "Add [Google ScaNN](https://github.com/google-research/google-research/tree/master/scann) as a new approximate nearest neighbor (ANN) backend for the `vector_similarity` skip index in `MergeTree` tables, accessible via `TYPE vector_similarity('scann', ...)` syntax. ScaNN uses an IVF-based index with asymmetric hashing (LUT16) and exact reranking, complementing the existing HNSW backend. The implementation integrates three new contrib dependencies — `scann`, `highway` (portable SIMD), and `eigen` (header-only linear algebra) — without modifying any upstream submodule source files; all platform adaptation is in the corresponding `-cmake/` wrappers. **Supported distance functions:** `L2Distance`, `cosineDistance`, `dotProduct`. **New query-tuning settings:** - `scann_num_leaves_to_search` — number of partitions to search at query time (0 = use the index-configured default, typically `num_vectors^0.25`). - `scann_candidate_pool_size` — asymmetric hashing candidate pool size fed into the exact reranker (0 = automatic: `1000 × topK`). All training artifacts (partitioner proto, asymmetric hashing codebook, hashed dataset, datapoint-to-token mapping) are serialized to disk on index build via `SingleMachineSearcherConfig`, so search results are deterministic across server restarts without retraining. **Example:** ```sql CREATE TABLE tab ( id UInt64, vec Array(Float32), INDEX idx vec TYPE vector_similarity('scann', 'cosineDistance', 128) GRANULARITY 100000000 ) ENGINE = MergeTree ORDER BY id; -- Query-time tuning SELECT id, cosineDistance(vec, [...]) AS dist FROM tab ORDER BY dist LIMIT 10 SETTINGS scann_num_leaves_to_search = 80, scann_candidate_pool_size = 5000; ``` ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added Google ScaNN as a new backend for the vector_similarity index type (syntax: TYPE vector_similarity('scann', 'cosineDistance', N)). Supports L2Distance, cosineDistance, and dotProduct distance functions with deterministic results across restarts. Two new settings — scann_num_leaves_to_search and scann_candidate_pool_size — allow per-query tuning of recall/latency trade-offs. ### Related issue: Closes #103664. The issue mentions ScaNN, but the reply references SPANN — not sure if that was a typo. I went ahead and implemented ScaNN as originally requested, and would welcome any response.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105780",
          "createdAt": "2026-05-25T11:29:39Z",
          "updatedAt": "2026-08-13T06:46:39Z",
          "timestamp": "2026-08-13T06:46:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 27
          },
          "labels": [
            "submodule changed",
            "manual approve",
            "can be tested",
            "pr-experimental"
          ],
          "author": "lvzhipin03",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d70cc53498fc02fd4e77",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114496",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114496",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not rewrite arrayExists to has when the element and needle string types differ",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `arrayExists(x -> x = needle, arr)` returning a wrong result, or failing with `TOO_LARGE_STRING_SIZE`, when the array element and the needle are different string types (for example a `String` needle against an `Array(FixedString(N))` element). The `optimize_rewrite_array_exists_to_has` optimization, enabled by default, rewrote such a call to `has`, which does not compare zero-padded the way `=` does. ### Description `optimize_rewrite_array_exists_to_has` (on by default) rewrites `arrayExists(x -> x = c, arr)` into `has(arr, c)`. The two do not agree for a string-family pair whose types differ, so the same query returns different answers depending on the setting. On released 26.7.4.21 and on master, with `v = ['V0']` of type `Array(FixedString(3))`, `arrayExists(x -> x = 'V0\\0', v)` is `1` with the setting off and `0` with it on, while `toFixedString('V0', 3) = 'V0\\0'` is `1`. Over an `Array(LowCardinality(FixedString(N)))` element the rewrite turns a working query into `Code: 131 TOO_LARGE_STRING_SIZE`, rejecting valid input. <details> <summary>Measured on 26.7.4.21 and master</summary> | query on `v = ['V0']`, `Array(FixedString(3))` | setting off | setting on (default) | |---|---|---| | `arrayExists(x -> x = 'V0\\0', v)` | 1 | **0** | | `arrayExists(x -> x = 'V0\\0\\0', v)` | 1 | **0** | | `arrayExists(x -> 'V0\\0' = x, v)` | 1 | **0** | | `arrayExists(x -> x = 'V0abc', v)` (control) | 0 | 0 | | `arrayExists(x -> x = 'V0', v)` (control) | 1 | 1 | | `arrayExists(x -> x = tuple('V0\\0'), v)` on `Array(Tuple(FixedString(3)))` | 1 | **0** | | `arrayExists(x -> x = toFixedString('V0',4), [toFixedString('V0',3)])` | 1 | **0** | | `arrayExists(x -> x = 'V0\\0\\0', v)` on `Array(LowCardinality(FixedString(3)))` | 1 | **Code: 131** | | `arrayExists(x -> x = ['V0\\0'], [[toFixedString('V0',3)]])` | 0 | **1** | | `arrayExists(x -> x = map('k','V0\\0'), [map('k',toFixedString('V0',3))])` | 0 | **1** | </details> Root cause: the pass admits the rewrite whenever the element and needle have a common supertype. For `FixedString(N)` plus `String` that supertype exists, but reaching it casts the element to `String`, stripping its trailing NULs, while `equals` compares the pair zero-padded. A supertype is not sufficient for interchangeability, which is why this pass already declines for `NULL`, and `HasToInPass` for Date, Enum, IPv4 and floats. The fix declines the rewrite when the element and needle are both in the string family but not the same type, recursing through `Tuple`, `Array` and `Map` members because the divergence reproduces at any nesting depth. Over a constant array the container cases diverge the other way (`0` off, `1` on): there `has` compares raw `Field`s, which are not padded either. Identical types cannot diverge, so `String`/`String` and equal-width `FixedString` pairs keep the optimization, as do numeric, Date and Enum elements. It takes no position on what `has` should mean here; that belongs to #112953. Validated by an A/B over two binaries: every divergent cell goes from disagreeing to agreeing, the controls do not move, and a 298-test array/membership/FixedString sweep shows no regression against pristine master. The test's pinned direct `has`/`indexOf`/`countEqual` row is today's behaviour, held as a control; it moves when #112953 lands.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114496",
          "createdAt": "2026-08-12T14:44:50Z",
          "updatedAt": "2026-08-13T06:41:13Z",
          "timestamp": "2026-08-13T06:41:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c1eff5c874a2e91d960f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104350",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104350",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix correlated subquery + GROUP BY ROLLUP under group_by_use_nulls",
          "text": "A correlated subquery references an outer column that is a GROUP BY key of an outer query under `group_by_use_nulls = 1` with `WITH ROLLUP`/`CUBE`/`GROUPING SETS`. The actual values fed into the inner subquery are post-rollup `Nullable`s, but the analyzer used to leave the inner column reference at its non-Nullable type. The planner then built CAST wrappers and aggregate functions with the original signature, and at header inference (after decorrelation replaced `PLACEHOLDER` nodes with `INPUT` nodes of the actual `Nullable` types) the precomputed wrappers raised `Logical error: 'Bad cast from type DB::ColumnNullable to DB::ColumnVector<...>'`. The minimal reproducer from issue #91119 hits `createUInt8ToBoolWrapper`: ```sql SELECT (SELECT c0) FROM (SELECT 1::Bool) t0(c0) GROUP BY c0 WITH ROLLUP SETTINGS group_by_use_nulls = 1; ``` The same surface bug class also appears for plain numeric types (modulo, plus), aggregate functions over correlated outer columns (e.g. `(SELECT anyLastOrDefault(number))`), correlated `HAVING`/`WHERE` filter steps, and `Nullable` source columns. The root cause is in `QueryAnalyzer::resolveExpressionNode`: the `nullable_group_by_keys` walk stops at the first enclosing `QUERY` scope, so the outer scope where the column was defined is never consulted for an inner correlated reference. This change makes the walk continue past the inner `QUERY` scope when the `node` is a column whose source table expression is registered in a deeper outer scope (i.e. the column is correlated). The aggregate-function exclusion is now evaluated against each scope's `expressions_in_resolve_process_stack`, so an outer scope's `nullable_group_by_keys` applies even when the inner subquery is currently inside an inner aggregate (the inner aggregate operates on the outer's already-`Nullable` values). For correlated matches the existing node is mutated in place instead of being replaced with a clone of the GROUP BY key. The same `shared_ptr` is referenced by the enclosing `QueryNode::correlated_columns` list (registered via `addCorrelatedColumn`); cloning would leave the planner's `correlated_columns_set` pointing to the original non-`Nullable` node, and `PlannerActionsVisitor::visitColumn` (which uses type-sensitive equality) would then fail to identify the projected column as correlated and skip the `PLACEHOLDER` step that the decorrelation pass replaces. Closes #91119. Closes #106377. Closes https://github.com/ClickHouse/ClickHouse/issues/109509 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `Bad cast` exception in correlated subqueries that reference an outer column promoted to `Nullable` by `group_by_use_nulls = 1` with `GROUP BY` `WITH ROLLUP`/`CUBE`/`GROUPING SETS`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104350",
          "createdAt": "2026-05-08T08:49:47Z",
          "updatedAt": "2026-08-13T06:41:12Z",
          "timestamp": "2026-08-13T06:41:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 61
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a5c446fceeeafce2e9a1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110429",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110429",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix \"Cannot write to finalized buffer\" in MergeTreeDeduplicationLog::rotate",
          "text": "Fixes a server abort with the logical error `Cannot write to finalized buffer` that was hit by the stress test. `MergeTreeDeduplicationLog::rotate` finalized the current log writer and only afterwards created the writer for the new log file. If creating the new writer threw — a transient I/O error, or, in the CI failure, a memory-tracker fault injection hitting the `WriteBufferFromS3` allocation after `finalize` had already succeeded — `current_writer` was left pointing at the already finalized buffer while `stopped` was still `false`. The next write to the deduplication log then wrote to that finalized buffer and aborted. In the CI failure the write came from the background `MergeTreeCleanupThread`: ``` MergeTreeCleanupThread::iterate -> MergeTreeData::clearEmptyParts -> StorageMergeTree::dropPartNoWaitNoThrow -> MergeTreeDeduplicationLog::dropPart -> writeRecord -> WriteBuffer::write -> Logical error: 'Cannot write to finalized buffer' ``` The failure was reproduced on a disk that does not support writing with append (`s3_plain_rewritable`), where the deduplication log is rotated on every operation, so a failing `writeFile` during rotation is easy to hit. The fix makes `rotate` exception-safe: the writer for the new log file is opened first, before any state is changed. If it throws, nothing has changed and `current_writer` still points to the previous, live writer, so the log stays usable and the operation can be retried; only once the new writer is ready is the previous one finalized and swapped out. This preserves the invariant that `current_writer` is never a finalized buffer. Additionally (a review finding): making `rotate` exception-safe was not enough to honor the retry contract. `addPart` had already published the block IDs into the in-memory deduplication map (and written the `ADD` records) before the rotation ran, while `MergeTreeSink` commits the part only after `addPart` returns. An insert aborted by a rotation failure therefore left its block IDs published, and a client retry of the same insert was wrongly deduplicated against a part that never became active — silently dropping the data. Publication in `addPart` is now all-or-nothing: on any failure, the already published block IDs are removed from the in-memory map, and compensating `DROP` records are written to the still-live writer (best effort), so replaying the log on server startup does not re-publish them either. Additionally (a second review finding): the compensating `DROP` writes above assumed `current_writer` was still a usable, live writer. That only holds when the failure came from `rotate` itself, inside `rotateAndDropIfNeeded`. If instead one of the `ADD` records failed to write directly, `WriteBuffer::next` cancels the buffer on any exception, so `current_writer` was left canceled and refused any further writes — the compensating `DROP` records were silently lost in that case, reopening the same wrongly-deduplicated-retry window. The rollback now rotates to a fresh writer first whenever `current_writer` is canceled; `rotate` already tolerates an already-canceled `current_writer`, so this is also safe (if redundant) when it is still the live writer from before a failed `rotate` call. Additionally (a third review finding, caught by the CI's own gtests): the previous fix for the canceled-writer case introduced a new abort. `rotate`'s rollback called `current_writer->finalize()` unconditionally on the previous writer, but `WriteBuffer::finalize` disallows calling it on an already-canceled buffer — it throws a `LOGICAL_ERROR`, which aborts the process immediately in debug and sanitizer builds (before the surrounding `try`/`catch` even runs). `rotate` and `shutdown` now skip `finalize` when `current_writer` is already canceled, since a canceled buffer has nothing left to flush anyway. Additionally (a fourth review finding): the all-or-nothing rollback in `addPart` still narrowed the deduplication window on a failure. `LimitedOrderedHashMap::insert` evicts the oldest entry once the map is at capacity, and the rollback only erased the block IDs the failing call itself had published — it did not restore entries evicted by those `insert` calls. With a small deduplication window, a failed insert could therefore evict an unrelated, already-active part's block ID from memory, and a retry of that unrelated block ID would then be wrongly accepted instead of deduplicated, until the log was replayed from disk. `addPart` now defers all `deduplication_map.insert` calls until the durable writes and the rotation have both succeeded, so the in-memory map is never mutated on a path that might still need to roll back. Additionally (a fifth review finding): `rotate` logged and suppressed failures to `finalize`/`sync` the previous log file's writer. Since `rotateAndDropIfNeeded` is the success boundary of `addPart`, an insert whose `ADD` records had just been written to that file was still treated as durably recorded: `addPart` published the block IDs and `MergeTreeSink` committed the part, even though the only on-disk `ADD` records may never have reached durable storage. After a restart the deduplication log forgot the committed insert, so a client retry of the same block was wrongly accepted and duplicated the data. `rotate` now rethrows such a failure after switching over to the new writer — so the log itself stays usable — and `addPart` treats the insert as failed: its rollback writes compensating `DROP` records into the freshly opened log file and the insert is aborted instead of committing a part the deduplication log may have already forgotten. Additionally (a sixth review finding): even after the rollback stopped re-publishing rolled-back block IDs, it still narrowed the deduplication window across a server restart. The rollback wrote compensating `DROP` records, but on startup `loadSingleLog` replays the log in order, so the failed insert's `ADD` record is applied — evicting the oldest committed block ID from the bounded in-memory map — before the `DROP` erases the rolled-back one. With `deduplication_window = 1`, a committed `block1` followed by a failed insert of `block2` replayed as `ADD block1`, `ADD block2`, `DROP block2` and ended with an empty map, so after a restart a retry of `block1` was wrongly accepted and its data duplicated (larger windows dropped the oldest committed block the same way). Rolled-back inserts are now written with a distinct `CANCEL` record instead of a `DROP`, and replay cancels each `CANCEL` against its matching preceding `ADD` and skips both before applying the remaining records — so a rolled-back insert consumes no deduplication-window slot on replay and the reloaded state matches the live in-memory state exactly. Additionally (a seventh review finding): even with rolled-back inserts written as `CANCEL` records, log retention still over-counted them. `dropOutdatedLogs` decides which older log files are redundant by summing each file's raw record count from the newest backwards until the deduplication window is covered, but a cancelled `(ADD, CANCEL)` pair contributes nothing to the reconstructed map. A failed multi-block insert could therefore make retention treat those transient records as consumed deduplication-window slots and drop an older log that still held live, committed block IDs. With `deduplication_window = 2`, committed `block1` / `block2` in the first log and a failed `addPart({\"block3\", \"block4\", \"block5\", \"block6\"}, ...)` whose rollback wrote four `CANCEL` records into the second log, the first restart rebuilt the map correctly but, seeing the inflated raw counts, rotated and dropped the first log; a second restart then replayed only the `CANCEL`-only log and forgot `block1` / `block2` — wrongly accepting, and duplicating, a retry of those committed inserts. Retention accounting is now based on the records that survive cancel-pair elimination: `applyRecords` recomputes each log's `entries_count` from only the non-cancelled records on replay, and `addPart`'s rollback undoes the retention count of the rolled-back `ADD` records (and does not count the `CANCEL` records), so the live and replayed accounting stay consistent and a rolled-back insert never shrinks the retained history. Additionally (an eighth review finding): the same durability boundary left `dropPart` inconsistent. It wrote a `DROP` record, erased the block ID from the in-memory map, and rotated, for each covered block ID one at a time; now that `rotate` can rethrow a failure to `finalize`/`sync` the previous log file, a multi-block drop could throw partway through — after some of the dropped part's block IDs had been erased but before the rest — leaving the in-memory map in a partial state that its caller `StorageMergeTree::dropPartNoWaitNoThrow` never repairs (it has already taken the part out of the active set and does not retry the drop). The remaining block IDs then stayed published against a part that no longer exists, and any non-durable `DROP` records resurrected the erased ones after a restart. `dropPart` is now transactional like `addPart`: it collects every covered block ID, writes all the `DROP` records, rotates, and only then erases the block IDs from the map, so a failed drop leaves every covered block ID published (the deduplicating, safe direction) instead of a half-applied mixture. Additionally (a ninth review finding): the reworked `dropPart` was still not all-or-nothing on disk. `writeRecord` flushes every record, so when the write of one `DROP` record failed partway through a multi-block drop, the records written before it were already durable while no block ID had been erased from the in-memory map. After a restart, replaying that durable prefix erased only part of the failed drop: a block ID whose `DROP` record reached the disk stopped deduplicating while its siblings still did — a half-applied drop that matches neither the live map (which kept every covered block ID) nor a completed drop, and that the caller `StorageMergeTree::dropPartNoWaitNoThrow` never repairs. Worse, the failed write left `current_writer` canceled, and a canceled buffer silently discards all further writes, so the `ADD` records of later, successfully committed inserts would never reach the disk either and those inserts would be forgotten after a restart, wrongly accepting and duplicating their retries. `dropPart` now mirrors `addPart`'s rollback: on any failure it writes a compensating `CANCEL` record for each `DROP` record that was written (rotating to a fresh writer first when the failed write canceled the current one) and undoes their retention count, and replay cancels a `CANCEL` against the most recent preceding un-cancelled `ADD` or `DROP` of the same block ID — so a failed drop keeps every covered block ID published both live and across a restart, and the log stays usable and durable afterwards. Additionally (a tenth review finding): the persisted `CANCEL` record had no downgrade contract. A server from before this change replays every log record as either an erase (`DROP`) or an insert (anything else), so after a downgrade with rollback records already on disk, the `CANCEL` written for a rolled-back insert replayed as an insert: the never-committed block ID stayed published, and a client retry of the failed insert was wrongly deduplicated — silently dropping its data. The rollback records are now encoded so that every server version, old or new, replays them with the correct net effect, with no need for a format version. The rollback of a failed insert is written as a plain `DROP` record carrying a reserved part name (`cancel`, which can never collide with a real part name — and the part name of a `DROP` record is never parsed by any server version): an older server replays it as the erase that unpublishes the never-committed block ID, so a retry of the failed insert is accepted, while a server with this change recognizes the marker and cancels the `(ADD, DROP)` pair out of the replay entirely, preserving the deduplication-window and retention guarantees above. The `CANCEL` operation remains only as the rollback of a failed drop, where it carries the real, parseable part name: an older server replays it as the insert that restores the still-published block ID — exactly the rollback's net effect, with only the entry's position in the eviction order diverging. A downgraded server is therefore never worse off on these logs than it would have been running on logs it produced itself. Additionally (an eleventh review finding): two remaining bookkeeping steps could still throw at a point where the `Cannot write to finalized buffer` state or a broken retry contract would return, this time on a memory-allocation (`std::bad_alloc`) rather than an I/O failure. First, `rotate` registered the new log file in `existing_logs` (a `std::map`, whose `emplace` allocates a node) only after finalizing the previous writer; if that allocation threw, `current_writer` was left pointing at the finalized old writer — the exact abort this pull request eliminates — and no rollback path could detect it, because the buffer is finalized, not canceled. `rotate` now performs that registration before finalizing the old writer, so the only remaining throwing bookkeeping step runs while the old writer is still live and usable; the switch-over to the new writer (an integer store and a `unique_ptr` move) is non-throwing. Second, `addPart` deferred `deduplication_map.insert` until after the `ADD` records were durable, but that insert is itself a throwing, evicting step: `LimitedOrderedHashMap::insert` allocates and, when the map is full, evicts the oldest entry before inserting. An allocation failure there — after the records were already durable — propagated an exception without writing any compensating rollback records, so a client retry could be deduplicated against a part that never committed (and, in the window-full case, an unrelated committed block ID was evicted from the live map). Publication is now split so that nothing which can throw runs after durability: the block IDs are inserted up front with a new `LimitedOrderedHashMap::insertWithoutEviction` — strongly exception-safe, and crucially never evicting, so a failure before the durable writes is rolled back with a plain non-allocating `erase` that never drops an unrelated block ID — and the deduplication window is enforced afterwards with `LimitedOrderedHashMap::trimToMaxSize`, which only pops the oldest entries and so cannot throw at a point where the insert could no longer be rolled back. `insert` and `setMaxSize` are expressed in terms of these two primitives, so their behavior is unchanged. Additionally (a twelfth review finding): two more failure modes, on the load and accounting paths. First, `loadSingleLog` appended each record and its originating log number to two parallel vectors as two separate steps; if the second append threw (for example `std::bad_alloc` while loading a large deduplication log) after the first had succeeded, the vectors were left with different lengths, and because `load` tolerates and still replays whatever was read, `applyRecords` — which indexes the two in parallel — then read past the end of the shorter one, turning an allocation failure at startup into undefined behavior. The two appends are now atomic: if appending the log number throws, the just-pushed record is removed again (`pop_back` on a non-empty vector never throws), so the two vectors always stay in lockstep. Second, the retention fix above had repurposed each log file's single record count to mean *records surviving cancel-pair elimination*, but that same field is also the log-growth threshold that drives rotation and compaction. A rolled-back operation's records net to zero surviving records, so repeated transient failures could append arbitrarily many raw rollback pairs while the count stayed at zero — the newest logs never reaching the rotation threshold, growing without bound, and forcing `load` to materialize a number of records proportional to the number of failures rather than to the deduplication window. Each log file now carries two counts: the raw `entries_count` (every physical record, which drives rotation, so a rollback-heavy log still rotates and stays bounded) and `effective_entries_count` (only the records that survive cancel-pair elimination, which drives retention, as in the seventh finding above). `addPart` and `dropPart` count every physical record — including their compensating rollback records — towards the raw count and never decrement it, undoing only the effective count of the rolled-back records; `applyRecords` recomputes both from the replayed stream. Additionally (a thirteenth review finding, following up on the twelfth): splitting the raw and effective counts bounded the growth of a single log file but not the *number* of files. `dropOutdatedLogs` cannot reclaim the `(ADD, rollback)` and `(DROP, CANCEL)` record pairs a rolled-back operation leaves behind — the rollback record sits in a newer file while the record it cancels sits in an older file that is still retained for other, live block IDs, and retention only ever drops an oldest prefix — so under repeated transient write or fsync failures those cancelled-out pairs, and the log files holding them, would still accumulate without bound, and every restart would replay a number of records proportional to the number of failures. A new `compact` step rewrites the whole live deduplication state — which the in-memory map already holds exactly — into a single fresh log file and drops every older file, once the raw record count across all files exceeds the effective (surviving-record) coverage by more than a couple of rotation intervals (which only happens once rolled-back operations have piled up, since the two counts are equal in normal operation). Written in the map's insertion order, the snapshot replays to the identical state, so discarding the accumulated history is safe. It runs at the end of a successful `addPart`/`dropPart` and after `load`, so both the retained files and the load-time replay stay bounded by the deduplication window regardless of how many failures preceded them. `compact` is best effort and never throws — it finalizes the snapshot before removing any old file, and on any failure leaves the existing files and writer untouched. Additionally (a fourteenth review finding): the compaction above was skipped entirely on disks that do not support writing with append (for example `s3_plain_rewritable`, the disk on which the original abort reproduced), on the assumption that its every-operation rotation already keeps the log small. It does not. `rotateAndDropIfNeeded` rotates on every operation there, so each rolled-back insert or drop still leaves a newer log file whose effective (surviving-record) count is zero — its rollback record cancels an `ADD`/`DROP` in an older file that is still retained for other, live block IDs — and `dropOutdatedLogs`, which can only drop an oldest prefix, cannot reclaim any of them. The retained files, and the records `load` replays on every restart, therefore still grew with the number of failures, unbounded, in exactly the regime this pull request targets. Compaction now runs on such disks too: because the finalized snapshot file cannot be reopened for appending there, `compact` writes the live snapshot to a fresh durable file and then starts the next operation in another fresh, empty file, instead of reopening the snapshot — reducing the retained history to the snapshot on any disk. Additionally (a fifteenth review finding, two more accounting issues). First, the all-or-nothing rollback discounted the rolled-back `ADD` (or `DROP`) records from the log file's `effective_entries_count` — the count `dropOutdatedLogs` uses for retention — only once, after the whole rollback loop, as `effective_entries_count -= written`. But `writeRecord` flushes per record, so a compensating write can throw partway through the loop after some records are already durable, and the post-loop decrement was then skipped entirely, leaving the count as if none of the rolled-back records had been cancelled. A replay of the partially written stream, however, cancels out exactly the records whose compensating record reached disk, so the inflated live count could make `dropOutdatedLogs` drop an older log that still held committed block IDs, and a restart then forgot those committed inserts and wrongly accepted — and duplicated — their retries. The discount now happens once per successfully written compensating record, right after it is durable, so a mid-loop failure discounts only the records that reached disk, matching what a replay reconstructs. Second, on a disk without append support (such as `s3_plain_rewritable`) the compaction bounded the number of retained files against operations and failures but not against restarts alone: every rotation, including the one in `load`, starts a fresh file, and `dropOutdatedLogs` can never reclaim a zero-record file that sits after the file holding the live state (a committed file or a compaction snapshot), because it only drops an oldest prefix. So each restart with no new operations left one more empty file behind — `snapshot`, `empty1`, `empty2`, … — and the retained files, and the records `load` had to replay, grew as O(number of restarts). `load` now removes the trailing zero-record log files before its first rotation (only without append support — with append support the last file is reopened and reused instead), so a restart is idempotent: the file holding the live state plus exactly one fresh writer file. Additionally (a sixteenth review finding): the compaction's failure cleanup could leave a stale snapshot behind that corrupts the next replay. `compact` writes its snapshot at a log number one past `current_log_number`; if the snapshot is made durable but the compaction then fails before switching over to it (for example the reopen of the snapshot for appending throws), the cleanup tried to remove the orphan snapshot and treated a failure to remove it as harmless. It is not: the durable snapshot then survives at a *higher* log number than the older files the server keeps appending to, so a later successful insert commits newer block IDs into an older file, and on the next restart `load` replays the stale snapshot last — after those newer records — resurrecting evicted block IDs and forgetting committed ones (a retry can then be wrongly deduplicated, or a committed insert forgotten, after the restart). The cleanup now goes through `neutralizeOrphanLog`, which removes the orphan file and, if the removal fails too, overwrites it with an empty file — an empty log replays as a no-op regardless of its log number, so it can no longer corrupt the reconstructed state, preserving a consistent numbering boundary instead of leaving an invisible higher-numbered log. Additionally (a seventeenth review finding, two replay/cleanup issues). First, `applyRecords` cancelled out the record pairs left by a rolled-back operation by matching a rollback record to the most recent preceding un-cancelled record of the same block ID, regardless of whether it was an `ADD` or a `DROP`. That pairing is too loose on the failed `dropPart` + failed `sync` path, where the rollback `CANCEL` can survive on disk while the `DROP` it was meant to undo never reached durable storage: matching by block ID alone then cancelled the committed `ADD` of that block instead, forgetting a still-published block ID after a restart and wrongly accepting a duplicate retry. Each rollback record is now paired only with a preceding record of the exact kind it undoes — a `CANCEL` with a real `DROP`, a cancelled-add `DROP` with an `ADD` — so a rollback record whose target was lost leaves the unrelated committed record untouched. Second, when `compact` could not remove some of the old, superseded log files, it treated the lingering file as harmless because the snapshot replays after it. That holds for the set of block IDs but not for their FIFO order: a lingering pre-snapshot file replays to a stale intermediate order, and the snapshot's `ADD` records on top do not refresh the position of an already-present key, so the next insert after a restart could evict a different committed block than the live process would. `compact` now neutralizes an un-removable old file (retrying the removal and, failing that, emptying it) so the snapshot alone determines the reloaded state and its eviction order. Additionally (an eighteenth review finding, refining the seventeenth): pairing a rollback record only by record kind and block ID is still too loose, because a block ID can be reused across part generations — committed as one part, dropped, and committed again as another. On the failed-drop + failed-fsync path, where the `DROP` a `CANCEL` undoes never reached durable storage while the `CANCEL` did, replaying `ADD partA`, `DROP partA`, `ADD partB`, `CANCEL partB` let the `CANCEL` consume the older generation's committed `DROP`: the surviving stream became `ADD partA`, `ADD partB`, and the reconstructed map kept the stale first generation (a repeated insert of a present block ID is a no-op). Dropping the current generation after the restart then no longer covered the block ID, so a legitimate reinsert after that drop was wrongly deduplicated. A `CANCEL` carries the part name of the `DROP` it rolls back, so it is now paired only with a preceding `DROP` of the same block ID *and* the same part name — the exact record generation it was written to undo — and a stray `CANCEL` whose target was lost pairs with nothing. The cancelled-add marker still pairs with an `ADD` by block ID alone (its part name field holds the reserved marker), which is sufficient: a second `ADD` of the same block ID can only be written once the live map no longer holds it, so any older surviving `ADD` is followed by a surviving `DROP` that erases it on replay regardless of the pairing. Additionally (a nineteenth review finding, two recovery-path issues). First, a canceled log writer was only healed while rolling back the operation that canceled it: if that rollback's rotation failed too (the disk still down), the writer stayed canceled, and the first retry after the disk recovered still failed with `Cannot write to canceled buffer` before reaching any recovery code — only that failed retry's own rollback reopened a writer, so the first retry always failed in the double-fault case. Second, when a file left behind by a failed compaction could neither be removed nor overwritten with an empty file, the failure was logged and forgotten, and the server carried on appending newer committed records to an older, lower-numbered file while a stale higher-numbered snapshot survived on disk — which a later restart would replay last, forgetting the newer committed block IDs. Both `addPart` and `dropPart` now start with a `prepareToWrite` step that runs before anything is written: it heals a canceled writer by rotating to a fresh one up front (so the first retry after recovery succeeds), and it retries neutralizing any such pending file, failing closed — the operation throws, retryably, with nothing written — for as long as one remains on disk, because failing an insert loudly is recoverable while silently deduplicating wrongly after a restart is not. Additionally (a twentieth review finding): that fail-closed barrier was process-local, so it did not survive a restart. `prepareToWrite` refused to write while a stale file was pending neutralization, but a crash — or a clean shutdown — before a retry neutralized it lost that knowledge, and the next `load` replayed the stale files as ordinary history, which can resurrect evicted block IDs or rebuild a diverged eviction order and silently deduplicate wrongly later. The barrier is now persisted as an on-disk marker, written durably before a compaction starts and cleared (the file removed or, failing that, overwritten empty) only once the on-disk history is provably consistent again — the compaction finished with a clean cleanup, its failure was fully rolled back, or a later retry neutralized every pending file. If `load` finds the marker still active, the previous run died inside that window, so it discards the whole on-disk history and starts afresh instead of replaying it: losing at most one deduplication window of best-effort insert deduplication (a client retry may be accepted again — a visible duplicate) is strictly safer than replaying history that is known to be possibly inconsistent, and unlike a precise per-file recovery it does not depend on reconstructing which of the compaction's steps had completed when the process died. If a file or the marker can neither be removed nor emptied even then, `load` throws, failing closed. The marker needs no format version: it is the deduplication log file with number 0 — a number no real log can ever get — holding a single rollback record that every server version replays as a no-op, so a downgraded server simply treats it as an ordinary, harmless log file. Additionally (a twenty-first review finding): the restart barrier of that marker could still be bypassed when the server restarted with deduplication disabled. A load with `non_replicated_deduplication_window = 0` deliberately leaves the marker alone — nothing is replayed or written while deduplication is disabled — but re-enabling deduplication with `ALTER TABLE ... MODIFY SETTING` then just reopened the newest fenced-off log file for appending, without consulting the marker. New committed records landed on top of exactly the stale history the marker fences off, and the next restart — acting on the still-active marker — discarded them together with the stale ones, wrongly accepting (and duplicating) a retry of an insert committed after the re-enable. Re-enabling deduplication now acts on the marker the same way a load with deduplication enabled would have: it discards the suspect history and clears the marker before anything new is written, so the records committed from then on survive the next restart. Additionally (a twenty-second review finding): changing the deduplication window with `ALTER TABLE ... MODIFY SETTING` is the one path that can rotate the log or reopen a writer without an insert or drop in front, and it bypassed the fail-closed recovery barrier those operations run first. After a failed compaction whose cleanup also failed completely, the stale files awaiting neutralization are known only to the running process (compaction had already unregistered them), so re-enabling deduplication in the same process discarded only the registered files: it cleared the on-disk marker while the stale files kept their content, and its subsequent rotation truncated the oldest of them in place; a restart before the next insert or drop then replayed the surviving stale files as ordinary history with the oldest records missing, forgetting committed block IDs and wrongly accepting — duplicating — their retries. Even while deduplication stayed enabled, the setter's rotation could reclaim the one stale file pending neutralization while the on-disk marker stayed active with nothing left to clear it, so the next restart discarded the records committed after the `ALTER` and wrongly accepted their retries. The setter now runs the same recovery barrier first whenever the new window is non-zero: the stale files are neutralized precisely — without discarding the consistent history — and the marker is cleared only once none remains, failing the `ALTER` closed (retryably) while the disk does not allow it; the discard on the re-enable transition therefore fires only when the marker was left by a previous process, whose precise in-process knowledge is gone. Additionally (a twenty-third review finding): the all-or-nothing insert contract stopped at the deduplication log's own boundary. `MergeTreeSink::commitPart` publishes the block IDs via `addPart` and only then makes the part active (`renameTempPartAndAdd` followed by `transaction.commit`); if one of those later steps threw, the part never became active but its block IDs stayed durably published, so a client retry of the same insert was silently deduplicated against a part that does not exist - both in the same process and after a restart. The sink now unpublishes them on that path with a best-effort `dropPart` (itself all-or-nothing per the earlier findings) before rethrowing; if even the drop fails - for example on the same broken disk that failed the commit - the block IDs stay published, which is no worse than before, and the original error is not masked. A new stateless test injects a failure between the publication and the part commit through the new `merge_tree_sink_fail_part_commit_after_dedup` failpoint and verifies that a retry of the failed insert is inserted rather than deduplicated (the test fails without the rollback). Gtests inject a `writeFile` failure during rotation, a `next()` (flush) failure on the currently open writer, and separately an fsync failure on the previous writer's `sync` while rotating away from it, and verify in all cases that the log remains usable afterwards, that a retry of the failed insert is accepted (and only then deduplicates as usual), and that unrelated, already-committed block IDs are not evicted from memory by the failed insert — both within the same process and, including the eviction check, after reloading the log from disk on a restart. A further gtest fails a four-block insert with an injected `sync` failure and restarts twice, verifying that the log file holding the committed block IDs is not dropped by the retention pass, so the committed inserts still deduplicate after the second restart. A final gtest drops a range covering two block IDs while injecting an fsync failure into the rotation that the first `DROP` triggers, and verifies that the failed drop erases neither block ID, so a retry of either is still deduplicated. Another gtest fails the write of the second of two `DROP` records mid-drop and verifies — both live and after reloading the log on a restart — that neither covered block ID is forgotten, and that a record written after the failed drop is durable (the rollback rotated away from the canceled writer, which would otherwise silently discard it). A final gtest replays the on-disk logs left behind by a rolled-back insert and a rolled-back drop using the exact pre-change replay logic (`DROP` = erase, anything else = insert, part names parsed) and verifies that an older server would not consider the rolled-back insert's block ID published (no silently dropped retry after a downgrade) and would keep both block IDs of a rolled-back drop published. A unit test exercises the new `LimitedOrderedHashMap` primitives directly (an `insertWithoutEviction` keeps every entry, including the oldest, present and looked up correctly while the map temporarily exceeds its limit; a following `trimToMaxSize` evicts the oldest in FIFO order), and a further gtest verifies that a successful insert still evicts the oldest block ID to enforce the deduplication window, so splitting publication into a non-evicting insert and a later trim does not change the observable windowing. A final gtest fails an insert on an fsync error and then commits a single insert into the rollback-only log file, verifying that the file rotates once its raw size (its rollback record plus the new record) reaches the threshold — which it would not if rotation were driven by the surviving-record count, letting a rollback-heavy log grow without bound. A final gtest accumulates many rolled-back inserts on an fsync-failing disk and verifies that a restart compacts them into a single log file while preserving the live deduplication state, rather than retaining every accumulated file. A final gtest repeats that accumulation on a simulated disk without append support and verifies that a restart likewise compacts the accumulated files back down to the snapshot (rather than retaining one file per failure), confirming the bound holds in the every-operation-rotation regime too. A further gtest interrupts a failed insert's rollback partway — injecting an fsync failure into the rotation and then a flush failure after only some of the compensating records are written — and verifies, after a restart, that the older log holding an unrelated committed block ID is not dropped by retention, so that block still deduplicates. A final gtest restarts several times on a disk without append support with no operations in between and verifies that the retained log file count stays bounded (the live-state file plus one fresh writer file) rather than growing by one empty file per restart. A final gtest makes a compaction snapshot durable and then fails both the reopen of the snapshot for appending and its removal during cleanup, verifying that the orphan snapshot is left as an empty file (so it cannot outrank the live state on the next replay) and that a healthy restart still reconstructs the live deduplication state exactly. A further gtest constructs the on-disk state of a failed drop whose `DROP` record was lost to an fsync failure while its `CANCEL` survived, and verifies that the covered block ID still deduplicates after a restart rather than being cancelled out by the stray `CANCEL`. A related gtest repeats that construction with a block ID reused across two part generations and verifies that, after a restart, the stray `CANCEL` does not cancel the older generation's committed `DROP`: the current generation stays in the map, dropping it clears the block ID, and a legitimate reinsert after the drop is accepted. A final gtest accumulates rollback garbage and then restarts on a disk that can rewrite files but cannot unlink them, verifying that compaction empties every pre-snapshot log file (so the snapshot alone determines the reloaded eviction order) and that the live state survives. A further gtest injects a double fault — a record write fails and the rollback's rotation fails as well — and verifies that the first retry after the disk recovers already succeeds (the canceled writer is healed up front) and that the retried insert survives a restart. A final gtest makes a compaction snapshot durable, fails its completion, and defeats the entire cleanup (the orphan can neither be removed nor emptied), verifying that inserts then fail closed while the stale orphan is on disk, that the committed state still deduplicates, that the first insert after the disk recovers neutralizes the orphan and succeeds, and that the state stays exact across a healthy restart. Two further gtests restart the process while a failed compaction's stale files are still on disk — once with the orphan snapshot left behind before the switch-over, once (on a disk that can create but not destroy files) with unremovable pre-snapshot files left behind after it — and verify that the new process discards the history instead of silently replaying the stale files, and that deduplication works normally from the fresh history on. Another gtest restarts with deduplication disabled while the marker is still active, re-enables it through the window-size setter, and verifies that the marker is cleared and the stale history discarded before anything is written, and that an insert committed after the re-enable still deduplicates across a further restart instead of being thrown away with the stale files. A further gtest replays the same-process trace — a failed compaction whose cleanup fails completely, the disk healing, and the window set to 0 and back to a non-zero value in the same process — and verifies that re-enabling neutralizes the pending stale files precisely (only the snapshot remains on disk), and that the committed block still deduplicates after a restart before any further write. (The on-disk record definitions were moved into their own header, `MergeTreeDeduplicationLogRecord.h`, so that the regression tests can also be compiled against the merge-base sources by the `Bugfix validation (unit tests)` CI job; the tests that exercise machinery introduced by this fix are compiled only when that header is present.) CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=a63a425c48d3123d28bef2ad60511c9fc0583468&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a logical error `Cannot write to finalized buffer` (an abort in debug and sanitizer builds) in the non-replicated `MergeTree` deduplication log when log rotation fails, for example after a transient I/O error on the deduplication log's disk. Failed deduplication-log writes, flushes, syncs, and compactions now roll back safely, so retries of failed inserts or part drops are not wrongly deduplicated against parts that never committed, committed block IDs are not forgotten after a restart, and rollback-heavy histories are compacted instead of growing without bound.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110429",
          "createdAt": "2026-07-14T18:08:47Z",
          "updatedAt": "2026-08-13T06:36:43Z",
          "timestamp": "2026-08-13T06:36:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8797f44163db35dd959e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112498",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112498",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed a crash when reading a Parquet file with inconsistent bloom filter metadata. Such files could also silently return fewer rows than they should. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ## Problem Reading a Parquet file whose bloom filter metadata is inconsistent segfaults the server. Seen in production: 25 crashes over 12 days on one instance, all with the same stack, while an hourly job read an Iceberg lake on S3. ``` (version 26.2.1.525 (official build), architecture: aarch64) Received signal 11 Signal description: Segmentation fault Address: 0xfffab740826f. Access: <not available>. Address not mapped to object. 2.0. inlined from base/base/../base/unaligned.h:12: unsigned int unalignedLoad<unsigned int>(void const*) 2. src/Processors/Formats/Impl/Parquet/Reader.cpp:906: DB::Parquet::Reader::BloomFilterLookup::findAnyHash(...) 3. src/Storages/MergeTree/KeyCondition.cpp:780: DB::mayExistOnBloomFilter(...) 4. src/Storages/MergeTree/KeyCondition.cpp:4127: DB::KeyCondition::checkInHyperrectangle(...) 5. src/Processors/Formats/Impl/Parquet/Reader.cpp:935: DB::Parquet::Reader::applyBloomAndDictionaryFilters(RowGroup&) 6. src/Processors/Formats/Impl/Parquet/ReadManager.cpp:117: DB::Parquet::ReadManager::finishRowGroupStage(...) ``` Consequences: - The server process dies, so every query on that instance fails, not just the Parquet one. The affected instance was single-replica, so each crash was a full outage until the pod restarted. - The crashing query retries and crashes again, once per retry. - When the out-of-range bloom filter block happens to land inside the read buffer instead of unmapped memory, there is no crash: the row group is pruned on unrelated bytes and rows go missing silently. `SELECT count() FROM file(...) WHERE s = '123456'` returned `0` for a value that is present. - Reproducible on `master`, and on every version since 25.11, when the v3 Parquet reader became the default (`input_format_parquet_bloom_filter_push_down` has defaulted to `1` since 25.5). ## Root cause `Reader::processBloomFilterHeader` takes the bloom filter bitset size from the file (`BloomFilterHeader.numBytes`, validated only for sign and 32-byte alignment) and derives the byte range of each 32-byte bloom filter block from it, at `bloom_filter_offset + header_size + block_idx * 32`. It never checks that `header_size + numBytes` fits inside the byte range it registered for the bloom filter — `ColumnMetaData.bloom_filter_length` when the file declares it, otherwise the \"next known offset\" upper bound computed in `initializePrefetches`. `Prefetcher::splitRange` was the only guard, and it missed the case in two independent ways: - `if (start < range.start || length > range.end - start)` **underflows**: when `start > range.end`, `range.end - start` wraps around, so the comparison is false and a subrange past the end of the range passes. - The check runs only while the parent range is still in state `HasRange`. The bloom filter header range (registered with `likely_to_be_used = true`) and the bloom filter data range start at the same file offset, and the data range is normally smaller than `min_bytes_for_seek`, so starting the header prefetch coalesces the data range into the same read task and flips it to `HasTask`. `splitRange`'s tail path then computes `req->task_offset = subranges[i].first - task->offset` with no validation at all. This is the path a real S3 read takes. `Prefetcher::getRangeData` guarded the resulting span with `chassert` only, which is compiled out in release builds, so it returned a `std::span` pointing outside `task->buf`, and `findAnyHash`'s `unalignedLoad<UInt32>` read unmapped memory. ## Fix Each of the three layers now fails closed: 1. `Reader::processBloomFilterHeader` rejects a header whose `header_size + numBytes` exceeds the bloom filter extent the file declared, with `INCORRECT_DATA` naming `input_format_parquet_bloom_filter_push_down=0` as the escape hatch — the same shape as the two bloom filter validation errors already there. The extent is remembered in the new `ColumnChunk::bloom_filter_data_bytes`, set at both `registerRange` call sites. 2. `Prefetcher::splitRange` makes the `HasRange` check underflow-safe and applies the equivalent check against the read task's byte range on the coalesced `HasTask` path, before touching refcount or `RequestState`s so throwing stays clean. 3. `Prefetcher::getRangeData` turns the buffer-bounds `chassert` into a real check against `task->length` (the invariant that holds for both the `buf` and the zero-copy `cached_region` paths), so a bookkeeping mistake anywhere surfaces as an error instead of an out-of-bounds read. Regression test: `04654_parquet_bloom_filter_bitset_out_of_bounds` reads a 1649-byte fixture whose `s` column `BloomFilterHeader` claims a 1 GiB bitset while its column metadata declares 272 bytes of bloom filter data; the read must report `INCORRECT_DATA`, and the same file still reads correctly with push-down off. The fixture is small on purpose, so its bloom filter stays under the default seek threshold and the test drives the coalesced path that production takes. Each of the three layers was also removed on its own and the test re-run, confirming none of them is dead code: without layer 1 the request is rejected by layer 2 (`Subrange out of bounds: [460964200, 460964232) not in read task [395, 1352)`), without layers 1 and 2 by layer 3, and with all three removed the reader aborts on the buffer-bounds assertion in a Debug build. Note that a file whose `bloom_filter_length` excludes the serialized header (`parquet.thrift` specifies that it includes it) is now rejected rather than read. Those reads were already either crashing or silently over-pruning; `input_format_parquet_bloom_filter_push_down=0` reads such files without their bloom filters. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1303` (included in `26.8` and later) - Backported to: `26.7.4.25` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112498",
          "createdAt": "2026-07-30T00:31:49Z",
          "updatedAt": "2026-08-13T06:35:36Z",
          "timestamp": "2026-08-13T06:35:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "tiandiwonder",
          "state": "closed",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4886901e9c3ebf91cf21",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100185",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100185",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add assertions in IColumn::mutate to verify deep unique ownership of sub-columns",
          "text": "Investigation for https://github.com/ClickHouse/ClickHouse/issues/99920 This is a draft PR with diagnostic assertions to help find a COW violation bug where columns stored in `HashJoin` have their `allocatedBytes()` change because internal sub-columns are modified by someone who treats immutable columns as mutable. ## What was found The bug manifests as `data->allocated_size != debug_allocated_size` exception in debug builds. Investigation narrowed it down to: 1. **Removing the single-chunk optimization** in `Squashing.cpp:70-72` fixes the bug — this optimization passes chunks through without `IColumn::mutate`, so columns remain shared with Memory engine blocks 2. The **modification happens between** `addBlockToJoin` calls — during pipeline processing of the next chunk 3. The mutation is a **reallocation** (reserve/grow), not a data write — `protect()` (mprotect) doesn't catch it 4. The assertions in this PR (`IColumn::mutate` deep ownership check) **do NOT fire** — meaning the mutation happens outside the `IColumn::mutate` path entirely. Someone treats an immutable column as mutable without going through `mutate`. ## Reproducer The test `04049_fix_hashjoin_shared_column_allocated_size` reproduces the exception in debug builds. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100185",
          "createdAt": "2026-03-20T11:00:24Z",
          "updatedAt": "2026-08-13T06:34:50Z",
          "timestamp": "2026-08-13T06:34:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 21
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1b2d29f3c1615aa88531",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114575",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114575",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112498 to 26.6: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114575",
          "createdAt": "2026-08-13T02:35:48Z",
          "updatedAt": "2026-08-13T06:20:59Z",
          "timestamp": "2026-08-13T06:20:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d8eff095dc50bd974c95",
        "signalId": "github:ClickHouse/ClickHouse:issue:114406",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114406",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Lossy codec (SZ3) on a sorting-key column silently produces mis-sorted parts after a merge (`Sort order of blocks violated` in debug)",
          "text": "🕵 A lossy codec is accepted on a column that is part of the sorting key. Because the loss is applied on every write, the values that come back from a part are not the values the merge sorted, so a merged part can end up out of order. In a debug build the merge aborts with `Sort order of blocks violated`; in a release build nothing is reported and the part - together with its primary key index - is silently mis-sorted. **Reproduction** (release build, `clickhouse local` is enough): ```sql SET allow_experimental_codecs = 1; CREATE TABLE t (i Float64 CODEC(SZ3('ALGO_INTERP_LORENZO', 'REL', 0.01)), f Int64) ENGINE = MergeTree ORDER BY i SETTINGS min_bytes_for_wide_part = 0; INSERT INTO t SELECT number / 80, number FROM numbers(200000); INSERT INTO t SELECT number / 80 + 0.5, number FROM numbers(200000); OPTIMIZE TABLE t FINAL; SELECT countIf(i < prev) AS descending_pairs, min(i - prev) FROM (SELECT i, lagInFrame(i) OVER (ORDER BY _part_offset) AS prev FROM t) WHERE prev > 0; ``` ``` ┌─descending_pairs─┬──────min(minus(i, prev))─┐ │ 1 │ -0.2436889648437699 │ └──────────────────┴──────────────────────────┘ ``` A single `INSERT` produces a sorted part here; the descending pair appears only after the merge. **In CI** this shows up as the AST fuzzer `Logical error: 'Sort order of blocks violated for column number 0, left: Float64_12.415000000000006, right: Float64_12.168837890625007. Chunk 25, rows read 199842.'` in `CheckSortedTransform` during a background merge. The fuzzer reaches it by rewriting a table's sorting-key column to `Float64 CODEC(SZ3('ALGO_INTERP_LORENZO', 'REL', 0.01))`, and the reported neighbours differ by ~2 %, which is the codec's relative error bound rather than any data property. The signature has hit at least 10 CI runs on unrelated pull requests in the last 30 days, including `master`. Suggested direction: reject a lossy codec (`SZ3`, and any future lossy codec) on a column used in the sorting key, primary key or partition key, the same way other structurally unsafe codecs are rejected at DDL time. Silently producing a mis-sorted part is the worst outcome, because index analysis then skips rows that do match. Related: https://github.com/ClickHouse/ClickHouse/issues/111139 Related: https://github.com/ClickHouse/ClickHouse/pull/113575",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114406",
          "createdAt": "2026-08-12T00:56:46Z",
          "updatedAt": "2026-08-13T06:17:45Z",
          "timestamp": "2026-08-13T06:17:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "clickgap-analyzed",
            "culprit-pr-pinned"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a32fef3979e260672969",
        "signalId": "github:ClickHouse/ClickHouse:issue:113182",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113182",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "`optimize_inverse_dictionary_lookup` rewrite inside a correlated EXISTS makes decorrelation fail: 48 \"Cannot decorrelate query, because 'DelayedCreatingSets' step is not supported\"",
          "text": "## Describe the problem A valid query with a correlated `EXISTS` subquery fails with exception 48 (`NOT_IMPLEMENTED`) at pure default settings when the subquery's `WHERE` contains a `dictGet(...) >= <const>` comparison. `optimize_inverse_dictionary_lookup` (default `1`) rewrites the `dictGet('d', 'attr', key) >= c` predicate into `key IN __set_...`. When the predicate sits inside a correlated subquery, the injected set adds a `DelayedCreatingSets` step to the subquery plan, and the correlated-subquery decorrelation then refuses the plan: ``` Code: 48. DB::Exception: Cannot decorrelate query, because 'DelayedCreatingSets' step is not supported. (NOT_IMPLEMENTED) ``` Both settings involved are on by default (`optimize_inverse_dictionary_lookup = 1`, `allow_experimental_correlated_subqueries = 1`), so an ordinary query errors out of the box. Disabling the rewrite (`optimize_inverse_dictionary_lookup = 0`) makes the same query run fine and return the correct result. This is the only setting that matters: `correlated_subqueries_use_in_memory_buffer = 0` does not help, and the same failure occurs when the `EXISTS` is used as a scalar (e.g. `o.g >= (EXISTS ...)`), not just in `WHERE`. ## How to reproduce Version: `26.8.1.653` (public master build; also reproduced on `26.8.1.561`). All settings at defaults. ```sql CREATE TABLE i_src (id UInt64, val UInt32) ENGINE=MergeTree ORDER BY id; INSERT INTO i_src SELECT number, number % 97 FROM numbers(500); CREATE DICTIONARY i_dict (id UInt64, val UInt32 DEFAULT 0) PRIMARY KEY id SOURCE(CLICKHOUSE(TABLE 'i_src' DB 'default')) LIFETIME(0) LAYOUT(FLAT()); CREATE TABLE i_t (k UInt32, g Int32, s String) ENGINE=MergeTree ORDER BY k; INSERT INTO i_t SELECT number, number % 7, toString(number % 5) FROM numbers(100); SELECT o.k FROM i_t AS o WHERE EXISTS ( SELECT 1 FROM i_t AS i WHERE i.s = o.s AND dictGet('i_dict', 'val', toUInt64(abs(g)) % 500) >= 76); ``` Observed (deterministic, 20/20 runs): ``` Code: 48. DB::Exception: Cannot decorrelate query, because 'DelayedCreatingSets' step is not supported. (NOT_IMPLEMENTED) ``` Expected: the query executes and returns the rows for which the `EXISTS` holds (here an empty result — confirmed by running with `SETTINGS optimize_inverse_dictionary_lookup = 0`, which succeeds). Settings analysis: - REQUIRED: none beyond defaults. `optimize_inverse_dictionary_lookup = 0` is the only toggle that avoids the error. - Incidental: `correlated_subqueries_use_in_memory_buffer` (both values fail), the exact comparison constant, table sizes, dictionary layout. ## Additional context Mechanism: `InverseDictionaryLookupPass` rewrites the `dictGet` comparison into a membership test against an internally built set (`__set_...`). Building that set inside the correlated subquery introduces a `DelayedCreatingSets` plan step, and the decorrelation of correlated subqueries does not support that step, so planning aborts with 48. Either the pass should be skipped inside correlated subqueries it would break, or the decorrelator should learn to handle `DelayedCreatingSets`. Same \"optimization-pass output escapes into an unsupported context\" family as the `ColumnSet`-reaches-`FINAL`-merge manifestation of the same pass, but a distinct trigger and error code. Related: https://github.com/ClickHouse/ClickHouse/issues/112030 Related: https://github.com/ClickHouse/ClickHouse/issues/99500",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113182",
          "createdAt": "2026-08-03T20:23:51Z",
          "updatedAt": "2026-08-13T06:17:33Z",
          "timestamp": "2026-08-13T06:17:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "bug",
            "comp-query-optimizer",
            "comp-query-analyzer",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a63bd98332e125532a17",
        "signalId": "github:ClickHouse/ClickHouse:issue:112908",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112908",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Constant projection + ORDER BY ALL + LIMIT n BY + LIMIT WITH TIES: exception 10 \"Not found column 1_UInt8 in block\" at default settings",
          "text": "A valid table-free query combining a constant projection, `ORDER BY ALL`, `LIMIT n BY`, and `LIMIT ... WITH TIES` throws an exception at **default settings**: ```sql SELECT 1 FROM numbers(10) ORDER BY ALL LIMIT 1 BY number LIMIT 4 WITH TIES ``` ``` Code: 10. DB::Exception: Not found column 1_UInt8 in block. There are only columns: __table1.number. (NOT_FOUND_COLUMN_IN_BLOCK) ``` All four ingredients are required — removing any one produces the correct result on 26.8.1.561: - without `WITH TIES`: `SELECT 1 ... LIMIT 1 BY number LIMIT 4` → OK - non-constant projection: `SELECT number ... WITH TIES` → OK - explicit key instead of `ORDER BY ALL`: `SELECT 1 ... ORDER BY number ... WITH TIES` → OK - without `LIMIT 1 BY number` → OK Deterministic (20/20), no settings involved, no tables involved. `WITH TIES` needs the sort-description columns in the stream to compare ties; the constant projection under `ORDER BY ALL` + `LIMIT BY` apparently drops the constant from the block header the tie-comparison expects. Found while fuzzing plan-shape combinations; the same shape fails identically through `Distributed` and parallel-replica reads (the local exception is the root).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112908",
          "createdAt": "2026-08-01T14:44:00Z",
          "updatedAt": "2026-08-13T06:17:28Z",
          "timestamp": "2026-08-13T06:17:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "comp-query-analyzer",
            "comp-query-execution",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "KochetovNicolai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:91d3b8b71ed1e56f2d7e",
        "signalId": "github:ClickHouse/ClickHouse:issue:112903",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112903",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "serialize_query_plan = 1: WITH ROLLUP / WITH CUBE / Join-engine lookup join in a distributed subquery fails with NOT_IMPLEMENTED \"Method serialize is not implemented\" — no fallback",
          "text": "`serialize_query_plan = 1` sends the shard-side fragment as a serialized query plan. Three query-plan steps have no `serialize` implementation, and there is no fallback to the text path — any query whose shipped fragment contains one of them fails with `NOT_IMPLEMENTED` at execution time, while the same query succeeds with `serialize_query_plan = 0`: - `Rollup` (`GROUP BY ... WITH ROLLUP` inside the distributed subquery) - `Cube` (`GROUP BY ... WITH CUBE`) - `JoinStepLogicalLookup` (a lookup join against a `Join`-engine table inside the distributed subquery) Notably `GROUP BY GROUPING SETS ((h), ())` — which subsumes `ROLLUP` semantically — serializes fine, as does `WITH TOTALS`; the holes are specifically the `Rollup`/`Cube` transform steps and the lookup-join step. **How to reproduce** On a server with any single-shard cluster whose replica is the server itself (e.g. `test_shard_localhost` from the standard test configs), version 26.8.1.561: ```sql CREATE TABLE t_r (h UInt32) ENGINE = MergeTree ORDER BY h; INSERT INTO t_r SELECT number % 10 FROM numbers(1000); CREATE TABLE dist_t_r AS t_r ENGINE = Distributed(test_shard_localhost, currentDatabase(), t_r); -- OK: returns 11 SELECT count() FROM (SELECT sum(h) FROM dist_t_r GROUP BY h WITH ROLLUP) SETTINGS serialize_query_plan = 0, prefer_localhost_replica = 0; -- Code: 48. DB::Exception: Method serialize is not implemented for Rollup: While executing Remote. (NOT_IMPLEMENTED) SELECT count() FROM (SELECT sum(h) FROM dist_t_r GROUP BY h WITH ROLLUP) SETTINGS serialize_query_plan = 1, prefer_localhost_replica = 0; -- Code: 48 ... Method serialize is not implemented for Cube ... SELECT count() FROM (SELECT sum(h) FROM dist_t_r GROUP BY h WITH CUBE) SETTINGS serialize_query_plan = 1, prefer_localhost_replica = 0; -- lookup-join variant: CREATE TABLE j_r (h UInt32, name String) ENGINE = Join(ANY, LEFT, h); INSERT INTO j_r SELECT number % 10, concat('n', toString(number % 10)) FROM numbers(10); -- OK with serialize_query_plan = 0; with 1: -- Code: 48 ... Method serialize is not implemented for JoinStepLogicalLookup ... SELECT count() FROM (SELECT t.h, j.name FROM dist_t_r AS t ANY LEFT JOIN j_r AS j USING (h)) SETTINGS serialize_query_plan = 1, prefer_localhost_replica = 0; ``` Deterministic: 20/20 failures for both the `Rollup` and the `JoinStepLogicalLookup` shape; 20/20 success with `serialize_query_plan = 0`. Required settings: `serialize_query_plan = 1` plus a genuinely remote read (`prefer_localhost_replica = 0` here — with the local-replica shortcut the plan is never serialized, which is why a plain local read hides the bug). Everything else is at defaults. **Expected**: either these steps get a `serialize` implementation, or `serialize_query_plan` falls back to sending the fragment as text (the way `make_distributed_plan` refuses gracefully with a code 344 `not serializable for remote execution` *before* execution). Failing a valid query at execution time on a plain production `Bool` setting is neither. Related: https://github.com/ClickHouse/ClickHouse/issues/112167 Related: https://github.com/ClickHouse/ClickHouse/issues/112079",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112903",
          "createdAt": "2026-08-01T13:37:49Z",
          "updatedAt": "2026-08-13T06:17:24Z",
          "timestamp": "2026-08-13T06:17:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "bug",
            "comp-query-optimizer",
            "comp-distributed",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [
            "yakov-olkhovskiy"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e0612dabffe5fc6afe84",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86768",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86768",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Feature: Enable overlay databases for server.",
          "text": "Enables server-side `Overlay` databases. An `Overlay` database is a read-only facade that exposes the union of the tables of several underlying databases, resolving each table name through the listed sources in order (the first source that has the table wins). DDL on the facade is rejected — it has no storage of its own — while `SELECT` and pass-through `INSERT` resolve to the underlying source table. Reading or writing through the facade requires the corresponding grant on both the facade database and the underlying source, and the facade's row policies are combined with the source's. Previously an `Overlay` database existed only as the implicit default database of `clickhouse-local`. Closes: https://github.com/ClickHouse/ClickHouse/issues/52764 <!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow creating `Overlay` databases. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86768",
          "createdAt": "2025-09-05T21:38:13Z",
          "updatedAt": "2026-08-13T06:17:06Z",
          "timestamp": "2026-08-13T06:17:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 107
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested",
            "pr-autogenerated-docs"
          ],
          "author": "AlyHKafoury",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e751fd745781e54c6976",
        "signalId": "github:ClickHouse/ClickHouse:issue:110281",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:110281",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Join-order cardinality estimation runs before MergeTree partition/PK analysis and ignores pruned parts",
          "text": "## Describe the unexpected behaviour When column statistics are enabled (`use_statistics = 1`, the default), the join-order optimizer estimates relation cardinalities over **all** active parts of a MergeTree table, ignoring partition/PK pruning that execution will later perform. A query whose `WHERE` prunes to a 1,000-row partition of a 5,001,000-row table is planned as if it read ~2,500,500 rows. The error is identical under `use_statistics_cache = 0` and `= 1` — this is not a stale-cache problem; the cache-miss path lands on the same all-parts set. There is also an asymmetry: with statistics *disabled*, `estimateReadRowsCount` falls back to `ReadFromMergeTree::selectRangesToRead()` and gets pruning-aware `selected_rows`. Measured on the same query: | setting | fact-side estimate | best plan cost | |---|---|---| | `use_statistics = 1` (default) | 2,500,500 rows | 100000 | | `use_statistics = 0` | 1,000 rows (exact) | 1000 | So enabling column statistics currently makes the plan 100× worse by the optimizer's own cost model on partition-pruned queries. ## How to reproduce Verified on the official release `26.7.1.448` @ `cff54e151b69` (macOS arm64) and on `master` @ `5a9528b4db5` (2026-07-13, macOS arm64 debug build) — identical estimates on both. Probes were run on a freshly started server, twice per cache setting (bit-identical), with `collect_hash_table_stats_during_joins = 0` to exclude the runtime-feedback loop. ```sql CREATE TABLE fact (p UInt8, id UInt64) ENGINE = MergeTree PARTITION BY p ORDER BY id SETTINGS refresh_statistics_interval = 1; -- fast background stats refresh CREATE TABLE dim (id UInt64) ENGINE = MergeTree ORDER BY id SETTINGS refresh_statistics_interval = 1; SET materialize_statistics_on_insert = 1; INSERT INTO fact SELECT 1, number FROM numbers(5000000); -- NDV(id) = 5M INSERT INTO fact SELECT 2, number % 10 FROM numbers(1000); -- NDV(id) = 10 INSERT INTO dim SELECT number FROM numbers(100000); -- wait a few seconds for the background statistics refresh, then: SELECT count() FROM fact AS f INNER JOIN dim AS d ON f.id = d.id WHERE f.p = 2 SETTINGS use_statistics_cache = 0, collect_hash_table_stats_during_joins = 0; -- run again with use_statistics_cache = 1 (same estimates), -- and with use_statistics = 0 (pruning-aware fallback, see table above) ``` Observe the `optimizeJoin` trace (`--send_logs_level=trace`): ```text -- use_statistics_cache = 0 (identical under = 1, minus the Loading lines): a1.fact: Loading statistics optimizeJoin: estimate statistics fact: 2500500 rows, columns: [id: 2500500, p: 2] optimizeJoin: Estimated statistics for Filter f: 2500500 rows, columns: [__table1.p: 2, __table1.id: 2500500] a1.dim: Loading statistics optimizeJoin: estimate statistics dim: 100000 rows, columns: [id: 100315] JoinOrderOptimizer: Optimized join order in 0.18 ms, best plan cost: 100000, estimated cardinality: 100000 ``` The fact side is estimated at 2,500,500 rows — exactly 5,001,000 × 1/NDV(p) = 1/2, the all-parts selectivity model. A pruning-aware estimate could not exceed 1,000. NDV(id) is estimated at ~2,500,500 vs 10 actual in the surviving partition (250,000×). Estimates are bit-identical for `use_statistics_cache = 0` and `1` across repeated runs; `collect_hash_table_stats_during_joins = 0` excludes the runtime-feedback loop. Meanwhile execution does prune — `EXPLAIN indexes = 1` on the same query: ```text └──Join (JOIN FillRightFirst) │ f[2500500] ⋈ d[100000] <-- planner: all-parts estimate │ ... ├──ReadFromMergeTree (a1.fact) │ Parts: 1 | Granules: 1 <-- execution: pruned to the p=2 part │ Indexes: │ Min-Max │ Condition: (p in [2, 2]) │ Parts: 1/6 │ Granules: 1/612 │ Partition │ Condition: (p in [2, 2]) │ Parts: 1/1 ``` The same `EXPLAIN` output contains both the pruning-blind estimate on the Join node and the pruned read on the leaf. ## Plan damage (counterfactual) Same data restricted to p=2 (`CREATE TABLE fact_p2 ...; INSERT ... WHERE p=2`), same join without the `WHERE`: ```text optimizeJoin: estimate statistics fact_p2: 1000 rows, columns: [id: 10] optimizeJoin: estimate statistics dim: 100000 rows, columns: [id: 100315] JoinOrderOptimizer: Optimized join order in 0.01 ms, best plan cost: 996.86, estimated cardinality: 996 ``` versus `best plan cost: 100000, estimated cardinality: 100000` for the identical rows behind the pruned `WHERE p = 2`. The chosen build side flips from `dim` (100,000 rows) to the small fact side (1,000 rows), and the optimizer's own cost drops ~100× — i.e. today the pruning-blind estimate makes the optimizer reject the plan it would itself prefer given correctly-scoped cardinalities (the `use_statistics = 0` probe above confirms this on the original table, not just the counterfactual). On real partitioned fact tables (time-series with a sharp `WHERE` on the partition key) the same mechanism inflates every join-order decision. ## Root cause (hypothesis) At the time `optimizeJoin` requests the estimator, `ReadFromMergeTree::analyzed_result_ptr` appears to be unset, so `getParts()` falls back to `prepared_parts` (all active parts) (src/Processors/QueryPlan/ReadFromMergeTree.h, `getParts()`); `MergeTreeData::getConditionSelectivityEstimator` then folds statistics over that set. The statistics branch of `estimateReadRowsCount` (src/Processors/QueryPlan/Optimizations/optimizeJoin.cpp) returns before the `selectRangesToRead()` fallback that would have run the index analysis. `use_statistics_cache` is irrelevant to the scope: the background `refreshStatistics()` task also folds over all active parts, so cache hit and miss agree — both wrong-scoped. It is possible the current pass ordering is deliberate (index analysis is expensive, and join reordering changes which filters reach which table) — if so, this issue is a request to document that trade-off and close the gap, not a claim of an oversight. A fix presumably needs pass-ordering (make partition/PK analysis results available to join planning) plus composing statistics over surviving parts only. Note that once that happens, the table-wide `cached_estimator` *would* diverge from the pruned scope and needs a part-set-aware design (e.g. cache decoded per-part statistics and fold per query). ## Related - #55065 (column statistics umbrella) - #78441 (wrong build-side swap with `WHERE`) - #85126 (unknown `LIKE` selectivity) - #62870 (physical partition pruning through JOIN condition — different: that one is about execution, this one is about estimation scope)",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/110281",
          "createdAt": "2026-07-13T14:42:22Z",
          "updatedAt": "2026-08-13T06:16:49Z",
          "timestamp": "2026-08-13T06:16:49Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "potential bug"
          ],
          "author": "skuznetsov-clickhouse",
          "state": "closed",
          "assignees": [
            "skuznetsov-clickhouse"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9927809d732682d17eb0",
        "signalId": "github:ClickHouse/ClickHouse:issue:88957",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:88957",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Unexpected syntax error when using `paste join` with `IcebergS3Cluster`",
          "text": "### Company or project name _No response_ ### Describe the unexpected behaviour ```sql SELECT * FROM icebergS3Cluster('replicated_cluster', 'http://minio:9000/warehouse/data2', 'admin', 'password') PASTE JOIN ( SELECT * FROM iceberg('http://minio:9000/warehouse/data3', 'admin', 'password') ) AS t1 ORDER BY tuple(*) ASC FORMAT Values ``` ``` Elapsed: 0.016 sec. Received exception from server (version 25.8.10): Code: 62. DB::Exception: Received from localhost:9000. DB::Exception: Received from clickhouse1:9000. DB::Exception: You must not specify ANY or ALL for PASTE JOIN.. (SYNTAX_ERROR) ``` ### Which ClickHouse versions are affected? 25.8 ### How to reproduce Run query ```sql SELECT * FROM icebergS3Cluster('replicated_cluster', 'http://minio:9000/warehouse/data2', 'admin', 'password') PASTE JOIN ( SELECT * FROM iceberg('http://minio:9000/warehouse/data3', 'admin', 'password') ) AS t1 ORDER BY tuple(*) ASC FORMAT Values ``` If using this syntax, it will work ```sql SELECT * FROM ( SELECT * FROM icebergS3Cluster('replicated_cluster', 'http://minio:9000/warehouse/data2', 'admin', 'password') ) AS t1 PASTE JOIN ( SELECT * FROM iceberg('http://minio:9000/warehouse/data3', 'admin', 'password') ) AS t2 ORDER BY tuple(*) ASC FORMAT Values ``` But I saw in tests, that both types of query are supported. ### Expected behavior Correct result of Paste Join. ### Error message and/or stacktrace _No response_ ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/88957",
          "createdAt": "2025-10-24T13:17:11Z",
          "updatedAt": "2026-08-13T06:15:25Z",
          "timestamp": "2026-08-13T06:15:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "unexpected behaviour",
            "comp-datalake",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "alsugiliazova",
          "state": "open",
          "assignees": [
            "thevar1able"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fc77203d68aa7d5834ee",
        "signalId": "github:ClickHouse/ClickHouse:issue:80759",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:80759",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "TCP server threads metrics (i.e. TCPThreads) is broken for <protocols> (over <tcp_port>/...)",
          "text": "In case of <protocols> is used over <tcp_port> (e.t.c.), the name of the server is different and the code in AsynchronouseMetrics.cpp fails to match them. Suggestion - add proper ServerType::Type enum into the ProtocolServerAdapter, and use it over matching by \"port name\".",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/80759",
          "createdAt": "2025-05-23T21:31:37Z",
          "updatedAt": "2026-08-13T06:15:14Z",
          "timestamp": "2026-08-13T06:15:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "easy task",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "azat",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b3db76fa7ab7b9aa91fc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110784",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110784",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix memory tracker leak when a parent tracker throws MEMORY_LIMIT_EXCEEDED",
          "text": "`MemoryTracker::allocImpl` increments `amount` optimistically at each level and then recurses into the parent tracker. When some ancestor (typically the total server tracker) threw `MEMORY_LIMIT_EXCEEDED`, only the throwing tracker reverted its own counter. All descendant trackers that had already incremented (thread, query, user) kept the rejected amount forever, together with the already-applied `CurrentMetrics` updates. Under sustained memory pressure this drifts `MemoryTracking` and the query/user accounting upward and can produce spurious `MEMORY_LIMIT_EXCEEDED` errors for queries that use little memory. Now every level undoes its own increment when the allocation fails anywhere in the chain, and the side effects (peak update, memory profiler trace, metric update) are committed only after the whole parent chain has accepted the allocation. As a consequence, rejected allocations no longer emit `TraceType::Memory` profiler events, no longer advance the profiler step, and no longer raise `peak_memory_usage`. Concurrent speculative charges may still affect these values. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix an accounting leak in memory tracking: when an allocation was rejected by a parent memory tracker (for example, the server-wide memory limit), the thread, query, and user level trackers kept the rejected amount. The error accumulated over time, inflating `MemoryTracking` and per-query/per-user memory usage, and could lead to spurious `MEMORY_LIMIT_EXCEEDED` errors.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110784",
          "createdAt": "2026-07-17T06:57:18Z",
          "updatedAt": "2026-08-13T06:11:12Z",
          "timestamp": "2026-08-13T06:11:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "seva-potapov",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3d65929f8010b6450f8b",
        "signalId": "github:ClickHouse/ClickHouse:issue:114598",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114598",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "PREWHERE Map subcolumn costing performs full-table metadata I/O before pruning",
          "text": "### Company or project name _No response_ ### Describe the situation After upgrading from 26.6.1 to 26.7, queries using Map subcolumns on a MergeTree became very slow on their first execution. This appears related to [#110623](https://github.com/ClickHouse/ClickHouse/pull/110623). PREWHERE costing now aggregates requested subcolumn sizes across every active part before partition/primary-key pruning. This causes sequential full-table-scale metadata work. Local disks incur per-part prefix/file reads; with an S3-backed MergeTree, these become remote requests. One example query touching S3-backed MergeTree: Profile event | First run | Second run -- | -- | -- QueryPlanBuildMicroseconds | 48,590,413 | 69,982 QueryPlanOptimizeMicroseconds | 48,573,662 | 39,756 SharedPartsLockHoldMicroseconds | 48,561,615 | 28,923 S3ReadMicroseconds | 46,118,276 | 124,920 ReadBufferFromS3InitMicroseconds | 46,358,270 | 0 S3ReadRequestsCount | 1,270 | 11 During the first run, query progress remains at zero for most of the time because this occurs during planning. The identical second run is fast because the per-part subcolumn-size metadata is cached. cc @Avogar ### Which ClickHouse versions are affected? 26.7 ### How to reproduce _ ### Expected performance _No response_ ### Related issues and pull requests _No response_ ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114598",
          "createdAt": "2026-08-13T06:06:47Z",
          "updatedAt": "2026-08-13T06:07:33Z",
          "timestamp": "2026-08-13T06:07:33Z",
          "metrics": {
            "reactions": 1,
            "comments": 0
          },
          "labels": [
            "performance"
          ],
          "author": "starpact",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:be1c654e6e126c7ea4a6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114525",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114525",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Optimize merges of the text index",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved performance of merges of text indexes. ### Additional context A few optimizations: - The main one: reducing overhead on deserialization of embedded and small postings caused by the allocation of the bitmap - Removed unneeded conversion to roaring bitmap on build of the output posting list - Used specialized sort cursor and batch sorting strategy for merging of text index segments",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114525",
          "createdAt": "2026-08-12T17:15:12Z",
          "updatedAt": "2026-08-13T06:03:59Z",
          "timestamp": "2026-08-13T06:03:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4d02003fda7c631a98e2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112688",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112688",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Push subcolumn reads into subqueries",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/75538 Related: https://github.com/ClickHouse/ClickHouse/issues/92455 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Read only the requested subcolumns of columns exported by subqueries, CTEs and views instead of the whole columns. For example, `SELECT data.a FROM (SELECT * FROM table)` with a `JSON` column `data` now reads only the subcolumn `data.a` from the table. Controlled by the new setting `optimize_push_subcolumns_into_subqueries` (enabled by default). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) --- When a subcolumn of a column exported by a subquery is requested, it is resolved into the `getSubcolumn` function over the whole column: `SELECT data.a FROM (SELECT * FROM test)` reads the whole `data` column from the table and extracts the subcolumn afterwards. The new query tree pass `PushSubcolumnsIntoSubqueries` adds the subcolumn to the subquery projection and replaces the `getSubcolumn` function with a reference to it. If the whole column is not used anywhere else, it is then removed from the subquery projection by the subsequent `RemoveUnusedProjectionColumns` pass, so only the subcolumn is read from the table. The pushdown works through several levels of subqueries, CTEs and views, and also applies to `Tuple`, `Nullable`, `Array` and other types with subcolumns. `UNION ALL` subqueries are supported: the subcolumn is added to every branch at the same position, all-or-nothing, and only when the branch types match exactly (the subcolumn of the least supertype is not guaranteed to be the supertype of the branch subcolumns). The `DISTINCT`, `INTERSECT` and `EXCEPT` modes deduplicate or match rows over all projection columns, so such subqueries are not rewritten (nor are recursive CTEs). The pushdown is skipped when it could change the query result: - when the subquery uses `DISTINCT`, `GROUP BY`, aggregate functions, or `ORDER BY ... WITH FILL` (an added projection column would change the result or would not be valid); - when the subquery is on a side of a `JOIN` that can be filled with default values for non-matched rows (`getSubcolumn` of a default value is not always equal to the default value of the subcolumn type, e.g. for the `null` subcolumn of `Nullable` columns); a shared subquery node (an ordinary CTE referenced several times) found in any such position is ineligible for all of its occurrences; - in the outer query with aggregation, an occurrence of `getSubcolumn` is only replaced when it is evaluated before the aggregation step (`WHERE`, `JOIN ON`, arguments of aggregate functions) or when the whole expression is an aggregation key; - when the column types diverge, e.g. under `join_use_nulls` or `group_by_use_nulls`; - when the whole column is also read in the outer query, including uses by correlated subqueries and uses under a different exported name of the same physical column (`SELECT tup AS x, tup FROM t`, or a trivial `ALIAS` storage column next to its base column): the subquery would then read both the whole column and the subcolumn from the table, while extracting the subcolumn from the already read column is cheaper. For a `UNION ALL` target the alias equivalence of exported names is detected in every branch (the pushdown is applied to all branches or to none), so a single branch exporting the same physical column under another name that stays alive blocks the rewrite. The decision counts only sibling subcolumns that are validated as actually pushable into the target (in a side-effect-free dry run): when two subcolumns of the same exported column are requested and only one of them can be pushed (e.g. the other is shadowed by a same-named storage column inside the subquery), nothing is pushed, since the unpushable sibling keeps the whole column alive. - for references of a reused `MATERIALIZED` CTE (`enable_materialized_cte`): the temporary table serves all references of the CTE, so pruning the parent column there would require proving that no reference needs the whole column. Single-use materialized CTEs are inlined by the analyzer and take the ordinary subquery path (subcolumn reads over materialized CTEs are currently broken on master independently of this change, see https://github.com/ClickHouse/ClickHouse/issues/113623). Subqueries spelled with an alias list, e.g. `SELECT x.a FROM (SELECT json FROM t) AS s(x)`, are supported. Trivial `ALIAS` columns (an `ALIAS` whose body is just another column of the same table, possibly chained) exported by a subquery are followed down to the underlying storage column. A subquery can also export a derived subcolumn (e.g. `SELECT json.a AS x FROM ...` over a deeper subquery keeps a `getSubcolumn(json, 'a')` projection expression); reading a subcolumn of such an export composes the paths (`a` + `b` -> `a.b`), so the pushdown continues through derived exports down to the base table. Conversely, when no name of such an alias-equivalent class of exports remains referenced in the outer queries, the never-referenced sibling exports are dead together with the replaced ones (all of them are removed by `RemoveUnusedProjectionColumns`), so they do not block the pushdown through the deeper levels of subqueries; the alias-equivalent classes are taken from every branch of a `UNION ALL` target.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112688",
          "createdAt": "2026-07-30T23:44:26Z",
          "updatedAt": "2026-08-13T05:56:00Z",
          "timestamp": "2026-08-13T05:56:00Z",
          "metrics": {
            "reactions": 1,
            "comments": 15
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3bc054cad36d0e2869d1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111932",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111932",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support the NetCDF format",
          "text": "Adds support for the [NetCDF](https://www.unidata.ucar.edu/software/netcdf/) format, a self-describing binary format for multidimensional arrays that is the standard way climate, weather, oceanographic and other scientific data is distributed. There is a lot of public data in it (ERA5, CMIP, NOAA, Copernicus), and until now the only way to query it with ClickHouse was to convert it first. Reading is supported for the three \"classic\" versions of the format: CDF-1, CDF-2 (64-bit offset) and CDF-5 (64-bit data), and writing produces CDF-2 or CDF-5, whichever the data needs. The files are parsed directly, so there is no new dependency. A NetCDF-4 file, which is an HDF5 file with a different data model on top, is recognized and reported with a message that says how to convert it. **Data model.** Every variable of a file becomes a column, and the rows enumerate the Cartesian product of all the dimensions that the variables use; a variable that does not use some of them is repeated along them. The `to_dataframe` method of `xarray` produces a table with the same columns and the same set of rows, though possibly in a different order: `to_dataframe` puts the dimensions in alphabetical order by default, while ClickHouse keeps the order of the dimensions of the variables. So a file with the dimensions `time`, `lat`, `lon` and the variables `time(time)`, `lat(lat)`, `lon(lon)`, `temperature(time, lat, lon)` reads as a table with four columns and `time * lat * lon` rows. The classic format has no string type, so a `char` variable is read as a String whose length is the last dimension of the variable, when that dimension serves only as the length of the strings; a dimension that anything else in the file uses as a real axis stays in the row space, and such a `char` variable is read as one character per row. Only the variables that a query needs are read, the number of rows comes from the header (so `count()` does not read any data), and an input that cannot be seeked is read into memory instead. **Writing.** Every column becomes a variable over a single dimension named `row`, so a file written by ClickHouse is read back with the same column names and the same rows; the types come back as the closest types of the classic format (a `FixedString` as a `String`, an `Enum` or a `LowCardinality` column as the type it wraps, dates and times as plain numbers, a `Nullable` column as its base type unless read with `input_format_netcdf_fill_value_as_null`). The version of the format is chosen automatically: CDF-5 when a column needs a type that only CDF-5 has or takes more than 4 GiB, CDF-2 otherwise. A `Nullable` column is written with the `_FillValue` attribute, and a column with dates or times gets the `units` attribute of the CF conventions, so `xarray` decodes it back into timestamps. **Settings.** `input_format_netcdf_fill_value_as_null` reads the values equal to the `_FillValue` (or `missing_value`) attribute of a variable as NULL, which is how the CF conventions mark missing data such as sea surface temperature over land. `input_format_netcdf_add_dimension_columns` adds a column with the index along every dimension that has no coordinate variable of the same name, for the files that have none. **Testing.** The test files in `tests/queries/0_stateless/data_netcdf` were written by the netCDF C library, not by this code. Beyond the two functional tests, during development the reader was compared value by value against an independent reference implementation for all three format versions, the output was checked by opening it with the netCDF C library for every supported type family, and 321 truncated and bit-flipped files were fed to the reader — all of them either parsed or failed with a clean error, with no crashes. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `NetCDF` format for both reading and writing. It is a self-describing binary format for multidimensional arrays, widely used for climate, weather and other scientific data. Reading supports all three classic versions of the format (CDF-1, CDF-2 and CDF-5), and writing produces CDF-2 or CDF-5, whichever the data needs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111932",
          "createdAt": "2026-07-26T04:47:02Z",
          "updatedAt": "2026-08-13T05:54:18Z",
          "timestamp": "2026-08-13T05:54:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 25
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:792e071e1dd95d9a75e0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105987",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105987",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Constant filter folding under materialize",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allows to constant-fold filters through materialize wrappers. Closes https://github.com/ClickHouse/ClickHouse/issues/78166#issuecomment-2758343581",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105987",
          "createdAt": "2026-05-27T17:00:57Z",
          "updatedAt": "2026-08-13T05:53:36Z",
          "timestamp": "2026-08-13T05:53:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yariks5s",
          "state": "open",
          "assignees": [
            "vdimir",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9f6ab285ee72f6181f0f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114027",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114027",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Drop stale totals and extremes ports in MergingAggregatedStep",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/issues/113708 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required. The only shape found to reach this logical error needs `inject_random_order_for_select_without_order_by`, documented as only useful for testing and development, so no released configuration is known to be affected. It is fuzzer-reachable (BuzzHouse randomizes it), which is how the abort was found. ### Description Found 2026-08-08 by `AST fuzzer (amd_debug)` (STID 0993-250f, [report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112384&sha=32e19e343e8bbf3cc8bb41f5bdaad2471caa7c05&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)) on PR 112384 as a bystander: `Logical error: Block structure mismatch in function connect between Limit and MergingAggregatedBucketTransform`. It aborts in `Port.cpp:22` `connect` from `MergingAggregatedMemoryEfficientTransform.cpp:577`, so at pipeline construction, not execution. `MergingAggregatedStep::transformPipeline` inherited the pipeline's totals and extremes ports, then attached its merging transforms to them. `addMergingAggregatedMemoryEfficientTransform` uses the `StreamType`-less `Pipe::addSimpleTransform` overload, so the getter also runs on `totals_port`. The transform it creates has an empty input header, while that port still carries the aggregate-state header, hence the mismatch. Where the transform is not attached to those ports, the stale port survives into `TotalsHavingStep`, which requires it to be null. The fix mirrors `AggregatingStep.cpp:398-399`, the other half of two-stage aggregation: drop the current totals and extremes, which this step invalidates and are recalculated afterwards. One call before the branch point covers all three merge branches, the extremes stream and all five construction sites. The other caller needs no drop: it builds pipes from `SourceFromNativeStream`, so neither port exists. Reachability decides the category: two `Merge` children must reach the step still carrying a totals port, and the setting named above is the only producer found among candidates measured with an instrumented `hasTotals`. The guard is worth adding anyway, since the state is constructible today. The test asserts the build-time shapes and each arm's merge branch with `EXPLAIN PIPELINE`, plus values against a hand-derivable ground truth. Two unrelated defects, identical on both sides of this change: `Chunk info was not set for chunk in GroupingAggregatedTransform`, owned by #113266, and the same setting over `merge()` loses a child's rows. An earlier revision of this line attributed the first to `MaterializingTransform` dropping chunk info; that is wrong, since it only calls `detachColumns`/`setColumns` and the `ISimpleTransform` default swaps `chunk_infos` through. Issue 113708 reports this guard firing over `merge()` through a set-operation plan with no `MergingAggregatedStep`; that did not reproduce, so this PR does not close it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114027",
          "createdAt": "2026-08-09T10:34:40Z",
          "updatedAt": "2026-08-13T05:53:27Z",
          "timestamp": "2026-08-13T05:53:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3bd1262c1f01dba3adf8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106231",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106231",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Implement CREATE HANDLER: SQL-defined HTTP handlers",
          "text": "Implements [#100000](https://github.com/ClickHouse/ClickHouse/issues/100000): create and manage custom HTTP handlers from SQL, without editing the server configuration file. New SQL statements: - `CREATE HANDLER [IF NOT EXISTS] name [PROTOCOL p] URL [PREFIX|REGEXP] '/x' [METHODS (GET, POST)] [TYPE query] AS <query>` - `ALTER HANDLER name ...` — partial update of any subset of clauses - `DROP [IF EXISTS] HANDLER name` Handlers are matched after configuration-defined handlers, in the lexicographical order of their names. The URL is matched without the `?` query string and `#` fragment; exact/prefix URLs are checked for ambiguity at create/alter time. The query is parsed (not analyzed) at creation, can be parameterized (URL params, form variables, headers, and named regexp capture groups), and an `INSERT` handler reads its data from the HTTP body. Handlers are persisted in a local or Keeper storage, mirroring named collections, configured via the `query_rules_storage` config section, and kept in sync across replicas. Access control adds `CREATE HANDLER`, `ALTER HANDLER` and `DROP HANDLER` grants. Introspection adds the `currentHandler` and `currentRequestURL` functions, the `http_handler_name` and `http_request_url` columns in `system.query_log`, and the `system.handlers` table. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `CREATE HANDLER`, `ALTER HANDLER` and `DROP HANDLER` statements to define custom HTTP handlers from SQL, persisted in a local or Keeper storage. Added the `currentHandler` and `currentRequestURL` functions, the `system.handlers` table, and the `http_handler_name` / `http_request_url` columns in `system.query_log`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106231",
          "createdAt": "2026-06-01T12:09:45Z",
          "updatedAt": "2026-08-13T05:52:33Z",
          "timestamp": "2026-08-13T05:52:33Z",
          "metrics": {
            "reactions": 1,
            "comments": 79
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f193920f44a48f1a21f5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113192",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113192",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject shorthand setting changes carrying a value for the whole query tree",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113025 The valueless `SETTINGS name` form stands for `name = true`, and the SQL parser always writes Bool `true` for it, so the `shorthand` flag paired with any other value is a parser-impossible shape that can only arrive from the AST JSON dialect. #113025 made `BaseSettings::checkShorthandChange` reject it, but that only covers the settings applied through `BaseSettings`. `SettingsChanges` are also consumed raw — the `Join` engine reads `persistent` and its other settings directly, `EXPLAIN` settings never consult a `BaseSettings` schema, and dictionary and data-lake settings have similar readers — and every such reader would execute the carried value for a change that claims to be valueless. Instead of duplicating the check in each raw reader, reject the shape once for the whole query tree in `executeQueryImpl`, for ASTs deserialized from the JSON dialect. The check deliberately runs after `query_for_logging` is prepared, so the exception is logged with the AST masked rather than with the raw JSON text — the same reason the shape is not rejected at deserialization. Because the tree-wide check has no settings schema, it fires before the per-setting type check, so the crafted payload in `04665_valueless_setting_ast_json_and_secret_parts` now reports `BAD_ARGUMENTS` (carries a value despite claiming the valueless form) instead of `TYPE_MISMATCH`; the genuine valueless forms and their error codes are unchanged, and the masked-logging assertion still holds. This addresses the remaining review finding on #113025. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reject a setting change that is marked as the valueless `SETTINGS name` form but carries a value other than `true` for every consumer of settings (e.g. the `Join` engine or `EXPLAIN` settings), not only for the settings applied through `BaseSettings`. Such a change can only be produced by the experimental AST JSON dialect.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113192",
          "createdAt": "2026-08-03T21:58:16Z",
          "updatedAt": "2026-08-13T05:51:59Z",
          "timestamp": "2026-08-13T05:51:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f713347a6bf07c67e4e1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114594",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114594",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "[Draft/Prototype] Add mutations_restrict session setting as a mutation safety catch",
          "text": "> ⚠️ **Experimental prototype — heavy LLM assistance.** > > This PR is an early prototype built with substantial LLM assistance > (GitHub Copilot CLI). Early draft, sharing for searchability and > to use CI support. ## Summary Adds a new session-level `UInt64` setting **`mutations_restrict`** that acts as a safety catch against accidental mutations. Modeled on the existing \"block all DDL\" switch `allow_ddl`, but scoped to mutation-producing statements and shaped as tiers like `readonly` / `mutations_sync`: | Value | Effect | |-------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | `0` | No restriction (default; existing behavior). | | `1` | Reject `ALTER TABLE` forms that produce a `system.mutations` entry: `UPDATE`, `DELETE`, `MATERIALIZE INDEX`/`PROJECTION`/`COLUMN`/`STATISTICS`/`TTL`, `APPLY DELETED MASK`, `APPLY PATCHES`, `DROP INDEX`/`PROJECTION`/`STATISTICS`, `RENAME COLUMN`, `DROP COLUMN` on a physical column, non-metadata `MODIFY COLUMN`, `CLEAR ... IN PARTITION`. Metadata-only `ALTER`s, `INSERT`, and `SELECT` still succeed. Lightweight `DELETE`/`UPDATE` still succeed. | | `2` | Additionally reject standalone lightweight `DELETE` and `UPDATE`. | The setting complements the server-level `disable_insertion_and_mutation` (which is global and also blocks `INSERT`). Unlike that setting, `mutations_restrict` is a normal `Settings` entry, so it can be: - set per session with `SET mutations_restrict = N;`, - set per user via profile, - pinned system-wide or per user with a `<readonly/>` constraint (same pattern as `allow_drop_detached`). Typical intended usage: users pin `1` (or `2`) as a soft default in their profile and lower it in-session when they need to mutate; admins pin it as `<readonly/>` for hard enforcement. ## Rationale There is currently no per-session safety catch for accidental mutations. Options in-tree today: - `disable_insertion_and_mutation` (server-scoped, blocks inserts too), - RBAC (`ALTER UPDATE`, `ALTER DELETE`, ... as separate grants — heavy for a \"guard against fat-fingering\" workflow), - `readonly = 1` (blocks everything, not just mutations). None of these fit \"let me use this session but refuse to run heavy mutations by mistake\". ## Where the checks live - **Tier 1** — `InterpreterAlterQuery::execute` after the `CommandSegments` are built and the existing `validateMutationsAllowed` runs. Uses `MutationCommands::hasNonEmptyMutationCommands`, so the schema-dependent cases (e.g. `MODIFY COLUMN` where the type change is a metadata-only conversion) are resolved correctly — the interpreter is the earliest point where the mutation-vs-metadata question has been answered. - **Tier 2** — `InterpreterDeleteQuery::execute` and `InterpreterUpdateQuery::execute`, at the top of `execute`. The alternative of gating in `ContextAccess::checkAccessImplHelper` alongside `allow_ddl` was considered and rejected: it would require introducing mutation-specific `AccessFlags`, is a larger refactor, and cannot distinguish schema-dependent cases. ## What's not covered (by design) - `SYSTEM ...` commands. - Partition manipulation (`ATTACH`/`DETACH`/`DROP PARTITION`). - `OPTIMIZE`, `TRUNCATE`, `CREATE`/`DROP TABLE`. - `INSERT`. These do not appear in `system.mutations`. Documented as such in the setting docstring. ## Tests - `tests/queries/0_stateless/04869_mutations_restrict_safety_catch.sql` — tier 0/1/2 behavior + in-session override + mutation vs metadata-only ALTER distinction. - `tests/integration/test_mutations_restrict/` — profile pin + `<readonly/>` enforcement, plus a \"soft default\" profile that demonstrates in-session override. ## Docs - Setting reference under `docs/reference/settings/session-settings/mutations.mdx` will regenerate from the `DECLARE(...)` docstring in `src/Core/Settings.cpp`. - Non-autogenerated section added to `docs/reference/statements/alter/index.mdx` cross-referencing the new setting. ## Known open questions (please review) 1. **Naming.** `mutations_restrict` vs `restrict_mutations` vs `mutation_safety_catch`. Chose `mutations_restrict` to group with the `mutations_*` family and match the \"higher = more restrictive\" direction of `readonly`. 2. **Tier semantics.** Currently tier 1 blocks `ALTER TABLE ... UPDATE` even when `alter_update_mode` would rewrite it to a lightweight update. Rationale: it's syntactically an `ALTER`. Alternative: move the tier-1 check to run after the lightweight rewrite decision, so it only fires on the heavy path. 3. **Standalone `DELETE FROM ...` heavy path.** For engines where `supportsDelete()` returns true (e.g. `KeeperMap`, `RocksDB`) the standalone `DELETE FROM` still calls `table->mutate` — this is currently only blocked at tier 2 (via the top-of-execute check). Should it also be blocked at tier 1? (Unclear whether the produced mutation is user-observable in the same \"system.mutations\" sense.) 4. **`SettingsChangesHistory.cpp`.** Added to the `26.8` block; the correct release for master at merge time may differ. ## Not tested locally The sandbox that authored this branch has no C++ toolchain, so I could not build, run stateless tests, or run the integration test locally. Relying on CI. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new session setting `mutations_restrict` as a safety catch against accidental mutations. Value `1` rejects `ALTER TABLE` forms that would create an entry in `system.mutations`; value `2` additionally rejects standalone lightweight `DELETE` and `UPDATE`. Can be set per session, per user profile, or pinned system-wide with a `<readonly/>` constraint.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114594",
          "createdAt": "2026-08-13T05:04:28Z",
          "updatedAt": "2026-08-13T05:47:48Z",
          "timestamp": "2026-08-13T05:47:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [],
          "author": "ringerc",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a49970f2ee33ec571546",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114576",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114576",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112498 to 26.7: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.25` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114576",
          "createdAt": "2026-08-13T02:36:15Z",
          "updatedAt": "2026-08-13T06:34:50Z",
          "timestamp": "2026-08-13T06:34:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "closed",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:552416210340df5e6861",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:103182",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:103182",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Lightweight Updates v2",
          "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Lightweight `UPDATE` patch parts now use a new v2 on-disk format sorted by (`sorting_key..., _block_number, _block_offset`) and applied with a new merging algorithm. Peak memory is bounded by the largest equal-sort-key run instead of the full patch, and updates that cross merge boundaries no longer fall back to in-memory Join apply. Old-format patch parts remain readable. During a rolling upgrade from a version before 26.8, keep `patch_parts_version = 'v1'` or use the `compatibility` setting until all replicas are upgraded. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/103182",
          "createdAt": "2026-04-20T17:45:11Z",
          "updatedAt": "2026-08-13T05:46:52Z",
          "timestamp": "2026-08-13T05:46:52Z",
          "metrics": {
            "reactions": 2,
            "comments": 8
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b8402bb8d0ed3e586062",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114577",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114577",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Mongo queries: keep whole documents in a JSON column",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/68493 Related: https://github.com/ClickHouse/ClickHouse/issues/58394 Adds two ways to run MongoDB queries against ClickHouse: - a **wire protocol endpoint** (`mongo_port`), so MongoDB drivers and tools such as `pymongo` and `mongosh` can connect to ClickHouse as if it were a MongoDB server; - a **query dialect** (`SET dialect = 'mongo'`), so MongoDB shell syntax can be sent over the usual ClickHouse interfaces. This continues @scanhex12's https://github.com/ClickHouse/ClickHouse/pull/68493 - it holds all of its commits - and changes how a collection is stored, which is what the review of that pull request asked for (https://github.com/ClickHouse/ClickHouse/pull/68493#discussion_r3771132310): a collection the endpoint creates keeps whole documents in one `JSON` column rather than a column inferred per field of the first document. ```sql CREATE TABLE db.users (`_id` String, `json` JSON) ENGINE = MergeTree ORDER BY `_id` ``` The document goes into the `JSON` column, whose paths are its fields, and its object id into the `_id` column, which is the primary key. A document may then hold a field that no document before it had, may leave out a field another one has, and `$exists` answers what it means. A table that was created in ClickHouse keeps its own columns, and a field of a query names the column of the same name there, so an application can also be pointed at a table that already holds the data. A read of the documents as they are stored selects them next to the types of their paths, so a date reads back as a BSON date and an integer as an `int32`/`int64` of its width. **This is a draft**: the storage change is done and verified by hand against `pymongo` (insert, `find` with filters, projections, `$exists`, nested paths, sort, limit, `count`, `distinct`, `aggregate`, `delete`, `createCollection`, a read by `_id`, the round trip of a date), but the integration tests of `test_mongo_protocol` still pin the previous shape and have to be reworked, and four things of a document collection are an explicit error rather than an answer: an `update` of it, an index over a path, the `mongo` dialect over it, and the element-wise match of an array by equality or `$in`. They are the next commits; the limitations are listed in `docs/en/interfaces/mongo.md` in the meantime. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a MongoDB-compatible wire protocol endpoint (the `mongo_port` server setting) and a MongoDB query dialect (`SET dialect = 'mongo'`) for basic collection operations. A collection the endpoint creates keeps whole documents in a `JSON` column.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114577",
          "createdAt": "2026-08-13T02:49:29Z",
          "updatedAt": "2026-08-13T05:45:23Z",
          "timestamp": "2026-08-13T05:45:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e6262f5f507ccdf22788",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114566",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114566",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve pruned statistics planning benchmark and comments",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110283 ### Summary Follow up on review feedback from #110283: - clarify the range-analysis, exact-empty, STREAM, and statistics-cache contracts; - document what `ConditionSelectivityEstimator::isStale()` compares; - remove cross-component implementation references from optimizer comments; - make the performance workload exercise statistics from 950 of 1000 parts instead of producing noisy 2–3 ms measurements. This PR does not change production behavior. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry: Not for changelog. ### Testing - `git diff --check` - XML validation with `xmllint` - local hostile diff review The release performance comparison is left to CI. 🕵️‍♂️",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114566",
          "createdAt": "2026-08-13T01:12:41Z",
          "updatedAt": "2026-08-13T05:44:39Z",
          "timestamp": "2026-08-13T05:44:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "skuznetsov-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:bb82afd4889173e85265",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110781",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110781",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add experimental support for reading Iceberg v3 deletion vectors",
          "text": "Resolves: https://github.com/ClickHouse/ClickHouse/issues/107502 **Changelog category (leave one):** - New Feature **Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):** This PR adds experimental read support for Iceberg v3 deletion vectors stored in Puffin files. Supported: - Read Iceberg v3 deletion vector metadata from manifest entries: - `referenced_data_file` - `content_offset` - `content_size_in_bytes` - Read Puffin deletion vector blobs from object storage. - Validate DV blob length, magic bytes, and CRC. - Decode Roaring64 deletion vectors. - Apply deletion vectors during reads. - Support mixed reads with: - v2 parquet position delete files - v3 Puffin deletion vectors - Support table functions and table engines through the common Iceberg read path, including local/S3/Azure object storage. - Add an experimental setting: - `allow_experimental_iceberg_deletion_vectors` Not supported: - Writing Iceberg v3 deletion vectors. - Updating existing deletion vectors. - Compaction/rewrite of deletion vectors. - Container-level lazy decoding of Roaring bitmaps. - Hard memory limit on decoded Roaring bitmap memory. - Full Iceberg v3 feature support beyond deletion vector reads. Notes: - The feature is disabled by default and must be enabled with `allow_experimental_iceberg_deletion_vectors = 1`. - When disabled, existing behavior is preserved. - The implementation decodes the DV blob into a Roaring bitmap and applies it with a streaming iterator during data reads. **Documentation entry for user-facing changes** - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110781",
          "createdAt": "2026-07-17T05:46:20Z",
          "updatedAt": "2026-08-13T05:41:34Z",
          "timestamp": "2026-08-13T05:41:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "linjiayu1025-collab",
          "state": "open",
          "assignees": [
            "asya-ch"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1739eabc6b3e5ee9fd3d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109925",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109925",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Warm statistics estimator on commit to avoid first-SELECT cold load",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/109454 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Warm the table statistics estimator right after an INSERT commits so the first `SELECT` does not synchronously cold-load and rebuild column statistics from disk in the query planner. ### Description Helper PR against `stats-on-insert-size-threshold`, folding in the root-fix direction discussed on #109454. Opened at @egor-click's request. Root fix (instead of gating the planner around the cold load): - Retain the full per-column `ColumnsStatistics` that INSERT already builds on the in-memory part (`IMergeTreeDataPart::statistics_cache` + `setStatistics`/`tryGetCachedStatistics`/`hasCachedStatistics`). Disk-loaded parts leave it empty and keep using the on-disk files. - Warm `MergeTreeData::cached_estimator` in `Transaction::commit` once committed parts become Active, so the first `SELECT` skips the synchronous planner cold-load. - The estimator builder now copies statistics into its own aggregate (`cloneEmpty()` + `merge`) before merging, so it never mutates the part-owned `ColumnStatistics` that `loadStatistics()` returns. This also protects `refreshStatistics` and the what-if estimator, which both route through `addStatistics`. Hardening applied (item 1 of the 3 reported on #109454): - Commit-time warming is gated on committed parts actually carrying retained in-memory statistics (`hasCachedStatistics`), not on `use_statistics_cache` alone. Without this gate, tables whose statistics are NOT materialized on insert (large tables above `materialize_statistics_on_insert_max_table_size`, or the setting off) would loop all active parts and do the full on-disk cold-load synchronously on every INSERT. Measurements (debug, this branch's base): - Smoke (MergeTree, tdigest+uniq+minmax+countmin on 2 cols): server-side planner cold-load marker = 0 across SELECT-after-INSERT / repeat / SELECT-after-2nd-INSERT; commit-warm fired 3x; no crash / LOGICAL_ERROR / sanitizer. - Gate verified both directions: implicit-stats table warms on commit under the INSERT query id; `auto_statistics_types=''` + `materialize_statistics_on_insert=0` skips warming under the INSERT id. - A/B on the exact `join_convert_outer_to_inner` query, same binary/data: fix = 0 planner cold-loads on first SELECT-after-INSERT; baseline (`use_statistics_cache=0`) = 4 cold-loads (2 parts x 2 tables). Remaining follow-ups (reported on #109454, not in this PR): incremental warm instead of O(parts) rebuild per commit; bound `statistics_cache` retention by a table-size budget or drop after first warm; and the `disk_connections_hard_limit` interaction the bot flagged for remote-storage inserts.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109925",
          "createdAt": "2026-07-09T16:30:04Z",
          "updatedAt": "2026-08-13T05:28:57Z",
          "timestamp": "2026-08-13T05:28:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:42e1ca865c228314529e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114561",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114561",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix MATERIALIZED columns frozen at their INSERT-time value after ALTER UPDATE",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/100613 Closes: https://github.com/ClickHouse/ClickHouse/issues/50237 Related: https://github.com/ClickHouse/ClickHouse/pull/99281 ## Problem A `MATERIALIZED` column computed from another `MATERIALIZED` column is never recalculated by `ALTER TABLE ... UPDATE`. It keeps the value `INSERT` computed, forever, no matter how many mutations run. ```sql CREATE TABLE t (x Int32, m1 Int32 MATERIALIZED x + 1, m2 Int32 MATERIALIZED m1 + 1) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t (x) VALUES (10); -- m1 = 11, m2 = 12 ALTER TABLE t UPDATE x = 20 WHERE 1; ALTER TABLE t UPDATE x = 30 WHERE 1; SELECT x, m1, m2 FROM t; -- 30 31 12 <- m2 frozen at its INSERT-time value -- expected: 30 31 32 ``` Concretely, on `master`: - `m2` and everything downstream of it are stale for the lifetime of the table. - A projection over a recalculated `MATERIALIZED` column keeps answering from pre-mutation data: `SELECT sum(m1)` returns `500500` through the projection versus `1000500500` without it. - A `minmax` skip index over such a column prunes every granule, so a filter that matches all rows returns `0`. - With `apply_mutations_on_fly = 1` and the mutation still pending, the same query returns different values depending on its select list: `SELECT m2` gives `12` while `SELECT x, m1, m2` gives `22`. - Also with a pending mutation, reading a `MATERIALIZED` column whose expression mentions an `EPHEMERAL` column fails outright with `Code: 47 ... Missing columns: 'e' while processing: 'x + e' ... (UNKNOWN_IDENTIFIER)`. ## Root cause `MutationsInterpreter::prepare` collected the `MATERIALIZED` columns to recalculate from **direct** dependencies only. It analysed each `MATERIALIZED` default expression and recorded an edge `dependency -> materialized column`, but only when `updated_columns` already contained the dependency: ```cpp for (const auto & dependency : required_columns) if (updated_columns.contains(dependency)) column_to_affected_materialized[dependency].push_back(column.name); ``` For `m2 MATERIALIZED m1 + 1` the dependency is `m1`, which the user did not update, so `m2` never entered `affected_materialized`, and the guard added by #99281 skipped it: ```cpp if (column.default_desc.kind == ColumnDefaultKind::Materialized && affected_materialized.contains(column.name) && column.default_desc.expression) ``` Before #99281 every `MATERIALIZED` column was recalculated in a single stage, so `m2` *was* rewritten — but from the pre-stage `m1`, i.e. always one mutation behind. That second half is #50237. Both halves come from the same place: the affected set is not transitive, and even a transitively dependent column was evaluated in the same stage as the column it reads. The remaining symptoms are the same \"direct dependencies only\" assumption in the neighbouring consumers. The skip-index, projection and dependency-analysis predicates only looked at `updated_columns`, `changed_columns` and `patch_updated_columns`, none of which ever names a recalculated `MATERIALIZED` column. `AlterConversions::addColumnsRequiredForMaterialized`, which decides what an on-fly mutation must read, walked one level and then required the dependency to be an updated column, so a read set of `{m2}` stayed `{m2}`, `filterMutationCommands` dropped the `UPDATE x` command because nothing overlapped, and the stale value was returned. That same function analysed default expressions against `getAllPhysical`, which excludes `EPHEMERAL` columns, hence the `UNKNOWN_IDENTIFIER`. ## Solution In `src/Interpreters/MutationsInterpreter.cpp`: 1. Record **every** edge of the `MATERIALIZED` dependency graph, including `MATERIALIZED -> MATERIALIZED` edges. 2. `getAffectedMaterializedByLevel` walks that graph from the columns a command updates and groups the reachable columns by **longest-path** dependency level. 3. Emit one mutation stage per level. Stage *N* reads the output of stage *N-1*, so an expression sees the recalculated value of the column it depends on instead of the stale on-disk one — this is what closes #50237, not merely refreshing the value. 4. `validateUpdateColumns` receives the same transitive closure, so an `UPDATE` reaching a `MATERIALIZED` **key** column through another `MATERIALIZED` column is rejected with `CANNOT_UPDATE_COLUMN`, as the direct case already was. Without this, step 3 would rewrite a sorting-key column and break the part's sort order. 5. Feed the recalculated columns into the skip-index, projection and dependency analysis. In `src/Storages/MergeTree/AlterConversions.cpp`: 6. `addColumnsRequiredForMaterialized` follows the chain, pulling in an intermediate `MATERIALIZED` column when an updated column is reachable through it. 7. `EPHEMERAL` columns enter its analysis set, as `MutationsInterpreter::prepare` already does, and a `MATERIALIZED` column reading one is excluded from the walk because it cannot be recalculated anyway. The pre-existing `EPHEMERAL` exclusion that #99281 added is preserved: such a column is skipped before its incoming edges are recorded, so the walk never reaches it. The exclusion removes that column from the graph, not the subgraph behind it — a column that also reads an updated column directly (`mc MATERIALIZED x + e`, `mc2 MATERIALIZED mc + x`) is still recalculated, from the stored `mc`, which is well defined because its own expression reads only stored columns. Items 5-7 are pre-existing defects rather than regressions of this change; item 5 is reproducible on `master` for a *directly* recalculated column too. They are fixed here because otherwise items 1-3 would turn a uniformly stale table into an inconsistent one: correct on disk, wrong through a projection, and different per select list. ### Tests `tests/queries/0_stateless/04869_transitive_materialized_column_update.sql` covers nine cases, one per clause of the fix: chain, diamond (longest-path placement), `EPHEMERAL` must-not-recalculate, `EPHEMERAL` with converging paths, `MATERIALIZED` key column, projection (direct and transitive), skip index, on-fly chain, on-fly `EPHEMERAL`. Each assertion was verified to be load-bearing by reverting only its own clause: the chain freezes at `12 13`, the diamond gives `5 6 60 26` when a re-reached column is placed at its first level instead of its deepest, the projections return `500500` / `501500`, the skip index returns `0`, and the on-fly reads return `12` / `13` or throw. `04044_mutation_ephemeral_materialized`, the test #99281 added, still passes. The projection and skip-index cases keep an untouched column `y`, because a mutation that rewrites every column goes through `MutateAllPartColumns`, which rebuilds everything and hides the bug. ### Notes for reviewers - **Behaviour change:** `UPDATE` of a column that transitively feeds a `MATERIALIZED` key column now throws `CANNOT_UPDATE_COLUMN` where it previously succeeded and left the key stale. The direct equivalent already throws on `master`, so this makes the two consistent. - Tables with `MATERIALIZED -> MATERIALIZED` chains now get one mutation stage per dependency level; everything else keeps the single stage that existed before. - More projections and skip indices are rebuilt than before. That is the point of item 5, but it does make such mutations more expensive. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `MATERIALIZED` columns computed from another `MATERIALIZED` column keeping their `INSERT`-time value forever after `ALTER TABLE ... UPDATE`. Projections and skip indices over a recalculated `MATERIALIZED` column were also left stale, returning wrong results.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114561",
          "createdAt": "2026-08-13T00:41:12Z",
          "updatedAt": "2026-08-13T05:23:32Z",
          "timestamp": "2026-08-13T05:23:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6843a08d0088ca70228b",
        "signalId": "github:ClickHouse/ClickHouse:issue:114595",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114595",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "S3Queue (multi-server, hash-ring): a stale per-server Processed cache silently skips a re-appeared object after its tracked-file TTL expires",
          "text": "### Company or project name ClickHouse QA (internal durability testing) ### Describe what's wrong In a multi-server `S3Queue`/`ObjectStorageQueue` setup with `enable_hash_ring_filtering = 1`, a per-server in-memory \"Processed\" cache silently skips a re-appeared object after that object's tracked-file record has been aged out of Keeper by a different server. The row is never ingested — no error, no failed file. The pieces: - `ObjectStorageQueueIFileMetadata::trySetProcessing()` returns `false` immediately when the in-memory per-server `local_file_statuses` entry for a path is `Processed`, **without reading Keeper** (the Keeper check in `setProcessingImpl` is only reached after this in-memory guard). So a server that has already processed key `K` will not even ask Keeper about `K` again. - `tracked_file_ttl_sec` cleanup removes a file's `/processed` node from Keeper, and it removes the file from `local_file_statuses` **only on the server that runs the cleanup** (`ObjectStorageQueueMetadata`, the TTL cleanup path). The other servers keep their stale in-memory `Processed` entry. - `enable_hash_ring_filtering` routes each key to exactly one server (`filterOutForProcessor` → `chooseServer(hash(path))`). Put together: server X processes `K` (X's cache = `Processed`, Keeper `/processed/K` created). Server Y later runs the TTL cleanup and deletes `/processed/K` from Keeper (and Y's own cache entry, which never had `K`). The object at key `K` is then rewritten with new content (a legitimate same-key overwrite). The hash ring still routes `K` to X. X consults its in-memory cache, sees `Processed`, and skips `K` — never reaching Keeper, where the record no longer exists. The rewritten, producer-acked content is silently dropped. ### Does it reproduce on the most recent release? Yes — reproduced 2/2 on a current `26.8.1.1` master build. ### How to reproduce Two ClickHouse servers sharing one Keeper and one bucket, each with an `S3Queue` table on the same `keeper_path` and a local target table via a materialized view: ```sql CREATE TABLE q (id UInt64, val String) ENGINE = S3Queue('http://minio:9000/bucket/*', 'JSONEachRow') SETTINGS mode = 'unordered', enable_hash_ring_filtering = 1, tracked_file_ttl_sec = 4, processing_threads_num = 1; ``` On server X set the cleanup interval high (e.g. `cleanup_interval_min_ms = cleanup_interval_max_ms = 3600000`); on server Y set it low (e.g. `500`/`1000`) so Y deterministically drives the TTL cleanup. (These are per-server local settings, not part of the shared table metadata.) 1. Upload objects; observe (via each server's own target table) which keys the hash ring routes exclusively to X. Pick such a key `K`. 2. Wait for Y's cleanup to age out `/processed` (confirm the `/processed` children are gone via `system.zookeeper`). 3. Re-upload key `K` with new content (a sentinel row). Confirm the object with the new content is present in the bucket. 4. The sentinel never appears in either server's target table — silent loss. Control (proves it is the stale in-memory cache, not a genuine Keeper state): repeat with a different X-routed key, but restart server X (clearing its in-memory `local_file_statuses`) before the re-upload — the rewritten object is then ingested. ### Expected behavior Once a file's tracked record has been aged out of Keeper, a re-appeared object with that key must be eligible for processing by whichever server owns it. The per-server in-memory `Processed` cache must not short-circuit past a Keeper record that no longer exists — `trySetProcessing` should confirm against Keeper before declining a file whose local state is `Processed`, or the TTL cleanup must invalidate the entry on all servers, not only the cleaning one. ### Error message and/or stacktrace _No error_ — the file is silently not processed; `system.s3queue`/`s3queue_log` show no failure. ### Additional context - In-memory short-circuit: `src/Storages/ObjectStorageQueue/ObjectStorageQueueIFileMetadata.cpp` (`trySetProcessing`, the `Processed` early-return before any Keeper request). - Cache removal only on the cleaning server: `src/Storages/ObjectStorageQueue/ObjectStorageQueueMetadata.cpp` (TTL cleanup `local_file_statuses.remove`). - Hash-ring routing: `filterOutForProcessor` / `chooseServer`. - Distinct from #112341 (single-server Processed-while-discarding on MV detach — no ageout, no cross-server cache asymmetry), #111537 (power-loss at-least-once → at-most-once), #112041 (DatabaseReplicated recovery drops the registry), and #109751 (Azure listing continuation-token truncation). This is specifically the stale per-server `Processed` cache winning over an aged-out Keeper record under hash-ring routing. - Reachability: `enable_hash_ring_filtering` is the multi-server distribution mode and `tracked_file_ttl_sec` is the documented tracked-files bound; the trigger is a same-key overwrite.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114595",
          "createdAt": "2026-08-13T05:21:03Z",
          "updatedAt": "2026-08-13T05:21:03Z",
          "timestamp": "2026-08-13T05:21:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:73d407c8e3af932e1721",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113347",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113347",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Write wide integers in Parquet as Decimal",
          "text": "Previously UInt128/UInt256/Int128/Int256 were written as `FIXED_LEN_BYTE_ARRAY`, but as little-endian. Since little-endian is not lexicographically sortable, there were no statistics and no way to prune row groups and pages. This was not intentional, but rather an accidental artifact from when Parquet serialization was first introduced in ClickHouse. However, we cannot just change the endianness and break existing code/data. Here instead we switch to `DECIMAL` encoding in Parquet, which is big-endian. This enables row group and page pruning for querying for queries that use filtering on wide integer columns. The encoding is a little bit akward with 17 or 33 bytes per integer, but it enables other data consumers to have a correct data interpretation hint. Since the new data cannot be read by older versions, this new serialization mode is behind `output_format_parquet_wide_integer_as_decimal` setting, which defaults to `0` to preserver backward compatibility. The new version can still read old data (but not the other way around). <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add opt-in Parquet serialization of UInt128/UInt256/Int128/Int256 as DECIMAL to enable row group and page pruning.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113347",
          "createdAt": "2026-08-04T15:40:06Z",
          "updatedAt": "2026-08-13T05:18:16Z",
          "timestamp": "2026-08-13T05:18:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "bobrik",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8e20a51e3cd4732e7997",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106215",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106215",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enable read_in_order_use_virtual_row by default",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/52624 When reading in order of the primary key over many parts (the `optimize_read_in_order` optimization), the merging pipeline opens one reader per part to feed the `MergingSortedTransform`, and keeps a `CompressedReadBuffer` per column resident for every one of them. With a large number of parts this makes even `ORDER BY pk LIMIT n` consume memory proportional to the number of parts, even though only a few of them can contribute to the result. The original report describes a query needing ~42 GB for this reason (1803 parts × 3 columns × 8 MB compress blocks). The `read_in_order_use_virtual_row` optimization lets `MergingSortedTransform` reprioritize sources using primary key values taken from the sparse index, so the parts that cannot contribute to the answer are never read and their readers are never created. The optimization already existed but was disabled by default; this change enables it. Measured on a table with 300 parts and an 8 MB compress block size, the peak memory of `SELECT * FROM t ORDER BY k LIMIT 10` drops from ~700 MB to a few MB. Related issues: #53101, #53287, #7413 (closed), and #89638 (open) describe the same underlying behavior. Note: the more aggressive `read_in_order_use_virtual_row_per_block` (which also disables `read_in_order_use_buffering` and the preliminary merge) is intentionally left disabled by default, since it trades some parallelism for the extra memory savings. **Update.** The first CI run showed consistent 1.3x–5x slowdowns in read-in-order perf tests (`distinct_in_order`, `optimize_sorting_for_input_stream`, `read_in_order_many_parts`, `monotonous_order_by`, `redundant_functions_in_order_by`) on both architectures. The cause: a source that starts with a virtual row is not read until the merge reaches its key, so the merge requests deferred sources strictly one by one and reading degrades to a single thread (locally, `SELECT DISTINCT` over 100M rows ran at the same speed with `max_threads=1` and `max_threads=16`). This was addressed by letting deferred sources read ahead: without a `LIMIT` a window of `number of threads` deferred sources, and with a `LIMIT` only sources provably needed soon (those whose virtual row is not greater than the key at which the merge must stop within an already-read chunk). **Update 2.** The second perf run and review showed that the \"provably needed\" rule for the `LIMIT` case was wrong in both directions: - it was unbounded: with many overlapping parts, every deferred source whose virtual row preceded the merge stop row was scheduled at once, which could re-create the O(parts) resident-reader memory issue this PR fixes; - it was still too conservative: with a selective filter (`SELECT * FROM hits WHERE UserID = ... ORDER BY pk LIMIT 100`), filtered chunks rarely reach the merge, the proof never arrives, and reading degrades to a single stream again — `order_by_read_in_order` showed 3.3x–4.7x and `lazyMaterialization` 1.9x–2.5x slowdowns on both architectures. Both are fixed by making the read-ahead window unconditional: the merge always keeps the next `number of threads` deferred sources (in the order it will need them) reading ahead, with or without a `LIMIT`. Under a `LIMIT` the read-ahead is speculative — the merge may finish before reaching a prefetched source — but the waste is bounded by O(threads) chunks, because a source past its virtual row only buffers one chunk in its output port until the merge consumes it. Peak reader memory stays bounded by the window instead of the number of parts: the 300-part `LIMIT 10` query above runs with ~7 MB peak memory and reads only the window's first chunks, vs ~1.3 GB and a full read without the optimization. Pending read-ahead is dropped as soon as the merge finishes, so a query that already produced its `LIMIT` does not start new reads. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable `read_in_order_use_virtual_row` by default. When reading in order of the primary key (e.g. `ORDER BY primary_key LIMIT n`) over a table with many parts, only the parts that can actually contribute to the result are read (plus a read-ahead window of at most `max_threads` parts that keeps reading parallel), which significantly reduces peak memory consumption. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106215",
          "createdAt": "2026-05-31T22:32:25Z",
          "updatedAt": "2026-08-13T05:06:22Z",
          "timestamp": "2026-08-13T05:06:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7f587c49ff1de92e534f",
        "signalId": "github:ClickHouse/ClickHouse:issue:114593",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114593",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "RESTORE: a mid-attach I/O error leaves an un-rolled-back durable data prefix; the suggested allow_non_empty_tables retry duplicates it",
          "text": "### Company or project name ClickHouse QA (internal durability testing) ### Describe what's wrong An I/O error partway through `RESTORE` leaves already-attached parts Active and read-visible, with no rollback, while `RESTORE` reports failure — and the server's own error message then steers the operator into silently duplicating that residue. `StorageReplicatedMergeTree::attachRestoredParts` attaches a backup's parts in a bare per-part loop of `sink->writeExistingPart(part, /*deduplicate_part*/ false)` with no try/compensation around it. Each part commits by renaming its `tmp_restore_*` staging directory into the final active name. If the rename of the k-th part fails (for example a transient filesystem error), the parts committed before it stay Active and readable, but `RESTORE` fails and reports the table could not be restored. The failed operation has left a durable, read-visible data prefix. The natural next step makes it worse. Because the table is now non-empty, a retry of the same `RESTORE` fails with `CANNOT_RESTORE_TABLE ... use allow_non_empty_tables=true`, and following that suggestion re-attaches every part with fresh block numbers (dedup off), duplicating the rows that the failed restore had already left behind. The operator is guided by the server's own message into duplicating the residue of a failed restore. ### Does it reproduce on the most recent release? Yes — reproduced on a current `26.8.1.1` master build; `attachRestoredParts` is a bare per-part `writeExistingPart` loop with no rollback in the current source. ### How to reproduce On a `ReplicatedMergeTree` table with several partitions (so RESTORE attaches several parts): ```sql CREATE TABLE rc (id UInt64, v String) ENGINE = ReplicatedMergeTree('/clickhouse/tables/rc','r0') PARTITION BY (id % 3) ORDER BY id; INSERT INTO rc SELECT number, toString(number) FROM numbers(300); -- 300 rows, 3 parts BACKUP TABLE rc TO File('rc'); DROP TABLE rc SYNC; CREATE TABLE rc (id UInt64, v String) ENGINE = ReplicatedMergeTree('/clickhouse/tables/rc','r0') PARTITION BY (id % 3) ORDER BY id; -- empty target ``` Inject a filesystem I/O error (`EIO`) on the rename that commits the **second** restored part (the second `tmp_restore_*` rename), then: ```sql RESTORE TABLE rc FROM File('rc'); -- fails SELECT count() FROM rc; -- 100 (a durable, read-visible prefix of a failed RESTORE) ``` Retry as the server's error message suggests: ```sql RESTORE TABLE rc FROM File('rc') SETTINGS allow_non_empty_tables = 1; -- succeeds SELECT count() FROM rc; -- 400 (300 unique rows + 100 duplicated from the residue) ``` Verified with a deterministic control-gated harness that arms the `EIO` on exactly the second restored part's commit rename. Controls: a clean `RESTORE` into an empty table round-trips exactly 300 rows; a single-part backup under the same fault leaves nothing (no prefix) and the retry restores exactly 300 — so the prefix is attributable to the multi-part attach fan-out. ### Expected behavior A `RESTORE` that fails partway through should leave the table unchanged (no durable, read-visible prefix from a failed restore), so that a retry restores exactly the backup's contents without duplication. At minimum, the residue of a failed `RESTORE` should not be counted as pre-existing data that the operator is then told to `allow_non_empty_tables` over. ### Error message and/or stacktrace The `RESTORE` fails with the injected `filesystem error: in rename: Input/output error`; the retry into the now-non-empty table fails with `CANNOT_RESTORE_TABLE` and the message suggesting `allow_non_empty_tables=true` (`src/Backups/RestorerFromBackup.cpp`, `throwTableIsNotEmpty`). ### Additional context - No-rollback attach loop: `StorageReplicatedMergeTree::attachRestoredParts` — per-part `writeExistingPart(part, deduplicate_part=false)`, no compensation. - Distinct from #104464 (RESTORE leaves an orphan **empty** table after a mid-restore **crash** — opposite residue: no data; this finding is a no-crash I/O failure that leaves a **data** prefix plus retry duplication). #104464 being accepted as a bug is precedent that \"RESTORE is not atomic on failure\" is treated as a defect. - Distinct from #114266 (duplication via the emptiness guard reading local state on a lagging replica, where RESTORE **succeeds**) and #112149 (INSERT-sink cancellation during an ambiguous commit — data loss, not residue). - The `allow_non_empty_tables` duplication is documented for legitimately pre-existing data (`src/Backups/RestoreSettings.h`). The defect reported here is the un-rolled-back durable prefix from a **failed** restore; the documented duplication is the consequence the operator is steered into, not the primary bug.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114593",
          "createdAt": "2026-08-13T05:02:42Z",
          "updatedAt": "2026-08-13T05:02:42Z",
          "timestamp": "2026-08-13T05:02:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:97ab6c43b32d516b67a8",
        "signalId": "github:ClickHouse/ClickHouse:issue:114592",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114592",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Lightweight UPDATE: a mid-commit failure leaves a read-visible partial update and a retry double-applies",
          "text": "### Company or project name ClickHouse QA (internal durability testing) ### Describe what's wrong A multi-partition lightweight `UPDATE` is neither atomic on failure nor idempotent on retry. `UPDATE t SET c = ... WHERE ...` over a table with N partitions creates one patch part per partition and commits them sequentially (`MergeTreeSinkPatch::finishDelayedChunk` for `MergeTree`, `ReplicatedMergeTreeSinkPatch::finishDelayed` for `ReplicatedMergeTree`), each with `deduplicate = false` hard-set. If the commit of the k-th partition's patch part fails (for example a filesystem error while renaming the staging directory into place), the patch parts for partitions `1..k-1` are already Active and read-visible, yet the `UPDATE` returns an error. Two problems follow: 1. The failed `UPDATE` has partially applied and that partial state is immediately visible on `SELECT` — with no retry at all. 2. Because `deduplicate = false` is hard-set for patch parts, a client that retries the same `UPDATE` (the natural response to an error) re-applies the change to the partitions that already have it. There is no idempotency: the retried patch is a fresh part with a fresh block number. The net effect is a silent split double-apply: the partition that received the residue ends at `+2` while the partition that did not ends at `+1`, when the user issued a single `SET x = x + 1`. This is distinct from #112133 (please read the distinction before deduplicating on the title): #112133 concerns the mutation-entry commit — a single Keeper `tryMulti` whose response is lost, misreported as failure, so a retry creates a second mutation entry. That path applies a whole N-partition mutation as one atomic znode (all-or-nothing; no partial application — this is exactly why heavy `ALTER TABLE ... UPDATE` does not exhibit this). The finding here is a different code path (the per-partition patch-part INSERT sinks), a different fault (a clean, unambiguous commit failure — no lost Keeper response needed), and has a facet #112133 does not: the partial application is read-visible after the failed op even with zero retries. ### Does it reproduce on the most recent release? Yes — reproduced on a current `26.8.1.1` master build; the mechanism is present in the current source. It reproduces at both the wide-part layout and the default (compact) part layout. ### How to reproduce ```sql CREATE TABLE t (id UInt64, x UInt64) ENGINE = MergeTree PARTITION BY (id % 2) ORDER BY id SETTINGS enable_block_number_column = 1, enable_block_offset_column = 1; INSERT INTO t SELECT number, 0 FROM numbers(200); -- 100 rows in each of 2 partitions, x = 0 ``` Inject a filesystem I/O error (`EIO`) on the rename that commits the **second** partition's patch part (the second `tmp_insert_patch-*` → `patch-*` rename), then run: ```sql UPDATE t SET x = x + 1 WHERE x < 1000000; -- errors with a filesystem_error ``` Observe that the update failed but one partition already reflects the change: ```sql SELECT id % 2 AS p, sum(x) FROM t GROUP BY p ORDER BY p; -- p=0 -> 100 (patch committed before the fault: durable residue of a failed UPDATE) -- p=1 -> 0 ``` Now retry the same statement (fault cleared, as a client would): ```sql UPDATE t SET x = x + 1 WHERE x < 1000000; -- succeeds SELECT id % 2 AS p, sum(x) FROM t GROUP BY p ORDER BY p; -- p=0 -> 200 (double-applied) -- p=1 -> 100 ``` This was verified with a deterministic control-gated harness that arms the `EIO` on exactly the second patch-part commit rename. Controls: with no fault the single `UPDATE` yields a uniform `+1` on both partitions; a single-partition table under the same fault leaves no residue and the retry yields a uniform `+1` (so the split is attributable to the multi-partition fan-out). ### Expected behavior A failed `UPDATE` should apply to no partition (atomic), and a retry of the same `UPDATE` should be idempotent — the final state should be `x + 1` on every partition regardless of the mid-commit failure and the retry. ### Error message and/or stacktrace ``` Code: 1001. std::exception. Code: 1001, type: std::__1::filesystem::filesystem_error, e.what() = filesystem error: in rename: Input/output error [...patch-...] ``` ### Additional context - Commit loops: `src/Storages/MergeTree/MergeTreeSinkPatch.cpp` (`finishDelayedChunk`, per-partition loop) and `src/Storages/MergeTree/ReplicatedMergeTreeSinkPatch.cpp` (`finishDelayed`, per-partition loop). - `deduplicate = false` is hard-set for patch parts, and a deduplicated patch part is itself treated as a `LOGICAL_ERROR` (\"Patch part {} was deduplicated. It's a bug\"), so idempotent retry is unattainable by design — atomicity is the only possible containment, and it is absent. - Lightweight `UPDATE` is a beta feature; its documentation states the query waits for patch-part creation before returning but says nothing about failure atomicity or retry idempotency. - Reachability: the only non-default requirement is `enable_block_number_column` / `enable_block_offset_column`, which are the documented prerequisites for lightweight updates, not an experimental gate.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114592",
          "createdAt": "2026-08-13T05:02:41Z",
          "updatedAt": "2026-08-13T05:02:41Z",
          "timestamp": "2026-08-13T05:02:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:34c6df1c87404641f10a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114559",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114559",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Never cancel a query before its `max_execution_time` has elapsed",
          "text": "`CancellationChecker::appendTask` derived a query's cancellation deadline from the current time truncated to whole milliseconds, and then aligned it up to the 100 ms grid the worker batches deadlines on. The alignment normally hides the truncation, but when the deadline already sits on the grid - one query in a hundred - it adds no padding, and the deadline lands up to 1 ms before `now + max_execution_time`. The query is then cancelled ahead of its own timeout and fails with a self-contradictory message: ``` Code: 159. DB::Exception: Timeout exceeded: elapsed 999.672 ms, maximum: 1000 ms. (TIMEOUT_EXCEEDED) ``` The current time is now rounded up to a whole millisecond instead, which is the direction the code around it already promises: *\"Tasks may be cancelled slightly later than their exact timeout, but never before\"*. Only the `throw` overflow mode was affected: in `break` mode the checker re-reads the elapsed time before acting on the deadline, so an early wake-up does nothing. This is what makes `02122_join_group_by_timeout` flaky - it asserts that a `max_execution_time = 1` cancellation takes at least 1000 ms, and `query_duration_ms` is measured from a stopwatch that starts before the deadline is computed. The failure shows up as `query_duration 1` becoming `query_duration 0`, always on one of the two `throw`-mode queries and never on the `break`-mode ones. In the server logs of both `Fast test` runs below, the query was cancelled early - at 999.672 ms and at 999.630 ms of its 1000 ms timeout. The deadline computation moved into `CancellationChecker::taskDeadlineMs` so that the invariant can be tested directly: the new unit test checks every sub-millisecond phase of `now` against every phase of the grid, for a range of timeouts, and fails with the old truncation. CI reports of the flaky test: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-113021-93e2851fddb5add7bad5b28bcef780be2f6ff7bb&sha=bef88e3d96ecb0f57a0f8d81f6b7ee2117db19fd&name_0=MergeQueueCI&name_1=Fast%20test - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-110451-2f02abe52a79c02fb554728f14d8fa847dcbc1e6&sha=2eba84a8a12a21c36516158d4fbe9530d3f0af8c&name_0=MergeQueueCI&name_1=Fast%20test The truncation is as old as the checker itself. Before deadlines were aligned to a grid in https://github.com/ClickHouse/ClickHouse/pull/95563 the deadline was plainly `floor(now) + timeout`, i.e. early for every query rather than for one in a hundred; and the assertion that catches it was only sharpened from `round(query_duration_ms / 1000) BETWEEN 1 AND 60` to `query_duration_ms BETWEEN 1000 AND 60000` in daf1b4980b1, which is why the flake is recent. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A query could be cancelled up to 1 ms before its `max_execution_time` had passed, failing with a self-contradictory `Timeout exceeded: elapsed 999.672 ms, maximum: 1000 ms`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114559",
          "createdAt": "2026-08-13T00:25:45Z",
          "updatedAt": "2026-08-13T04:58:10Z",
          "timestamp": "2026-08-13T04:58:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "thevar1able"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fb7dc234c889e42aa0c9",
        "signalId": "github:ClickHouse/ClickHouse:issue:95963",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:95963",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Idempotency key for DDL operations",
          "text": "### Company or project name _No response_ ### Use case Retries after network connection errors are really complex in distributed systems. There are many operations that are quite risky to be retried such as `EXCHANGE TABLE` and we do not have easy way to check that previous operation succeeded or not for example, when disconnect was result of completely disappearing node and we have query log local to this node. ### Describe the solution you'd like - Add optional query setting `idempotency_key` that can be passed from a client with a query similar to `insert_deduplication_token`. - When it is passed operations that change state(or go through keeper) will commit the `idempotency_key` to a special node in keeper. - idempotency_key work the same way as insert_deduplication_token - it will be cleaned up with the similar window logic. - Client libraries can set this key automatically and retry on network error preserving `idempotency_key`. It will make most retries completely safe. - Server will check if key exists before running query and also include it during operation commit to keeper. This way you can also cancel one of 2 parallel EXCHANGE TABELS so only one will succeed. ### Describe alternatives you've considered Other option is to pass some key similar to query_if from a client and on error check that this query was not running before. Saving this information somehow to keeper with operations that change metadata in keeper will allow to later check if we had one already finished. But we do not have guarantee that there is none inflight with unreachable node, that is still connected to keeper. So it will add more complexity if we can not store a tombstone saying this operation should not succeed after you checked. ### Additional context Language clients suffer a lot of disconnects as sometimes there is a race for connections to be not closed for up to 5 minutes: https://github.com/ClickHouse/ClickHouse/issues/92046 They can't easily detect if this case was there i.e. no bytes were sent from client and just ordinary disconnect when network is bad.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/95963",
          "createdAt": "2026-02-04T16:27:50Z",
          "updatedAt": "2026-08-13T04:48:48Z",
          "timestamp": "2026-08-13T04:48:48Z",
          "metrics": {
            "reactions": 2,
            "comments": 1
          },
          "labels": [
            "feature"
          ],
          "author": "qoega",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:105f93f9bc28fd8fada8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109454",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109454",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Materialize column statistics on INSERT for small tables by default",
          "text": "Cost-based join reordering relies on column statistics (in particular the number of distinct values of join keys) to avoid pathological plans. Column statistics are auto-declared by default (`auto_statistics_types = 'minmax, uniq'`), but they are only materialized during merges (`materialize_statistics_on_merge`) — so a freshly bulk-loaded table has no usable statistics at query time until it happens to be merged. This bites the TPC-H benchmark. At scale factor 40 on a 32 GB machine, Q5 and Q8 time out (>100s, spilling to disk): - Without the NDV of the low-cardinality `nationkey` (25 distinct values), the optimizer cannot see that joining `customer` and `supplier` on `nationkey` (via the transitive `c_nationkey = s_nationkey`) produces a ~19 billion row many-to-many intermediate, and it materializes that as the hash-join build side. - The dimension tables that carry the decisive join keys are small (SF40: `customer` 509 MiB, `supplier` 32 MiB, `region`/`nation` tiny), so materializing their statistics on insert is cheap — but the fact tables (`lineitem` 11 GiB, `orders` 2.5 GiB) are large and should not pay per-insert statistics cost. This PR enables `materialize_statistics_on_insert` by default, bounded by a new setting `materialize_statistics_on_insert_max_table_size` (default 25 GiB): tables whose current size is at or below the threshold build statistics at insert time, larger tables skip it and materialize during merges as before. The size check is an `O(1)` atomic load of the table's current active size; `0` disables the limit. Measured on TPC-H SF40 with the server memory capped at 28.8 GiB (0.9 × 32 GiB), out of the box (no manual `MATERIALIZE STATISTICS`): | Query | Before | After | | --- | --- | --- | | Q5 | >200s, 8.6 GB spilled to disk | 0.5s, 636 MiB, no spill | | Q8 | 143s, 4.35 GB spilled to disk | 0.6s, 1.37 GiB, no spill | The per-insert cost is gated correctly: inserting into a table already above the threshold builds no statistics (`MergeTreeDataWriterStatisticsCalculationMicroseconds = 0`), while an uncapped insert of the same block spends the usual time. ## Validation on other benchmarks To check the effect more broadly, I compared having statistics available at load time (this change) against not having them (`use_statistics=0`, the pre-change planning state), on the ClickBench `/versions` query sets, on the same freshly-loaded data under the 32 GiB cap: - **TPC-DS (SF40, 103 queries):** 4 queries go from OOM/timeout to succeeding (q24, q51, q71, q93); many large speedups (q74 98×, q66 31×, q7 22×, q83 18×, …); ~35% faster in aggregate; no genuine regressions. - **JOB / IMDB (113 queries):** net ~12% faster; wins up to ~8× (q1, q91, q92). The dataset is small enough that nothing fails either way; a couple of queries regress mildly (q55, q26). - **Coffee Shop (fact 359M rows + 2 dimensions, 17 queries):** wins up to ~50× (q8 42×, q11 and q16 go from timeout to 1.4s / 63s, q13 15×); one query regresses ~2.3× (q17). Overall, making dimension statistics available out of the box is strongly net-positive — many multi-× speedups and several queries rescued from OOM/timeout. A small number of queries regress because the cost-based join reorder occasionally makes a worse choice with statistics than without; that is a pre-existing optimizer-quality concern that enabling statistics by default surfaces more often, not a regression in the materialization itself. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Column statistics are now materialized on `INSERT` by default when the table's current active size plus the written block size is at most the new `materialize_statistics_on_insert_max_table_size` setting (default 25 GiB); the check is per written block, so a first bulk load into an empty table may still materialize statistics for each written block. This gives the cost-based join optimizer accurate estimates for freshly-loaded dimension tables and avoids pathological join orders (for example, TPC-H Q5 and Q8 no longer time out at scale factor 40), while large established fact tables keep materializing statistics during merges. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1307` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109454",
          "createdAt": "2026-07-05T21:20:22Z",
          "updatedAt": "2026-08-13T04:46:37Z",
          "timestamp": "2026-08-13T04:46:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 44
          },
          "labels": [
            "pr-performance",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0610b7cabe1da14fc583",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113984",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113984",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix logical error on filter push-down with a differently-typed same-name column",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix a logical error (`Unexpected return type`) when a filter over a view whose column type differs from the underlying storage (e.g. `information_schema.tables`, where `engine` is `Nullable(String)`, over `system.tables`, where it is `String`) is pushed down to the storage. Minimal reproducer: ```sql SELECT 1 FROM (SELECT 1 FROM information_schema.tables WHERE indexHint(toString(engine))); -- Code: 49. DB::Exception: Unexpected return type from toString. Expected Nullable(String). Got String. (LOGICAL_ERROR) ``` The predicate is typed against the view header, where `engine` is `Nullable(String)`. When `SourceStepWithFilter::applyFilters` rebuilds the filter with `ActionsDAG::buildFilterActionsDAG`, the name-based input replacement substituted the storage column `engine String` for the input while the parent `FUNCTION` nodes were rebuilt with their existing `function_base` — leaving the rebuilt DAG (including the DAG captured inside `indexHint`) internally inconsistent: `toString` still declared `Nullable(String)` while returning `String`. Evaluating it over the candidate-tables block in `getFilteredTables` then failed the return-type assertion in `ExpressionActions`. Two fixes: - `ActionsDAG::buildFilterActionsDAG`: skip a name-based input replacement when it would change the input type. Keeping the original input is safe — the subtree is then simply not evaluated over the storage columns, and the filter is still applied upstream. - `canEvaluateSubtree` in `VirtualColumnUtils`: match the allowed inputs by type as well as by name, as a fail-close guard at the evaluation site, so a mismatched predicate subtree is not pushed down. The exact-type check exposed a lying push-down sample in `StorageSystemTables`: both `detail::getFilteredTables` and `ReadFromSystemTables::applyFilters` declared `uuid` as `String` while the real column of `system.tables` / `system.detached_tables` is `UUID`, which would have silently disabled the `uuid` prefilter fast path. The samples now declare `UUID`, and a test pins the fast path via the `SelectedRows` profile event. Found by the AST fuzzer on two unrelated PRs (pre-existing on `master`; the same error also appears in `master` stress tests going back months): [AST fuzzer (amd_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113636&sha=bea222ed980b6db96146b4fe56dc968fa2280dd1&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29). Closes: https://github.com/ClickHouse/ClickHouse/issues/113982 Related: https://github.com/ClickHouse/ClickHouse/pull/113636 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1195` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113984",
          "createdAt": "2026-08-08T21:39:30Z",
          "updatedAt": "2026-08-13T04:46:36Z",
          "timestamp": "2026-08-13T04:46:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:310bf1f1710760cdf664",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101783",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101783",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support specifying or auto-assigning Parquet field_ids for output columns",
          "text": "### What this PR does Two related knobs for writing Parquet files with column `field_id` metadata — needed for Apache Iceberg compatibility, which identifies columns by `field_id` rather than by name. Both knobs can be used independently or together. #### 1. Explicit per-column overrides New setting `output_format_parquet_column_field_ids : Map(String, Int32)`. ```sql SET output_format_parquet_column_field_ids = '{\"col_a\": 10, \"col_b\": 20, \"col_c\": 30}'; INSERT INTO FUNCTION file('data.parquet') SELECT 1::UInt32 AS col_a, 'x'::String AS col_b, 42::Int64 AS col_c; ``` Each output column gets the id specified in the map. The map must cover every output column (unless auto-assign below is enabled to fill the gaps), and ids must be unique; violations raise `BAD_ARGUMENTS` with specific messages. #### 2. Auto-assign (Iceberg writer convention) New setting `output_format_parquet_auto_assign_field_ids : Bool`, default `false`. When enabled, every output column is assigned a unique sequential `field_id` starting at `1`, matching the convention used by Apache Iceberg writers. ```sql SET output_format_parquet_auto_assign_field_ids = 1; INSERT INTO FUNCTION file('iceberg.parquet') SELECT 1 AS a, 'x' AS b, 42 AS c; -- a -> 1, b -> 2, c -> 3 ``` #### 3. Mixed mode The two settings compose: override-map entries win for the columns they mention; auto-assign fills the remaining columns with the smallest unused positive ids. ```sql SET output_format_parquet_auto_assign_field_ids = 1, output_format_parquet_column_field_ids = '{\"b\": 1}'; INSERT INTO FUNCTION file('mixed.parquet') SELECT 1 AS a, 2 AS b, 3 AS c; -- b -> 1 (override), a -> 2, c -> 3 (auto-assign skipping 1) ``` ### Implementation - `src/Core/FormatFactorySettings.h` — declare both settings. - `src/Formats/FormatSettings.h` — carry a `std::vector<std::pair<String, Int32>>` of overrides and the `auto_assign_field_ids` bool; no string parsing on the hot path. - `src/Formats/FormatFactory.cpp` — convert the `Map` setting into the pair vector, with structural validation (tuple shape, string keys, integer values in `Int32` range). - `src/Processors/Formats/Impl/ParquetBlockOutputFormat.cpp` — `buildColumnFieldIds()` resolves overrides against the actual output header, auto-assigns unused ids if the flag is on, and rejects unknown columns, duplicate ids, and non-covering overrides. ### Test `tests/queries/0_stateless/04321_parquet_column_field_ids.sh` writes Parquet files with `clickhouse-local`, reads the `field_id` metadata back with `pyarrow`, and covers: - Explicit per-column overrides. - Auto-assign only. - Mixed override + auto-assign. - Writing with neither setting (no `field_id`s, unchanged behavior). - Nested types: auto-assign and dotted-path overrides for `Array.element`, `Tuple` subfields, and `Map.key`/`Map.value`. - Geo columns written as WKB under GeoParquet (the `Point`/`Array`/`Tuple` shape collapses to a single top-level field). - Error cases: unknown column, non-covering map (top-level and nested), duplicate id, non-integer value, negative id, and a dotted top-level name colliding with a nested path (in both auto-assign and override modes). Closes: https://github.com/ClickHouse/ClickHouse/issues/58753 ### Changelog category (leave one): - New Feature ### Changelog entry: Added `output_format_parquet_column_field_ids` (`Map(String, Int32)`) to set explicit Parquet `field_id` overrides per column and `output_format_parquet_auto_assign_field_ids` (`Bool`) to auto-assign sequential `field_id`s to every output column, matching the Apache Iceberg writer convention. The two settings can be combined. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches the Parquet write path and schema generation to emit `field_id` metadata, which could affect interoperability and schema expectations; defaults are preserved unless new settings are enabled. > > **Overview** > Adds two new Parquet output settings to control column `field_id` metadata for Iceberg compatibility: `output_format_parquet_column_field_ids` (explicit per-column overrides) and `output_format_parquet_auto_assign_field_ids` (sequential auto-assignment starting at 1). > > `FormatFactory` now parses/validates the `Map` setting into typed overrides up-front, while `ParquetBlockOutputFormat` resolves the final `field_id` mapping (including error checks for unknown columns, duplicates, negative ids, and incomplete coverage when auto-assign is off) and explicitly rejects using these settings when a datalake writer provides its own column-id mapping. > > Documentation and a new stateless test are added to verify correct `field_id` emission via `pyarrow` and expected failure modes. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit ce5f371493174b5b0ee58307dfc05f5df32af0b2. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101783",
          "createdAt": "2026-04-04T18:49:42Z",
          "updatedAt": "2026-08-13T04:43:18Z",
          "timestamp": "2026-08-13T04:43:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 43
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested"
          ],
          "author": "Onyx2406",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:01cc1da32adb39cd88e3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114451",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114451",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Show the real per-query verdict in the performance comparison report",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/111992 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description A `Performance Comparison` shard that PASSED renders a list of red rows, and every affected query is listed twice. Reported in #111992: shard `arm_release, master_head, 1/6` was `OK`, its own message said `11 unstable`, and the report showed 22 red rows that had to be hand-verified. Two causes, both in the block that builds the `Check Results` rows: - The rows exist only to carry a per-query \"query history\" link, and each was built with a hardcoded `FAIL`, discarding the `slower`/`unstable` verdict into `info`. The renderer already classifies both natively, so passing the real one through is enough. - `compare.sh` emits one row per side (`array join map('old', left, 'new', right)`, named `<query> #<idx>::old` / `::new`), so 11 queries became 22 rows. Only the candidate side is now listed, and only its displayed name is stripped. `compare.sh` is unchanged: its per-side rows are what goes into the CI database, so changing the emission would fork that history and break every existing history link. Nothing is dropped, only displayed once. This cannot turn a passing shard red: `Check Results` is created with an explicit status, and praktika aggregates child statuses only when none is given. A test pins that guard, and the slower-query gate that decides the job verdict is untouched. One more line in `json.html`: `getStatusClass` had no `slower` case and fell through to gray, while the status filter and the row ordering both already treat it as a failure. Tests: `ci/tests/test_perf_check_results_children.py`, driven by two trimmed real shard artifacts - the one from #111992, and a failing `release_base` shard as the negative control, which must stay red. The parent's own message states the query count, an independent oracle for the row count. The CI database rows are byte-identical before and after, and the history links still resolve.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114451",
          "createdAt": "2026-08-12T09:11:17Z",
          "updatedAt": "2026-08-13T04:42:48Z",
          "timestamp": "2026-08-13T04:42:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "maxknv",
            "egor-click"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:98220b0bc20f584746a1",
        "signalId": "github:ClickHouse/ClickHouse:issue:101474",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:101474",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Vector similarity: support TurboQuant quantization",
          "text": "### Company or project name _No response_ ### Use case [TurboQuant, a quantization method](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/) that perhaps could be applied to ClickHouse to improve Approximate Nearest Neighbor (ANN) searches. technical details are somewhat beyond my expertise ### Describe the solution you'd like New TurboQuant quantization option for ANN indexes, something similar to the existing quantization QBit ### Describe alternatives you've considered _No response_ ### Additional context - [Blog post](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/) - [Paper](https://arxiv.org/pdf/2504.19874)",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/101474",
          "createdAt": "2026-04-01T08:55:13Z",
          "updatedAt": "2026-08-13T04:41:17Z",
          "timestamp": "2026-08-13T04:41:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "feature"
          ],
          "author": "Wachynaky",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:afa6ab637e51bf626413",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107960",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107960",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cache vended credentials for REST catalogs",
          "text": "Now, new vended credentials are requested on each metadata request. This PR adds an (optional) cache for creds with configurable TTL. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add option to cache vended credentials for REST catalogs; add a setting `vended_credentials_cache_ttl` (seconds). 300 by default. 0 means no caching.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107960",
          "createdAt": "2026-06-19T12:17:13Z",
          "updatedAt": "2026-08-13T04:40:44Z",
          "timestamp": "2026-08-13T04:40:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement",
            "manual approve",
            "can be tested"
          ],
          "author": "zvonand",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:da4ad1df43fa21a0ac4d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114585",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114585",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #107028 to 26.5: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114585",
          "createdAt": "2026-08-13T03:41:21Z",
          "updatedAt": "2026-08-13T05:48:10Z",
          "timestamp": "2026-08-13T05:48:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "kssenii",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c6ac2bebf8c1db4cf9ce",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114568",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114568",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `/dialect`, `/lang`, `/language` client commands",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added client-side commands `/dialect`, `/lang`, `/language` to `clickhouse-client` and `clickhouse-local` — an equivalent of `SET dialect = '...'` that is handled before parsing, so it also works when the current dialect cannot express a `SET` query (for example, to leave `kusto` or `promql`). While a non-default dialect is active, it is displayed in the prompt after the server display name, e.g. `clickhouse-cloud (polyglot) :)`. The command without an argument prints the current dialect. A plain `SET dialect = '...'` also updates the prompt, since the prompt now renders the dialect from the client-side session settings on every line. Implementation notes: - The prompt is now rendered by `getPrompt` on every line from a template that keeps the `{display_name}` placeholder: the current dialect is appended to the display name in parentheses (with a separating space if the display name is non-empty), and the `:) ` smiley is appended at the end, preserving the previous behavior for the default dialect. - The command works in `clickhouse-client`, `clickhouse-local`, and the embedded client (SSH/web terminal). - Documentation: a new \"Switching the SQL dialect\" section in the client page. - Test: `04869_client_dialect_command.expect`, verified locally along with manual checks against a running server (prompt display, alias forms, escaping from `kusto` back to `clickhouse`, and that a Kusto query parses after `/dialect kusto`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114568",
          "createdAt": "2026-08-13T01:59:34Z",
          "updatedAt": "2026-08-13T04:36:32Z",
          "timestamp": "2026-08-13T04:36:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:16974b180c5619267907",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114106",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114106",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add table schema to SQLInsert output",
          "text": "Adds the `output_format_sql_insert_include_table_schema` setting for the `SQLInsert` output format. When enabled, `SQLInsert` prepends a `CREATE TABLE` statement derived from the query result schema, making the generated SQL directly replayable into an empty database. The table definition uses the result column names and types with `MergeTree` and `ORDER BY tuple()`. Source-table metadata such as keys, default expressions, codecs, TTLs, and indexes is not preserved. The setting is disabled by default and cannot be enabled together with `output_format_sql_insert_use_replace`. The change includes embedded documentation and stateless coverage for replaying the generated SQL, complex types, identifiers, empty results, batching, and incompatible settings. Closes: https://github.com/ClickHouse/ClickHouse/issues/84736 ### Changelog category: - Improvement ### Changelog entry: The `SQLInsert` output format can now prepend a `CREATE TABLE` statement with result column names and types. Enable `output_format_sql_insert_include_table_schema` to produce a ClickHouse SQL script that recreates the table and loads the exported data.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114106",
          "createdAt": "2026-08-10T04:55:57Z",
          "updatedAt": "2026-08-13T04:31:01Z",
          "timestamp": "2026-08-13T04:31:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "mishok2503",
          "state": "open",
          "assignees": [
            "nikitamikhaylov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:da2b4e6dcd326457276d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107436",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107436",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Automatic LowCardinality serialization based on uniq statistics",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/85145 --> This PR adds automatic `LowCardinality` serialization for `MergeTree`, analogous to the existing automatic `Sparse` serialization. A `String` or `FixedString` column whose declared type is **not** `LowCardinality` can now be stored on disk in dictionary-encoded form when it has a `uniq` statistic with a low cardinality estimate. The encoding is transparent: the column keeps its declared data type, flows through the query pipeline as a (non-native) `ColumnLowCardinality`, and is materialized to a full column at the boundaries that require it. **Motivation:** save disk space and read I/O for low-cardinality string columns without requiring users to declare the column type as `LowCardinality(...)`. **How it works** - A new per-part serialization kind `LOW_CARDINALITY` (shown as `LowCardinality` in `system.parts_columns.serialization_kind`). - The decision is made at write time from the **existing** `uniq` statistic — there are **no changes to statistics serialization**. In `MergeTreeDataWriter`, a `String`/`FixedString` column whose `uniq` estimate does not exceed the new `MergeTree` setting `max_uniq_number_for_low_cardinality` (default `0` = disabled), and which was not already chosen for `Sparse`, is stored with `SerializationLowCardinality`. The column must therefore declare `STATISTICS(uniq)` and have `materialize_statistics_on_insert` enabled. - A new `is_native` flag on `ColumnLowCardinality` makes `getDataType` report the nested type, so the \"column class matches data type\" invariant holds while such a column flows through the pipeline. It is materialized to a full column in `removeSpecialRepresentations` (aggregation, joins, squashing, blocks), in the merge-input step of `IMergingAlgorithm`, in the sorting and sparse-removing transforms, and in `IFunction`. The `optimize_functions_to_subcolumns` pass skips columns stored as non-native `LowCardinality`, since the dictionary encoding does not expose the data type's regular subcolumns. - The kind is independent of sparse serialization: `SerializationInfo` entries for eligible columns are created in the writer, in merges, in mutations and in the per-table serialization hints even when `ratio_of_defaults_for_sparse_serialization = 1` disables `Sparse` entirely. **Notes** - Opt-in: disabled by default and requires an explicit `uniq` statistic, so existing tables are unaffected. Requires `allow_experimental_statistics`. - Reading a subcolumn of an encoded column (for example `s.size`) is not supported and reports `NOT_IMPLEMENTED`: the subcolumn's streams do not exist in a dictionary-encoded part. Making it read the full column and extract the subcolumn in memory needs a dedicated serialization and is left for a separate change. - Extracts and reimplements the automatic `LowCardinality` part of the draft #85145 against current `master`, without the statistics-serialization refactoring. That refactoring is still unmerged (#85145 and #109128 are open drafts). Validated locally with a server built from this branch: kind selection for `Wide` and `Compact` parts, `String` and `FixedString`, query correctness, functions, `GROUP BY`/`ORDER BY`/`DISTINCT`/`LIMIT BY`/window functions, joins, `IN` subqueries, `INSERT SELECT`, `UNION ALL`, `OPTIMIZE FINAL` (LowCardinality preserved through merges), mutations, `ALTER RENAME`/`MODIFY COLUMN`, `DETACH`/`ATTACH`, `PREWHERE`, input/output formats with and without `allow_special_serialization_kinds_in_output_formats`, and tables with mixed `Default` and `LowCardinality` active parts. The three tests also pass with `ratio_of_defaults_for_sparse_serialization` forced to `0.0` and `1.0`, with `enable_block_offset_column`/`enable_block_number_column`, with compact parts, and over repeated fully randomized runs. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added automatic `LowCardinality` serialization: a `String`/`FixedString` column that has a `uniq` statistic and low cardinality can be stored in dictionary-encoded form on disk while keeping its declared data type. Controlled by the new `MergeTree` setting `max_uniq_number_for_low_cardinality` (disabled by default). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107436",
          "createdAt": "2026-06-14T04:04:43Z",
          "updatedAt": "2026-08-13T04:27:33Z",
          "timestamp": "2026-08-13T04:27:33Z",
          "metrics": {
            "reactions": 1,
            "comments": 6
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1a73a6287e724dc345bb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101039",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101039",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add leader election for non-replicated MergeTree on shared storage",
          "text": "Add leader election for non-replicated MergeTree tables on shared object storage (currently `S3`; `Azure` is implemented but not yet enabled, pending test coverage), enabling active/standby failover without external coordination (no Keeper). Uses conditional writes (`If-Match` / `If-None-Match`) on object storage to maintain a lease file with JSON content `{\"version\":1,\"leader_id\":\"...\",\"timestamp\":...}`. The leader renews its lease periodically; followers monitor and claim leadership when the lease expires. New MergeTree settings: - `leader_election` (Bool, default false) — enable leader election - `leader_election_heartbeat_interval` (Seconds, default 10) — lease renewal interval - `leader_election_session_timeout` (Seconds, default 30) — lease expiry threshold; must be at least 3x the heartbeat interval When not the leader, inserts, merges, mutations, and DDL are blocked; background data processing is skipped. Participating nodes should keep their clocks synchronized (NTP) to within `leader_election_session_timeout`; the conditional-write protocol always prevents split-brain, but excessive clock skew can cause leadership churn. Closes https://github.com/ClickHouse/ClickHouse/issues/91613 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `leader_election` setting for non-replicated MergeTree tables on shared object storage, enabling active/standby failover using conditional writes without external coordination. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) Three new MergeTree-level settings are added: - `leader_election` (Bool) — When enabled on a non-replicated MergeTree table stored on an `S3` object storage disk (`Azure` is implemented but rejected at table creation until it has test coverage), multiple ClickHouse instances sharing the same data path will elect a single leader. Only the leader performs writes, merges, and mutations. Followers act as read-only replicas and automatically claim leadership when the current leader's lease expires. - `leader_election_heartbeat_interval` (Seconds, default 10) — How often the leader renews its lease and followers check for an expired lease. - `leader_election_session_timeout` (Seconds, default 30) — How long a lease remains valid without renewal. Must be at least 3x `leader_election_heartbeat_interval`. Example: ```sql CREATE TABLE shared_table (x UInt64) ENGINE = MergeTree ORDER BY x SETTINGS leader_election = true, leader_election_heartbeat_interval = 10, leader_election_session_timeout = 30; ``` <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > High risk because it introduces new leader-election coordination and modifies core `MergeTree` write/DDL/drop and catalog cleanup paths; bugs could cause write unavailability or accidental deletion/incorrect visibility on shared storage. > > **Overview** > Adds a **Beta** `leader_election` mode for non-replicated `MergeTree` tables on shared S3/Azure, using a conditional-write lease file to ensure only one instance performs inserts/merges/mutations while followers remain read-only and can take over on failure. > > Introduces new MergeTree settings (`leader_election`, `leader_election_heartbeat_interval`, `leader_election_session_timeout`) with validation, new `MergeTreeLeaderElection*` metrics/events, and follower part-refresh + leader takeover sync to load new parts and advance block counters before enabling writes. > > Hardens destructive operations for shared-storage tables: blocks most `ALTER`/partition-mutation operations and `RENAME` under leader election, gates background processing on leadership, and updates drop/cleanup logic (`dropSkipsDataDirectoryCleanup`, `DatabaseCatalog`/`StorageTableProxy` fail-closed behavior) to avoid recursive deletion or hangs when tables are shared or cannot be materialized. Adds documentation and integration/stateless tests covering failover, metrics, validation, and rejection cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 78c1252915ba512bebf94a147932158e1bacecf0. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101039",
          "createdAt": "2026-03-28T21:56:21Z",
          "updatedAt": "2026-08-13T04:27:15Z",
          "timestamp": "2026-08-13T04:27:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 87
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4937a87f4de5253ec230",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114310",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114310",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113563 to 26.6: Do not analyze the shared row policy AST in place in `Merge`",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113563 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31486886135/job/93764217930)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114310",
          "createdAt": "2026-08-11T11:50:56Z",
          "updatedAt": "2026-08-13T04:26:26Z",
          "timestamp": "2026-08-13T04:26:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-backport",
            "pr-critical-bugfix"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [
            "KochetovNicolai",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b328aa3a829211a8ca48",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111720",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111720",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Prepare changelog for 26.8",
          "text": "Automated daily preparation of `CHANGELOG.md` for the upcoming 26.8 release. Every day the `NightlyChangelog` CI job appends the raw changelog entries for the pull requests newly merged into `master` (generated with `utils/changelog/changelog.py`) as one commit, and edits them following `.claude/skills/edit-changelog/SKILL.md` as a separate commit, so both the raw and the edited state stay reviewable. The point up to which entries were generated is recorded as a `Changelog-generated-up-to:` trailer in the generate commits. This pull request stays a draft until the release. The release manager finalizes it manually: fills in the release date and the presentation/video links (the `FIXME` placeholders), reviews the entries, and marks it ready. ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111720",
          "createdAt": "2026-07-24T04:09:28Z",
          "updatedAt": "2026-08-13T04:18:24Z",
          "timestamp": "2026-08-13T04:18:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "do not test",
            "pr-not-for-changelog"
          ],
          "author": "clickhouse-gh[bot]",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6ce1ce4dacaa0333d286",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114564",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114564",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112921 to 26.7: Evaluate randomHadamardTransform once for a constant vector",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112921 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31655073318/job/94307646536)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114564",
          "createdAt": "2026-08-13T00:58:05Z",
          "updatedAt": "2026-08-13T04:08:09Z",
          "timestamp": "2026-08-13T04:08:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:28d50e0e81f77a25ede8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109369",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109369",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support Variant type in aggregate functions (sum, avg, min, max, ...)",
          "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Aggregate functions that do not handle the `Variant` type natively (`sum`, `avg`, `min`, `max`, `argMin`, `argMax`, `quantile`, `stddevPop`, ...) can now be applied to `Variant` arguments. Previously they failed, e.g. `Illegal type Variant(Decimal(7, 2), Float64) of argument for aggregate function sum`. The aggregate functions that do accept a `Variant` argument natively (`count`, `any`, `groupArray`, `groupConcat`, the `uniq` family, ...) now skip the rows where the `Variant` holds a NULL value, like they skip the NULL values of a `Nullable` argument (only the row skipping matches `Nullable`: the result type and the empty-set result of the function are preserved) -- previously `count` counted those rows, `any` could return NULL from a group that has non-NULL values, `groupArray` stored the NULLs, and the `uniq` family counted NULL as a distinct value. Set `aggregate_functions_skip_variant_nulls = 0` (or `SET compatibility = '26.7'`) to restore the previous behavior. Persisted `AggregateFunction(f, Variant(...))` states (materialized views, `AggregateFunction` columns) keep the same type name, layout and serialization in both modes, so states written before and after the upgrade merge freely but mean different things -- the old ones aggregated the NULL rows, the new ones skip them -- and a merged result over the mix follows neither rule consistently. To keep the pre-`26.8` results over such data, set `aggregate_functions_skip_variant_nulls = 0` before inserting new data; the NULL rows already aggregated into the historical states cannot be removed by re-merging, only by re-inserting the data from the source. ### Description Such functions now aggregate over the least common supertype of the variants, wrapped in `Nullable`: ``` f(variant) == f(CAST(variant AS Nullable(supertype(T1, ..., TN)))) ``` `Nullable` preserves the implicit NULLs of the `Variant`, which the aggregation skips. **How it works.** `AggregateFunctionFactory::get()` resolves the function normally first; only if that fails for a `Variant` argument with an \"unsupported argument type\" error — `ILLEGAL_TYPE_OF_ARGUMENT`, or `NOT_IMPLEMENTED` for the creators that use it instead (`rankCorr`, `mannWhitneyUTest`, `kolmogorovSmirnovTest`, `largestTriangleThreeBuckets`) — does it fall back to a new `AggregateFunctionVariantAdapter`, which casts the `Variant` argument column(s) to `Nullable(supertype)` on the fly and forwards everything else to the nested function (whose state layout it shares). If the supertype route is not possible either, the original error is reported unchanged. Functions that already accept `Variant` (`count`, `any`, `uniq`, `groupArray`, ...) are unaffected. **Supertype.** `getLeastSupertype` is strict and has no common type for e.g. `Decimal`/`Int64` ↔ `Float64` (no lossless conversion). For an aggregate whose result is a floating-point value computed by arithmetic over its input (the `sum`/`avg`/variance/... families), a mix of numeric variants with no lossless common supertype falls back to `Float64` — but only when the user has opted into the lossy numeric supertype with the `allow_lossy_numeric_supertype` setting, and under the same promotion rule that setting applies at type inference for `if`/`multiIf`/`coalesce`/`ifNull`/`array`/`map` (all-numeric variants with at least one floating-point member; an integer-only mix such as `Variant(Int64, UInt64)` is not promoted). With the setting off (the default) the adapter handles only lossless supertypes and the original `ILLEGAL_TYPE_OF_ARGUMENT` error — which points at the setting when enabling it would help — is reported unchanged. Reconstruction of an already declared state type such as `AggregateFunction(sum, Variant(Int64, Float64))` always allows the promotion, so a state validated when it was declared never becomes unreadable because of the setting's current value: this covers the background paths that have no query context (table load / `ATTACH` at startup, merges) and, through `from_declared_state_type`, every query-time path that names the state type explicitly — a column type, a `CAST` target, a decoded binary type, and the nested function of `-Merge` over such a state. The fallback is deliberately **not** applied to exact/order-based aggregates (`min`, `max`, `argMin`, `argMax`, `any`, `quantileExact`, ...) even under the setting: a lossy `Float64` cast would silently return wrong results for them (two distinct integers above 2^53 collapse to the same `Float64`), so they keep reporting the original `ILLEGAL_TYPE_OF_ARGUMENT` when there is no lossless common supertype. Whether a function is float-promoting is declared by `AggregateFunctionProperties::is_float_promoting` (set at registration, consulted through the factory), fail-closed: a function that does not set it is treated as not float-promoting. Clean supertypes are preserved regardless of the function and the setting (`Variant(UInt8, UInt32) -> UInt32`, `Variant(Int32, Float64) -> Float64`, `Variant(Date, DateTime) -> DateTime`), so `min`/`max` still work over such variants. **Combinators.** The adapter is applied as the outermost wrapper (the combinator recursion resolves nested functions without it), so combinators compose in the usual order, e.g. `Adapter(Null(If(sum)))`. `-If`, `-State`/`-Merge` and `GROUP BY` work; state serialization for `sumState` uses the supertype (`AggregateFunction(sum, Nullable(Float64))`), so distributed / two-phase aggregation is unchanged. Only the argument positions the function actually rejects are adapted: `argMin`/`argMax` (and the `*ArgMin`/`*ArgMax` combinators) accept a `Variant` in the returned \"arg\" position and reject it only in the comparable key, so the \"arg\" keeps its original `Variant` type while just the key is cast to `Nullable(supertype)`. **Examples.** ```sql SET allow_lossy_numeric_supertype = 1; SELECT sum(v), toTypeName(sum(v)) FROM values('v Variant(Decimal(7, 2), Float64)', 1.5, 2.5, NULL, 10); -- 14 Nullable(Float64) ``` **Scope.** Only top-level `Variant` arguments whose common supertype can be wrapped in `Nullable` are handled — the adapter uses `Nullable` to carry the implicit NULLs of the `Variant`. A `Variant` whose common supertype is a container type (`Array`/`Tuple`/`Map`) is therefore still rejected, even for an orderable aggregate such as `min`/`max` over `Variant(Array(UInt8), Array(UInt16))` (supporting it would require tracking the `Variant`'s NULLs separately from the value column). A `Variant` nested inside `Tuple`/`Array`, and the `Dynamic` type, are likewise still rejected as before. These are all natural follow-ups. The result is `Nullable` because a `Variant` value can always be NULL. `singleValueOrNull` is deliberately excluded from the adapter (`AggregateFunctionProperties::is_distinctness_sensitive`): its contract keys on how many distinct values there are, and the cast to `Nullable(supertype)` collapses `Variant` values that are distinct because their alternative types differ (`1::UInt8` vs `1::UInt64`, which `uniq` counts as 2), so it would silently return a non-NULL value where the contract requires NULL — including inside the `x = ALL (SELECT ...)` rewrite. It keeps reporting `ILLEGAL_TYPE_OF_ARGUMENT` for a `Variant` argument, unchanged from before. The `-Distinct` combinator makes any combined function distinctness-sensitive in the same way (`sumDistinct` must deduplicate the genuine `Variant` values, not their casts to the supertype), so the combinator propagates the property through `AggregateFunctionFactory::tryGetProperties` (`IAggregateFunctionCombinator::isDistinctnessSensitive`) and every `...Distinct` form over a `Variant` argument keeps its original error too. `groupArrayInsertAt` and `groupArraySorted` do not claim native `Variant` support either: their generic implementations keep the state as `Field`s, which loses the original alternative type of a `Variant` value on ingest and reinfers the first compatible one on output (`1::UInt8` and `1::UInt64` collapse), and `groupArraySorted` would order by `Field` comparison instead of `Variant` order. Like `sum` / `avg`, they go through the adapter over the least common supertype of the variants, and reject a `Variant` (or `Dynamic`) argument at resolution when there is no lossless supertype. **Error codes.** `BAD_ARGUMENTS` is deliberately not part of the \"unsupported argument type\" retry signal, because creators use it for genuine semantic failures — above all invalid parameters, as in `kolmogorovSmirnovTest('bogus')` — and retrying on it would report the unrelated type error of the original, unadapted call instead of the parameter error. The few creators that rejected argument *types* with `BAD_ARGUMENTS` now reject them with `ILLEGAL_TYPE_OF_ARGUMENT`, like every other function, which keeps them adaptable to a `Variant` argument: `analysisOfVariance`, `studentTTest`, `studentTTestOneSample`, `welchTTest`, `meanZTest` and `boundingRatio` report `ILLEGAL_TYPE_OF_ARGUMENT` instead of `BAD_ARGUMENTS` when applied to an argument of an unsupported type. **`count` over a `Variant`.** `count` accepts a `Variant` natively, without the adapter, but `count(expr)` counts the not-NULL values of its argument, and a `Variant` row can hold a NULL value (`isNull` is true for it). A `Variant` is not `Nullable`, so the `Null` combinator -- and with it `AggregateFunctionCountNotNullUnary` -- is never applied to it, and the native path counted every row. `count` over a single `Variant` argument now resolves to `AggregateFunctionCountNotNullVariant`, which counts the rows whose local discriminator is not `NULL_DISCRIMINATOR`; its state is the same single counter and normalizes to `AggregateFunction(count)`, so the state forms stay byte-compatible with the plain `count` state. **The NULL-skipping contract of the natively-supported functions.** The other functions that accept a `Variant` natively are held to the same contract by `AggregateFunctionVariantNull`, a wrapper mirroring the `Null` combinator: the rows where a `Variant` argument is NULL are skipped, like the `Null` combinator skips the NULL values of `Nullable` arguments (only the row skipping is mirrored; the result-type promotion and the all-NULL-group result of the combinator are deliberately not). Without it, `any` returned NULL from a group that has non-NULL values, `groupArray` stored the NULLs its documentation promises to remove, `groupConcat` concatenated them as data, and the uniq family counted NULL as a distinct value. The wrapper is applied at the factory's leaf resolution point (`getImpl`), which every path funnels through -- the top-level native resolution as well as nested functions reconstructed from declared `AggregateFunction(f, Variant(...))` state types -- so the state layouts always match. Unlike the `Null` combinator, the wrapper never changes the result type, the state layout or the serialization of the nested function: a function that accepts a `Variant` natively already accepted it before this change, with exactly the nested result type and state representation, and both are persisted (in `AggregateFunction(f, Variant(...))` columns and in the schemas of the materialized views reading them), so changing either would break an upgrade. An all-NULL group produces the empty nested state and returns the nested function's empty-set result (`groupConcat` over a `Variant` keeps returning `String`, and an empty string for an all-NULL group); for functions whose result is the `Variant` itself (`any`, `argMin`, ...) an all-NULL group still reports NULL, because a `Variant` represents NULL on its own. Window functions keep handling their argument types themselves, so the `RESPECT NULLS` forms and `estimateCompressionRatio` still see the NULL rows, and `count` keeps its dedicated implementation above and declares `AggregateFunctionProperties::skips_variant_nulls`. **Compatibility of the NULL-skipping contract.** The functions above accepted a `Variant` argument before this change, so skipping its NULL rows is a change of their behavior and is gated by a compatibility setting: `aggregate_functions_skip_variant_nulls = 0` (implied by `compatibility` below `26.8`) restores the previous behavior, where those rows were aggregated as ordinary values. The setting only controls which rows are added to a state -- the result type, the state layout and the serialization are identical in both modes, so an `AggregateFunction(f, Variant(...))` state written by either mode stays readable and mergeable by the other. For the same reason the setting cannot change the meaning of a state that has already been written: a state written by an older version keeps the values that went into it, including the ones that came from the NULL rows, and no state-type version can retroactively tell the two apart because their bytes are identical. The consequence for existing data is spelled out in the changelog entry and in the documentation: states written before and after the upgrade merge freely but follow different accumulation rules, so a deployment that must keep the old results over persisted states has to keep writing with `aggregate_functions_skip_variant_nulls = 0` until the historical data is re-inserted. It is deliberately not consulted when there is no query context (a background operation, or a table loaded at startup). The one state that replays raw values into a nested function -- the distinct-key history of the `-Distinct` combinator -- keeps the rule by construction: the nested function of `-Distinct` over a `Variant` argument is resolved with the NULL skipping disabled, so the replay always aggregates the NULL keys a history contains (exactly like the versions that wrote states before the contract existed, so a stored `countDistinctState` over a `Variant` reads back to the same result under either value of the setting), and the NULL rows of newly aggregated data are skipped in front of the history instead, by an outer `AggregateFunctionVariantNull` over the combined function under the same setting. The aggregation of `Variant` arguments, the NULL-skipping contract and this compatibility note are documented in the `Variant` docs (the embedded documentation block in `DataTypeVariant.cpp` and `docs/reference/data-types/variant.mdx`). Added `tests/queries/0_stateless/04504_variant_aggregate_functions.sql`, `tests/queries/0_stateless/04644_variant_aggregate_state_setting_independence.sql`, `tests/queries/0_stateless/04652_variant_count_skips_nulls.sql` `tests/queries/0_stateless/04657_variant_aggregate_functions_skip_nulls.sql` `tests/queries/0_stateless/04692_variant_aggregate_nulls_compatibility_setting.sql` and `tests/queries/0_stateless/04817_variant_distinct_states_setting_independence.sql`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109369",
          "createdAt": "2026-07-03T21:55:20Z",
          "updatedAt": "2026-08-13T04:04:51Z",
          "timestamp": "2026-08-13T04:04:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 24
          },
          "labels": [
            "pr-backward-incompatible",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7596ff121e0df2ce9ff1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104431",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104431",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallelize reads from a single Parquet file in StorageFile, again",
          "text": "Reverts ClickHouse/ClickHouse#104359",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104431",
          "createdAt": "2026-05-08T22:04:53Z",
          "updatedAt": "2026-08-13T03:55:52Z",
          "timestamp": "2026-08-13T03:55:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 68
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6773563c9a54d205e1a2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114552",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114552",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Raise the log-wait timeout in `test_keeper_session_loss_direct_read` to fix flakiness",
          "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Raise the log-wait timeout in the integration test `test_storage_kafka/test_keeper_session_loss_direct_read.py` to fix its flakiness. The test waits for the `StorageKafka2` activation task to notice the expired Keeper session after a network partition. The task's check period is one minute, and on an overloaded CI host (sanitizer builds, busy schedule pool) the log line has been seen missing the 180-second window, failing the test with `retry_failed`. `wait_for_log_line` returns as soon as the line appears, so a healthy run does not pay for the extra margin. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1305` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114552",
          "createdAt": "2026-08-12T22:48:38Z",
          "updatedAt": "2026-08-13T03:54:19Z",
          "timestamp": "2026-08-13T03:54:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:162bbfa45b2183d873e4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114380",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114380",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Clarify `aiSimilarity` cosine similarity range in its documentation",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110777 ### Changelog category (leave one): - Documentation (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1306` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114380",
          "createdAt": "2026-08-11T20:21:22Z",
          "updatedAt": "2026-08-13T03:54:15Z",
          "timestamp": "2026-08-13T03:54:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "davidmenggx",
          "state": "closed",
          "assignees": [
            "george-larionov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b097d6e1e86e5a44d029",
        "signalId": "github:ClickHouse/ClickHouse:issue:114455",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114455",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Function `resetSerialID` to reset/remove a `generateSerialID` series",
          "text": "### Company or project name _No response_ ### Use case `generateSerialID(series_identifier)` (introduced in 25.x) creates a named auto-increment counter whose state is persisted as a znode in [Zoo]Keeper under series_keeper_path. There is currently no SQL-level way to reset or delete a series once it has been created — the state is effectively permanent for the lifetime of the Keeper data. This matters in particular for: 1. **Multi-tenant applications:** a common pattern is one series per tenant (e.g. generateSerialID('serial_<tenant_uuid>')) to produce per-tenant sequential IDs. When a tenant is deleted, its series znode remains forever. Over time this accumulates orphaned znodes and consumes the max_autoincrement_series budget with dead series. 2. **ClickHouse Cloud / managed environments**: users have no access to Keeper at all, so they cannot clean up series state out-of-band the way a self-hosted user could with zkCli/keeper-client. Whatever series they create is permanent unless support intervenes manually. 3. **Testing / re-ingestion workflows:** dropping and recreating a table does not reset its associated series, which is surprising — reloading data into a fresh table continues numbering from the old counter, with no way to start from 0 again under the same series name. ### Describe the solution you'd like A function or statement to delete a series and its Keeper state, gated by an appropriate grant, e.g.: ``` sql SELECT resetSerialID('my_series'); -- resets counter to 0 (or deletes the znode) ``` or, arguably more idiomatic as a DDL-like/system operation rather than a function with side effects: ``` sql SYSTEM DROP SERIAL ID 'my_series'; -- or SYSTEM RESET SERIAL ID 'my_series' [TO <value>]; ``` ### Describe alternatives you've considered - Manual znode removal via keeper-client — works for self-hosted, impossible for ClickHouse Cloud users. - Using a new series name after tenant deletion / reload — leaks znodes and consumes max_autoincrement_series indefinitely. ### Additional context Docs: https://clickhouse.com/docs/reference/functions/regular-functions/other-functions#generateSerialID",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114455",
          "createdAt": "2026-08-12T09:37:29Z",
          "updatedAt": "2026-08-13T03:50:56Z",
          "timestamp": "2026-08-13T03:50:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "feature",
            "easy task"
          ],
          "author": "Yonatan-Dolan",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:219dac6c7bde41746d5d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96844",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96844",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Columns Cache",
          "text": "Implement the deserialized columns cache that keeps columns in memory. In comparison to the uncompressed cache, this saves time on deserialization and copying. Closes #82335. The cache identifies entries by table UUID, so it is only active for tables in databases that assign UUIDs (`Atomic` and `Replicated`); tables in `Ordinary` databases are silently excluded. The feature is gated behind the `EXPERIMENTAL` settings `use_columns_cache`, `enable_reads_from_columns_cache`, and `enable_writes_to_columns_cache`. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Implement an experimental columns cache that keeps deserialized columns in memory for `MergeTree` tables. In comparison to the uncompressed cache, this saves time on deserialization and copying. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96844",
          "createdAt": "2026-02-13T21:00:45Z",
          "updatedAt": "2026-08-13T03:49:57Z",
          "timestamp": "2026-08-13T03:49:57Z",
          "metrics": {
            "reactions": 5,
            "comments": 77
          },
          "labels": [
            "pr-performance",
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:37007efa858229f2ea01",
        "signalId": "github:ClickHouse/ClickHouse:issue:114588",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114588",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Column named like an array subcolumn (a.size0) added via ALTER: old parts silently return the subcolumn value instead of the DEFAULT, and merges materialize the wrong values",
          "text": "**Describe what's wrong** After `ALTER TABLE ... ADD COLUMN` adds a column whose name collides with a generated array subcolumn (e.g. a column named `a.size0` next to an `Array` column `a`), reads of that column from parts created **before** the ALTER silently return the **subcolumn's computed value** (the array size) instead of the added column's DEFAULT. Parts created after the ALTER return the stored value, so one `SELECT` returns a mix of two different meanings for the same identifier — and a subsequent merge (`OPTIMIZE ... FINAL`) **materializes the wrong values permanently** into the new part. **Does it reproduce on the most recent release?** Reproduces on current master, `26.8.1.1068` and `26.8.1.1194` (official builds). Deterministic: 20/20 runs. **How to reproduce** ```sql CREATE TABLE t_sub (a Array(UInt64)) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t_sub VALUES ([0]), ([1, 2]); ALTER TABLE t_sub ADD COLUMN `a.size0` UInt64 DEFAULT 7; INSERT INTO t_sub (a, `a.size0`) VALUES ([3], 7); SELECT a, `a.size0` FROM t_sub; -- [0] 1 <- wrong: the array-size subcolumn, expected the DEFAULT 7 -- [1,2] 2 <- wrong: expected 7 -- [3] 7 <- correct (part written after the ALTER) ``` The implicit default has the same problem: ```sql CREATE TABLE t_sub3 (a Array(UInt64)) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t_sub3 VALUES ([10, 20, 30]); ALTER TABLE t_sub3 ADD COLUMN `a.size0` UInt64; SELECT `a.size0` FROM t_sub3; -- 3 <- wrong: expected 0 (the type default) ``` And the merge bakes the wrong values in, so the corruption survives the mixed-parts state: ```sql OPTIMIZE TABLE t_sub FINAL; SELECT a, `a.size0` FROM t_sub ORDER BY a; -- [0] 1 <- now stored physically -- [1,2] 2 -- [3] 7 ``` A control table created with the column from the start returns the stored values everywhere. **Expected behavior** A declared physical column must shadow the generated subcolumn consistently: reads from parts that predate the ALTER should evaluate the column's DEFAULT (`7`, or `0` for the implicit default), exactly as they do for any other added column name, and merges must materialize those DEFAULT values. Related: https://github.com/ClickHouse/ClickHouse/issues/113420 (the same name-collision family: `ALTER MODIFY COLUMN` with mixed-type parts makes every SELECT throw `AMBIGUOUS_COLUMN_NAME`; this issue is the silent wrong-result face under `ADD COLUMN`) Related: https://github.com/ClickHouse/ClickHouse/issues/87375 Related: https://github.com/ClickHouse/ClickHouse/issues/90219 Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114588",
          "createdAt": "2026-08-13T03:45:57Z",
          "updatedAt": "2026-08-13T03:45:57Z",
          "timestamp": "2026-08-13T03:45:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "potential bug"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:189264eee95c82886296",
        "signalId": "github:ClickHouse/ClickHouse:issue:114587",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114587",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "arrayIntersect overflow guard tests isInteger on a Nullable type, so it never fires",
          "text": "### Describe what's wrong **`arrayIntersect` on `Nullable` integer arrays of different widths silently matches values that do not survive the narrowing cast to the common element type. `arrayIntersect([toNullable(1)], [toNullable(257)])` returns `[1]`; the non-Nullable form `arrayIntersect([1], [257])` correctly returns `[]`. Pre-existing — not introduced by this PR.** - **Root cause:** arrayIntersect.cpp:398 sets `nested_init_type` to the array's nested type WITHOUT removing `Nullable`, and line 401 then calls `isInteger(nested_init_type)`. `WhichDataType(Nullable(UInt32))` is `TypeIndex::Nullable`, so `isInteger` (and `isDate`/`isDateTime`/`isDateTime64`) return false and the whole overflow-mask block is skipped for every `Nullable` argument. <details> <summary>Analysis details (evidence, affected locations, impact)</summary> **Why we believe this is a bug:** `executeImpl` (arrayIntersect.cpp:463) computes the common subtype with `getMostSubtype`, which NARROWS (`Nullable(UInt16)` + `Nullable(UInt8)` -> `Nullable(UInt8)`), then `castColumns` casts the wider argument down. `prepareArrays` is meant to catch the values that changed under that cast by building `arg.overflow_mask` (arrayIntersect.cpp:395-415), and the element loop skips masked elements (arrayIntersect.cpp:711). For a `Nullable` array the guard at line 401 never lets that happen, so the truncated value is inserted and matched as if it were the original. **Affected locations:** - [`src/Functions/array/arrayIntersect.cpp:401`](https://github.com/ClickHouse/ClickHouse/blob/2485f1496f7dff/src/Functions/array/arrayIntersect.cpp#L401) — overflow-mask guard: isInteger on a still-Nullable nested type - [`src/Functions/array/arrayIntersect.cpp:398`](https://github.com/ClickHouse/ClickHouse/blob/2485f1496f7dff/src/Functions/array/arrayIntersect.cpp#L398) — nested_init_type keeps its Nullable wrapper **Impact:** Wrong results from `arrayIntersect` whenever the arguments are `Nullable` arrays of different integer widths (or Date/DateTime widths) and a value in the wider argument aliases a value in the narrower one modulo the narrower type. Reachable from ordinary SQL over two `Array(Nullable(...))` table columns; `Nullable` array columns are common, and the non-Nullable behaviour that users would infer from `00930_arrayIntersect.sql` is the opposite. </details> ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `2485f1496f7dff`). ### How to reproduce [▶ Run on ClickHouse Fiddle](https://fiddle.clickhouse.com/6be77c3a-66e5-40e6-9596-b3c75b04e31c) <details> <summary>Reproducer</summary> ```sql -- Test: a value that does not survive the cast to the common element type is not in the intersection, -- also when the arrays are Nullable. DROP TABLE IF EXISTS t_04869; SELECT arrayIntersect([1], [257]); SELECT arrayIntersect([toNullable(1)], [toNullable(257)]); SELECT arrayIntersect([-100], [156]); SELECT arrayIntersect([toNullable(-100)], [toNullable(156)]); DROP TABLE IF EXISTS t_04869; CREATE TABLE t_04869 (a Array(Nullable(UInt32)), b Array(Nullable(UInt8))) ENGINE = Memory; INSERT INTO t_04869 VALUES ([1024, 1031], [0, 7]), ([256, 300], [0, 44]); SELECT arraySort(arrayIntersect(a, b)) FROM t_04869 ORDER BY a; DROP TABLE IF EXISTS t_04869; ``` </details> ### Expected behavior ``` [] [] [] [] [] [] ``` ### Error message and/or stacktrace ``` [] [1] [] [156] [0,44] [0,7] ``` <details> <summary>Suggested fix</summary> Strip `Nullable` before the type test, e.g. compute `auto init_nested = removeNullable(nested_init_type);` and gate on that. `callFunctionNotEquals` is already handed the null-stripped nested columns (arrayIntersect.cpp:389-392), so the types passed with them at lines 408-409 must be null-stripped too, and the existing `removeNullable(overflow_mask)` at line 412 then becomes redundant rather than load-bearing. </details> <details> <summary>Additional context</summary> **Open risks:** - `arrayUnion` and `arraySymmetricDifference` use `getLeastSupertype` and therefore never narrow — verified `arrayUnion([toNullable(1)], [toNullable(257)])` = `[1,257]`. No sibling call site to fix. Found during automated review of [PR #113021](https://github.com/ClickHouse/ClickHouse/pull/113021). Severity P1 · Finding `h_pr113021_001` </details> cc @alexey-milovidov (author of #113021) ### Additional context _Generated by ClickGap / Claude._",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114587",
          "createdAt": "2026-08-13T03:44:35Z",
          "updatedAt": "2026-08-13T03:44:35Z",
          "timestamp": "2026-08-13T03:44:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "clickgapai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d08b4f1e27240614c23f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105823",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105823",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Iceberg: skip eager getContent() JSON serialization when metadata log is disabled",
          "text": "`AvroForIcebergDeserializer::getContent(row_index)` was passed as an argument to `insertRowToLogTable` and serialized a manifest entry to JSON on every call, even when `iceberg_metadata_log_level` (default `none`) discarded the result. Hoist the level check above the five call sites so the payload is only built when the log will consume it. CPU profile of a query over ~300k pruned manifest entries (`max_threads=1`): ~59% of CPU was JSON serialization (`writeJSONString`, `SerializationTuple::serializeTextJSON`, `SerializationObjectPool::getOrCreate` and the resulting allocator churn). ``` ┌─samples─┬───pct─┬─function─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ 3278 │ 13.15 │ DB::SerializationObjectPool::getOrCreate(wide::integer<128ul, unsigned int>, std::__1::function<DB::ISerialization* ()>) │ │ 2797 │ 11.22 │ DB::WriteBuffer::write(char) │ │ 2679 │ 10.75 │ syscall │ │ 1283 │ 5.15 │ DB::writeJSONString(char const*, char const*, DB::WriteBuffer&, DB::FormatSettings const&) │ │ 1009 │ 4.05 │ DB::SerializationTuple::serializeTextJSON(DB::IColumn const&, unsigned long, DB::WriteBuffer&, DB::FormatSettings const&) const │ │ 737 │ 2.96 │ je_sallocx │ │ 715 │ 2.87 │ DB::Iceberg::ManifestFileIterator::processRow(unsigned long) │ │ 635 │ 2.55 │ DB::SerializationNamed::getHash(std::__1::shared_ptr<DB::ISerialization const> const&, std::__1::basic_string<char, std::__1::char_traits<char>, std::__1::allocator<char>> const&, DB::ISerialization::Substream::Type) │ │ 632 │ 2.53 │ DB::SerializationObjectPool::getOrCreate(wide::integer<128ul, unsigned int>, std::__1::function<DB::ISerialization* ()>)::$_0::operator()(DB::ISerialization const*) const │ │ 558 │ 2.24 │ CurrentMemoryTracker::allocImpl(long, bool) │ │ 536 │ 2.15 │ DB::Iceberg::(anonymous namespace)::deserializeFieldFromBinaryRepr(std::__1::basic_string<char, std::__1::char_traits<char>, std::__1::allocator<char>>, std::__1::shared_ptr<DB::IDataType const>, bool) │ │ 499 │ 2 │ DB::SerializationTuple::create(std::__1::vector<std::__1::shared_ptr<DB::SerializationNamed const>, std::__1::allocator<std::__1::shared_ptr<DB::SerializationNamed const>>>, bool) │ │ 463 │ 1.86 │ │ │ 444 │ 1.78 │ DB::(anonymous namespace)::writeTraceInfo(DB::TraceType, int, siginfo_t*, void*) │ │ 444 │ 1.78 │ je_malloc │ │ 420 │ 1.68 │ je_nallocx │ │ 383 │ 1.54 │ DB::DataTypeTuple::doGetSerialization(DB::SerializationInfoSettings const&) const │ │ 381 │ 1.53 │ auto DB::Field::dispatch<DB::Field::create(DB::Field const&)::'lambda'(auto&), DB::Field const&>(auto&&, DB::Field const&) │ │ 310 │ 1.24 │ DB::Iceberg::IcebergSchemaProcessor::tryGetFieldCharacteristics(int, int) const │ │ 304 │ 1.22 │ rtree_read │ │ 250 │ 1 │ CurrentMemoryTracker::free(long) │ │ 226 │ 0.91 │ DB::Field::~Field() │ │ 220 │ 0.88 │ auto DB::Field::dispatch<DB::Field::create(DB::Field&&)::'lambda'(auto&), DB::Field&>(auto&&, DB::Field&) │ │ 195 │ 0.78 │ itoa(long, char*) │ │ 188 │ 0.75 │ DB::SerializationTuple::supportsPooling() const │ │ 177 │ 0.71 │ je_sdallocx │ │ 172 │ 0.69 │ DB::ISerialization::pooled(wide::integer<128ul, unsigned int>, std::__1::function<DB::ISerialization* ()>) │ │ 166 │ 0.67 │ DB::SerializationArray::serializeTextJSON(DB::IColumn const&, unsigned long, DB::WriteBuffer&, DB::FormatSettings const&) const │ │ 165 │ 0.66 │ DB::SerializationNamed::~SerializationNamed() │ │ 155 │ 0.62 │ operator new[](unsigned long) │ └─────────┴───────┴──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` related to #104169 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Avoid eager JSON serialization of Iceberg manifest entries when `iceberg_metadata_log_level` is below the call site's threshold.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105823",
          "createdAt": "2026-05-26T06:14:12Z",
          "updatedAt": "2026-08-13T03:41:06Z",
          "timestamp": "2026-08-13T03:41:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-performance",
            "manual approve",
            "can be tested"
          ],
          "author": "starpact",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c98c66e2583c8de67887",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114584",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114584",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #107028 to 26.3: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114584",
          "createdAt": "2026-08-13T03:40:48Z",
          "updatedAt": "2026-08-13T03:40:57Z",
          "timestamp": "2026-08-13T03:40:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "kssenii",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:01e168a8d612ef62e99b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114583",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114583",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #107028 to 25.8: Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/107028 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31663540960/job/94333230705)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114583",
          "createdAt": "2026-08-13T03:40:10Z",
          "updatedAt": "2026-08-13T03:40:19Z",
          "timestamp": "2026-08-13T03:40:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "kssenii",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:07d22071f7067c8e6501",
        "signalId": "github:ClickHouse/ClickHouse:issue:114582",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114582",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Failed INSERT INTO s3(...) PARTITION BY leaves a durable read-visible prefix; default hive strategy silently duplicates it on retry",
          "text": "### Company or project name ClickHouse QA (durability testing) ### Describe what's wrong A failed `INSERT INTO FUNCTION s3(...) PARTITION BY <key>` (and the equivalent object-storage table engines) is **not atomic and leaves a durable, read-visible prefix of the partitions it had already written**, while the statement reports failure. `PartitionedSink` finalizes one object per partition value **sequentially**, with no rollback of the partitions committed before the one that fails: ```cpp // src/Storages/PartitionedSink.cpp:120-125 void PartitionedSink::onFinish() { for (auto & [_, sink] : partition_id_to_sink) { sink->onFinish(); // each StorageObjectStorageSink::onFinish -> finalizeBuffers -> write_buf->finalize() (the durable PUT / CompleteMultipartUpload) } } ``` Each child `StorageObjectStorageSink::onFinish` (`src/Storages/ObjectStorage/StorageObjectStorageSink.cpp:89-119`) finalizes (PUTs) its own object. If the PUT of partition *k* throws, the exception propagates out of the loop; the partitions finalized **before** *k* are already durable and immediately listable/readable, and nothing deletes them or records that only a prefix landed (`cancelBuffers` only aborts the not-yet-finalized buffers). So the error's implied \"nothing happened\" contract is a lie. The consequence on the **default** partition strategy is silent duplication on a natural retry. `file_like_engine_default_partition_strategy` defaults to `HIVE` (`src/Core/Settings.cpp:7754`), and `HiveStylePartitionStrategy::getPathForWrite` appends a fresh `generateSnowflakeID()` to every filename: ```cpp // src/Storages/IPartitionStrategy.cpp:383 path += std::to_string(generateSnowflakeID()) + \".\" + Poco::toLower(file_format); ``` so every attempt writes new object keys. A client that treats the failed INSERT as \"nothing happened\" and re-submits it therefore **duplicates every partition that survived the first attempt** — there is no dedup for object-storage writes and no exists-check that can catch it (the keys differ each time). A client that does *not* re-submit is left with silently half-applied state. The `WILDCARD` strategy (an explicit `{_partition_id}` placeholder in the path) fails safe instead of duplicating: the retry recomputes the same keys and the exists-check rejects it (`Code: 36 ... already exists, enable s3_truncate_on_insert`), which is fail-closed but still leaves the client permanently half-applied and unable to tell which partitions survived without listing the bucket. ### Does it reproduce on the most recent release? Yes — reproduced deterministically on 26.8.1.1. The code paths (`PartitionedSink::onFinish`, `StorageObjectStorageSink::finalizeBuffers`, `HiveStylePartitionStrategy::getPathForWrite`, and the `HIVE` default) are all present on current `master`. ### How to reproduce A durability rig reproduces both facets 3/3 with controls (MinIO behind a fault proxy that returns a non-retryable `403 AccessDenied` on **one** partition's PUT — the classic least-privilege / transient-denial to a single prefix — and a single clean server). `N = 300` rows over 3 partition values `p = number % 3`. **Facet 1 — a failed INSERT leaves a read-visible prefix (WILDCARD path, deterministic):** 1. `INSERT INTO FUNCTION s3('http://.../base/{_partition_id}/data.parquet', ..., 'Parquet', 'id UInt64, p UInt64') PARTITION BY p SELECT number AS id, number % 3 AS p FROM numbers(300)`, with the proxy denying PUTs to the key of the partition that finalizes **last**. 2. The INSERT fails (`Code: 499 ... S3_ERROR`), yet the two partitions finalized before it are durable and queryable through the object store, while the faulted partition is absent. - **Clean control** (proxy disarmed): the 3-partition INSERT round-trips all 300 rows. - **Isolation control** (same fault, but `INSERT ... PARTITION BY p ... WHERE p = <faulted>` — a single partition, no fan-out): the INSERT fails and **nothing** is durable, proving the surviving prefix in step 2 is the fan-out's doing. **Facet 2 — the default HIVE strategy silently duplicates on retry:** 1. Same INSERT into a non-wildcard path (default `partition_strategy = hive`), denying PUTs to one partition's `p=<v>/` prefix. 2. The INSERT fails; a `SELECT count()` over the whole prefix returns **200** (the two surviving partitions). 3. The client re-submits the identical INSERT (its natural response to a failed op). It succeeds, and `SELECT count()` over the prefix now returns **500** — 300 unique rows plus the **200-row surviving prefix duplicated** (fresh snowflake filenames, no dedup, no exists-check). ### Expected behavior A failed `INSERT ... PARTITION BY` should not silently leave a durable, read-visible subset of its partitions with the statement reporting failure. Either the partitioned write should roll back the partitions it already committed on failure (so the error's \"nothing happened\" contract holds), or the surviving prefix should be recorded/surfaced so the outcome is not silently half-applied — and in particular, on the **default** hive strategy a natural client retry of a failed partitioned INSERT should not silently duplicate the rows of the partitions that survived the first attempt. Honest caveat: object-storage writes through the `s3`/`azureBlobStorage` table functions and engines are not transactional, and at-least-once export is a known property. What this report isolates is the specific, undocumented combination — a failed multi-partition INSERT leaves a **query-visible** durable prefix, and the **default** hive strategy makes an ordinary retry duplicate that prefix with no dedup and no way to be idempotent — which is a data-integrity footgun distinct from the concurrent-writer clobber already filed as #112419. ### Error message and/or stacktrace `Code: 499. DB::Exception: ... (S3_ERROR)` on the faulted INSERT; `Code: 36 ... Object in bucket ... already exists (...enable s3_truncate_on_insert)` on the wildcard retry. ### Additional context Found by the ClickFawkes durability framework under its `failed_op_durable_residue` lens (a client op fails but a durable read-visible prefix survives, un-rolled-back). Reproduced via `--mode objectstorage-partition-prefix-residue`.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114582",
          "createdAt": "2026-08-13T03:38:15Z",
          "updatedAt": "2026-08-13T03:38:15Z",
          "timestamp": "2026-08-13T03:38:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d723a8bb594e58f071cc",
        "signalId": "github:ClickHouse/ClickHouse:issue:114581",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114581",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "IN (SELECT ...) inside a higher-order-function lambda always evaluates to 0 when the query is a derived table or the lambda is in WHERE",
          "text": "**Describe what's wrong** An `IN (SELECT ...)` predicate inside a higher-order-function lambda (`arrayExists`, `arrayFilter`, `arrayMap`, ...) always evaluates to `0` when the enclosing SELECT is used as a derived table (or when the lambda sits in an outer `WHERE`). The identical expression at the top level returns the correct result. **Does it reproduce on the most recent release?** Reproduces on current master, `26.8.1.1068` and `26.8.1.1194` (official builds). Deterministic: 20/20 runs. **How to reproduce** No tables needed: ```sql SELECT arrayExists(x -> x IN (SELECT 2), [2]); -- 1 (correct) SELECT * FROM (SELECT arrayExists(x -> x IN (SELECT 2), [2])); -- 0 (wrong: the same expression, only wrapped in a derived table) ``` More shapes of the same mechanism: ```sql SELECT arrayFilter(x -> x IN (SELECT 2), [1, 2, 3]); -- [2] (correct) SELECT * FROM (SELECT arrayFilter(x -> x IN (SELECT 2), [1, 2, 3])); -- [] (wrong) SELECT arrayMap(x -> x IN (SELECT '2'), [2, 3]); -- [1,0] (correct) SELECT * FROM (SELECT arrayMap(x -> x IN (SELECT '2'), [2, 3])); -- [0,0] (wrong) -- WHERE context loses rows: SELECT count() FROM (SELECT 1 AS k) WHERE arrayExists(x -> x IN (SELECT 1), [k]); -- 0 (wrong: expected 1) WITH tm1 AS (SELECT arrayExists(x -> x IN (SELECT 2), [2])) SELECT * FROM tm1; -- 0 (wrong) ``` All at default settings; `query_plan_enable_optimizations = 0` does not cure it, so it looks like the set for the lambda-captured `IN` is not built/bound when the expression is resolved inside a subquery scope, rather than a plan-optimization issue. Possibly related observation: with `enable_analyzer = 0` the derived-table form fails outright with an exception `Code: 47` `UNKNOWN_IDENTIFIER`, where the required column is spelled `... in(x, _subquery1) ...` but the available column is `... in(x, _subquery2) ...` — the same set-identity confusion visible in the old analyzer. **Expected behavior** Wrapping a SELECT in a derived table (or moving the expression into `WHERE`) must not change the value of `IN (SELECT ...)` inside a lambda: all the wrapped forms above should return the top-level results (`1`, `[2]`, `[1,0]`, `1`). Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114581",
          "createdAt": "2026-08-13T03:36:46Z",
          "updatedAt": "2026-08-13T03:36:46Z",
          "timestamp": "2026-08-13T03:36:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "potential bug"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ab74897db717bf9bd2f3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107028",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107028",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix data race on FileCacheQueryLimit::query_map causing LOGICAL_ERROR",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/106364 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a rare `LOGICAL_ERROR` \"Attempt to release query context that does not exist\" and the accompanying server crash when reading `MergeTree` tables through a filesystem cache disk created with `enable_filesystem_query_cache_limit = 1`. ### Description `FileCacheQueryLimit::query_map` (a `std::unordered_map`) was reachable from two different cache locks with no single lock serializing it: - the read `tryGetQueryContext` (`query_map.find`) runs under `CacheStateGuard::Lock` (called from `FileCache::doTryReserve`, once per reservation); - the writes `getOrSetQueryContext` (`query_map.emplace`) and `removeQueryContext` (`query_map.erase`) run under `CachePriorityGuard::WriteLock` (called from `FileCache::getQueryContextHolder` and `~QueryContextHolder`). Since the read takes a different mutex than the writes, a `find()` can run concurrently with an `emplace()`/`erase()` that rehashes the table and invalidates buckets and iterators. That is undefined behaviour on `std::unordered_map` and surfaces as either `Attempt to release query context that does not exist` (the map is corrupted so a present key looks absent in `removeQueryContext`) or a server crash while reading a half-rehashed bucket. Both were seen together in the same CI run on `Stateless tests (amd_debug, sequential)`. The split appeared when `tryGetQueryContext` was moved from the priority write lock to the cheaper cache state lock, leaving the map reachable from two locks at once. Fix: guard all three `query_map` accessors with a dedicated leaf mutex, so the map has a single owning lock regardless of which cache lock the caller holds. The mutex is taken last and never nests another cache lock under it, so it does not change the existing lock order, and the optimization of keeping the read off the priority write lock is preserved. Reproduced with a concurrent reader/writer witness over an S3-backed cache disk with `enable_filesystem_query_cache_limit = 1`: 2073 overlapping accesses to `query_map` without the fix, 0 with it. ### Follow-up: last-holder release leak (#109508) @ Algunenano found a related issue by code analysis (#109508): the last-holder decision in `~QueryContextHolder` still dropped this holder's own reference to the context outside the cache lock, and `removeQueryContext` gated the erase on `use_count() > 2` before that drop. `use_count()` is not a synchronization primitive, so this is a TOCTOU. A query with parallel read streams has several holders for the same `query_id` (each `CachedOnDiskReadBufferFromFile` creates its own), so `use_count > 2`. When two holders release at the same time both observe the shared count, both skip the erase, and after both drop their reference only the map entry remains and is never removed. That orphans `query_map[query_id]` for the lifetime of the cache, so a later query reusing the same `query_id` picks up stale per-query limit state. Fix: drop this holder's reference inside `removeQueryContext` under `query_map_mutex` and erase only once the map entry is the sole owner (`use_count() == 1`). Every reference change to the context is then serialized by the lock that also guards `getOrSetQueryContext`, so the check-and-erase is atomic with respect to concurrent holders. Added `FileCacheTest.QueryLimitConcurrentReleaseNoLeak`, which drives the concurrent-release interleaving deterministically (fails on the previous `use_count() > 2` logic, passes now). Related: https://github.com/ClickHouse/ClickHouse/issues/109508 <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.1.840` (included in `26.7` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107028",
          "createdAt": "2026-06-10T16:02:03Z",
          "updatedAt": "2026-08-13T03:34:12Z",
          "timestamp": "2026-08-13T03:34:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a0ef91fbb1f46fe58b81",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110210",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110210",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Web UI prototype for framing formats",
          "text": "Prototype changes to the Web UI (`programs/server/play.html`) to test and showcase the framing formats feature. This is a draft for experimentation and review of the client-side experience, not intended to merge as-is. It builds on the framing-formats server work in #110127 (this branch is based on that PR, so the diff includes those commits until it merges into `master`). What the prototype does: - Streams every query with the `EventStream` framing format over HTTP (`framing_output_format=EventStream`, `send_logs_level=trace`), decoding data / progress / log / profile_events / exception packets as they arrive - so progress and logs work regardless of the output format (including `FORMAT Pretty`, binary formats via base64, and images). - Shows realtime CPU, memory, and disk usage like `clickhouse-client`, aggregated per host (total and max/host), switching to peak RAM when the query finishes. The meter state is owned by each tab: the CPU counters in `profile_events` packets are per-packet increments, so a query running in a background tab keeps accumulating them, and reopening the tab continues the meter from its live values (rather than restarting near zero). - Merges elapsed time, realtime metrics, and rows/bytes read stats into a single progress area with the progress-bar gradient rendered behind all of them; the text is tinted by a mask clipped from the same gradient so it stays readable over the fill. - Displays server logs in real time, colored like the client, with a Logs button available even on exception; the log view stays fast for 100k+ lines while keeping native browser search, bounding retention to the most recent 100k lines (older lines are dropped and counted in a marker line). The full response text kept for the tab/history snapshot is also capped as it is collected (at ~100 KB, the size above which the snapshot is dropped), so a single large framed data/log burst is not retained in full only to be discarded later. A failed run whose reply outgrew that cap is persisted as a compact snapshot that keeps the exception carrier in the form its framing kind replays (the terminal `event: exception` block, the `{\"packet\":\"exception\",...}` line, or the plain `{\"exception\":...}` line, captured at stream-read time since the capped text stops growing before the terminal exception arrives), so reopening or reloading a failed tab still shows the error reason. - Base64-decodes every `data`/`totals`/`extremes` payload (the `EventStream` wire encodes each block as a single base64 `data:` field and the `Content-Type` carries `payload=base64`), deciding the rendering by the output format: the default format is reassembled into the table, an image format (or bytes carrying an image signature) is rendered as an inline image once the stream completes, and any other format is decoded and shown incrementally as raw text. - Uses a framing-compatible default format for queries without an explicit `FORMAT` clause: the framed request asks for `JSONCompactStringsEachRowWithNamesAndTypes` (the framing rejects the in-band-progress `JSONStringsEachRowWithProgress`) and the client reassembles the compact rows into the table renderer's shapes; on the server side (#110127) the compact family emits totals and extremes under framing, so `WITH TOTALS` and extremes-based column coloring keep working on the default path. When such a format is shown as raw text instead (an explicit `FORMAT JSONCompactEachRow`), only the `data` packets are concatenated, so the rendered text is exactly the plain output of that format rather than one carrying the totals and extremes rows the format itself drops. - Propagates a framed `exception` packet as a query failure even when the HTTP status was already 200 (so `Run all` stops at the failed statement) - and, when a query chooses its own `JSONEachPacket*` framing that fails before the 200 OK header (coming back as a non-200 `application/x-ndjson` packet stream ending with a `{\"packet\":\"exception\",...}` line), shows those packets verbatim and records `framing_kind = 'ndjson_packets'` for replay rather than rendering the whole stream as one opaque error string; retries framing-incompatible explicit formats (e.g. `FORMAT JSONEachRowWithProgress`, `FORMAT Template`) once without framing for read-only queries (read-only-ness is resolved with a CTE-aware lexer walk - `WITH y AS (SELECT 1) INSERT INTO t SELECT * FROM y` is a write - shared with `Run all`'s grouping, so such a statement is also a barrier there and never runs in parallel with the reads that follow it; the port with regression coverage lives in `src/Parsers/tests/gtest_play_query_is_read_only.cpp`); a retried `JSON*EachRowWithProgress` format that itself reports a failure in-band - a trailing `{\"exception\":...}` object while the HTTP status stays 200 (`http_write_exception_in_output_format`) - is detected as a failure too, keyed off the output format rather than only a user-chosen `JSONEachPacket*` framing, so `Run all` stops after such a retried query fails; and does not add its own framing to a query that sets its own `framing_output_format` to a real framing choice (the response is then dispatched by content type, so the requested packets are shown verbatim); a query that sets `framing_output_format = 'None'` is refused client-side instead, since this page's rendering depends on framing (values the page cannot know upfront - `= DEFAULT`, a reset to the session/server default, and query-parameter placeholders like `= {fmt:String}` - are classified conservatively the same way and refused, rather than sent with a request shape the response might not match); likewise a standalone `SET framing_output_format = ...` is refused, because it would change the setting for the whole session (with a `session_id`) while the page keeps adding its own framing per request - a query-level `SETTINGS framing_output_format = ...` clause is the supported way to choose framing for one query. A query that carries its own `framing_output_format` is also refused for download (the setting would override the download's chosen `default_format`, so the file would be the framing packet stream), and if such a query fails, its history snapshot - the raw `JSONEachPacket*` packet stream - is replayed as raw text on tab-switch/reload rather than as one opaque error string. - Pins `framing_output_format=None` on every request that expects an unframed response (the plain/chart request, the compatibility retry, the download, and the panel/server-status/completion queries), rather than only omitting the setting - otherwise a framing carried by the connection URL or by the HTTP session behind it would frame those responses too. A query-level `SETTINGS framing_output_format = ...` clause is applied after the URL parameters, so a query that intentionally chooses a framing still overrides the pin. - Records the framing kind (`event_stream` / `ndjson_packets` / none) with each result snapshot and keys the history/tab replay off it, instead of guessing from the payload's first bytes - so a raw result whose text happens to start with `event:` (e.g. `SELECT 'event: data' FORMAT RawBLOB`) or `{\"packet\":` is not reparsed as a framing stream after a tab switch or reload. A snapshot recorded as `ndjson_packets` is replayed as raw text regardless of whether the run succeeded and of its underlying output format, matching the live path - so a successful user-framed `JSONEachPacket*` result whose format has its own restore path (a table or a `JSONCompactColumns` chart) is not reparsed as that format's JSON. The snapshot also records whether the framed stream was truncated, so replay keeps the live fail-closed behavior for images: a cut-off framed `FORMAT PNG` response that showed only its error live reopens from history / Back / Forward as that error too, never as a partially decoded picture (a failed but complete stream - a terminal `exception` packet after the payload - still renders its collected image, as live). - Detects both the query's `FORMAT` clause and its `framing_output_format` setting with the WASM lexer rather than a raw text match, so a mention inside a string literal or a comment - e.g. `SELECT 'FORMAT JSONCompactColumns'` - does not make the page silently opt out of its own framing. The detection is positional, not keyword-adjacent: a settings context is recognized by its `name = value` list grammar (a column merely named `settings` does not open one), and a `FORMAT` clause candidate must follow a token that ends an expression (so in `WITH 1 AS format SELECT format JSONCompactColumns SETTINGS max_threads = 1` both `format` words are identifiers, not a clause). The download reuses the same `FORMAT`-clause detection (now returning the clause span) to strip only a real trailing `FORMAT` clause from the download query, leaving text or ordinary SQL like `SELECT 'FORMAT TSV' AS s` untouched. Both detectors also accept a quoted spelling of the name - the server parses setting names and `FORMAT` names with identifier parsers, so a backquoted `framing_output_format` or format name is real - comparing by the unquoted name. Regression coverage for both lives in `src/Parsers/tests/gtest_play_detect_explicit_format.cpp` and `gtest_play_detect_framing_setting.cpp` (ports of the token walking onto the real `DB::Lexer`). - Dispatches on the response output format case-insensitively. Format names are case-insensitive in ClickHouse (`FormatFactory` looks them up by their lowercased name), while `X-ClickHouse-Format` echoes the identifier exactly as the `FORMAT` clause spelled it, so `FORMAT jsoncompactcolumns` used to lose the chart renderer, a lowercased default format lost the table renderer, and the late in-band exception probes did not recognize `FORMAT xml` / `FORMAT json` / `FORMAT jsoneachrowwithprogress` - a query failing after its `200 OK` header was then reported as a success and `Run all` continued past it. Every dispatch now compares a lowercased copy of the format name. - On the server side (#110127), framed pulling `SELECT` queries now also end with the documented final `progress` packet carrying `result_rows` / `result_bytes` / `memory_usage`, matching the native protocol and the no-result path. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Prototype Web UI changes to test framing formats (draft). Related: https://github.com/ClickHouse/ClickHouse/pull/110127",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110210",
          "createdAt": "2026-07-13T04:17:42Z",
          "updatedAt": "2026-08-13T03:34:09Z",
          "timestamp": "2026-08-13T03:34:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 23
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f5ea670e4ced93e2d30c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114312",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114312",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix nested disk S3 I/O ignoring inner RESOURCE throttler at CachedObjectStorage delegation points",
          "text": "Fix nested disk S3 I/O ignoring inner RESOURCE throttler at DiskObjectStorage entry points ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fix nested disk RESOURCE bandwidth limiting: when a cache disk wraps an S3 disk, S3 I/O is now throttled by the inner disk's RESOURCE instead of being ignored. ### Description When a cache disk (e.g. `cached_s3`) wraps an S3 disk (e.g. `s3_inner`), creating a RESOURCE on the inner S3 disk had no effect — S3 I/O was never throttled. This happened because `DiskObjectStorage` always used the outer disk's resource name (`getReadResourceName()`), but the outer cache disk's local I/O never enters `IOSchedulingScope` (which only exists in `ReadBufferFromS3` and `WriteBufferFromS3`). For example, with this disk configuration: ```xml <s3_inner> <type>s3</type> <endpoint>http://minio:9001/root/data/</endpoint> </s3_inner> <cached_s3> <type>cache</type> <disk>s3_inner</disk> </cached_s3> ``` The user creates a RESOURCE on the inner S3 disk and a table on the cache policy: ```sql CREATE RESOURCE io_s3 (WRITE DISK s3_inner, READ DISK s3_inner); CREATE WORKLOAD all SETTINGS max_bytes_per_second = 10000000 FOR io_s3; CREATE TABLE t ... SETTINGS storage_policy = 'cached_s3'; ``` Before the fix, `INSERT INTO t ... SETTINGS workload = 'all'` would bypass `io_s3` entirely — `system.scheduler` showed zero `dequeued_cost` for the `io_s3` resource. After the fix, S3 I/O is correctly metered through the inner disk's RESOURCE. The fix pushes the inner (wrapped) disk's resource name at three `DiskObjectStorage` entry points: - `createObjectStorageTransaction` - `createObjectStorageTransactionToAnotherDisk` - `prepareRead` Each now uses `wrapped_disk->getReadResourceName()` when a wrapped disk is present, falling back to `getReadResourceName()` for non-nested disks. Only one source file is modified. Non-nested disks are unaffected — the ternary falls through to the same `getReadResourceName()` call as before.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114312",
          "createdAt": "2026-08-11T11:53:05Z",
          "updatedAt": "2026-08-13T03:33:33Z",
          "timestamp": "2026-08-13T03:33:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "Binnn-MX",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d44e2dc0aa0fca8b6742",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:83505",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:83505",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add early short-circuit evaluation for OR/AND in the analyzer to prevent unnecessary scalar subquery execution",
          "text": "<!--- Disable AI PR formatting assistant: true --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add early short-circuit evaluation for logical OR/AND expressions in the new analyzer. When resolving OR/AND functions, if any argument is a decisive constant (truthy for OR, falsy for AND), the entire expression is replaced with the constant value before argument resolution, preventing scalar subqueries in non-decisive branches from being executed. Resolves #83017 ### Details This PR adds early short-circuit evaluation for logical OR/AND expressions in `QueryAnalyzer::resolveFunction()`. When resolving OR/AND functions, the analyzer now checks if any argument is a decisive constant: - For `OR`: if any argument is a truthy constant (1, true, \"true\"), replace the entire expression with 1 - For `AND`: if any argument is a falsy constant (0, false, \"false\"), replace the entire expression with 0 This optimization happens **before** argument resolution, preventing scalar subqueries in non-decisive branches from being executed at all. Resolves #83017",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/83505",
          "createdAt": "2025-07-09T04:57:56Z",
          "updatedAt": "2026-08-13T03:31:23Z",
          "timestamp": "2026-08-13T03:31:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "fhw12345",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:16dbd8eefa1efc6054fa",
        "signalId": "github:ClickHouse/ClickHouse:issue:114579",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114579",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Lazy FINAL with `optimize_aggregation_in_order` merges one group at a time (25000 `mergeBlocks` calls and 50000 log lines for a 35000-row table)",
          "text": "### Describe the situation When `query_plan_optimize_lazy_final` and `optimize_aggregation_in_order` are both on, the aggregation that the lazy `FINAL` replacement builds is merged **one group at a time**: `Aggregator::mergeBlocks` is called once per distinct key, and each call writes two log lines (a `Trace` \"Merging partially aggregated blocks\" and a `Debug` \"Merged partially aggregated blocks for bucket #-1\"). On a 35000-row table with 25000 distinct keys that is 25000 merges and 50000 log lines for a single query. A plain in-order `GROUP BY` of the same size does 7 merges, so this is specific to the lazy `FINAL` shape rather than to aggregation in order in general. The cost is dominated by the logging, so it scales with how verbose the server is configured to be. On a debug/sanitizer build with the log level CI uses, the query below took **525 s**, of which `LoggerElapsedNanoseconds` attributes **285 s** to the logger; the same query with `optimize_aggregation_in_order = 0` took ~3 s. This is not a correctness problem - the results match - and it is not new; it reproduces on builds well before the report below. ### How to reproduce Any recent `master`. With `clickhouse-local`: ```sql CREATE TABLE lf (k UInt64, version UInt64, is_deleted UInt8, v UInt64) ENGINE = ReplacingMergeTree(version, is_deleted) ORDER BY k; INSERT INTO lf SELECT number, 1, 0, number FROM numbers(20000); INSERT INTO lf SELECT number, 2, if(number % 10 = 0, 1, 0), number * 2 FROM numbers(10000, 15000); SELECT count(), sum(v) FROM lf FINAL WHERE k % 7 != 6 SETTINGS max_threads = 4, max_block_size = 8192, query_plan_optimize_lazy_final = 1, max_rows_for_lazy_final = 10000000, min_filtered_ratio_for_lazy_final = 0, optimize_aggregation_in_order = 1; ``` Run it with `--send_logs_level=trace` and count the merges: ``` optimize_aggregation_in_order = 1 -> 25000 \"Merging partially aggregated blocks\" lines optimize_aggregation_in_order = 0 -> 0 ``` The plan shows the replacement's own aggregation (`GROUP BY k` with `argMax` states) under `LazyReadReplacingFinal`; with aggregation in order it emits one chunk per group into the merge stage. ### Expected performance The merge stage should batch groups the way it does for an ordinary in-order `GROUP BY` (7 merges for 50000 groups), instead of one merge per group. ### Additional context Found while triaging a test timeout in https://github.com/ClickHouse/ClickHouse/pull/111459, where the randomized `optimize_aggregation_in_order = 1` setting made a small lazy `FINAL` test query take ~525 s in every flaky check. It is unrelated to that pull request - it reproduces with the feature under test switched off and on builds that predate it - and the test there now pins the setting, but the underlying pathology is worth fixing. Related: https://github.com/ClickHouse/ClickHouse/issues/113704",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114579",
          "createdAt": "2026-08-13T03:29:56Z",
          "updatedAt": "2026-08-13T03:29:56Z",
          "timestamp": "2026-08-13T03:29:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d72a3f02d6813ca8bb08",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112788",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112788",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow COMMENT after all the other column modifiers",
          "text": "The parser accepted the modifiers of a column declaration in one fixed order only, so `COMMENT` had to be written before `CODEC`, `STATISTICS`, `TTL` and per-column `SETTINGS`: ```sql CREATE TABLE t (x UInt64 CODEC(ZSTD) COMMENT 'text') ENGINE = Memory; -- Code: 62. DB::Exception: Syntax error: failed at position 38 (COMMENT). -- Expected one of: STATISTICS, TTL, PRIMARY KEY, SETTINGS, ... ``` while the same declaration with the two modifiers swapped was accepted. There is no ambiguity here, only an accident of how the parser was written, and it is annoying: nothing in the syntax hints that the comment has to go first, and the same is true for `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN`. Now the modifiers that follow the type - `COMMENT`, `CODEC`, `STATISTICS`, `TTL`, `COLLATE`, `PRIMARY KEY` and `SETTINGS` - are parsed in a loop, so they can be written in any order, and each of them at most once. In `CREATE TABLE`, everything that was accepted before is still accepted, and `ASTColumnDeclaration::formatImpl` keeps printing the modifiers in the canonical order, so `SHOW CREATE TABLE` output does not change. The `ALTER` surface is made consistent instead of merely permissive. `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN` applies the declared modifiers - `COMMENT`, `CODEC`, `STATISTICS`, `TTL` and per-column `SETTINGS` - and these can now be written in any order. Per-column `SETTINGS` in `ADD COLUMN` and a declared `STATISTICS` in `ADD COLUMN` / `MODIFY COLUMN` used to be silently dropped and are now applied (and validated), like in `CREATE`: `ADD COLUMN` sets the declared statistics on the new column, and `MODIFY COLUMN` replaces the explicit statistics of the column. The column-declaration `STATISTICS` in these `ALTER` commands honors the `allow_statistics` setting and requires the same `ALTER ADD STATISTICS` / `ALTER MODIFY STATISTICS` access rights, like the dedicated `ADD/MODIFY/DROP STATISTICS` commands, and is excluded from the comment-only fast path in the distributed DDL routing. It is also gated on engine support (new `IStorage::supportsStatistics`, true for the `MergeTree` family): storages that reject the dedicated `ADD/DROP/MODIFY STATISTICS` commands, such as `Memory` or `Distributed`, reject the column-declaration spelling with the same `NOT_IMPLEMENTED` error instead of accepting it through the generic column alter. `StorageAlias` forwards `supportsStatistics` to its target table, and `StorageProxy` forwards it (together with `supportsTTL`) to the nested table, so support does not depend on whether the table is addressed directly, through an `Alias`, or through the lazy-loading proxy of a database with `lazy_load_tables = 1`. `COLLATE` and `PRIMARY KEY` in `ADD COLUMN` / `MODIFY COLUMN` now throw an exception instead of being silently ignored: some spellings of them (with a type present) used to parse successfully and do nothing, which was a bug, not a feature - per-column `PRIMARY KEY` cannot be altered at all. This also fixes a formatting round-trip: `formatImpl` prints `COLLATE` after `COMMENT` and after `TTL`, while the parser used to accept `COLLATE` only immediately after the type or after `NULL`/`NOT NULL`, so the formatted result of `x String COLLATE utf8_bin COMMENT 'text'` did not parse back: ```sql SELECT formatQuery(formatQuery('CREATE TABLE t (a String COLLATE utf8_bin COMMENT \\'a comment\\') ENGINE = Memory')); -- before: Code: 62. DB::Exception: Syntax error: failed at position 48 (COLLATE). ``` ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The modifiers of a column declaration - `COMMENT`, `CODEC`, `STATISTICS`, `TTL`, `COLLATE`, `PRIMARY KEY` and per-column `SETTINGS` - can now be written in any order in `CREATE TABLE`, and each of them at most once. Previously only one fixed order was accepted, and, for example, `x UInt64 CODEC(ZSTD) COMMENT 'text'` was a syntax error. In `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN`, the supported modifiers - `COMMENT`, `CODEC`, `STATISTICS`, `TTL` and per-column `SETTINGS` - can also be written in any order; per-column `SETTINGS` in `ADD COLUMN` and a declared `STATISTICS` in `ADD COLUMN` / `MODIFY COLUMN` are now applied instead of being silently dropped, and `COLLATE` and `PRIMARY KEY` in these `ALTER` commands now throw an exception instead of being silently ignored.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112788",
          "createdAt": "2026-07-31T18:32:53Z",
          "updatedAt": "2026-08-13T03:25:41Z",
          "timestamp": "2026-08-13T03:25:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 21
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9b0537f049977dc92167",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114560",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114560",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113534 to 26.7: Fix a mixed JOIN ON condition evaluated over mismatched column types for a dictionary",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113534 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114423 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31653108878/job/94302376144)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114560",
          "createdAt": "2026-08-13T00:26:20Z",
          "updatedAt": "2026-08-13T03:23:47Z",
          "timestamp": "2026-08-13T03:23:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6ccf38dffc5fa3840b15",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113448",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113448",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Honor `use_statistics_cache` when loading per-part statistics",
          "text": "`IMergeTreeDataPart::getEstimates` (introduced in #87241) builds the per-part statistics estimates used by part pruning and by `system.parts_columns`. Before this change it always consulted a per-part estimates cache and ignored the `use_statistics_cache` setting (introduced in #88670): a session that sets `use_statistics_cache = 0` still got the cached values on this path. That made the opt-out inconsistent with the selectivity-estimator path, which already honors `use_statistics_cache`. This change makes `IMergeTreeDataPart::getEstimates(bool use_cache)` honor the setting: when `use_statistics_cache = 0` it loads statistics directly from disk and neither reads nor populates the per-part cache, matching the selectivity-estimator path. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `use_statistics_cache = 0` now also bypasses the per-part statistics estimates cache used by part pruning and `system.parts_columns`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113448",
          "createdAt": "2026-08-05T10:03:58Z",
          "updatedAt": "2026-08-13T03:20:42Z",
          "timestamp": "2026-08-13T03:20:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "zoomxi",
          "state": "closed",
          "assignees": [
            "hanfei1991"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f8228612d342b91084b5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:68493",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:68493",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Mongo queries (wire protocol + dialect)",
          "text": "Adds two ways to run MongoDB queries against ClickHouse: - a **wire protocol endpoint** (`mongo_port`), so MongoDB drivers and tools such as `pymongo` and `mongosh` can connect to ClickHouse as if it were a MongoDB server; - a **query dialect** (`SET dialect = 'mongo'`), so MongoDB shell syntax can be sent over the usual ClickHouse interfaces. A Mongo database maps onto a ClickHouse database, a collection onto a table, a document onto a row, and a nested field onto an `a.b` column. A collection created by the first `insert` gets one column per field of the first inserted document. The supported commands are `insert`, `find`, `count`, `update`, `delete`, `create`, `drop`, `createIndexes`, `listDatabases`, `listCollections`, `isMaster` and `saslStart`; only the `PLAIN` authentication mechanism is supported, because it is the only one that provides the cleartext password ClickHouse needs. Both are experimental and cover a subset of MongoDB. They are meant for pointing an existing MongoDB application at ClickHouse without rewriting its queries, not as a MongoDB replacement. The limitations are listed in `docs/en/interfaces/mongo.md`. Task: https://github.com/ClickHouse/ClickHouse/issues/58394 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a MongoDB-compatible wire protocol endpoint (the `mongo_port` server setting) and a MongoDB query dialect (`SET dialect = 'mongo'`) for basic collection operations.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/68493",
          "createdAt": "2024-08-17T07:09:19Z",
          "updatedAt": "2026-08-13T03:18:54Z",
          "timestamp": "2026-08-13T03:18:54Z",
          "metrics": {
            "reactions": 1,
            "comments": 25
          },
          "labels": [
            "can be tested",
            "pr-experimental"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ec19198f56375cb980f4",
        "signalId": "github:ClickHouse/ClickHouse:issue:114498",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114498",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Add a `string_bounds` column statistic for String columns",
          "text": "### Company or project name ClickHouse ### Use case Column statistics currently give string predicates almost nothing: - Range predicates (`url < 'https://m'`, `tenant BETWEEN 'a' AND 'b'`) fall through to the hard-coded `default_cond_range_factor = 0.33`, because `tdigest`/`basic`/`minmax` estimation is numeric-only. - `LIKE` / `ILIKE` always get the hard-coded `default_like_factor = 0.1`, whether the pattern matches 30% of rows or zero. - Equality against an impossible constant (e.g. `WHERE code = 'USA'` on a column where every value is 2 bytes long, or a constant outside the column's value range) is still estimated as if it could match. - `StatisticsPartPruner` only supports numeric MinMax, so parts can never be skipped based on statistics for string predicates — even on tables naturally clustered by a string column (tenant/customer IDs), unless the column is in the primary key or a skipping index. - There are no min/max string length or alphabet (all-ASCII) statistics for width estimates or future execution fast paths. All of these gaps are served by _bounds_, not frequencies — a statistic that is tiny, cheap to build, and losslessly mergeable. ### Describe the solution you'd like A new opt-in, string-specific column statistic, `string_bounds(K)` (default `K = 16`), storing per column per part: - truncated **min/max value bounds** — first `K` bytes of the smallest/largest string, each with an explicit exactness state (exact vs. truncated), compared bytewise (`memcmp` order); - exact **min/max string byte length**; - an **all-ASCII flag**; - a non-null row counter for coverage accounting. ```sql CREATE TABLE t (k UInt64, url String STATISTICS(basic, uniq_v2, string_bounds(16))) ENGINE = MergeTree ORDER BY k; ALTER TABLE t ADD STATISTICS url TYPE string_bounds(16); ALTER TABLE t MATERIALIZE STATISTICS url; ``` Initial uses: 1. **Impossibility detection** for `=` / `IN`: constants outside the value bounds or length bounds estimate to zero (before `mcv`/`countmin`/NDV machinery runs). 2. **Range estimation** for `<`, `<=`, `>`, `>=`, `BETWEEN` on strings: deterministic 0/all classification at the bounds, with an optional (setting-gated) interpolated point estimate in between. 3. **`LIKE 'prefix%'` / `startsWith`**: convert extractable fixed prefixes into ranges (reusing the existing `KeyCondition` prefix machinery) instead of the constant `default_like_factor`. 4. **Part pruning**: extend `StatisticsPartPruner` to string columns via the per-part bounds. Properties that make this the cheapest member of the statistics family: ~60 bytes per column per part, one `memcmp`-dominated pass per block to build, and exact, order-independent merging with no error terms — unlike `mcv`/`histogram`, merging never degrades accuracy. Scope for v1: `String`, `Nullable(String)`, and `LowCardinality` wrappers; explicit opt-in only (not in `auto_statistics_types`). `FixedString(N)` (which needs zero-padded comparison normalization) and nullable part pruning are follow-ups. ### Describe alternatives you've considered - **Extending the existing numeric `minmax` statistic to strings** — rejected: it stores exact typed values, whereas string bounds must be truncated with explicit exactness states, and the length/alphabet fields have no home there. - **Relying on `mcv` / `countmin` for strings** — those cover per-value frequencies of heavy hitters; they cannot answer range, prefix, or impossibility questions, and don't support part pruning. - **Putting length bounds / ASCII flag into `basic`** — viable later (its serialization allows additions), but keeping v1 a self-contained opt-in statistic leaves the default write path untouched. - **Storing full (untruncated) min/max values** — rejected: a single pathological long string would balloon a payload loaded during planning for every selected part; truncated bounds with exactness markers are the approach proven by Parquet (`is_min/max_value_exact`) and ORC (`lowerBound`/`upperBound`). ### Additional context DuckDB's per-segment string statistics drew my attention to this area in ClickHouse, which can benefit from zonemap-style min/max metadata applied to strings; similar truncated string min/max statistics also exist in Parquet and ORC.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114498",
          "createdAt": "2026-08-12T14:46:41Z",
          "updatedAt": "2026-08-13T03:17:29Z",
          "timestamp": "2026-08-13T03:17:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "feature",
            "st-need-info"
          ],
          "author": "cv4g",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:bc3bbc9932293a15401c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113107",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113107",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Disable the sampling query profiler under Memory Sanitizer",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/100242 Related: https://github.com/ClickHouse/ClickHouse/issues/106591 Every `MemorySanitizer` report produced in CI since 2026-05-28 is truncated to a single line: ``` ==117==WARNING: MemorySanitizer: use-of-uninitialized-value MemorySanitizer: nested bug in the same thread, aborting. ``` No stack trace, no `SUMMARY`, no origin — so no MSan failure can be located. In the CI database, `msan` checks recorded 1009 reports with a usable stack in 2026-04 and 0 in 2026-07, while `asan`, `tsan` and `ubsan` reports are unaffected and still carry full stacks. The regression window matches #100242 (merged 2026-05-28), which enabled the sampling query profiler for sanitizer builds. Printing an MSan report takes the sanitizer runtime seconds, because it symbolizes every frame through an `llvm-symbolizer` subprocess, and it holds `ScopedErrorReportLock` for the whole time. A profiler signal delivered to that thread meanwhile runs instrumented code in the handler and trips a second MSan check; compiler-rt treats a same-thread re-entry into the report lock as unrecoverable, writes `nested bug in the same thread, aborting.` with `CatastrophicErrorWrite` and calls `internal__exit`, so the report that was already being written to `MSAN_OPTIONS=log_path` stops after its header line. Reproduced with the `arm_msan` binary built from master `3885f7b7bbf1`, running the same query against the same server three times and forcing a sanitizer report with `max_allocation_size_mb=32 allocator_may_return_null=0`: | configuration | result | | --- | --- | | profiler on (default `global_profiler_real_time_period_ns`) | `nested bug in the same thread, aborting.`, 1-line report | | `<trace_log remove=\"1\"/>` (no profiler at all) | full 31-line report with stack | | `trace_log` on, `global_profiler_*_period_ns = 0`, per-query profiler off | full 31-line report with stack | So it is the profiler timer signals, not the trace collector. This turns `QUERY_PROFILER_SUPPORTED` off under MSan, next to the existing TSan-on-macOS exclusion. The trace collector stays enabled: the memory profiler, `trace_profile_events` and `SYSTEM INSTRUMENT` all feed it from ordinary code rather than from a signal handler. Other sanitizers keep the profiler, because their checks fire only on genuinely invalid accesses, which the handler does not perform. `00974_query_profiler` and `01569_query_profiler_big_query_id` assert that samples reach `system.trace_log`, so they get `no-msan` back. The other profiler and `trace_log` tests either already carry `no-msan`, exercise the memory profiler / `trace_profile_events` / `SYSTEM INSTRUMENT` (unaffected), or make no assertion on sample counts. This does not fix the uninitialized read that the `BuzzHouse (amd_msan)` run below hit — that one is unidentifiable as reported. It makes the next occurrence diagnosable. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=3885f7b7bbf1b8791021e1c97ca2e4eea132de04&name_0=MasterCI&name_1=BuzzHouse%20%28amd_msan%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113107",
          "createdAt": "2026-08-03T14:09:30Z",
          "updatedAt": "2026-08-13T03:10:52Z",
          "timestamp": "2026-08-13T03:10:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4f342652d946842423b8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110127",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110127",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Framing formats: multiplex data, totals, extremes, progress, logs, and profile events in the HTTP response stream",
          "text": "A framing format multiplexes different response parts of the query in a single stream: chunks of data, totals and extremes, progress packets, profile events (metrics), and server logs — everything that the native protocol supports. This allows rich data exchange in the HTTP protocol. Framing formats are independent of output formats: they encapsulate bytes produced by any output format, by separating and potentially encoding these chunks of bytes. The concatenation of the payloads of all `data`, `totals`, and `extremes` packets is exactly what the output format would have produced without framing. Auxiliary packets (progress, logs, profile events, exceptions) are represented as JSON. The framing format is selected by the new query setting `framing_output_format` (BETA tier). It currently applies to the HTTP protocol and is ignored for other interfaces. The implemented framing formats: - `None` — transparently routes everything applicable (data, totals, extremes, progress) to the output format, and ignores everything that is not applicable (metrics, logs), so everything works as it is by default. - `EventStream` — frames packets as HTTP server-sent events (`text/event-stream`). It integrates with the HTTP protocol and throws an exception when not applicable. - `JSONEachPacketBase64` — every packet is a JSON object on a separate line; the formatted data is base64-encoded (suitable for binary output formats). - `JSONEachPacketString` — every packet is a JSON object on a separate line; the formatted data is put into a string. Example: ``` $ curl \"http://localhost:8123/?framing_output_format=EventStream\" -d \"SELECT number FROM numbers(3) FORMAT JSONEachRow\" event: progress data: {\"read_rows\":\"3\",\"read_bytes\":\"24\",\"total_rows_to_read\":\"3\",\"elapsed_ns\":\"684907\"} event: data data: {\"number\":0} data: {\"number\":1} data: {\"number\":2} event: profile_events data: {\"host_name\":\"localhost\",\"current_time\":\"2026-07-11 22:38:21\",\"thread_id\":\"0\",\"type\":\"increment\",\"name\":\"SelectedRows\",\"value\":\"3\"} ... ``` Implementation: a framing format works as a multiplexor. The output format writes into the framing format's payload buffer, and `IOutputFormat` notifies the framing format on packet boundaries (under its writing mutex), which wraps everything accumulated since the previous boundary into a packet of the corresponding kind. Progress is routed through the existing throttled concurrent-progress path. Server logs (with the `send_logs_level` setting) and profile events reuse `InternalTextLogsQueue` and `ProfileEvents::getProfileEvents` — the same mechanisms as the native protocol — newly attached for HTTP queries. Exceptions are always written as the last packet of the stream (regardless of `http_write_exception_in_output_format`), so the client can always parse the response as a stream of packets. Parallel formatting is not used when framing is enabled, because the framing format needs to know the packet boundaries. Processing of multiple queries at once is out of scope of the first implementation, but the design allows it: every packet can be extended with the information about the query index along multiple queries. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add framing formats, selected by the new setting `framing_output_format`: they multiplex different response parts of the query in a single HTTP response stream — chunks of data, totals and extremes, progress packets, profile events, server logs, and exceptions. Implemented framing formats: `None` (default, everything works as before), `EventStream` (HTTP server-sent events), `JSONEachPacketBase64`, and `JSONEachPacketString` (a JSON object per packet with base64-encoded or string data). - [x] Documentation entry for user-facing changes: `docs/en/interfaces/framing-formats.md`",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110127",
          "createdAt": "2026-07-11T22:45:02Z",
          "updatedAt": "2026-08-13T03:10:18Z",
          "timestamp": "2026-08-13T03:10:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5633d41167002b520c90",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112932",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112932",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reimplement the KQL (Kusto) dialect on a lexer, an AST and AST translation",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/61742 --> ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reimplemented the experimental KQL (Kusto) dialect. It now has its own lexer and parser and builds an AST directly, instead of translating a token stream into SQL text and reparsing it. This fixes an expression-injection hole in the string operators (`contains`, `has`, ...), several results that silently disagreed with Kusto (`7 / 2`, `substring` with a negative start, `bin`), and 28 functions that were registered but did nothing. The supported subset is smaller and documented; anything outside it is now rejected by name instead of being mistranslated. ### Documentation entry for user-facing changes `docs/guides/clickhouse/kusto-query-language.mdx` --- ## Why The dialect was contributed in 2022 and abandoned by its authors; the rewrite they promised in #62668 was closed unmerged. Since then it has been emergency-disabled once (#59305) and gated behind `allow_experimental_kusto_dialect` (#74224), and #61742 (\"Remove KQL support if this code will not be fixed\") has stayed open. The implementation translated a KQL token stream directly into ClickHouse SQL **text** and reparsed it — 298 `fmt::format` sites, 41 places that re-lexed generated text, and no AST anywhere in the function layer. Everything below follows from that: | | before | after | |---|---|---| | `T \\| where s contains \"x') OR 1 = 1 OR ilike(s, 'y\"` | injected the expression | needle is an `ASTLiteral`; no injection is representable | | `'50x' contains '50%'` | `true` — needle was pasted into a LIKE pattern | `false` | | `substring('abcdefg', -3, 2)` | `'ab'` | `'ef'` | | `7 / 2` | `3.5` | `3` | | `format_timespan(col, 'hh:mm')` | `stoi: no conversion` — `std::stoi` on generated SQL | rejected by name | | `series_fir(...)` | `Function series_fir does not exist` (one of 28 stubs) | rejected by name | | `search 'x'` | `Unknown table expression identifier 'search'` | `'search' is not a supported KQL operator` | | `let` bindings | `static thread_local`, leaked between queries | scoped to one parse | ## What is here Two commits: the removal, then the reimplementation. ``` src/Parsers/Kusto/ KQLLexer KQL's own tokens: !in, =~, .., timespans (2.5h), verbatim @'...', datetime(...). A bad literal is an Error token carrying a reason, so nothing downstream asks isValidKQLPos() — the core of #61742. KQLAST the tabular level only. Scalar expressions are ClickHouse AST directly; a parallel expression hierarchy would add only conversions. KQLParser recursive descent. `let` bindings are an ordinary member. KQLTranslator each operator fills a still-empty clause of the select being built, or wraps what exists so far in a subquery and starts a new one. KQLFunctions name -> builder returning an ASTPtr. src/Functions/Kusto/ kqlDivide `7 / 2` is 3 in KQL and 3.5 in SQL — the choice depends on operand kqlBin types, so it is made by an IFunctionOverloadResolver during analysis rather than guessed from how the argument was spelled. ``` KQL now has its own entry point (`parseKQLQuery`) and never touches the SQL tokenizer. Because `ClientBase` already catches, the parser can simply throw — so there is no `tryParseKQLQuery`, and no `catch` anywhere in the new code. **`src/Parsers/Kusto` no longer needs its carve-out from check 19** (\"do not catch exceptions in src/Parsers\"), and this PR deletes it. `src/Parsers/Kusto` goes from 13017 lines to 3967, plus 337 lines of runtime functions. ## Scope Deliberately smaller than before, and written down in the guide. A construct is either translated with the semantics Kusto documents, or rejected by name. Rejected in this PR: `search`, `parse`, `mv-apply`, `lookup`, `evaluate`, `invoke`, `facet`, `top-nested`, `make-series` and friends; the `series_*`, `bag_*`/`pack_*` and `ipv4_*` families; `parse_url`, `parse_csv`, `parse_json`, `toscalar`, `format_timespan`; and `dynamic` **objects** (arrays map onto ClickHouse `Array`). Three known divergences are documented rather than papered over: subtracting two datetimes yields seconds instead of a timespan, `project-rename` moves the renamed column to the end, and `union` needs compatible schemas. ## Testing `04670`–`04673` cover the pipeline, the string operators (including the injection cases), the scalar semantics that used to be wrong, and 47 constructs that must be rejected. A 2500-query random fuzz found no crashes or logical errors. The 302 conformance cases that were deleted with the old implementation are being reintroduced separately, on top of this branch.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112932",
          "createdAt": "2026-08-01T16:56:06Z",
          "updatedAt": "2026-08-13T03:08:13Z",
          "timestamp": "2026-08-13T03:08:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:594b2f410296ce555dad",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114557",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114557",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Unattended `continue-pr-auto` skill and the `continue-all-prs` driver",
          "text": "Adds the unattended `continue-pr-auto` skill and the `utils/continue-all-prs.sh` driver that runs it over the open pull requests, with a status bar renderer (`utils/continue-all-prs-status.py`) and an `utils/exclude-authors.txt` list. `continue-pr` keeps its interactive behavior (it may still ask about ambiguous conflicts and unclear reviewer suggestions); `continue-pr-auto` never asks, resolves conflicts autonomously, and always pushes, so `disable-model-invocation: true` keeps interactive sessions on `continue-pr`. Both skills now determine the repository they run in and pick whom to ping about CI failures unrelated to the pull request from it: `@groeneai` in `ClickHouse/ClickHouse`, `@oranjeai` in `ClickHouse/clickhouse-private`. Every comment the unattended skill posts is prefixed with 🕵 so automated comments are identifiable. The branch was stale, so it also merges current `master`, and it includes the `continue-pr` change from the pull request below (that change touches the same lines, so both are here to keep this branch based on `master` directly). Related: https://github.com/ClickHouse/ClickHouse/pull/114556 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1304` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114557",
          "createdAt": "2026-08-12T23:20:59Z",
          "updatedAt": "2026-08-13T03:03:28Z",
          "timestamp": "2026-08-13T03:03:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:19b9a3648f9a135eba1c",
        "signalId": "github:ClickHouse/ClickHouse:issue:114512",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114512",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "`PREWHERE <bare column>` with `FINAL` throws THERE_IS_NO_COLUMN when PREWHERE is deferred after FINAL",
          "text": "🕵️ ## Describe what's wrong An explicit `PREWHERE` whose expression is a plain column reference fails when PREWHERE is deferred until after `FINAL`: ``` Code: 8. DB::Exception: Cannot find column `b` in source stream, there are only columns: [other]. (THERE_IS_NO_COLUMN) ``` This is reachable **with stock default settings** — no `SETTINGS` clause required. The deferral is triggered automatically by a row policy on a non-sorting-key column, because `apply_row_policy_after_final` defaults to `1` and PREWHERE must run after the policy: ```sql CREATE TABLE t (k UInt64, b UInt8, other UInt8) ENGINE = ReplacingMergeTree() ORDER BY k; INSERT INTO t SELECT number, number % 2, number % 3 FROM numbers(20); CREATE ROW POLICY pol ON t USING other = 1 TO ALL; SELECT count() FROM t FINAL PREWHERE b; -- Code: 8. THERE_IS_NO_COLUMN ``` The same failure occurs without any row policy by opting into the deferral directly, on every FINAL-capable engine tested (`ReplacingMergeTree`, `CoalescingMergeTree`, `AggregatingMergeTree`, `SummingMergeTree`): ```sql SELECT count() FROM t FINAL PREWHERE b SETTINGS apply_prewhere_after_final = 1; ``` ## What makes it fire The deferred filter step resolves its column against the source stream, but nothing adds the PREWHERE expression's own required columns to the set of columns read from the table. So it only succeeds *by accident*, when the column is already in the SELECT list: | projection | `PREWHERE b` (plain `UInt8`) | `PREWHERE n.null` (subcolumn) | |---|---|---| | `SELECT count()` | **fails** (stream: `[]`) | **fails** (stream: `[]`) | | `SELECT k` | **fails** (stream: `[k]`) | **fails** (stream: `[k]`) | | `SELECT b` | ok (`b` happens to be in the stream) | **fails** | | `SELECT *` | ok (`b` happens to be in the stream) | **fails** (subcolumns are not in the stream) | Two further properties, both consistent with \"the required columns are never requested\": - **Only a bare column reference fails.** Any wrapping makes it work, because the wrapping expression's DAG declares the column as an input: `NOT b`, `b = 1`, `toUInt8(b)` and `b AND k > 5` all succeed where bare `b` does not. - **A subcolumn never works**, for any projection, since subcolumns are not present in the source stream even for `SELECT *`. Controls that behave correctly: the same query without `FINAL`, the same query with `apply_prewhere_after_final = 0`, and the same table without a row policy. ## Root cause `ReadFromMergeTree::deferFiltersAfterFinalIfNeeded` (`src/Processors/QueryPlan/ReadFromMergeTree.cpp:2509`) moves the PREWHERE aside for later: ```cpp if (query_info.prewhere_info && isPrewhereDeferredAfterFinal()) deferred_prewhere_info = query_info.prewhere_info; ``` `isPrewhereDeferredAfterFinal` (`:2500`) returns true for `apply_prewhere_after_final`, or whenever `isRowPolicyDeferredAfterFinal` does — which for a non-sorting-key policy is the default: ```cpp /// PREWHERE must run after the row policy, so deferred row policy defers PREWHERE as well return context->getSettingsRef()[Setting::apply_prewhere_after_final] || isRowPolicyDeferredAfterFinal(); ``` Once deferred, the PREWHERE's required columns still have to be read from the table so the later filter step can evaluate them, and that does not appear to happen for a bare column reference. ## Does it reproduce on the most recent release? Reproduced on `master`, version 26.8.1.1226 (debug build). ## How to reproduce Full matrix script: `tmp/repro_prewhere_subcol/repro.sh`. Minimal case is the 4-statement default-settings snippet above. ## Expected behaviour `SELECT count() FROM t FINAL PREWHERE b` returns the count of matching rows, exactly as `SELECT count() FROM t FINAL PREWHERE NOT b` and `... PREWHERE b = 1` already do. ## Additional context Adjacent but distinct from #109703 (\"Do not move conditions to PREWHERE when PREWHERE is deferred after FINAL\"), which stopped the *optimizer* from moving `WHERE` into a deferred PREWHERE. That fix is why `WHERE b` works here while an explicit, user-written `PREWHERE b` still takes the broken path — the deferral machinery itself was not made safe, only the automatic move into it. The deferral feature comes from #91065 (\"Apply row policies and PREWHERE after FINAL\"). Found by a PREWHERE correctness fuzzer (`tmp/fuzz_prewhere/`) that compares every PREWHERE variant against the same filter with all PREWHERE paths disabled; this surfaced as `apply_prewhere_after_final` disagreeing with the ground truth on `PREWHERE n.null`, and minimising showed the subcolumn was incidental — a plain `UInt8` column fails the same way.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114512",
          "createdAt": "2026-08-12T16:01:41Z",
          "updatedAt": "2026-08-13T02:59:32Z",
          "timestamp": "2026-08-13T02:59:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e42980f8c316faf1ca0d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109455",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109455",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Native Google Cloud Storage integration (google-cloud-cpp)",
          "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental native Google Cloud Storage integration built on the Google Cloud C++ SDK (`google-cloud-cpp`, the GCS JSON API), as an alternative to the existing S3-compatibility path. Enable it with the new `use_native_gcs` setting for the `gcs` table function and the `GCS` table engine, or use it as a MergeTree storage disk via `object_storage_type: gcs`. ### Documentation entry for user-facing changes Documented in `docs/reference/functions/table-functions/gcs.mdx` (new \"Native GCS integration\" section). --- ## Description ClickHouse currently talks to GCS **only through its S3-compatible XML API** (the AWS SDK against a `storage.googleapis.com` endpoint, plus a cluster of GCS quirk-toggles in the S3 client). This PR adds a first-class native backend on top of the already-vendored `google-cloud-cpp` storage client. **Opt-in, non-breaking.** Everything is gated behind the new experimental setting `use_native_gcs` (default `false`). With it off, `gcs()` / `ENGINE = GCS` behave exactly as before (S3-compatibility, HMAC keys). With it on, they route through the native client. ### What's added - `ObjectStorageType::GCS` + `GCSObjectStorage : IObjectStorage` (`src/Disks/DiskObjectStorage/ObjectStorages/GCS/`), with `ReadBufferFromGCS`/`WriteBufferFromGCS` over the SDK's `ObjectReadStream`/`ObjectWriteStream` (ranged seeks, resumable uploads), native listing, `RewriteObject`-based copy, `DeleteObject`, `GetObjectMetadata`. - Registered as the `gcs` object-storage type + a `gcs` disk-type alias, so MergeTree data can live natively on GCS (`type: object_storage, object_storage_type: gcs`). - `StorageGCSConfiguration` (reuses `StorageS3Configuration`'s argument parsing) and a `TableFunctionGCS` / `ENGINE = GCS` selector that picks native vs S3-compat by the `use_native_gcs` setting. - Auth mirrors the compat surface on the native side: Application Default Credentials, service-account JSON, the `google_adc_*` OAuth refresh-token flow (reusing `IO/GCPOAuth`), anonymous / `NOSIGN`, and a REST endpoint override for the GCS emulator. - `USE_GOOGLE_CLOUD` is now exposed in `system.build_options`. ### Testing - Every changed/new translation unit compiles against the real headers and the vendored SDK (`USE_GOOGLE_CLOUD=1`, `USE_AWS_S3=1`). - Unit test for endpoint parsing (`gtest_gcs_endpoint`). - Stateless test for the setting wiring (`04502_use_native_gcs_setting`). - Integration test `test_native_gcs` runs against a `fake-gcs-server` emulator (a new `with_gcs` helper in `cluster.py`): `gcs()` INSERT/SELECT + glob with `use_native_gcs=1`, and a MergeTree-on-GCS-disk round trip. Note: `fake-gcs-server` speaks the GCS API (unlike minio, which is S3-only). The module is skipped on builds without the SDK. ### Known follow-ups - `GCSObjectStorage::iterate()` currently materializes the full listing (a lazy paginating iterator like S3's is a future optimization). - The `google_adc_*` refresh-token access token is minted eagerly (no auto-refresh yet for long-lived disks). - `readSmallObjectAndGetObjectMetadata` uses the base default (relevant only for a future native Iceberg-over-GCS path, which is not wired here). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109455",
          "createdAt": "2026-07-05T22:03:13Z",
          "updatedAt": "2026-08-13T02:51:58Z",
          "timestamp": "2026-08-13T02:51:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "submodule changed",
            "pr-experimental",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:79ff8fa6d1f6e4766f0b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114574",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114574",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #112498 to 26.5: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114574",
          "createdAt": "2026-08-13T02:35:19Z",
          "updatedAt": "2026-08-13T02:35:26Z",
          "timestamp": "2026-08-13T02:35:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7987715e1ab68e617e48",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114573",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #112498 to 26.3: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114573",
          "createdAt": "2026-08-13T02:34:42Z",
          "updatedAt": "2026-08-13T02:34:49Z",
          "timestamp": "2026-08-13T02:34:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7ac92f2c515ce72ccd61",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112940",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112940",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reapply \"Make `postgresql` and `PostgreSQL` engine work against a ClickHouse instance\", with a fix for the cancel-request logical error",
          "text": "Reapply #110760, reverted in #112935, together with a fix for the bug that motivated the revert. Related: https://github.com/ClickHouse/ClickHouse/pull/110760 Related: https://github.com/ClickHouse/ClickHouse/pull/112935 Related: https://github.com/ClickHouse/ClickHouse/issues/84085 ### Why it was reverted, and why the bug is not in the reverted change Every `Stress test` job on master started failing with ``` Logical error: 'Query context must be created after authentication' DB::Session::makeQueryContextImpl @ src/Interpreters/Session.cpp:690 DB::PostgreSQLHandler::cancelRequest @ src/Server/PostgreSQLHandler.cpp:867 DB::PostgreSQLHandler::startup @ src/Server/PostgreSQLHandler.cpp:737 ``` CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=cd6fd72e2b29fa607cffe0595e347452f8dcba9b&name_0=MasterCI&name_1=Stress%20test%20%28amd_debug%29 A PostgreSQL client cancels a running statement by opening a *second* connection and sending a `CancelRequest` on it, carrying the `(process id, secret key)` pair the server handed out in `BackendKeyData`. By the protocol that connection never authenticates - the secret key is the credential - so `PostgreSQLHandler::cancelRequest` has no authenticated session, and the `session->makeQueryContext()` it called throws `LOGICAL_ERROR`. That has been true since sessions were introduced (`51ffc334573`), and it is reachable by any PostgreSQL client that cancels a statement: `psql` on Ctrl-C, or the `pgx` driver in issue #84085, which reports this very error. #110760 did not introduce it, it only made CI reach it: with ClickHouse acting as a libpq/pqxx client against itself, `PostgreSQLSource` cancels the remote statement through `PQcancel`, so a stress run that cancels a query now sends a cancel request to the server's own PostgreSQL port. ### The fix (second commit) - `cancelRequest` cancels the query through the process list directly, which needs no session, via a new `ProcessList::sendCancelToQueryOfAnyUser`. Only queries whose id has the `postgres:<connection id>:<secret key>` shape the server itself assigns can be named this way, and the secret key is what makes such an id unguessable - the same credential PostgreSQL relies on. - The raw `KILL QUERY` result the old code wrote straight into the client socket is gone with it. The protocol expects no answer at all to a cancel request, so that was protocol garbage. - The secret key never matched anything either: `BackendKeyData` was sent while `secret_key` was still zero, and every statement then re-randomized it, so the query id a cancel request resolved to was never the id of a running query - cancellation over the PostgreSQL protocol has never worked. The key now belongs to the connection, as in PostgreSQL, where it identifies the backend rather than one statement, and every statement of the connection runs under `postgres:<connection id>:<secret key>`. Statements of one connection run one after another, so reusing the id is safe. - New test `04669_postgresql_protocol_cancel_request`, which drives a real `psql` and cancels it with `SIGINT` - that is what makes `psql` send a `CancelRequest`. The `-- ping` half of issue #84085 (a comment-only statement is not recognized as an empty query) is not addressed here. ### Privileges A self-connect needs nothing beyond access to the table being read. `system.databases`, `system.tables` and `system.columns`, which the emulated `pg_namespace`, `pg_class` and `pg_attribute` are views over, are readable by every user even with `access_control_improvements.select_from_system_db_requires_grant` enabled, and their rows are filtered by the reader's own grants - so the catalog shows a user exactly the relations that user may read, as `pg_catalog` does in PostgreSQL. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `postgresql` table function and `PostgreSQL` table engine can now be used to connect to another ClickHouse server over the PostgreSQL protocol (when a table name is used; the `query(...)` variant is not supported yet). Added the PostgreSQL-compatibility functions `format_type` and `current_setting`. Fixed cancellation over the PostgreSQL protocol: a cancel request used to log a logical error and never cancelled anything.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112940",
          "createdAt": "2026-08-01T18:51:27Z",
          "updatedAt": "2026-08-13T02:34:26Z",
          "timestamp": "2026-08-13T02:34:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 20
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:01f842fd2bdf69432575",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114572",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114572",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #112498 to 25.8: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114572",
          "createdAt": "2026-08-13T02:34:01Z",
          "updatedAt": "2026-08-13T02:34:09Z",
          "timestamp": "2026-08-13T02:34:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7d8f97c0280f3292b674",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113023",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113023",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Rewrite `length(arrayFilter(f, arr))` to `arrayCount(f, arr)`",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/2120 --> `arrayFilter` builds an array of the elements that pass the predicate, and `length` then throws that array away and keeps only its size. `arrayCount` computes the same number without materializing anything. ```sql SELECT length(arrayFilter(x -> (x >= 2), arr)); -- becomes SELECT arrayCount(x -> (x >= 2), arr); ``` The rewrite is a new query tree pass in the spirit of the existing `RewriteArrayExistsToHasPass`, controlled by the new `optimize_rewrite_array_filter_length_to_array_count` setting (on by default). A new node is built rather than rewriting the `arrayFilter` node in place, because the filtered array can be referenced elsewhere in the query tree, where it is still an array. `arrayCount` counted into `UInt32`, which silently wrapped for a row whose array has more than `4294967295` matching elements, while `length(arrayFilter(...))` counts through `ColumnArray::Offset` and is exact. It now counts into `UInt64` and returns `UInt64`: this fixes the overflow of `arrayCount` itself, and makes the rewrite value-preserving for arrays of any size. The result type of `arrayCount` therefore changes from `UInt32` to `UInt64`; the values are unchanged. Because `length` and `arrayCount` now have the same result type, the rewritten node needs no `CAST`, and the pass only rewrites when the two result types match. Since this changes the stable public signature of an existing function, the previous behavior is kept behind the new compatibility setting `array_count_legacy_uint32_result` (default `false`): setting it to `true` restores the `UInt32` result type, and it is wired into the settings changes history, so `compatibility = '26.7'` (or older) restores it automatically on the servers where it is set. For distributed queries, the setting (like any setting) is forwarded from the initiator to the shards, so a query initiated by a server that has it set behaves exactly as before on every shard. The one case it does not cover is a rolling upgrade with a *not-yet-upgraded initiator*: an old server does not know the setting and cannot forward it, so type-sensitive expressions evaluated locally on already-upgraded shards (for example, `byteSize(arrayCount(...))`) observe `UInt64` there, while the initiator still converts the top-level result to its own `UInt32` header. To keep such queries fully unchanged during the upgrade, set `array_count_legacy_uint32_result = 1` on the upgraded servers for the users under which shard-side queries execute (with an interserver `secret` configured that is the initiator's current user; otherwise it is the user from the cluster definition or the `remote` table function - the simplest robust approach is to enable it for all users of the upgraded servers), and remove it once the whole cluster is upgraded. This is documented in the `arrayCount` documentation and in the setting's description, and both directions of the mixed-version scenario are pinned by the integration test `test_backward_compatibility/test_array_count_return_type.py`. In legacy mode the rewrite does not fire, because the pass requires the result types to match. Measured with the setting toggled on the same binary (release, aarch64), best of five, before the `CAST` was removed: | query | setting off | setting on | | |---|---|---|---| | one 30M-element array, half the elements pass | 0.38 s | 0.33 s | 1.15x | | one 30M-element array, all elements pass | 0.40 s | 0.31 s | 1.29x | | one 10M-element `String` array | 0.40 s | 0.37 s | 1.08x | | 5M rows x 32-element arrays, unpredictable predicate | 1.42 s | 1.29 s | 1.10x | | 1M rows x 8-element `String` arrays | 0.25 s | 0.22 s | 1.14x | | 5M rows x 32-element arrays, half the elements pass | 1.41 s | 1.48 s | 0.95x | The gain grows with the array length, which is where the copy that `arrayFilter` does starts to cost something. The last row loses 5%: for short arrays the copy is cheap, `ArrayCountImpl` counts the filter with a scalar loop while `arrayFilter` copies with vectorized code, and the cast to `UInt64` added a column of its own - that cast is no longer generated. I did try replacing that scalar loop with `countBytesInFilter`, but its vectorized path is `__SSE2__` only, so on the aarch64 machine I measured on it made no difference and I dropped it rather than commit something I could not verify. It may be worth doing separately, measured on x86. Found while profiling https://github.com/ClickHouse/ClickHouse/issues/2120, where `length(arrayFilter(x -> (x >= 2), arrayEnumerateUniq(...)))` is applied to a single array of 60 million elements. Related: https://github.com/ClickHouse/ClickHouse/issues/2120 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `arrayCount` now returns `UInt64` instead of `UInt32`, so that it is exact for arrays with more than 4294967295 matching elements; set the new setting `array_count_legacy_uint32_result` to `true` (or use `compatibility = '26.7'`) to restore the previous result type. During a rolling upgrade, set `array_count_legacy_uint32_result = 1` on the upgraded servers for the users under which shard-side queries execute (the simplest robust approach is to enable it for all users of the upgraded servers), so that distributed queries initiated by not-yet-upgraded servers (which cannot forward the setting) keep the previous behavior on upgraded shards; remove it after the upgrade is complete. In addition, `length(arrayFilter(func, arr))` is now rewritten to `arrayCount(func, arr)`, which counts the matching elements instead of building an array of them only to take its size - up to 1.29x faster on long arrays, controlled by the `optimize_rewrite_array_filter_length_to_array_count` setting.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113023",
          "createdAt": "2026-08-02T18:39:16Z",
          "updatedAt": "2026-08-13T02:22:02Z",
          "timestamp": "2026-08-13T02:22:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "pr-performance",
            "pr-backward-incompatible"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e37f4cc31b35c18c925b",
        "signalId": "github:ClickHouse/ClickHouse:issue:112586",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112586",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Range condition in JOIN ON against a single-row side is not used for index analysis (no part pruning) since the logical join step became default",
          "text": "### Company or project name ClickHouse Inc. Found while comparing two server versions on an internal test cluster, on a reporting workload that derives its date window from a joined single-row subquery. ### Describe the situation When a range predicate on a `MergeTree` primary/partition key column is expressed in the `ON` clause of a `JOIN` against a single-row constant side, the condition is no longer turned into a `KeyCondition`. Index analysis reports `Condition: true` and every part is read, instead of pruning. The logical join step turns the inequality `ON` into a `cross` join plus a `Filter (Post Join Actions)` sitting above the join, and that filter never reaches index analysis. The same predicate written as a plain `WHERE` (or as scalar subqueries in `WHERE`) prunes correctly, and results are identical either way, so this is purely lost part/granule pruning. This is a common shape for \"current week / current month\" reporting queries, where the bounds are computed once in a CTE and joined in: ```sql WITH bounds AS (SELECT toStartOfWeek(today(), 1) AS lo, today() AS hi) SELECT ... FROM big_table AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi ``` On a real workload (a ~365M row table partitioned by month) this was the difference between reading 0 parts and reading all 302 parts, and between **1.4s and 35-68s (~25x)** for a query whose correct answer was an empty result set. ### Which ClickHouse versions are affected? The pruning is lost whenever the logical join step is used, i.e. whenever `query_plan_use_new_logical_join_step = 1`. **The regression is 26.5.** That is where https://github.com/ClickHouse/ClickHouse/pull/104017 made the setting `Obsolete`, hardcoding it to `true`. Up to and including 26.4 it was a normal `Production` setting, so anyone hitting this could simply set it to `0` and get pruning back. From 26.5 on there is **no setting-level mitigation at all** and the only fix is to rewrite the query. For completeness on the earlier history: the declared default went `false` -> `true` in 25.2 (https://github.com/ClickHouse/ClickHouse/pull/74909), so a 25.2..26.4 user who never touched the setting also saw no pruning. But that was recoverable, and deployments that pinned `query_plan_use_new_logical_join_step = 0` (ClickHouse Cloud among them, via its `compatibility` profile) kept correct pruning all the way through 26.4 and only lost it on upgrade past 26.5. Verified with official builds: | Version | Default behavior | With `query_plan_use_new_logical_join_step = 0` | |---|---|---| | 26.3.1.876 | `Parts: 14/14` (no pruning) | `Parts: 1/14` (prunes) | | 26.7.1.2043 | `Parts: 14/14` | setting is `Obsolete`, no effect | | 26.8.1.48 (master, `d38cb514a590e60af07777d8f48c34afdf20aa1f`) | `Parts: 14/14` | setting is `Obsolete`, no effect | `SETTINGS compatibility = '26.4'` does **not** restore the pruning on master either. ### How to reproduce Works in `clickhouse local`, no special settings needed. ```sql CREATE TABLE repro_join_prune (d Date, v UInt32) ENGINE = MergeTree PARTITION BY toYYYYMM(d) ORDER BY d; INSERT INTO repro_join_prune SELECT toDate('2025-01-01') + number, number FROM numbers(400); -- (1) bounds arrive via JOIN ON: no pruning on 25.2+ EXPLAIN indexes = 1 WITH bounds AS (SELECT toDate('2025-06-01') AS lo, toDate('2025-06-10') AS hi) SELECT count() FROM repro_join_prune AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi; -- (2) same predicate as a plain WHERE: prunes correctly on every version EXPLAIN indexes = 1 SELECT count() FROM repro_join_prune AS t WHERE t.d >= toDate('2025-06-01') AND t.d <= toDate('2025-06-10'); -- (3) same predicate as scalar subqueries: also prunes correctly on every version EXPLAIN indexes = 1 SELECT count() FROM repro_join_prune AS t WHERE t.d >= (SELECT toDate('2025-06-01')) AND t.d <= (SELECT toDate('2025-06-10')); ``` All three queries return `10`, so results are correct in every case. Plan for query (1) on master. Note that the range predicate is present as a `Filter (Post Join Actions)` but index analysis still gets `Condition: true`: ``` Aggregating │ Keys: │ Aggregates: count() │ Skip merging: 0 └──Filter (Post Join Actions) │ Filter column: d >= '2025-06-01' AND d <= '2025-06-10' └──Join (JOIN FillRightFirst) │ t[400] ⋈ system.one[1] │ Type: cross | Strictness: all | Algorithm: HashJoin │ Result rows: 400 │ Output: │ Left: __join_result_dummy, d │ Right: hi, lo ├──ReadFromMergeTree (default.repro_join_prune) │ Read type: Default │ Parts: 14 | Granules: 14 │ Output: d │ Indexes: │ Min-Max │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ Partition │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ PrimaryKey │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ Ranges: 14 └──ReadFromSystemOne ``` Plan for query (2) on master, for comparison: ``` Aggregating │ Keys: │ Aggregates: count() │ Skip merging: 0 └──Filter ((WHERE + Change column names to column identifiers)) │ Filter column: d >= '2025-06-01' AND d <= '2025-06-10' └──ReadFromMergeTree (default.repro_join_prune) Read type: Default Parts: 1 | Granules: 1 Output: d Indexes: Min-Max Keys: d Condition: and((d in (-Inf, 20249]), (d in [20240, +Inf))) Parts: 1/14 Granules: 1/14 Partition Keys: toYYYYMM(d) Condition: and((toYYYYMM(d) in (-Inf, 202506]), (toYYYYMM(d) in [202506, +Inf))) Parts: 1/1 Granules: 1/1 PrimaryKey Keys: d Condition: and((d in (-Inf, 20249]), (d in [20240, +Inf))) Parts: 1/1 Granules: 1/1 Search Algorithm: binary search Ranges: 1 ``` To see the \"good\" plan for query (1), run it on 26.4 or earlier with `SETTINGS query_plan_use_new_logical_join_step = 0`. ### Expected performance Query (1) should prune the same way queries (2) and (3) do, since the join is against a single-row constant side and the `ON` conditions are a plain range over the sorting/partition key. It did prune before the logical join step became the default, and it still prunes on 26.4 and earlier with `query_plan_use_new_logical_join_step = 0`. Note that a `CROSS JOIN` with the same predicate moved into `WHERE` does **not** prune on any version tested, including 26.4 with the old join step. That looks like a separate, pre-existing limitation rather than part of this regression, but it may share a root cause. ### Related issues and pull requests Caused by: https://github.com/ClickHouse/ClickHouse/pull/104017 Related: https://github.com/ClickHouse/ClickHouse/pull/74909 ### Additional context The following settings were tried on 26.6+ and none of them restore the pruning: - `query_plan_merge_filter_into_join_condition = 1` - `query_plan_convert_join_to_in = 1` - `allow_general_join_planning = 1` - `query_plan_filter_push_down = 1` - `query_plan_convert_outer_join_to_inner_join = 1` - `query_plan_optimize_join_order_limit = 10` - `use_join_disjunctions_push_down = 1` - `compatibility = '26.4'` The only workaround we found is a query rewrite: duplicate the bounds as a plain `WHERE` on the scanned table. The `JOIN` can stay in place, so the rewrite is semantics-preserving, and it restored the real workload from 35-68s to 1.3s. ```sql WITH bounds AS (SELECT toDate('2025-06-01') AS lo, toDate('2025-06-10') AS hi) SELECT count() FROM repro_join_prune AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi WHERE t.d >= toDate('2025-06-01') AND t.d <= toDate('2025-06-10'); ``` Because the escape-hatch setting is now obsolete, users upgrading from 26.4 (or from any version where they had pinned `query_plan_use_new_logical_join_step = 0`) to 26.5+ hit this as a hard performance regression with no setting-level mitigation. <!-- ch-version-info:start --> ### Version info - Resolved by: #113484 - Backported to: `26.7.4.24` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112586",
          "createdAt": "2026-07-30T12:41:56Z",
          "updatedAt": "2026-08-13T02:20:30Z",
          "timestamp": "2026-08-13T02:20:30Z",
          "metrics": {
            "reactions": 2,
            "comments": 6
          },
          "labels": [
            "performance"
          ],
          "author": "fm4v",
          "state": "closed",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2b025642b8c411e9667e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100160",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100160",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Feature: Paimon minmax index",
          "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Supports Paimon min-max index pushdown by setting `use_paimon_minmax_index_pruning=1` ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100160",
          "createdAt": "2026-03-20T08:23:32Z",
          "updatedAt": "2026-08-13T02:19:03Z",
          "timestamp": "2026-08-13T02:19:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "submodule changed",
            "can be tested",
            "pr-experimental"
          ],
          "author": "JiaQiTang98",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e59ec6b41d567c423310",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:103158",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:103158",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix TSAN data race in executeASTFuzzerQueries clearing caller's transaction",
          "text": "The server-side AST fuzzer (`executeASTFuzzerQueries`, active only when `ast_fuzzer_runs > 0`) used to clear the transaction on the caller's query and session `Context` with no lock held: ```cpp context->getQueryContext()->getSessionContext()->setCurrentTransaction(NO_TRANSACTION_PTR); context->setCurrentTransaction(NO_TRANSACTION_PTR); ``` `Context::setCurrentTransaction` writes `merge_tree_transaction` unsynchronized, while other threads read the same field under the shared `Context::mutex` via `Context::createCopy`. Under `RESTORE ASYNC` with `ast_fuzzer_runs > 0`, `RestorerFromBackup` background workers keep copying the context while the fuzzer finish callback overwrites the transaction pointer, producing the TSAN race (STID 2604-385d read side, 3336-2c6d write side, seen on `Stress test`/`BuzzHouse (experimental, serverfuzz, *_tsan)`). ## Fix 1. Move the transaction reset onto the fuzz session-context copy, so the caller context is never mutated. 2. Skip fuzzed `BACKUP`/`RESTORE` queries. An async backup/restore returns from `executeQuery` immediately while `BackupsWorker` keeps the (fuzz) context alive for background workers that read it via `Context::createCopy`; mutating that escaped copy afterwards reintroduces the same race. The fuzzer has negligible value on backup/restore (no real backup target), so skipping them is safe. Tests: `04104_ast_fuzzer_preserves_caller_transaction` (caller transaction preserved across `ast_fuzzer_runs > 0`) and `04305_ast_fuzzer_skips_backup_restore` (server stays alive and data intact with fuzzed backup/restore). The whole change lives in fuzzer-only code gated on the `ast_fuzzer_runs` CI/stress-test setting, never a production user path, so this is categorized as a CI fix. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/103158",
          "createdAt": "2026-04-20T13:48:10Z",
          "updatedAt": "2026-08-13T02:17:00Z",
          "timestamp": "2026-08-13T02:17:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 20
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bb2371ee27c02a9f47f3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114070",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114070",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Convert stateless tests that modify the server's data on disk to integration tests",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113978 Related: https://github.com/ClickHouse/ClickHouse/pull/114056 Related: https://github.com/ClickHouse/ClickHouse/pull/114057 Stateless tests are not allowed to modify the server's data on disk at all (https://github.com/ClickHouse/ClickHouse/pull/113978#pullrequestreview-4891978370): the stateless suite runs against arbitrary server configurations — object storage, shared merge tree, encrypted disks — where the local part layout either does not exist or does not mean what the test assumes, and modifying it corrupts shared state. This PR converts all stateless tests that do such manipulation into integration tests, grouped by scenario, and adds a style check that rejects new ones: **Extended existing modules:** - `test_broken_projections` ← `02254_projection_broken_part`, `04511_check_table_broken_projection_columns` - `test_broken_tmp_txn_version_startup` ← `04104_transaction_version_metadata_dummy_tid_load`, `04492_attach_as_replicated_clears_tmp_txn_version`, `04493_attach_part_clears_tmp_txn_version`, `04507_packed_part_stale_txn_version_guard` - `test_lost_part` ← `02369_lost_part_intersecting_merges`, `02370_lost_part_intersecting_merges`, `04215_replicated_missing_covered_part_on_start` **New modules:** - `test_corrupted_part_files` (files inside part directories damaged/removed; loading, fetching, and `CHECK TABLE` reactions) ← `02253_empty_part_checksums`, `02255_broken_parts_chain_on_start`, `02444_async_broken_outdated_part_loading`, `04151_unique_key_sst_rebuild_on_load`, `04235_corrupted_columns_substreams_detection`, `04323_text_index_marks_empty_part`, `04506_packed_part_fetch_checksum` - `test_mutations_with_tampered_parts` (mutations over parts with old-version emulation or corrupted index files) ← `04401_dynamic_mutation_old_part_no_substreams_file`, `04412_mutation_old_part_basic_map_partial`, `04426_mutate_repair_corrupted_missing_idx_checksums`, `04427_mutate_some_columns_drop_index_corrupted_idx`, `04428_mutate_corrupted_text_index_multistream`, `04431_mutate_corrupted_index_sibling_owns_file` - `test_attach_tampered_detached_parts` (detached part directories manipulated before `ATTACH` / `DROP DETACHED`) ← `04063_drop_detached_part_with_try_n_suffix`, `04246_materialize_index_force_recalc` (the canned `part_25.8.tar.gz` moves into the module), `04402_mutate_all_columns_preserve_legacy_idx_minmax`, `04403_mutate_preserve_legacy_idx_packed_minmax`, `04404_mutate_rebuild_legacy_idx_minmax`, `04425_mutate_mixed_legacy_idx_minmax` - `test_backups_from_disk` (backups placed or tampered with directly on the `backups` disk) ← `04054_backup_restore_validate_entry_paths`, `04495_backup_metadata_version_overflow`, `02864_restore_table_with_broken_part`, `03001_restore_from_old_backup_with_matview_inner_table_metadata`, `03214_backup_and_clear_old_temporary_directories`, `03231_old_backup_without_access_entities_dependents`. The canned backup zips move from `tests/queries/0_stateless/backups/` into the module, and `helpers/install_predefined_backup.sh` is replaced by a module helper. - `test_server_metadata_files` (direct manipulation of on-disk table metadata `.sql` and the `flags/` directory) ← `03001_matview_columns_after_modify_query`, `04329_create_or_replace_force_drop_flag`, `04545_attach_projection_part_offset_setting_disabled` The conversions are faithful: every `.reference` output became an exact assertion, the originals' explanatory comments are preserved, and `DETACH`/`ATTACH` stayed `DETACH`/`ATTACH` except where the original explicitly emulated a server restart (`02255`, `04215` — \"on start\" scenarios now use a real server restart). Notable finds along the way: - `02444_async_broken_outdated_part_loading` contained a quoting bug since its introduction: `rm -f \"$path/*.bin\"` never expanded the glob, so the \"broken\" outdated part was never actually broken. The converted test applies the corruption for real, and the assertions still hold. - `02255_broken_parts_chain_on_start` dropped a nonexistent table (`projection_broken_parts_1`) in cleanup instead of its own tables. Verified locally against a fresh master binary: all new tests pass under the integration runner, and the tampered-parts modules were additionally cross-checked against a pre-fix binary where the bug-detecting assertions fire as intended. The style check (`server_data_manipulation_in_stateless_tests` in `ci/jobs/check_style.py`, folded in from https://github.com/ClickHouse/ClickHouse/pull/114056) flags stateless `.sh` tests that fetch a server-side filesystem path from a system table (`system.parts`, `system.detached_parts`, `system.projection_parts`, `system.tables`, `system.disks`, `system.server_settings` with `name = 'path'`) and also run file-modifying shell commands (`rm`/`cp`/`mv`/`dd`/`truncate`/`ln`/`chmod`/`touch`/`mkdir`/`tar`, `sed -i`, output redirection into a variable-derived path, `clickhouse-disks write`/`remove`/...). With all offenders converted here, the exclusion list holds only one documented false positive (`04326_disks_app_read_checksums`, which writes an `mktemp` scratch file under `CLICKHOUSE_TMP`) and `04630_merge_over_stale_packed_tmp_dir`, whose conversion is pending in https://github.com/ClickHouse/ClickHouse/pull/114057. Verified: the check is clean on this branch and catches each converted test if its stateless original is restored. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114070",
          "createdAt": "2026-08-09T20:47:29Z",
          "updatedAt": "2026-08-13T02:15:28Z",
          "timestamp": "2026-08-13T02:15:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:df58bc7809bfee4b7374",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113902",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113902",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: `break` does not always return a partial result",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112483 --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Both `timeout_overflow_mode` descriptions promise unconditionally that `break` returns \"the partial result, as if the source data ran out\", so a user picking `break` to avoid errors can still get one. `QueryStatus::checkTimeLimit` returns false rather than throwing under `break`, meaning \"stop and return what you have\". That works where a truncated output is a smaller valid one, not where it would be a wrong one, so several sites stop instead of truncating. A function computing one scalar value (`arrayFold`, `geohashesInBox`, string replace) stops without producing it, surfacing `TIMEOUT_EXCEEDED` only when not interrupted mid-pipeline, since a pipeline absorbs the throw into a clean cancellation. An incomplete `Memory`-table mutation (`StorageMemory::mutate`) leaves the table unchanged and reports it. A quorum write cannot report partial success at all, so it ignores the false return and does not stop at `max_execution_time`: it waits out `insert_quorum_timeout`, then reports `UNKNOWN_STATUS_OF_INSERT`. The mutation and dictionary-load waits are the same, hence \"some operations\". The `INSERT` code differs by executor: `QUERY_WAS_CANCELLED` when the caller pushes data block by block, as over the native protocol; \"may\", because an in-pipeline throw pre-empts that translation, and a pipeline completed with its own input source raises nothing. Retention is an engine property, so none is promised. That sentence also made `max_estimated_execution_time` a subject of `break`, but `ExecutionSpeedLimits.cpp:82` gates that check on `throw`, so the estimate is never evaluated. That gate predates this change, so only the text is corrected. The leaf setting has no estimate and reaches no `INSERT` or quorum write, so those clauses are omitted there. The sentence is shared boilerplate appearing 6 times in `Settings.cpp`; only these 2 are on the elapsed-time path, so the other 4 threshold modes are unchanged. Hedging is required in both directions: `FillingTransform` and `sleep` do return partial output. No behaviour changes. Raised by `clickhouse-gh[bot]` in https://github.com/ClickHouse/ClickHouse/pull/112483#discussion_r3686684259. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1300` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113902",
          "createdAt": "2026-08-07T23:52:27Z",
          "updatedAt": "2026-08-13T02:12:43Z",
          "timestamp": "2026-08-13T02:12:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-documentation",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ae2be7fbab4bb3656dcf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110943",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110943",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix ILLEGAL_TYPE_OF_ARGUMENT on nullable MySQL spatial columns",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/110933 Related: https://github.com/ClickHouse/ClickHouse/pull/108944 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `Nested type LineString cannot be inside Nullable type (ILLEGAL_TYPE_OF_ARGUMENT)` when reading a nullable MySQL spatial column (`LINESTRING`, `POLYGON`, `MULTILINESTRING`, `MULTIPOLYGON`, `MULTIPOINT`, or the generic `GEOMETRY`) through the `mysql` table function or the `MySQL` table/database engines. Such columns now fall back to `Nullable(String)` holding the value as MySQL returns it (a 4-byte SRID prefix followed by the WKB payload). `POINT` keeps mapping to `Nullable(Point)`. ### Description `convertMySQLDataType` maps the MySQL spatial types to the `Array`/`Variant`-based ClickHouse geometric types (`LineString`, `Polygon`, `MultiLineString`, `MultiPolygon`, `MultiPoint`, `Geometry`) and then unconditionally wrapped the result in `Nullable`. Those types return `canBeInsideNullable() == false`, so a nullable MySQL spatial column threw at schema-inference time and broke `DESCRIBE`/`CREATE`/`SELECT`. Nullable is MySQL's default for spatial columns and the geometry mapping is on by default, so this broke out of the box (regression from #108944). The fix guards the wrap with `canBeInsideNullable()`; a nullable column whose mapped type cannot be inside `Nullable` falls back to `Nullable(String)`, matching what the query-result-set overload already does for `MYSQL_TYPE_GEOMETRY`. The guard is a property test rather than a list of type names, so a future mapping that cannot be inside `Nullable` is covered too. `Point` deliberately does not take that fallback: it is a `Tuple`, so it can be inside `Nullable` and keeps mapping to `Nullable(Point)`. Added an integration test (`test_storage_mysql/test.py::test_mysql_nullable_geometry`) covering nullable spatial columns through both the `mysql` table function and the `MySQL` table engine. It asserts the inferred `Nullable(String)` schema, the exact bytes of a fallback value, that a MySQL NULL reads back as NULL (not a silently-defaulted empty geometry), and that a nullable `POINT` still reads back as `Nullable(Point)` so an over-broad fix would be caught. The `mysql_datatypes_support_level` setting text and the MySQL engine docs are updated to match, including the previously missing `MULTIPOINT` row in the type-mapping table. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1302` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110943",
          "createdAt": "2026-07-18T15:36:57Z",
          "updatedAt": "2026-08-13T02:12:39Z",
          "timestamp": "2026-08-13T02:12:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 24
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:437fc514d4cec225ac63",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113557",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113557",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check table name length on RENAME DATABASE unconditionally",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/101747 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/101747 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `RENAME DATABASE` accepting a target name long enough to make the database's tables impossible to drop. The table name length check now runs regardless of `check_table_dependencies`, and it also covers detached tables, which were never checked and so were affected even at default settings. Closes #101747. ### Description `RENAME DATABASE` could leave every table in the database permanently undroppable: `DROP TABLE` then failed with `Code: 1001 ... in rename: File name too long`, recoverable only by renaming the database back to a shorter name. Two independent holes caused it. The length check sat inside the dependency-check guard, so `SET check_table_dependencies = 0` skipped it. The two invariants are unrelated: the check enforces a filesystem limit on `metadata_dropped/{db}.{table}.{uuid}.sql`, while the guard expresses a preference about dependency validation. Detached tables were never checked at all, and this half needed **no** non-default setting. `DETACH TABLE` moves the table out of the attached-tables map, so the existing loop could not see it; `ATTACH TABLE` does not re-check the length and re-attaching binds the table to the new database name. A database whose only table is detached has an empty attached-tables map, so the loop had nothing to iterate even at default settings. The check now runs in its own loop over both containers, before the symlink removal, the metadata `moveFile` and `updateDatabaseName`, so a rejection leaves no partial rename. `DatabaseReplicated::renameDatabase` delegates here before writing to ZooKeeper and is covered; `DatabaseOrdinary` has no override and throws `NOT_IMPLEMENTED`. A `RENAME DATABASE` that previously succeeded under `check_table_dependencies = 0` can now return `ARGUMENT_OUT_OF_BOUND`. The newly rejected renames are exactly those that would have produced undroppable tables (proof below), and renaming a database *shorter* only raises the limit, so the documented recovery path is unaffected. No setting default changes, so `SettingsChangesHistory.cpp` is not updated. New test `04700_rename_database_name_length_gate` has 13 arms: 5 that flip, and 8 controls including an accept/reject pair on the same database length that differ only by table name length. <details><summary>Why making the check unconditional does not over-reject</summary> `computeMaxTableNameLength` returns `min(N - 13, max(0, N - 42 - esc_db))` where `N` is the filesystem name limit, `esc_db` the escaped database name length. Since `42 > 13`, the second term always binds, so ``` reject <=> esc_db + esc_tbl > N - 42 ``` The dropped-metadata filename built by `DatabaseCatalog::getPathForDroppedMetadata` is `{db}.{table}.{uuid}.sql`, i.e. `esc_db + 1 + esc_tbl + 1 + 36 + 4` bytes, so ``` does not fit <=> esc_db + esc_tbl + 42 > N <=> esc_db + esc_tbl > N - 42 ``` The same inequality. The check rejects exactly the pairs whose table would be undroppable, and no others. Measured at `N = 255`: | esc_db | limit | esc_tbl | verdict | drop segment | droppable? | |---|---|---|---|---|---| | 200 | 13 | 2 | accept | 244 | yes | | 200 | 13 | 20 | reject | 262 | no | | 211 | 2 | 2 | accept | 255 | yes | | 214 | 0 | 2 | reject | 258 | no | Rows 1 and 2 are the discriminating pair in the test: same database length, opposite verdicts, so no length-blind rule satisfies both. Row 3 is the exact-boundary control, and the test drops that table afterwards to prove an accepted rename really leaves it droppable. `max_to_drop` is unconditionally the smaller term, so splitting the limit per engine would return the same value and ship dead code. Warning instead of refusing was also considered and rejected: the rename would still create undroppable tables, and the three other call sites of the check all throw `ARGUMENT_OUT_OF_BOUND`. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113557",
          "createdAt": "2026-08-05T19:30:13Z",
          "updatedAt": "2026-08-13T02:12:34Z",
          "timestamp": "2026-08-13T02:12:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "diegomestre2",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:975a58604db124330ebe",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114520",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114520",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Require full field consumption on every Regexp escaping rule",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/108091 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the `Regexp` input format silently discarding the unparsed rest of a matched field, which stored a truncated or fabricated value instead of reporting an error: `v=-1` into `UInt64` read `0`, and `v=2020-01-01junk` into `Date` read `1975-07-14` under the `JSON` rule. Malformed fields are now rejected under the `Escaped`, `CSV` and `JSON` rules, matching the formats those rules are documented as behaving like. Under `Raw`, a matched field containing a tab is now read as a whole instead of being truncated at the tab. ### Description **The problem.** A matched capture group is one complete field, and `Regexp` has no delimiter after it, so when a type's parser stopped early the bytes it left behind were discarded and the partial parse reported as success. `v=-1` into `UInt64` gave `0`, `v=1abc` gave `1`, and `v=2020-01-01junk` into `Date` under the `JSON` rule gave `1975-07-14`, which is not even a prefix of the input. `TSV`, `TSKV`, `CSV`, `JSONEachRow` and `CustomSeparated` all reject these same bytes, because there the leftover fails the delimiter that follows. **Why the existing check was not enough.** #108091 added this check for the `Quoted` rule only; it now applies to every rule. `CSV` and `JSON` first skip trailing whitespace, because the formats they mirror accept it, so `v=1 ` still reads `1` while `v=1 abc` is rejected. **A second, more reachable mechanism.** `Raw` is the default rule, so it needs no setting at all, and it fails differently: its reader stops at a tab, so `v=abc<TAB>junk` into `String` read `abc`. A tab is legal inside a `String`, so rejecting that field would be a new bug; it is now read as a whole instead, while a tab that cannot belong to the value (`v=1<TAB>junk` into `UInt64`) is rejected. A field equal to the null representation keeps the old reader wherever a null-aware one is selected, so `Raw`'s null tokens are unchanged. The check stays inside `Regexp`: the two other formats sharing this deserialization helper have delimiters, so for them a leftover is normal. Validation: run unchanged against unpatched master the new test fails, nearly every error assertion reporting a missing error; the exceptions are the pre-existing `Quoted` and `Raw` controls. It passes 50/50 under randomized settings. Also noticed, not changed here: `Values` accepts `v=1 ` while `Regexp`+`Quoted` rejects it, a divergence predating this PR.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114520",
          "createdAt": "2026-08-12T16:38:11Z",
          "updatedAt": "2026-08-13T02:12:21Z",
          "timestamp": "2026-08-13T02:12:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d7921b803ba815e8485d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113140",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113140",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix cubic complexity of planning a JOIN with a `merge` table",
          "text": "Planning a query that joins a `Merge` table with something else could take minutes and hundreds of megabytes of memory, and it was interruptible neither by `max_execution_time` nor by `KILL QUERY`, because all of it happens inside `QueryPlan::optimize`. `ReadFromMerge::createChildrenPlans` builds a separate child query tree for every source table. When the `Merge` table takes part in a JOIN, this goes through `replaceTableExpressionAndRemoveJoin`, which rebuilds the projection list from the required column names. Every column there was resolved by its own `QueryAnalysisPass` run, and every such run rebuilt `AnalysisTableExpressionData` for all the columns of the `Merge` table. The structure of a `Merge` table is the union of the structures of its source tables, so the cost of planning was cubic in the number of source tables. All the required columns are now resolved with a single `QueryAnalysisPass` run. For 40 tables with 20 distinct columns each, a `SELECT *` joined with `merge` goes from 18 seconds down to about a second in a release build; the results are byte-identical before and after. This is what made the `Hung check` fail in a stress test: the AST fuzzer produced `SELECT * FROM (SELECT number AS a FROM numbers(11)) AS t1 PASTE JOIN merge('system', '') AS t2 LIMIT 294` for `02933_paste_join.sql`, and the query spent more than 16 minutes inside `QueryPlan::optimize` under TSan. https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=dd1e53dfe2715c22d4b7e63c2459bc9eec994581&name_0=MasterCI&name_1=Stress%20test%20%28arm_tsan%29 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed cubic complexity of query planning for a JOIN with a `Merge` table over many tables with different structures. Such queries could previously spend minutes in query planning without responding to `max_execution_time` or `KILL QUERY`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113140",
          "createdAt": "2026-08-03T15:52:16Z",
          "updatedAt": "2026-08-13T02:01:14Z",
          "timestamp": "2026-08-13T02:01:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9ee1924ddd27b6a4003c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114541",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114541",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add ProfileEvents and CurrentMetrics for fiber stacks",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/73510 Related: https://github.com/ClickHouse/ClickHouse/issues/72968 Fibers are used for asynchronous communication with remote replicas, and a stack is the only thing a fiber consists of, so allocating the stack is the whole cost of creating a fiber. Nothing observed it: neither the number of allocations, nor the amount of memory they hold, nor the time they take. When distributed queries became slow because every fiber went through `mmap`, `mprotect` and `munmap` (https://github.com/ClickHouse/ClickHouse/issues/72968), it had to be found with a profiler. `FiberStack::allocate` and `FiberStack::deallocate` now account for: - profile events `FiberStackAllocs` and `FiberStackAllocBytes`; - profile events `FiberStackAllocNanoseconds` and `FiberStackFreeNanoseconds` — in nanoseconds, because a single allocation normally takes less than a microsecond and would be truncated to zero; - metrics `FiberStacks` and `FiberStackBytes`. This is the observability part of https://github.com/ClickHouse/ClickHouse/pull/73510, which was closed because the allocator it proposed became obsolete after https://github.com/ClickHouse/ClickHouse/pull/79147. For a two-shard query, it looks like this: ``` SELECT count() FROM remote('127.0.0.{1,2}', system.one) SETTINGS async_socket_for_remote = 1, prefer_localhost_replica = 0 allocs: 6 bytes: 1966080 -- 6 stacks of 320 KiB alloc_ns: 26393 free_ns: 10700 ``` ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added observability for the stacks of fibers, which are used for asynchronous communication with remote replicas: profile events `FiberStackAllocs`, `FiberStackAllocBytes`, `FiberStackAllocNanoseconds`, `FiberStackFreeNanoseconds`, and metrics `FiberStacks`, `FiberStackBytes`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114541",
          "createdAt": "2026-08-12T20:47:15Z",
          "updatedAt": "2026-08-13T01:57:40Z",
          "timestamp": "2026-08-13T01:57:40Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:36b71199837eabcf0dc5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:97254",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:97254",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Push join key filters into MergeTree index during recursive CTE evaluation",
          "text": "When a recursive CTE joins against a `MergeTree` table, each recursion step previously scanned the entire table, because the join condition (e.g. `ON e.from_id = t.current_id`) was not pushed into `MergeTree`'s key condition. This change analyzes the recursive query's join tree to find equi-join conditions between the CTE working table and real tables. Before each step it reads the join-key values from the working table and injects them as an `IN (...)` predicate directly into the `WHERE` clause of each recursive `QueryNode` that joins against the CTE (the original `WHERE`/`HAVING`/`QUALIFY` clauses are snapshotted at construction and restored after each step, so nothing accumulates across steps). The planner then lowers this predicate into `ReadFromMergeTree`'s key condition. The injected set is bounded by the new setting `recursive_cte_max_in_filter_cardinality`, and injection fails closed to a plain scan whenever it could otherwise change results — when the generated set would exceed `max_rows_in_set`/`max_bytes_in_set`, cannot be materialized under `max_memory_usage`, or the `IN` predicate cannot be resolved for the join-key type. Parallel replicas are disabled for the recursive step queries to avoid stale cached `GLOBAL JOIN` tables returning incorrect results; the forcing mode `allow_experimental_parallel_reading_from_replicas = 2` raises `SUPPORT_IS_DISABLED` for the recursive part rather than silently downgrading. For a table with ~1M rows and 10 recursion steps, `read_rows` drops from ~10M to ~120. Closes: https://github.com/ClickHouse/ClickHouse/issues/75026 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `WITH RECURSIVE` queries that use `ON` or comma equi-joins against `MergeTree` tables can now use the primary key index, reducing the number of rows read per recursion step. The optimization is controlled by the new setting `recursive_cte_max_in_filter_cardinality`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) No separate documentation page is needed — this is a transparent performance improvement. The only user-visible addition is the setting `recursive_cte_max_in_filter_cardinality`, which is documented at the source level in `src/Core/Settings.cpp`. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches recursive CTE execution and query settings, and dynamically injects `additional_table_filters`, which could affect correctness/performance for some JOIN patterns despite being scoped to recursive steps. > > **Overview** > Improves recursive CTE performance by detecting equi-join conditions between the CTE working table and real tables, then **pushing join-key values into `additional_table_filters`** (built as `IN (...)` predicates) on each recursive step so MergeTree reads can use the primary-key index; user-provided `additional_table_filters` are merged rather than overwritten. > > Also **disables parallel replicas** for recursive CTE step queries to avoid incorrect results from reused cached GLOBAL JOIN tables, and adds a stateless test (`03924_recursive_cte_join_index`) asserting both correct output and low `read_rows` for explicit `INNER JOIN` and comma-join forms. > > <sup>Written by [Cursor Bugbot](https://cursor.com/dashboard?tab=bugbot) for commit 37f26180bfa2189247b6862712551c7ecc41ed18. This will update automatically on new commits. Configure [here](https://cursor.com/dashboard?tab=bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/97254",
          "createdAt": "2026-02-18T07:38:36Z",
          "updatedAt": "2026-08-13T01:56:02Z",
          "timestamp": "2026-08-13T01:56:02Z",
          "metrics": {
            "reactions": 1,
            "comments": 77
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7b3f9a669f02d1f2fa5c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113244",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113244",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix set index on an expression with a Nullable operand",
          "text": "A `SELECT` filtered on an expression that matches a `set` skip index raises `LOGICAL_ERROR`: ```sql CREATE TABLE events (id UInt32, INDEX b_set intDiv(id, 1000) TYPE set(100) GRANULARITY 1) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO events SELECT number FROM numbers(5000); SELECT count() FROM events WHERE intDiv(id, (SELECT 1000)) = 3; -- Code: 49. Unexpected return type from equals. Expected Nullable(UInt8). Got UInt8. ``` `SELECT 1000` is Nullable(UInt16), while the underlying column type of the index is plain UInt32, so when MergeTreeIndexConditionSet rewrites the predicate it loses the Nullable type and just saves that it is a UInt. However the equals function still expects a Nullable, and this mismatch throws an error in `ExpressionActions.cpp:executeAction`. ============================================ No `Nullable` appears anywhere in the schema or the query. A scalar subquery is typed `Nullable` because it may return no rows, and that is enough. In practice this shows up when the divisor or bucket size comes from a lookup or settings table instead of a literal; the typing is identical. Explicit wrappers reach the same failure: `nullIf(1000, 0)`, `toNullable(1000)`, `materialize(toNullable(1000))`. `MergeTreeIndexConditionSet` matched a query subexpression to an index key column by name only. Names are computed from constant-folded arguments and never include types, so `intDiv(id, (SELECT 1000))` renders exactly like the index expression `intDiv(id, 1000)` while carrying `Nullable(UInt32)`. The whole subtree was then replaced by the granule column while the enclosing `equals` was re-added with its already-resolved `IFunctionBase`, keeping the return type it was resolved with. `ExpressionActions::execute` binds inputs by name without checking types, so the granule supplied a non-`Nullable` column and the declared return type no longer matched what execution produced. The rebuilt condition is now typed from the granule side throughout. `atomFromDAG` binds the key column input to the type the granule block holds instead of the query-side type, and re-resolves the atom's function against the granule-side argument types instead of reusing the query-side `IFunctionBase`. The index therefore keeps pruning on these queries rather than being skipped: `EXPLAIN indexes = 1` reports the same 5/50 granules for `t % nullIf(19, 0) = 16` as for the plain `t % 19 = 16`. Two cases still fall back to `UNKNOWN_FIELD`, which leaves the granule unpruned and sends the query through the regular filter path. A function that cannot be re-resolved from `FunctionFactory` by name — internal casts, lambdas, parametric functions — and argument types the function rejects. And a key column that is already an `INPUT` of the filter DAG, which cannot be re-typed, because a second input under the same name would be left unbound where `ExpressionActions::execute` maps each name to one block column. The regression test covers both granule filtering paths, since `secondary_indices_enable_bulk_filtering` selects between `getPossibleGranules` and `mayBeTrueOnGranule`, and each reaches the same actions independently. An index whose expression is itself `Nullable` is unaffected in either direction, because the types agree and nothing is re-typed. Reproduces on the official release `26.8.1.727`. In builds with assertions enabled it aborts the server, which is how it surfaced in CI — 14 occurrences in 90 days across unrelated pull requests and every build configuration, through `tests/queries/0_stateless/01786_explain_merge_tree.sh` once the AST fuzzer wraps the modulus in a `Nullable` expression. Closes: https://github.com/ClickHouse/ClickHouse/issues/113234 Related: https://github.com/ClickHouse/ClickHouse/issues/113233 Related: https://github.com/ClickHouse/ClickHouse/issues/89802 Related: https://github.com/ClickHouse/ClickHouse/pull/111830 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `LOGICAL_ERROR` when reading a table with a `set` skip index defined on an expression and the query repeats that expression with a `Nullable` operand, for example `WHERE intDiv(id, (SELECT 1000)) = 3` over `INDEX b_set intDiv(id, 1000) TYPE set(100)`. A scalar subquery is enough to trigger it, since it is typed `Nullable`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113244",
          "createdAt": "2026-08-04T07:24:53Z",
          "updatedAt": "2026-08-13T01:52:52Z",
          "timestamp": "2026-08-13T01:52:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "george-larionov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c4e05588d5c10fcbfa8f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109896",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109896",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use_constant_folding_in_index_analysis issues",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109893 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix three issues with `use_constant_folding_in_index_analysis`: a logical error `Invalid partition key size` that could abort a `SELECT` using a normal projection when the table sets `part_minmax_index_columns = 'with_block_number_offset'`; wrong results (silently dropped rows) for a filter on a modulo partition key such as `PARTITION BY id % 200`; and a crash (data race) when a `text` index with a `sparseGrams` tokenizer is queried with a `LIKE` predicate over many partitions with `max_threads > 1`. ### Description Three issues in the `use_constant_folding_in_index_analysis` path. **1. Logical error `Invalid partition key size`** (fuzzer STID 2677-496b, [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=106734&sha=6220f1e884aa1a725214970d897d217a75ba8c6a&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20parallel%29)). In `selectPartsToRead`, `minmax_idx_condition` comes from a normal projection's metadata, whose partition key is empty, but the loop passed the *parent* part's partition (size 1) to per-partition specialization, so `MergeTreePartition::getID` aborted against the size-0 key. Fix: projection parts are not specialized; they get the unsubstituted condition. **2. Wrong results for modulo partition keys** ([repro](https://fiddle.clickhouse.com/3a7a3f59-b256-420f-908f-09459e46aa8d)). The stored value uses the backward-compatible `moduloLegacy` rewrite (8-bit: `moduloLegacy(-199, 200) = 57`) while the filter evaluates modern `modulo` (16-bit: `-199`). Matching the modern predicate against the modern key substituted the stored legacy value, turning `id % 200 < 0` into `57 < 0` and pruning parts that do hold matching rows (498 rows became 356). Fix: match against the legacy-adjusted key, so a modern `modulo` node no longer matches. Non-modulo keys fold unchanged. **3. Crash with a stateful `sparseGrams` tokenizer** ([repro](https://fiddle.clickhouse.com/c8e3cc24-f366-4413-835e-4bb7afc0d0d8)). With folding on, the skip-index condition is rebuilt per partition inside a thread pool, and every condition got the index's single tokenizer as a raw pointer. `sparseGrams` advances a mutable iterator, so concurrent builds corrupted its state (SIGSEGV in `SparseGramsTokenizer::nextInStringLike`). Fix: add `ITokenizer::isStateful` and give each condition a private clone; stateless tokenizers stay shared. Aggregators clone too, for symmetry with `MergeTreeIndexAggregatorText`, but that is defence in depth rather than a fixed race: each part-writer owns its aggregator, so writes never share one tokenizer concurrently. Tests in `04510_projection_partition_minmax_key_size`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109896",
          "createdAt": "2026-07-09T14:01:22Z",
          "updatedAt": "2026-08-13T01:48:45Z",
          "timestamp": "2026-08-13T01:48:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 28
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:483dcb9fa1d72e58bcd9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114521",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114521",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix Iceberg query failure after MODIFY COLUMN to Nullable",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/85029 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `LOGICAL_ERROR` when filtering an Iceberg column that `ALTER TABLE ... MODIFY COLUMN` made `Nullable`. Closes #85029. ### Description Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/85029#issuecomment-5264517370; his analysis comment there carries the reproducer. ```sql CREATE TABLE t (id Int64, s String) ENGINE = IcebergLocal('lake/'); INSERT INTO t SELECT number, toString(number) FROM numbers(5); ALTER TABLE t MODIFY COLUMN id Nullable(Int64); SELECT * FROM t WHERE id > 3; -- Unexpected return type from greater. Expected Nullable(UInt8). Got UInt8 ``` Root cause: in `IcebergSchemaProcessor::getSchemaTransformationDag` the node emitted for an existing field id is chosen from the Iceberg `type` string alone. Nullability lives in the separate `required` key, which `getFieldType` folds into ClickHouse nullability. `MODIFY COLUMN ... Nullable` flips only `required`, leaving type string and name unchanged, so neither the cast nor the alias branch fired and the transform forwarded the input node built from the old flag. Its output header said `Int64` while consumers were told `Nullable(Int64)`. Only PREWHERE observes this: the fallback `FilterTransform` is built against that header, so `greater` yields `UInt8` where the planner expects `Nullable(UInt8)`. Without the pushdown the filter sits above the reader, past the cast in `getColumnFromBlock`, so `SELECT *` and `count()` were correct. The fix emits the missing cast, gated twice: inside the equal-type comparison, leaving the `allowPrimitiveTypeConversion` allow-list and spec-prohibited conversions untouched; and only for the legal relaxation of required to optional, so an externally written reverse pair keeps its current passthrough instead of gaining a `Nullable(T) -> T` cast that would reject rows holding NULL. Both excluded directions measure byte-identical. Validated on unpatched and patched binaries for `Int64`, `Int32`, `String`, `Float64`, `Date32` and `DateTime64`, with a rename composed on top and mixed pre/post-`ALTER` files. The new test is 50/50 green; the Iceberg and Delta suites show no failure absent unpatched.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114521",
          "createdAt": "2026-08-12T16:38:44Z",
          "updatedAt": "2026-08-13T01:42:40Z",
          "timestamp": "2026-08-13T01:42:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1e0c7415993dc0cb3d38",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112875",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112875",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Azure: retry transient authentication failures on the remaining object-storage call sites",
          "text": "> **Series:** **#112871** → #112875 (this), #112872, #112873, #112874, #112876 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Retry a transient Azure `AuthenticationException` (credential/token-acquisition failure, e.g. during the managed-identity/RBAC propagation window) on object-storage call sites that previously had no retry at all: blob existence/metadata/size probes, single and batch deletes, the post-upload verification, the container existence check, and ADLS Gen2 writes. Previously a single transient credential failure at any of these points failed the whole query or background job. --- Follow-up to #112871, which handles the HTTP half of the problem: the Azure SDK `RetryPolicy` now retries a transient 403 for every client (`retry_options.StatusCodes.insert(Forbidden)` in `AzureBlobStorageCommon.cpp`). What the SDK cannot retry is `Azure::Core::Credentials::AuthenticationException`: it is thrown by the credential/token layer *around* the transport, so the retry policy — which only sees HTTP responses and `TransportException` — never observes it. The read/download/write buffer loops already absorb it via their `catch (...)` retry, but these call sites had no retry of any kind: - `AzureObjectStorage::exists` and `getObjectMetadata` / `getObjectMetadataIfExists` — one-shot `GetProperties()` - `AzureObjectStorage::removeObjectImpl` (blob and ADLS) and the batch `SubmitBatch` - `ReadBufferFromAzureBlobStorage::tryGetFileSize` / `getRemoteFileMetadata` - `WriteBufferFromAzureBlobStorage::finalizeImpl` post-upload verification — the upload already succeeded, a transient credential failure must not turn a good write into a failure - `containerExists` in the client factory — the first request many flows issue - `WriteBufferFromAzureDataLakeStorage::runWithRetries` — retried HTTP errors but let `AuthenticationException` escape on the first attempt A single transient credential failure at any of them surfaced as a query/job failure. Changes: - New `retryAzureOnAuthError()` (`src/IO/AzureBlobStorage/retryAzureOnAuthError.h`): bounded exponential backoff around `AuthenticationException` **only**. HTTP-status retries deliberately stay in the SDK retry policy (#112871), so retry layers don't stack multiplicatively. `getBlobPropertiesWithRetry()` lives in the same header for the `GetProperties()` sites. - Route the call sites above through it; widen their `Azure::Storage::StorageException` catches to the base `Azure::Core::RequestFailedException` (a behavior-preserving superset; NotFound handling unchanged). - ADLS Gen2 `runWithRetries` additionally retries `AuthenticationException` within the same budget (shared give-up/backoff tail for both failure kinds). - `readBigAt`: guard against a null body stream in the download response (mirrors the existing guard in `initialize()`). - Tests (`test_azure_403_handling`): a one-shot injected `AuthenticationException` on the direct metadata path (no ClickHouse-level retry loop anywhere) must be absorbed — this test fails without this PR; a permanent one must still fail with the real auth error after the bounded budget. Compared to the previous revision of this PR: all 403/HTTP retrying has been removed from the helper. After #112871 that is the SDK retry policy's job; doing it here as well would stack a second retry loop on top of the SDK's (up to 3 × 11 attempts on a permanent 403). This PR is now strictly about the auth exception the SDK cannot see. The per-object batch re-issue loop is gone for the same reason (an RBAC 403/auth failure is per-principal, i.e. all-or-nothing across a batch; batch-level retry covers it). Not covered by tests: the ADLS Gen2 write path (azurite cannot emulate ADLS/OneLake endpoints). Stacked on #112871 (branch base); to be merged after it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112875",
          "createdAt": "2026-08-01T06:53:55Z",
          "updatedAt": "2026-08-13T01:40:18Z",
          "timestamp": "2026-08-13T01:40:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "arsenmuk",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:171744994f092000ded0",
        "signalId": "github:ClickHouse/ClickHouse:issue:110933",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:110933",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "convertMySQLDataType wraps Array/Variant-based geo types in Nullable, throwing ILLEGAL_TYPE_OF_ARGUMENT on nullable MySQL spatial columns",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ ### Describe what's wrong Reading a MySQL table that has a NULLABLE spatial column (`LINESTRING`/`POLYGON`/`MULTILINESTRING`/`MULTIPOLYGON`/generic `GEOMETRY`) via the `mysql()` table function, the `MySQL` table engine, or the `MySQL` database engine throws `Nested type LineString cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT)` at schema-inference time, so `CREATE`/`DESCRIBE`/`SELECT` fail entirely. Nullable is the MySQL default for a spatial column, and the `geometry` mapping is enabled by default, so this triggers out of the box. **Root cause:** src/DataTypes/convertMySQLDataType.cpp:171-172 applies `std::make_shared<DataTypeNullable>(res)` to `res` regardless of whether `res` can be inside Nullable. Before this PR the non-Point spatial types produced `String` (canBeInsideNullable=true), so `Nullable(String)` was valid; the new geo types (`Array`/`Variant` based) cannot be inside `Nullable`. **Why we believe this is a bug:** fetchTablesColumnsList (src/Databases/MySQL/FetchTablesColumnsList.cpp:116-119) calls convertMySQLDataType with is_nullable = external_table_functions_use_nulls (default true) && IS_NULLABLE='YES' -> convertMySQLDataType maps the spatial type to a geo type (e.g. `LineString` = `Array(Point)`) at src/DataTypes/convertMySQLDataType.cpp:143 -> then unconditionally wraps it at line 172 `res = std::make_shared<DataTypeNullable>(res)` -> DataTypeNullable ctor throws because DataTypeArray/DataTypeVariant::canBeInsideNullable() == false. **Affected locations:** - `src/DataTypes/convertMySQLDataType.cpp:172` — unconditional Nullable wrap of a geo-typed res - `src/DataTypes/convertMySQLDataType.cpp:143` — linestring -> LineString (Array(Point)) branch; same for polygon/multilinestring/multipolygon/geometry at 147/151/155/163 - `src/Databases/MySQL/FetchTablesColumnsList.cpp:119` — passes is_nullable=true for nullable MySQL columns when external_table_functions_use_nulls (default true) **Impact:** By default (geometry mapping on), any MySQL integration over a table with a nullable spatial column becomes unusable - schema inference throws and the table cannot be created/described/read. This is a regression: those columns previously mapped to `Nullable(String)` and worked. Affects the `mysql()` table function, `MySQL` table engine, and `MySQL` database engine. ## Assumptions The bot recorded these claims it could not verify directly from source. A maintainer ✅ confirms; ❌ flags a wrong premise (the bot should rework or close the finding). - [ ] **A plain MySQL spatial column (declared without NOT NULL) is reported as IS_NULLABLE='YES' in information_schema.columns** - *Why unverifiable:* the test harness could not connect to the MySQL container in this sandbox - *Falsifiable test:* CREATE TABLE t (g LINESTRING) in MySQL, then read information_schema.columns.IS_NULLABLE for column g (expected 'YES'); then DESCRIBE mysql(...) from ClickHouse (expected exception). ### Does it reproduce on most recent release? Likely yes — see testability note in additional context. ### How to reproduce ```python Run: pytest tests/integration/test_storage_mysql/test.py -k test_mysql_nullable_geometry -v -s (with a reachable mysql80 container). OR reproduce the invariant directly on any server: SELECT CAST(NULL AS Nullable(LineString)); -- throws ILLEGAL_TYPE_OF_ARGUMENT, whereas SELECT CAST(NULL AS Nullable(String)); succeeds. ``` ### Expected behavior ``` DESCRIBE returns rows including 'ls\\tLineString' and 'geo\\tGeometry' (schema inference succeeds). ``` ### Error message and/or stacktrace ``` Could not run end-to-end (MySQL container unreachable in sandbox). Type-invariant confirmed via SQL: `CREATE TABLE t (x Nullable(LineString)) ENGINE=Memory` -> Code: 43. DB::Exception: Nested type LineString cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT); same for Polygon/MultiLineString/MultiPolygon/Geometry. ``` ### Additional context **Open risks:** - The umbrella `Geometry` (Variant) also cannot be inside Nullable - same failure for a nullable generic GEOMETRY column. - The MySQL database engine (DatabaseMySQL) shares fetchTablesColumnsList and is affected identically. **Suggested fix:** When the mapped type cannot be inside Nullable (e.g. `!res->canBeInsideNullable()`), either skip the Nullable wrap for geo types (geo columns are inherently non-Nullable in ClickHouse) or fall back to `Nullable(String)` for nullable spatial columns. **Analysis details:** Confidence HIGH | Severity P1 | Testability: `INTEGRATION_TEST` Found during automated review of [PR #108944](https://github.com/ClickHouse/ClickHouse/pull/108944). ### CI Proof Bug confirmed by `Integration tests (amd_llvm_coverage, 2/8)` on [proof PR #110924](https://github.com/ClickHouse/ClickHouse/pull/110924). <details><summary>Reproducer test</summary> ```sql def test_mysql_nullable_geometry(started_cluster): # A nullable MySQL spatial column (MySQL's default when NOT NULL is omitted) must not # break schema inference. With the `geometry` mapping on by default, convertMySQLDataType # maps `linestring` to `LineString` (Array(Point)) and then wraps it in Nullable, which # throws ILLEGAL_TYPE_OF_ARGUMENT because Array/Variant-based geo types cannot be inside Nullable. table_name = \"test_mysql_nullable_geometry\" node1.query(f\"DROP TABLE IF EXISTS {table_name}\") conn = get_mysql_conn(started_cluster, cluster.mysql8_ip) drop_mysql_table(conn, table_name) with conn.cursor() as cursor: cursor.execute( f\"\"\" CREATE TABLE `clickhouse`.`{table_name}` ( `id` int NOT NULL, `ls` linestring, `geo` geometry, PRIMARY KEY (`id`)) ENGINE=InnoDB; \"\"\" ) conn.commit() # Expected (correct behavior): DESCRIBE succeeds and reports the geo types. # Actual (bug): raises 'Nested type LineString cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT)'. result = node1.query( f\"DESCRIBE mysql('mysql80:3306', 'clickhouse', '{table_name}', 'root', '{mysql_pass}')\" ) assert \"LineString\" in result, result drop_mysql_table(conn, table_name) conn.close() ``` </details> <details><summary>CI log excerpt</summary> ``` 2026-07-18T10:37:18.7126776Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_TESTS_SERVER_BIN_PATH to /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/clickhouse 2026-07-18T10:37:18.7127449Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_BINARY to /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/clickhouse 2026-07-18T10:37:18.7128117Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_TESTS_CLIENT_BIN_PATH to /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/clickhouse 2026-07-18T10:37:18.7128634Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_USE_OLD_ANALYZER to 0 2026-07-18T10:37:18.7129004Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_USE_DISTRIBUTED_PLAN to 0 2026-07-18T10:37:18.7129369Z [2026-07-18 10:37:18] Setting environment variable CLICKHOUSE_USE_DATABASE_DISK to 0 2026-07-18T10:37:18.7129977Z [2026-07-18 10:37:18] Setting environment variable PYTEST_CLEANUP_CONTAINERS to 1 2026-07-18T10:37:18.7130398Z [2026-07-18 10:37:18] Setting environment variable JAVA_PATH to /usr/lib/jvm/java-11-openjdk-amd64/bin/java 2026-07-18T10:37:18.7130811Z [2026-07-18 10:37:18] Setting environment variable COMPLIANCE_RESULT_FILE to 2026-07-18T10:37:18.7131220Z /home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/promql_compliance_result.json 2026-07-18T10:37:18.7131649Z [2026-07-18 10:37:18] Setting environment variable LLVM_PROFILE_FILE to it-%4m.profraw 2026-07-18T10:37:18.7132392Z [2026-07-18 10:37:18] Run command: [pytest test_storage_mysql/test.py::test_mysql_nullable_geometry --report-log-exclude-logs-on-passed-tests --tb=short -n 1 2026-07-18T10:37:18.7133099Z --dist=loadfile --session-timeout=1200 --report-log=/home/ubuntu/actions-runner/_work/ClickHouse/ClickHouse/ci/tmp/pytest_retries.jsonl 2026-07-18T10:37:18.7133681Z --log-file=/home/ubuntu/actions-runner/_work/ClickHouse/Cli ``` </details> --- _ClickGapAI · Confidence: HIGH · Severity: P1 · Finding: `h_pr108944_001`_ <!-- ch-version-info:start --> ### Version info - Resolved by: #110943 - Merged into: `26.8.1.1302` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/110933",
          "createdAt": "2026-07-18T11:04:41Z",
          "updatedAt": "2026-08-13T01:36:30Z",
          "timestamp": "2026-08-13T01:36:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "comp-mysql"
          ],
          "author": "clickgapai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9262ebdd9da3b336dace",
        "signalId": "github:ClickHouse/ClickHouse:issue:107334",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:107334",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "DeltaLake Logical error: 'No version found in table state snapshot'",
          "text": "### Describe the bug Seems not related to https://github.com/ClickHouse/ClickHouse/issues/102037 Reading a persistent DeltaLake (or Iceberg) table through a code path that does **not** call `updateExternalDynamicMetadataIfExists()` reaches `IDataLakeMetadata::iterate()` with a storage snapshot whose `datalake_table_state` is empty. `iterate()` treats this as impossible and throws `LOGICAL_ERROR`, which aborts the server: ``` <Fatal> : Logical error: 'No version found in table state snapshot'. DeltaLakeMetadataDeltaKernel::iterate (DeltaLakeMetadataDeltaKernel.cpp:331) -> DataLakeConfiguration<…, DeltaLakeMetadata>::iterate (DataLakeConfiguration.h:294) -> StorageObjectStorageSource::createFileIterator (StorageObjectStorageSource.cpp:287) -> ReadFromObjectStorageStep::createIterator / initializePipeline ``` Found by BuzzHouse on master `26.6.1.694` via `SELECT * FROM d2.t284` on a `DeltaLakeS3` table that had been through many ALTER/mutation/cluster-function operations. ## Root cause A persistent data-lake table's in-memory metadata gets its snapshot version pinned (`setDataLakeTableState`) only inside `updateExternalDynamicMetadataIfExists()` (`StorageObjectStorage.cpp:397`). The analyzer calls it on the table it resolves (`IdentifierResolver.cpp:335`), so a *direct* SELECT is fine. But several read paths build the child storage snapshot straight from `getInMemoryMetadataPtr()` **without** that call: - `StorageMerge` — `StorageMerge.cpp:730` (and `:414`) - cluster/distributed table-function reads on the executing node - (the materialized-view target path used to have this bug; it was fixed by adding the update call — `StorageMaterializedView.cpp:400`) When the table is reached first through one of these, the snapshot carries no `datalake_table_state`, and `iterate()` throws instead of falling back. The throw is also asymmetric with its sibling in the *same file*: `DeltaLakeMetadataDeltaKernel::prepareReadingFromFormat()` (`DeltaLakeMetadataDeltaKernel.cpp:505-511`) handles the identical missing-version case **gracefully** by falling back to the version from settings. `iterate()` should do the same. Iceberg is identical: `IcebergMetadata::iterate()` (`IcebergMetadata.cpp:1023`) throws `\"Can't extract iceberg table state from storage snapshot\"` under the same condition. ## Relationship to prior findings Same bug *class* as the Iceberg over-assertions filed in #107316 — a `LOGICAL_ERROR` guarding an invariant that is actually reachable from normal SQL. Here the invariant is \"the storage snapshot always carries a pinned table-state version\", which is false for read paths that skip `updateExternalDynamicMetadataIfExists`. ### How to reproduce With clickhouse binary in the PATH and the server running, run this script: ```bash set -u BIN=clickhouse # 26.6.1.694, has delta-kernel TABLE=contrib/delta-kernel-rs/kernel/tests/data/table-without-dv-small P=\"/var/lib/clickhouse/user_files\" rm -rf \"$P\" && mkdir -p \"$P\" cp -r \"$TABLE\" \"$P/delta1\" ABS=\"$P/delta1/\" echo \"=== control: direct SELECT pins state, works ===\" \"$BIN\" --client --multiquery \" SET allow_experimental_delta_kernel_rs=1; CREATE TABLE ok ENGINE = DeltaLakeLocal('$ABS'); SELECT count() FROM ok; \" echo \"=== repro: merge() is the SOLE first access -> aborts ===\" \"$BIN\" --client --multiquery \" SET allow_experimental_delta_kernel_rs=1; CREATE TABLE t ENGINE = DeltaLakeLocal('$ABS'); SELECT * FROM merge(currentDatabase(), '^t\\$') ORDER BY 1 LIMIT 5; \" echo \"exit=$? (134 = SIGABRT = reproduced)\" ``` ### Error message and/or stacktrace Stack trace: ``` <Fatal> : Logical error: 'No version found in table state snapshot'. <Fatal> : Format string: 'No version found in table state snapshot'. <Fatal> : Stack trace (when copying this message, always include the lines below): 0. /ClickHouse/contrib/llvm-project/libcxx/include/__exception/exception.h:115:14: Poco::Exception::Exception(String&&, int) @ 0x0000000027021373 1. /ClickHouse/src/Common/Exception.cpp:139:7: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x0000000014bdab29 2. /ClickHouse/src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000d189cd6 3. /ClickHouse/src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000d189798 4. /ClickHouse/src/Common/Exception.h:189:77: DB::Exception::Exception<>(int, FormatStringHelperImpl<>) @ 0x0000000014bd9279 5. /ClickHouse/src/Storages/ObjectStorage/DataLakes/DeltaLakeMetadataDeltaKernel.cpp:331:15: DB::DeltaLakeMetadataDeltaKernel::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) const @ 0x0000000018ab9f3d 6. /ClickHouse/src/Storages/ObjectStorage/DataLakes/DataLakeConfiguration.h:294:34: DB::DataLakeConfiguration<DB::StorageLocalConfiguration, DB::DeltaLakeMetadata>::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) @ 0x0000000016eaa901 7. /ClickHouse/src/Storages/ObjectStorage/StorageObjectStorageSource.cpp:287:36: DB::StorageObjectStorageSource::createFileIterator(std::shared_ptr<DB::StorageObjectStorageConfiguration>, DB::StorageObjectStorageQuerySettings const&, std::shared_ptr<DB::IObjectStorage>, std::shared_ptr<DB::StorageInMemoryMetadata const>, bool, std::shared_ptr<DB::Context const> const&, DB::ActionsDAG::Node const*, DB::ActionsDAG const*, DB::NamesAndTypesList const&, DB::NamesAndTypesList const&, std::vector<std::shared_ptr<DB::ObjectInfo>, std::allocator<std::shared_ptr<DB::ObjectInfo>>>*, std::function<void (DB::FileProgress)>, bool, bool, bool) @ 0x00000000189951a0 8. /ClickHouse/src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:170:24: DB::ReadFromObjectStorageStep::createIterator() @ 0x0000000020405658 9. /ClickHouse/src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:98:5: DB::ReadFromObjectStorageStep::initializePipeline(DB::QueryPipelineBuilder&, DB::BuildQueryPipelineSettings const&) @ 0x00000000204048f1 10. /ClickHouse/src/Processors/QueryPlan/ISourceStep.cpp:20:5: DB::ISourceStep::updatePipeline(std::vector<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>, AllocatorWithMemoryTracking<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>>>, DB::BuildQueryPipelineSettings const&) @ 0x00000000202cf75e 11. /ClickHouse/src/Processors/QueryPlan/QueryPlan.cpp:214:47: DB::QueryPlan::buildQueryPipeline(DB::QueryPlanOptimizationSettings const&, DB::BuildQueryPipelineSettings const&, bool) @ 0x0000000020342ec9 12. /ClickHouse/src/Storages/StorageMerge.cpp:1223:31: DB::ReadFromMerge::buildPipeline(DB::ReadFromMerge::ChildPlan&, DB::QueryProcessingStage::Enum) const @ 0x000000001eed7f9b 13. /ClickHouse/src/Storages/StorageMerge.cpp:582:32: DB::ReadFromMerge::initializePipeline(DB::QueryPipelineBuilder&, DB::BuildQueryPipelineSettings const&) @ 0x000000001eed755d 14. /ClickHouse/src/Processors/QueryPlan/ISourceStep.cpp:20:5: DB::ISourceStep::updatePipeline(std::vector<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>, AllocatorWithMemoryTracking<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>>>, DB::BuildQueryPipelineSettings const&) @ 0x00000000202cf75e 15. /ClickHouse/src/Processors/QueryPlan/QueryPlan.cpp:214:47: DB::QueryPlan::buildQueryPipeline(DB::QueryPlanOptimizationSettings const&, DB::BuildQueryPipelineSettings const&, bool) @ 0x0000000020342ec9 16. /ClickHouse/src/Interpreters/InterpreterSelectQueryAnalyzer.cpp:399:34: DB::InterpreterSelectQueryAnalyzer::buildQueryPipeline() @ 0x000000001a693cb0 17. /ClickHouse/src/Interpreters/InterpreterSelectQueryAnalyzer.cpp:364:29: DB::InterpreterSelectQueryAnalyzer::execute() @ 0x000000001a69389e 18. /ClickHouse/src/Interpreters/executeQuery.cpp:1824:40: DB::executeQueryImpl(char const*, char const*, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum, std::unique_ptr<DB::ReadBuffer, std::default_delete<DB::ReadBuffer>>&, boost::intrusive_ptr<DB::IAST>&, std::shared_ptr<DB::ImplicitTransactionControlExecutor>, std::function<void ()>, DB::QueryResultDetails&) @ 0x000000001aa21fff 19. /ClickHouse/src/Interpreters/executeQuery.cpp:2196:11: DB::executeQuery(std::basic_string_view<char, std::char_traits<char>>, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum) @ 0x000000001aa1ad8a 20. /ClickHouse/src/Server/TCPHandler.cpp:832:68: DB::TCPHandler::runImpl() @ 0x000000001fbcacb5 21. /ClickHouse/src/Server/TCPHandler.cpp:3075:9: DB::TCPHandler::run() @ 0x000000001fbe84e4 22. /ClickHouse/base/poco/Net/src/TCPServerConnection.cpp:40:3: Poco::Net::TCPServerConnection::start() @ 0x000000002708af8e 23. /ClickHouse/base/poco/Net/src/TCPServerDispatcher.cpp:115:42: Poco::Net::TCPServerDispatcher::run() @ 0x000000002708b4d2 24. /ClickHouse/base/poco/Foundation/src/ThreadPool.cpp:207:14: Poco::PooledThread::run() @ 0x000000002705443f 25. /ClickHouse/base/poco/Foundation/src/Thread_POSIX.cpp:341:27: Poco::ThreadImpl::runnableEntry(void*) @ 0x000000002705280f 26. start_thread @ 0x000000000009caa4 27. __GI___clone3 @ 0x0000000000129c6 ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #102033 - Merged into: `26.7.1.448` (included in `26.7` and later) - Backported to: `26.5.7.46` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/107334",
          "createdAt": "2026-06-12T13:57:01Z",
          "updatedAt": "2026-08-13T01:36:06Z",
          "timestamp": "2026-08-13T01:36:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "bug",
            "fuzz",
            "comp-datalake"
          ],
          "author": "PedroTadim",
          "state": "closed",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2e04f7a18e2795ce45eb",
        "signalId": "github:ClickHouse/ClickHouse:issue:102037",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:102037",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Logical error: Can't extract iceberg table state from storage snapshot for table location A (STID: 2606-4a47)",
          "text": "_Important: This issue was automatically generated and is used by CI for matching failures. DO NOT modify the body content. DO NOT remove labels._ Test name: Logical error: Can't extract iceberg table state from storage snapshot for table location A (STID: 2606-4a47) CI report: [Stress test (amd_msan)](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=96682&sha=bc37844d817ee88bc66bf40870fb4d9c5373d488&name_0=PR&name_1=Stress%20test%20%28amd_msan%29) Failing test history: [cidb](https://play.clickhouse.com/play?user=play&run=1#V0lUSAogICAgOTAgQVMgaW50ZXJ2YWxfZGF5cwpTRUxFQ1QKICAgIHRvU3RhcnRPZkRheShjaGVja19zdGFydF90aW1lKSBBUyBkYXksCiAgICBjb3VudCgpIEFTIGZhaWx1cmVzLAogICAgZ3JvdXBVbmlxQXJyYXkocHVsbF9yZXF1ZXN0X251bWJlcikgQVMgcHJzLAogICAgYW55KHJlcG9ydF91cmwpIEFTIHJlcG9ydF91cmwKRlJPTSBjaGVja3MKV0hFUkUgKG5vdygpIC0gdG9JbnRlcnZhbERheShpbnRlcnZhbF9kYXlzKSkgPD0gY2hlY2tfc3RhcnRfdGltZQogICAgQU5EIHRlc3RfbmFtZSA9ICdMb2dpY2FsIGVycm9yOiBDYW4nJ3QgZXh0cmFjdCBpY2ViZXJnIHRhYmxlIHN0YXRlIGZyb20gc3RvcmFnZSBzbmFwc2hvdCBmb3IgdGFibGUgbG9jYXRpb24gQSAoU1RJRDogMjYwNi00YTQ3KScKICAgIC0tIEFORCBjaGVja19uYW1lID0gJ1N0cmVzcyB0ZXN0IChhbWRfbXNhbiknCiAgICBBTkQgdGVzdF9zdGF0dXMgSU4gKCdGQUlMJywgJ0VSUk9SJykKICAgIEFORCAocHVsbF9yZXF1ZXN0X251bWJlciA9IDAgT1IgYmFzZV9yZWYgSU4gKCdtYXN0ZXInKSkKICAgIEFORCB0ZXN0X2NvbnRleHRfcmF3IExJS0UgJyVFeGNlcHRpb246JScKR1JPVVAgQlkgZGF5Ck9SREVSIEJZIGRheSBERVNDCg==) Test output: ``` Log files: clickhouse-server.err.log, stderr.log Error: Logical error: 'Can't extract iceberg table state from storage snapshot for table location /var/lib/clickhouse/user_files/t_test_pxpyjno7_28122/'. --- Stack trace: __pthread_kill @ 0x00000000000969fd gsignal @ 0x0000000000042476 __lgamma_r_finite@GLIBC_2.15 @ 0x00000000000287f3 src/Common/Exception.cpp:60:5: DB::abortOnFailedAssertion(String const&, std::basic_string_view<char, std::char_traits<char>>, void* const*, unsigned long, unsigned long) @ 0x000000002bc3f621 src/Common/Exception.cpp:93:13: DB::handle_error_code(String const&, std::basic_string_view<char, std::char_traits<char>>, int, bool, std::vector<void*, std::allocator<void*>> const&) @ 0x000000002bc4178a src/Common/Exception.cpp:146:19: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x000000002bc4211e ./src/Common/Exception.h:171:100: DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000b75d1ae ./src/Common/Exception.h:57:54: DB::Exception::Exception(PreformattedMessage&&, int) @ 0x000000000b75bb0d ./src/Common/Exception.h:189:77: DB::Exception::Exception<String const&>(int, FormatStringHelperImpl<std::type_identity<String const&>::type>, String const&) @ 0x000000000b798f4a src/Storages/ObjectStorage/DataLakes/Iceberg/IcebergMetadata.cpp:1037:15: DB::IcebergMetadata::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) const @ 0x000000003d64f43e ./src/Storages/ObjectStorage/DataLakes/DataLakeConfiguration.h:276:34: DB::DataLakeConfiguration<DB::StorageLocalConfiguration, DB::IcebergMetadata>::iterate(DB::ActionsDAG const*, std::function<void (DB::FileProgress)>, unsigned long, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DB::Context const>) @ 0x0000000034ed9d1c src/Storages/ObjectStorage/StorageObjectStorageSource.cpp:259:36: DB::StorageObjectStorageSource::createFileIterator(std::shared_ptr<DB::StorageObjectStorageConfiguration>, DB::StorageObjectStorageQuerySettings const&, std::shared_ptr<DB::IObjectStorage>, std::shared_ptr<DB::StorageInMemoryMetadata const>, bool, std::shared_ptr<DB::Context const> const&, DB::ActionsDAG::Node const*, DB::ActionsDAG const*, DB::NamesAndTypesList const&, DB::NamesAndTypesList const&, std::vector<std::shared_ptr<DB::ObjectInfo>, std::allocator<std::shared_ptr<DB::ObjectInfo>>>*, std::function<void (DB::FileProgress)>, bool, bool, bool) @ 0x000000003d25bab6 src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:156:24: DB::ReadFromObjectStorageStep::createIterator() @ 0x0000000056ed3981 src/Processors/QueryPlan/ReadFromObjectStorageStep.cpp:85:5: DB::ReadFromObjectStorageStep::initializePipeline(DB::QueryPipelineBuilder&, DB::BuildQueryPipelineSettings const&) @ 0x0000000056ed0902 src/Processors/QueryPlan/ISourceStep.cpp:20:5: DB::ISourceStep::updatePipeline(std::vector<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>, std::allocator<std::unique_ptr<DB::QueryPipelineBuilder, std::default_delete<DB::QueryPipelineBuilder>>>>, DB::BuildQueryPipelineSettings const&) @ 0x0000000056ac3ac0 src/Processors/QueryPlan/QueryPlan.cpp:211:47: DB::QueryPlan::buildQueryPipeline(DB::QueryPlanOptimizationSettings const&, DB::BuildQueryPipelineSettings const&, bool) @ 0x0000000056bf17ae src/Interpreters/MutationsInterpreter.cpp:1776:37: DB::MutationsInterpreter::addStreamsForLaterStages(std::vector<DB::MutationsInterpreter::Stage, std::allocator<DB::MutationsInterpreter::Stage>> const&, DB::QueryPlan&) const @ 0x0000000044b55e3a src/Interpreters/MutationsInterpreter.cpp:1831:20: DB::MutationsInterpreter::execute() @ 0x0000000044b62cde src/Storages/ObjectStorage/DataLakes/Iceberg/Mutations.cpp:166:72: DB::Iceberg::writeDataFiles(DB::MutationCommands const&, std::shared_ptr<DB::Context const>, std::shared_ptr<DB::StorageInMemoryMetadata const>, DB::StorageID, std::shared_ptr<DB::IObjectStorage>, String, DB::FileNamesGenerator&, DB::Iceberg::IcebergPathResolver const&, std::optional<DB::FormatSettings> const&, std::optional<DB::ChunkPartitioner>&, Poco::SharedPtr<Poco::JSON::Object, Poco::ReferenceCounter, Poco::ReleasePolicy<Poco::JSON::Object>>) @ 0x000000003d83743d src/Storages/ObjectStorage/DataLakes/Iceberg/Mutations.cpp:611:31: DB::Iceberg::mutate(DB::MutationCommands const&, std::shared_ptr<DB::Context const>, std::shared_ptr<DB::StorageInMemoryMetadata const>, DB::StorageID, std::shared_ptr<DB::IObjectStorage>, DB::DataLakeStorageSettings const&, DB::Iceberg::PersistentTableComponents const&, String const&, std::optional<DB::FormatSettings> const&, std::shared_ptr<DataLake::ICatalog>) @ 0x000000003d8309e6 src/Storages/ObjectStorage/DataLakes/Iceberg/IcebergMetadata.cpp:586:5: DB::IcebergMetadata::mutate(DB::MutationCommands const&, std::shared_ptr<DB::StorageObjectStorageConfiguration>, std::shared_ptr<DB::Context const>, DB::StorageID const&, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DataLake::ICatalog>, std::optional<DB::FormatSettings> const&) @ 0x000000003d63b34e ./src/Storages/ObjectStorage/DataLakes/DataLakeConfiguration.h:154:27: DB::DataLakeConfiguration<DB::StorageLocalConfiguration, DB::IcebergMetadata>::mutate(DB::MutationCommands const&, std::shared_ptr<DB::Context const>, DB::StorageID const&, std::shared_ptr<DB::StorageInMemoryMetadata const>, std::shared_ptr<DataLake::ICatalog>, std::optional<DB::FormatSettings> const&) @ 0x0000000034ede9b6 src/Storages/ObjectStorage/StorageObjectStorage.cpp:781:20: DB::StorageObjectStorage::mutate(DB::MutationCommands const&, std::shared_ptr<DB::Context const>) @ 0x000000003d15f44a src/Interpreters/InterpreterAlterQuery.cpp:303: DB::(anonymous namespace)::runCommandSegments(std::vector<std::variant<DB::AlterCommands, DB::MutationCommands, std::vector<DB::PartitionCommand, std::allocator<DB::PartitionCommand>>, std::vector<DB::ASTAlterCommand const*, std::allocator<DB::ASTAlterCommand const*>>>, std::allocator<std::variant<DB::AlterCommands, DB::MutationCommands, std::vector<DB::PartitionCommand, std::allocator<DB::PartitionCommand>>, std::vector<DB::ASTAlterCommand const*, std::allocator<DB::ASTAlterCommand const*>>>>>&, std::shared_ptr<DB::IStorage> const&, std::shared_ptr<DB::Context const> const&) src/Interpreters/InterpreterAlterQuery.cpp:447:24: DB::InterpreterAlterQuery::executeToTable(DB::ASTAlterQuery const&) @ 0x000000004471b19c src/Interpreters/InterpreterAlterQuery.cpp:345:16: DB::InterpreterAlterQuery::execute() @ 0x0000000044707b25 src/Interpreters/executeQuery.cpp:1815:40: DB::executeQueryImpl(char const*, char const*, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum, std::unique_ptr<DB::ReadBuffer, std::default_delete<DB::ReadBuffer>>&, boost::intrusive_ptr<DB::IAST>&, std::shared_ptr<DB::ImplicitTransactionControlExecutor>, std::function<void ()>, DB::QueryResultDetails&) @ 0x00000000456817e5 src/Interpreters/executeQuery.cpp:2203:11: DB::executeQuery(String const&, std::shared_ptr<DB::Context>, DB::QueryFlags, DB::QueryProcessingStage::Enum) @ 0x000000004566eb09 src/Server/TCPHandler.cpp:823:68: DB::TCPHandler::runImpl() @ 0x00000000551b54b8 src/Server/TCPHandler.cpp:2974:9: DB::TCPHandler::run() @ 0x000000005522e104 base/poco/Net/src/TCPServerConnection.cpp:40:3: Poco::Net::TCPServerConnection::start() @ 0x0000000069704c60 base/poco/Net/src/TCPServerDispatcher.cpp:115:42: Poco::Net::TCPServerDispatcher::run() @ 0x0000000069705c9a base/poco/Foundation/src/ThreadPool.cpp:205:14: Poco::PooledThread::run() @ 0x00000000695a459a ./base/poco/Foundation/src/Thread_POSIX.cpp:341:27: Poco::ThreadImpl::runnableEntry(void*) @ 0x000000006959d811 start_thread @ 0x0000000000094ac3 __clone3 @ 0x00000000001268d0 ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #102033 - Merged into: `26.7.1.448` (included in `26.7` and later) - Backported to: `26.5.7.46` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/102037",
          "createdAt": "2026-04-08T10:00:31Z",
          "updatedAt": "2026-08-13T01:36:04Z",
          "timestamp": "2026-08-13T01:36:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "testing",
            "fuzz"
          ],
          "author": "KochetovNicolai",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3124a7f31cebeb7a7303",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:102033",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:102033",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix LOGICAL_ERROR crash in IcebergMetadata::iterate when datalake_table_state is missing",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fix `LOGICAL_ERROR` exceptions when reading `Iceberg` or `DeltaLake` data lake tables through paths that can reach the read pipeline without a pinned `datalake_table_state`, such as concurrent `Iceberg` metadata updates or `merge` reads over `DeltaLake` tables. ### What this PR does **Problem:** `IcebergMetadata::iterate` and `isDataSortedBySortingKey` throw a `LOGICAL_ERROR` (`Can't extract iceberg table state from storage snapshot`) when `datalake_table_state` is not set in the storage metadata snapshot. The same family of failures hits `DeltaLake` (`No version found in table state snapshot`). This raises an exception (and, in stress tests, a server abort under `abort_on_logical_error`) — tracked as STID `2606-4a47`, a chronic failure with master and PR hits across many sanitizer/build variants over the last 30 days. `PR #101577` fixed one such path (materialized views accessing `Iceberg` target tables), but the failure continued on master because other code paths can reach the read pipeline without `datalake_table_state`: - Concurrent metadata updates create a TOCTOU race between `updateExternalDynamicMetadataIfExists` setting the state (via `setInMemoryMetadata`) and the next `getInMemoryMetadataPtr` reading it. - Cluster table functions and other paths that bypass the analyzer/interpreter reach `read`/`iterate` without having had the state pinned. - `merge` reads over `DeltaLake` tables reach the read step with a snapshot that lacks the pinned state. **Fix:** Pin a single, internally consistent data lake snapshot at the source instead of papering over the missing state deep inside the callees. This follows the reviewer's guidance to *fix all the places to create the table state snapshot instead of ad-hoc fallbacks in the callees*: - `StorageObjectStorage::read` pins one coherent `datalake_table_state` **before** `prepareReadingFromFormat`, so the requested columns, the field-id mapping used to list/read files, and the read-in-order sorting key all come from the same data lake snapshot. For cluster table functions it first calls `lazyInitializeIfNeeded`, otherwise `update`, so `getTableStateSnapshot` does not hit an uninitialized-metadata assertion. When `shouldReloadSchemaForConsistency` is set, columns and sorting key are rebuilt from the same state via `buildStorageMetadataFromState` so they cannot diverge. - `StorageObjectStorageSource::createFileIterator` repopulates `datalake_table_state` from `getTableStateSnapshot` before `iterate`, as a safety net for the remaining callers. - `ReadFromObjectStorageStep::getDataOrder` reads the sorting key from the pinned `storage_snapshot->metadata` rather than re-deriving it. The `LOGICAL_ERROR` assertions inside `iterate` and `isDataSortedBySortingKey` are preserved untouched as safety nets — they should now be unreachable on the fixed paths. A test-only failpoint `datalake_simulate_missing_table_state` strips the pinned state from the snapshot right before the read step, so the regression test reproduces the missing-state condition deterministically. ### Tests - `04305_iceberg_missing_table_state.sh` — deterministic reproducer using the `datalake_simulate_missing_table_state` failpoint; a non-trivial read (`sum`) and an `ORDER BY` query exercise both the `iterate` and `isDataSortedBySortingKey` sites. - `04306_delta_lake_merge_missing_table_state.sh` — `merge` read over a `DeltaLake` table, the `DeltaLake` variant of the same bug. - `04141_iceberg_concurrent_no_logical_error.sh` — concurrent readers and a writer producing fresh snapshots, exercising the TOCTOU race in the regular planner path. Closes: https://github.com/ClickHouse/ClickHouse/issues/102037 Closes: https://github.com/ClickHouse/ClickHouse/issues/107334 Related: https://github.com/ClickHouse/ClickHouse/issues/93278 <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.1.448` (included in `26.7` and later) - Backported to: `26.5.7.46` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/102033",
          "createdAt": "2026-04-08T09:17:23Z",
          "updatedAt": "2026-08-13T01:36:03Z",
          "timestamp": "2026-08-13T01:36:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 36
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8f5f41e99551e48f2d9d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101512",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101512",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `date_time_overflow_behavior` being ignored for integer and float casts to DateTime64/Time64",
          "text": "Fixes `date_time_overflow_behavior` being silently ignored when casting integer and `Float32`/`Float64` values to `DateTime64` and `Time64`. Previously, overflowing values could skip the `VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE` exception in `throw` mode and skip clamping to the target boundary in `saturate` mode. The fix routes native and wide integer sources (`Int8`–`Int256`, `UInt8`–`UInt256`) as well as `Float32`/`Float64` through overflow-aware, scale-aware transforms, and scales `Float32` inputs in the `Float64` domain so the result is no longer distorted by source-float precision. `Decimal*` and `BFloat16` sources are out of scope here — their scale-aware overflow handling belongs in `DataTypesDecimal.cpp`, not in `FunctionsConversion.h` — and are left for a follow-up. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `date_time_overflow_behavior` for integer and `Float32`/`Float64` casts to `DateTime64` and `Time64`: overflowing values now throw `VALUE_IS_OUT_OF_RANGE_OF_DATA_TYPE` in `throw` mode and clamp to the target boundary in `saturate` mode; `Float32` inputs are scaled in `Float64` precision, correcting previously imprecise results. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101512",
          "createdAt": "2026-04-01T14:42:05Z",
          "updatedAt": "2026-08-13T01:35:01Z",
          "timestamp": "2026-08-13T01:35:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 25
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yariks5s",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6f615451ec3a6cf7ba7c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113045",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113045",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Perform the `Too many parts` check once per INSERT query",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109000 `MergeTreeSink::onStart` and `ReplicatedMergeTreeSink::onStart` perform the `parts_to_throw_insert` / `parts_to_delay_insert` check, and `onStart` is the only place where an `INSERT` may be rejected with `TOO_MANY_PARTS` - rejecting a query that has already written a part is not acceptable, as the comment right above the call says. Since #109000 a plain `INSERT` writes through up to `max_insert_threads` sinks running in parallel, and every one of them runs `onStart` (`ExceptionKeepingTransform::prepare` returns `Ready` for the `Start` stage of every sink, whether or not it ever receives a chunk). The `onStart` calls of the different sinks are not ordered with respect to each other's writes: a sink that the executor schedules late runs its check *after* another sink of the same query has already committed a part, counts that part, and rejects the query in the middle of it. The `INSERT` fails with `TOO_MANY_PARTS` even though the table was below the threshold when the query started - and even when the table was empty and the only part counted is the one the query wrote itself. Note that #109000 exposed the flaw rather than introduced it. The parallel sink fan-out itself (`sink_stream_size = max_insert_threads` in `InsertDependenciesBuilder`) has been in place for the `INSERT SELECT` path since `031a9d5a528`, and the default of `max_insert_threads` changed from `1` to `0` (auto, the number of CPU cores) in 26.8 - so an `INSERT SELECT` into a MergeTree table could already be rejected by a stream that counted a part the query had written itself. Before #109000 the plain-`INSERT` pipeline passed a hardcoded `/*max_insert_threads*/ 1` to the builder and always had a single sink, which is why the flaky tests below - all of them plain `INSERT ... VALUES` - only started failing when that changed. The check is now shared by all the sinks of one query: the first sink to reach `onStart` runs it and the others wait until it is finished, so it runs exactly once and strictly before any of the sinks writes anything. This restores the behaviour that the single-sink pipeline had. Only the rejection is shared: the `parts_to_delay_insert` backpressure is applied by every sink before each of its blocks, including the first one, so a parallel `INSERT` keeps the same per-block throttling as a single-sink one (a sink that receives no data does not delay). The gates live in a small per-query registry (`InsertStartGates`) keyed by the physical destination table id. `InsertDependenciesBuilder` creates the registry and takes each sink's gate from it next to `setRuntimeData` / `setHasDependentMaterializedViews`, so the sinks of all the streams that write into the same table share one gate - including the branches of different materialized views converging on the same target table. An `Alias` destination forwards the write through a nested `INSERT` per parallel branch and the real check runs inside those nested inserts, so `AliasSink` receives the registry and threads it into the nested `InterpreterInsertQuery`, whose builder inherits it instead of creating its own. `Distributed` and `Buffer` destinations are always kept single-stream by #109000, so their own pre-write checks are unaffected - but their in-query nested `INSERT`s into the underlying tables are not. The *local* writes of a `Distributed` destination run through nested `INSERT`s into the underlying table: one per incoming block on the `writeToLocal` path (the direct local write of a background-send insert with `prefer_localhost_replica`) and one per local replica job on the foreground path. Each of them used to create its own gates, so the second block of a multi-block `INSERT` counted the part the first block had already committed on the local target and failed the query mid-way; `DistributedSink` also receives the registry and threads it into both nested `INSERT` paths (the new test `04826_insert_distributed_local_too_many_parts_gate` covers a two-block insert into a local target with `parts_to_throw_insert = 1` on both paths). The `parallel_distributed_insert_select` rewrite of an `INSERT ... SELECT` between two `Distributed` tables writes the local shards the same way - one nested `InterpreterInsertQuery` per destination shard this server belongs to - and a server can be local for several destination shards of one cluster, so those sibling nested `INSERT`s write into the same underlying table and now share one registry as well (the new test `04847_insert_parallel_distributed_insert_select_local_shards` covers a two-local-shard cluster with `parts_to_throw_insert = 1` on the target, plus the non-parallel quorum variant of the same topology). The direct writes of a `Buffer` destination run through nested `INSERT`s into its destination table as well: one per block that bypasses the buffer by exceeding the max thresholds, and one per flush the buffer runs by threshold from within the query - so a later direct write of a multi-block `INSERT` used to count the part an earlier one had committed on the destination; `BufferSink` also receives the registry and threads it through `flushBuffer` / `writeBlockToDestination` into those nested `INSERT`s, while the background flush keeps creating fresh gates per flush as before (the new test `04827_insert_buffer_direct_write_too_many_parts_gate` covers both paths with `parts_to_throw_insert = 1` on the destination). A `TimeSeries` destination forwards the write through nested `INSERT`s into its inner target tables, so `TimeSeriesSink` also receives the registry and threads it into them, sharing the check across sibling branches converging on the same `TimeSeries` table. Those nested pipelines are now created in `onStart` rather than in the sink's constructor: the registry reaches the sink only after `StorageTimeSeries::write` has returned it, so building them in the constructor handed them an empty registry and the sharing held only by the accident of every sink of the query being constructed before any of them writes anything. A `WindowView` forwards every chunk of every branch into a fresh sink of its inner table, so `PushingToWindowViewSink` also receives the registry and threads it into `StorageWindowView::writeIntoWindowView`, which sets the query's gate of the inner table on the sinks it creates - sharing the check across the branches and across the successive chunks of one branch. This path cannot be covered by a test yet: a window view currently cannot receive inserted data at all - the view dependency of the source table is never registered and a direct `INSERT` into a window view fails with `std::out_of_range` - both since before this change (they reproduce on 26.7), see https://github.com/ClickHouse/ClickHouse/issues/113493. A related single-in-flight-part contract exists one level up: a non-parallel quorum insert (`insert_quorum >= 2` or `'auto'`, with `insert_quorum_parallel = 0`) permits a single in-flight quorum part per table, which is why its `max_insert_threads` fan-out is kept single-stream. But the single sink stream is still duplicated into one branch per dependent materialized view, and with `parallel_view_processing = 1` those branches ran concurrently - two views converging on one `ReplicatedMergeTree` target raced two in-flight quorum parts of one `INSERT` against each other, so a sibling branch could fail with `UNSATISFIED_QUORUM_FOR_PREVIOUS_WRITE` or conflict on the `/quorum/status` node. The serialization such an insert needs is derived from the collected sink graph, not from the settings alone (`InsertDependenciesBuilder::computeQuorumStreamRequirements`): only writes reaching a `ReplicatedMergeTree` table are quorum writes, and only two of them racing on the same table conflict. The `INSERT SELECT` fan-out - which had no quorum guard at all - is kept single-stream only when a branch of the write may produce a quorum part: a reachable `ReplicatedMergeTree` target, or a forwarding storage (an `Alias`, a `Distributed`, a `Buffer`, a `WindowView`, a `TimeSeries`) that hides its physical destination, in which case the probe fails closed. An insert whose write graph never reaches a replicated table keeps its `max_insert_threads` fan-out even under a global quorum profile. Only the branches reachable in the executable graph are counted: a view branch pruned from the graph - e.g. a view whose dropped target table is ignored by `ignore_materialized_views_with_dropped_target_table` - never creates a sink, so it does not force the serialization (the new test `04825_insert_quorum_ignored_broken_view_fanout` covers a broken ignored view attached to a plain `MergeTree` destination). The dependent views of such an insert are pushed sequentially only when two branches converge on the same replicated table (a sequential branch blocks in `commitPart` until the quorum of its part is satisfied, so the next branch starts with the quorum node already gone) or when a hidden write target makes such a convergence impossible to rule out - branches writing to distinct replicated tables keep running concurrently, since they do not share a `/quorum/status` node. An `Alias` destination hides its target's views behind the nested `INSERT` its sink runs, and that nested `INSERT` observes the same settings and derives the same requirements. The new test `04828_insert_timeseries_too_many_parts_gate` exercises non-parallel quorum inserts through two views converging on one `TimeSeries` table whose data table is replicated - the shape that raced `UNSATISFIED_QUORUM_FOR_PREVIOUS_WRITE` before `TimeSeries` was included in the hidden-target probes. The new test `04817_insert_quorum_sequential_dependent_views` covers the `INSERT SELECT` fan-out deterministically via `EXPLAIN PIPELINE` and exercises quorum inserts through two materialized views converging on one replicated target with `parallel_view_processing = 1`; the new test `04823_insert_quorum_graph_derived_serialization` covers the fan-out kept for a plain `MergeTree` destination, the fan-out dropped when a dependent materialized view targets a replicated table, and quorum inserts through two views writing to two distinct replicated tables. `parallel_view_processing = 0` gets the same treatment on the `INSERT SELECT` path: `buildInsertPipeline` already kept a plain `INSERT` single-stream when the destination is a forwarding storage that hides its target's dependent-view graph behind the nested `INSERT` its sinks run (`serial_hidden_views`), but `addInsertToSelectPipeline` still passed the raw `max_insert_threads` to the builder - which sees no views for such a destination - so an `INSERT SELECT` into an `Alias` fanned out into several `AliasSink`s whose hidden materialized views ran concurrently across sibling branches despite the setting. The guard is now mirrored there; an `Alias` in front of a target with dependent views is the only topology whose fan-out survived to that point, since `Distributed`, `Buffer` and `TimeSeries` destinations are already collapsed to one stream by the builder's `supportsParallelInsert` check. The new test `04869_insert_select_alias_hidden_views_serial` covers the single stream with a hidden dependent view, and the fan-out kept both without dependent views and with `parallel_view_processing = 1`. This is the cause of a family of flaky tests that started failing on `master` on 2026-08-01, right after #109000 was merged, each of them rejecting an `INSERT` that counts the part it has written itself: | test | failing query | reported parts | | --- | --- | --- | | `02458_relax_too_many_parts` | `INSERT INTO test VALUES (6, 'a')` | 3 with `parts_to_throw_insert = 3` | | `02280_add_query_level_settings` | `INSERT INTO table_for_alter VALUES (1, '1')` into an empty table | 1 with `parts_to_throw_insert = 1` | | `00980_merge_alter_settings` | `INSERT INTO table_for_alter VALUES (1, '1')` into an empty table | 1 with `parts_to_throw_insert = 1` | | `02015_async_inserts_5` | `INSERT INTO async_inserts VALUES` | 1 with `parts_to_throw_insert = 1` | The server log of the `02458_relax_too_many_parts` failure below shows the mechanism directly: 26 sinks are created for the query, its part `all_5_5_0` is renamed into place at `01:46:59.792095`, and the query is rejected with `Too many parts (3 ...)` at `01:46:59.798713` - six milliseconds after it committed the part it is being counted for. Report of that failure (`MergeQueueCI`, `Fast test`): https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=gh-readonly-queue/master/pr-112777-508dbed54a148ce6ae2dabe8ff91f4c7842bb1a0&sha=c71ec1f59e4b52ffe87c96e12743488cec72a7e0&name_0=MergeQueueCI&name_1=Fast%20test Corresponding pull request: https://github.com/ClickHouse/ClickHouse/pull/112777 The new test `04692_insert_too_many_parts_check_once_per_query` covers both halves: an `INSERT` with a large `max_insert_threads` that brings a table from one part up to `parts_to_throw_insert = 2` has to be accepted, and `ProfileEvents['DelayedInserts']` of a single-row `INSERT` has to match the number of blocks the query actually writes (one; two when the query writes through two materialized views converging on the same target table) rather than the number of streams it writes through - also when it writes through an `Alias`. `TimeSeries` is also classified as a forwarding storage in the deduplication-safety probes (`storageDeduplicatesBlocksOnInsert`, `storageRebuildsDeduplicationIdsOnInsert`, `forwardedInsertReachesDependentView`, `forwardedInsertHidesDependentView`, `forwardedInsertHidesDependentViewForwardingToSeparateContext`), since `TimeSeriesSink` forwards a plain block into its nested `INSERT`s and drops the outer chunk's `DeduplicationInfo`. This is fail-close classification only, not a reachable fix: `StorageTimeSeries` does not override `IStorage::supportsParallelInsert`, so an insert whose graph reaches a `TimeSeries` table is already single-stream and the probes cannot change the outcome - and writes arriving through `TimeSeriesSink` turn out not to deduplicate on the inner tables at all, so no rows are lost when two views converge on one `TimeSeries` table. The clauses keep the classification correct if `StorageTimeSeries` ever starts supporting parallel inserts. The new test `04846_insert_mv_timeseries_dedup_parallel_views` pins the no-row-loss half end to end: two identical materialized views converging on one `TimeSeries` table with a deduplicating data table (`non_replicated_deduplication_window`) under `parallel_view_processing = 1` land both branches in full - nothing is deduplicated between the sibling branches of one query. #### Relation to #114016 While this pull request was open, #114016 landed on `master` with a narrower fix for the same root cause: it evaluates the check on the query thread in the `MergeTreeSink` / `ReplicatedMergeTreeSink` constructor and rethrows it from `onStart`. Merging `master` resolves that in favour of the shared gate, which subsumes it - a sink whose `onStart` runs late skips the check instead of re-running it - and additionally covers the sinks that are created *during* execution, which the constructor placement cannot: the nested `INSERT`s of `Alias`, `Distributed`, `Buffer`, `WindowView` and `TimeSeries`. Everything else from #114016 is kept, including the `merge_tree_sink_on_start_random_sleep` failpoint and its test `04826_parallel_insert_sinks_too_many_parts_self_race`, which the shared gate also makes pass. Related: https://github.com/ClickHouse/ClickHouse/pull/114016 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a spurious `TOO_MANY_PARTS` error for an `INSERT` executed with `max_insert_threads` greater than one. The `Too many parts` check was performed by every parallel writing stream, so a stream that started after another one had written a part counted that part and rejected the query, even when the table was below the threshold - or empty - when the query started.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113045",
          "createdAt": "2026-08-03T01:08:11Z",
          "updatedAt": "2026-08-13T01:30:59Z",
          "timestamp": "2026-08-13T01:30:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:190e6979fcf88386fc9a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113207",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113207",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add the `logsql` dialect: LogsQL, the query language of VictoriaLogs",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Added an experimental `logsql` dialect: LogsQL, the log query language of VictoriaLogs, can now be used to query ClickHouse tables with logs. Enable it with `SET allow_experimental_logsql_dialect = 1, logsql_table = '<table>', dialect = 'logsql'` (plus `logsql_database`, `logsql_time_column`, and `logsql_message_column` as needed). LogsQL queries are translated at parse time into ordinary `SELECT` queries, so text filters expect `String`-backed log columns, while the `math` pipe applies raw arithmetic and needs numeric operand columns (`String` field values are not coerced to numbers). Documentation entry for user-facing changes. - [ ] Documentation is written (mandatory for new features) Implementation notes: - Nearly the whole language is supported: the filters (words, phrases, prefixes, `field:value`, `:=`, `~`, comparisons, `range`, `in` with subqueries, `contains_any`/`contains_all` with subqueries, `seq`, `len_range`, `string_range`, `ipv4_range`, `ipv6_range`, `pattern_match*`, `contains_common_case`/`equals_common_case`, `json_array_contains_any`, `i()`, full `_time:` filters with `day_range`/`week_range`, `_stream:{...}` selectors), and the pipes `fields`, `delete`, `copy`, `rename`, `limit`, `offset`, `sort`/`first`/`last` (including `rank` and `partition by`), `stats` (with `by`-buckets, per-function `if`, `switch`, `rate`), `where`, `uniq`, `top`, `math`, `extract`, `extract_regexp`, `format`, `unpack_json`, `unpack_logfmt`, `join`, `union`, `unroll`, `running_stats`, `total_stats`, `len`, `hash`, `coalesce`, `decolorize`, `split`, `unpack_words`, `time_add`, `sample`, `generate_sequence`, `field_values`, `json_array_len`, `json_array_concat`, `replace`, `replace_regexp`, `pack_json`, `pack_logfmt`. On typed (non-`String`) columns, filters that emit string functions fail with a type error instead of applying VictoriaLogs' schemaless per-value matching. The numeric stats functions (`sum`, `avg`, `median`, `quantile`, `stddev`, `rate_sum`) parse the numeric value out of every field value and skip the values that are not numbers, like VictoriaLogs, which computes its stats in `float64`; a numeric column is therefore aggregated with `Float64` precision and not exactly. - The dialect is modeled on the existing `kusto`/`prql`/`promql`/`polyglot` dialects: a parser under `src/Parsers/LogsQL/` builds a ClickHouse `ASTSelectQuery` directly, so all downstream machinery (analyzer, distributed queries, query log) works as usual. Relational pipes map to real SQL constructs: `join` to `LEFT`/`INNER JOIN ... USING`, `union` to `UNION ALL`, `running_stats`/`total_stats` and the `rank`/`partition by` clauses to window functions. - The lexer and the grammar mirror the reference implementation in VictoriaLogs (`lib/logstorage` of `VictoriaMetrics/VictoriaLogs`, Apache 2.0), including Go-style string literals, compound tokens (`foo-bar.com:123/x`), comments with `#`, and the exact filter/pipe syntax. The test queries are reused from the VictoriaLogs parser tests: 675 of its 878 valid parser-test queries run end-to-end against a ClickHouse table, and 605 of its 612 invalid queries are rejected. - The implicit table is configured by the `logsql_database`/`logsql_table` settings (like `promql_database`/`promql_table`). The special fields `_time` and `_msg` are mapped to columns via `logsql_time_column`/`logsql_message_column`, so real tables (e.g. `system.text_log` with `event_time`/`message`) can be queried after setting those. - Word and phrase filters translate to `hasToken`/`match` with token boundaries, so tables with `tokenbf_v1` indexes on the message column benefit from index analysis. - Documented deviations from VictoriaLogs: word boundaries follow ClickHouse tokenization (ASCII alphanumerics; VictoriaLogs also treats `_` and non-ASCII letters as word characters); numeric comparison filters parse `String` field values per row (rows with non-numeric values do not match) and compare typed numeric columns exactly, but the per-row parsing of field values understands plain numeric text only (decimal integers, floats, scientific notation, `inf`/`nan`) - LogsQL-only spellings such as `10KiB`, `1h`, `0x10`, or `1_000` are supported in query literals but are treated as non-numeric when stored in a field value, whereas VictoriaLogs parses them there too; equality with a non-numeric value compares with the native column type instead of VictoriaLogs' string semantics; `count(field)` counts non-NULL values instead of non-empty strings; `pattern_match*` placeholders are approximated with regular expressions (VictoriaLogs uses a greedy matcher with extra boundary rules); `seq()` matches substrings in order without word-boundary checks. - The only constructs that remain `NOT_IMPLEMENTED` (with clear errors naming the construct) are the ones with no ClickHouse meaning: VictoriaLogs storage introspection (`block_stats`, `blocks_count`, `query_stats`, `value_type`, `histogram` buckets, `set_stream_fields`, `stream_context`), features that require the dynamic set of fields of a schemaless store (`facets`, `field_names`, wildcard field selectors like `foo*:filter`, stats over all fields like `sum(*)`, `unpack_json` without a `fields` list), `unpack_syslog` (a stateful wall-clock-dependent parser), and `collapse_nums` (its number-boundary rules require lookaround which RE2 lacks).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113207",
          "createdAt": "2026-08-04T00:00:57Z",
          "updatedAt": "2026-08-13T01:27:34Z",
          "timestamp": "2026-08-13T01:27:34Z",
          "metrics": {
            "reactions": 5,
            "comments": 5
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a5fc478bef7832310e13",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:79509",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:79509",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Async insert parallel parsing",
          "text": "Parses the data of one batch of asynchronous inserts with several threads when the batch is flushed, controlled by the new setting `async_insert_parse_threads` (`0` by default, which keeps the previous behaviour). Addresses https://github.com/ClickHouse/ClickHouse/issues/74162 ### Motivation With `wait_for_async_insert = 1` the client waits for the whole batch to be flushed, and for text formats most of that time is spent parsing the accumulated data in a single thread. That makes the observed `INSERT` latency proportional to the size of the batch, which is what makes `wait_for_async_insert = 1` unusable for some workloads. ### How it works The entries of a batch are split into `async_insert_parse_threads` contiguous ranges. Every range is parsed by its own `StreamingFormatExecutor` with its own input format (the format holds mutable parsing state and cannot be shared), on a thread pool bounded server-wide by the new `max_async_insert_parsing_thread_pool_size` (100 by default). The calling thread parses one of the ranges itself instead of only waiting. This is a pool of its own rather than the format parsing pool, because some input formats (`Parquet`, `ArrowStream`, ...) parallelize their own work on the format parsing pool: a range waiting for such a nested task while holding a thread of that very pool would deadlock once every thread of it is held by a range. Afterwards the ranges are concatenated back into a **single** chunk with a **single** `DeduplicationInfo`, and the per-entry bookkeeping (the `system.asynchronous_insert_log` elements, the deduplication tokens, the rows and bytes reported back to the waiting clients) is replayed by the flushing thread in the original order of the entries. So only parsing is parallelized: the flush still pushes one block into the insert pipeline, the number of parts written per flush does not change, and no shared state is written from more than one thread. The tasks run through `ThreadPoolCallbackRunnerLocal`, which attaches them to the thread group of the flush, so profile events, memory and CPU accounting of the parsing are attributed to the flush query instead of being lost. The number of ranges is clamped to the number of entries in the batch and to the size of the pool, so a batch of 3 inserts never uses more than 3 threads. `0` and `1` mean the flushing thread parses everything and the pool is not touched at all. ### Measurements 2000 asynchronous inserts of 500 `JSONEachRow` rows each, collected into one batch of one million rows and flushed explicitly; the numbers are the `AsyncInsertFlush` entry of `system.query_log`, two runs per value (aarch64, 96 cores, shared machine): | `async_insert_parse_threads` | flush `query_duration_ms` | rows per ms | `UserTimeMicroseconds` | | --- | --- | --- | --- | | `0` | 738, 680 | 1355, 1471 | 634 ms, 627 ms | | `2` | 457, 600 | 2188, 1667 | 629 ms, 643 ms | | `5` | 331, 323 | 3021, 3096 | 628 ms, 638 ms | | `10` | 293, 272 | 3413, 3677 | 632 ms, 625 ms | So the flush of a one-million-row batch goes from ~709 ms to ~327 ms with 5 threads (2.2x) and to ~283 ms with 10 (2.5x), and `UserTimeMicroseconds` stays at ~630 ms throughout - the CPU time of the parsing is accounted to the flush query no matter how many threads did it, and parallelizing it adds no measurable CPU overhead. An earlier measurement of the first version of this change, on a different data set, showed the same effect on latency (~2763 ms against ~663 ms for 5 threads): https://pastila.nl/?007f67ed/3ce4613dcb3ed6f2f8cfc00d6e8bc906#en58FLI4Ilg9Nk2Vp6b8XQ== - but there `UserTimeMicroseconds` collapsed from 2.3 s to 15 ms, because that version did not attach the parsing tasks to the thread group of the flush, so their profile events were lost. That is what the flat `UserTimeMicroseconds` column above verifies is fixed. ### Testing `tests/queries/0_stateless/04869_async_insert_parse_threads.sh` flushes explicitly built batches with 0, 1, 4 and 16 parse threads and checks that all the rows are inserted with their column defaults, that one flush still produces exactly one part, and that a row which fails to parse fails only its own insert (`ParsingError` in `system.asynchronous_insert_log`) while the other entries of the batch are inserted. `async_insert_parse_threads` is randomized in `tests/clickhouse-test`, so the `Stateless tests (AsyncInsert)` job exercises the parallel path across the whole suite. Locally, the 83 stateless tests whose name contains `async_insert` were run twice against the same binary, once with `async_insert_parse_threads = 4` forced for every query and once with `0`, and the two failure sets are identical - 12 failures in both, all of them gaps in the minimal local server config used for the run (no second shard on `127.0.0.2`, no `pandas` for the `.python` tests, and `Settings['async_insert']` not being recorded because the CI `users.d` default is absent). With those gaps filled, `02481_async_insert_dedup` and `02481_async_insert_dedup_token` - the coverage that matters for the rebuilt `DeduplicationInfo` - both pass with 4 parse threads. `02187_async_inserts_all_formats` (tagged `long`) exceeds the 600 s harness timeout on that machine with 4 threads and with 0 alike, so that one is the machine, not the setting; it sends one insert per batch anyway, which clamps to a single range and the plain serial path. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added the setting `async_insert_parse_threads` that specifies how many threads parse the data of one batch of asynchronous inserts when it is flushed. It reduces the latency of `INSERT`s that wait for the flush (`wait_for_async_insert = 1`). Disabled by default. ### Documentation entry for user-facing changes",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/79509",
          "createdAt": "2025-04-23T21:31:36Z",
          "updatedAt": "2026-08-13T01:26:55Z",
          "timestamp": "2026-08-13T01:26:55Z",
          "metrics": {
            "reactions": 2,
            "comments": 24
          },
          "labels": [
            "pr-performance",
            "manual approve",
            "can be tested"
          ],
          "author": "ilejn",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f4055343166f349638f2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114275",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114275",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a Keeper server never joining the cluster when background snapshot IO is enabled",
          "text": "Related: https://github.com/ClickHouse/NuRaft/pull/123 ## Problem With `nuraft_use_bg_thread_for_snapshot_io` enabled, a Keeper server added to the cluster never joins. The leader's catch-up loop waits indefinitely for a response that cannot arrive, so the configuration change is never committed. The setting is off by default, which is why this has gone unnoticed. ## Root cause `sync_log_to_new_srv` builds one of two requests for the joining server — a snapshot request when the peer is below the retained log, or a log-pack `sync_log_request` otherwise — and then dispatches on the snapshot IO mode: ```cpp if (!params->use_bg_thread_for_snapshot_io_) { srv_to_join_->send_req(srv_to_join_, req, ex_resp_handler_); } else { snapshot_io_mgr::instance().invoke(); } ``` Only the *snapshot* branch hands its work to the IO thread. Nothing ever enqueues a log pack with `snapshot_io_mgr` — `create_sync_snapshot_req` is its only producer — so with background snapshot IO on, invoking the thread does not send the log-pack request. It is dropped. Observed on the unfixed build as the new server absent from the cluster configuration and stuck at log index 1 while the leader was at 12. ## Solution Dispatch on what was actually built rather than on the IO mode. The snapshot branch leaves the request null exactly when it has queued the work with the IO thread; in every other case — including the log-pack branch in *both* IO modes — there is a request and it must be sent. Covered by `log_sync_with_async_snapshot_io_test` in NuRaft's `learner_new_joiner_test`, verified to fail against the exact pre-patch code while the four existing tests in that binary keep passing. The submodule is pinned to the pull request branch so that the Keeper integration tests, `nightly_keeper` and Jepsen exercise it here. **The pointer must be moved to the merge commit before this pull request is merged.** ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a Keeper server never joining the cluster when `nuraft_use_bg_thread_for_snapshot_io` is enabled: the log batch sent to a joining server was silently dropped, so the membership change never completed.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114275",
          "createdAt": "2026-08-11T05:52:43Z",
          "updatedAt": "2026-08-13T01:24:03Z",
          "timestamp": "2026-08-13T01:24:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "submodule changed"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8264119ffa7aae7fa6e9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110180",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110180",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check access rights in EXPLAIN QUERY TREE and EXPLAIN SYNTAX",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/78938 `EXPLAIN QUERY TREE` and `EXPLAIN SYNTAX` (in the analyzer) resolve the query and dump table metadata such as column names and types, but unlike `EXPLAIN PLAN` they do not build a query plan. The `SELECT` access check that the planner performs in `prepareBuildQueryPlanForTableExpression` was therefore skipped, so a user with no privileges could read the column names and data types (including `Enum` element lists) of tables they are not allowed to access, while `SELECT`, `EXPLAIN PLAN` and `EXPLAIN PIPELINE` are correctly rejected. This adds a `SELECT` access check for every table referenced anywhere in the query tree (including tables inside subqueries in expressions such as `WHERE x IN (SELECT ... FROM t)`). The check mirrors the planner: it validates access to the columns that are actually read, with the trivial-count fallback (access is granted if at least one column is accessible) for queries that read no specific column, e.g. `SELECT count() FROM t`. As a result, a user with a column-level grant sees the same behavior as for a plain `SELECT` (`EXPLAIN QUERY TREE SELECT granted_col FROM t` is allowed, `... other_col ...` is denied). `buildQueryTree` only builds the tree; table identifiers are bound to storages and columns are resolved by the query analysis pass. The check runs on the resolved tree: on `query_tree` directly when the passes already resolved it, or on a throwaway resolved copy otherwise (e.g. `EXPLAIN QUERY TREE run_passes = 0`, which intentionally dumps the unresolved tree, so the tree that gets dumped is unchanged). ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `EXPLAIN QUERY TREE` and `EXPLAIN SYNTAX` not checking access rights on the referenced tables, which allowed a user without privileges to read table column names and data types. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110180",
          "createdAt": "2026-07-12T19:53:08Z",
          "updatedAt": "2026-08-13T01:18:57Z",
          "timestamp": "2026-08-13T01:18:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-must-backport",
            "pr-critical-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9b75a00941210600e184",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113382",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113382",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix division by zero in `ReadFromMergeTree` when `StorageMerge` truncates the number of streams",
          "text": "The `AST fuzzer (amd_debug)` crashed the server with `Integer divide by zero` on ```sql SELECT id FROM merge(currentDatabase(), '^t$') ORDER BY id DESC SETTINGS max_streams_to_max_threads_ratio = 1073741824, max_threads = 4 ``` The planner computes `max_streams = max_threads * max_streams_to_max_threads_ratio = 2^32`, which passes the existing overflow check (it only rejects values that do not fit into `size_t`). `ReadFromMerge::createPlanForTable` then passed `UInt32(streams_num)` to `storage->read`, truncating `2^32` to `0`. The child `ReadFromMergeTree` divides by the requested number of streams in `spreadMarkRangesAmongStreamsWithOrder` (`(info.sum_marks - 1) / num_streams`), so debug and sanitizer builds crash with SIGFPE (release builds happen to eliminate the dead division, since its only use is in a loop that runs zero times). `IStorage::read` takes `size_t num_streams`, so the cast was a pointless leftover — this change removes it and adds a regression test that crashes debug builds without the fix. Additionally (per review), `StorageMerge` multiplied the requested number of streams by `max_streams_multiplier_for_merge_tables` (clamped to the number of selected tables) with unchecked `Float64` -> `size_t` casts, while the planner only bounds-checks `max_threads * max_streams_to_max_threads_ratio`. Both multiplications now go through a shared helper that throws `PARAMETER_OUT_OF_BOUND` (like the planner check) when the product does not fit into `size_t`, with a second regression test. Also (per review), `spreadMarkRangesAmongStreams` and `spreadMarkRangesAmongStreamsWithOrder` clamp an unnecessarily large number of streams down to the amount of data, but the check itself computed `num_streams * min_marks_for_concurrent_read`, which wraps around for huge stream counts. The clamp was then skipped and the ordered path went on to `split_parts_and_ranges.reserve(num_streams)`, throwing `std::length_error` in every build — reachable both directly on a `MergeTree` table and, now that the stream count is no longer truncated, through a `Merge` table. The comparison now uses a division, which is equivalent (`min_marks_for_concurrent_read` is always at least one) and cannot overflow, with a third regression test. Also (per review, and confirmed by the AST fuzzer on this PR), a streaming read (`FROM ... STREAM`) bypasses the `spreadMarkRanges*` helpers entirely: `groupPartitionsByStreams` created one `MergeTreeCommitOrderSequentialSource` per requested stream with no clamp, so a huge `max_threads * max_streams_to_max_threads_ratio` product threw `std::length_error` from `pipes.reserve` (or would exhaust memory for merely absurd values). Since a streaming read has no marks by which the stream count could be clamped, such values are now rejected with `PARAMETER_OUT_OF_BOUND` (with a cap of one million, mirroring the limit on the number of threads in `MergeTreeReadPool`), with a fourth regression test. Found by the AST fuzzer on an unrelated PR: [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113010&sha=59507af89c33f62942e855b50b3e66704cbcf2e9&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29) Related: https://github.com/ClickHouse/ClickHouse/pull/113010 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a crash (arithmetic exception in debug and sanitizer builds) when reading from a `Merge` table with a very large `max_streams_to_max_threads_ratio`: the number of streams was truncated to 32 bits, so multiples of `2^32` became zero. Also, throw an exception instead of undefined behavior when the product of the number of streams and `max_streams_multiplier_for_merge_tables` exceeds the range of `size_t`, fix a `std::length_error` when reading a `MergeTree` table with a huge `max_streams_to_max_threads_ratio`, and reject huge stream counts on streaming reads (`FROM ... STREAM`) instead of trying to create a source per stream.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113382",
          "createdAt": "2026-08-04T20:03:50Z",
          "updatedAt": "2026-08-13T01:18:16Z",
          "timestamp": "2026-08-13T01:18:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9910e35ff5bbd5a5b1ef",
        "signalId": "github:ClickHouse/ClickHouse:issue:112444",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112444",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "[RFC] Query Acceleration: derived structures for transparent query rewrite",
          "text": "## Summary ClickHouse needs a first-class way to maintain **derived, query-specific data structures** next to a source table and use them **transparently** during planning — without making those structures a second system of record, and without forcing applications to query a side table. **Query Acceleration** is that framework: 1. User creates an **accelerator** bound to a source table and a set of indexed columns. 2. Background work **materializes** derived data from active source parts, with bounded resources and full ops visibility. 3. On a normal `SELECT` against the **source**, the optimizer asks registered accelerators whether they can improve the plan. 4. If matched, the plan uses the derived structure to produce **candidates / read hints**, then gathers the final rows from the source. If not matched, the ordinary source plan remains correct. The first concrete family is an experimental **ANN (approximate nearest-neighbor) index** for vector top‑k / bounded range queries. This issue is about the **acceleration principle and framework contract**; family-specific algorithms and knobs belong in per-family docs and follow-up issues. --- ## Motivation ### Problem Many expensive query patterns share the same operational needs: - The **source table** must stay the durable, mutable dataset (insert / merge / mutate as today). - A **derived structure** (index, sketch, layout, partial aggregate, …) would make a known shape much cheaper. - Construction should be **asynchronous**, resource-bounded, and observable — not folded only into foreground inserts or ad-hoc ETL. - Freshness may **lag** source changes for a while; queries must still be **correct** under partial coverage. - Applications should keep writing SQL against the **source**, not against an internal index table. Today’s alternatives each miss part of that story: | Approach | Gap | |---|---| | Full scan / exact evaluation only | Correct but does not scale for selective ranked or specialized filters | | Native secondary indexes tied to part metadata | Limited independence of lifecycle, algorithm choice, and ops surface | | Manual side tables / external systems | Dual-write or ETL; no transparent rewrite; correctness and freshness are the user’s problem | ### Why a framework (not one-off indexes) If each specialized structure re-implements catalog binding, rebuild, plan matching, hybrid fallback, and telemetry, the product fragments. Query Acceleration factors the common contract so a new family only owns: - on-disk / in-memory format and build; - query-shape recognition and cost; - how to search the derived data and map hits back to source rows; - family-specific settings and metrics. Candidate future families (design directions, not commitments): text / sparse retrieval, geospatial, join layouts, aggregation state, domain-specific predicates. --- ## Goals 1. **Source is authoritative.** User DML and ordinary reads target the source. Accelerators are recomputable from the source; they are not independently backed up as a second truth. 2. **Transparent acceleration.** Create once; query the source with natural SQL. No application dual-path to an index table. 3. **Correct under lag.** Incomplete materialization must not omit required rows from the query contract. Uncovered regions use a family-defined **exact / residual** path (or decline the rewrite), not silent omission. 4. **Declarative selection.** Matchers offer or decline with stable reasons. The planner never blindly forces an accelerator. 5. **Isolation of concerns.** Framework: catalog, lifecycle hooks, matching plumbing, coverage snapshot semantics, SYSTEM commands, shared telemetry. Family: format, build, search, cost, algorithm settings. 6. **Operable.** Coverage, backlog, jobs, failures, and per-query outcomes are first-class (system tables + logs). 7. **Fail closed.** Capability mismatch, corrupt lineage/mapping, and execution failures raise errors. They are not silently reclassified as “uncovered” or swapped to another interface. ## Non-goals - Turning an accelerator into a general-purpose user-facing table (`SELECT`/`INSERT`/`ALTER` as primary API). - Guaranteeing that every source query shape is accelerated. - Making approximate families exact by default (approximation, when used, is an explicit family semantic). - Shipping many families in the first GA; the first family validates the path, not the entire roadmap. --- ## Principle: how query acceleration works ### Lifecycle of an accelerator ```text CREATE ACCELERATOR acc ON source (cols) ENGINE = <Family>(...) │ ▼ Background scheduler compare active source parts ↔ recorded lineage build / compact / publish derived parts │ ▼ Query on source optimizer extracts a plan shape matcher: match or decline (reason_code) │ matched│ declined ▼ ▼ Hybrid / hinted plan Ordinary source plan accelerated regions + residual for uncovered │ ▼ Gather payload from source by identity ``` **Users never query the accelerator as a data table.** Lifecycle is `CREATE` / `DROP` / `DETACH` / `ATTACH` / rename, plus `SYSTEM` controls for refresh, start/stop builds, and bounded sync for tests and rollout. ### Four layers | Layer | Responsibility | |---|---| | **Catalog & binding** | Accelerator is catalog-visible, bound to one source and indexed columns; dependency prevents accidental drop when checks are enabled | | **Materialization** | Async jobs reconciling source parts with derived parts; resource pools and per-source limits | | **Planning & execution** | Shape recognition → matcher offer → rewrite that uses derived data only where coverage and capability allow | | **Operations & telemetry** | Live state, job history, query decision log, coverage and fallback outcomes | ### Match kind today: `ReadHint` The first rewrite pattern is **candidate-first / read-hint**: 1. Derived structure returns a bounded set of **source-row identities** (and optional scores). 2. Optional global selection (e.g. top‑k or quota) merges accelerated and residual producers. 3. Pipeline **gathers** full row payloads from the source using those identities (hints), instead of scanning the whole selected range for ranking. Other match kinds (e.g. partial aggregate, join probe) are possible later when a family needs them; they should not be invented without a concrete consumer. ### Fail-closed matching Matchers decline with stable `reason_code`s (shape unsupported, metric/config mismatch, no ready parts, competing native index preferred, …). Optional strict settings can turn “would fall back to full source plan” into an exception for tests and benchmarks, so silent non-acceleration is visible. ### Observability contract Operators should answer four questions without reading code: 1. What accelerators exist, and what is their **coverage / backlog / health**? 2. What **background jobs** are running or failing? 3. Did this query **match**, and if not, **why**? 4. Was the outcome fully accelerated, partial residual, full fallback, or failed? Shared objects (names may gain family-specific companions): - `system.accelerators` - `system.accelerator_jobs` / `accelerator_job_log` - `system.accelerator_query_log` (final outcomes + stage/decision rows) --- ## Core abstractions (implementation anchors) | Concept | Role | |---|---| | `IAccelerator` | Catalog storage base: source binding, family/impl, observability snapshot; exposes matcher + scheduler; rejects direct user R/W by default | | `IAcceleratorMatcher` | Plan-time: plan shape → offer (`ReadHint`) or decline with `reason_code` | | `IAcceleratorScheduler` | Background entry: snapshot + schedule jobs | | Optimizer glue | Extract shape, enumerate matchers, inject hybrid / hinted plan when accepted | | Family package | Build, format, search, cost, residual evaluator, family system tables | Framework code must stay family-agnostic enough that a **second family** does not require forking catalog DDL, SYSTEM commands, or query-log infrastructure. --- ## First implementation (brief) The first family is **`ANNIndex` / `ReplicatedANNIndex`**: approximate nearest-neighbor indexes over a vector column on a MergeTree-family source. - **Purpose:** accelerate source queries of the form ranked distance top‑k and certain bounded distance ranges. - **User model:** create accelerator → query **source** → optional `SYSTEM SYNC ACCELERATOR` for tests/rollout. - **Gate:** experimental setting for create/attach. - **Validates the framework:** async part-level builds, hybrid covered/uncovered execution, match/decline reasons, job and query telemetry, replication of derived parts as a separate lifecycle from the source. Algorithm choice, index parameters, filtered-search modes, and ANN-only ops details are **out of scope for this epic issue**; see the ANN accelerator guide and family follow-ups. --- ## Known framework gaps | Area | Gap | Direction | |---|---|---| | Families | Only one family exists | Land checklist + second family to prove generality | | Match kinds | Only `ReadHint` | Add kinds only with a real consumer | | Multi-accelerator choice | Limited arbitration when several could match | Cost-based selection + deterministic tie-break | | Distributed / parallel replicas | Accelerated path not universally available | Explicit policy per execution mode; residual ordinary plan | | Strict shapes | Some query features force full source plan | Document; extend rewrite only when ownership of predicates is clear | | Coupling risk | First-family types can leak into optimizer entry points | Keep matcher/plan-shape APIs generic; confine family pipelines | | GA readiness | Experimental gates, format stability, runbooks | Soak first family; freeze public contracts | --- ## Roadmap ### Phase A — Framework contract solid with first family as reference 1. Freeze DDL / SYSTEM / shared system-table and log contracts. 2. Document and test hybrid coverage, residual semantics, and fail-closed matching. 3. Golden `EXPLAIN` / outcome codes for match, partial residual, and full fallback. 4. Reduce first-family leakage into shared optimizer types. 5. Performance of shared paths: materialization scheduling, candidate production, payload gather. ### Phase B — Production readiness of the framework + first family 1. Format / upgrade story for derived parts. 2. Narrow experimental gates after soak. 3. Runbooks: coverage lag, degraded state, build storms, fallback rate alerts. 4. Clear distributed and parallel-replicas policy. ### Phase C — Second family (proves the framework) Pick one high-ROI shape (e.g. geospatial, constrained join, or repeated aggregation). Acceptance: lands **without** forking catalog, SYSTEM, or shared query-log infrastructure; only family package + thin matcher/optimizer registration. --- ## One-line pitch > Query Acceleration maintains derived structures beside a MergeTree source table and rewrites matching source queries transparently, with asynchronous builds, hybrid residual evaluation under partial coverage, and first-class operational telemetry — starting with an experimental ANN family for vector search.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112444",
          "createdAt": "2026-07-29T14:26:22Z",
          "updatedAt": "2026-08-13T01:17:56Z",
          "timestamp": "2026-08-13T01:17:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [],
          "author": "fastio",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:27a5e0edfaeaad45b53c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111867",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111867",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use Gaussian centroids for truncated QBit codes",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109405 Related: https://github.com/ClickHouse/ClickHouse/pull/110911 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve reduced-precision `L2DistanceTransposedQuantized`, `cosineDistanceTransposedQuantized`, and `dotProductTransposedQuantized` by reconstructing `p < 8` `QBit(Int8)` codes with Gaussian conditional-mean prefix centroids. Full precision (`p = 8`) remains bit-exact; reduced-precision approximate distances may change, and recall/latency gains are workload-dependent. ### Motivation The existing reduced-precision path selects one middle fine Lloyd-Max reconstruction level for every truncated prefix. That value is not the conditional mean of the complete Gaussian interval represented by the prefix. This change uses the conditional mean `(phi(lo) - phi(hi)) / (Phi(hi) - Phi(lo))` which minimizes scalar MSE for that prefix under the standard-normal source model of the existing Lloyd-Max codec. The 127 positive values are evaluated from the existing `Float32` boundaries at high precision, rounded once to `Float32`, stored as hexadecimal literals, and mirrored exactly for negative prefixes. This keeps the hot distance loop as one LUT lookup and avoids platform-dependent libm work during LUT initialization. ### Validation - A new stateless regression fails on the exact official `26.7.1.1315` binary for all 20 reduced-precision centroid/sign checks and passes its `p = 8` control. The candidate passes all 20 checks and the control. - An independent all-raw probe covers every raw byte and every `p = 1..8`: expected level counts, finiteness, symmetry, prefix-block invariance, index-order monotonicity, independent Gaussian means (maximum 0 ULP), and bit-exact legacy `p = 8` reconstruction. - An engine probe covers all raw bytes and precisions, non-strided and `QBit(Int8, 16, 8)`, `used_dims` 8/16, dot/L2/cosine, and `optimize_qbit_distance_function_reads` 0/1. Partial-read modes match bitwise; bounded SimSIMD tolerances are used for L2/cosine. - Debug `programs/clickhouse` build passes. Updated `04504_transposed_distance_quantized` and new `04628_qbit_lloyd_max_prefix_centroid` match their references with empty stderr in clean `clickhouse local` paths. On 103,000 source-disjoint Nomic embedding vectors (768 dimensions, 200 queries, four randomized-Hadamard seeds), scalar coordinate MSE decreases at every changed precision. The transformed-space retrieval proxy is deliberately reported separately because it is not a production ClickHouse latency benchmark: | `p` | scalar MSE delta | recall@10 delta (pp) | hit@1 delta (pp) | |---:|---:|---:|---:| | 1 | -3.83% | 0.000 | 0.000 | | 2 | -3.49% | -0.250 | -0.375 | | 3 | -0.89% | +0.538 | +0.750 | | 4 | -15.17% | +1.738 | +1.625 | | 5 | -17.70% | +1.300 | +1.875 | | 6 | -24.08% | +0.800 | +0.125 | | 7 | -44.32% | +0.588 | +0.250 | | 8 | unchanged | 0.000 | 0.000 | There is no universal retrieval improvement claim: `p = 2` regresses slightly in this proxy. There is also no latency claim until a matched Release-build p50/p95 benchmark is available. The intentional compatibility boundary is numerical output at `p < 8`; function signatures, storage, and `p = 8` results are unchanged.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111867",
          "createdAt": "2026-07-25T01:52:18Z",
          "updatedAt": "2026-08-13T01:17:55Z",
          "timestamp": "2026-08-13T01:17:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-improvement",
            "can be tested",
            "v26.7-must-backport"
          ],
          "author": "skuznetsov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:43edeaf4e1fcaa16be72",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110615",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110615",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support TLS/SSL for PostgreSQL connections",
          "text": "Adds TLS/SSL support to all PostgreSQL integrations. Until now, ClickHouse could not establish an encrypted connection to a PostgreSQL server that enforces SSL, nor verify the server certificate: `postgres::formatConnectionString` only emitted `dbname`/`host`/`port`/`user`/`password`/`connect_timeout`, so `libpq`'s `sslmode`/`sslrootcert`/`sslcert`/`sslkey` could never be set. The bundled `libpq` is already compiled with SSL support (`USE_SSL` is defined through `pg_config_manual.h`), so this change is purely about exposing the options — no build change is required. The design follows the model agreed in the review discussion (https://github.com/ClickHouse/ClickHouse/pull/110615#issuecomment-5116481882) and already implemented for MySQL in https://github.com/ClickHouse/ClickHouse/pull/112070: - `sslmode` (`disable`, `allow`, `prefer`, `require`, `verify-ca` or `verify-full`) can be specified anywhere: as a named collection key or as a trailing `key = value` argument after the positional arguments. When unset, the `libpq` default of `prefer` applies. - `sslrootcert` (CA certificate), `sslcert` (client certificate) and `sslkey` (client private key) are **paths to server-local files**. They are only accepted from a named collection defined in the server configuration file (or from a dictionary defined there) and cannot be overridden in a query: the server opens the files with its own privileges, so a path taken from SQL would let anyone who can define a PostgreSQL source probe the local filesystem and authenticate with a client certificate they are not allowed to read themselves. - `sslrootcert_pem`, `sslcert_pem` and `sslkey_pem` accept the **literal contents** of the corresponding file (copy-pasteable, also usable in ClickHouse Cloud where there is no server filesystem to reference). They are accepted from anywhere — a query, a named collection created with SQL, an override of a configuration-defined collection — and are masked in logs and `SHOW` queries like a password. `libpq` can only load credentials from files, so the contents are materialized into a private temporary file (mode 0600, the permission `libpq` requires of a key file) whose lifetime is tied to the connection pool that uses it. These parameters work for the `PostgreSQL` table engine, the `postgresql` table function, the `PostgreSQL` and `MaterializedPostgreSQL` database engines, the `MaterializedPostgreSQL` table engine, and `PostgreSQL` dictionaries. The dictionary source previously accepted an `sslmode` key but silently ignored it; it is now honored. The integration test `test_postgresql_ssl` enables TLS on a PostgreSQL server at runtime, requires SSL via `pg_hba.conf`, and checks positive and falsifying cases on every surface (wrong CA rejected, client certificate required, contents overriding a configured path, restart survival, masking). A stateless test pins the `[HIDDEN]` masking and the rejection of paths from SQL on every surface. Closes: https://github.com/ClickHouse/ClickHouse/issues/80787 Related: https://github.com/ClickHouse/ClickHouse/pull/112070 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support TLS/SSL connections to PostgreSQL for the `PostgreSQL` table engine, the `postgresql` table function, the `PostgreSQL` and `MaterializedPostgreSQL` database engines, and `PostgreSQL` dictionaries: `sslmode` plus the certificates and the key, given either as literal contents (`sslrootcert_pem`, `sslcert_pem`, `sslkey_pem`; masked like passwords) or as paths (`sslrootcert`, `sslcert`, `sslkey`; accepted only from a named collection defined in the server configuration file).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110615",
          "createdAt": "2026-07-15T19:38:24Z",
          "updatedAt": "2026-08-13T01:17:38Z",
          "timestamp": "2026-08-13T01:17:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 54
          },
          "labels": [
            "pr-feature",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a63df289bdc5c1551d45",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113021",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113021",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Build the `arrayIntersect` hash map from the smallest argument",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/2120 --> `arrayIntersect` filled its hash map from every argument and then rescanned the first one. For `arrayIntersect(a, b)` with 30 million and 5 million elements that is 35 million insertions into a map sized for the union of both, followed by 30 million lookups. A value that is missing from any one of the arguments cannot be in the intersection, so it is enough to fill the map from the smallest argument and to only look the other ones up. The map then stays as small as the smallest argument, which is what decides the speed once it no longer fits in cache. Which argument seeds the map does not affect the result, so it is chosen once for the whole column and short arrays pay nothing per row. Measured with a baseline built from unmodified sources in the same build directory (release, aarch64), best of three: | query | baseline | this PR | | |---|---|---|---| | `length(arrayIntersect(aa, bb))`, 30M vs 5M elements | 5.25 s, 1.94 GiB | 3.00 s, 1.26 GiB | 1.75x | | `length(arrayIntersect(bb, aa))`, 5M vs 30M elements | 3.95 s, 1.94 GiB | 2.14 s, 1.26 GiB | 1.85x | | 2 x 8M-element arrays | 1.59 s, 1.03 GiB | 1.26 s, 0.66 GiB | 1.26x | | 3 x 4M-element arrays | 0.86 s, 0.62 GiB | 0.75 s, 0.62 GiB | 1.15x | | 5M rows x 8-element arrays | 1.01 s | 1.05 s | 0.96x | | 1M rows x 4-element `String` arrays | 0.26 s | 0.26 s | 1.00x | | `arrayUnion`, 2 x 8M-element arrays | 1.41 s | 1.44 s | 0.98x | | `arraySymmetricDifference`, 2 x 8M-element arrays | 1.39 s | 1.42 s | 0.98x | The small cases lose 2-4%, reproducibly across best-of-seven runs, even though the new code executes fewer instructions for them (9.97 G against 10.38 G on the 5M x 8 case) - it looks like code layout, the element loop is now instantiated for both the filling and the looking-up argument. I left it as it is: a few percent on arrays that fit in cache in exchange for 1.75x on the arrays where this function actually gets slow. The performance tests run on quieter machines than the one I measured on, so they are the better judge of the small cases. Reordering the arguments requires the counter to be exact, and that also fixes a wrong result. A value repeated in a later argument used to be counted twice, so it could reach the \"present in every argument\" count while being absent from an argument in between: ```sql SELECT arrayIntersect([1], [2], [1, 1]); -- was [1], now [] SELECT arrayIntersect([1, 2], [2], [1, 1, 2]); -- was [1,2], now [2] SELECT arraySymmetricDifference([1], [2], [1, 1]); -- was [2], now [2,1] SELECT arraySymmetricDifference([1], [2, 2]); -- was [1], now [2,1] ``` For `arraySymmetricDifference` two arguments are enough, because it reads the counter directly instead of rescanning the first argument: the two copies of `2` took the counter to 2, which the old code read as \"present in both arguments\" and left out of the result. A value is now counted for an argument only when it was present in every argument before it, so a counter equal to the number of arguments means exactly \"present everywhere\". `arrayUnion` is unaffected, it only ever asked whether the counter was non-zero. The counter used to be exact - `if (*value == arg_num) ++(*value);`, written together with the comment above it in 54407002993. It was relaxed to `<= arg_num` in d29f0d4c966, the commit that added `arrayUnion` in https://github.com/ClickHouse/ClickHouse/pull/68989, because that mode takes every key whose counter is non-zero, and with the exact condition a value that first appears in argument `k > 0` is inserted as `0`, never incremented, and dropped from the union. What the relaxation also allowed - a second increment inside the same array whenever the counter lags behind `arg_num` - is the wrong result above. This PR does not restore the exact condition in place, it removes the need for the relaxation: `arrayUnion` no longer looks at the counter at all, because every key of the map is present in at least one of the arguments by construction. While in the same loop, `arrayUnion` and `arraySymmetricDifference` no longer look every key up again through `map.find` while iterating the map itself, which was one redundant lookup per key. Related: https://github.com/ClickHouse/ClickHouse/issues/2120 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `arrayIntersect` and `arraySymmetricDifference` no longer treat a value repeated inside a single argument as if it appeared in several arguments. Queries relying on the previous behavior can now return different results: `arrayIntersect([1], [2], [1, 1])` returns `[]` instead of `[1]`, `arrayIntersect([1, 2], [2], [1, 1, 2])` returns `[2]` instead of `[1, 2]`, and `arraySymmetricDifference([1], [2], [1, 1])` returns `[2, 1]` instead of `[2]`. For `arraySymmetricDifference` two arguments are already enough: `arraySymmetricDifference([1], [2, 2])` returns `[2, 1]` instead of `[1]`. A value is now counted for an argument only when it was present in every argument before it, so the result contains exactly the values present in all of the arguments. As part of the same change, `arrayIntersect` builds its hash table from the smallest argument rather than from all of them, which makes it up to 1.85x faster and use a third less memory when the arguments differ a lot in size. `arrayUnion` is not affected. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1295` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113021",
          "createdAt": "2026-08-02T18:10:30Z",
          "updatedAt": "2026-08-13T03:45:12Z",
          "timestamp": "2026-08-13T03:45:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "pr-backward-incompatible",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f12f677fed4ab67d7fe6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114527",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114527",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix flaky `test_url_reconnect`",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110130 `test_redirect_url_storage/test.py::test_url_reconnect` blocks the outgoing connections to the HDFS datanode web port with a silent `DROP` and heals the network only after the client has demonstrably hit a connect timeout, waiting up to 30 seconds for a fresh `connect timed out` line in the server log. That budget is not always enough, because the query does not necessarily open a new connection at all: the earlier tests in the module leave keep-alive connections to `hdfs1:50075` in the HTTP connection pool, and a silent `DROP` kills an already established connection only by the receive timeout - `http_receive_timeout`, 30 seconds by default - rather than by the one-second connect timeout. The first try then dies after 30 seconds with a bare `Timeout`, and the first `connect timed out` line appears only on the next try, about 31 seconds into the query, just past the budget: ``` 16:57:35.001 executeQuery: select sum(cityHash64(id)) from url('http://hdfs1:50075/... 16:58:05.013 ReadWriteBufferFromHTTP: ... Error: Timeout. Failed at try 2/10. 16:58:06.115 ReadWriteBufferFromHTTP: ... Error: Timeout: connect timed out: 172.16.1.2:50075. Failed at try 3/10. ``` The fix drops the connection cache before installing the rule, so the query has to open a fresh connection and the first try fails by the connect timeout, as the test expects. The wait budget is also raised to 60 seconds: the loop exits as soon as the line shows up, so a generous budget costs nothing on a healthy run. Seen in an unrelated pull request: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110130&sha=7acc19de7fb0c97f7ea0fb7c6c234930d188ac66&name_0=PR&name_1=Integration%20tests%20%28amd_msan%2C%205%2F8%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1294` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114527",
          "createdAt": "2026-08-12T17:49:11Z",
          "updatedAt": "2026-08-13T01:17:29Z",
          "timestamp": "2026-08-13T01:17:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:76e260babe4c29b2d502",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114556",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114556",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Ask `@oranjeai` instead of `@groeneai` in the private repository",
          "text": "The `continue-pr` skill hardcoded `@groeneai` as the person to ping about CI failures that are unrelated to the pull request being worked on. That handle is only correct for `ClickHouse/ClickHouse`; when the skill runs in `ClickHouse/clickhouse-private`, the request should go to `@oranjeai` instead. Step 1 now detects the repository (`gh repo view --json nameWithOwner`, falling back to `git remote get-url origin`), remembers it as `$REPO` for the `gh` commands, API URLs, and GraphQL arguments further down, and step 4 picks the reviewer from it. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1297` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114556",
          "createdAt": "2026-08-12T23:18:07Z",
          "updatedAt": "2026-08-13T01:17:25Z",
          "timestamp": "2026-08-13T01:17:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a98966174cb81babba3e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114400",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114400",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not let artifact-collection rows discard a bugfix-validation verdict",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113397 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description @ nikitamikhaylov asked me to investigate why `Bugfix validation` fails on #113397 and suggest improvements ([comment](https://github.com/ClickHouse/ClickHouse/pull/113397#issuecomment-5258844349)). This fixes the harness defect I found while answering that. Related: #113397, @ tiandiwonder's PR, which this one does not close. There, both functional bugfix-validation jobs report FAIL even though the bug reproduced on both arches: the regression-test row is `OK`, and the only FAIL row is `Scraping system tables`. An ordering defect. `invert_bugfix_validation_status` decides the verdict, then the COLLECT_LOGS stage appended artifact-collection rows with a bare `extend_sub_results`, which re-derives the parent status from its children (`praktika/result.py`), so one failed system-table dump overwrote the verdict with `FAIL`. `new_tests_check.py` reads that status with strict `is_success`, so `any_bugfix_validation_passed` found no validating arch and blocked the PR with \"No per-arch Bugfix Validation job validated the bug\". Those rows come from `prepare_logs`, after the verdict, so they can never be part of it. The fix attaches them through a helper that restores the captured status on a labelled bugfix-validation job. The rows stay visible in the report, they just stop voting. Restoring the captured status rather than forcing `OK` keeps all three verdicts: reproduction `OK`, no-repro `SKIPPED`, inconclusive `ERROR`. An ordinary job still reddens. The defect reproduces in-process, and six new cases in `ci/tests/test_bugfix_validation_inverter.py` pin an ordering that had no coverage. Each was checked against a mutated fix so its redness is attributable: dropping the restore fails the three verdict cases, restoring the bare call fails the AST case, and applying the restore unconditionally fails the ordinary-job control. One records why the rows must not be labelled `LOG_CHECK` instead, which would count a broken dump as a reproduction. Only the ordering is fixed here. The dump itself failed because a graceful-stop timeout left the server holding its status-file lock, which is separate. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1299` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114400",
          "createdAt": "2026-08-12T00:02:39Z",
          "updatedAt": "2026-08-13T01:17:21Z",
          "timestamp": "2026-08-13T01:17:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3a19af718e555b62dfa1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112921",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112921",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Evaluate randomHadamardTransform once for a constant vector",
          "text": "`randomHadamardTransform` did not use the default implementation for constant arguments, so it called `convertToFullColumnIfConst` on its first argument and ran the transform for every row of the block even when the vector was a constant. It also returned a plain `ColumnArray` rather than a `ColumnConst`, so the analyzer could not fold the call to a literal either (`resolveFunction.cpp` only folds when the executed column is a `ColumnConst`). This matters for the intended vector search usage, where the query vector has to be rotated the same way as the stored vectors: ```sql WITH randomHadamardTransform([...]) AS target SELECT id FROM t ORDER BY cosineDistanceTransposedQuantized(vec_quantized, target, 2, 384) LIMIT 100 ``` On 2 million 768-dimensional vectors (`QBit(Int8, 768, 128)`, 2 bits and 384 dimensions, 64-core machine, warm cache): | | before | after | |---|---|---| | transform written inline in the query | 2.744 s | **0.144 s** | | transform hidden in a scalar subquery (computed once) | 0.148 s | 0.126 s | The whole difference was the transform of the constant query vector being repeated for every row; after the change the inline form costs the same as the hand-rolled workaround. The fix sets `useDefaultImplementationForConstants` and declares `seed` and `output_dims` in `getArgumentsThatAreAlwaysConstant`, so an all-constant call is executed on a single row and wrapped in a `ColumnConst`. The constant check for `seed` and `output_dims` now comes from the base class, with the same `ILLEGAL_COLUMN` error code as before. Related: https://github.com/ClickHouse/ClickHouse/issues/103466 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `randomHadamardTransform` of a constant vector is now evaluated once instead of once per row. This speeds up vector search queries that rotate the query vector with `randomHadamardTransform` before comparing it against a `QBit` column.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112921",
          "createdAt": "2026-08-01T15:44:42Z",
          "updatedAt": "2026-08-13T01:17:17Z",
          "timestamp": "2026-08-13T01:17:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-performance",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v26.7-must-backport"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f1a9cd7393210f53d735",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114434",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114434",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Preserve the RabbitMQ broker log in integration tests",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114415 Related: https://github.com/ClickHouse/ClickHouse/pull/113610 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description When a RabbitMQ container never becomes available, `wait_rabbitmq_to_start` raises and every test in the module fails from that one event. On master @ `e4d3698282346fcd7e496ef668ff517106c84ac5` that cost all 21 tests of `test_storage_rabbitmq/test_system_stop.py` from one fixture failure ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=e4d3698282346fcd7e496ef668ff517106c84ac5&name_0=MasterCI&name_1=Integration%20tests%20%28arm_binary%2C%20distributed%20plan%2C%201%2F4%29)). Four such events in 21 days, two on master. The artifact that would say why does not survive the run. `rabbitmq.conf` sets `log.file = /rabbitmq_logs/rabbit.log` and no `log.console`, so everything goes to that file and nothing to the console, which is also why the `docker logs` capture is empty. `/rabbitmq_logs` was declared `tmpfs`, so the log died with the container. `cluster.py` has always exported `RABBITMQ_LOGS` and `RABBITMQ_LOGS_FS=\"bind\"` and created the host directory, but no compose file consumed them after 545bb5bae91233313868ee4d98d7e6dc6efe20db replaced the bind mount with tmpfs. RabbitMQ was the only broker whose `*_LOGS` export was dead. `rabbit.log` has zero hits in the failing run's whole artifact bundle. This restores the mount in the same `${VAR_FS:-tmpfs}` / `${VAR:-}` form the siblings use, so an unset `RABBITMQ_LOGS` still degrades to tmpfs. `/var/lib/rabbitmq` stays a tmpfs: it holds Mnesia state, and persisting it would change test semantics. `test_system_stop.py` passes 21/21 and `test.py` 58/58 with the mount, each producing a `rabbit.log` of over 100 KB under `_instances*/rabbitmq/logs/`, which the integration job already tars. The new `test_rabbitmq_broker_log_is_collected` asserts that log exists and is non-empty; reverting the hunk alone leaves the rest of the module passing and reddens only that test, on an empty log directory. This neither stops the abort nor explains why the Erlang node failed to register with epmd. #114415 recreates a hung container; this makes the next occurrence diagnosable. #113610 fixed only the second diagnostics loop of `wait_rabbitmq_to_start` and is not a container-start fix.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114434",
          "createdAt": "2026-08-12T08:11:44Z",
          "updatedAt": "2026-08-13T01:17:12Z",
          "timestamp": "2026-08-13T01:17:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:68f4d61a515aa5895742",
        "signalId": "github:ClickHouse/ClickHouse:issue:101747",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:101747",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "checkTableNameLengthUnlocked in renameDatabase gated behind dependency-checking condition, bypassed with check_table_dependencies=0",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _Retrospective finding from a historical scan of [PR #79488](https://github.com/ClickHouse/ClickHouse/pull/79488) (merged 2025-04-25). Confirmed on current codebase — close with a note if already fixed._ ### Describe what's wrong RENAME DATABASE to a longer name bypasses table name length validation when check_table_dependencies=0, creating tables that cannot be dropped (filesystem error: File name too long) **Root cause:** DatabaseAtomic.cpp:725-729: checkTableNameLengthUnlocked is placed inside the `if (check_ref_deps || check_loading_deps)` block, but the name length check is independent of dependency checking and should be unconditional **Why we believe this is a bug:** DatabaseAtomic::renameDatabase (DatabaseAtomic.cpp:715) → check_ref_deps/check_loading_deps computed from settings (line 723-724) → if-block at line 725 gates BOTH dependency check AND name length check → when check_table_dependencies=0, the entire block is skipped including checkTableNameLengthUnlocked at line 729 **Affected locations:** - `src/Databases/DatabaseAtomic.cpp:729` — checkTableNameLengthUnlocked inside dependency-checking if-block **Impact:** When a user renames a database to a longer name with check_table_dependencies=0, existing tables may end up with names exceeding filesystem limits. These tables become undropable (DROP TABLE and DROP DATABASE both fail with 'File name too long' error), requiring manual intervention to recover. ### Does it reproduce on most recent release? Yes — confirmed on current `master` (commit `18dfe15a9516`). ### How to reproduce ```sql DROP DATABASE IF EXISTS db_short_79488; CREATE DATABASE db_short_79488 ENGINE=Atomic; CREATE TABLE db_short_79488.aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa (x UInt8) ENGINE=MergeTree ORDER BY x; -- With check_table_dependencies=0, RENAME should STILL check table name length -- but currently bypasses it (the check is inside the dependency-checking if-block) SET check_table_dependencies=0; RENAME DATABASE db_short_79488 TO db_short_79488_with_a_much_longer_name_reduces_table_len; -- { serverError ARGUMENT_OUT_OF_BOUND } -- Cleanup: rename back if the above succeeded (bug), then drop SET check_table_dependencies=0; RENAME DATABASE db_short_79488_with_a_much_longer_name_reduces_table_len TO db_short_79488; -- { serverError UNKNOWN_DATABASE } DROP DATABASE IF EXISTS db_short_79488; ``` [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/9a5afb99-69f7-4a1c-b97e-d21efa069052) ### Expected behavior ``` No output (RENAME should fail with ARGUMENT_OUT_OF_BOUND error code 69) ``` ### Error message and/or stacktrace ``` The query succeeded but the server error '69' was expected (query: RENAME DATABASE db_short_79488 TO db_short_79488_with_a_much_longer_name_reduces_table_len; -- { serverError ARGUMENT_OUT_OF_BOUND }). ``` ### Additional context **Open risks:** - Other database engines inheriting from DatabaseOnDisk may have similar issues if they override renameDatabase **Suggested fix:** Move the checkTableNameLengthUnlocked loop outside and before the `if (check_ref_deps || check_loading_deps)` block, so the name length check runs unconditionally regardless of dependency-checking settings. **Analysis details:** Confidence HIGH | Severity P1 | Testability: `STATELESS_SQL` Found during automated review of [PR #79488](https://github.com/ClickHouse/ClickHouse/pull/79488). --- _ClickGapAI · Confidence: HIGH · Severity: P1 · Finding: `h_pr79488_001`_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/101747",
          "createdAt": "2026-04-04T04:52:38Z",
          "updatedAt": "2026-08-13T01:16:05Z",
          "timestamp": "2026-08-13T01:16:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "comp-database-engines"
          ],
          "author": "clickgapai",
          "state": "closed",
          "assignees": [
            "mstetsyuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d9911500f277e88894ff",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112876",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112876",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Azure: fix ranged copy corruption and harden the copy path",
          "text": "> **Series**: #112871 -> **#112876** (this), #112872, #112873, #112874, #112875 **Problem.** A ranged Azure copy corrupts the destination — native `CopyFromUri` copies the whole source blob (ignoring offset/size) and the single-part read+write fallback reads from byte 0 — and a transient 403/auth after the server-side copy starts spawns a second writer to the same blob. Failure scenario: ``` copyFile(src, offset=X>0, size=S): native CopyFromUri -> copies the ENTIRE src (ignores X,S) -> corrupt single-part fallback -> reads [0,S) not [X,X+S) -> corrupt transient 403 after StartCopyFromUri -> read+write fallback -> 2 writers to dest ``` **Fix.** Gate native copy on `offset==0`; seek the fallback via `LimitSeekableReadBuffer(offset, total_size)`; only *start* the copy in the guarded block and poll the same operation on transient errors (no mechanism switch); retry the read+write upload writes through the shared retry helper. **Changes.** - `IO/AzureBlobStorage/copyAzureBlobStorageFile`: `offset==0` native-copy gate; seek the single-part fallback; poll-same-operation on transient errors; retry upload writes. - `tests/integration/{test_backup_restore_azure_blob_storage,test_backup_restore_s3}`: incremental Log-family / ranged-copy regression tests. ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a ranged Azure Blob Storage copy corrupting the destination during incremental backups: server-side native copy ignored the requested offset and size, and the single-part read-and-write fallback read from the start of the source. Native copy is now used only for a full-object copy, the fallback seeks to the correct offset, and a transient error after a server-side copy has started no longer starts a second writer to the destination.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112876",
          "createdAt": "2026-08-01T06:54:19Z",
          "updatedAt": "2026-08-13T01:11:55Z",
          "timestamp": "2026-08-13T01:11:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-critical-bugfix"
          ],
          "author": "arsenmuk",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:410223a69f3864862553",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114542",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114542",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: improve data warehousing diagram rendering",
          "text": "Replaces the shared data warehousing diagram with an opaque dark-canvas version that includes internal padding. This keeps the artwork consistent in light and dark themes and prevents edge cropping without CSS or component changes. The existing asset path remains valid for the English page and all generated locale pages. Validated in the local Mintlify preview in both light and dark themes. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Updated the data warehousing architecture diagram for consistent theme rendering. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1292` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114542",
          "createdAt": "2026-08-12T21:01:08Z",
          "updatedAt": "2026-08-13T01:07:00Z",
          "timestamp": "2026-08-13T01:07:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation",
            "pr-synced-to-cloud"
          ],
          "author": "dhtclk",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9884169a43595d432f58",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113899",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113899",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Avoid scans for constant sort keys",
          "text": "## What `isAlreadySorted` now returns early, right after the sort descriptors are resolved, when every sort key is a `ColumnConst`. Writing a `MergeTree` part whose sorting keys are all constant no longer walks the block doing adjacent-row comparisons to confirm an ordering that is constant by construction. Collation validation still runs before the early return, and any block that mixes constant and non-constant keys keeps the existing comparator path — only the all-constant case takes the new path. ## Why it helps When the sorting-key columns are constant across a block — a common shape when a leading `ORDER BY` column is fixed per part (time-ordered or batched ingestion, per-source or per-partition writes) — the sortedness check was doing a full comparison pass to reach a foregone conclusion. Returning as soon as the keys are known-constant removes that pass. Measured on `MergeTree inserts with constant and mixed sorting keys`, 64K/256K/1M rows (paired medians, co-measured on both trees): | Metric (1M rows) | Before | After | Δ | | --- | ---: | ---: | ---: | | Constant-key sortedness check, 1 key | 1106 µs | 2 µs | **−99.8%** | | Constant-key sortedness check, 4 keys | 4449 µs | 4 µs | **−99.9%** | | End-to-end insert latency, 4 keys | 24548 µs | 20201 µs | **−17.7%** | | Insert CPU, 4 keys | 20732 µs | 16301 µs | **−21.4%** | Smaller block sizes land in the same range (e.g. 1-key 64K: 66 µs → 2 µs). The sortedness check collapses to a near-constant cost, and that saving carries into a full insert as the end-to-end latency and CPU gains. Blocks that aren't all-constant take the unchanged comparator path. ## Testing Stateless coverage exercises single and multiple constant keys plus a non-constant suffix (the mixed case that must keep comparing). The declared correctness check passed. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Skip the redundant sortedness scan when writing MergeTree parts whose sorting keys are all constant. --- Contributed by [Perfloop](https://app.perfloop.ai): the numbers above were co-measured on both trees and independently re-verified before submission — the full public record is at [case_8ebwekrder](https://app.perfloop.ai/t/oss/case_8ebwekrder). Replies from this account are human-approved, and a human operator is accountable for this contribution.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113899",
          "createdAt": "2026-08-07T23:19:06Z",
          "updatedAt": "2026-08-13T01:04:33Z",
          "timestamp": "2026-08-13T01:04:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "perfloop-agent",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c53ca6e3834ec1897283",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113266",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113266",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not inject random ORDER BY into queries planned to an intermediate stage",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110188 With `inject_random_order_for_select_without_order_by = 1`, `InjectRandomOrderIfNoOrderByPass` wrapped every top-level query into `SELECT * FROM (...) ORDER BY rand()`, including queries that are planned only up to an intermediate stage. For a `Merge` table with a `Distributed` child, the other children are planned to `WithMergeableState`; the wrapper made such a child plan complete its aggregation, so blocks without `AggregatedChunkInfo` reached `MergingAggregatedTransform`, and the query failed with a logical error (exception) `Chunk info was not set for chunk in MergingAggregatedTransform`. Now the injection pass is only added when the query is processed to stage `Complete`, so intermediate-stage plans (`Merge`/`Distributed` children) are left untouched, while the injection still applies to the user-facing query. Minimal reproducer (found by BuzzHouse on [this report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110188&sha=cdc23648d36878b136bf2c8d87ee9b0a29651412&name_0=PR&name_1=BuzzHouse%20%28arm_asan_ubsan%29), unrelated to that PR's changes — it reproduces on master): ```sql CREATE TABLE t_local (x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t_local SELECT number FROM numbers(100); CREATE TABLE t_dist (x UInt64) ENGINE = Distributed('test_shard_localhost', currentDatabase(), 't_local'); SET inject_random_order_for_select_without_order_by = 1; SELECT count() FROM merge(currentDatabase(), '^t_(local|dist)$'); -- LOGICAL_ERROR before this fix ``` ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a logical error (exception) `Chunk info was not set for chunk in MergingAggregatedTransform` when the setting `inject_random_order_for_select_without_order_by` is enabled and an aggregation query reads from a `Merge` table containing a `Distributed` child: the random `ORDER BY rand()` wrapper is no longer injected into queries planned only up to an intermediate stage.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113266",
          "createdAt": "2026-08-04T08:30:54Z",
          "updatedAt": "2026-08-13T01:01:26Z",
          "timestamp": "2026-08-13T01:01:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a4ed5752bd0ddc703e3f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:79490",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:79490",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Real cache size",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Added a new filesystem cache setting `use_real_disk_size` (disabled by default). When enabled, cached-file sizes are accounted in filesystem allocation units (block size) instead of the written byte count, so the `FilesystemCacheSize` metric, the eviction-size metrics, and `current_size` in `system.filesystem_cache_settings` use block-aligned accounting. Without it, a cached file smaller than one block (for example, a 1-byte file) is accounted as its written size while occupying a whole block on disk, which can make the reported cache size much smaller than the actual on-disk usage. The accounting is a block-aligned approximation of the physical size, not the exact allocated size. ### Documentation entry for user-facing changes - [x] Documentation is written Documented the `use_real_disk_size` filesystem cache setting in `docs/en/operations/storing-data.md`. It accounts cache usage in filesystem allocation units (block size), so `FilesystemCacheSize`, the eviction-size metrics, and `current_size` in `system.filesystem_cache_settings` use block-aligned accounting (a block-aligned approximation of the physical on-disk usage). The per-segment `size` in `system.filesystem_cache` is not affected by this setting and stays the segment's logical range size; for a partially downloaded segment, use `downloaded_size` for the bytes actually written. <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/79490",
          "createdAt": "2025-04-23T15:44:34Z",
          "updatedAt": "2026-08-13T00:59:05Z",
          "timestamp": "2026-08-13T00:59:05Z",
          "metrics": {
            "reactions": 1,
            "comments": 35
          },
          "labels": [
            "pr-improvement",
            "manual approve",
            "can be tested"
          ],
          "author": "codeworse",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cf9d719449194d2bf9f1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:73776",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:73776",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enable dynamic evaluation of whether a short-circuit function's argument should be lazily executed",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Lazy execution in short-circuit evaluation is not entirely cost-free. The `maskedExecute` method incurs additional overhead for filtering and expanding columns. In certain scenarios, this overhead can exceed the cost of fully evaluating the expression. In this PR, we have added runtime performance profiling for functions. Based on the collected runtime performance data, we can evaluate the cost of using lazy evaluation versus direct full computation for a function, and dynamically select the execution strategy with the lower cost. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> > Information about CI checks: https://clickhouse.com/docs/en/development/continuous-integration/ #### CI Settings (Only check the boxes if you know what you are doing): - [ ] <!---ci_set_required--> Allow: All Required Checks - [ ] <!---ci_include_stateless--> Allow: Stateless tests - [ ] <!---ci_include_stateful--> Allow: Stateful tests - [ ] <!---ci_include_integration--> Allow: Integration Tests - [ ] <!---ci_include_performance--> Allow: Performance tests - [ ] <!---ci_set_builds--> Allow: All Builds - [ ] <!---batch_0_1--> Allow: batch 1, 2 for multi-batch jobs - [ ] <!---batch_2_3--> Allow: batch 3, 4, 5, 6 for multi-batch jobs --- - [ ] <!---ci_exclude_style--> Exclude: Style check - [ ] <!---ci_exclude_fast--> Exclude: Fast test - [ ] <!---ci_exclude_asan--> Exclude: All with ASAN - [ ] <!---ci_exclude_tsan|msan|ubsan|coverage--> Exclude: All with TSAN, MSAN, UBSAN, Coverage - [ ] <!---ci_exclude_aarch64|release|debug--> Exclude: All with aarch64, release, debug --- - [ ] <!---ci_include_fuzzer--> Run only fuzzers related jobs (libFuzzer fuzzers, AST fuzzers, etc.) - [ ] <!---ci_exclude_ast--> Exclude: AST fuzzers --- - [ ] <!---do_not_test--> Do not test - [ ] <!---woolen_wolfdog--> Woolen Wolfdog - [ ] <!---upload_all--> Upload binaries for special builds - [ ] <!---no_merge_commit--> Disable merge-commit - [ ] <!---no_ci_cache--> Disable CI cache",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/73776",
          "createdAt": "2024-12-24T06:56:51Z",
          "updatedAt": "2026-08-13T00:57:54Z",
          "timestamp": "2026-08-13T00:57:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "lgbo-ustc",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5df38a8d3bd9672b73af",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113534",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113534",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix a mixed JOIN ON condition evaluated over mismatched column types for a dictionary",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112831 A mixed join condition is a cross-side non-equi residual in the `ON` clause, for example `ON (t.key = d.key) AND (t.a * 10 < d.a)`. Only the hash family evaluates one: `HashJoin` passes `TableJoin::getMixedJoinExpression` into `AddedColumns::additional_filter_expression` and applies it while matching, resolving the right columns it needs **by name** against the stored right blocks. `buildPhysicalJoinImpl` drops the `join_use_nulls` conversion of the right columns from the right-side expression when the right side is a prepared storage, because such a storage delivers its columns already converted, and for a key-value storage it merges the conversion back into the right-side DAG under the original column names. But `DirectKeyValueJoin` declines a mixed condition, so a join onto a dictionary that carries one runs an ordinary algorithm which reads the dictionary as a stream - and is handed a right side whose `a` is `Nullable(UInt32)` under the very name the mixed condition declared as `UInt32`. `buildAdditionalFilter` creates the column from the declared type and fills it with `insertFrom` from the stored one, so a `ColumnUInt32` was filled from a `ColumnNullable`. ```sql CREATE TABLE dsrc (key UInt64, a UInt32) ENGINE = Memory; INSERT INTO dsrc VALUES (1, 100), (2, 20); CREATE DICTIONARY dict (key UInt64, a UInt32) PRIMARY KEY key SOURCE(CLICKHOUSE(TABLE 'dsrc')) LIFETIME(0) LAYOUT(FLAT()); CREATE TABLE t (key UInt64, a UInt32) ENGINE = Memory; INSERT INTO t VALUES (1, 1), (2, 2), (3, 3); SELECT count(), sum(d.a) FROM t LEFT ANY JOIN dict AS d ON (t.key = d.key) AND (t.a * 10 < d.a) SETTINGS allow_experimental_join_condition = 1, join_use_nulls = 1, join_algorithm = 'hash'; ``` Only key 1 satisfies the residual, so the answer is `3, 100`. A release build returned `3, 120`, having matched key 2 as well; a sanitizer build aborted on the `IColumn::insertFrom` type assertion. The fix keeps the conversion in the right-side expression when a key-value storage carries a mixed condition, so the join applies `join_use_nulls` itself exactly as it does for an ordinary table and the mixed condition keeps the non-`Nullable` columns it was built on. `StorageJoin` is unaffected - it rejects a mixed condition outright with `INCOMPATIBLE_TYPE_OF_JOIN`. `HashJoin` now also compares the type and not only the name when it resolves those columns, so a future divergence throws where both types are still known instead of reading a column through a mismatched `IColumn` interface. The shape has been reachable with an explicit `join_algorithm = 'hash'` for as long as the mixed condition and the prepared-storage handling have coexisted. https://github.com/ClickHouse/ClickHouse/pull/112831 made `DirectKeyValueJoin` decline a mixed condition, which put it on the default `join_algorithm` path and into the new `04667_mixed_join_condition_algorithm_selection` test, so the assertion started firing in most stress runs from 2026-08-04 - STID `2508-30f6` and `2508-2e0c`, which differ only by the concurrent-hash-join wrapper in the stack. No tracking issue was open for either STID. CI report the fix was written from: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=0f1b923e6247ef3d8409635a69738c7f3f02b84a&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results and, in a build with assertions enabled, an aborted assertion for a `JOIN` onto a dictionary whose `ON` clause has a non-equi condition over both tables, such as `ON (t.key = d.key) AND (t.a * 10 < d.a)`, when `join_use_nulls` is enabled. The condition was evaluated over a column read through a mismatched type, so rows could match arbitrarily.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113534",
          "createdAt": "2026-08-05T17:17:26Z",
          "updatedAt": "2026-08-13T00:57:05Z",
          "timestamp": "2026-08-13T00:57:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v26.7-must-backport"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d7fed54daa2bd584c81a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114537",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114537",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Resolve the query status per call in functions that check for cancellation\"",
          "text": "Reverts ClickHouse/ClickHouse#113456 Needs a better solution",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114537",
          "createdAt": "2026-08-12T19:52:40Z",
          "updatedAt": "2026-08-13T00:56:46Z",
          "timestamp": "2026-08-13T00:56:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6e2d88c6ccc691530b6f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109710",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109710",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "[DRAFT] Remove CatBoost integration",
          "text": "One fine day ... Requires ~https://github.com/ClickHouse/ClickHouse/pull/80363~ https://github.com/ClickHouse/ClickHouse/pull/109999 first. Removes the CatBoost integration: the `catboostEvaluate` function, the `system.models` table, the `SYSTEM RELOAD MODEL(S)` queries and the matching `SYSTEM RELOAD MODEL` privilege, together with their documentation, tests and the bundled model artifacts. The `CANNOT_LOAD_CATBOOST_MODEL` and `CANNOT_APPLY_CATBOOST_MODEL` error codes are intentionally kept so their numbers are not reused. Refs: - https://github.com/ClickHouse/ClickHouse/issues/70771 - https://github.com/ClickHouse/ClickHouse/issues/45052 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Removed the CatBoost integration. The `catboostEvaluate` function, the `system.models` table and the `SYSTEM RELOAD MODEL(S)` queries are no longer available. Queries, views and grants referring to them have to be migrated before the upgrade: in particular, revoke `SYSTEM RELOAD MODEL` from all users and roles beforehand, because an access entity whose grants still mention the removed privilege cannot be parsed on startup. Also let any in-flight `SYSTEM RELOAD MODEL(S) ON CLUSTER` entries drain from the distributed DDL queue before the upgrade: upgraded hosts can no longer parse such entries and will finish them with an error status instead of executing them. The same applies to persisted table metadata: a table whose column `DEFAULT`/`MATERIALIZED`/`ALIAS` expressions, secondary indices, constraints, projections, partition/order/sample keys, or TTL expressions still reference `catboostEvaluate` cannot be loaded after the upgrade (`UNKNOWN_FUNCTION`), so drop or rewrite such definitions beforehand. The revoke also has to cover every other carrier of serialized grants: a `GRANT SYSTEM RELOAD MODEL` line in the `grants` section of `users.xml` must be removed as well, because it can no longer be parsed and the server then refuses to load the users configuration at startup; on clusters with replicated (Keeper-backed) access storage, revoke the privilege before upgrading any replica, because with the default `access_control_improvements.throw_on_invalid_replicated_access_entities = 0` an upgraded replica that reads a user or role still carrying `SYSTEM RELOAD MODEL` from ZooKeeper cannot parse it and silently drops that entity from its in-memory copy (set `throw_on_invalid_replicated_access_entities = 1` to fail loudly instead); and access backups (`BACKUP ... ACCESS`) taken while the grant was still present cannot be restored with `RESTORE ... ACCESS` on the new version, so re-create such backups after the revoke.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109710",
          "createdAt": "2026-07-07T20:39:00Z",
          "updatedAt": "2026-08-13T00:56:28Z",
          "timestamp": "2026-08-13T00:56:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-backward-incompatible",
            "hold",
            "pr-autogenerated-docs"
          ],
          "author": "rschu1ze",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:045f30b31107c62a71d8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114528",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114528",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #102033 to 26.5: Fix LOGICAL_ERROR crash in IcebergMetadata::iterate when datalake_table_state is missing",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/102033 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114385 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31624414958/job/94207002824) <!-- ch-version-info:start --> ### Version info - Merged into: `26.5.7.46` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114528",
          "createdAt": "2026-08-12T18:07:36Z",
          "updatedAt": "2026-08-13T01:35:58Z",
          "timestamp": "2026-08-13T01:35:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "SmitaRKulkarni",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:45358ff2b1a7c8ad7abd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105429",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105429",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Release pull request for branch 26.5",
          "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches security-sensitive internal/DDL execution gating by replacing `query_kind` checks with a new server-set flag, which could affect ON CLUSTER/replication/backup behavior if mis-propagated. Also tightens Arrow/Native input validation, which may reject previously-accepted malformed inputs and impact ingestion edge cases. > > **Overview** > Introduces a new `Context` flag `is_ddl_or_on_cluster_internal` (server-set and non-spoofable) and switches multiple code paths from `ClientInfo::QueryKind::SECONDARY_QUERY` to this flag for *security-sensitive* decisions, including internal backup/restore gating, Replicated DB DDL handling, `ON CLUSTER`/UUID-macro allowances, and context creation in DDL/replication/system operations. > > Hardens Arrow ingestion by reading geo metadata from the Arrow schema, validating BinaryArray offsets/lengths against buffer bounds (and handling absent/empty buffers), and avoiding `mutable_data()` usage in geo parsing; adds regression tests for corrupted Arrow offsets and geo metadata. > > Improves error classification for Variant deserialization under Native format (reports `INCORRECT_DATA` instead of `LOGICAL_ERROR`), adds an integration test for malformed Native Variant payloads, and updates planner behavior to disable/forbid parallel replicas when `additional_table_filters` is used without `serialize_query_plan` (with new stateless tests). > > Minor: CI style job skips test-number gap checks on release/backport branches, MySQL protocol test fixes dotnet working dir, and release metadata (version/contributors) is updated. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 1a682e2727943363cb077a81015b9242ab666852. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105429",
          "createdAt": "2026-05-20T13:36:34Z",
          "updatedAt": "2026-08-13T00:53:43Z",
          "timestamp": "2026-08-13T00:53:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "release"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8eba08f566c8ec4a2e3d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113681",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113681",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Replace the per-bucket hash map in `timeSeries*ToGrid` with a sorted-append sample array",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replaced the per-bucket hash map inside the `timeSeries*ToGrid` aggregate functions with a flat sorted array of samples: sample ingestion becomes an O(1) append for in-order inputs (the overwhelmingly common case) and the per-bucket copy-and-sort at finalization is gone. ### Description The `timeSeries*ToGrid` functions kept each bucket's samples in an `absl::flat_hash_map<timestamp, value>`: every `add()` paid a hash-map emplace, and the order-dependent functions (`rate`, `increase`, `delta`, `changes`, `resets`) copied and sorted every bucket at finalization. Yet the input is almost perfectly ordered — samples come from MergeTree tables sorted by `(id, timestamp)`; an instrumented probe on a 32-thread read of a 62.5-billion-sample table counted **1 out-of-order add in 1,474,559,998**. The bucket is now a flat, memory-tracked vector of `(timestamp, value)` pairs: O(1) append while timestamps ascend, in-place max on an equal timestamp, and a rare out-of-order add just clears a `sorted` flag — normalization (sort + max-dedup) runs lazily, only for buckets that actually saw disorder. `merge()` is a linear merge of sorted runs with an append fast path for disjoint time ranges. `forEachSample` now guarantees ascending order, so the copy-and-sort buffers are deleted from the rate/delta/changes aggregators. The wire format and `FORMAT_VERSION`s are unchanged; `deserialize()` assumes no order of incoming pairs (old peers send hash-map iteration order), so mixed-version clusters interoperate — verified in both directions. Duplicate timestamps keep the larger value with the old `std::max` argument order. Measured on the 62.5-billion-sample table: a 30-day `sum by(...)(rate(...))` over 25,600 series drops ~6% of total query CPU (1106 s -> 1037 s); wall time and peak memory move little (the scan dominates the critical path, and raw sample storage dominates the state either way). The structural point is what this enables: the sorted buffer is the prerequisite for O(1) per-segment summaries in the rate family (follow-up), which is where the ~44 GiB state peaks of such queries actually go away. All query results are fingerprint-identical. Tests: a new stateless test (shuffled and duplicate-timestamp inputs in both orders, NaN at duplicated timestamps, `-Merge` of unsorted in-memory states, interleaved parts, two-level merges, serialized-state merges through `remote('127.0.0.{1,2}', ...)`, an `AggregatingMergeTree` roundtrip, a fixed state literal in old-peer wire order); all 40 existing timeseries/PromQL stateless tests pass byte-identically; the perf test gains an ingestion-heavy scenario (50M rows, 10k series). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113681",
          "createdAt": "2026-08-06T13:46:38Z",
          "updatedAt": "2026-08-13T00:48:51Z",
          "timestamp": "2026-08-13T00:48:51Z",
          "metrics": {
            "reactions": 1,
            "comments": 6
          },
          "labels": [
            "pr-performance",
            "comp-promql"
          ],
          "author": "nikitamikhaylov",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a200226f3c9c694185a6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105714",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105714",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add INSERT ... RETURNING (non-atomic, user-supplied SELECT)",
          "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `INSERT ... RETURNING (SELECT ...)` to run a user-supplied SELECT after INSERT and return its result in one round-trip. Closes #21697. **Backward-incompatible change:** `RETURNING` is now a reserved keyword and can no longer be used as an unquoted implicit column alias (e.g. `SELECT 1 returning`). Use `AS returning` or a quoted identifier instead. --- Closes #21697 ## Summary Draft implementation of `INSERT ... RETURNING` for ClickHouse. The INSERT runs first using the existing pipeline; if it succeeds, a parenthesized `RETURNING (SELECT ...)` subquery runs in the same session and its result set is returned to the client in one round-trip. This is intentionally **not** Postgres-style atomic `RETURNING`. The user supplies the trailing SELECT and owns the `WHERE` clause used to identify inserted rows. ## Syntax For `INSERT VALUES` and `INSERT FORMAT`, `RETURNING` appears **before** the data clause (parser constraint: bytes after `VALUES`/`FORMAT` are raw input): ```sql INSERT INTO t (id, name) RETURNING (SELECT * FROM t WHERE id = 123) VALUES (123, 'foo'); INSERT INTO t RETURNING (SELECT count() FROM t WHERE batch_id = 'X') FORMAT JSONEachRow {\"batch_id\":\"X\", ...} ``` For `INSERT SELECT`, `RETURNING` appears **after** the source query: ```sql INSERT INTO t SELECT * FROM src RETURNING (SELECT count() FROM t WHERE batch_id = 'X'); ``` The parenthesized subquery is required (distinct from Postgres column-list `RETURNING id, name`). ## Semantics - INSERT runs first with its existing pipeline untouched. - If INSERT throws, the RETURNING SELECT does not run. - If INSERT succeeds, a select interpreter is built for the trailing subquery using the same Context; its pipeline is returned as the statement result. - The trailing SELECT inherits session, user, and settings. It can reference any table and use the full SELECT grammar. - Client receives **one result set** (the RETURNING SELECT's), not an INSERT ack plus a result set. ## Limitations (documented, not bugs) - **Not atomic.** Concurrent writers can land rows between INSERT completion and SELECT execution. - **No automatic \"the rows I just inserted.\"** User provides the SELECT and owns its `WHERE`. - **Replication lag is visible** on Replicated tables unless consistency settings are used. - **Incompatible with `async_insert=1`.** Returns `NOT_IMPLEMENTED`. - **No special MV interaction.** MVs fire on INSERT; RETURNING SELECT is a normal SELECT. - If INSERT succeeds but RETURNING SELECT fails, inserted data is **not rolled back**. ## Backward-incompatible change `RETURNING` is added to the parser's reserved-keyword list (same list as `FROM`, `WHERE`, `SETTINGS`, etc.) to prevent it from being consumed as an implicit table alias in `INSERT ... SELECT ... FROM tablename RETURNING (...)`. As a side-effect, `RETURNING` can no longer be used as an unquoted implicit alias anywhere in ClickHouse SQL: ```sql -- Before: worked (returning was a plain identifier) SELECT 1 returning SELECT * FROM t returning -- After: must use AS or quoting SELECT 1 AS returning SELECT * FROM t AS returning ``` ## Implementation - **Parser** (`ParserInsertQuery.cpp`): parse `RETURNING` + parenthesized SELECT; placement rules above. - **AST** (`ASTInsertQuery.h`): `returning_select` field; clone/format updated. - **Interpreter** (`InterpreterInsertQuery.cpp`): reject `async_insert`; wrap completed INSERT pipeline with RETURNING via `DelayedSource`. - **Pipeline** (`buildInsertReturningPipeline.cpp`): run INSERT to completion, then build SELECT pipeline; native push inserts swap to RETURNING SELECT after data is received. - **executeQuery** / **TCPHandler** / **LocalConnection** / **ClientBase**: inlined-data wrap, native-protocol push path, query-log accounting, process-list propagation. - **Docs**: section added to `insert-into.md` ## Tests added - `tests/queries/0_stateless/04266_insert_returning_21697.sql` — core `INSERT ... RETURNING` semantics, delayed planning/failure behavior, query-cache rejection paths, transactional rollback checks (`implicit_transaction=1` + explicit transaction), and parser keyword regression coverage (`SELECT 1 returning` rejected; explicit/quoted alias accepted). - `tests/queries/0_stateless/04267_insert_returning_native_format.sh` — native TCP `INSERT FORMAT` + external data + `RETURNING`, including `input(...)` + `FORMAT` placement/round-trip stability. - `tests/queries/0_stateless/04410_insert_returning_trailing_settings.sql` — source-side trailing `SETTINGS` scoping/precedence, nested source settings handling, source-side query-cache settings rejection (`NOT_IMPLEMENTED`), and a plain `INSERT ... SELECT` regression proving nested source subquery settings stay local (non-`RETURNING`). ## Alternatives considered - **`AND SELECT` / `THEN SELECT` keywords** — parser ambiguity or poor discoverability for Postgres migrants. - **Atomic Postgres-style RETURNING** — much larger surface (formats, MVs, distributed, async); deferred as a possible follow-up. - **Multi-statement `;`** — out of scope. ## Open questions for maintainers 1. Is the non-atomic, user-supplied-SELECT model the right v1 framing? 2. Confirm canonical placement rules (before data for VALUES/FORMAT, after source for INSERT SELECT). 3. Restrict trailing SELECT to read-only queries only? (Current implementation: yes, via SELECT parser.) 4. Confirm single-result-set protocol behavior across HTTP/native/JDBC. ## Test plan - [x] Stateless test `04266_insert_returning_21697` - [x] `INSERT VALUES` + `RETURNING` with row filter - [x] `INSERT SELECT` + `RETURNING` with aggregate - [x] `RETURNING` against a table other than INSERT target - [x] INSERT failure prevents `RETURNING` subquery from running - [x] `NOT_IMPLEMENTED` when `async_insert=1` - [x] Native TCP `INSERT FORMAT` + `RETURNING` (`04267_insert_returning_native_format`) - [x] `input(...)` source with `FORMAT` before `RETURNING` (parser/formatter round-trip + execution) - [x] Source-side `SETTINGS` precedence and nested-source scoping (`04410_insert_returning_trailing_settings`) - [x] Plain `INSERT ... SELECT` keeps nested source subquery `SETTINGS` local (no non-`RETURNING` leakage) - [x] Source-side query-cache settings rejection in `INSERT ... RETURNING` source `SETTINGS` - [x] `02310_profile_events_insert` local `clickhouse-local` profile-events check is non-fatal in CI (`InsertedRows` is not emitted consistently there) - [x] Transactional rollback when delayed `RETURNING` planning fails (`implicit_transaction=1` and explicit transaction) - [x] `RETURNING` keyword parser regression (`SELECT 1 returning` rejected; `SELECT 1 AS returning` and quoted alias accepted) - [x] `04266` rollback subtests stabilized by avoiding fixed database-name collisions in flaky/parallel runs (use `t_insert_returning_tx` in the current test DB) - [ ] Replicated-table non-atomic behavior (integration test, future)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105714",
          "createdAt": "2026-05-24T01:08:43Z",
          "updatedAt": "2026-08-13T00:44:27Z",
          "timestamp": "2026-08-13T00:44:27Z",
          "metrics": {
            "reactions": 2,
            "comments": 32
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "tylerhannan",
          "state": "open",
          "assignees": [
            "nikitamikhaylov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5755b6bd7499a94d54d8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:93833",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:93833",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "TimeSeries: enable dynamic remote-write routing via URL table prefix",
          "text": "This PR addresses https://github.com/ClickHouse/ClickHouse/issues/93831 ### Changelog category (leave one): - Improvement ### Changelog entry Support dynamic table routing for Prometheus `remote-write` based on the request URL, allowing a single handler to ingest data into multiple TimeSeries tables while preserving backward compatibility with fixed-table configuration. ### Documentation entry for user-facing changes - [x] Documentation is written #### Motivation In large-scale observability deployments, users often maintain many TimeSeries tables with similar schemas but different retention or ownership. The existing Prometheus `remote-write` integration requires one handler per table, with the table name hard-coded in `config.xml`, making configuration and operations difficult to scale. This change allows routing the target TimeSeries table dynamically from the request URL, enabling a single `remote-write` handler to serve multiple tables while keeping the previous behavior unchanged by default. #### Parameters - **`enable_table_name_url_routing`** (handler option, default: `false`) When enabled, the handler expects URLs in the form `/{database}/{table}/...` and extracts the target table from the first two path segments. - **`prometheus_remote_write_dynamic_routing_enabled`** (TimeSeries table setting, default: `false`) Enables dynamic table routing for the specified TimeSeries table. Requests targeting tables without this setting enabled will be rejected. #### Example usage ```xml <handlers> <remote_write> <url>regex:^/[^/]+/[^/]+/write$</url> <handler> <type>remote_write</type> <enable_table_name_url_routing>true</enable_table_name_url_routing> </handler> </remote_write> </handlers> ``` Request example: `POST /{db_name}/{table_name}/write` Enable dynamic routing on the target TimeSeries table: ```sql ALTER TABLE db.ts MODIFY SETTING prometheus_remote_write_dynamic_routing_enabled = 1; ``` <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds request-path-driven table selection to the Prometheus `remote_write` handler and gates writes with a new per-table setting, so misconfiguration or unexpected URLs could redirect writes or cause new 4xx/5xx failures. > > **Overview** > Enables Prometheus `remote_write` to **dynamically route writes to different `TimeSeries` tables** by extracting `{database}/{table}` from the request URL when `enable_table_name_url_routing` is set on the handler. > > Dynamic routing is **opt-in and additionally gated per target table** via new `TimeSeries` setting `prometheus_remote_write_dynamic_routing_enabled` (default `false`); requests to tables without it error out. Docs and integration tests were updated with the required URL regex and validation coverage for both enabled/disabled cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit d0e40205e663a54acd9edfa31186b6790c8f5443. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/93833",
          "createdAt": "2026-01-09T16:10:59Z",
          "updatedAt": "2026-08-13T00:37:35Z",
          "timestamp": "2026-08-13T00:37:35Z",
          "metrics": {
            "reactions": 2,
            "comments": 14
          },
          "labels": [
            "pr-improvement",
            "manual approve",
            "can be tested",
            "comp-promql"
          ],
          "author": "niyue",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d519393f01eb30141295",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106006",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106006",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix LOGICAL_ERROR in DatabaseReplicatedDDLWorker with max_replication_lag_to_enqueue=0",
          "text": "Reject `max_replication_lag_to_enqueue = 0` at parse time everywhere it can be supplied. `0` was never useful: it asks the post-recovery unsynced check `max_log_ptr + max_replication_lag_to_enqueue <= new_max_log_ptr` to be trivially true (the ZooKeeper counter is monotonically non-decreasing), which mis-flags a caught-up replica as unsynced and then trips `chassert(our_log_ptr < max_log_ptr)` in `DatabaseReplicatedDDLWorker::initAndCheckTask`, aborting the server on debug and sanitizer builds. Switching the setting type from `UInt64` to `NonZeroUInt64` lets the parser reject `0` from every source: `CREATE DATABASE ... SETTINGS`, the `<database_replicated>` server-config block, `ATTACH` replay, and upgrade-time replay of existing database metadata. The post-recovery comparison stays at the original `<=` because the bad value can no longer reach the worker. The smallest valid value is `1`. New regression test `04298_98823_replicated_zero_lag_assertion` checks both the rejection (`A setting's value has to be greater than 0`, BAD_ARGUMENTS) and the smallest valid value (`1`) going through the post-recovery DDL path that previously aborted. Adapted from the analysis in #100043. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reject `max_replication_lag_to_enqueue = 0` on `Replicated` databases. The value made the post-recovery unsynced check trivially true and aborted the server with `LOGICAL_ERROR` in debug and sanitizer builds. `0` is now rejected at parse time with `BAD_ARGUMENTS` from every source (`CREATE DATABASE ... SETTINGS`, `<database_replicated>` server-config block, `ATTACH` replay, and upgrade-time replay of existing metadata). The smallest valid value is `1`. Closes #98823. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106006",
          "createdAt": "2026-05-27T19:40:08Z",
          "updatedAt": "2026-08-13T00:37:31Z",
          "timestamp": "2026-08-13T00:37:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 20
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:435406e519e2e5cfe9c3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:102499",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:102499",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Added Parquet Shredded VARIANT Support to ParquetReaderv3",
          "text": "This PR introduces the parquet Shredded VARIANT standard to CLickHouse. duckdb has this already and the main benefit is that it can read less bytes from disk. This is of course critical for a main parquet use case, which is remote reads from S3. Benchmark results for JSONBench reading from file() on NVME on the final binary are below. | query | method | wall s | user CPU s | OSReadBytes | |---|---|---:|---:|---:| | q1_collection_counts | parquet_variant_shredded | 0.330 | 0.08 | 4.1 MB | | | parquet_variant_unshredded | 0.620 | 0.69 | 133.9 MB | | | parquet_string | 0.750 | 1.25 | 133.6 MB | | | parquet_json | 4.830 | 4.57 | 131.8 MB | | q2_collection_users | parquet_variant_shredded | 0.420 | 0.22 | 17.4 MB | | | parquet_variant_unshredded | 0.660 | 0.83 | 133.7 MB | | | parquet_string | 0.830 | 2.08 | 133.3 MB | | | parquet_json | 4.910 | 5.26 | 131.2 MB | | q3_hourly_events | parquet_variant_shredded | 0.370 | 0.13 | 9.4 MB | | | parquet_variant_unshredded | 0.640 | 0.78 | 133.7 MB | | | parquet_string | 0.810 | 1.93 | 133.3 MB | | | parquet_json | 4.730 | 4.64 | 131.2 MB | | q4_first_posts | parquet_variant_shredded | 0.400 | 0.13 | 19.2 MB | | | parquet_variant_unshredded | 0.670 | 1.13 | 133.3 MB | | | parquet_string | 0.800 | 1.93 | 133.0 MB | | | parquet_json | 4.910 | 5.23 | 130.9 MB | | q5_activity_span | parquet_variant_shredded | 0.400 | 0.15 | 19.7 MB | | | parquet_variant_unshredded | 0.670 | 1.06 | 133.6 MB | | | parquet_string | 0.780 | 1.92 | 133.4 MB | | | parquet_json | 4.840 | 5.00 | 131.2 MB | ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add Parquet shredded VARIANT support, including read/write paths and subcolumn-aware read optimizations for semi-structured data. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/102499",
          "createdAt": "2026-04-12T14:56:07Z",
          "updatedAt": "2026-08-13T00:35:07Z",
          "timestamp": "2026-08-13T00:35:07Z",
          "metrics": {
            "reactions": 3,
            "comments": 31
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested"
          ],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "scanhex12"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a43a2ec23341bf02261d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114523",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114523",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong results of the partial aggregation strategy in distributed query plans",
          "text": "The partial+merge aggregation strategy of `make_distributed_plan` rewrites a final aggregation into a partial `AggregatingStep` plus a memory-efficient `MergingAggregatedStep`. The memory-efficient merge consumes each input as a stream of two-level buckets in ascending order, but the rewrite kept `should_produce_results_in_order_of_bucket_number = false` on the partial step, so its multi-stream output was combined into the single exchange stream in arbitrary order. When a bucket arrived after the merge had already emitted it, the merge emitted it a second time: duplicated `GROUP BY` keys with split aggregate states (for ClickBench Q08, a wrong `COUNT(DISTINCT UserID)`). The fix, one commit each: - Make the row count estimation look through `LogicalExchangeStep`. The `#if CLICKHOUSE_CLOUD` guard around it was needed when the exchange steps existed only in the private repo; since they are in the public repo, the guard only made every aggregation over a distributed read fall back to the Shuffle strategy in non-cloud builds (and masked this bug there). - Validate the bucket delivery order in `GroupingAggregatedTransform`: an unannounced late bucket now fails with a `LOGICAL_ERROR` exception instead of silently duplicating groups. - Wait for delayed buckets in `GroupingAggregatedTransform` when the consumer needs the output in bucket order, and release the announcements of finished inputs. Covered by a gtest. - Build the partial step with `should_produce_results_in_order_of_bucket_number = true` when the merge is memory-efficient (`AggregatingStep::cloneAsPartial`). The partial aggregation then produces a single bucket-ordered stream per worker, the same contract the classic distributed path establishes on shards. Covered by a stateless test that reproduces the duplicated keys on the code without the fix (about 80% of single runs, 8 repetitions, 50k rows, ~0.5s). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix duplicated `GROUP BY` keys in the partial aggregation strategy of `make_distributed_plan`, and allow this strategy for aggregations over a distributed read. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114523",
          "createdAt": "2026-08-12T16:51:05Z",
          "updatedAt": "2026-08-13T00:31:42Z",
          "timestamp": "2026-08-13T00:31:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3782b8bf92417b6d982d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111770",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111770",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Disable uniq, uniq_v2 for high cardinality column types",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111291 This PR adds optional all-distinct materialization for automatic cardinality statistics. It is intended for column types where building a full cardinality sketch can be expensive and the data is commonly close to unique. New user-visible MergeTree settings: - `auto_statistics_assume_floats_distinct` - `auto_statistics_assume_long_strings_distinct` - `auto_statistics_long_string_distinct_min_length` - `auto_statistics_long_string_distinct_probe_rows` The settings are disabled by default. When enabled, eligible automatic `uniq` / `uniq_v2` statistics can be represented by an `assumed_all_distinct` implementation that estimates cardinality from non-NULL row counts instead of scanning all values into a distinct sketch. ### Implementation notes The implementation keeps the user-facing policy separate from the low-level statistics mechanics: - `StatisticsAssumedAllDistinct` is a lightweight `IStatistics` implementation that stores only cardinality. - `StatisticsUniqStringProbe` owns the stateful probing logic for long String / FixedString columns. - `UniqAssumedAllDistinctPolicy` centralizes the build/merge decisions, replacement logic, and compatibility rules. - `ColumnStatistics` remains responsible for orchestration only: it asks the policy for optional build/merge decisions and applies them. - The merge path now uses an optional `AssumedAllDistinctMergeDecision`, matching the existing `decideBuild` style and avoiding a disabled “plan” state. ### Tests Added coverage for: - Float columns with the assumption enabled/disabled. - Long and short String columns. - Nullable values. - Explicit `assumed_all_distinct` materialization. - Materialization on insert, materialization via mutation, and merge behavior. ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added MergeTree settings to materialize automatic `uniq`/`uniq_v2` statistics for Float and long String columns with an `assumed_all_distinct` cardinality model.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111770",
          "createdAt": "2026-07-24T10:54:30Z",
          "updatedAt": "2026-08-13T00:31:40Z",
          "timestamp": "2026-08-13T00:31:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "cv4g",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f503ba3f476b654c27d9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113553",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113553",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix data race when tracing profile events",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a theoretical data race in profile events when using the `trace_profile_events_list` setting to write stack traces of certain profile events to `system.trace_log`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113553",
          "createdAt": "2026-08-05T18:58:10Z",
          "updatedAt": "2026-08-13T00:30:12Z",
          "timestamp": "2026-08-13T00:30:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-must-backport",
            "pr-critical-bugfix"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8089da79898ea55d93c5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:76595",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:76595",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallel Replicas: a setting to disable for queries with multiple tables",
          "text": "Queries with a `JOIN` read the non-leftmost side in full on every replica, which can make parallel replicas slower than a plain single-node execution. This adds a kill switch so such queries can be excluded from parallel replicas without disabling parallel replicas altogether. New setting `parallel_replicas_for_queries_with_multiple_tables` (default `true`, i.e. the pre-existing behaviour). When set to `false`, parallel replicas are not used for a query that joins multiple tables: the decision is applied in `findParallelReplicasQuery` (query/table selection) and in `buildJoinTreeQueryPlan`, and it is propagated into subquery table expressions — including `UNION` table expressions and their branches, which are planned by independent `Planner` instances — and into the `IN` subqueries collected into prepared sets, which are also planned by independent `Planner` instances built from the subqueries' own contexts. The legacy (pre-analyzer) interpreter respects the setting as well: when `parallel_replicas_only_with_analyzer = 0` allows task-based parallel replicas there, `InterpreterSelectQuery` applies the same kill switch before the storage read. `ARRAY JOIN` does not count as a join between tables, and a `UNION` query without a `JOIN` is not affected: each `UNION` branch is an independent single-table read, so parallel replicas remain applicable to it. ### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Added the setting `parallel_replicas_for_queries_with_multiple_tables` (default `true`) to control whether parallel replicas are used for queries joining multiple tables. When disabled, parallel replicas are not used for queries with `JOIN`, where the non-leftmost side is read in full on every replica. The setting does not affect a `UNION` query without a `JOIN` (each branch is an independent single-table read) nor `ARRAY JOIN`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/76595",
          "createdAt": "2025-02-21T17:53:10Z",
          "updatedAt": "2026-08-13T00:28:54Z",
          "timestamp": "2026-08-13T00:28:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 14
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "devcrafter",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:51a69c43d4cc30e424db",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114423",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114423",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #113534 to 26.7: Fix a mixed JOIN ON condition evaluated over mismatched column types for a dictionary",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113534 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31567499495/job/94022263699)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114423",
          "createdAt": "2026-08-12T06:02:32Z",
          "updatedAt": "2026-08-13T00:26:07Z",
          "timestamp": "2026-08-13T00:26:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-clickhouse",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:531f3e9a0cf1a1ae1efd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114054",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114054",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Use libdeflate for gzip/zlib/deflate compression and decompression\"",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114045 Related: https://github.com/ClickHouse/ClickHouse/pull/108074 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Revert #108074 (use libdeflate for gzip/zlib/deflate): its streaming decompressor suspends only at DEFLATE block boundaries, so decompressing an HTTP request body whose single DEFLATE block spans the whole stream (the shape zlib-ng produces at level 1, which is the default in the official .NET SDK) was `O(n^2)` in the compressed block size and buffered the whole block in memory. A 22 MB gzip body that 26.6 ingested in 0.3 s took 15 s on 26.7. This returns the gzip/zlib/deflate paths to zlib-ng, restoring 26.6 behavior. The `gzip`/`deflate` compression levels `10-12` that 26.7 introduced stay accepted and are clamped to `9`, the maximum level supported by `zlib`, so configurations written against 26.7 keep working. ### Details This is a plain `git revert` of the merge commit of #108074, with the following manual adjustments: - `tests/queries/0_stateless/04230_iceberg_optimize_metadata_not_initialized_104711.sh` and `04240_iceberg_supports_parallel_insert_mv_104891.sh` keep their current (post-#108074) form: they now corrupt Iceberg metadata files directly instead of relying on `output_format_compression_level=11` being rejected, which works regardless of the compression backend. - The relocated doc `docs/reference/statements/select/into-outfile.mdx` documents the compression level range as `1-12` for `gzip`/`deflate` with the clamp, since the old `docs/en` tree no longer exists for the revert to apply to. The text for `http_zlib_compression_level` is changed in the `DECLARE` docstring in `src/Core/Settings.cpp`, because `docs/reference/settings/session-settings/http.mdx` is inside an `AUTOGENERATED` region and hand-editing it fails the `No direct edits to generated or read-only docs` check. **Backward compatibility.** 26.7 shipped `libdeflate` and, with it, `gzip`/`deflate` compression levels `10-12` on every output surface that routes through `CompressionMethod`: `INTO OUTFILE ... LEVEL`, `output_format_compression_level`, `http_zlib_compression_level` and the gRPC `output_compression_level`. A plain revert would narrow the accepted range back to `1-9`, so existing queries and settings profiles would start failing with `Invalid compression level` or `deflateInit2 failed: stream error`. To avoid that, `getCompressionLevelRange` keeps returning `1-12` for `gzip`/`zlib` and `createWriteCompressedWrapper` clamps levels above `9` to zlib's maximum. Levels above `12` are still rejected, as in 26.7. Two tests are restored in a backend-agnostic form instead of being deleted with the feature: - `04842_parquet_gzip_compression_level` round-trips a mixed-compressibility dataset through Parquet with `output_format_parquet_compression_method='gzip'` at levels `1, 3, 6, 9, 12`, keeping non-default-level coverage of `output_format_compression_level` for Parquet. - `04843_outfile_compression_level_range` covers the accepted `INTO OUTFILE ... LEVEL` range for `gzip`/`deflate`, the clamp of `10-12`, and the rejection of `13`. A follow-up PR will re-introduce libdeflate with the streaming decompressor fixed to suspend at symbol granularity (linear time, bounded memory).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114054",
          "createdAt": "2026-08-09T17:34:24Z",
          "updatedAt": "2026-08-13T00:24:29Z",
          "timestamp": "2026-08-13T00:24:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "submodule changed",
            "v26.7-must-backport"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1e596887b90a9fa980b0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114057",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114057",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Convert test 04630_merge_over_stale_packed_tmp_dir to an integration test",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113978 Related: https://github.com/ClickHouse/ClickHouse/pull/111652 Test `04630_merge_over_stale_packed_tmp_dir` (added in https://github.com/ClickHouse/ClickHouse/pull/111652) copies a packed part directory with plain `cp` to simulate a stale `tmp_merge_` directory left by an interrupted merge. Modifying table data on disk is not allowed in stateless tests at all: the stateless suite runs against arbitrary server configurations (object storage, shared merge tree, encrypted disks), where direct filesystem manipulation is either meaningless or destructive — this test corrupted shared S3 blob reference counts in stress runs (https://github.com/ClickHouse/ClickHouse/pull/113978#pullrequestreview-4891978370) before it was pinned to the local disk in https://github.com/ClickHouse/ClickHouse/pull/113978. This PR moves the scenario to the integration test `test_packed_io::test_merge_over_stale_packed_tmp_dir`, where the cluster environment is fully controlled and the part directory can be copied safely inside the container. The test logic and all assertions are unchanged: the merge must reclaim the leftover `tmp_merge_all_1_2_1` directory (verified via the `Removing stale temporary directory` log message and the directory being gone), must not seed any data from it into the new part, and the resulting packed part must pass `CHECK TABLE`. Verified locally: the new integration test passes. A style check preventing new stateless tests from modifying the server's data directory is added separately in https://github.com/ClickHouse/ClickHouse/pull/114056. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114057",
          "createdAt": "2026-08-09T17:59:57Z",
          "updatedAt": "2026-08-13T00:20:37Z",
          "timestamp": "2026-08-13T00:20:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:81a81a86dd57a2ae4752",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110594",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110594",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add aiFilter for natural-language boolean filtering via LLMs.",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/110352 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `aiFilter` function that evaluates a natural-language condition against text with an LLM and returns `UInt8` for use in `WHERE`, `PREWHERE`, and `JOIN ... ON`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110594",
          "createdAt": "2026-07-15T16:49:31Z",
          "updatedAt": "2026-08-13T00:19:39Z",
          "timestamp": "2026-08-13T00:19:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-feature",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "ylw510",
          "state": "closed",
          "assignees": [
            "george-larionov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a4098fee74f119013695",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114085",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114085",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Delay background mutations by a bounded random amount in stress tests",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113925 Related: https://github.com/ClickHouse/ClickHouse/pull/113225 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Description A part written before an `ALTER` keeps its own older type (or misses the column) until the background mutation rewrites it. Reads over such parts are a distinct code path: columns resolve with the part's type while metadata and subcolumn entries carry the new one. Our tests pin their results with `mutations_sync = 2`, so the corpus systematically avoids this state, and in ordinary runs background mutations close the window in milliseconds. That is how the regression fixed by #113925 passed the full pre-merge CI of #113225: the bug lived only in the pending-mutation state, which nothing exercised until an unrelated test added the fixture two days later. This PR makes the stress test hold that window open routinely: - New failpoint `mutate_task_random_sleep_in_prepare`: a random 0-3 s sleep at the top of `MutateTask::prepare`, covering plain, replicated and shared `MergeTree` mutations (all funnel through `MutateTask`). - `stress.py` enables it via `SYSTEM ENABLE FAILPOINT` right after the smoke check (never in upgrade check, where the old binary may not know the failpoint) and disables it in `prepare_for_hung_check` next to `SYSTEM STOP THREAD FUZZER`, so pending mutations drain at full speed before the hung check. Disabling a registered failpoint that was never enabled is a documented no-op, so the cleanup is idempotent. With the failpoint on, every test that ALTERs without waiting reads unrewritten parts under the stress runner's random settings and the server-side AST fuzzer's query mutations - the exact combination that found the #113925 crash within four hours once a fixture existed, applied now to the whole corpus on every stress run. **Validation.** Built and measured on a `RelWithDebInfo` build: a `mutations_sync = 2` `UPDATE` takes 0.12 s with the failpoint off, 0.65-1.67 s with it on, and is fast again after `SYSTEM DISABLE FAILPOINT` (including a second, no-op disable). `stress.py` passes `python3 -m py_compile`. The delay is bounded, so `mutations_sync = 2` tests finish at most a few seconds later per mutation, and `prepare_for_hung_check` removes the delay entirely before the hung check. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1288` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114085",
          "createdAt": "2026-08-10T00:07:13Z",
          "updatedAt": "2026-08-13T00:19:34Z",
          "timestamp": "2026-08-13T00:19:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ea3ded2d405d5f844862",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110283",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110283",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Compose join-order statistics over parts surviving partition/PK pruning",
          "text": "Closes: [https://github.com/ClickHouse/ClickHouse/issues/110281](<https://github.com/ClickHouse/ClickHouse/issues/110281>) ### Changelog category (leave one): * Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](<https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md>) of the changes that goes into CHANGELOG.md): Join-order cardinality estimation now composes column statistics over the parts surviving partition/PK pruning instead of all active parts. Previously a query pruned to a small partition was planned against table-wide statistics (observed 2500x row overestimation flipping the hash-join build side), and disabling statistics paradoxically produced a better plan. Also fixed: lazy FINAL was silently disabled for MergeTree relations under a JOIN because the join-order optimizer's memoized index-analysis result was mistaken for an applied projection. ### Documentation entry for user-facing changes Not needed: no user-visible interface changes; the fix makes existing settings (`use_statistics`, `use_statistics_cache`, `query_plan_optimize_lazy_final`) behave as documented. --- **Symptom.** With `use_statistics = 1` (default), `estimateReadRowsCount` in `optimizeJoin.cpp` estimated relation sizes over **all** active parts, ignoring partition/PK pruning. On the reproducer from ClickHouse/ClickHouse#110281 (fact: 5M rows in `p=1` + 1k rows in `p=2`, dim: 100k, `WHERE p = 2`): fact side estimated at 2,500,500 rows (= 5,001,000 × 1/NDV(p), a 2,500× error; NDV error 250,000×), identical under `use_statistics_cache = 0/1`. With statistics *disabled* the index-based fallback calls `selectRangesToRead()` and estimates 1,000 rows exactly — i.e. enabling statistics made the plan \\~100× worse by the optimizer's own cost model. **Root cause.** At `optimizeJoin` time `analyzed_result_ptr` is not populated yet, so `ReadFromMergeTree::getParts()` falls back to `prepared_parts` (all parts) when building the `ConditionSelectivityEstimator`. The table-wide `cached_estimator` (populated by the background `refreshStatistics()` task over all active parts) matched that wrong scope, which masked the cache/query-scope divergence. **Fix.** 1. `estimateReadRowsCount` obtains the partition/PK analysis result up front (`getAnalyzedResult()` or `selectRangesToRead()` — exactly what the non-statistics fallback branch already did) and passes it to a new `getConditionSelectivityEstimator(required_columns, analyzed_result)` overload, so statistics are folded over `parts_with_ranges`. 2. `MergeTreeData::getConditionSelectivityEstimator` returns the table-wide cached estimator only when the requested part set matches the set the cache was built over (allocation-free `isStale(RangesInDataParts)` overload; the shared_ptr is copied under `stats_mutex`, the comparison runs outside — the estimator is immutable once published). Without this, fixing (1) would make the previously-masked cache scope divergence real. 3. `optimizeLazyFinal`: the \"projection was applied\" guard checked for a mere non-null analysis result. `selectRangesToRead()` memoizes its result on the reading step, so join-order estimation (the no-statistics fallback before this PR, the statistics path as well after it) made lazy FINAL silently bail for any MergeTree relation under a JOIN. The guard now checks `readFromProjection()`. 4. `optimizeLazyFinal` also re-ran `selectRangesToRead()` unconditionally after the guard, repeating the full part/PK/skip-index analysis that join-order estimation had already memoized (observed: two `SelectExecutor` \"Key condition\" passes over the same parts in one stats-enabled `ReplacingMergeTree FINAL JOIN` query). It now reuses the memoized result — analysis passes drop 2 → 1; PK conditions are pushed in the first optimization pass, before both consumers, so the repeated analysis was provably identical. 5. Early index-analysis exits that prove the read empty now preserve the exact-zero invariant; join ordering short-circuits to zero rather than degrading to unknown when no estimator exists (`WHERE 0`: `f ⋈ d` before, `f[0] ⋈ d[0]` after). 6. With throwing `max_rows_to_read` or `max_rows_to_read_leaf`, join estimation uses a non-memoized range analysis without row-limit checks. The final read repeats analysis after `optimizeReadInOrder`; it remains limited unless InOrder is selected. **Scope note.** Statistics are composed over surviving *whole parts*; mark-range granularity inside a surviving part is not used (follow-up material, see the issue). Filtering by skip indexes deferred to the scan by `use_skip_indexes_on_data_read = 1` is likewise not reflected in join-order estimates: join order is fixed before those results exist. The fix relies only on partition/PK analysis remaining pre-scan. **Verification** (macOS arm64 debug build @ `5a9528b4db5`): * Reproducer from the issue: fact estimate 2,500,500 → **1,000** rows under both cache settings; the pruned query's plan cost becomes bit-identical (996.86) to the physically-pruned counterfactual table; build side flips to the small side. * Unpruned query with a warm cache still hits the cache (0 `Loading statistics` events) — no regression from the part-set check. * `04516_join_order_estimation_pruned_parts` (fails before the fix: `f[50500]` on 26.7.1.448; passes after: `f[1000]`), with an NDV oracle probe (`f[100]` = 1000 × 1/NDV(id) — derivable only from pruned column statistics, the index fallback yields `f[1000]`, so a silently degraded statistics path fails the test) and a PK-only pruning case (no partitioning; a part fully excluded by the primary key). * `04517_lazy_final_join_with_statistics` (new): `InputSelector` present for single-table FINAL and for FINAL under a JOIN with statistics on and off, plus result correctness. * `04518_statistics_cache_pruned_scope` (new): warms the table-wide cache via the real background refresh (deterministic retry on `LoadedStatisticsMicroseconds = 0`, the same oracle as `03707_statistics_cache`), then asserts: pruned query bypasses the cache and gets pruned-scope estimates; the unpruned cache hit is preserved afterwards. * Locally green: lazy_final suite (`03990`, `03991`; `03988`/`04092`/`04093` skipped as `no-debug`), `03707_statistics_cache`, `03279_join_choose_build_table_{,auto_}statistics`, `03788_statistics_part_pruning*`, `02864_statistics_*`, explain-pretty/join-order/estimate suites. **Cost/risk notes.** * The statistics path normally runs index analysis at planning time and **memoizes it** on the reading step (`selectRangesToRead()` writes `analyzed_result_ptr`), so execution reuses the analysis instead of re-running it. This matches what the no-statistics fallback already did before this PR; the new part is that the statistics path joins that behavior. Throwing read limits are the deliberate exception: estimator analysis stays local until `optimizeReadInOrder` determines whether the final read is exempt. * Pruned queries bypass the table-wide estimator cache and fold per-part statistics per query (`Loading statistics`). A part-set-keyed or per-part-decoded statistics cache is follow-up work (discussed in ClickHouse/ClickHouse#110281). * `tests/performance/join_planning_pruned_statistics.xml` (new): planning-only `EXPLAIN` benchmarks over 100- and 1000-part fact tables — per-part fold (cold cache, selective/unselective), a *genuinely warm* table-wide cache (`refresh_statistics_interval = 1` + warm-up; verified: zero statistics loads on the unpruned hit) exercising the O(parts) part-set comparison and the pruned bypass, and a planning-only lazy FINAL join over a 200-part ReplacingMergeTree guarding the duplicate-analysis regression. An end-to-end execution query is kept as a separate benchmark. CI perf compares merge-base vs head on a Linux release build. * If the pre-analysis pass ordering was a deliberate trade-off, happy to hear maintainer context — the issue discusses this explicitly.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110283",
          "createdAt": "2026-07-13T15:11:51Z",
          "updatedAt": "2026-08-13T00:19:28Z",
          "timestamp": "2026-08-13T00:19:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v26.4-must-backport"
          ],
          "author": "skuznetsov-clickhouse",
          "state": "closed",
          "assignees": [
            "fkastrati"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3269d9ccf30936818d67",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113947",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113947",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "MySQL: reject an empty TLS contents override also when the collection stores contents",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112070 Related: https://github.com/ClickHouse/ClickHouse/pull/110615 The empty-override rejection in `StorageMySQL::getSSLParams` only ran when the base named collection stored a credential *path*: with the credential stored in the contents form (`ssl_ca_pem` / `ssl_cert_pem` / `ssl_key_pem`), `get_path` returned at the empty-path fast path before looking at `isQueryOverridden`, so a query could pass `ssl_ca_pem = ''` and silently strip the collection-provided CA or client certificate — e.g. disable the verification of the server certificate — on every MySQL surface using that collection. The check is now hoisted above the fast path, so an empty contents override is rejected regardless of the form the stored credential has. The same defect in the PostgreSQL counterpart was found by the AI review in #110615, where it is fixed the same way; this is the twin fix for the MySQL surface introduced in #112070. The existing empty-override test gains the contents-storing-collection case. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a defect in the `MySQL` integrations where overriding a TLS credential of a named collection with an empty `ssl_ca_pem`/`ssl_cert_pem`/`ssl_key_pem` value was accepted when the collection stored the credential in the contents form, silently dropping the configured CA or client certificate instead of rejecting the override. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1289` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113947",
          "createdAt": "2026-08-08T14:20:06Z",
          "updatedAt": "2026-08-13T00:19:22Z",
          "timestamp": "2026-08-13T00:19:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:475e296c28c178179b89",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114495",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114495",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Targeted tests: skip previously failed tests that no longer exist",
          "text": "The targeted functional-test jobs replay tests that failed in previous CI runs of the PR (a 30-day CIDB window). When such a test is deleted or renamed on master in the meantime — e.g. a flaky test removed instead of deflaked — its name still comes back from CIDB, `clickhouse-test` matches zero tests, prints `No tests were run.` and exits with code 1, turning the whole job red for weeks on every PR that ever saw the test fail. The same failure mode existed for orphan data files and was fixed in #104097; this is the deleted-test variant. Seen on #110958: `Stateless tests (arm_asan_ubsan, targeted)` kept replaying `04648_geohashes_in_box_cancellation`, which was deleted from master on 2026-08-09 (d5bc95515d9e \"Remove recently added cancellation tests, they are flaky\"). CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110958&sha=7acb5ba0dddf5862fe1cd118b17899a272708202&name_0=PR&name_1=Stateless%20tests%20%28arm_asan_ubsan%2C%20targeted%29 The fix filters the CIDB result in `Targeting.get_previously_failed_tests` to tests that still exist in the checkout: a file with a known test extension under `tests/queries/0_stateless/` for stateless jobs, the test directory under `tests/integration/` for integration jobs. Skipped names are logged, not silently dropped. Unit tests added in `ci/tests/test_find_tests.py`. Related: https://github.com/ClickHouse/ClickHouse/pull/110958 Related: https://github.com/ClickHouse/ClickHouse/pull/104097 ### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1290` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114495",
          "createdAt": "2026-08-12T14:44:06Z",
          "updatedAt": "2026-08-13T00:19:18Z",
          "timestamp": "2026-08-13T00:19:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f1eed79e551b50de7c0c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114469",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114469",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep heavy check in `ColumnArray` only in debug build",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) Fixes #114105. Introduced in #112504. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1291` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114469",
          "createdAt": "2026-08-12T11:47:45Z",
          "updatedAt": "2026-08-13T00:19:13Z",
          "timestamp": "2026-08-13T00:19:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "CurtizJ",
          "state": "closed",
          "assignees": [
            "scanhex12"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a693789a5ed732e28e9d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110493",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110493",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix MaterializedPostgreSQL database/table with ON CLUSTER",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/58726 A `MaterializedPostgreSQL` database or table created with `ON CLUSTER` gets the *same* ClickHouse UUID on every replica, because the UUID is generated once on the initiator and shipped in the DDL. With `materialized_postgresql_use_unique_replication_consumer_identifier = 1` the replication slot name was derived from that UUID and the publication name from the PostgreSQL database/table name, so every replica computed the same names and fought over one PostgreSQL slot and publication: a single replica replicated, the others failed in the background with `pqxx::unique_violation` or could not hold the slot, leaving divergent data. The fix hashes the persistent per-server `ServerUUID` together with the object UUID (via `makeUUIDv4FromHash`, which keeps the result within PostgreSQL's identifier length limit) and derives both the slot and the publication name from it, so each replica gets its own pair and replicates independently. Both UUIDs are persistent, so `ATTACH` keeps reusing the same slot. The default behaviour (setting disabled) is unchanged, and since the change is in the shared `PostgreSQLReplicationHandler` it covers both the database and the table engine. A user-managed `materialized_postgresql_replication_slot` has one fixed name that every replica would share, so combining it with this setting is contradictory and is now rejected for a freshly supplied engine definition — a `CREATE`, or an `ATTACH` that carries a full definition, so it cannot be bypassed through `ATTACH ... ON CLUSTER`. A replay of an already stored definition (startup, `RESTORE`, short-syntax `ATTACH`) still accepts it, so existing deployments keep starting. ## Upgrade of existing deployments Salting the names is a rename of the generated identity, so `adoptLegacyReplicationIdentityIfNeeded` — the mechanism that already handles the schema-aware rename — now also adopts the pre-salt slot and publication when the salted ones do not exist. Replication resumes from the same slot position: no re-snapshot into the already-populated nested tables, nothing orphaned. The pre-salt slot name embeds the object's own ClickHouse UUID, so its existence proves ownership. The publication name is schema-blind and does not, so it is adopted only when it publishes *exactly* the tables this engine replicates: the configured `materialized_postgresql_tables_list`, or otherwise the tables the engine actually replicated in the previous run (their nested tables exist on disk) — not the live PostgreSQL schema, which may have grown since `CREATE DATABASE` with tables that are never replicated without an explicit `ATTACH TABLE`. Adoption is skipped for a database that never materialized a single nested table: there is nothing to preserve and its empty on-disk set cannot prove any publication's ownership, so it synchronizes under the current identity and the legacy objects are named in the log for the operator. ## Fail-closed attach paths `pgoutput` resolves publication membership from the historic catalog snapshot at each change's LSN, so streaming through a publication that was created or narrowed after those changes were written silently drops them — unrecoverably, once `confirmed_flush_lsn` advances. Attach therefore no longer papers over a broken PostgreSQL-side state; it fails with a retryable error, so startup keeps retrying and an operator can repair the conflict or rebuild the replica. This holds for every identity, not only for the pre-salt rename: - **Publication gone behind a surviving slot** — previously recreated silently, skipping every change committed while it was missing. - **Slot gone behind a surviving publication**, as after `pg_upgrade` (it keeps publications but not replication slots) or an operator dropping the slot — previously an in-place re-snapshot, whose rows are materialized with `_sign = 1` and `_version = 1` and therefore neither delete rows that disappeared from PostgreSQL nor override rows already at a higher version, silently leaving the replica stale. A clean rebuild recovers every row. A user-managed slot is excluded, since a missing one is already reported as a configuration error. - **Both gone** while the replica already holds data from a previous run. - **Publication drift under the current identity** — `ALTER PUBLICATION ... DROP TABLE`, an altered `publish` set (`CREATE PUBLICATION` defaults to `insert, update, delete, truncate`), or, on PostgreSQL 15 and newer, a row filter or a column list on a published table. The published set is compared by exact `(schema, table)` pair, because in the single-schema modes both the publication table list and the WAL consumer key tables by bare name, so a publication rewritten from `foo.a` to `bar.a` would otherwise resume and replay `bar.a`'s changes into the ClickHouse table for `foo.a`. Extra published tables stay tolerated — the identity already owns the publication name — unless a foreign-schema table's bare name collides with a replicated one. A database that has not materialized a single nested table is exempt throughout: there is nothing a snapshot could make stale, so it bootstraps through whatever survived. If an interrupted first synchronization left both the publication and the slot behind, the leftover slot is dropped so the initial synchronization can run instead of resuming and then failing forever with `UNKNOWN_TABLE` on a nested table that was never created (a user-managed slot, which this engine may not recreate, is refused with an explicit message). Relatedly, the set of tables to materialize on attach now comes from the existing publication — the engine's persisted table set, which the code always claimed to assume correct — instead of the live schema. Previously a PostgreSQL table created after `CREATE DATABASE` made every attach retry fail with `UNKNOWN_TABLE` on its missing nested table, so a whole-schema database could never resume replication after a restart; this reproduces on master with the setting disabled. Published-but-not-materialized tables are reported in a warning that names the recovery (`ATTACH TABLE`), because that state is indistinguishable from a table that was in the original definition but never completed its first snapshot; persisting the configured table set is deferred to the separate persistent-replication-state change, together with the `rebuild required` marker discussed in the review. ## Tests The new `test_postgresql_replica_database_engine_on_cluster` suite covers the fix — two-node `ON CLUSTER` databases and tables, each replica getting its own slot and publication and synchronizing independently, and both removed by `DROP ... ON CLUSTER` — the adoption of a reconstructed pre-salt identity (including a whole-schema database whose PostgreSQL schema has grown, and foreign publications that must not be adopted), every fail-closed case above, the bootstrap paths for never-synchronized runs and for a fresh full-definition `ATTACH TABLE`, and the rejection of a user-managed slot on `CREATE` and on a full-definition `ATTACH ... ON CLUSTER` of either engine. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `MaterializedPostgreSQL` databases and tables created with `ON CLUSTER`. When `materialized_postgresql_use_unique_replication_consumer_identifier` is enabled, the PostgreSQL replication slot and publication names are now unique per server, so replicas no longer collide on a shared replication slot and publication, which previously left all but one replica unable to replicate. Also fixed a whole-schema `MaterializedPostgreSQL` database failing to resume replication after a server restart when a new table had been created in PostgreSQL, and made the attach fail closed instead of silently losing changes when the publication or the replication slot is missing. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110493",
          "createdAt": "2026-07-15T02:39:06Z",
          "updatedAt": "2026-08-13T00:15:34Z",
          "timestamp": "2026-08-13T00:15:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:133f9e276c13eee5da19",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114551",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114551",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support quoted identifiers in PromQL selectors",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support quoted metric and label names in PromQL selectors. ### Description Prometheus 3 supports UTF-8 metric and label names. Prometheus's own end-to-end test uses selectors such as: - `{\"http.requests\", \"service.name\"=\"api-server\", instance=\"0\", group=\"canary\"}` AWS CloudWatch also documents quoted OpenTelemetry selectors such as: - `{\"http.server.active_requests\", \"@resource.service.name\"=\"myservice\"}` These names are also used in OpenTelemetry metrics and user-facing PromQL products. For example, SigNoz documents queries such as: - `{\"system.cpu.utilization\", \"service.name\"=\"frontend\"}` - `sum by (\"k8s.pod.name\") (rate({\"container.cpu.utilization\", \"k8s.namespace.name\"=\"ns\"}[5m]))` ClickHouse currently accepts only identifier-style tokens for selector names. A quoted selector identifier is rejected by the grammar before query evaluation. This change: - accepts quoted metric and label names in selectors - treats a standalone quoted selector name as an `__name__` matcher - validates and unquotes the identifier - keeps non-legacy names quoted when serializing the query tree References: - [Prometheus UTF-8 guide](https://prometheus.io/docs/guides/utf8/) - [Prometheus end-to-end UTF-8 selector test](https://github.com/prometheus/prometheus/blob/main/promql/promqltest/test_test.go#L166-L191) - [AWS CloudWatch PromQL examples](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-PromQL-Querying.html) - [OpenTelemetry HTTP metric conventions](https://opentelemetry.io/docs/specs/semconv/http/http-metrics/) - [OpenTelemetry Prometheus compatibility survey](https://opentelemetry.io/blog/2024/prometheus-compatibility-survey/) - [SigNoz PromQL UTF-8 guide](https://signoz.io/docs/userguide/write-a-prom-query-with-new-format/) ### Tests - Regenerated the ANTLR parser artifacts. - Added `gtest_PromQLParser` coverage for quoted metric names, quoted label names, and invalid identifiers. - Ran focused syntax checks for the generated parser and touched PromQL sources. - Ran focused grammar checks for quoted metric and label selectors.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114551",
          "createdAt": "2026-08-12T21:52:00Z",
          "updatedAt": "2026-08-13T00:09:52Z",
          "timestamp": "2026-08-13T00:09:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "submodule changed",
            "can be tested"
          ],
          "author": "fallintoplace",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:40e13f08dafd3fa31699",
        "signalId": "github:ClickHouse/ClickHouse:issue:80787",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:80787",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "MaterializedPostgreSQL DB Engine: Support for TLS",
          "text": "### Company or project name _No response_ ### Describe the unexpected behaviour As of now, it is impossible to connect to a PostgreSQL with enforced SSL connections, with or without verifying the certificates ### How to reproduce Setup a PostgreSQL with required SSL (e.g. with zalando postres operator) Try to create a MaterializedPostgreSQL Database Connection will fail - because clickhouse won't use SSL. ### Expected behavior I propose adding pgsslmode into the Connection creation and adding settings to set the certificates when pgsslmode=verify or pgsslmode=verify-full. In the code there's already a connection string builder, which should be capable of adding these settings. ### Error message and/or stacktrace _No response_ ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/80787",
          "createdAt": "2025-05-25T09:07:21Z",
          "updatedAt": "2026-08-13T00:06:52Z",
          "timestamp": "2026-08-13T00:06:52Z",
          "metrics": {
            "reactions": 1,
            "comments": 0
          },
          "labels": [
            "help wanted",
            "unfinished code"
          ],
          "author": "martin31821",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d1493321ce7ee6bdab18",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114131",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114131",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use a continuous primary-key range for whole-metric PromQL selectors of TimeSeries tables",
          "text": "A PromQL selector over a `TimeSeries` table filters the samples table with `id IN (SELECT id FROM tags WHERE <matchers>)`. For a metric with tens of thousands of series, `KeyCondition` runs its single-threaded generic exclusion search with the whole set: 284 ms per selector on a 62-billion-row part with 1.9M marks (503 ms at 8.1M marks), and rule-style queries evaluate up to 5 selectors. With the two-component id layout `Tuple(UInt64, UUID)` the canonical id generator derives the first component from the metric name alone, so all series of one metric form one continuous primary-key range. When a selector matches a whole metric — verified by metadata checks plus one `LIMIT 1` probe on the tags table, which also detects out-of-range ids left by an earlier `ALTER ... MODIFY SETTING id_generator` — the generated WHERE additionally carries `id >= tuple(hash(name), min) AND id <= tuple(hash(name), max)` and the inner query sets `use_index_for_in_with_subqueries_max_values = 1`. Index analysis uses the range; the `IN` stays for exact row filtering, so both emissions return identical rows on any data. Any failed check emits today's SQL unchanged. The range can select a few extra boundary granules (+5 of 135,005 marks on a 30-day scan). ## Measured effect tsbench PromQL suite: 62.455B samples / 361,432 series, 1.9M-mark part; Ryzen 9950X (16C/32T); baseline = clean master 9b6a2d7346f. Cold medians of 3 interleaved rounds: | query | master | this PR | delta | |---|--:|--:|--:| | s07 (30m range) | 2.22 s | 1.23 s | −44.8% | | r03 (rule, 3 selectors) | 5.02 s | 3.07 s | −38.8% | | s11 (24h range) | 4.30 s | 2.98 s | −30.8% | | r02 (25.6k-series instant) | 3.04 s | 2.17 s | −28.7% | | s06 (24h instant) | 5.04 s | 3.61 s | −28.4% | | full 24-query suite, cold geomean | 950 ms | 847 ms | **−10.8%** | Selectors that do not match a whole metric fall back and are unaffected (r05, s05: ±0.3%). Probe cost on non-firing selectors: ~2–4 ms each (r07: 66 → 73 ms); single-component id layouts never reach the probe. The removed cost grows with mark count, so the effect is larger at `index_granularity_bytes = 262144`. ## Tests `04836_time_series_selector_whole_metric_pk_range`: fires for whole-metric selectors (plan carries the range, `IN` retained), falls back byte-identically for label-filtered, regex, custom-generator, and ALTERed-`id_generator` history cases; both tuple layouts; full `prometheusQuery`/`prometheusQueryRange` results compared. `tests/performance/promql_selector_pk_range.xml`: 20,000-series metric, range path plus fallback control. Related: #113768 (open) — removes no-op casts in the same generated SELECT; complementary, each stands alone. --- ### Changelog category (leave one): - Performance Improvement ### Changelog entry: PromQL selectors that match all series of one metric now filter the samples table of a `TimeSeries` table with a continuous primary-key range on `id` during index analysis instead of a large `id IN <set>` condition, when the id layout is a two-component tuple with the canonical id generator. Removes the dominant single-threaded index-analysis cost of selector-heavy PromQL queries: up to −45% cold latency on dashboard and rule query shapes, −11% cold geomean over the full suite on a 62-billion-sample table. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114131",
          "createdAt": "2026-08-10T10:07:19Z",
          "updatedAt": "2026-08-13T00:05:26Z",
          "timestamp": "2026-08-13T00:05:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-performance",
            "comp-promql"
          ],
          "author": "nikitamikhaylov",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:03e79af83e0e8e279862",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110972",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110972",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support parallel replicas for Merge tables and the merge() table function",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/67770 Related: https://github.com/ClickHouse/ClickHouse/pull/95128 Queries over `Merge` tables and the `merge` table function can now be executed with parallel replicas, gated behind a new setting `parallel_replicas_allow_merge_tables` (default off). Previously such queries always ran on a single replica ([#67770](https://github.com/ClickHouse/ClickHouse/issues/67770)): on the CI Logs cluster, `SELECT count() FROM text_log WHERE message LIKE '%test%'` reads at 100 GB/sec using 10 replicas, while the same query through `merge()` reads at 10 GB/sec on one replica. The implementation follows the same approach as parallel replicas over views (`parallel_replicas_allow_view_over_mergetree`) rather than the per-table-coordinator UNION expansion attempted in the earlier PR [#95128](https://github.com/ClickHouse/ClickHouse/pull/95128): the whole first stage of the query (including aggregation) is offloaded to the replicas, and reading from every underlying `MergeTree` table is coordinated by a single reading coordinator, where each underlying table forms its own data stream (the same `stream_id` multiplexing that already serves `UNION ALL` views over `MergeTree` and projection splits). - On secondary replicas, the child reading steps created by `ReadFromMerge` pick up the coordination callbacks from the query context naturally, one announcement per underlying table. - On the initiator, `createLocalPlanForParallelReplicas` finds the `ReadFromMerge` step and switches it into a mode where each child reading step is created as a local parallel replicas reading step wired to the initiator's coordinator (`ReadFromMerge::enableParallelReplicasLocalPlan`). - A storage created by a table function does not exist on remote replicas under its generated id (`_table_function` database), so for `merge(...)` the remote reading step is created without a \"main table\"; otherwise the tables-status check would mark every remote replica as unusable. Eligibility (checked on the initiator and re-checked on the replicas): every underlying table must be a `MergeTree` table, and non-replicated underlying tables additionally require `parallel_replicas_for_non_replicated_merge_tree`. A `Merge` table with a non-`MergeTree` child, no children at all, or `FINAL` falls back to regular single-replica execution, because any child that cannot be coordinated at the level of parts and mark ranges would be read in full by every replica and duplicate its rows in the result. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support parallel replicas for `Merge` tables and the `merge` table function (enabled by the setting `parallel_replicas_allow_merge_tables`): reading from every underlying `MergeTree` table, and the whole first stage of the query, is now distributed across the replicas of the cluster. Closes [#67770](https://github.com/ClickHouse/ClickHouse/issues/67770). ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110972",
          "createdAt": "2026-07-19T06:54:53Z",
          "updatedAt": "2026-08-13T00:00:44Z",
          "timestamp": "2026-08-13T00:00:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 32
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:66a7af35423db54affd5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109252",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109252",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make JSONExtract honour cast_string_to_date_time_mode when parsing DateTime values",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `JSONExtract` now honours `cast_string_to_date_time_mode` when converting string JSON values to `DateTime`/`DateTime64`, consistently with `CAST`. Closes #109126. --- Extracting a string JSON value into `DateTime`/`DateTime64` is a string-to-type cast, but `JSONExtract` keyed its parsing mode off `date_time_input_format` (an input-format parsing setting), while the equivalent `CAST` honours `cast_string_to_date_time_mode`. With `date_time_input_format = 'basic'` and `cast_string_to_date_time_mode = 'best_effort'` (reproduced on current master): ```sql SELECT JSONExtract('{\"date\":\"2020-01-01 00:00:00.123Z\"}', 'date', 'DateTime64(3)'); -- 1970-01-01 00:00:00.000 (silently returns default) SELECT toDateTime64('2020-01-01 00:00:00.123Z', 3); -- 2020-01-01 00:00:00.123",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109252",
          "createdAt": "2026-07-03T03:40:20Z",
          "updatedAt": "2026-08-13T00:00:24Z",
          "timestamp": "2026-08-13T00:00:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "Utkal059",
          "state": "open",
          "assignees": [
            "george-larionov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:00fb660aa78cd98b69cb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114059",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114059",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix quadratic gzip/zlib streaming decompression of single-block streams",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114045 Related: https://github.com/ClickHouse/ClickHouse/pull/108074 Related: https://github.com/ClickHouse/ClickHouse/pull/114054 Related: https://github.com/ClickHouse/libdeflate/pull/6 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `O(n^2)` decompression of gzip/zlib/deflate streams whose single DEFLATE block spans the whole stream, the shape zlib-ng produces at compression level 1 (the default of the official .NET SDK): a 24 MiB gzip HTTP body that took 11 s to ingest in 26.7 now takes 0.12 s, identical to a multi-block body, with memory bounded by the buffer size instead of the block size. ### Details The root cause: the streaming decompressor added to `contrib/libdeflate` kept its resume checkpoint only at DEFLATE block boundaries. On input exhaustion or output-buffer overflow it rolled back to the start of the current block, so for a stream that is one giant block, every socket refill re-decoded the block from its beginning — quadratic time — and the whole block's output had to be buffered before any byte was exposed — unbounded memory. The work happens in `copyDataImpl` while assembling the query, before the query starts, so it was invisible in `system.processes` and `query_duration_ms`. The fix (https://github.com/ClickHouse/libdeflate/pull/6, vendored here as a submodule bump) re-takes the checkpoint at every symbol boundary in the generic decode loop, plus two new decompressor fields (`in_block`, `block_is_final`) that let a resumed call jump straight back into symbol decoding — the litlen/offset decode tables already persist in the decompressor between calls. The fastloop is untouched: suspensions can only trigger from the generic loop, which always runs between the fastloop and input exhaustion. A suspension now loses at most one partially decoded symbol, so decompression is linear-time with bounded memory regardless of block structure. `LibdeflateInflatingReadBuffer` needed no logic changes, only comment updates: its \"grow the output buffer\" path is now reachable only when a single match (≤ 258 bytes) or stored block (≤ 64 KiB) exceeds the free output region. Verification: - Standalone harness: 5000+ randomized round-trips (zlib streams at all levels with random flush points; crafted static- and dynamic-Huffman single-block streams with matches reaching the full 32 KiB window; input chunking down to 1 byte; random output regions), zlib inflate as decode oracle, one-shot decoder as cross-check, and a progress assertion on every suspension. - End-to-end with the issue's reproducer shape (24 MiB single-block gzip body, `Transfer-Encoding: chunked`, local release build): **11.25 s before → 0.12 s after**, now identical to a multi-block body of the same content (0.12 s). Verified the vendored zlib-ng really emits this shape: at level 1 it produces 1 block for a 4 MiB input (stock zlib: 80 blocks). - No throughput regression on normal multi-block streams: 685 → 697 MB/s on a 256 MB zlib-level-6 stream. New tests: - `LibdeflateSingleBlock.StreamingConsumesInputWithinBlock` pins the linearity contract (a suspension inside a Huffman block leaves at most a few bytes unconsumed). Against the pre-fix library it fails immediately, with 35 MB left unconsumed on a 32 MiB single-block stream. - `LibdeflateInflateTest.RoundTripFromZlibNgLevelOneLarge` decodes a large stream produced by the vendored zlib-ng at level 1 — the real-encoder shape of the issue. - `LibdeflateInflateTest.SingleDeflateBlockSpanningWholeMember` decodes crafted single-block gzip/zlib members through `LibdeflateInflatingReadBuffer`. - `04836_http_gzip_single_deflate_block` ingests a crafted single-block gzip body over HTTP with chunked transfer encoding, end to end. Because the pre-fix decoder is correct, just quadratic, the assertion is on time: the body is one 48 MB block and the request is capped at 15 s. Ingest time for a single-block body of a given size, release build, `master`'s `contrib/libdeflate` vs this branch's: 10 MB - 1.8 s vs 0.05 s; 20 MB - 6.9 s; 40 MB - 27.3 s vs 0.20 s; 48 MB - 39.1 s vs 0.25 s. So the limit sits 2.6x below the pre-fix time and 60x above the fixed one, and the clean quadratic growth means the pre-fix margin does not depend on the machine speed. The payload is 120 lines of 400 KB so that line parsing and the `MergeTree` write do not dominate. Note: `Bugfix validation (unit tests)` reports `ERROR` (inconclusive, non-blocking) here for a structural reason: the fix itself lives in `contrib/libdeflate`, and the job cannot populate its merge-base \"before\" worktree at the merge-base submodule revision, so it refuses to build a before-binary that would validate the wrong submodule code. The functional test above is what validates the fix. Note: this PR and the revert #114054 are alternatives on master — if the revert merges first, this branch will be updated to re-apply the integration together with the fix; if this merges first, the revert can be closed (the 26.7 backport of the revert can proceed independently).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114059",
          "createdAt": "2026-08-09T18:13:29Z",
          "updatedAt": "2026-08-12T23:54:43Z",
          "timestamp": "2026-08-12T23:54:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "submodule changed"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:281fadd974bd6394a159",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114643",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114643",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-land aggregate function `gini` in the `sum` family",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/113763 Related: https://github.com/ClickHouse/ClickHouse/pull/113868 Related: https://github.com/ClickHouse/ClickHouse/pull/112280 --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New aggregate function `gini`, which calculates the [Gini coefficient](https://en.wikipedia.org/wiki/Gini_coefficient) of a column of finite, non-negative numeric values. The result ranges from `0` (all values equal) towards `1` as inequality grows; for a sample of `n` values the maximum is `(n - 1) / n`. `NaN` values are skipped and infinite values are rejected. The function returns `Float64` and consumes `O(n)` memory. ### Description Re-lands the `gini` function that #113868 reverted, following the option-2 spec @ Manerone gave in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5217284116 and confirmed in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5279278069. It re-adds #112280's function with exactly three items removed, and touches no file under `src/AggregateFunctions/Combinators/`: - the `getArgumentsThatCanBeOnlyNull` override, - `.returns_default_when_only_null = true` on registration, - the `argument_type->onlyNull()` branch in the creator. The property was doing the work. It made `AggregateFunctionFactory::getImpl` skip its only-null guard and build a real `gini` instance over `Nullable(Nothing)`, so `gini(NULL)` returned `Float64` `nan` where every `sum`-family function folds to `Nullable(Nothing)`. Without it the fold happens in the `Null` combinator before any other combinator is applied, and the creator's only-null branch becomes unreachable, which is also why `createAggregateFunctionSum` has no equivalent. `gini` now registers exactly like `sum`. Runtime `Nullable` handling is a separate axis, via `getOwnNullAdapter`, and is unchanged: `gini` over `[1, NULL, 3]` still returns `0.25`. Validated on three binaries: post-revert master (`gini` absent), a build carrying #112280's version, and this branch. Every literal-`NULL` cell on this branch equals the `sum` value measured on the same binary, and all non-`NULL` results are unchanged from #112280. No documentation files are touched. The page body is generated from the `FunctionDocumentation` block in `AggregateFunctionGini.cpp`, which this PR restores, so the docs autogeneration workflow fills the page from that source. cc @Manerone",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114643",
          "createdAt": "2026-08-13T13:34:59Z",
          "updatedAt": "2026-08-13T16:18:37Z",
          "timestamp": "2026-08-13T16:18:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "Manerone"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:59ec49e8653024057294",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112679",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112679",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Parallelize delete manifests reads",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description] Queries on Iceberg tables with many delete files now start faster. ClickHouse reads and decodes the delete manifest files concurrently instead of one at a time, so their storage reads overlap. The new setting `iceberg_delete_manifest_decode_concurrency` (default `4`) controls how many are decoded at the same time.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112679",
          "createdAt": "2026-07-30T23:03:10Z",
          "updatedAt": "2026-08-13T16:18:26Z",
          "timestamp": "2026-08-13T16:18:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-performance",
            "submodule changed"
          ],
          "author": "asya-ch",
          "state": "open",
          "assignees": [
            "scanhex12"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:df5ef56d61c6782ded0d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114644",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114644",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Revert \"NATS: add inline credentials setting\"\"",
          "text": "Reverts https://github.com/ClickHouse/ClickHouse/pull/114178, restoring the `nats_credentials` setting of the `NATS` table engine (originally added in https://github.com/ClickHouse/ClickHouse/pull/110733), and fixes the reason of the original revert: the possibility of referencing arbitrary server paths through `nats_credential_file`. `nats_credential_file` is a path on the server filesystem: the server opens it with its own privileges, and during authentication the credentials are sent to `nats_url`, which comes from the same query. So a path taken from SQL lets anyone who can define a `NATS` source probe the local filesystem (the error text distinguishes a missing file, a permission error, and a file without a seed), and exfiltrate files the server can read to a NATS server they control (the part of the file before the seed line is sent verbatim in the `CONNECT` frame's `jwt` field). See the analysis in https://github.com/ClickHouse/ClickHouse/pull/110733#discussion_r3750203599. This hole predates #110733 (`nats_credential_file` exists since 24.2), but with the inline `nats_credentials` setting restored, the path form is no longer needed in SQL at all. The restriction, modelled on `StorageMySQL::getSSLParams` (`e700bbec4c84c585`): `nats_credential_file` is accepted only from a named collection defined in the server configuration file, or as `nats.credential_file` in the server configuration itself. Every SQL spelling — the `SETTINGS` clause, an engine-argument override `NATS(collection, nats_credential_file = ...)`, or a named collection created by `CREATE NAMED COLLECTION` — throws `BAD_ARGUMENTS` with a message directing to `nats_credentials`. Replacing a configured path with inline `nats_credentials` from the query remains allowed (the path itself is not used then). Loading from previously-validated metadata (server startup, force-restore, and short-syntax `ATTACH`) is exempt via `isLoadingFromExistingMetadata`, so existing tables keep working after an upgrade; a user-issued full `ATTACH TABLE` query is still checked. `ALTER TABLE ... MODIFY SETTING` is not a bypass: the `NATS` engine does not support settings alter. Tests: a new stateless test `04891_nats_credential_file_path_restriction` covers all rejected SQL spellings and the accepted configuration-file sources, using a `NATS` named collection added to the stateless-test server configuration (`tests/config/config.d/named_collection.xml`, with `nats_url = '127.0.0.1:1'`, so passing the validation surfaces as `CANNOT_CONNECT_NATS`); `04665_nats_credentials_named_collection` is updated — the `nats_credential_file` query-override direction is now rejected. Related: https://github.com/ClickHouse/ClickHouse/pull/110733 Related: https://github.com/ClickHouse/ClickHouse/pull/114178 Related: https://github.com/ClickHouse/ClickHouse/issues/85213 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `NATS` table engine accepts credentials inline in the new `nats_credentials` setting (the same payload as a `.creds` file), and no longer accepts `nats_credential_file` from SQL: the path is a reference to a file on the server filesystem, which the server opens with its own privileges, so it can only be specified in a named collection defined in the server configuration file, or as `nats.credential_file` in the server configuration itself. Tables created before this restriction keep working after an upgrade.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114644",
          "createdAt": "2026-08-13T14:03:24Z",
          "updatedAt": "2026-08-13T16:18:19Z",
          "timestamp": "2026-08-13T16:18:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f14867f75281faf858d8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114184",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114184",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Implementing generic block nested loop join",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Implemented generic block nested loop join",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114184",
          "createdAt": "2026-08-10T16:04:43Z",
          "updatedAt": "2026-08-13T16:18:08Z",
          "timestamp": "2026-08-13T16:18:08Z",
          "metrics": {
            "reactions": 2,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:a42764577dcdebf4237b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114661",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114661",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Document that PREWHERE filters one join input before the JOIN\"",
          "text": "Reverts ClickHouse/ClickHouse#114484 - it is too low-quality, sorry. CC @PedroTadim @dhtclk",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114661",
          "createdAt": "2026-08-13T15:53:35Z",
          "updatedAt": "2026-08-13T16:18:03Z",
          "timestamp": "2026-08-13T16:18:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "rschu1ze",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:07f18f03ddbe9f0c8115",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114656",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114656",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "docs: convert merge loop code block to Steps component",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replace the numbered code block describing the background merge loop on the AWS performance page with a `<Steps>` component for clearer rendering. <!-- mintlify-agent-attribution --> --- Generated by Mintlify Agent. Requested by: shaun.struwig@clickhouse.com via Slack Mintlify session: slack_1785790244.316179_D0B0U0V3Z88",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114656",
          "createdAt": "2026-08-13T15:16:40Z",
          "updatedAt": "2026-08-13T16:17:56Z",
          "timestamp": "2026-08-13T16:17:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "mintlify[bot]",
          "state": "open",
          "assignees": [
            "Blargian"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:d0d70df24b744d309d9e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114665",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114665",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow overriding the HTTP method for SELECT through the url table function and URL engine",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/62352 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `http_method='POST'` — as a key-value argument of the `url` table function and the `URL` table engine, or through the `http_method`/`method` named collection keys — now applies to `SELECT` queries: reads use `POST` instead of the default `GET`, for servers that accept only `POST`. Schema inference follows the configured method. `PUT` keeps its write-only meaning: a `SELECT` through a configuration with `http_method='PUT'` still uses `GET`. ### Implementation notes - `IStorageURLBase::getReadMethod()` now returns `POST` when the configured `http_method` is `POST`. `PUT` still applies to writes only (pre-signed upload URLs, #44326), and `INSERT` behavior is unchanged: `POST` by default, `PUT` when configured. - The write path no longer mutates the shared `http_method` member when defaulting to `POST` — the storage instance is shared between queries, so persisting the write default would have flipped subsequent reads of the same table from `GET` to `POST` (and raced concurrent reads). - Schema inference (`getTableStructureAndFormatFromData` and the URL read-buffer iterator) uses the same effective read method instead of hardcoded `GET`, for both `url()`/`URL` and `urlCluster`. - The `http_method = '...'` (or `method = '...'`) key-value argument is parsed by `StorageURL::evalArgsAndCollectHeaders` alongside `headers(...)`, kept in the `CREATE` AST (so it survives `SHOW CREATE TABLE` and DETACH/ATTACH), validated to be `POST`/`PUT`, and rejected when the URL scheme dispatches to another backend (`file://`, `s3://`, ...), mirroring `headers(...)`. - In the analyzer, the argument's left-hand identifier is excluded from column resolution (`skipAnalysisForArguments`), like `headers(...)`. - Documentation for the `url` table function and the `URL` engine is updated. The new stateless test `04869_url_select_http_method` asserts the method actually used on the wire via `system.query_log.http_method`: default `GET`, `POST` override for both the data and the schema-inference requests, `PUT`-configured `SELECT` staying on `GET`, the `URL` engine path, rejection of unsupported methods, and persistence of the argument in the table DDL. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01XvXEBdaJTCCnUdunCf8icr",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114665",
          "createdAt": "2026-08-13T16:17:38Z",
          "updatedAt": "2026-08-13T16:17:38Z",
          "timestamp": "2026-08-13T16:17:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "valerypetrov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:4d0d8d7635a81e14446a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114609",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114609",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix flaky 04780_json_subcolumn_index_match_not_quadratic",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113289 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Requested in https://github.com/ClickHouse/ClickHouse/pull/112705#issuecomment-5276141544. Related: https://github.com/ClickHouse/ClickHouse/pull/113289, which added the test. Two independent test defects. The engine fix itself is fine. **1. The oracle read a counter that does not exist on sanitizer builds.** `ProfileEvents['MemoryAllocatedWithoutCheckBytes']` is written only from the allocation paths that skip the memory-limit check, whose only non-Darwin producer is `src/Common/malloc.cpp`, compiled under `#if USE_JEMALLOC`. jemalloc is disabled for every sanitizer except UBSan, and the throwing `operator new` goes through `CurrentMemoryTracker::allocThrow` instead, so on `asan_ubsan`, `tsan` and `msan` the counter has no writer: on an ASAN build a query peaking at 76 MB reports 256 bytes through it. The test's `-gt 0` guard then aborted the first arm, which is the `no query_log row for plain dotted arm` failure. The test now reads `system.query_log.memory_usage`. Peak memory is tracked by `operator new` accounting, which is not sanitizer-gated, so it is defined on every flavour. The untracked-memory limit stays pinned to 0 so allocations reach the query tracker immediately, not in deferred batches whose flush points depend on concurrent load. **2. The last `EXPLAIN` asserted a master-only default.** `Parts: 0 | Granules: 0` comes from `describeActions`, gated by `actions`, which only master force-enables (`set_default_pretty_explain_settings`, absent from 26.5 and 26.6), so the two open backports of #113289 fail deterministically. Fixed by requesting `actions = 0` and dropping that line. `Granules: 0/1000`, the pruning under test, comes from `describeIndexes` and is unaffected. I could not confirm the 150% bound is too tight under sanitizers. It is sound on a quantity that measures the matcher, so it stays. Validation: 80 of 80 runs pass on debug and ASAN with randomization on and off, and 10 of 10 on ASAN under six concurrent load generators. On a pre-#113289 binary the oracle reddens on all four arms at 2.372x to 2.483x, against 1.0001x with the fix present.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114609",
          "createdAt": "2026-08-13T09:13:35Z",
          "updatedAt": "2026-08-13T16:17:19Z",
          "timestamp": "2026-08-13T16:17:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:491cc344bd2c2f89bfef",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111830",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111830",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject truncated/incomplete AI text function responses",
          "text": "The AI text functions (`aiGenerate`, `aiClassify`, `aiExtract`, `aiTranslate`) could silently return a truncated answer. When a provider stops early — most commonly by hitting the `max_tokens` limit — it reports this in the response, but the shared base class `FunctionBaseAI` never inspected the signal and returned whatever partial text came back. For example, `aiGenerate('Write three sentences about the ocean.', map(..., 'max_tokens', '5'))` returned the fragment `The ocean covers more than` with no error. This change normalizes each provider's native stop reason (OpenAI `finish_reason`, Anthropic `stop_reason`) into a canonical `FinishReason` enum, so the base class can make a single completeness decision without knowing the provider dialect: - `Truncated` (token/context limit) → throws `AI_PROVIDER_RESPONSE_TRUNCATED`. - `ContentFilter` (OpenAI `content_filter`, Anthropic `refusal`) and `ToolCall` → throws `AI_PROVIDER_RESPONSE_INCOMPLETE`. - `Complete` (natural end, or a caller stop sequence such as Anthropic `stop_sequence`) and `Unknown` (unrecognized reason) are accepted, so benign non-`stop` reasons are not misclassified as truncation. The rejection is thrown inside the existing per-row `try`, so it is non-retriable (retrying would hit the same limit) and honors `ai_function_throw_on_error`: with `1` the exception propagates; with `0` the row becomes the column default. `aiEmbed` is unaffected (embeddings have no finish reason). Integration tests in `test_ai_functions` cover truncation (throw + graceful), content-filter, an accepted unknown reason, and the Anthropic `stop_sequence` (must not throw) and `max_tokens` (must throw) cases. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): AI text functions (`aiGenerate`, `aiClassify`, `aiExtract`, `aiTranslate`) now reject truncated or otherwise incomplete provider responses (e.g. when the model hits the `max_tokens` limit) instead of silently returning partial output. Behavior follows the `ai_function_throw_on_error` setting.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111830",
          "createdAt": "2026-07-24T17:24:29Z",
          "updatedAt": "2026-08-13T16:16:56Z",
          "timestamp": "2026-08-13T16:16:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "george-larionov",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:703f2556f93840ebb4fa",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110838",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110838",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add introspection TCP port",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ClickHouse server now has an introspection port. This is a native protocol TCP listener that starts before the server begins attaching tables and stops only after the tables' detach completes. During these windows, an operator can connect to it with `clickhouse client` and run queries such as `SHOW PROCESSLIST`, `SELECT * FROM system.stack_trace`, or `SYSTEM INSTRUMENT ADD 'QueryMetricLog::startQuery' SLEEP ENTRY 0.5`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110838",
          "createdAt": "2026-07-17T11:07:23Z",
          "updatedAt": "2026-08-13T16:16:42Z",
          "timestamp": "2026-08-13T16:16:42Z",
          "metrics": {
            "reactions": 1,
            "comments": 16
          },
          "labels": [
            "pr-feature"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "evillique"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f72ab03ef566612f968c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114602",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114602",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: regenerate reference documentation from source",
          "text": "This pull request is opened automatically by the nightly documentation autogeneration workflow. It regenerates settings, functions, table and database engines, data types, formats, table functions, window functions, system tables, and asynchronous metrics from the structured documentation embedded in the ClickHouse source and exposed through the corresponding `system.*` tables. The generator preserves page frontmatter and hand-written content outside the `{/*AUTOGENERATED_START*/}` / `{/*AUTOGENERATED_END*/}` regions. Some pages are fully generated below their frontmatter. Do not edit generated content by hand -- edit the structured documentation in the defining source code instead; the next nightly run regenerates the pages. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114602",
          "createdAt": "2026-08-13T08:02:58Z",
          "updatedAt": "2026-08-13T16:15:22Z",
          "timestamp": "2026-08-13T16:15:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-autogenerated-docs"
          ],
          "author": "clickhouse-gh[bot]",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:34006a575acf220073c5",
        "signalId": "github:ClickHouse/ClickHouse:issue:58242",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:58242",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Non-constant right hand side of IN",
          "text": "``` milovidov-desktop :) SELECT number FROM numbers(10) WHERE number % 2 IN (number % 3, number % 5) Received exception: Code: 47. DB::Exception: Missing columns: 'number' while processing query: 'number % 3', required columns: 'number' 'number': While processing (number % 2) IN (number % 3, number % 5). (UNKNOWN_IDENTIFIER) ``` It can be done in the following way: - support non-constant case for IN [array] as full scan (`has`); - interpreting non-constant IN (set) in the same way as IN [array].",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/58242",
          "createdAt": "2023-12-27T11:35:09Z",
          "updatedAt": "2026-08-13T16:15:22Z",
          "timestamp": "2026-08-13T16:15:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "feature",
            "comp-sql-syntax",
            "warmup task"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "novikd"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:4f6f9d1b39c42aae5fc9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104993",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104993",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Handle non-constant RHS for `IN`",
          "text": "Fix `IN` and `NOT IN` expressions with non-constant right-hand side operands that reference columns from the current row. # old analyzer Previously, the old analyzer tried to build a standalone `Set` for expressions such as `number % 2 IN (number % 3, number % 5)`, which made the right-hand side unable to resolve `number` and produced an `UNKNOWN_IDENTIFIER` exception. This change rewrites such expressions to row-wise `has` expressions instead. ```sql --- error was reported for the query below for old analyzer as mentioned in #58242 SET enable_analyzer = 0; SELECT number FROM numbers(10) WHERE number % 2 IN (number % 3, number % 5) ORDER BY number; ``` ### Out of scope: bare source-column RHS under the old analyzer Under the old analyzer (`enable_analyzer = 0`), a bare column of the `FROM` source as the right-hand side, such as `x IN (arr)` where `arr` is a column of the current row, still fails with `UNKNOWN_TABLE`: `MarkTableIdentifiersVisitor` rewrites `x IN ident` into `x IN (SELECT * FROM ident)` before source columns are collected, so the expression never reaches the row-wise rewrite. The new analyzer resolves the same query as a column and succeeds; tests `04234` and `04812` pin this divergence explicitly. Closing it needs either reordering that visitor after source columns are known, or falling back from a table to a column when the table does not exist, plus a compatibility decision for `x IN t` when a column shadows an existing table name - that is tracked in the review discussion and left out of this PR on purpose. # new analyzer The new analyzer already handled the basic non-constant right-hand side case, but some tuple and NULL cases still failed. This change fixes: * tuple-typed right-hand side expressions produced by functions other than tuple * tuple left-hand side membership checks that previously tried to create `Nullable(Tuple(...))` * `NULL` operands in non-constant tuple right-hand side operands, where the old cast target could become `Nullable(Nothing)` Examples: ```sql --- Before this fix, the new analyzer treated the tuple-typed if RHS as a single tuple value and failed with a type error; after this fix, it expands the tuple value one level for scalar IN, so the query returns 1 SELECT number IN (if(number >= 0, tuple(number, number + 1), tuple(0, 0))) FROM numbers(1); ``` ```sql --- tuple in tuple, user exepects some rows to match, however, error like `Cannot create column with type 'Nullable(Tuple(UInt8, UInt8))' because Nullable Tuple type is not allowed` will be reported before this fix SELECT number, (1, 1) IN ((number % 3, number % 2), (2, 2)) FROM numbers(6) ORDER BY number; ``` ```sql --- user expects `NULL` but error like `Conversion from UInt8 to Nothing is not supported` will be reported before this fix SELECT x IN (y, 1) FROM ( SELECT materialize(NULL) AS x, materialize(2) AS y ); ``` Issue: https://github.com/ClickHouse/ClickHouse/issues/58242 ### Changelog category (leave one): - Bug Fix ### Changelog entry: - Fix IN and NOT IN expressions with non-constant right-hand side operands referencing columns from the current row, and align new analyzer tuple right-hand side handling with existing ClickHouse IN semantics. Under the old analyzer, a bare source column as the right-hand side (`x IN (arr)`) still resolves as a table name and stays out of scope. This closes [#58242](https://github.com/ClickHouse/ClickHouse/issues/58242)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104993",
          "createdAt": "2026-05-15T04:11:24Z",
          "updatedAt": "2026-08-13T16:15:21Z",
          "timestamp": "2026-08-13T16:15:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 27
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "niyue",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3a4e9480426b076e169c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:103706",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:103706",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "ReaderExecutor: pipeline-based read orchestration",
          "text": "Builds on top of #103234 (ReadPipeline). Gated by `SET use_reader_executor = 1` (experimental, default off). This commit replaces the matryoshka of `ReadBuffer` wrappers with a `ReaderExecutor` that owns offset mapping, cache decisions, prefetch and decryption in one place. Caches plug in via a uniform `ICacheProvider` / `ICacheHandle` API. ## What this brings **Zero-copy page-cache hits.** Reads served from `PageCache` reach the caller as a `shared_ptr` into the mmap'd cell — no `memcpy` between cache and working buffer (compare to `CachedInMemoryReadBufferFromFile`'s per-block copy). **Connection reuse across remote read calls.** Sequential reads against the same S3/Azure/HTTP object reuse an open buffer instead of issuing a new HTTP request per window. `SourceBufferLimit` caps live connections globally (`max_remote_read_connections`, default 1000); over-limit reads fall back to stateless open-read-close. Visible via `system.remote_read_connections`. **Coordinated cache + prefetch decisions.** The matryoshka design had each layer making independent choices (prefetch ahead in `AsynchronousBoundedReadBuffer`, decide what to cache in `CachedOnDiskReadBufferFromFile`, gather across objects in `ReadBufferFromRemoteFSGather`). The executor sees the whole window: - prefetches only what isn't already cache-hit - retains source over-read only when paired with a live connection (drops it otherwise) - knows which bytes are user-requested vs cache-fill, so it doesn't account cache-fill bytes against the user's read budget **Per-object cache identity, not per-pipeline.** `DiskCacheProvider` derives `FileCacheKey` / `FileCacheOriginInfo` per `StoredObject` (etag-keyed caching for `StorageObjectStorageSource`, segment-key-type classification for `Data` vs `System` queues). The old single-key-per-executor approach silently shared cache identity across all objects in a gather-mode read. **Foundation for new cache backends.** New caches (vector cache, distributed cache, etc.) implement one interface (`ICacheProvider::lookup`) and plug into the chain without touching read paths. ## Observability - `system.remote_read_connections` — open connections (path, query id, position, elapsed) - `system.reader_executor_log` — per-executor stats (cache hit/miss/populate bytes, prefetch hits/cancellations, source request count, decrypt time, …) - `ProfileEvents`: `ReaderExecutor*` counters and microsecond timers; `LiveSourceBuffer*` counters - `HistogramMetrics`: cache get/populate/source-read/prefetch-wait latencies ## Server settings - `reader_executor_prefetch_pool_size` (default 8) — shared prefetch thread pool size - `reader_executor_prefetch_queue_size` (default 0 = `pool_size * 10`) — pool queue depth - `max_remote_read_connections` (default 1000) — live-buffer slot count, reload-able ## Internals - `Rope` / `RopeNode` / `OwnedRopeBuffer` — refcounted buffer chains with mixed provenance (owned, page-cache-pinned), built-in cursor (`peek`/`advance`/`tryRewind`), sorted-on-insert nodes, disjoint-interval coverage tracking. - `ICacheProvider` / `ICacheHandle` — uniform cache API (lookup → status/get/put). Two implementations: `PageCacheProvider` (file-level, zero-copy), `DiskCacheProvider` (per-object, FileCache-backed). - `ISourceReader` — stateless range-read interface. `LocalSourceReader`, `ObjectStorageSourceReader`, `BufferSourceReader` (adapter for backup / `BufferCreator`). - `OffsetMap` — replaces `ReadBufferFromRemoteFSGather`'s gather logic with a logical-to-(object, object-offset) lookup; supports single-object unknown-size (S3 `HEAD` without `Content-Length` → streams to EOF). - `ReaderExecutor` — owns: position, offset map, cache chain, prefetch handle, live buffer, over-read tail, source-buffer slot. - `PipelineReadBuffer` — thin `ReadBufferFromFileBase` exposing the executor through `BufferBase::set/next/seek` (so legacy callers see no change). - `PrefetchThreadPool` — shared bounded pool, returns `nullptr` on overflow (sync fallback), task cancellation is race-free. ## Testing ~3,700 lines of tests (≈46% of the PR), 130+ gtest cases: - `gtest_rope` — 31 cases (append / peek / advance / tryRewind / slice / copyTo / coverage queries / shift). - `gtest_reader_executor` — 41+ cases (cache-chain combinations, per-object lookup, prefetch races, EOF release, slot leak fixes, unknown-size streaming, …). - `gtest_filecache` — 22 cases including over-read / bypass-mode / partial coverage scenarios. - `gtest_read_pipeline` — known/unknown size selection through the executor. - Functional: `04262_reader_executor_observability` plus opt-outs (`SET use_reader_executor = 0`) on tests that intentionally exercise the legacy path. ## Load tesing the simpliest cache API\" ``` ┌─────────┬────────────┬──────────┬───────┬────────────────────────┐ │ regime │ legacy QPS │ exec QPS │ QPS Δ │ per-query-type latency │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ cold │ 0.252 │ 0.320 │ +27% │ −12% (faster) ✅ │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ warm │ 0.407 │ 0.252 │ −38%¹ │ +27% (slower) ❌ │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ partial │ 0.247 │ 0.251 │ +2% │ +37% (slower) ❌ │ └─────────┴────────────┴──────────┴───────┴────────────────────────┘ ``` changed cache API (stream aware): ``` Full clean verdict — 9e9abe9d (26.6.1.1) ┌─────────────┬─────────────────────────┬───────────────────────────────────────────────────┐ │ Regime │ Executor vs legacy │ Note │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ cold │ −15% (win) │ 0 slot fail; coalesced reads │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ warm │ −8% (win) │ 0 slot fail; cache reads healthy │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ⚠️ partial │ +7% │ S3-source over-read / churn on uncached half │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ❌ populate │ +17% (worst) │ 50% connection reuse on cache-fill path — churn │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ stress │ 27 ok / 1 OOM vs 4 / 11 │ controlled degradation = net win (bounded, alive) │ └─────────────┴─────────────────────────┴───────────────────────────────────────────────────┘ stress details: ┌───────────────────────┬───────────┬──────────┐ │ │ legacy │ executor │ ├───────────────────────┼───────────┼──────────┤ │ queries finished (ok) │ 4 │ 27 │ ├───────────────────────┼───────────┼──────────┤ │ failed │ 11 │ 1 │ ├───────────────────────┼───────────┼──────────┤ │ OOM │ 11 │ 1 │ ├───────────────────────┼───────────┼──────────┤ │ connection reuse │ 86.8% │ 98.8% │ ├───────────────────────┼───────────┼──────────┤ │ connection resets │ 233,618 │ 38,581 │ ├───────────────────────┼───────────┼──────────┤ │ peak TCP recv-buffer │ 4,152 MiB │ 813 MiB │ ├───────────────────────┼───────────┼──────────┤ │ TCP sockets │ 17,333 │ 1,548 │ ├───────────────────────┼───────────┼──────────┤ │ cgroup mem │ 24.4 GiB │ 24.6 GiB │ └───────────────────────┴───────────┴──────────┘ sha -- 494fcfe8825c ┌─────────┬──────────────┬──────────────┬──────────────┐ │ regime │ baseline QPS │ executor QPS │ exec/base │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ cold │ 0.207 │ 0.292 │ 1.41× (+41%) │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ warm │ 0.330 │ 0.307 │ 0.93× (−7%) │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ partial │ 0.375 │ 0.420 │ 1.12× (+12%) │ └─────────┴──────────────┴──────────────┴──────────────┘ ``` That is good outcome. I need to optimize the work with caches. Need to stream data from available cache segments without any additional costs. sha -- dd4ae00fdaac (26.7.1.1) ``` ┌──────────┬──────────────┬──────────────┬──────────────┐ │ regime │ baseline QPS │ executor QPS │ exec/base │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ cold │ 0.230 │ 0.328 │ 1.43× (+43%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ warm │ 0.244 │ 0.302 │ 1.24× (+24%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ partial │ 0.332 │ 0.325 │ 0.98× (−2%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ populate │ 0.348 │ 0.384 │ 1.10× (+10%) │ └──────────┴──────────────┴──────────────┴──────────────┘ ``` First build with every regime at parity or better. The warm coordination-CPU tax and the populate cache-write regression are both gone; populate over-read down to 14%. Run under production-default networking (`disk_connections_rcvbuf=204800`, `max_remote_read_connections=1000`), with `reader_executor_use_long_connections=1` set explicitly (its default flipped to 0 on this build). Long-connection hit rate 99.5–100% on all regimes. sha -- fcccec1c7e6c (26.7.1.1) — equal-workload methodology ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 959 s │ 705 s │ 1.36× (+36%) │ │ warm │ 665 s │ 450 s │ 1.48× (+48%) │ │ partial │ 592 s │ 540 s │ 1.10× (+10%) │ │ populate │ 552 s │ 516 s │ 1.07× (+7%) │ │ evict │ 601 s │ 565 s │ 1.06× (+6%) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Methodology fix: both arms now run the identical query multiset (`--iterations`, sequential — no `--randomize`), metric = total wall time. The earlier windowed-QPS numbers systematically penalized the faster arm (a free benchmark worker draws a new random query, so the faster arm attracts more heavy queries) — warm was reported −9…−14% but is actually **+48%**; `evict` = new regime with the FileCache as a transit buffer (continuous eviction, no cross-query hits). The executor is faster in all five cache regimes; the 128-thread stress arm now completes without stuck queries (previously required `KILL QUERY`), with 4× fewer sockets than legacy. No correctness errors or crashes across the campaign since `c089b7e5`. Known remaining costs (per-query analysis): (1) cold/partial S3 over-read 2.2–2.8× from prefetch speculation; (2) on mixed cache regimes small queries regress 2-3× — `cache_get` returns zero bytes for data that is present (`FileSegmentWait` on DOWNLOADING segments, then reads from source anyway) plus block-granular request amplification on narrow columns (~90 MiB requested for a 5 MiB column), paid at S3 first-byte latency with ~90% of prefetches cancelled. sha -- 018c0257567b (26.7.1.1) — plan-look-ahead window build ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 810 s │ 619 s │ 1.31× (+31%) │ │ warm │ 477 s │ 450 s │ 1.06× (+6%) │ │ partial │ 785 s │ 518 s │ 1.52× (+52%) │ │ populate │ 532 s │ 478 s │ 1.11× (+11%) │ │ evict │ 579 s │ 581 s │ 1.00× (par) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Same equal-workload methodology (identical 186-query multiset per arm). Since the workload is fixed, executor wall times are directly comparable across builds: vs `fcccec1c` the executor improved on cold (705→619 s, −12%), populate (516→478 s, −7%) and partial (540→518 s, −4%) — the plan-window changes helped. Baseline wall times swing between rounds (legacy path + master merge + day variance), so cross-build conclusions should use the executor-vs-executor columns, not the ratios. Also corrected with the fixed workload: the true equal-work S3 over-read is ~2.0× on cold and ~2.3× on partial (the earlier 2.8–2.9× figures were inflated by the windowed methodology — the faster arm simply ran more queries per window). Stress arm self-completed again (4th consecutive build). No `Code: 33`, no crashes. sha -- 2b43930ebbcf (26.8.1.1) — read-path log cut + settings-history dedup ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 496 s │ 315 s │ 1.57× (+57%) │ │ warm │ 223 s │ 226 s │ 0.99× (par) │ │ partial │ 326 s │ 254 s │ 1.28× (+28%) │ │ populate │ 242 s │ 240 s │ 1.01× (par) │ │ evict │ 281 s │ 304 s │ 0.92× (−8%) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Two rounds this cycle. The morning round (head `a6465d03`) exposed a hot-read-path logging tax: `PipelineReadBuffer` logged one `Trace` message per window advance (6.3M messages per warm arm) plus ~1M per-buffer `Debug` \"Created\" messages, costing 795 s of thread time per arm (baseline: 6.7 s) — roughly 14% of the CPU budget on cache-served regimes, which showed as warm/populate 0.95×. The same-day fix (`da34b431`) is verified by this round: logger time dropped 795 → 170 s, allocation volume dropped 5.85 → 1.94 TB ≈ baseline's 1.84 TB (most of the long-standing \"executor allocates ~4× more\" observation was log-message formatting), and executor `UserTime` is now below baseline on warm. Populate flipped to parity-plus; warm is at 0.99× with the residual logger time still 34× baseline — a few more sites may be worth demoting. S3 over-read: populate 1.01× and evict 1.03× (the executor no longer over-fetches where the cache absorbs writes), cold 1.20× while making 7× fewer GETs at ~11 MiB each (this is what buys the 1.57× cold wall), partial 1.38× — the remaining cost item. The evict 0.92× turned out to be regime noise: a clean repeat pair on the settled cluster came back 323/279 s = 1.16× in the executor's favor, matching the morning round's 1.17× (evict has a history of single-pair swings — 0.78× → 1.18× between two runs of one build earlier). Verdict: executor ≥ baseline in all five regimes. Environment notes for cross-entry comparison: starting with this cycle the staging instance's default profile `compatibility` was bumped `25.12` → `26.8`, which changes effective defaults for both arms — all walls in this entry are ~2× faster than in previous entries for that reason; compare ratios, not walls, across entries. Relatedly, the settings-history dedup in this head means Cloud `compatibility` no longer resurrects the pre-release 8 MiB `reader_executor_plan_look_ahead_max_window` introduction default on private builds. Zero failed queries, no `Code: 33`, no crashes across all 20 arms of both rounds. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental `ReaderExecutor` for pipeline-based read orchestration with unified cache API (`PageCacheProvider`, `DiskCacheProvider`), connection reuse for remote reads, shared prefetch pool, and `system.remote_read_connections` / `system.reader_executor_log` observability tables. Enable with `SET use_reader_executor = 1`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/103706",
          "createdAt": "2026-04-29T07:27:30Z",
          "updatedAt": "2026-08-13T16:14:37Z",
          "timestamp": "2026-08-13T16:14:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:3d9c0c691fed3f68c56e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114615",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114615",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "diff-review skill second edition",
          "text": "Reworks the `diff-review` skill in `.claude/skills/diff-review` — the local in-browser diff review that Claude Code serves before a commit or a PR. Heads-up: I tailored this to how *I* like to review, so please treat it as a suggestion rather than a standard. Check the branch out, run it on one of your own diffs, and keep it only if you find it helpful. What changed: - **Multi-round review.** Sending a round no longer ends the review: the server stays up, the agent works on the comments it was handed while you keep reading, and each comment turns green in the page as it gets addressed. - **Comments are durable state.** They are written to the `--out` file as you type, so they survive a server restart or a page reload, and after the agent's fixes move the code they are relocated by their anchor line instead of pointing at the wrong place. Resolutions, replies and dismissals are part of that state. - **Explicit review range.** `--staged`, `--head <sha>` and `--committed` alongside `--base`, so a review of recorded work never picks up local edits; the header always states which range is on screen. - **Navigation.** Directory tree with per-file status, `+a −d` counts and open-comment badges, a path filter, an all-files page, whole-file mode with expandable folded regions, a Comments pane listing open and already-addressed comments, draggable sidebar and split, keyboard shortcuts with a `?` overlay, and light / dark / system themes. - **Tests.** `ui.html` is split into ES modules under `ui/`, covered by `ui_test.mjs`, and `persist_test.mjs` exercises the whole persistence round-trip against real servers on a throwaway repository. - Vendored `@pierre/diffs` bumped from 1.2.12 to 1.3.5. Local agent tooling only — nothing in the server, the build or the tests. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114615",
          "createdAt": "2026-08-13T09:54:07Z",
          "updatedAt": "2026-08-13T16:14:27Z",
          "timestamp": "2026-08-13T16:14:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8c65c2a6a8fdcd9cd903",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109252",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109252",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make JSONExtract honour cast_string_to_date_time_mode when parsing DateTime values",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `JSONExtract` now honours `cast_string_to_date_time_mode` when converting string JSON values to `DateTime`/`DateTime64`, consistently with `CAST`. Closes #109126. --- Extracting a string JSON value into `DateTime`/`DateTime64` is a string-to-type cast, but `JSONExtract` keyed its parsing mode off `date_time_input_format` (an input-format parsing setting), while the equivalent `CAST` honours `cast_string_to_date_time_mode`. With `date_time_input_format = 'basic'` and `cast_string_to_date_time_mode = 'best_effort'` (reproduced on current master): ```sql SELECT JSONExtract('{\"date\":\"2020-01-01 00:00:00.123Z\"}', 'date', 'DateTime64(3)'); -- 1970-01-01 00:00:00.000 (silently returns default) SELECT toDateTime64('2020-01-01 00:00:00.123Z', 3); -- 2020-01-01 00:00:00.123",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109252",
          "createdAt": "2026-07-03T03:40:20Z",
          "updatedAt": "2026-08-13T16:14:17Z",
          "timestamp": "2026-08-13T16:14:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "Utkal059",
          "state": "open",
          "assignees": [
            "george-larionov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ae88db30b2b340683bd9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104437",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104437",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add spatial_bbox skip index for MergeTree geometry columns",
          "text": "Part of making ClickHouse fastest spatial analytical engine on Earth https://github.com/bacek/chgeos/blob/main/BENCHMARK.md ;) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Adds `spatial_bbox` skip index for MergeTree geometry columns. The index stores a bounding box per granule and skips granules whose geometry cannot intersect the query geometry, reducing work for spatial predicates. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104437",
          "createdAt": "2026-05-08T22:48:54Z",
          "updatedAt": "2026-08-13T16:14:03Z",
          "timestamp": "2026-08-13T16:14:03Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "bacek",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1f7d1e925c572ed8ac1b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114466",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114466",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: internationalize master",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114466",
          "createdAt": "2026-08-12T10:49:59Z",
          "updatedAt": "2026-08-13T16:13:22Z",
          "timestamp": "2026-08-13T16:13:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 45
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "locadex-agent[bot]",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:72196530ea36853d6873",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114659",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114659",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Record TopK-filtered granules in the query condition cache",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Granules fully filtered by the dynamic threshold of an `ORDER BY ... LIMIT n` (TopK) read are now recorded in the query condition cache, so repeat runs of such queries skip them at the mark-selection stage instead of re-reading and re-filtering most of the table. On the ClickBench Q24/Q26 shapes (`SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10`), a warm run drops from ~95M rows / ~12000 granules read to ~1.3M rows / ~160 granules, and from ~15–30 ms to ~8 ms on a 96-core machine. The `ORDER BY ... LIMIT n` (TopK) optimization pushes a dynamic `__topKFilter` threshold into the `MergeTree` read as a PREWHERE, dropping rows that cannot beat the running top-N. But the granules it emptied were never recorded in the query condition cache: the PREWHERE write path in `MergeTreeSelectProcessor::read` rejects non-deterministic conditions, and `__topKFilter` is one. Only the downstream WHERE `FilterTransform` wrote entries, at whole-chunk granularity, which learns almost nothing (a single surviving row voids the attribution of the whole chunk) — measured on hits, warm runs still selected 11766 of 12348 granules and re-read ~95M of 100M rows. The original out-of-tree build of this optimization (PR #81944, commit `2e2f594308f4` plus the follow-up making the TopN dynamic filters deterministic and reusable for the cache) recorded them and skipped ~98% of granules on warm runs — that is the 2–3x Q24/Q26 gap measured in issue #114639. Recording these granules is sound: for a fixed plan and data, the running threshold only tightens, so a granule none of whose rows survive the filter contains no row that could reach the final top-N, regardless of the threshold trajectory. The entries are keyed with the TopK plan salt (`TopKFilterInfo::condition_hash`: sort column, type, `LIMIT`, direction, `NULLS` direction, collation locale, number of sort columns, and the part-set snapshot), mirroring what the WHERE write path already does since #104478/#110507 — only the same TopK plan over the same part set ever reuses them. Changes: - `isDeterministicAllowingTopKFilter` (two identical static copies in `updateQueryConditionCache.cpp` and `ReadFromMergeTree.cpp`) moved into `VirtualColumnUtils`. - The TopK salt is plumbed to the reader via `MergeTreeReaderSettings::query_condition_cache_top_k_salt`. - The PREWHERE write path accepts a `__topKFilter`-bearing condition when the salt is present, folding the salt into the cache key. Any other non-deterministic condition is still never cached. - The PREWHERE consult path (`filterPartsByQueryConditionCache`) applies the salt exactly when the PREWHERE contains `__topKFilter`, so the keys match the write side; deterministic user PREWHERE conditions keep using plain, unsalted entries shared with non-TopK queries. - `04217_query_condition_cache_topk` and `04242_query_condition_cache_topk_collate` pin exact cache entry counts, which double (each TopK plan now writes a WHERE entry and a PREWHERE entry per part). - New test `04891_query_condition_cache_topk_prewhere_granules` isolates the PREWHERE write path with a WHERE-less TopK query (on current master such a query writes no cache entries at all), asserts that the ASC and DESC plans do not share granule decisions, and that a warm run selects fewer marks. - New performance test `topk_query_condition_cache.xml` with the ClickBench Q24/Q26 shapes over `hits_100m_single`. Measured on the 100M-row single-part ClickBench `hits` (96-core aarch64, hot, interleaved runs): warm Q24/Q26 at 0.007–0.009 s, on par with the original bench-opt build (0.009–0.011 s), against 0.015–0.03 s for current master; warm runs read 159 of 12348 granules against master's 11766. A first run with an empty cache shows no measurable overhead (0.016–0.019 s with the mechanism on and off). Closes: https://github.com/ClickHouse/ClickHouse/issues/114639 Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/110507",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114659",
          "createdAt": "2026-08-13T15:46:11Z",
          "updatedAt": "2026-08-13T16:12:48Z",
          "timestamp": "2026-08-13T16:12:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:ffd8a11027dfc975c4d9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114596",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114596",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix target access checks for the Alias engine",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> `Alias` table now requires `SHOW COLUMNS` on its target when a new definition is submitted. This prevents users from using `Alias` to reveal a target table's schema. The trivial `count` optimization is disabled when the user has no `SELECT` privilege on the target, allowing the normal read path to enforce access checks. `rows` and `bytes` statistics are also hidden unless the user has `SHOW TABLES` on the target. Closes: https://github.com/ClickHouse/ClickHouse/issues/114511 ### Changelog category (leave one): - Critical Bug Fix (crash, data loss, RBAC) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an RBAC bypass in the `Alias` table engine that allowed users without privileges on the target table to reveal its schema, row count, size, and existence.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114596",
          "createdAt": "2026-08-13T05:30:42Z",
          "updatedAt": "2026-08-13T16:12:31Z",
          "timestamp": "2026-08-13T16:12:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-must-backport",
            "can be tested",
            "pr-critical-bugfix"
          ],
          "author": "nauu",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6d3a8fb2dd19d57efaaf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114635",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114635",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113742 to 26.7: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114635",
          "createdAt": "2026-08-13T12:50:24Z",
          "updatedAt": "2026-08-13T16:12:25Z",
          "timestamp": "2026-08-13T16:12:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c2a4c48d404620e76d13",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111451",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111451",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Release pull request for branch 26.7",
          "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111451",
          "createdAt": "2026-07-22T17:40:24Z",
          "updatedAt": "2026-08-13T16:12:25Z",
          "timestamp": "2026-08-13T16:12:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "release"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fdac17bd946c86e2ecd1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114472",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114472",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the patch release version bump increasing across recoveries",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113834 Related: https://github.com/ClickHouse/ClickHouse/pull/113528 --> Scheduled patch releases stopped incrementing the patch number — successive releases on a branch reused the same `vX.Y.P.*` line (e.g. `26.6.2.81`, `26.6.2.158`, `26.6.2.160`), because the post-release version bump was lost whenever a release was interrupted after the tag push, and every later recovery skipped it too. Prepare now classifies a run from the ref and the branch-tip version file into two flags: `is_recovery` (this run re-publishes an existing release rather than creating one) and `is_late_recovery` (the branch has already advanced to a newer release). The deferred bump is gated on `not is_late_recovery`, so a normal run and a current-release recovery complete the interrupted bump, while a superseded recovery never rewrites the version backwards. `stage_bump` refuses to write a version older than the branch tip, and asserts `is_recovery` on an empty bump — a fresh release must advance the version, a recovery may legitimately find it already done. Split out of #113834 — this is the version-bump half with a single push. Handling a non-fast-forward push (a backport moving the branch mid-release) is left to #113834. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114472",
          "createdAt": "2026-08-12T12:04:19Z",
          "updatedAt": "2026-08-13T16:12:22Z",
          "timestamp": "2026-08-13T16:12:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "leshikus",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8a24454c0f1d38f8ee76",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110230",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110230",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix S3Queue shutdown hanging on a streaming pipeline stuck inside a blocking call",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/103126 ## Problem `S3Queue`/`AzureQueue` shutdown blocks for as long as an in-flight streaming pipeline stays stuck inside a blocking call — a stalled object storage read, a mutex convoy on the shared file iterator, or an executor thread parked in `epoll_wait`. Consequences observed in production (a pipeline wedged for two hours): - A synchronous `DROP TABLE` on the queue table hangs, and everything serialized behind it hangs too. - Plain server shutdown / `DETACH TABLE` hangs the same way. Independently, the periodic metadata cleanup deletes `processing/` znodes and bucket locks older than `persistent_processing_node_ttl_seconds` (default 3600) by mtime alone, even while the owning streaming execution is still running in the same process. The execution's commit then fails with `Coordination::Exception: Transaction failed ... (No node)`, and the freed bucket lock can be acquired by another server concurrently (Ordered-mode correctness violation). ## Root cause `StorageObjectStorageQueue::shutdown` sets `shutdown_called` and calls `task->deactivate()`, which blocks until the in-flight `streamToViews` returns. The streaming pipeline runs `CompletedPipelineExecutor::execute` with no cancel callback, so all shutdown handling is cooperative at chunk boundaries inside `ObjectStorageQueueSource::generateImpl` (what #103126 merged). A pipeline parked inside a blocking call never reaches a chunk boundary, so shutdown blocks unboundedly. The cleanup TTL (`ObjectStorageQueueMetadata::cleanupPersistentProcessingNodes`) is meant to reap nodes orphaned by dead servers, but it has no notion of ownership: a node whose owner is alive in this very process is deleted just the same once its mtime passes the TTL. ## Background: why #103126's first draft dropped executor-level cancellation, and why it is safe now The first draft of #103126 used exactly this approach (`setCancelCallback` in `streamToViews`), and the review rejected it — not as wrong in principle, but because two correctness gaps had no answer at the time: 1. **Silent success after cancel → data loss** ([review comment](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3193288074): \"reading is safe, but AFAICS writing the data is not\"). `executor->cancel()` can land after the source finished reading but before the sink finalized; `CompletedPipelineExecutor::execute` then returns without an exception, and `commit(insert_succeeded=true)` marks files `Processed` whose rows were never written. 2. **Forced failed-commit → duplicate inserts.** The draft's counter-measures (\"skip commit when `shutdown_called`\", then a `cancel_was_triggered` flag) were each found racy ([1](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3195586146), [2](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3196025042)): a cancel landing between files forces an already-fully-`Processed` batch into the failed-commit path, so those files are reset and re-read — duplicating rows when deduplication is off. Two more factors sealed it: Kafka's identical fix (#100388, `setCancelCallback` + skip offset commits on cancel) had been reverted 18 days earlier (#101646) after making `test_kafka_commit_on_block_write` flaky, and the [proposed alternative](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3193311270) — a chunk-boundary throw inside `generateImpl`, where per-file state is precisely known — was much smaller to review. That is what merged. It solves \"shutdown waits for the in-flight file to reach EOF\" (minutes), but not \"the pipeline never reaches a chunk boundary at all\" (hours, this PR's production case). Both gaps are closed here, on top of the machinery the merged #103126 itself introduced: - Gap 1: when the cancel callback has fired, `streamToViews` **always** throws instead of committing success — the silent-success path no longer exists. - Gap 2: the cancel gate requires `table_is_being_dropped || is_deduplication_v2`, so a forced retry is either moot (drop — no retry) or absorbed by deduplication; the worst case is an extra read of the file, which the #103126 review itself deemed acceptable ([comment](https://github.com/ClickHouse/ClickHouse/pull/103126#discussion_r3213509835)). The failure direction also flips: the draft's races erred toward *not committing what was written* (loss), this PR's over-triggering errs toward *reset-for-retry under dedup* (at most a re-read). - The Kafka precedent does not transfer: Kafka's offset semantics have no deduplication, so cancellation must choose between loss and duplication; `S3Queue` since 26.2 has `deduplication_v2`, plus the `Cancelled`-not-`Failed` file state machine that the merged #103126 built — the very foundation that makes executor-level cancellation safe now. ## Solution 1. `streamToViews` installs `CompletedPipelineExecutor::setCancelCallback` (1 s poll), gated exactly like the chunk-boundary abort: `shutdown_called && (table_is_being_dropped || is_deduplication_v2)`. Without deduplication, a mid-file abort would duplicate already-inserted rows on retry, so dedup-off plain shutdown keeps processing the in-flight file to EOF, as before. `executor->cancel()` cancels the processors and wakes the polling queue, and the in-flight file lands in `Cancelled` state (reset for retry, not `Failed`) through the existing machinery. A cancelled pipeline that finishes without an exception is routed through the failed-commit path instead of being committed as successful, because the sink may not have finalized. **Scope of the cancellation.** `executor->cancel()` wakes the executor's own waits directly: `PollingQueue::finish` writes the self-pipe the driver thread sleeps on — which is exactly the shape of the two-hour production hang (the driver parked in `epoll_wait` with no worker running or blocked, frame-proven from `trace_log`). A worker inside a blocking storage call is not preempted; that call stays bounded by the client's per-socket-operation timeouts, and cancellation takes effect at the first check after it returns — what this PR removes is the previously unbounded continue-to-EOF/next-files work after that point. A blocking call that never returns (e.g. a dead connection outliving every socket timeout) needs an end-to-end per-request deadline; that is deliberately out of scope here and tracked as a follow-up. 2. Local executions register their node paths in a new `ObjectStorageQueueLocalActiveNodes` registry *before* creating them in Keeper (`trySetProcessing`, `prepareSetProcessingRequests`, bucket acquisition), backing off when registration is refused; the cleanup wraps its revalidate-and-remove window in a try-only removal lock that fails while the path is registered. Registration and the removal lock exclude each other under one mutex, so the cleanup can never delete a node owned by a live local execution, in any interleaving — the register-before-create ordering matters because a node deleted and recreated between the cleanup's revalidation read and its remove restarts at Keeper version 0, so no captured version can guard the removal; a recreated node is additionally recognized by its fresh mtime on the revalidation read. The registry is shared between sibling tables on the same keeper path, because `ObjectStorageQueueMetadataFactory` shares the whole `ObjectStorageQueueMetadata` object per `{zookeeper_name, zookeeper_path}`. Another server's cleanup can still delete a >TTL node of a live remote execution — inherent to the mtime-TTL design and unchanged here. Regression tests: a new `object_storage_queue_park_in_generate` failpoint parks a source mid-file until pipeline cancellation or failpoint disable; integration tests cover `DROP TABLE` with a stuck pipeline, both clauses of the cancel gate (`DETACH` with dedup on/off), live-node survival across TTL cleanup runs, and recreation inside the cleanup's collection-to-removal window (via a pauseable failpoint); a gtest pins the registry fencing protocol. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `S3Queue`/`AzureQueue` shutdown and `DROP TABLE` hanging for as long as a streaming pipeline stayed stuck inside a blocking call. Also fix the periodic metadata cleanup deleting the processing nodes and bucket locks of executions still running on the same server, which made their commits fail once `persistent_processing_node_ttl_seconds` elapsed. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110230",
          "createdAt": "2026-07-13T09:09:59Z",
          "updatedAt": "2026-08-13T16:11:22Z",
          "timestamp": "2026-08-13T16:11:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:4ef8d61ea91cb61f76ab",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96487",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96487",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix LazilyReadFromMergeTree optimization with ALIAS columns (#96452)",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry: Fix missing lazy read and top-K read optimizations (skip-index top-K and dynamic top-K filtering) when selecting `ALIAS` columns with `WHERE ... ORDER BY ... LIMIT` on MergeTree tables. Closes: https://github.com/ClickHouse/ClickHouse/issues/96452 ### Documentation entry for user-facing changes N/A",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96487",
          "createdAt": "2026-02-09T21:21:08Z",
          "updatedAt": "2026-08-13T16:11:13Z",
          "timestamp": "2026-08-13T16:11:13Z",
          "metrics": {
            "reactions": 1,
            "comments": 31
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "jayvenn21",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:64c8329847ca782ff7f7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100371",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100371",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `borrow_from_cache` object storage and `memory` metadata types",
          "text": "Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache using ephemeral `FileSegment`s. Each stored object is backed by a cache segment held alive via `FileSegmentsHolder`; when released, the cache reclaims the space. Add a new metadata type `memory` that keeps all file-to-blob mappings and directory structure entirely in memory with no persistence. On server restart all data is lost, which is acceptable for the intended use case of temporary tables. Configuration example: ``` disk(type=object_storage, object_storage_type='borrow_from_cache', metadata_type='memory', cache_name='some_cache') ``` ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A new object storage type and metadata type suitable for temporary tables. Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache and holds it from eviction. Add a new metadata storage type `memory` that keeps the mapping in memory. For example, this could allow creating temporary tables with custom engines in the Cloud without using S3.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100371",
          "createdAt": "2026-03-22T15:18:44Z",
          "updatedAt": "2026-08-13T16:10:29Z",
          "timestamp": "2026-08-13T16:10:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 32
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2102cec24ba9ccdc970a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114664",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114664",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113450 to 26.7: Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114664",
          "createdAt": "2026-08-13T16:09:00Z",
          "updatedAt": "2026-08-13T16:09:46Z",
          "timestamp": "2026-08-13T16:09:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "JiaQiTang98",
            "groeneai"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:0e4120f8821184a1a66d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114658",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114658",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix inconsistent AST formatting for a subquery argument of the view table function",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114004 (auto-closes the issue when this PR is merged into the default branch) --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114004 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed formatting of a subquery argument of the `view` and `viewIfPermitted` table functions when the enclosing query has a trailing `SETTINGS` clause. Such a query was formatted as `view((SELECT ...))`, which cannot be parsed back, so the query failed the internal format-parse-format check and raised `Inconsistent AST formatting`. ### Description `ASTQueryWithOutput::formatImpl` sets `parent_has_trailing_settings` so an inner `ASTSelectWithUnionQuery` parenthesizes its individual SELECTs, keeping the re-parser from consuming the trailing `SETTINGS` into the last SELECT. The flag is inherited down the format frame, so it also reached the argument of `view` / `viewIfPermitted`. There the parentheses are both redundant and rejected: the closing paren of the table function already terminates the select, and `ViewLayer::parse` bails out on a parenthesized lone select (`ExpressionListParsers.cpp:2981-2991`). The formatted text therefore did not parse back. The trailing `SETTINGS` clause is the trigger. Without it nothing sets the flag and the output is already correct. This clears the flag at the two places that cross the `view` argument boundary, which is what three other boundary owners already do: `ASTSubquery.cpp:98`, `ASTCreateQuery.cpp:1113` and `ASTAlterQuery.cpp:1032`. `ViewLayer` is the only producer of a bare-select function argument and serves exactly these two functions, so the two sites are the complete set. The issue frames parser-versus-formatter as the fork. This takes the formatter side, on the grounds that the parentheses carry no meaning in this position and that the same reset is the established pattern; the parser side would widen the accepted grammar to fix an output bug. The call is his to close. One output change beyond the broken shape: a multi-select argument such as `view(SELECT 1 UNION ALL SELECT 2)` also loses its branch parentheses. That form parsed before and parses now. Validation: `DESCRIBE TABLE view(SELECT 1) SETTINGS input_format_orc_use_fast_decoder = 0` aborts a debug server with `Logical error: 'Inconsistent AST formatting'` before the change and returns `1 UInt8` after it. The new test fails on the unpatched binary and passes 50/50 under randomized settings on the patched one. The `formatQuery` / round-trip / `1941` stateless family was run on both binaries and the failure sets are byte-identical.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114658",
          "createdAt": "2026-08-13T15:46:03Z",
          "updatedAt": "2026-08-13T16:09:26Z",
          "timestamp": "2026-08-13T16:09:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:4b221ab89ee070cb3b74",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113450",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels",
          "state",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113450",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113337 Related: https://github.com/ClickHouse/ClickHouse/pull/113425 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables that contain a nullable `ARRAY` or `MAP` column. Such a table could not be read at all, because the schema mapper wrapped the composite type in `Nullable`, which ClickHouse forbids, so both `DESC` and `SELECT` failed with `Nested type Array(Nullable(Int32)) cannot be inside Nullable type`. A nullable composite column is now mapped to a non-`Nullable` composite type and a `NULL` value is read as an empty one. Closes #113337. ### Description This takes over #113425 by @zlareb1, who closed it and asked me to carry it forward. The diagnosis and the fixture are his; the source change uses the in-tree capability gate rather than deleting the wrap. **What breaks.** Paimon columns are nullable by default, so a plain `CREATE TABLE paimon.default.t (f ARRAY<INT>)` from Spark produces an unreadable table. `Paimon::DataType::parse` applied its `if (nullable)` wrap in the composite branch as well as the scalar one, and `DataTypeArray`/`DataTypeMap` return `canBeInsideNullable() == false`, so `DataTypeNullable`'s constructor threw. This happens while parsing the schema, so it takes out the whole table rather than one column. Affects `paimonS3`/`paimonLocal`/`paimonAzure`, their `*Cluster` variants, the `Paimon*` engines and the REST catalog. The engines need `allow_experimental_paimon_storage_engine`; the table functions do not. **The change.** The two inner wraps become a single `makeNullableSafe`, which wraps only when the type permits it. Neither the Iceberg nor the DeltaLake schema processor wraps a composite in `Nullable` (Iceberg gates on `canBeInsideNullable()`; DeltaLake keeps the wrap in its scalar branches only), so this aligns Paimon with them. The scalar wrap is untouched, so inner nullability survives: the fixture reads as `Array(Nullable(Int32))` and `Map(String, Nullable(Int32))`. The gate, rather than deletion, keeps the wrap available for a future `ROW`, whose `DataTypeTuple` does permit it. A `NULL` composite reads as an empty one, the Parquet reader's documented behaviour, so the two become indistinguishable. That is forced by the type system and matches Iceberg, DeltaLake, Arrow, ORC and Avro. **Validation.** New test `04757_paimon_nullable_composite_types` over @zlareb1's fixture fails on master with the error above and passes with the fix. The ten existing Paimon tests are identical on both binaries. Restoring either wrap individually reddens the new test on its own message. 50/50 randomized runs pass. `Types.h` is absent on 25.8, so 26.3 through 26.7 are affected; `must-backport` labels look appropriate.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113450",
          "createdAt": "2026-08-05T10:05:17Z",
          "updatedAt": "2026-08-13T16:09:09Z",
          "timestamp": "2026-08-13T16:09:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-backports-created",
            "pr-must-backport-synced",
            "v26.5-must-backport"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "JiaQiTang98"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:097e01064b042d32dcea",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114663",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114663",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113450 to 26.6: Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114663",
          "createdAt": "2026-08-13T16:08:31Z",
          "updatedAt": "2026-08-13T16:09:08Z",
          "timestamp": "2026-08-13T16:09:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "JiaQiTang98",
            "groeneai"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:fab449c9af18f760b59c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114580",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114580",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: Duplicate TLS argument rejection and positional-arity stripping untested",
          "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #110615](https://github.com/ClickHouse/ClickHouse/pull/110615). That PR: (1) Adds TLS/SSL to every PostgreSQL integration: `sslmode` plus certificate/key either as server-local paths (`sslrootcert`/`sslcert`/`sslkey`, config-only) or as literal contents (`*_pem`, accepted from SQL, materialized into `TemporarySecretFile` and masked as secrets). New code: … **1. Duplicate TLS argument rejection and positional-arity stripping untested** `src/Storages/StoragePostgreSQL.cpp:726`, `src/Databases/PostgreSQL/DatabasePostgreSQL.cpp:573` **Risk:** `StoragePostgreSQL::extractSSLParamsFromArguments` strips trailing TLS `key = value` pairs and rejects repeats at `StoragePostgreSQL.cpp:726-727`; the stripped list then feeds the arity check at `DatabasePostgreSQL.cpp:573`. Risk if broken: a repeated `sslmode = 'require', sslmode = 'disable'` … **Unique vs PR tests:** 04820 formats queries with distinct TLS keys and checks masking plus path rejection; 04846 checks query-tree masking; test_postgresql_ssl exercises real handshakes through named collections. None repeats a TLS key (the `specified more than once` branch) and none combines the maximum positional … **Tags:** `-- Tags: no-fasttest` — `no-fasttest`: the PostgreSQL integration is not built in the fast test build; the PR's own 04820 uses the same tag cc @alexey-milovidov (author of #110615) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114580",
          "createdAt": "2026-08-13T03:32:59Z",
          "updatedAt": "2026-08-13T16:08:46Z",
          "timestamp": "2026-08-13T16:08:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e2dfba1b81b0d58955c7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110997",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110997",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix data race on DataTypeAggregateFunction version during Native serialization",
          "text": "Related: found by the `arm_tsan` and `azure, amd_tsan` Stress tests (STID 3977-4818, ThreadSanitizer data race). No existing issue. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results reading `AggregateFunction` states after one client requested them at an older protocol revision. The serialization version chosen for that one response was written onto the data type object shared by the whole table, so it stayed there: every later query read the states at that version, and for `sumMap` over `Decimal32` that returns wrong values, while the column also lost the version in `system.columns` and on the wire. The same in-place write was a data race between concurrent queries serializing such a column in the `Native` format. ### Description A single `DataTypeAggregateFunction` instance is shared across query result blocks: it lives once in the table's column description and is aliased by shallow column copies. `NativeWriter`/`NativeReader` called `setVersionToAggregateFunctions`, which walked to the leaf type and wrote its `mutable version` field in place. Two concurrent `Native` serializations of the same aggregate-function-typed column then raced on that field. Reports: * https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109496&sha=62eeb400eafe86cedff9b5a4e9a36aa722633cda&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29 * https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b59441bd06c2fcb6b103a30874528cc398afc723&name_0=MasterCI&name_1=Stress%20test%20%28azure%2C%20amd_tsan%29 Both racing stacks are `setVersionToAggregateFunctions` -> `DataTypeAggregateFunction` version setter, via `NativeWriter::write` -> `TCPHandler::processOrdinaryQuery`/`sendData`. Fix: instead of mutating the shared type object, replace the versioned leaf with a copy that carries the version via the constructor (the same way the binary-encoding decode path builds versioned types). The in-place `setVersion`/`updateVersionFromRevision` mutators and the `mutable` qualifier are removed so the object is immutable after construction. This also removes a latent issue that was worse than the race itself. `NativeWriter` passes `if_empty = false` for a client older than `DBMS_MIN_REVISION_WITH_AGGREGATE_FUNCTIONS_VERSIONING`, which unconditionally forced version 0 onto the shared type. Every later query then kept that 0 (`if_empty = true` sees a version already set), so it advertised a type name without a version while serializing version-0 states, and the receiving client - deriving the version from the server revision - deserialized them as version 1. #### Preserving custom type names Because the leaf is now replaced rather than mutated, the rebuilt tree is what the caller ends up with, so the rebuild must not lose anything. Rebuilding the wrappers through `transformTypesRecursively` recreates `Array`/`Tuple`/`Map` via `make_shared` and drops custom type names: | type | expected | with a naive rebuild | |---|---|---| | `Nested(x AggregateFunction(sumMap, ...))` | preserved | `Array(Tuple(...))` | | `SimpleAggregateFunction(anyLast, AggregateFunction(sumMap, ...))` | preserved | `AggregateFunction(...)` | Both are user-visible: the type is sent to the client over `Native`, and on `ATTACH` it becomes the column type in the table metadata. Losing the `SimpleAggregateFunction` name is worse than cosmetic - `AggregatingSortedAlgorithm` and `SummingSortedAlgorithm` recognise such a column by `dynamic_cast` on that very name object, so the column would silently start merging as a plain aggregate function state. So `setVersionToAggregateFunctions` walks the type itself, over exactly the wrappers `transformTypesRecursively` descended into, and returns the original pointer when no leaf changes. `Nullable` is among them: a state cannot be directly inside `Nullable`, but a `Tuple` can, and `Nullable(Tuple(AggregateFunction(...)))` is reachable with `enable_nullable_tuple_type`. A custom name can also sit on the wrapper rather than on the leaf, as in `SimpleAggregateFunction(anyLast, Array(AggregateFunction(...)))`, so a rebuilt wrapper carries the customization of the original too. `DataTypeCustomNamePtr` becomes a `shared_ptr` so a copy of a type can carry the very same custom name object, via the new `IDataType::cloneCustomization`. `Nested` is rebuilt with its custom name kept in sync with the new element types, directly rather than through `createNested`: the latter derives the type from the printed name, and version 0 is deliberately not printed, so a name round trip would turn a leaf explicitly pinned to version 0 back into an unversioned one using the latest version. `callOnNestedSimpleTypes` had no other caller and is removed, so `transformTypesRecursively` (shared with schema inference) is left untouched. ### Testing * `gtest_aggregate_function_version_race` covers the shared-object mutation, nested types, both custom-name cases above, and stress-tests concurrent version assignment over one shared type object. * `04612_aggregate_function_version_custom_type_names` round-trips both types through `Native` (which assigns the version in the writer and again in the reader) and through `DETACH`/`ATTACH`. * `04613_aggregate_function_version_not_sticky` is the one that fails on `master` HEAD. The two tests above assert output that is byte-identical to `master` by design, so neither can. It asks for a `Native` response at a revision below the one that introduced versioning, then checks the column again: on `master` the version 0 forced for that one response stays on the shared type, so a later plain `SELECT finalizeAggregation(s)` reads `([1,2],[10.5,20.25])` back as `([1],[10.5])` and `system.columns` loses the version. The revision is pinned explicitly and the type name the response carries is asserted, so the test cannot pass without that path having run - checked against both ways of it not running, a request that fails and a version assignment that does nothing. * The serialized version bytes are unchanged; the `Native` wire type names are byte-identical to `master` for both types. * 472 related stateless tests (`simple_aggregate`, `nested`, `native`, `geo`, `point`, `polygon`, `aggregate_function`) were run against this build and against a `master` build on the same server config: the failure sets are identical, i.e. no test fails only with this change.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110997",
          "createdAt": "2026-07-19T14:41:39Z",
          "updatedAt": "2026-08-13T16:08:38Z",
          "timestamp": "2026-08-13T16:08:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 24
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1ec129f1637a8a017d9f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114662",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114662",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113450 to 26.5: Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114662",
          "createdAt": "2026-08-13T16:08:00Z",
          "updatedAt": "2026-08-13T16:08:37Z",
          "timestamp": "2026-08-13T16:08:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "JiaQiTang98",
            "groeneai"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:12039a7b706ed64c4c13",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114525",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114525",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Optimize merges of the text index",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved performance of merges of text indexes. ### Additional context A few optimizations: - The main one: reducing overhead on deserialization of embedded and small postings caused by the allocation of the bitmap - Removed unneeded conversion to roaring bitmap on build of the output posting list - Used specialized sort cursor and batch sorting strategy for merging of text index segments",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114525",
          "createdAt": "2026-08-12T17:15:12Z",
          "updatedAt": "2026-08-13T16:08:10Z",
          "timestamp": "2026-08-13T16:08:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:faef1ca028cf61d13d52",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111457",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111457",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Compare read-in-order virtual row on its covered sort-key prefix",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/106740 Closes: https://github.com/ClickHouse/ClickHouse/issues/106630 Related: https://github.com/ClickHouse/ClickHouse/pull/110725 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a wrong result (mis-ordered merge) for read-in-order queries with a virtual row when `distinct-in-order` or `LIMIT BY` widens the read to a longer sort-key prefix than `ORDER BY` set it up for, and when a key column fixed by the filter is skipped by `ORDER BY` (e.g. `WHERE b = 1 ORDER BY a, c` on key `(a, b, c)`). The virtual row announced a wrong merge boundary: in release builds the merge could be silently mis-ordered, in debug builds the boundary assertion fired, and a `Nullable` key column after the skipped one threw the `Virtual row has different type` exception. The virtual row optimization now stays enabled in these cases. ### Description Alternative to https://github.com/ClickHouse/ClickHouse/pull/110725: instead of dropping the virtual row conversion when the in-order read prefix changes (and disabling it for skipped key columns), keep the optimization enabled and compare the virtual row only on the sort-key prefix it validly covers. **Root cause.** The read-in-order virtual row announced a wrong merge boundary in two ways: 1. *Widened prefix.* `optimizeReadInOrder` builds the virtual row conversion for the prefix `ORDER BY` needs (e.g. `CounterID`). A later optimization (`optimizeDistinctInOrder`, `optimizeLimitByInOrder`) re-requests the read with a longer prefix (`CounterID, EventDate`), but the `pk_block` width was derived from the conversion's input count, so the extra sort column was default-filled with `0` in `setVirtualRow`. In reverse order `0` understates the real values, so a real row exceeded the announced boundary: `Virtual row boundary violated in MergingSortedAlgorithm ... the virtual row announced UInt64_0 but the source then produced UInt64_1` in debug builds, a silently mis-ordered merge in release builds. 2. *Skipped fixed key.* For key `(a, b, c)` and `WHERE b = 1 ORDER BY a, c`, the fixed key `b` is skipped without an `ORDER BY` counterpart, but the conversion DAG indexed key columns densely, mapping `c` onto key column `b` (visible in `EXPLAIN actions=1`: input `b` aliased to `__table1.c`). The wrong value tripped the boundary check; a wrong type (`Nullable` key) threw the `Virtual row has different type` logical error even in release builds. Moreover, index values of the columns after the skipped key are semantically unusable: the index describes pre-filter data, so the entry `(5, 0, 9)` does not bound the filtered row `(5, 1, 3)` projected to `(a, c)`. **Fix.** - The virtual row conversion outputs only the sort-description prefix it can announce exactly: index values while the key prefix is contiguous, plus constants for fixed columns from `ORDER BY`. A skipped fixed key column ends the index-backed part: the index entry at a mark boundary may hold a filtered-out value for it, so its later components bound nothing in the filtered stream. A column fixed by the filter that stays in `ORDER BY` keeps disabling the virtual row, as before this fix. - The merge compares a virtual row only on the covered prefix and places it first on a covered-prefix tie (equivalent to treating the uncovered columns as minus infinity in the merge order, without materializing any values). The covered prefix is derived from the pk block column names in `MergingSortedAlgorithm` and carried per cursor in `SortCursorImpl::sort_prefix_limit`, honored by the generic `SortCursor::greaterAt`. A truncated virtual row can only occur with a multi-column sort description (its coverage is at least the first column), which always uses the generic cursor, so the single-column specialized queues are unaffected; the JIT comparator is bypassed when a truncated cursor participates. - `ReadFromMergeTree::readInOrder` reads index values for the whole used sorting-key prefix instead of only the conversion inputs. The in-order merges inside the read step sort by the full prefix, and index values are exact bounds for it even after filtering (a filter only removes rows), so these merges always see fully covered virtual rows. - `ReadFromMergeTree::requestReadingInOrder` drops the conversion only when a re-request makes it unsound: a prefix narrower than the one it was built for (the conversion could lose its inputs), or one not fully backed by the primary index. - `setVirtualRowConversions` builds the conversion with `project_inputs` so a raw index column cannot shadow a same-named conversion output when the merge looks sort columns up by name (matters with the old analyzer). This handles all key types uniformly — e.g. a descending `String` column after a skipped key keeps the optimization even though the type has no greatest value to pad with. Verified on release and debug builds (the boundary assertion is compiled only into debug builds): the previously aborting repros now return correct results, and `EXPLAIN` keeps `Virtual row conversions` for the widened and skipped-key reads. The test is based on the one from https://github.com/ClickHouse/ClickHouse/pull/110725, extended with checks that the optimization stays enabled, a descending skipped-key case, a `String` case in both directions, and a fixed key kept in `ORDER BY` (still disabled, as before).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111457",
          "createdAt": "2026-07-22T18:29:12Z",
          "updatedAt": "2026-08-13T16:06:50Z",
          "timestamp": "2026-08-13T16:06:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:709cfdae8869dce45ebc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114220",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114220",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.6: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114220",
          "createdAt": "2026-08-10T20:07:09Z",
          "updatedAt": "2026-08-13T16:05:26Z",
          "timestamp": "2026-08-13T16:05:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3932e025ce570cf40502",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114300",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114300",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "TimeSeries: store all tags in the `tags` column",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): TimeSeries: store all tags in the `tags` column The `tags` column of the tags target table now contains all the tags, including the `__name__` tag with the metric name and the tags with dedicated columns from the `tags_to_columns` setting, so the map alone fully identifies a time series. The default id generator now hashes just `tags`, i.e. now it looks like `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(tags)))` (instead of `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(metric_name, all_tags)))` )",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114300",
          "createdAt": "2026-08-11T10:29:30Z",
          "updatedAt": "2026-08-13T16:05:22Z",
          "timestamp": "2026-08-13T16:05:22Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-promql"
          ],
          "author": "vitlibar",
          "state": "open",
          "assignees": [
            "nikitamikhaylov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d2ee8ff803108df31453",
        "signalId": "github:ClickHouse/ClickHouse:issue:114026",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114026",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "`rewrite_in_to_join`: `PREWHERE x IN (subquery)` throws `Unknown function exists` while the `WHERE` spelling works (also hit via `make_distributed_plan`)",
          "text": "**Describe what's wrong** With `rewrite_in_to_join = 1`, any non-constant `IN (subquery)` predicate placed in `PREWHERE` makes the query fail with `Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION)`. The same predicate in `WHERE` works and returns the correct result. `NOT IN`, tuple `IN`, and the same query with an explicit `JOIN` fail the same way; `GLOBAL IN` and scalar subqueries (`PREWHERE s = (SELECT ...)`) are unaffected. This is hit at default settings by users of `make_distributed_plan`: since `allow_experimental_correlated_subqueries` defaults to `1`, `adjustSettingsForMakeDistributedPlan` (`src/Core/SettingsQuirks.cpp`) silently force-enables `rewrite_in_to_join`, so `SELECT ... FROM distributed_table PREWHERE x IN (subquery) SETTINGS make_distributed_plan = 1` fails with the same bogus `Unknown function exists` instead of a declared unsupported-feature error (the `WHERE` spelling of that distributed query gets the honest `Code: 48. Correlated subqueries are not supported with remote tables`). **How to reproduce** Reproduces on current master (26.8.1.1008) and on every version tried back to 26.6, with `clickhouse local` (20/20 deterministic; the `WHERE` control arm is 20/20 correct): ```sql CREATE TABLE t (k UInt64, s String) ENGINE = MergeTree ORDER BY k; INSERT INTO t SELECT number, toString(number % 2) FROM numbers(1000); SELECT count() FROM t PREWHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1; -- Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION) SELECT count() FROM t WHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1; -- 500 (correct) ``` Required: `rewrite_in_to_join = 1` (or `make_distributed_plan = 1`, which force-enables it), the analyzer (`enable_analyzer = 1`, the default; the old analyzer is unaffected), a non-constant left-hand side, and the predicate in `PREWHERE`. **Mechanism** (from the stack trace on 26.8.1.1008) 1. Under `rewrite_in_to_join`, the analyzer rewrites `x IN (subquery)` into a special `exists(...)` `FunctionNode` (`src/Analyzer/Resolve/resolveFunction.cpp`, \"Replace IN (subquery)\"). `exists` is resolved by a dedicated special-function path — it is not registered in `FunctionFactory`. 2. `PREWHERE` resolution then clones the predicate tree and runs `ReplaceColumnsVisitor` to undo JOIN-changed column types (`src/Analyzer/Resolve/QueryAnalyzer.cpp`, \"Expressions in PREWHERE with JOIN should not change their type\"). That visitor calls `rerunFunctionResolve` on every resolved `FunctionNode` it visits. 3. `rerunFunctionResolve` (`src/Analyzer/Utils.cpp`) resolves ordinary functions through `FunctionFactory::instance().get(name, context)`. It already special-cases `grouping` (also resolved outside the factory) but not `exists`, so the lookup throws `UNKNOWN_FUNCTION`: ``` 4. src/Functions/FunctionFactory.cpp:89: DB::FunctionFactory::getImpl(String const&, ...) 5. src/Functions/FunctionFactory.cpp:108: DB::rerunFunctionResolve(DB::FunctionNode*, ...) 6. src/Analyzer/InDepthQueryTreeVisitor.h:62: DB::QueryAnalyzer::resolveQuery(...) ``` **Expected behavior** The query returns the same result as its `WHERE` spelling (or, where the rewritten form is genuinely unsupported, a declared `SUPPORT_IS_DISABLED`/`NOT_IMPLEMENTED` error naming the feature — not `Unknown function exists`). Related: https://github.com/ClickHouse/ClickHouse/issues/109476 Related: https://github.com/ClickHouse/ClickHouse/issues/109326 Related: https://github.com/ClickHouse/ClickHouse/issues/102630 Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114026",
          "createdAt": "2026-08-09T10:21:45Z",
          "updatedAt": "2026-08-13T16:04:02Z",
          "timestamp": "2026-08-13T16:04:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug",
            "comp-joins",
            "comp-query-optimizer",
            "comp-query-analyzer"
          ],
          "author": "zlareb1",
          "state": "closed",
          "assignees": [
            "KochetovNicolai"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:0e14495ad253ec99ccba",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114067",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114067",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix(Analyzer): skip rerunFunctionResolve for 'exists' nodes created by rewrite_in_to_join",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114026 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed `Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION)` thrown when a `PREWHERE` clause contains an `IN (subquery)` predicate and `rewrite_in_to_join = 1` (or `make_distributed_plan = 1`, which force-enables it) is set. The same query spelled with `WHERE` worked correctly. --- ### Problem With `rewrite_in_to_join = 1`, the analyzer rewrites `x IN (subquery)` into an `exists(...)` `FunctionNode` resolved via `FunctionExists`, a special function that is not registered in `FunctionFactory`. When `PREWHERE` is resolved, `ReplaceColumnsVisitor` calls `rerunFunctionResolve` on every `FunctionNode` in the predicate. `rerunFunctionResolve` (`src/Analyzer/Utils.cpp`) already special-cases `grouping` (also resolved outside the factory), but not `exists`, so it called `FunctionFactory::instance().get(\"exists\", context)` and threw `UNKNOWN_FUNCTION`. ### Fix Two parts: 1. `PREWHERE` is evaluated by the reading step and cannot execute a correlated subquery — the planner rejects one with `ILLEGAL_PREWHERE`. So the `rewrite_in_to_join` rewrite is now skipped while resolving a `PREWHERE` expression, and the plain `IN` is kept there. `PREWHERE x IN (subquery)` then returns the same result as its `WHERE` spelling, which is what the issue asks for. Subqueries nested inside `PREWHERE` still rewrite their own `IN` predicates. 2. `exists` is added to the special-case early return in `rerunFunctionResolve`, next to `grouping`. This matters for an explicitly written `PREWHERE EXISTS (correlated subquery)`, which is genuinely unsupported: it is now reported honestly as `ILLEGAL_PREWHERE` instead of `Unknown function exists`. ### Test `tests/queries/0_stateless/04820_rewrite_in_to_join_prewhere_exists.sql` covers `PREWHERE ... IN (subquery)` against the `WHERE` control arm, `NOT IN`, tuple `IN`, an `IN` nested in a subquery inside `PREWHERE`, and the explicit `PREWHERE EXISTS (...)` case asserting `ILLEGAL_PREWHERE`. ### Reproduction ```sql CREATE TABLE t (k UInt64, s String) ENGINE = MergeTree ORDER BY k; INSERT INTO t SELECT number, toString(number % 2) FROM numbers(1000); -- Threw: Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION) SELECT count() FROM t PREWHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; -- Worked (control) SELECT count() FROM t WHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; ```",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114067",
          "createdAt": "2026-08-09T20:05:33Z",
          "updatedAt": "2026-08-13T16:04:01Z",
          "timestamp": "2026-08-13T16:04:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "RohithPariki",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5cb7aec0f0f55367def1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113899",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113899",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Avoid scans for constant sort keys",
          "text": "## What `isAlreadySorted` now returns early, right after the sort descriptors are resolved, when every sort key is a `ColumnConst`. Writing a `MergeTree` part whose sorting keys are all constant no longer walks the block doing adjacent-row comparisons to confirm an ordering that is constant by construction. Collation validation still runs before the early return, and any block that mixes constant and non-constant keys keeps the existing comparator path — only the all-constant case takes the new path. ## Why it helps When the sorting-key columns are constant across a block — a common shape when a leading `ORDER BY` column is fixed per part (time-ordered or batched ingestion, per-source or per-partition writes) — the sortedness check was doing a full comparison pass to reach a foregone conclusion. Returning as soon as the keys are known-constant removes that pass. Measured on `MergeTree inserts with constant and mixed sorting keys`, 64K/256K/1M rows (paired medians, co-measured on both trees): | Metric (1M rows) | Before | After | Δ | | --- | ---: | ---: | ---: | | Constant-key sortedness check, 1 key | 1106 µs | 2 µs | **−99.8%** | | Constant-key sortedness check, 4 keys | 4449 µs | 4 µs | **−99.9%** | | End-to-end insert latency, 4 keys | 24548 µs | 20201 µs | **−17.7%** | | Insert CPU, 4 keys | 20732 µs | 16301 µs | **−21.4%** | Smaller block sizes land in the same range (e.g. 1-key 64K: 66 µs → 2 µs). The sortedness check collapses to a near-constant cost, and that saving carries into a full insert as the end-to-end latency and CPU gains. Blocks that aren't all-constant take the unchanged comparator path. ## Testing Stateless coverage exercises single and multiple constant keys plus a non-constant suffix (the mixed case that must keep comparing). The declared correctness check passed. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Skip the redundant sortedness scan when writing MergeTree parts whose sorting keys are all constant. --- Contributed by [Perfloop](https://app.perfloop.ai): the numbers above were co-measured on both trees and independently re-verified before submission — the full public record is at [case_8ebwekrder](https://app.perfloop.ai/t/oss/case_8ebwekrder). Replies from this account are human-approved, and a human operator is accountable for this contribution.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113899",
          "createdAt": "2026-08-07T23:19:06Z",
          "updatedAt": "2026-08-13T16:04:01Z",
          "timestamp": "2026-08-13T16:04:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "perfloop-agent",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ffbe971d7dfc5e3e6e95",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109299",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109299",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Respect `date_time_overflow_behavior` for numeric temporal casts",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/101131 Related: https://github.com/ClickHouse/ClickHouse/pull/110459 Related: https://github.com/ClickHouse/ClickHouse/pull/101512 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Makes `date_time_overflow_behavior` take effect for numeric casts to `Date`, `Date32`, `DateTime` and `Time`. The `throw` and `saturate` paths were dead code for these targets, so an out-of-range value was always silently reinterpreted or truncated. ### Description `ConvertImpl` instantiated the numeric-to-temporal transforms with the compile-time constant `default_date_time_overflow_behavior` (`ignore`) instead of the runtime setting, so their `throw` and `saturate` branches were never instantiated. Threading the runtime value in makes them reachable, exposing three latent defects there: the rejected value was formatted with a narrowing `static_cast<Int64>`, undefined for a huge, infinite or `NaN` float (floats are now widened to `double`); `NaN` passed every range comparison into a narrowing cast, so it is guarded explicitly; and the clamp narrowed to `time_t` before `std::min`, so a `UInt64` above `INT64_MAX` wrapped negative and gave `1970` instead of the maximum (now clamped in the source domain first). The wide integer types missed every branch of the `DateTime` dispatch and fell through to `convertNumericGeneral`, which truncates and ignores the setting; they now use the overflow-aware transforms, so `toDateTime32(toInt128(99999999999999999999999999))` saturates instead of returning `1970-02-04`. The numeric `Time` dispatch and its bounds already landed on the base as a714d76b341fa8c from an earlier round here. `convertFieldToType` (the `VALUES`/`IN` coercion path) ignored the setting too. Serving two kinds of caller, in `ignore` mode it keys on the exactness flag: an exact target (`DROP`/`OPTIMIZE PARTITION`, strict `IN`, `KeyCondition`, sharding key) gets the canonical storage value, so an unstorable literal raises `ARGUMENT_OUT_OF_BOUND` instead of addressing a clamped partition; value materialization clamps like `CAST` (`Date`/`Date32` now clamp where they returned `NULL`). `values()` is the exception: built with a default `FormatSettings`, it keeps clamping while `CAST` raises under `throw`. The test pins that divergence. `Date32` was the only temporal target that accepted a non-finite float, saturating `inf`/`NaN` to a boundary date in every mode while the other three raise `CANNOT_CONVERT_TYPE`; it now rejects them too, on both the `CAST` and the materialization side. PR 110459 (merged) is already reconciled in this branch.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109299",
          "createdAt": "2026-07-03T14:13:03Z",
          "updatedAt": "2026-08-13T16:03:53Z",
          "timestamp": "2026-08-13T16:03:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f3241e0779a5e96b567f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113022",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113022",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a type-aware Bloom filter index for JSON",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 The original design used `JSONAllValues` as the input to a Bloom filter. `JSONAllValues` serializes each value as text. It does not preserve the runtime type. This loss of type information is important for `Dynamic` values. ClickHouse can compare JSON values with different runtime types. Some type pairs can match after conversion. Other type pairs can return an exception. A Bloom filter that stores only text cannot safely model these rules. It can skip a granule that contains a match. It can also hide an exception. `JSONAllValues` remains useful for text search, but it is not a safe base for this index. This PR replaces that design with `jsonbf_v1`. The new index creates tokens from these components: - JSON path - Container role - Runtime type - Binary value The container role separates scalar values, array elements, and map values. The index also processes nested JSON objects and named tuple fields. For `Dynamic` values, the index stores type-presence and complex-value presence tokens. Query analysis uses exact value tokens only when the comparison is safe. If analysis cannot prove safety, ClickHouse reads the granule. Container roles are preserved through nested tuples, and nested casts are handled conservatively. The index supports: - Equality and typed `IN` conditions - `has`, `hasAny`, and `hasAll` for arrays - Typed map values by key - Nested JSON objects and tuples The index does not optimize range conditions or whole-container equality. It also rejects unsafe comparison paths, such as Decimal-to-Float comparisons. An unsupported `Dynamic` runtime type disables skipping for its granule. In a one-million-row JSONBench test, the index reduced reads from 123 granules to 12–15 granules. Selected string equality queries were 1.6–2.3 times faster. Indexed inserts were approximately 2.9 times slower in the checked-in performance test. A local performance test produced these median results: | Operation | Without index | With `jsonbf_v1` | Difference | |---|---:|---:|---:| | Numeric equality, matching value | 13.05 ms | 10.99 ms | 1.19 times faster | | Numeric equality, missing value | 13.25 ms | 10.80 ms | 1.23 times faster | | String equality | 24.49 ms | 12.90 ms | 1.90 times faster | | Array `has` | 24.71 ms | 12.78 ms | 1.93 times faster | | Insert | 20.51 s | 40.87 s | 1.99 times slower | The small performance test shows limited benefit for numeric scalar equality and larger improvements for string equality and array membership. The larger JSONBench data set benefits more because the index skips more granules. Token generation and Bloom-filter construction increase insert time. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds the `jsonbf_v1` data-skipping index for type-aware equality and array membership on `JSON` values.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113022",
          "createdAt": "2026-08-02T18:31:45Z",
          "updatedAt": "2026-08-13T16:03:15Z",
          "timestamp": "2026-08-13T16:03:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f1ae479edf20463a6357",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114646",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114646",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fail-closed allowlist of hypothetical index types for WHATIF",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114646",
          "createdAt": "2026-08-13T14:06:56Z",
          "updatedAt": "2026-08-13T16:02:42Z",
          "timestamp": "2026-08-13T16:02:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "yariks5s",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:4a93bd9b140b5539703a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113983",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113983",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allowlist the expected FileLog bad-path reattach error in the upgrade check",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113781 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Upgrade check (amd_release)` intermittently fails its `Error message in clickhouse-server.log` sub-test on one benign line: ``` <Error> StorageFileLog (test_1.filelog_bad_path_attach): The absolute data path should be inside `user_files_path`(/var/lib/clickhouse/user_files/) ``` No product defect: the server starts, nothing crashes, no data is affected. Root cause. `04202_filelog_attach_path_outside_user_files` ATTACHes a FileLog table whose path is outside `user_files_path`. `ATTACH` is `LoadingStrictnessLevel::ATTACH` (2), which is `>= SECONDARY_CREATE` (1), so the constructor takes the relaxed branch at `src/Storages/FileLog/StorageFileLog.cpp:195-198`: it logs at `<Error>` and returns instead of throwing `BAD_ARGUMENTS`. That branch is deliberate and is what the test covers, since refusing to load at reattach time would break server startup. The table then outlives the test: stress threads run with a fixed `--database=test_N` (`ci/jobs/scripts/stress/stress.py`), and `clickhouse-test` skips its per-test teardown whenever `--database` is set (`need_cleanup = not args.database`), so that shared database is never dropped. The upgrade restart re-attaches the table, the relaxed branch fires again, and the line lands in the scanned log, where the post-restart scrub in `tests/docker_scripts/upgrade_runner.sh` had no entry for it. Hence the intermittency: `04202` must land on a fixed-database thread. Change. One `grep -av` entry in that scrub's existing secondary pipe, plus a short rationale comment next to the sibling entries. The pattern requires the fixture table name and the message together, and (bare parens are literals in BRE) the `StorageFileLog (db.table):` prefix shape. No source change, no test change. Validation. The scan pipeline, extracted verbatim from the runner, was run under GNU grep 3.11 against the failing run's own 19.9 MB `clickhouse-server.upgrade.log`. With the entry the artifact is empty; with it deleted the output is byte-identical to the 189-byte `upgrade_error_messages.txt` CI produced, so the sub-test flips `FAIL` to `OK`. Eight negative controls still surface, including a table whose name merely ends with the fixture name (`prod.other_filelog_bad_path_attach`), which the required `.` separator keeps visible.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113983",
          "createdAt": "2026-08-08T21:17:46Z",
          "updatedAt": "2026-08-13T16:01:55Z",
          "timestamp": "2026-08-13T16:01:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b99c4b94111718ebf6e3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:80353",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:80353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Redis-wire protocol",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Add an opt-in Redis wire-protocol server backed by `Join` tables: a configured `redis.port` serves `GET`/`MGET` and `HGET`/`HMGET` point lookups (plus `AUTH`, `SELECT`, `PING`, `ECHO`) against ClickHouse tables mapped to Redis database numbers. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/80353",
          "createdAt": "2025-05-16T14:55:43Z",
          "updatedAt": "2026-08-13T16:01:26Z",
          "timestamp": "2026-08-13T16:01:26Z",
          "metrics": {
            "reactions": 4,
            "comments": 10
          },
          "labels": [
            "pr-feature",
            "comp-protocols",
            "manual approve",
            "can be tested"
          ],
          "author": "m4ttheux",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:fba28ce5933e09068657",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110104",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110104",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check for cancellation in AggregatingInOrderTransform",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107941 --> Related: https://github.com/ClickHouse/ClickHouse/issues/107941 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `optimize_aggregation_in_order` ignoring query cancellation. `AggregatingInOrderTransform` now checks for cancellation while aggregating a chunk, so a query stopped by `KILL QUERY` or by `max_execution_time` (in the default `timeout_overflow_mode = 'throw'`) stops promptly instead of running the whole chunk to completion. ### Description `AggregatingInOrderTransform::consume()` splits one input chunk into runs of equal keys in a loop. Query time and cancellation limits are only checked between pipeline steps (between `work()` calls), so a chunk with many distinct keys makes a single `consume()` call run for a long time (`O(distinct_keys)` iterations, each an `upper_bound` over the remaining rows) with no cancellation checkpoint. As a result a cancelled query (`KILL QUERY`, `max_execution_time`) using `optimize_aggregation_in_order` kept aggregating until the whole chunk was done; the connection thread then blocked in `PullingAsyncPipelineExecutor::cancel() -> ThreadFromGlobalPool::join()` waiting for that loop. The server-side AST fuzzer repeatedly hit this as `Hung check failed, possible deadlock found` (Stress test, all sanitizers), with `system.processes` showing `is_cancelled = 1` and `elapsed` far past the 90s hung-check window while the worker thread sat in `AggregatingInOrderTransform::consume -> Aggregator::executeImpl`. The loop now checks `isCancelled()` once per key interval (cheap) and returns early; the partial aggregation state is discarded because the pipeline is being torn down. This mirrors the existing per-loop cancellation checks in `WindowTransform` and `FillingTransform`. Scope: this covers cancellation that sets `is_cancelled` on the pipeline, i.e. `KILL QUERY` and `max_execution_time` in the default `timeout_overflow_mode = 'throw'`, which is what the reproduced hung check hit (`system.processes` showed `is_cancelled = 1`). The non-default `break` mode is a soft limit that returns a partial result and never sets `is_cancelled`; honoring it mid-chunk (as `FillingTransform` does via `process_list_element->checkTimeLimit()`) is a separate partial-result change, out of scope here. The regular `AggregatingTransform` behaves the same way. Regression test `04512_aggregation_in_order_cancellation` forces one long `consume()` over 40M distinct-key rows in a single chunk, `KILL QUERY ... SYNC` once every row is read: with the fix the KILL returns in a fraction of a second, without it it blocks for the several seconds the loop needs to finish.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110104",
          "createdAt": "2026-07-11T17:17:10Z",
          "updatedAt": "2026-08-13T16:00:26Z",
          "timestamp": "2026-08-13T16:00:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "yakov-olkhovskiy"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d6becb7598db54bad37f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109367",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109367",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Scheduler: reclaimable memory tracking and dynamic spilling",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/109064 --> Scheduler-side support for reclaimable memory tracking and dynamic spilling, built on top of the memory reservations subsystem. The design is described in #109064. Queries can report the portion of an allocation that can be spilled or discarded on request (`IAllocationQueue::setReclaimable`), which is aggregated bottom-up as a new `reclaimable` field on every scheduler node. Each `AllocationLimit` gains a soft limit: when a workload's allocated memory exceeds it and the subtree has reclaimable memory, the scheduler asks a victim to reclaim memory (`ResourceAllocation::spillAllocation`, mirroring `killAllocation`) instead of waiting for the hard `max_memory` limit to force a kill. The request is replied: the query finishes it with `IAllocationQueue::finishSpill` after issuing the decreases for the freed memory (zero reclaimable doubling as a decline), the reply travels to the root as `Update::spilled` on the same propagation path as the state it describes, and at most one spill request is outstanding per subtree until it arrives; a victim that leaves the queue counts as having replied. Victim selection is deterministic and matches the kill order (largest usage, least precedence, largest allocation), descends a single root-to-leaf path via reclaimable-filtered ordered sets, and is fail-close: with nothing reclaimable, or with no soft limit configured, behavior is exactly as before. The soft limit is configured per workload via two new settings: `max_memory_before_spill` (absolute) and `max_memory_to_spill_ratio` (a fraction of the workload's own `max_memory`), smaller wins. The new state is observable in `system.scheduler`: `reclaimable`, `spills`, and the effective `soft_limit`. This is the scheduler side only. `MemoryReservation::spillAllocation` is currently a no-op; the query side that reports reclaimable memory and reacts to spill signals is a separate change. Documentation for the scheduler-side workload settings (`max_memory_before_spill`, `max_memory_to_spill_ratio`) and a `Spilling reclaimable memory` section are included here; the set of operators that can spill will be documented together with the query-side reaction. Invariants for the new machinery are documented on `ISpaceSharedNode` (I1-I8). Added unit tests cover fail-close, largest-first selection, fair descent skipping unreclaimable subtrees, the gate held until the victim's reply, the reply deferred to a pending decrease, re-signalling until under the soft limit, clamping reclaimable on a shrink, a declining victim, a victim removed mid-spill, the soft-limit enable transition via `CREATE OR REPLACE WORKLOAD`, the settings (absolute, ratio, and smaller-wins combination), a concurrency stress, and a parametrized throughput test. Verified under ThreadSanitizer (145 scheduler/workload gtests, 0 data races). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added workload settings `max_memory_before_spill` and `max_memory_to_spill_ratio` for the experimental memory reservation scheduling. They configure a soft memory limit above which a workload's queries are asked to spill reclaimable memory, before the hard `max_memory` limit forces an eviction. This is the scheduler-side foundation: queries do not yet report reclaimable memory or react to spill requests, so the settings have no effect until the query-side change lands. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109367",
          "createdAt": "2026-07-03T21:49:24Z",
          "updatedAt": "2026-08-13T15:59:55Z",
          "timestamp": "2026-08-13T15:59:55Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "serxa",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:27a37630ad215391a7fd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114484",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114484",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Document that PREWHERE filters one join input before the JOIN",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required for a documentation change. ### Description `prewhere.mdx` did not mention `JOIN` at all. This is the documentation KochetovNicolai asked for when he closed issue 89097 as not-a-bug: \"We need to document this.\" One `<Note>`, next to the existing note that documents the same class of fact for `FINAL`. It states the rule and gives the equivalent explicit spelling as a filtered subquery. The `SELECT` clause list is annotated to agree with it. Measured on the example as it appears on the page: `PREWHERE b.y > 50` returns `[1,2,3,4]`, the filtered-subquery spelling the same, the `WHERE` spelling `[1,4]`. The `WHERE` result is unchanged whether `query_plan_filter_push_down` is on or off, which is why the note credits `WHERE` to the join result rather than to a fixed position, per PedroTadim's review. The note claims nothing about individual join kinds or strictnesses, since `FULL JOIN` and `ASOF INNER JOIN` both differ too, and it attributes the fill value to [join_use_nulls](https://clickhouse.com/docs/reference/settings/session-settings/join#join_use_nulls) rather than naming one. Related: https://github.com/ClickHouse/ClickHouse/issues/89097 Related: https://github.com/ClickHouse/ClickHouse/issues/114206 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1341` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114484",
          "createdAt": "2026-08-12T13:01:26Z",
          "updatedAt": "2026-08-13T15:59:44Z",
          "timestamp": "2026-08-13T15:59:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:83e074d55fcc6dd43f0a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101791",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101791",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "In case of trivial views, push whole outer query to shards.",
          "text": "### Changelog category: - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): In case of trivial views over distributed table push whole outer query to shards. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Description / Proposed Solution When a VIEW is defined over a Distributed table, ClickHouse traditionally executes it on the shards without enclosing outer query. This means filters and expressions declared in the outer query are evaluated on the coordinator after pulling raw data from shards. For views whose body is a plain SELECT (column references, *, or arbitrary expressions — but no aggregation, grouping, ordering, joins, window functions, or scalar subqueries) over a single Distributed table, we can do better: inline the view body as a subquery and hand the whole thing to StorageDistributed. Each shard then receives the full outer query with the view body inlined, evaluates it against its local table, and only ships the result back. A view qualifies as \"trivial\" if its inner query: - Has a single SELECT (no UNION) - Selects only column references, *, or expressions — but no window functions (require the full dataset) and no scalar subqueries in the SELECT list - Has no WITH, PREWHERE, GROUP BY, HAVING, QUALIFY, ORDER BY, LIMIT, LIMIT BY, DISTINCT, or ARRAY JOIN - Has no subqueries in the WHERE clause - Reads from exactly one table with no joins, no table functions, no FINAL, no SAMPLE - Is not a parameterized view and does not use SQL SECURITY DEFINER **The optimization can be disbaled by setting (enabled by default):** ``` SET optimize_trivial_view_pushdown_to_distributed = 0; ``` ### Example: Env setup: ``` create table x engine = MergeTree ORDER BY tuple() AS SELECT intDiv(number,100000) as a, number as b FROM numbers(1000000000); SET prefer_localhost_replica = 0; CREATE TABLE x_dist AS x ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), x); CREATE VIEW v_computed AS SELECT a + 1 AS x, b AS y FROM x_dist WHERE a != 0; ``` Performance: ``` :) SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: 6baedc55-c8e7-4b2b-9946-b0d828abaf25 ┌─plus(a, 1)─┬──────────sum(b)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └────────────┴─────────────────┘ 1 row in set. Elapsed: 17.199 sec. Processed 2.00 billion rows, 32.00 GB (116.29 million rows/s., 1.86 GB/s.) Peak memory usage: 38.58 MiB. :) SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 7c4f2854-3de2-4d40-b960-efcc44a7b26d ┌─────x─┬──────────sum(y)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └───────┴─────────────────┘ 1 row in set. Elapsed: 16.497 sec. Processed 2.00 billion rows, 32.00 GB (121.24 million rows/s., 1.94 GB/s.) Peak memory usage: 38.82 MiB. ``` Plan: ``` :) explain SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; EXPLAIN SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: e2d32d5f-7024-4e6a-b686-45c7c0d5f2ef ┌─explain──────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression (Before GROUP BY) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression (Before GROUP BY) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └──────────────────────────────────────────────────────────────────────────────────┘ :) explain SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; EXPLAIN SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 88213305-46a5-493e-862c-a80c675c9452 ┌─explain───────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes query planning/execution for a subset of views over `Distributed` tables and touches access checks/row policy enforcement and SQL SECURITY semantics, which can affect correctness and security-sensitive behavior. > > **Overview** > Adds a new default-on setting `optimize_trivial_view_pushdown_to_distributed` to inline *trivial* views over `Distributed` tables and push the full outer query down to shards, reducing coordinator-side filtering/processing and network transfer. > > Implements planner rewrites to swap the view table expression with an analyzed subquery, merge `FINAL`/`SAMPLE` modifiers, and preserve semantics by suppressing pushdown when the outer query contains non-deterministic functions, while also explicitly handling SQL SECURITY modes, row-policy injection/logging, and column-pruned privilege checks. > > Extends integration/stateless tests to cover modifier propagation, non-determinism suppression, row-policy enforcement, SQL SECURITY behavior, and interactions with `max_rows_to_read_leaf`. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 583e1e6c0e8e25081391d7a07af086c6f9888c6f. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101791",
          "createdAt": "2026-04-04T19:32:01Z",
          "updatedAt": "2026-08-13T15:59:36Z",
          "timestamp": "2026-08-13T15:59:36Z",
          "metrics": {
            "reactions": 3,
            "comments": 15
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "simonmichal",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:80cf796679bdfdc02d7c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114035",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114035",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not throw when comparing an Enum column with a non-member string literal",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111545 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `UNKNOWN_ELEMENT_OF_ENUM` being thrown when an `Enum` column is compared with a string literal that is not one of its members, even though `validate_enum_literals_in_operators` is disabled by default. Previously the query was rejected if the column was in the primary key or partition key, or merely carried a `minmax`, `set` or `bloom_filter` index, and `!=` was rejected for any `Enum` column. Closes [#111545](https://github.com/ClickHouse/ClickHouse/issues/111545). ### Description Closes: #111545 `validate_enum_literals_in_operators` is off by default and documents that non-member enum literals are then not validated, yet comparing an `Enum` column with one threw, and whether it threw depended on physical layout. Two independent causes: 1. **Index analysis.** `KeyCondition::extractAtomFromTree` converts the literal with the throwing `convertFieldToType`, leaving the `isNull()` handling right below it unreachable. Since `minmax` and `set` build a `KeyCondition` over the index expression, that one line served the primary key, the partition key and both indexes. `bloom_filter` repeated the defect at four of its own sites, as did the JSON-subcolumn helper shared by `bloom_filter` and `tokenbf_v1`. The conversions are now wrapped in a catch scoped with `isParseError` (following `ConditionSelectivityEstimator`), so a non-representable literal declines the atom (full scan, never a wrong answer) while fatal errors such as `MEMORY_LIMIT_EXCEEDED` still propagate. 2. **An inverted guard.** In `FunctionsComparison.h`, `!equals && not_equals` selects `notEquals` alone, yet the value it guards, `IsOperation<Op>::not_equals`, was written to serve `notEquals`, so the tolerant branch could only ever return 0. Deleting the guard is the fix: any rewrite that keeps a condition there breaks the four ordered operators, for which both traits are false. Two corrections to the report, both re-measured. `e < '4'` does not throw on a non-key column (the ordered operators all return 0), so there only `!=` was broken; and the column need not be in a key, since a `minmax`, `set` or `bloom_filter` index suffices. `e <=> '4'` and `isDistinctFrom(e, '4')` were affected too. No query that previously succeeded changes its answer, and declining loses no pruning: `LIKE` over an `Enum` key never pruned before, since the converted field is an `Int8` that the `like` atom handler rejects. The tests assert that each index is still used for representable literals, and that the setting still rejects every operator when enabled. <details><summary>Validation</summary> Every carrier measured on a debug build, before and after, each fixture carrying a positive control: | Carrier | Before | After | Control | |---|---|---|---| | primary key `ORDER BY e` | 691 | 0 | `= 'a'` -> 1 | | partition key, `LIKE '%Beta%'` | 691 | 500 | `use_partition_pruning = 0` -> 500 | | `minmax` index, no key | 691 | 0 | `use_skip_indexes = 0` -> 0 | | `set(0)` index, no key | 691 | 0 | `use_skip_indexes = 0` -> 0 | | `bloom_filter` index, no key | 691 | 0 | `use_skip_indexes = 0` -> 0 | | `bloom_filter` over `Array(Enum)`, `has` / `hasAny` / `hasAll` | 691 | 0 | `has(a, 'a')` -> 1 | | `bloom_filter` over `Map(Enum, ...)` keys, `mapContains` | 691 | 0 | `mapContains(m, 'a')` -> 1 | | `JSONAllPaths` `bloom_filter` / `tokenbf_v1` | 691 | 0 | `= 'y'` -> 1, index used | | `!=` on any table | 691 | row count | `NOT (e = 'zzz')` agrees | | `<=>`, `isDistinctFrom` | 691 | 0 / row count | member-literal arms | | `< <= > >=` on a non-key column | 0 | 0 | unchanged | | `IN ('4', 'a')` | 1 | 1 | unchanged (#72686) | Each of the four hunks was reverted independently: every one reddens its own carrier and only its own carrier. The tests assert index *usability* through `force_data_skipping_indices` rather than row counts alone, so neither an over-broad decline nor a missing one can pass silently; four mutations were run and each reddens a new assertion. `Nullable(Enum8)`, `Nullable(Enum16)` and `Enum16` keys covered, including a real `NULL` row (`NULL != '4'` stays `NULL`; `NULL <=> '4'` is 0). `LowCardinality(Enum)` is not covered because `DataTypeEnum` does not override `canBeInsideLowCardinality()`, so the type cannot be constructed. 50 randomized runs of both new tests passed 100/100; `01310_enum_comparison`, `03278_enum_in_unknown_value`, `04049_statistics_enum_invalid_value` and `04539_enum_string_search_dictionary` stay green. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114035",
          "createdAt": "2026-08-09T13:50:23Z",
          "updatedAt": "2026-08-13T15:59:32Z",
          "timestamp": "2026-08-13T15:59:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:8d8fd4799a0e485bd3c6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111973",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111973",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Let read-in-order propagate through SpillingHashJoin",
          "text": "`SpillingHashJoin::hasDelayedBlocks` was hardcoded to `true`, even in the `IN_MEMORY_JOIN` state where nothing is ever delayed. That is the flag gating the read-in-order-through-join and top-k-through-join optimizations, so wrapping a hash join for auto-spilling silently disabled both — and `max_bytes_ratio_before_external_join` defaults to `0.5`, which wraps every hash join. The in-tree comment in `topKThroughJoin.cpp` already describes this as the steady state. `IJoin` documents that \"SpillingHashJoin overrides `keepLeftPipelineInOrder` to forbid switching to GraceHashJoin at runtime\", but no such override existed, so the escape hatch the comment describes was never implemented. This implements it. `keepLeftPipelineInOrder` now pins the join to its in-memory algorithm, and `hasDelayedBlocks` reports `false` from that point on. Pinning is required for **correctness**, not just speed: once the plan drops a sort because the join preserves the left order, a later switch to `GraceHashJoin` would scatter rows by hash and silently return them in the wrong order. The optimizer asks before it commits (`findReadingStep` checks feasibility, and `keepLeftPipelineInOrder` is only called later, if reading in order actually turns out to be possible). So a new `IJoin::canKeepLeftPipelineInOrder` carries the question \"would you preserve the order if I asked?\", defaulting to `!hasDelayedBlocks()` so every other join keeps its current answer. `SpillingHashJoin` answers yes while still reporting delayed blocks, and only stops reporting them once actually pinned — if the optimization turns out not to apply, nothing is pinned and the delayed-block transforms are still built. **Trade-off:** a pinned join can no longer spill, so it holds the whole right side in memory and the memory tracker enforces the limit, exactly as it would with no auto-spill threshold configured. Since dropping the sort and then spilling would be a wrong-results bug, the only alternative is the conservative status quo of never propagating read-in-order through these joins. Both are now reachable: the new setting `query_plan_read_in_order_through_spilling_join` (default `1`) turns the optimization off again, and it is registered in `SettingsChangesHistory` with `previous_value = false`, so `compatibility` set to a version before 26.8 restores the old behavior. `ConcurrentHashJoin` does not spill on its own (its threshold argument only bounds preallocation), so `switchToGraceHashJoin` is the single place the invariant has to hold. The same gate applies to the first-pass `topKThroughJoin`: an `ORDER BY left_key LIMIT n` over a spill-capable `LEFT JOIN` now steps aside for the second-pass read-in-order plan instead of materializing a pushed-down `Sort` and `Limit`, and goes back to pushing them down when the setting is `0`. The handoff keeps the deferral's pre-existing conditions — most notably, `query_plan_join_swap_table` must be explicitly `false`, because under the default `auto` a later optimization may swap the join sides and invalidate the read-in-order plan; with `auto`, such queries keep the pushed-down `Sort` and `Limit`. Making this the default path exposed a pre-existing problem in read-in-order through a `JOIN`, which #110283 pinned down with `04516_join_order_estimation_pruned_parts` a few hours before this branch entered the merge queue. `max_rows_to_read` with `read_overflow_mode = 'throw'` is not checked against the rows a query reads, but against the rows the reading steps announce up front: `ReadProgressCallback::onProgress` substitutes `progress.total_rows_to_read` for `progress.read_rows` whenever the estimate is the larger of the two, and `ReadFromMergeTree` announces `min(part rows, InputOrderInfo::limit)`. `buildSortingDAG` dropped the limit at every `JoinStep`, so a plan reading in order through a `JOIN` always announced whole parts, and `ORDER BY left_key LIMIT 1` over a `LEFT JOIN` that reads 20 rows announced 1010 and failed `max_rows_to_read = 100`. Dropping the limit is right for a join that can filter the left stream, but a `LEFT ALL`/`LEFT ANY` join emits at least one row for every left row, so `n` output rows need at most `n` left rows - duplication only makes fewer left rows necessary. The limit now survives those joins and is dropped everywhere else (`INNER`, `SEMI`, `ANTI`, ...). `InputOrderInfo::limit` never truncates a read - it feeds the announcement, the read-pool task size, the `take_full_part` heuristic and `use_buffering` - so this cannot change results. The behavior was already reachable on `master` with `max_bytes_ratio_before_external_join = 0`, which makes `topKThroughJoin` defer to the second pass exactly as it now does by default. Related: https://github.com/ClickHouse/ClickHouse/pull/111248 Related: https://github.com/ClickHouse/ClickHouse/pull/111972 Related: https://github.com/ClickHouse/ClickHouse/pull/110283 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reading in order through a JOIN now also works when a hash join has an automatic spill-to-disk threshold configured (the default since 26.5). When this optimization applies, the join stays in memory so it preserves the left-side order; set `query_plan_read_in_order_through_spilling_join = 0` to restore the previous behavior, where such joins are not used for read-in-order and remain free to spill. As part of this, an `ORDER BY ... LIMIT` over a `LEFT JOIN` read in order no longer reports the whole table as the number of rows it intends to read, so it no longer trips `max_rows_to_read` with `read_overflow_mode = 'throw'` for a read that stops early.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111973",
          "createdAt": "2026-07-26T19:07:03Z",
          "updatedAt": "2026-08-13T15:59:11Z",
          "timestamp": "2026-08-13T15:59:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c36f0afe88439ba5348f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114463",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114463",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Report mutation cancellation instead of stopping nested pipelines silently",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107619 Related: https://github.com/ClickHouse/ClickHouse/pull/112152 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a `Not-ready Set is passed as the second argument for function 'in'` error when a mutation whose predicate contains `IN (subquery)` is cancelled, for example by `KILL MUTATION` or `DETACH DATABASE`, while the subquery's set is still being built. Related: #107619. ### Description Reported by @ alexey-milovidov on https://github.com/ClickHouse/ClickHouse/pull/112152#issuecomment-5180281674 after a `Stress test (amd_debug)` run aborted this way. `buildSetInplace` materializes the right side of `IN (subquery)` through a nested `CompletedPipelineExecutor` and polls the mutation's interactive-cancel callback. That callback returned a bare `bool`, so on cancellation the executor called `cancel()` and the poll loop returned normally: `Set::finishInsert` never ran, while `build()` had already moved the source plan out. `buildSetsForDAG` returns `void`, so `getMinMaxCountProjectionBlock` evaluated the filter in the same call and `FunctionIn` threw. A constant left-hand side (`1 IN (...)`) makes this the first materialization attempt: it maps to no key column, so primary-key analysis returns first. The callback now throws `ABORTED`, mirroring `MergeTask::checkOperationIsNotCanceled`, which reports the merge path the same way and is likewise used as a nested-pipeline callback. `MergeTreeBackgroundExecutor` already treats `ABORTED` as a normal outcome and logs it at DEBUG. Only the mutation installs this callback shape, so the installers in `TCPHandler` and `LocalConnection` are untouched and client cancellation still returns `QUERY_WAS_CANCELLED` quietly. Validation: `KILL MUTATION` and `DETACH DATABASE` mid-build abort the server before the change and are clean after it, and the table stays mutatable. New test `04865_cancel_mutation_in_subquery_minmax_projection`, 50 of 50 runs. The other cancellation routes and overflow modes checked are listed in the validation gate comment below. Out of scope: a caller that executes a filter DAG synchronously still does not check set readiness, and `buildSetInplace` still leaves the set unbuildable if some future path stops it silently. I found no input reaching either independently of this cancellation. Note for review: #113939 rewrites the same lambda for the S3 read path but keeps the interactive callback non-throwing, so it does not fix this. The two conflict textually; a resolution must keep both the throw and that PR's cancellation persistence.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114463",
          "createdAt": "2026-08-12T10:03:39Z",
          "updatedAt": "2026-08-13T15:58:55Z",
          "timestamp": "2026-08-13T15:58:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d016e112f7a41c336a1c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114522",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114522",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "QueryRunner follow-up: do not occupy threads eagerly",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): QueryRunner tables now start worker threads on demand and release them once idle, instead of occupying threads for the table's whole lifetime.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114522",
          "createdAt": "2026-08-12T16:42:37Z",
          "updatedAt": "2026-08-13T15:58:51Z",
          "timestamp": "2026-08-13T15:58:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e22b5b3f4b2204979eda",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112873",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112873",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Azure: log batch-delete events when SubmitBatch itself fails",
          "text": "> **Series**: #112871 -> **#112873** (this), #112872, #112874, #112875, #112876 **Problem.** When the batch `SubmitBatch` delete request itself fails, the per-object response loop is skipped, so zero `system.blob_storage_log` Delete events are recorded for the whole batch. Failure scenario: ``` removeObjectsBatchIfExists: SubmitBatch() throws (e.g. 403 on the batch endpoint) -> per-object GetResponse() loop skipped -> 0 Delete events logged ``` **Fix.** Record a Delete event per object on batch-level failure before rethrowing. **Changes.** - `Disks/…/AzureBlobStorage/AzureObjectStorage::removeObjectsBatchIfExists`: emit per-object `blob_storage_log` Delete events on `SubmitBatch` failure. - `tests/integration/test_azure_403_handling`: batch-delete-logging test + `configs/blob_log.xml`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed Azure batch object deletion recording no `system.blob_storage_log` Delete events when the batch request itself failed, leaving the whole batch unlogged. A Delete event is now recorded for each object on a batch-level failure.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112873",
          "createdAt": "2026-08-01T06:53:30Z",
          "updatedAt": "2026-08-13T15:58:20Z",
          "timestamp": "2026-08-13T15:58:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "arsenmuk",
          "state": "open",
          "assignees": [
            "SmitaRKulkarni"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b9745ace6d666853018c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105710",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105710",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject divergent same-named parts in parallel replicas coordinator",
          "text": "The parallel replicas coordinator deduplicated parts purely by `MergeTreePartInfo` (name + version), so when two replicas announced a same-named part holding genuinely different data the second announcement was silently merged into the first. The coordinator then dispatched ranges from the first replica's snapshot to the second replica, whose local part was different, producing an exception `Trying to get non existing mark N, while size is M` inside `MergeTreeIndexGranularityConstant::getMarkRows` — or, when the layouts happened to line up, potentially incorrect results. This is reachable with `parallel_replicas_for_non_replicated_merge_tree = 1` against a cluster whose members each have independent local `MergeTree` data: block numbers of a non-replicated `MergeTree` come from a node-local `SimpleIncrement`, so each member's local first part is named `all_1_1_0` but can store different data. The fix makes the coordinator validate the identity of same-named parts across announcements: - `RangesInDataPartDescription` now carries the underlying part's total mark count, a content fingerprint (`checksums.getTotalChecksumUInt128`), and a `part_name_identity` tri-state saying whether a part name identifies the same content on every cluster member. The fields are gated on a new parallel replicas protocol version (`DBMS_PARALLEL_REPLICAS_PROTOCOL_VERSION` bumped to 10). - `part_name_identity` is `ClusterWide` when the engine coordinates block numbers through Keeper (`ReplicatedMergeTree` and descendants) *or* when all of the table's disks keep their metadata in storage shared by every cluster member (`MetadataStorageType::Plain`, `PlainRewritable`, `StaticWeb`, `WebIndex`, `Keeper`) — there every member enumerates literally the same parts. It is `NodeLocal` for a plain `MergeTree` on ordinary local or per-node remote disks. - `InOrderCoordinator` and `DefaultCoordinator` snapshot these fields from the first announcement of each part and compare later announcements against the snapshot. A fingerprint mismatch raises `BAD_ARGUMENTS` naming the diverging part and pointing the user at `ReplicatedMergeTree` or at disabling `parallel_replicas_for_non_replicated_merge_tree`. - When the fingerprint is unavailable on either side (a replica running an older server version, or a part whose checksums are not loaded) and either side reports `NodeLocal` part names, the coordinator fails closed with `BAD_ARGUMENTS` instead of weakening the identity check — such same-named parts cannot be verified, and merging them blindly could return incorrect results. For `ClusterWide` part names the part name implies identical content, so a mark-count fallback keeps mixed-version clusters working during rolling upgrades. Only analyzed-view fields (`rows`, `ranges`) are deliberately NOT compared: per-replica PK and skip-index analysis can legitimately select different mark subsets from the same underlying part, and `ranges` is consumed in place as the coordinator dispatches work. The protocol bump also required pinning one existing consumer of the serializer. The `read_bucket` task parameter of a distributed read plan (`make_distributed_plan`) embeds a `RangesInDataPartsDescription` blob but travels as an opaque query parameter with no handshake to negotiate a version on, so producer and consumer both used the build's own `DBMS_PARALLEL_REPLICAS_PROTOCOL_VERSION` and would disagree about the field layout across a rolling upgrade. The blob is now pinned to `DBMS_PARALLEL_REPLICAS_DISTRIBUTED_READ_BUCKET_VERSION = 8`, the layout in effect when it was introduced. Surfaced by the AST fuzzer (STID `4920-51f2`) on PR #105706. The 5-replica `parallel_replicas` cluster used by `tests/queries/0_stateless/02275_full_sort_join_long.sql.j2` exposes the divergence whenever local `t2` parts differ in size across replicas. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105706&sha=84790f83ec78aee26b08ab6e0bc6712f6e4f1745&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29 Related: https://github.com/ClickHouse/ClickHouse/pull/105706 Coverage is via `gtest_parallel_replicas_coordinator`, which exercises both coordination modes and both announcement orderings: rejection on divergent mark counts and divergent fingerprints, acceptance of identical announcements and of divergent analyzed views over the same part, the fail-closed path for `NodeLocal` part names without a fingerprint, the mark-count fallback for `ClusterWide` part names in mixed-version clusters, and a serializer round trip across protocol versions 8, 9 and 10 that requires the reader to consume the whole buffer. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an exception `Trying to get non existing mark N, while size is M` and potential incorrect parallel-replica range assignment when `parallel_replicas_for_non_replicated_merge_tree = 1` is used on divergent local `MergeTree` data: the parallel replicas coordinator now validates same-named parts by a content fingerprint of the underlying part instead of merging announcements blindly, and fails closed when the fingerprint is unavailable and part names are not guaranteed to identify the same content on every cluster member. ### Documentation entry for user-facing changes - [x] Documentation is unchanged",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105710",
          "createdAt": "2026-05-23T23:10:40Z",
          "updatedAt": "2026-08-13T15:57:38Z",
          "timestamp": "2026-08-13T15:57:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 27
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f35ba396141fe397262d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110613",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110613",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add experimental PCO compression codec (linking the pcodec Rust crate)",
          "text": "Adds an experimental `PCO` compression codec that links the [pcodec](https://github.com/pcodec/pcodec) (`pco`) Rust crate — the reference implementation — rather than reimplementing it. `pco` is a lossless codec specialized for sequences of fixed-width numbers; on numeric columns with smooth, multimodal, or high-entropy distributions (measurements, timings, counters, identifiers) it often beats `Gorilla`/`FPC`/`ZSTD`/`ALP` on ratio. This is the Rust-library counterpart to the native C++ port explored in https://github.com/ClickHouse/ClickHouse/pull/106222. Instead of maintaining a C++ reimplementation, it links the upstream crate. ### Runtime CPU dispatch `pco` has no explicit SIMD: it relies on the compiler autovectorizing its per-batch loops, which only happens when the crate is compiled with the relevant instruction sets enabled (its `build.rs` warns to build with `-C target-feature=+avx2,+bmi1,+bmi2`). That does not work for ClickHouse, which compiles all Rust once at a fixed baseline micro-architecture level for a binary that must run on a range of CPUs. So `pco` is taken from a ClickHouse fork patched to add runtime CPU-feature dispatch of its hot loops — the same mechanism ClickHouse's C++ uses via `TargetSpecific.h`. The vectorizable loop bodies (decode `read_offsets`; encode `write_short_uints`/`write_uints`/`set_offsets`) are compiled at the crate baseline and again for `x86-64-v3` (AVX2 + BMI1/2 + FMA) and `x86-64-v4` (AVX-512), and the widest variant the running CPU supports is selected once and cached. On aarch64 the baseline already includes NEON, so the baseline body is used directly. Verified at baseline `-C target-cpu=x86-64`: the `read_offsets` v3 trampoline emits AVX2 (`ymm`) and v4 emits AVX-512 (`zmm`). The patch is merged into `ClickHouse/pcodec` (a fork of pcodec/pcodec in the ClickHouse organization) via its own PR: https://github.com/ClickHouse/pcodec/pull/1. The crate is added as the `contrib/pcodec` submodule and linked through a thin FFI wrapper crate (`rust/workspace/pco`). Its one not-yet-vendored dependency, `rand_xoshiro`, was added to `contrib/rust_vendor` in https://github.com/ClickHouse/rust_vendor/pull/72. ### Codec Supported types are all fixed-width numerics of 1/2/4/8 bytes via their underlying integer/float representation: `Int8`..`Int64`, `UInt8`..`UInt64`, `Float32`/`Float64`, and the types backed by them (`Date`, `DateTime`, `Decimal32`/`Decimal64`, `IPv4`, `Enum`, ...). The on-disk block stores a 2-byte header (element width + partial-tail byte count) followed by a raw partial-value tail (as in `Gorilla`/`FPC`) and the payload. The payload is a standalone `.pco` stream, wire-compatible with the reference pcodec implementation; when compression would not shrink a block the raw bytes are stored instead (a per-block \"stored\" flag), so the output never expands by more than the 2-byte header and `getMaxCompressedDataSize` is tight. The codec is gated behind `allow_experimental_codecs`. Because it needs the column type, it can only be specified per column: it is rejected in `TTL ... RECOMPRESS` and in the untyped compression settings that resolve codecs without a type. `CompressionCodecMultiple` propagates the experimental / column-type-requiring properties so a chain such as `CODEC(Delta, PCO)` is still gated, and a codec-only `ALTER TABLE ... MODIFY COLUMN x CODEC(PCO)` validates against the existing column type. Tests: `04512_pco_codec` round-trips every supported numeric and backed type (per-element verification, edge cases, codec chaining, compression-ratio check) and `04513_pco_codec_gating` covers the experimental gate, the `TTL RECOMPRESS` rejection, and the codec-only `ALTER`. The FFI wrapper has its own Rust unit tests (round-trip of all types, the no-expansion fallback, and fail-closed handling of malformed/mismatched streams). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new experimental compression codec `PCO`, which links the [pcodec](https://github.com/pcodec/pcodec) library (patched for runtime CPU dispatch), specialized for fixed-width numeric columns. It is wire-format compatible with pcodec `.pco` streams and is enabled with `allow_experimental_codecs`. ### Documentation entry for user-facing changes: - [x] Documentation is written (the `PCO` codec is documented in `docs/en/sql-reference/statements/create/table.md`, and the compression-frame method byte in `docs/en/interfaces/specs/NativeFormat.md`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110613",
          "createdAt": "2026-07-15T19:32:58Z",
          "updatedAt": "2026-08-13T15:57:23Z",
          "timestamp": "2026-08-13T15:57:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 15
          },
          "labels": [
            "submodule changed",
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c27165a1c1ca1734104e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113894",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113894",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #110019 to 26.6: Support some settings alter for OneLake catalog",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/110019 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/112549 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31221192737/job/93005936464)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113894",
          "createdAt": "2026-08-07T22:03:13Z",
          "updatedAt": "2026-08-13T15:56:43Z",
          "timestamp": "2026-08-13T15:56:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "open",
          "assignees": [
            "alesapin",
            "scanhex12"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:6a878d5abf35422c2c19",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108653",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108653",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support `GROUPS` frame mode for window functions",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> The query below applies the same `1 PRECEDING AND 1 FOLLOWING` bounds as a `ROWS`, a `RANGE`, and a `GROUPS` (this PR) frame. The `order` column contains duplicate and non-consecutive values, so the three modes cover different rows: ```sql CREATE TABLE wf_frame_groups (`order` UInt64, value UInt64) ENGINE = Memory; INSERT INTO wf_frame_groups FORMAT Values (10, 1), (10, 2), (20, 3), (30, 4), (30, 5); SELECT order, value, groupArray(value) OVER (ORDER BY order ROWS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS rows_frame, groupArray(value) OVER (ORDER BY order RANGE BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS range_frame, groupArray(value) OVER (ORDER BY order GROUPS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS groups_frame FROM wf_frame_groups ORDER BY order, value; ``` ```response ┌─order─┬─value─┬─rows_frame─┬─range_frame─┬─groups_frame─┐ │ 10 │ 1 │ [1,2] │ [1,2] │ [1,2,3] │ │ 10 │ 2 │ [1,2,3] │ [1,2] │ [1,2,3] │ │ 20 │ 3 │ [2,3,4] │ [3] │ [1,2,3,4,5] │ │ 30 │ 4 │ [3,4,5] │ [4,5] │ [3,4,5] │ │ 30 │ 5 │ [4,5] │ [4,5] │ [3,4,5] │ └───────┴───────┴────────────┴─────────────┴──────────────┘ ``` Each mode interprets the bounds differently: - `ROWS` counts physical rows, so the frame is at most three adjacent rows: the current row plus one on each side. - `RANGE` counts `order` values, so `1 PRECEDING` and `1 FOLLOWING` cover rows whose `order` is within 1 of the current row's. With gaps of 10, no neighbouring row qualifies, so the frame holds only the rows that share the current `order`. - `GROUPS` (added by this PR) counts peer groups, so `1 PRECEDING` and `1 FOLLOWING` always include the adjacent groups in full, whatever the gaps between `order` values. cc: @cwurm ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `GROUPS` frame mode for window functions (SQL:2011), e.g. `any(price) OVER (PARTITION BY symbol ORDER BY ts GROUPS BETWEEN CURRENT ROW AND 1 FOLLOWING)`. In a `GROUPS` frame the boundaries count whole peer groups — sets of rows that are equal on the `ORDER BY` key — so `N PRECEDING`/`N FOLLOWING` mean `N` peer groups before/after the current row's peer group, rather than physical rows (`ROWS`) or `ORDER BY` value distances (`RANGE`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108653",
          "createdAt": "2026-06-26T20:45:31Z",
          "updatedAt": "2026-08-13T15:56:02Z",
          "timestamp": "2026-08-13T15:56:02Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "nihalzp",
          "state": "open",
          "assignees": [
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ebfaf8bdcb693bfd2273",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114575",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114575",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112498 to 26.6: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114575",
          "createdAt": "2026-08-13T02:35:48Z",
          "updatedAt": "2026-08-13T15:55:58Z",
          "timestamp": "2026-08-13T15:55:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "closed",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:550299daa97dd287ae6b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108336",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108336",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Release pull request for branch 26.6",
          "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108336",
          "createdAt": "2026-06-23T22:33:53Z",
          "updatedAt": "2026-08-13T15:55:58Z",
          "timestamp": "2026-08-13T15:55:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "release"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:40109d596530534ad593",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114283",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114283",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add pre-hook to insert CI links into PR body",
          "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) -- Adds a `ci_links.py` pre-hook to the `PR` workflow that, on upstream `ClickHouse/ClickHouse` pull request runs, appends a `:ci_links:` block to the PR description with: - a link to the workflow report, and - a link to a GitHub search for the corresponding sync PR (`sync-upstream/pr/<number>`). The block is added only when it is not already present, so subsequent runs do not re-edit the PR body. Non-upstream / non-PR runs are skipped, and any failure is caught so it can never break the workflow. <!-- CI automatic block start :ci_links: --> --- Workflow [[PR](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114283&sha=latest&name_0=PR)] Sync PR [[sync-upstream/pr/114283](https://github.com/search?q=head%3Async-upstream%2Fpr%2F114283+org%3AClickHouse+type%3Apr&type=pullrequests)] <!-- CI automatic block end :ci_links: -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114283",
          "createdAt": "2026-08-11T07:48:23Z",
          "updatedAt": "2026-08-13T15:55:55Z",
          "timestamp": "2026-08-13T15:55:55Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:acc2f546e8ee287a8e00",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114660",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114660",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Stop SeaweedFS deleting live object folders in stateless CI",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113828 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Since 2026-08-12, `s3 storage` stateless jobs intermittently fail an ordinary INSERT with `Code: 499 ... Immediately after upload: Object ... suddenly disappeared (S3_ERROR)`. The object really was destroyed: SeaweedFS's asynchronous empty-folder cleaner counts a folder's entries and then deletes the folder without excluding a PUT arriving in between, so a write landing in that window is lost. From a failing master run's own artifacts: ``` 16:18:17.675651 EmptyFolderCleaner: deleting empty folder /buckets/test/test/jys 16:18:17.676290 PUT of the object SUCCEEDS, 282 bytes 16:18:17.678805 verifying HEAD -> 404 16:18:17.681060 WriteBufferFromS3: Nothing to abort (so ClickHouse did not remove it) ``` This started with #113828, which replaced MinIO with SeaweedFS. ClickHouse keys objects as `<prefix>/<3 chars>/<random>`, so the bucket holds thousands of shallow folders that empty and refill continuously, and that churn arms the race: in one job 34,819 of 45,578 cleaner log lines were deletions, peaking at 3,238 per minute. CIDB has 0 hits in the preceding 180 days, then hits on 2026-08-12 only. I set the filer's empty-folder cleanup delay past any job's lifetime, so no folder becomes eligible for deletion while the suite runs. The cleaner keeps running and keeps queueing; only its eligibility window moves, and the folders persist in an instance destroyed at job end. No `src/` change: `s3_check_objects_after_upload` caught a genuinely destroyed write. Retrying or relaxing it would have hidden a real lost object. Validated against the shipped script, arms differing only by these lines: without the change the cleaner deletes 13 folders and 0 of 12 survive; with it, 0 deletions and 12 of 12 survive while the queue still holds its 12 items past 3m03s, where the baseline drained at 2m03s. The race is upstream's, which already carves `.uploads` out of this same cleaner for the identical reason (`empty_folder_cleaner.go:247`); this only keeps CI out of its way.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114660",
          "createdAt": "2026-08-13T15:52:45Z",
          "updatedAt": "2026-08-13T15:55:40Z",
          "timestamp": "2026-08-13T15:55:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:949afbb194292e73b21c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114634",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114634",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113742 to 26.6: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114634",
          "createdAt": "2026-08-13T12:49:55Z",
          "updatedAt": "2026-08-13T15:53:21Z",
          "timestamp": "2026-08-13T15:53:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fbdfc9fee676f7f344b9",
        "signalId": "github:ClickHouse/ClickHouse:issue:96651",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:96651",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "NATS JetStream pull consumer does not recover after failover (or change server)",
          "text": "### Company or project name _No response_ ### Describe the unexpected behaviour When using ClickHouse ENGINE = NATS with JetStream (pull mode, durable consumer), the consumer stops fetching messages after a NATS cluster failover. The TCP connection is successfully re-established (PING/PONG continues), but the JetStream pull loop does not resume. Messages remain unprocessed until the table is manually restarted using: ``` DETACH TABLE <table>; ATTACH TABLE <table>; ``` ### Which ClickHouse versions are affected? ClickHouse server version 26.1.2 ### How to reproduce 1. Run ClickHouse with NATS JetStream consumer. 2. Ensure messages are being consumed normally. 3. Stop the NATS node currently serving as JetStream connect by Clickhouse. 4. Wait for Clickhouse reconnect to next node 5. TCP reconnect happens (PING/PONG visible) 6. Clickhouse send message connect like: ``` CONNECT {\"verbose\":false,\"pedantic\":false,\"user\":\"clickhouse\",\"pass\":\"mysuperPass\",\"tls_required\":false,\"name\":\"\",\"lang\":\"C\",\"version\":\"3.9.2\",\"protocol\":1,\"echo\":true,\"headers\":true,\"no_responders\":true} 7. Then Clickhouse send: ``` SUB _INBOX.PXZCFR8UL2QF2MTNAXU2OL.* 1 ``` 8. After that Clickhouse send only ping, without fetch message 9. After detatch/attach table: ``` DETACH TABLE nats_consumer; ATTACH TABLE nats_consumer; ``` clickhouse send: ``` PUB $JS.API.CONSUMER.INFO.test.clickhouse _INBOX.PXZCFR8UL2QF2MTNAXU2OL.9 0 SUB _INBOX.PXZCFR8UL2QF2MTNAXU47V.* 11 PUB $JS.API.CONSUMER.MSG.NEXT.test.clickhouse _INBOX.PXZCFR8UL2QF2MTNAXU47V.1 13 {\"batch\":128} ``` and consumption resumes immediately. ### Expected behavior _No response_ ### Error message and/or stacktrace I don't see any errors in the logs ### Additional context DDL: ``` --connect to nats CREATE TABLE IF NOT EXISTS nats_connect ( uuid String, send UInt8, metainfo String ) ENGINE = NATS SETTINGS nats_server_list = 'node1:4222,node2:4222,node3:4222,node4:4222', nats_subjects = 'test.>', nats_consumer_name = 'clickhouse', nats_format = 'JSONEachRow', nats_num_consumers = 1, nats_max_rows_per_message = 1, nats_username = 'clickhouse', nats_password = 'mysuperPass', nats_stream = 'test', nats_handle_error_mode = 'stream'; --main table to insert data from nats CREATE TABLE IF NOT EXISTS nats_data' ( uuid String, send UInt8, metainfo String, ts DateTime DEFAULT now() ) ENGINE = MergeTree() ORDER BY ts; --consumer CREATE MATERIALIZED VIEW nats_consumer TO nats_data AS SELECT uuid, send, metainfo FROM nats_connect; ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #112828 - Merged into: `26.8.1.1343` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/96651",
          "createdAt": "2026-02-11T10:11:19Z",
          "updatedAt": "2026-08-13T15:51:47Z",
          "timestamp": "2026-08-13T15:51:47Z",
          "metrics": {
            "reactions": 7,
            "comments": 2
          },
          "labels": [
            "unexpected behaviour"
          ],
          "author": "echohes",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:9add64ee360eebba0e08",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104965",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104965",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add server setting `additional_memory_tracking_per_thread`",
          "text": "Each thread accumulates up to `max_untracked_memory` (4 MiB by default) of allocations before reporting them to the server-wide `MemoryTracker`. With many threads, this unreported memory can sum to a large amount, causing the server's tracked memory usage to under-count actual consumption and leading to OOM. This PR introduces a server-level setting `additional_memory_tracking_per_thread` (default 4 MiB) and speculatively charges this amount to the server-wide `MemoryTracker` around every job executed in our `ThreadPool` workers. The global tracked memory becomes a safe upper bound on actual consumption. The reservation is charged on the server-wide (total) tracker only — deliberately not through the query's tracker chain. Query-level accounting feeds heuristics that compare memory deltas against byte thresholds (conversion of aggregation hash tables to two-level via `group_by_two_level_threshold_bytes`, spill-to-disk decisions), and phantom reservations of `num_threads * 4 MiB` trip those thresholds immediately: an earlier revision of this PR that charged the query tracker showed consistent slowdowns of GROUP BY queries in performance tests for exactly this reason. With the server-wide-only reservation, query-level and user-level accounting (`max_memory_usage`, `memory_usage` in `system.processes`) are unaffected. The speculative reservation uses the throwing path of the memory tracker, so when it would exceed the server memory limit the corresponding job is treated as failed with `MEMORY_LIMIT_EXCEEDED` — the same behavior as if the job itself had exceeded the limit. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new server setting `additional_memory_tracking_per_thread` (default 4 MiB) which speculatively reserves this amount on the server-wide memory tracker around every `ThreadPool` job. It compensates for the up to `max_untracked_memory` of un-reported allocations per thread, making the server's tracked memory a safe upper bound on actual consumption and reducing the risk of OOM with many concurrent threads. Query-level and user-level memory accounting are not affected. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASan - [ ] <!---ci_include_tsan--> All with TSan - [ ] <!---ci_include_msan--> All with MSan - [ ] <!---ci_include_ubsan--> All with UBSan - [ ] <!---ci_include_coverage--> All with Coverage - [ ] <!---ci_include_aarch64--> All with Aarch64 #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASan - [ ] <!---ci_exclude_tsan--> All with TSan - [ ] <!---ci_exclude_msan--> All with MSan - [ ] <!---ci_exclude_ubsan--> All with UBSan - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 #### Extra options: - [ ] <!---ci_set_arm--> Add tests with aarch64 builds - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104965",
          "createdAt": "2026-05-14T16:39:55Z",
          "updatedAt": "2026-08-13T15:50:37Z",
          "timestamp": "2026-08-13T15:50:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "pr-improvement",
            "memory"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8e18d31a5a0bd559bd27",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101273",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101273",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add geo aggregate functions #80186",
          "text": "### Functions for geospatial aggregation #80186 Add three SQL aggregate functions for polygon set operations and convex hull computation: - `groupPolygonUnion` — union of polygonal geometries in a group - `groupPolygonIntersection` — intersection of polygonal geometries in a group - `groupConvexHull` — convex hull of grouped point, linear, and polygonal geometries The functions support typed Geo inputs and `Geometry` variant columns, `State` / `Merge` combinators, binary aggregate-state serialization with writer/reader invariant checks and corruption guards, and well-defined empty-geometry semantics. Closes: https://github.com/ClickHouse/ClickHouse/issues/80186 ### Changelog category: - New Feature ### Changelog entry Added geospatial aggregate functions `groupPolygonUnion`, `groupPolygonIntersection`, and `groupConvexHull`. ### Documentation entry for user-facing changes - [x] Documentation is written in code",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101273",
          "createdAt": "2026-03-30T21:25:08Z",
          "updatedAt": "2026-08-13T15:50:16Z",
          "timestamp": "2026-08-13T15:50:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 36
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "zhemalb",
          "state": "open",
          "assignees": [
            "scanhex12"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:666729f462c45ae23517",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114645",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114645",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Speed up `IN (subquery)` set building by pre-deduplicating each `MergeTree` partition independently",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/108326 Related: https://github.com/ClickHouse/ClickHouse/pull/105126 The set for `IN (subquery)` is built by a single `CreatingSetsTransform`: all streams of the subquery are merged into one and every row is hashed serially, no matter how many threads read the data. If the partition expression of the subquery's table is a function of the subquery's output columns (the set is keyed on all of them), the reading will now emit each partition through a single port and each stream is deduplicated independently before the filling transform. Because a key then lives in exactly one stream, per-stream deduplication is complete, and the single filling transform only hashes unique rows — the serial part of the build shrinks from all rows to distinct rows, and the deduplication itself runs in parallel. ```sql CREATE TABLE t (a UInt64) ENGINE = MergeTree ORDER BY tuple() PARTITION BY a % 8; INSERT INTO t SELECT number % 1000000 FROM numbers(100000000); OPTIMIZE TABLE t FINAL; EXPLAIN PIPELINE SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t) SETTINGS allow_creating_set_partitions_independently = 1, max_threads = 8; ``` ```response (CreatingSets) DelayedPorts 9 → 8 (Expression) ExpressionTransform × 8 (Aggregating) Resize 1 → 8 AggregatingTransform (Expression) ExpressionTransform (Filter) FilterTransform (ReadFromSystemNumbers) NumbersRange 0 → 1 (CreatingSet) CreatingSetsTransform <- the single filling transform now hashes ~1M unique rows instead of 100M Resize 8 → 1 DistinctTransform × 8 <- new: parallel pre-deduplication on partition-disjoint streams (Expression) ExpressionTransform × 8 (ReadFromMergeTree) MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) × 8 0 → 1 <- per-partition reading (8 partitions → 8 streams) ``` On the table above (100M rows, 1M distinct keys, 8 balanced partitions; 64-core machine, average of 3 runs after 2 warm-ups): | query | off | on | speedup | |------------------------------------------------------------|--------|--------|---------| | `SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t)` | 0.643s | 0.113s | 5.7× | ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Speed up set building for `IN (subquery)` on partitioned `MergeTree` tables by keeping each partition's rows within a single stream and deduplicating each stream independently, so the single set-filling transform — previously hashing every row serially — only sees unique rows. This applies when the partition expression is a deterministic function of the subquery's output columns. The optimization is not applied when the largest partition holds more than twice the rows of the average partition; the new setting `force_creating_set_partitions_independently` (disabled by default) bypasses this check. Controlled by the new setting `allow_creating_set_partitions_independently` (enabled by default).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114645",
          "createdAt": "2026-08-13T14:04:03Z",
          "updatedAt": "2026-08-13T15:50:07Z",
          "timestamp": "2026-08-13T15:50:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-performance"
          ],
          "author": "nihalzp",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:948c55a6360ff7246040",
        "signalId": "github:ClickHouse/ClickHouse:issue:114004",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114004",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Inconsistent AST formatting: `view((SELECT ...))` in table-function arguments loses subquery parentheses and cannot be parsed back (STID: 1941-1bfa)",
          "text": "🕵 Found by `AST fuzzer (amd_debug, targeted, old_compatibility)` on an unrelated PR (https://github.com/ClickHouse/ClickHouse/pull/91993, which only touches hex encoding): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=91993&sha=eae3a3cbbfb5d81bed061634d53c701167dce04a&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The fuzzer produced a query containing `view((SELECT ...))` as an argument of a comparison function nested inside `file(...)` table-function arguments. Formatting the AST drops the parentheses around the `view` subquery argument (`view((SELECT ...))` → `view(SELECT ...)`), and the formatted text cannot be parsed back, so the format-consistency check fails with a logical error (`abortOnFailedAssertion` in the debug build): ``` Logical error: 'Inconsistent AST formatting: the query: DESCRIBE TABLE file(concat(xor(if(indexHint(and(greaterOrEquals(alias5764._time, toLowCardinality(toNullable(1048575)) + number), lessOrEquals(alias5764._time, view((SELECT toUInt16OrDefault(65536, equals(minus(materialize(NULL, arrayElement(b)), greater(lowCardinalityIndices('\\\\\\\\', toLowCardinality(NULL)), -1)), toInt8OrZero(2147483646)))))))), a <= toInt256OrDefault(-2147483647), materialize(isNullable(NULL))), and(b >= 257, 1.1754943508222875e-38 > b)), currentDatabase(), toFixedString(toNullable('_04513.orc'), assumeNotNull(256)))) SAMPLE 100 / 5 SETTINGS input_format_orc_use_fast_decoder = 0 cannot parse query back from DESCRIBE TABLE file(concat(xor(if(indexHint(and(greaterOrEquals(alias5764._time, toLowCardinality(toNullable(1048575)) + number), lessOrEquals(alias5764._time, view(SELECT toUInt16OrDefault(65536, equals(minus(materialize(NULL, arrayElement(b)), greater(lowCardinalityIndices('\\\\\\\\', toLowCardinality(NULL)), -1)), toInt8OrZero(2147483646))))))), a <= toInt256OrDefault(-2147483647), materialize(isNullable(NULL))), and(b >= 257, 1.1754943508222875e-38 > b)), currentDatabase(), toFixedString(toNullable('_04513.orc'), assumeNotNull(256)))) SAMPLE 100 / 5 SETTINGS input_format_orc_use_fast_decoder = 0'. ``` The two texts differ only in the parentheses around the `view` subquery: the original has `view((SELECT ...))`, the re-formatted text has `view(SELECT ...)`. On a recent master-based **release** build (`clickhouse local`) the asymmetry is observable in the opposite direction — the parenthesized form is the one that does not parse inside table-function arguments: ```sql -- parses fine (special `view` handling in expression context): SELECT lessOrEquals(t, view((SELECT 1))); -- parses fine (unparenthesized subquery): SELECT * FROM file(if(indexHint(lessOrEquals(t, view(SELECT 1))), 'a', 'b')); -- SYNTAX_ERROR (parenthesized subquery inside table-function arguments): SELECT * FROM file(if(indexHint(lessOrEquals(t, view((SELECT 1)))), 'a', 'b')); ``` So whether `view((SELECT ...))` / `view(SELECT ...)` parses depends on context (plain expression vs. table-function argument) and on build/compatibility configuration, while the formatter always emits the unparenthesized form. Either the parser should accept both forms in all contexts where `view` is accepted at all, or the formatter should print the form that is guaranteed to parse back in the surrounding context. Previous distinct manifestations of this dedup bucket (STID: 1941-1bfa) were fixed individually: #109501 (back-quoted numeric type name, fixed by #109720), #106850, #106358, #100131. This `view` parenthesization case is a new one. CIDB shows this check failing on 5 unrelated PRs in the last 30 days (112921, 110886, 113533, 91993, 111061) and never on master — expected, since the AST fuzzer only runs per-PR with random queries.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114004",
          "createdAt": "2026-08-09T02:45:03Z",
          "updatedAt": "2026-08-13T15:47:35Z",
          "timestamp": "2026-08-13T15:47:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "fuzz",
            "clickgap-analyzed",
            "culprit-pr-not-found"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6cbc4b9e675396819c89",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114285",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114285",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
          "text": "Files written first turn the namespace into a plain directory, which a catalog sharing the storage view (SeaweedFS) rejects with HTTP 500; the swallowed error left an orphaned metadata file that broke retries. Ensure the namespace before the first write; propagate failures except 404-then-create and 409 (REST) / AlreadyExists (Glue). Example: https://pastila.nl/?01f91a25/620869a81af815ba6927860efbc72af4#NhE2Nbzh3AHhUHhA5Fdkzw==GCM ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114285",
          "createdAt": "2026-08-11T08:20:50Z",
          "updatedAt": "2026-08-13T15:46:40Z",
          "timestamp": "2026-08-13T15:46:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "azat",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:6518bfca5aff21bf7cec",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114328",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114328",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enhance MetadataStorageFromMemory",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Pure refactoring change, doesn't affect any working part of code.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114328",
          "createdAt": "2026-08-11T13:44:39Z",
          "updatedAt": "2026-08-13T15:45:53Z",
          "timestamp": "2026-08-13T15:45:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-object-storage-disks"
          ],
          "author": "alesapin",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8a4f9d5378c8396005c1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114073",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114073",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `ORDER BY ... LIMIT` returning too few rows under a row policy",
          "text": "The top-K `ORDER BY <column> LIMIT n` optimization decides whether a query is filtered by looking at the plan-visible filters only: `where_clause = filter_step || getPrewhereInfo()`. A row policy is a reader-side filter that is not visible there, so a query filtered only by a policy took the unfiltered fast path: `MergeTreeDataSelectExecutor` enabled `perform_top_k_optimization` and narrowed the read to the marks holding the smallest values of the sort key, and the row policy then discarded all rows in those marks. The query returned fewer rows than the `LIMIT` - possibly none - even though later marks hold rows the policy keeps. ```sql CREATE TABLE t (key UInt64, INDEX mm_key key TYPE minmax GRANULARITY 1) ENGINE = MergeTree ORDER BY tuple() SETTINGS index_granularity = 8; INSERT INTO t SELECT number FROM numbers(300); CREATE ROW POLICY rp ON t FOR SELECT USING key >= 100 TO ALL; SELECT key FROM t ORDER BY key LIMIT 3; -- returned nothing, expected 100, 101, 102 SELECT count() FROM t; -- 200, correct ``` The optimization now counts `getRowLevelFilter()` as a filter, exactly like a visible `WHERE` or `PREWHERE`, so such a query takes the filtered path. The wrong result is reproducible with default settings since 25.12, when `use_skip_indexes_for_top_k` was enabled by default. Related: https://github.com/ClickHouse/ClickHouse/pull/110188 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `ORDER BY ... LIMIT` returning fewer rows than requested, possibly none, when a row policy was the only filter of the query and the sort column had a `minmax` skip index. The top-K optimization narrowed the read before the row policy was applied. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114073",
          "createdAt": "2026-08-09T21:19:31Z",
          "updatedAt": "2026-08-13T15:43:48Z",
          "timestamp": "2026-08-13T15:43:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c3f933b17e8f20e583d5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110144",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110144",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support per-authentication-method GRANTS clause in CREATE USER and ALTER USER",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/109117 Implements per-authentication-method grant limits, the first part of the linked issue: ```sql ALTER USER vasya ADD IDENTIFIED WITH password BY 'WYmdFyas8PftrHbHQQo8' VALID UNTIL '2026-12-31' GRANTS (SELECT ON db.table) ``` When a user logs in with such a method, the access rights of the session are the intersection of the user's access rights (including granted roles) with the listed elements. The clause never adds rights: a listed privilege that is not granted to the user stays unavailable. This provides a way to create tokens for applications: an additional credential with an expiration date and a limited set of grants, which is tied to the user — it is displayed in `query_log` and `processlist` as the user, stops working if the user is deleted, and is narrowed when the user loses grants. Details: - The clause is parsed after the per-method `VALID UNTIL`, works with `CREATE USER`, `ALTER USER [ADD] IDENTIFIED`, and `NOT IDENTIFIED`, is shown by `SHOW CREATE USER`, and persists through the SQL serialization of access entities (and therefore backups). Elements without a database name are bound to the current database when the query is interpreted. - The intersection erases all grant options (the clause cannot express them), so such sessions cannot `GRANT` anything. Role administration is denied entirely (fail-close), including per-role admin option. `EXECUTE AS`, `ALTER USER`, `CREATE USER` and similar escapes require the corresponding rights to be listed explicitly and granted to the user. - Reattaching to a named session (`session_id`) with a different credential re-applies the limit of the credential used by the new connection, so a limited credential cannot pick up the full rights of a session created by an unrestricted one. - The limit is captured at login: `ALTER USER` affects new sessions, not established ones (same as `VALID UNTIL`). - A new `auth_grants` column in `system.users` exposes the limit of each authentication method. - `ContextData`'s copy constructor now preserves `external_roles` and the new field: previously a recalculation of access rights on a copied context silently dropped external roles; for the new field that would mean silently widening the rights. Known limitations (consistent with `VALID UNTIL` and external roles): the limit is not propagated to other nodes of a cluster in distributed queries or `ON CLUSTER` DDL (the query is checked on the initiator), and the clause is not available in `users.xml`. The `CREATE TOKEN` syntactic sugar and the separate grant for self-service `ADD IDENTIFIED` mentioned in the issue are left for a follow-up. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Authentication methods in `CREATE USER` and `ALTER USER ... ADD IDENTIFIED` support a `GRANTS (SELECT ON db.table, ...)` clause which limits the access rights of sessions authenticated with that method to the intersection with the listed grants. This allows using additional credentials as tokens for applications: `ALTER USER vasya ADD IDENTIFIED WITH password BY '...' VALID UNTIL '2026-12-31' GRANTS (SELECT ON db.table)`. Closes [#109117](https://github.com/ClickHouse/ClickHouse/issues/109117). ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110144",
          "createdAt": "2026-07-12T05:16:00Z",
          "updatedAt": "2026-08-13T15:43:33Z",
          "timestamp": "2026-08-13T15:43:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9f8c2e15e77fe8240f9c",
        "signalId": "github:ClickHouse/ClickHouse:issue:113337",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113337",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Paimon tables with a nullable ARRAY or MAP column cannot be read: \"Nested type ... cannot be inside Nullable type\"",
          "text": "A Paimon table containing a nullable `ARRAY` or `MAP` column cannot be read at all: schema parsing wraps the composite type in `Nullable`, which ClickHouse forbids, so `DESC` and `SELECT` throw an exception before any data is read. Since Paimon columns are nullable by default (e.g. any Spark-created table with an array column), this makes fairly ordinary Paimon tables entirely unreadable — the exception is raised while parsing the table schema in `PaimonMetadata::create`, so it affects the whole table, not just queries touching the offending column. **How to reproduce** Write a Paimon table with a nullable array column (Paimon 1.1.1, Spark 3.5): ```sql -- spark-sql with org.apache.paimon:paimon-spark-3.5:1.1.1, -- spark.sql.catalog.paimon = org.apache.paimon.spark.SparkCatalog, -- spark.sql.catalog.paimon.warehouse = file:/tmp/paimon_wh CREATE TABLE paimon.default.t (f ARRAY<INT>) TBLPROPERTIES ('file.format'='parquet'); INSERT INTO paimon.default.t VALUES (array(1,2)); ``` Read it with ClickHouse (master, commit `7c826d816cd5`): ``` $ clickhouse local --query \"DESC paimonLocal('/tmp/paimon_wh/default.db/t')\" Code: 43. DB::Exception: Nested type Array(Nullable(Int32)) cannot be inside Nullable type. (ILLEGAL_TYPE_OF_ARGUMENT) ``` The same error occurs via `paimonS3` / `paimonAzure` and the `Paimon*` table engines. A nullable `MAP` column fails identically with `Nested type Map(...) cannot be inside Nullable type`. **Root cause** `Paimon::DataType::parse` in `src/Storages/ObjectStorage/DataLakes/Paimon/Types.h` wraps the result in `DataTypeNullable` when the Paimon field is nullable, including for the `ARRAY` and `MAP` branches: ```cpp if (real_type == \"ARRAY\") { ... type.clickhouse_data_type = std::make_shared<DataTypeArray>(nested_type.clickhouse_data_type); if (nullable) type.clickhouse_data_type = std::make_shared<DataTypeNullable>(type.clickhouse_data_type); } ``` `DataTypeNullable`'s constructor rejects composite nested types (`src/DataTypes/DataTypeNullable.cpp`), hence the exception. The `Nullable` wrap should be skipped for `ARRAY` and `MAP` (reading a null array/map as empty), which is how the Iceberg and Delta Lake type mappers handle nullable composites. The existing test fixtures (`tests/queries/0_stateless/data_minio/paimon_all_types`) never hit this because their array/map fields are declared non-nullable. Found during E2E tests: https://github.com/ClickHouse/clickhouse-private/pull/67431; the new type-matrix test works around it with `NOT NULL` composite columns for now. Caused by: https://github.com/ClickHouse/ClickHouse/pull/102343",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113337",
          "createdAt": "2026-08-04T14:18:21Z",
          "updatedAt": "2026-08-13T15:43:09Z",
          "timestamp": "2026-08-13T15:43:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "bug",
            "unfinished code",
            "experimental feature",
            "comp-datalake"
          ],
          "author": "zlareb1",
          "state": "closed",
          "assignees": [
            "scanhex12"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:39630a5f17f6089ead72",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112828",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112828",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Recover a NATS JetStream subscription closed by the broker",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/96651 Related: https://github.com/ClickHouse/ClickHouse/pull/103557 Related: https://github.com/ClickHouse/ClickHouse/pull/112464 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a `NATS` table with `nats_stream` set silently consuming nothing after the NATS server is restarted. The `JetStream` subscription is now re-established automatically instead of requiring `DETACH TABLE` and `ATTACH TABLE`. ### Description A `NATS` engine table reading from a `JetStream` stream stops consuming permanently once the NATS server is restarted. Nothing is logged, the connection reports healthy, and only `DETACH TABLE` plus `ATTACH TABLE` or a server restart recovers it. It is also how the flaky `test_nats_restore_failed_connection_without_losses_on_write` fails on master. Root cause: an asynchronous pull subscription renews its pull request only when a message is delivered, and a reconnect resends the `SUB` line without the outstanding pull request, so with nothing in flight when the server goes away the chain never restarts. The server does report this, answering the outstanding request with `409 Server Shutdown`, and the client then closes the subscription. ClickHouse missed it because `isSubscribed` only tests whether the subscription vector is non-empty, so the existing re-subscribe path was gated on a predicate that cannot see a dead subscription. This adds a per-subscription liveness check and consults it in the streaming task, which drops the subscriptions so the existing re-subscribe runs in the same iteration. Only `JetStream` consumers opt in: core NATS subscriptions are already restored by the client, and recovery drops buffered messages core NATS never redelivers. Validated with five new integration tests. Three restart the broker and fail on master, 9 of 9 repeats, before passing after the change; a fourth asserts a healthy consumer never re-subscribes, so it passes either way and exists to bound the cost. The fifth covers a restart of a table reading two subjects, which nothing covered before. Not covered: a broker loss leaving the subscription with no status at all, such as a hard kill, a partition, or a loss coinciding with a re-subscribe. That needs a local fetch timeout, which would also periodically tear down healthy subscriptions. #103557 targets the same defect from a connection-level reconnect counter, but no longer applies to this code and has no integration test. Close whichever you prefer. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1343` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112828",
          "createdAt": "2026-07-31T22:25:07Z",
          "updatedAt": "2026-08-13T15:51:45Z",
          "timestamp": "2026-08-13T15:51:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "antaljanosbenjamin",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bf48e872a588d6f3d0dd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106734",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106734",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "More settings to randomize",
          "text": "Extends the stateless-test randomizers in `tests/clickhouse-test` with settings added since they were last swept, and pins the tests that were implicitly relying on the old defaults. `SettingsRandomizer`: * Broadens `query_plan_optimize_join_order_algorithm` to cover `dpsub` and `dphyp` (always with `greedy` kept in the fallback chain, since exhausting the chain throws `EXPERIMENTAL_FEATURE_ERROR`), and randomizes `query_plan_optimize_join_order_max_searched_plans`. * Adds ~35 further query-level settings (`LIMIT BY` / `FINAL` / lazy-`FINAL` / runtime-filter / text-index / statistics / regexp-compilation toggles and their numeric companions). * Adds `use_statistics_for_part_pruning` → `use_statistics` and `use_projection_index_in_read_pools` → `optimize_use_projection_filtering` to `conditional_settings`, so a child setting is not enabled while its parent is off. * Randomizes `function_implementation` over the values the server actually supports, probed at startup: it globally filters every `ImplementationSelector`-backed function and each has a different arch-tag set, so a forced value that some function lacks would raise `NO_SUITABLE_FUNCTION_IMPLEMENTATION`. * A few settings are deliberately left commented out with the reason recorded in place, e.g. `use_constant_folding_in_index_analysis` (unsafe per-part min-max prune, and the trigger for https://github.com/ClickHouse/ClickHouse/issues/109893), `correlated_subqueries_use_in_memory_buffer`, `min_filtered_ratio_for_lazy_final` (https://github.com/ClickHouse/ClickHouse/issues/112332) and `max_streams_for_union_step`. `MergeTreeSettingsRandomizer`: * Randomizes `concurrent_part_removal_threshold_for_remote_disk`. * The implicit min-max index settings (`add_minmax_index_for_block_number_column`, `add_minmax_index_for_block_offset_column`, `part_minmax_index_columns`) are explicitly kept out: they are not transparent to queries. The first two add real `MINMAX` skipping indices named `auto_minmax_index__block_number` / `auto_minmax_index__block_offset`, which surface in `SHOW INDEXES`, `system.data_skipping_indices`, part checksums and the `Skip` section of `EXPLAIN indexes = 1`; `part_minmax_index_columns = with_block_number_offset` gives every table a part-level min-max index even with no partition key, so `EXPLAIN indexes = 1` grows an extra `Min-Max / Condition: true` block everywhere. This is the same reason their siblings `add_minmax_index_for_{numeric,string,temporal}_columns` were never randomized — all five are `isReadonlySetting`, i.e. part of the table schema. The rest of the diff pins the affected settings in the stateless tests whose output depends on them. Related: https://github.com/ClickHouse/ClickHouse/pull/107586 (fixes the `Invalid binary search result in MergeTreeSetIndex` logical error this randomization made much more likely to be hit; merged, so it is pulled in by the branch update) Related: https://github.com/ClickHouse/ClickHouse/issues/109893 Related: https://github.com/ClickHouse/ClickHouse/issues/112332 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Running CI a few times before merging.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106734",
          "createdAt": "2026-06-08T16:44:21Z",
          "updatedAt": "2026-08-13T15:43:04Z",
          "timestamp": "2026-08-13T15:43:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-ci"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:16433bda12eaf209a45c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112805",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112805",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not drop a named collection that a detached table still uses",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/96181 Related: https://github.com/ClickHouse/ClickHouse/issues/77366 Related: https://github.com/ClickHouse/ClickHouse/pull/110529 A table detached with a plain `DETACH TABLE` keeps its metadata file, so the server attaches it again on the next start. It is gone from `DatabaseCatalog` though, so `isTableExist` returns false for it, and the `check_named_collection_dependencies` check (added in #96181) treated its dependency as a stale leftover of a failed `CREATE TABLE`: it removed the dependency and let `DROP NAMED COLLECTION` succeed. The `ATTACH` replayed at startup then threw `NAMED_COLLECTION_DOESNT_EXIST`, which aborts loading the metadata, and the server did not start at all. This is how the `Stress test (arm_tsan)` job fails on master with `Cannot start clickhouse-server`: the AST fuzzer makes `04320_url_engine_dispatch_partition_and_format` leave its `URL(named_collection)` table detached (the test's `ATTACH TABLE` never reaches the server), and the test then drops the named collection. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=8aad759007771032aa94d7f4eee3e18103dd63c3&name_0=MasterCI&name_1=Stress%20test%20%28arm_tsan%29 ``` Application: Caught exception while loading metadata: Code: 722. DB::Exception: Waited job failed: Code: 695. DB::Exception: Load job 'load table test_3.04320_..._n' failed: Code: 669. DB::Exception: There is no named collection `04320_..._nc`: Cannot attach table `test_3`.`04320_..._n` from metadata file store/e5b/.../04320_..._n.sql from query ATTACH TABLE ... ENGINE = URL(`04320_..._nc`, format = 'JSON'). (NAMED_COLLECTION_DOESNT_EXIST) ``` The same signature accounts for 6 of the 8 `Cannot start clickhouse-server` failures with a missing named collection in the last 60 days (per `play.clickhouse.com`); the other two come from `03822_named_collection_drop_dependency_check`, which drops the collection with `check_named_collection_dependencies = 0` on purpose, and are the case #110529 handles by tolerating the missing collection at startup. ### Changes Per review feedback, the implementation is a simple in-memory bookkeeping in `NamedCollectionFactory` (an earlier revision inspected the metadata the detached table would be attached from, which required probing database disks and sweeping metadata directories after renames): - `DETACH TABLE` (and `DETACH DATABASE`, for every table inside) moves the dependencies of the table into a list of (collection, database, table) entries. - `DROP NAMED COLLECTION` is refused with `NAMED_COLLECTION_IS_USED` while an entry for the collection exists. - `ATTACH` does not remove the entry by itself: the dependencies are registered while the engine arguments are resolved, and the attach can still fail after that (an unknown format name, a failure creating the storage), leaving the table detached — the entry must keep protecting it. Instead, the entry is removed by the events that prove the metadata under that name is gone or harmless: `DROP TABLE`, `DETACH TABLE ... PERMANENTLY`, and `RENAME` of the (necessarily re-attached) table; `DROP DATABASE` removes the entries of the database's detached tables, and `RENAME DATABASE` re-keys them. The `DROP NAMED COLLECTION` check itself removes nothing: the table's existence in `DatabaseCatalog` is racy against in-flight attaches and detaches (the table can exist while nothing in the drop query has validated its live dependency), so the drop path is read-only and every recorded entry refuses the drop. - `DETACH TABLE ... PERMANENTLY` does not record an entry: a permanently detached table is not loaded at startup, so dropping a collection it references cannot break the server start. A later explicit `ATTACH` of such a table fails cleanly with `NAMED_COLLECTION_DOESNT_EXIST` and is recoverable by recreating the collection. - The list lives in memory only, which is consistent across a restart: a plainly detached table is attached again at the next start, where regular dependency tracking picks it up, and a permanently detached one records no entry at all. The list is deliberately imprecise in one direction: a stale entry may keep refusing the drop for a while after the detached table itself is gone, or after the table was attached back (until the table is dropped or renamed). In exchange, the `DROP NAMED COLLECTION` path performs no disk access at all. - A `RENAME TABLE` that moves a table between an `Ordinary` and an `Atomic` database changes the identity the dependency is keyed by: the move into `Atomic` assigns a fresh UUID to the table and the move out of it drops the UUID, while a dependency is keyed by the UUID for tables of `Atomic` databases and by the name for tables of `Ordinary` ones. The rename interpreter only knows the names, so the entry used to keep the identity the table had before the move and nothing found it afterwards: the detach recorded no entry and, for the `Atomic -> Ordinary` direction, the drop check even classified the entry of the still attached table as a leftover of a failed `CREATE` and dropped the collection from under it. `DatabaseOnDisk::renameTable` now re-keys the entries of the moved table to its new `StorageID`, where both identities are known. `EXCHANGE` is unaffected: it is only supported between two `Atomic` databases, where the UUIDs do not change. - A table of a database with `lazy_load_tables = 1` is attached as a `StorageTableProxy` and its real storage is built only on the first access, so the engine arguments are not resolved at load time and the dependency on the named collection they name stayed unregistered. `DROP NAMED COLLECTION` was then allowed while such a table still referenced the collection - breaking it at the first access with `NAMED_COLLECTION_DOESNT_EXIST` - and a `DETACH` of it had no dependency to move to the list of the detached ones, so the protection above silently disappeared in that mode. `DatabaseOrdinary::loadTableLazy` now registers the dependency straight from the metadata, via the new `tryGetUsedNamedCollectionName` helper. An identifier first argument counts as a collection reference only for the engines that resolve their arguments through named collections (a new `StorageFactory::StorageFeatures::supports_named_collections` flag) - for other engines an identifier means something else, e.g. a cluster name for `Distributed` - and for them the signal is time-stable: whether the collection currently exists is deliberately not checked, so the dependency of a collection that is missing at load time (say, after a drop with `check_named_collection_dependencies = 0`) protects it when it is recreated later. The exception is `Remote`/`RemoteSecure`, where the same identifier is also a valid positional argument - a cluster name - when the named-collection lookup does not resolve (marked by the new `StorageFeatures::named_collection_argument_is_ambiguous`; every other flagged engine treats an unknown collection as an error, not as a fallback to a positional form). Syntax alone cannot prove that such a table uses a collection, so the helper replicates the decision the engine's own argument parsing would make at the same moment - which is exactly what a non-lazy load of the same metadata does: the named-collection branch is taken only when a collection with that name exists, and only `key = value` overrides may follow the collection name, which a positional argument list never looks like. `MongoDB` and `MaterializedPostgreSQL` were the only flagged engines whose eager argument resolution did not pass the dependent table to `tryGetNamedCollectionWithOverrides` (they registered the dependency at the lazy load but not when the storage is built); they now pass it, so both load modes register the same dependency. `addDependency` ignores an exact duplicate, because the same dependency is registered again when the proxy is materialized. Tables created with `CREATE TABLE ... AS f(...)` need no lazy branch: a database with `lazy_load_tables = 1` deliberately loads them eagerly as a `StorageTableFunctionProxy` (see `DatabaseOrdinary::shouldLazyLoad`), and that load registers the dependency via `ITableFunction::getUsedNamedCollectionName`. - The cleanup of stale *active* dependencies (leftovers of a failed `CREATE TABLE`, pre-existing from #96181) no longer treats the table's absence from `DatabaseCatalog` alone as a proof of staleness: the dependency of an in-flight `CREATE`/`ATTACH` is registered while the engine arguments are resolved, before the table is committed to the catalog, and a concurrent `DROP NAMED COLLECTION` could prune it and drop the collection while the create later succeeds — recreating the broken metadata this PR fixes. The creating query holds the `DDLGuard` of the table name for the whole window between the registration and the commit, so the drop re-checks the table's existence under that guard before pruning: once the guard is acquired, no create is in flight, and the table's absence proves the entry is stale. (Entries with an empty database name come from dictionaries defined in the configuration files, which are not created through DDL; they are pruned as before.) The pruning removes only the exact stale entry (the collection and the recorded database, table and UUID): `CREATE TABLE ... UUID` can reuse the UUID of a failed create under a different table name, which the guard of the recorded name does not synchronize with, and removing everything under the UUID would erase the live dependency of such an in-flight create — the collection it uses could then be dropped from under the committed table. A new `create_table_pause_before_commit` failpoint keeps a create inside the window for the tests. ### Documented behavior impact `DROP NAMED COLLECTION` now rejects a collection that a detached table or a table in a detached database references, where it previously succeeded (and left a server that could not start). This is what `check_named_collection_dependencies` already promises - \"Check that DROP NAMED COLLECTION will not break tables that depend on it\" - so the documented behavior of the setting does not change, and setting it to `0` still allows the drop. No documentation update is needed. ### Verified locally (release build) - Before: `CREATE NAMED COLLECTION` + `CREATE TABLE ... ENGINE = URL(nc)` + `DETACH TABLE` + `DROP NAMED COLLECTION` succeeds, and the server then fails to start with the exact error chain above (exit code 210). Same with `DETACH DATABASE`. - After: the drop is refused, the collection stays, the table attaches back, and a restart of a server with the detached table present succeeds. - `04660_drop_named_collection_detached_table`, `04698_drop_named_collection_detached_after_rename`, and `04823_drop_named_collection_broken_attach` (all new, covering plain `DETACH TABLE` (blocks the drop), `DETACH TABLE ... PERMANENTLY` (does not block; the later `ATTACH` fails cleanly), `DETACH DATABASE`, `Ordinary` databases, renames of the table and of the database before and after the detach, stale dependencies of failed `CREATE TABLE`, and an `ATTACH` that fails after the dependencies were registered — the drop stays refused), and `04836_drop_named_collection_inflight_create` (new, runs `DROP NAMED COLLECTION` against a `CREATE TABLE` and an `ATTACH TABLE` paused between the dependency registration and the commit to the catalog: the drop blocks on the `DDLGuard` and is refused), and `04840_drop_named_collection_cross_engine_rename` (new, moves a table between an `Ordinary` and an `Atomic` database in both directions and then detaches the table and its database), and `04848_drop_named_collection_reused_uuid` (new, prunes the stale entry of a failed `CREATE TABLE ... UUID` while a create of a different table reusing the UUID is paused inside the window: the drop of the old collection succeeds, and the drop of the collection the new table uses stays refused; verified to fail without the fix), plus `03822_named_collection_drop_dependency_check` and `04003_named_collection_drop_dependency_check_dict` pass. - `test_named_collections/test.py::test_drop_while_used_by_lazily_loaded_table` (new integration test: a table using a named collection in an `Atomic` database with `lazy_load_tables = 1`, a server restart so the table comes back as a never-accessed proxy, and the drop refused both while it is attached and after `DETACH TABLE`). It is an integration test because the hole is only reachable once the in-memory list is empty, i.e. after a restart: without one, the `DETACH DATABASE` that precedes `ATTACH DATABASE` leaves its own entry behind and that entry refuses the drop on its own. Verified that it fails without the fix (the drop succeeds) and passes with it. - `test_drop_collection_recreated_under_lazily_loaded_table` (new integration test: the collection is dropped with `check_named_collection_dependencies = 0`, the server restarts while it is missing, and the drop of the recreated collection is refused; verified to fail without the fix), and `test_drop_not_used_by_lazily_loaded_distributed_table` (new integration test: a collection named after the cluster of a lazily loaded `Distributed` table is droppable; passes before and after, pinning the behavior). - `test_drop_not_used_by_lazily_loaded_remote_table` (new integration test: a collection named after the cluster of a lazily loaded `ENGINE = Remote(cluster, system, one)` table, created after the table, is droppable after a restart; verified to fail without the fix - the drop was refused with `NAMED_COLLECTION_IS_USED` by the unrelated table), and `test_drop_while_used_by_lazily_loaded_table_function` (new integration test pinning that a `CREATE TABLE ... AS bigquery(collection)` table in a database with `lazy_load_tables = 1` keeps blocking the drop after a restart, before the first access and after a `DETACH TABLE`: such tables are loaded eagerly as a `StorageTableFunctionProxy`, which re-registers the dependency; passes without any code change, confirming no lazy table-function branch is needed). ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the server failing to start after a named collection was dropped while a detached table (or a table in a detached database) still referenced it. `DROP NAMED COLLECTION` now counts detached tables as dependents and is refused with `NAMED_COLLECTION_IS_USED`, as `check_named_collection_dependencies` already does for attached tables.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112805",
          "createdAt": "2026-07-31T19:56:49Z",
          "updatedAt": "2026-08-13T15:42:37Z",
          "timestamp": "2026-08-13T15:42:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:404253a82bbbe7060db1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111852",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111852",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix(Silk): honor O_NONBLOCK in the fiber TLS BIO",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... --- This is a latent-interaction bug that appears only when **three** components are combined — none is wrong on its own: 1. **Silk fiber sockets** (#107680): the fiber-aware OpenSSL BIO `silkBioRead`/`silkBioWrite` submits io_uring I/O and parks the caller for the socket's send/receive timeout. That is correct for a blocking read — but the BIO never consults the fd's `O_NONBLOCK` flag, unlike OpenSSL's default socket BIO, which returns `EAGAIN` immediately when the flag is set. 2. **TLS** (`USE_SSL`): the affected path is the TLS BIO. The plain-socket staleness check uses a raw `recv(MSG_PEEK | MSG_DONTWAIT)` and never touches this code, so the problem is TLS-only. 3. **The connection pool's `SSL_peek` staleness probe** (#110402): before reusing a pooled connection, `DB::getSocketState` flips the fd non-blocking via `ScopedNonBlocking` (a raw `fcntl(F_SETFL, O_NONBLOCK)`, behind Poco's back) and calls `SSL_peek`, expecting an immediate `EAGAIN` — its comment reads *\"The socket is non-blocking, so this never blocks.\"* #110402 introduced this probe, replacing the previous `poll`-based keep-alive disconnect check. Only with all three present does that non-blocking `SSL_peek` route through the silk BIO, which ignores `O_NONBLOCK` and blocks for the receive timeout left on the socket by the previous request. Each component is individually correct; the fix lands on the silk BIO because it is the one whose behavior diverges from OpenSSL's socket-BIO contract (honoring `O_NONBLOCK`), while the TLS layer and the `#110402` probe are behaving as intended. On `master` today the silk BIO has no wired production consumer — it is infrastructure — so this three-way combination is not yet reachable in a shipped server; it was reproduced with downstream work that routes object-storage-disk connections through silk fiber sockets. **Impact.** Every borrow of a pooled TLS connection to an object-storage disk pays a timeout it should not. A server loading tables from an HTTPS object-store disk at startup does many such borrows and stalls — a deterministic ~13.5 s in a local reproduction, and an unbounded boot hang (never reaching \"Ready for connections\", no error logged) with production timeouts or a zero/unset receive timeout, where the wait becomes a deadline-less `future.wait()`. **Root-cause evidence** (local TLS-MinIO reproduction): a server-side request trace showed each request arriving only *after* its wait expired; the ~13.5 s decomposed exactly into the adaptive per-method receive timeouts paid in sequence (GET 500 ms + PUT 3000 ms + DELETE 10000 ms, `ConnectionTimeouts.cpp`); `ss` showed frozen `bytes_sent` through each stall; and a live backtrace was parked at `silkBioRead` ← `SSL_peek` ← `getSslSocketState` ← `isStale` ← `getConnection`. A control with silk sockets disabled does the same step in <15 ms. The plain-HTTP staleness probe uses a raw `recv(MSG_PEEK|MSG_DONTWAIT)` and is unaffected — the bug is TLS-specific. **Fix.** When the fd is non-blocking, `silkBioRead`/`silkBioWrite` do a direct `recv`/`send` with `MSG_DONTWAIT` and set the BIO retry flags (immediate `EAGAIN` → `SSL_ERROR_WANT_READ`), matching OpenSSL's default BIO. The fiber/io_uring path is unchanged for blocking sockets. The non-blocking state is read fresh from the fd on each call via `fcntl(F_GETFL)`, because `Poco::Net::SocketImpl::getBlocking()` is a cached flag the raw-`fcntl` probe never updates (and silk sockets reject `setBlocking(false)` outright). Adds a regression test (`SilkFiberSecureSocketTest.NonBlockingPeekDoesNotBlockOnIdleConnection`) that drives the real `getSocketState` path against an idle TLS connection with a 5 s receive timeout and asserts it returns in under 500 ms; without the fix it blocks the full timeout. Not for changelog: the silk fiber BIO is infrastructure with no in-tree production consumer yet, so no released user is affected. Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111852",
          "createdAt": "2026-07-24T20:53:25Z",
          "updatedAt": "2026-08-13T15:41:29Z",
          "timestamp": "2026-08-13T15:41:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "mstetsyuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:17deb20a6a363fdad73d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114409",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114409",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Revert the PromQL topk/limitk streaming plan and its shared-subquery materialization\"",
          "text": "Reverts ClickHouse/ClickHouse#114326 Depends on https://github.com/ClickHouse/ClickHouse/pull/113397",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114409",
          "createdAt": "2026-08-12T01:15:49Z",
          "updatedAt": "2026-08-13T15:40:32Z",
          "timestamp": "2026-08-13T15:40:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:49814beb3aab561e4652",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112717",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112717",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Rebuild subcolumn skip index on ALTER MODIFY COLUMN of the parent",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/112302 --> Related: https://github.com/ClickHouse/ClickHouse/pull/112302 `ALTER TABLE ... MODIFY COLUMN <parent>` that changes the on-disk representation of a subcolumn (e.g. a `Tuple`/`Nested`/`Map`/`JSON` element) did not rebuild a skip index defined on that subcolumn when the table uses **wide parts**. `MutationsInterpreter::prepare`, in the `READ_COLUMN` branch, decided whether an index must be rebuilt by comparing the mutated column name against the index's required columns using exact string equality. For an index on `t.a` and `ALTER ... MODIFY COLUMN t ...`, the required column is `t.a` while `command.column_name` is `t`, so the comparison never matched. The index was neither rebuilt nor dropped, and on a wide part the unchanged `skp_idx_*` files were hardlinked into the mutated part, leaving the index holding bytes serialized under the old subcolumn type. Depending on the index type and how the stale bytes decode under the new type, this produced silent wrong results, `CANNOT_READ_ALL_DATA`, or `MEMORY_LIMIT_EXCEEDED`. Compact parts are unaffected because a mutation rewrites the whole part regardless. The fix resolves each required column to the column it is stored in (via `getNameInStorage`) before comparing to the altered column name, mirroring the existing subcolumn-aware checks in `MergeTreeData::checkAlterIsPossible` and `MutationsInterpreter::validateUpdateColumns`. Minimal reproduction (silent wrong results before the fix): ```sql CREATE TABLE t (id UInt32, t Tuple(a Int32, b String), INDEX idx t.a TYPE minmax GRANULARITY 1) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 4, min_bytes_for_wide_part = 0, min_rows_for_wide_part = 0; INSERT INTO t SELECT number, (number, 'x') FROM numbers(16); ALTER TABLE t MODIFY COLUMN t Tuple(a Float32, b String); SELECT count() FROM t WHERE t.a >= 1; -- 0 (wrong, was) SELECT count() FROM t WHERE t.a >= 1 SETTINGS use_skip_indexes=0; -- 15 (ground truth) ``` ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix stale skip index after `ALTER TABLE ... MODIFY COLUMN` changes the type of a subcolumn (e.g. a `Tuple`/`Nested`/`Map`/`JSON` element) that the index is defined on, for tables using wide parts. The index is now rebuilt instead of being hardlinked unchanged, which previously could lead to wrong query results.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112717",
          "createdAt": "2026-07-31T09:40:47Z",
          "updatedAt": "2026-08-13T15:39:40Z",
          "timestamp": "2026-08-13T15:39:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "Avogar",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:294bec5925c7398c8711",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114636",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114636",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix logical error in text index lazy apply mode on cancelled queries",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114603 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed logical error `Multi-block postings must be compressed` in queries over tables with a text index in the lazy posting-list apply mode, when the query was canceled during the read.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114636",
          "createdAt": "2026-08-13T12:51:52Z",
          "updatedAt": "2026-08-13T15:39:26Z",
          "timestamp": "2026-08-13T15:39:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:17ad55077db2ade94434",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114478",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114478",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113289 to 26.7: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114478",
          "createdAt": "2026-08-12T12:07:07Z",
          "updatedAt": "2026-08-13T15:38:46Z",
          "timestamp": "2026-08-13T15:38:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll3",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:79b19faebaac75b7a9a8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108820",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108820",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow Distributed and Remote tables in Replicated databases",
          "text": "`database_replicated_allow_only_replicated_engine` rejects table engines that keep their own unreplicated data on disk in a `Replicated` database. Before this change, `Distributed` was also rejected because its optional background `INSERT` queue makes `storesDataOnDisk` return true, even though the queue is a transient send buffer and the table's actual data belongs to its destination shards. This change distinguishes table-owned on-disk data from auxiliary delivery state. The restriction applies to non-`ATTACH` `CREATE` queries when `database_replicated_allow_only_replicated_engine = 1`. | Engine or group | Result | Reason | |---|---|---| | `ReplicatedMergeTree` family and `SharedMergeTree` | Allow | The engine provides a replication or shared-storage contract. | | Writable non-replicated `MergeTree` with a local or remote storage policy | Reject | It owns unreplicated on-disk table data; a remote disk alone does not establish replication. | | Static read-only non-replicated `MergeTree` | Allow | It cannot create new table-owned data. | | `Log`, `TinyLog`, `StripeLog`, `Set`, `Join`, `EmbeddedRocksDB`, and database `File` | Reject | They own unreplicated on-disk table data. | | `MaterializedPostgreSQL` | Reject | It owns a local nested table. | | `Distributed`, `Remote`, and `RemoteSecure` | Allow | Their optional local queue is auxiliary delivery state rather than data of the table itself. | | `Memory`, `Buffer`, and `Null` | Allow | They do not own on-disk table data. | | `Merge`, `Alias`, and `View` | Allow | They are metadata-only. | | Views with inner tables | Depends | Each generated inner table is created and checked separately. | | External-storage engines, data lakes, external databases, and external queues | Allow | Their data is managed outside ClickHouse-owned table storage. | | Lazy `StorageTableProxy` | Reject conservatively | The nested storage is unknown without loading it. | `ATTACH` remains outside the existing gate. The `Distributed` background `INSERT` queue is local to the node accepting an insert and is not replicated. Users requiring acknowledgement only after data reaches the destination shards should set `distributed_foreground_insert = 1`. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow `Distributed`, `Remote`, and `RemoteSecure` tables in `Replicated` databases when `database_replicated_allow_only_replicated_engine` is enabled, while continuing to reject writable non-replicated `MergeTree` tables using local or remote storage policies.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108820",
          "createdAt": "2026-06-29T15:29:14Z",
          "updatedAt": "2026-08-13T15:38:05Z",
          "timestamp": "2026-08-13T15:38:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "UberDever",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:15305d31adc65e9a2c72",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111494",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111494",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Text index: add trivial count optimization",
          "text": "Currently, the text index direct read optimization deserialize the sparse index, dictionary block and postings when there is token that exists in the index. Once the postings is read from disk, it fills the newly created boolean virtual column with postings data. With this optimization, we aim to reduce reading postings from disk and creating a virtual column. Instead we can answer queries using the token metadata from the dictionary block for specific query patterns as follows:. 1. `SELECT count() FROM table WHERE hasToken(column, 'foo');` 2. `SELECT count() FROM table WHERE hasAnyTokens(column, ['foo', 'bar']);` 3. `SELECT count() FROM table WHERE hasAllTokens(column, ['foo', 'bar']);` For the 1. case, we can avoid reading postings at all and use the cardinality metadata stored in the dictionary block to answer the query. For 2. and 3. cases, we would still read the postings but can avoid creating a virtual column. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Returns `COUNT()` queries directly from the text index cardinality metadata.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111494",
          "createdAt": "2026-07-22T22:33:56Z",
          "updatedAt": "2026-08-13T15:37:49Z",
          "timestamp": "2026-08-13T15:37:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-performance"
          ],
          "author": "ahmadov",
          "state": "open",
          "assignees": [
            "Ergus",
            "CurtizJ",
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:280a8c385d7eec3dff60",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114422",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114422",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Convert only SEMI JOIN to IN in the convertJoinToIn optimization",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/101698 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results with `query_plan_convert_join_to_in = 1`. The optimization replaced a JOIN with `key IN (subquery)` for `ALL` and `ANY` strictness, but `IN` only tests membership: `ALL INNER JOIN` lost rows on a duplicated right key, and `ANY INNER JOIN` returned extra rows on a duplicated left key. Only `SEMI JOIN` is converted now. Closes #101698. ### Description `tryConvertJoinToIn` rewrites a JOIN into `key IN (subquery)` over the left side, emitting each matching left row exactly once: semi-join semantics. The gate excluded strictnesses by name rather than requiring that property, so it admitted two that break it in opposite directions. - **`ALL`**, the default and the reported bug, emits `left_count(k) * right_count(k)` rows, so a duplicated right key multiplies left rows and `IN` loses them: 3 rows off, 2 on. - **`ANY`** deduplicates the *left* side at the default `any_join_distinct_right_table_keys = 0`, so `IN` instead adds rows: 2 off, 5 on. The gate now admits only `Semi`; `Left` joins the kind gate since `SEMI` needs `LEFT`/`RIGHT`. Six more declines, each a divergence I measured off vs on: - a `Join` engine right side, whose declared kind and strictness the rewrite stops validating: on `Join(ANY, LEFT, id)` it threw `INCOMPATIBLE_TYPE_OF_JOIN` off, returned rows on; - an active `max_rows_in_join`, `max_bytes_in_join`, `max_rows_to_transfer` or `max_bytes_to_transfer`: the join bounds its stored right side, the set only its hash table, so a limit could stop being enforced; - a key whose type has dynamic structure, which `IN` rejects; - a key-value prepared right side (dictionary, `EmbeddedRocksDB`, any `IKeyValueEntity`), which the planner probes by key: `DirectKeyValueJoin` off, a 2000000-element set on. The `Join` engine check above tests the whole `PreparedJoinStorage`, covering both; - an algorithm the planner would pick ahead of hash, since being enabled is not being chosen: `partial_merge`, `prefer_partial_merge`, `grace_hash` or `auto` listed first streams or spills the right side the set materializes (1 off, 0 on each). `hash,partial_merge` and the stock `direct,parallel_hash,hash` still convert, so position decides; `full_sorting_merge` rejects `SEMI` and already falls through to hash. - a key expression consuming a left column the projection needs, e.g. `ON arrayJoin(l.tags) = r.tag` selecting `l.tags`, throwing `NOT_FOUND_COLUMN_IN_BLOCK`. #104809 fixes that hole from the `ALL` side, so I reuse its predicate. Requested by @ PedroTadim in https://github.com/ClickHouse/ClickHouse/issues/101698#issuecomment-5254250810",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114422",
          "createdAt": "2026-08-12T05:30:34Z",
          "updatedAt": "2026-08-13T15:37:31Z",
          "timestamp": "2026-08-13T15:37:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:12b87100f1ce577b1d45",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113208",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113208",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix array membership over an erased element holding one concrete type",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `has`, `indexOf`, `indexOfAssumeSorted`, `countEqual`, `mapContainsKey` and `mapContainsValue` missing rows that `=` matches when the array element type is type-erasing (`Dynamic` or `Variant`, including nested in `Tuple`): `has([1::UInt64]::Array(Dynamic), 1::UInt8)` returned 0 while `1::UInt64::Dynamic = 1::UInt8` is 1. The fix applies to a row whose own elements resolve to one concrete alternative and `equals` over that peeled pair succeeds; other rows keep the previous behaviour. ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/112953 At default settings these functions miss rows that scalar `=` matches, and `countEqual` breaks the `arrayCount(elem -> elem = x, arr)` equivalence its docs state. **Root cause.** `FunctionArrayIndex` tests equality with `IColumn::compareAt(...) == 0`, but for `ColumnDynamic`/`ColumnVariant` `compareAt` compares the variant type name (or discriminator) before the value: a sort order, not equality. Equal values under different variants never compare equal. **The change.** One dispatcher at both array entry points covers all six `FunctionArrayIndex` instantiations as constant, materialized and `LowCardinality`. For a type-erasing common type it peels an operand resolving to one concrete alternative and compares with the registered `equals`, folding each wrapper level's nullness at its own level so an outer NULL stays distinguishable from a nested one. `Tuple` recursion stops at `Array`/`Map`, also `compareAt`-based. **Scope, deliberately narrow, decided per ROW.** A block shares one flattened element column, so the decision is taken per row group: a row's answer follows from its own elements and the needle, not from which rows share its block. It declines, before any behaviour change, for a row mixing several concrete types, shared-variant rows, container alternatives, `LowCardinality` elements, and any pair `equals` rejects (the condition is `equals` succeeding, not the types, since comparability is partly value-dependent). Declined rows keep master's answer bit-for-bit. A `NULL` needle against a materialized erased array holding a NULL now matches: `has([NULL], NULL)` -> 1. No setting is added. **Validation.** New parallel-safe test `04706`: every cell asserted against an oracle in the same row, controls pinning each declined shape, and a group asserting three block partitions agree. A/B against pristine master over the 229 stateless tests reaching these functions: identical failure sets. 50/50: 100 OK. <details> <summary>Validation detail</summary> Every measurement pairs `SELECT lower(buildId())` against `readelf -n` on the serving binary. `#112953` touches this file but neither entry point. **Left for separate PRs**, each measured unchanged on both arms in a run where `has` itself moved 0 -> 1 on the same fixture: `hasAny`/`hasAll` (separate GatherUtils implementation); `arrayCompact`; container equality itself (`[1::UInt64::Dynamic] = [1::UInt8::Dynamic]` is 0 while the `Tuple` twin is 1, which is why this fix stops there); and the heterogeneous-row miss. **Mutations**, each rebuilt with the Build ID confirmed to move, the whole test re-run, then restored with it confirmed to return. Each reddens only what is named: | mutation | reddens | |---|---| | remove the new call | 15 lines; controls green | | guard on `Dynamic` only | the 2 `Variant` cells | | drop the NULL-vs-NULL match | the NULL cells | | bypass the dispatcher on the `Map` path | direct `has(map, key)` | | non-recursive, then one-level, `Tuple` recursion | the 3 `Tuple` cells; then `Tuple(Tuple(Dynamic))` | | decline constant arrays | the erased constant cell | | accept a row of several concrete types | the heterogeneous controls: the decline protects them | | never unwrap `Nullable` | the `Nullable(Tuple(Dynamic))` cells | | remove the `Map` cardinality normalisation | the 2 `Map` NULL-needle cells | | ignore a column's own nullness in the fold | the `LowCardinality`-needle cell | | decide per block, not per row group | the block-partition cells | **Performance** (debug, `max_threads=1`, median of 5): a 1000-element constant `Array(Dynamic)` over 2000 rows is unchanged at 0.023 s; a materialized 200k x 50 one goes 0.287 s -> **0.135 s**. **The updated reference** (`04338`, 2 lines) now enforces its own comment, \"`has(m, k)` must agree with `has(mapKeys(m), k)` for every row\": `has(map(NULL::Dynamic, 1), CAST(NULL, 'LowCardinality(Nullable(String))'))` was 0, the array path 1. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113208",
          "createdAt": "2026-08-04T00:05:50Z",
          "updatedAt": "2026-08-13T15:36:29Z",
          "timestamp": "2026-08-13T15:36:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:bcf62b84c8982f84366c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114273",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114273",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _This is a test-only PR — no source code changes. Please review: test quality, whether the claimed coverage gaps are real, and whether test output makes sense._ Adds test coverage for 1 untested code path, found during automated review of [PR #109744](https://github.com/ClickHouse/ClickHouse/pull/109744). That PR (1) The PR changes `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:403-421): the whole-column `input_format_null_as_default` short-circuit (`SerializationNullable::deserializeNullAsDefaultOrNestedTextCSV`) is now skipped for a bare, non-empty `Tuple` whose `tuple_ **1. A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested** `src/Processors/Formats/Impl/CSVRowInputFormat.cpp:413`, `src/DataTypes/Serializations/SerializationTuple.cpp:683` **Risk:** `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:409-413) now skips the whole-column `null_as_default` short-circuit for a bare `Tuple`, so the leading field is handed to `SerializationTuple::deserializeTextCSV`, whose per-element arm at src/DataTypes/Serializations/SerializationTuple.cpp:683 applies `null_as_default` to the FIRST ELEMENT — and when that element is itself a `Tuple`, one field stands for the whole nested element. This is exactly the sentence … **What's unique vs PR tests:** The PR's 04405_csv_tuple_leading_null_null_as_default.sql covers a leading `\\N` for a SCALAR first element and, for nested tuples, only `Tuple(Nullable(Int32), Tuple(Int32, Int32))` with `\\N,2,3` where the nested tuple's own fields are present; its comment states `a null inside a nested tuple is read as that whole nested element and is not covered`. This test covers precisely that: the `\\N` field is the field of a nested `Tuple` element (and of a `Nullable(Tuple)` element), the resulting one- … [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/e4ec131d-5685-4047-bd72-b61c347ba13b) cc @groeneai (author of #109744), @Avogar (merged/approved #109744) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114273",
          "createdAt": "2026-08-11T05:31:50Z",
          "updatedAt": "2026-08-13T15:36:17Z",
          "timestamp": "2026-08-13T15:36:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:12d4c389195ab4525167",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114398",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114398",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Accept signed PromQL @ timestamps",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Accept signed numeric timestamps in PromQL @ modifiers. ### Summary PromQL allows an optional `+` or `-` sign before a numeric timestamp in an `@` modifier. ClickHouse previously accepted only unsigned timestamps. This change accepts signed timestamps and preserves negative values in the PromQL query tree. It also keeps negative evaluation times safe for TimeSeries tables with non-negative timestamp types. For `DateTime` and `UInt32` storage, selector ranges are clipped at the Unix epoch instead of allowing negative timestamps to wrap to a large positive value. `DateTime64` continues to support pre-epoch timestamps. Examples: - `http_requests_total @ -100` - `http_requests_total @ +3.3e1` - `http_requests_total @ 1700000000` ### Changes - Allow an optional sign before numeric timestamps in the PromQL grammar. - Preserve negative timestamps in the query tree. - Clip selector ranges to the supported timestamp range for `DateTime` and `UInt32` TimeSeries storage. - Keep negative timestamps available for signed `DateTime64` storage. ### Tests - Added parser coverage for negative timestamps. - Added parser coverage for explicitly positive scientific timestamps. - Added integration coverage for `UInt32`, `DateTime`, and `DateTime64` TimeSeries timestamp types. - Covered negative timestamps before the Unix epoch and positive timestamps whose lookback range crosses the epoch. - Regenerated the ANTLR parser artifacts. References: - [Prometheus query basics](https://prometheus.io/docs/prometheus/latest/querying/basics/)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114398",
          "createdAt": "2026-08-11T23:40:48Z",
          "updatedAt": "2026-08-13T15:35:58Z",
          "timestamp": "2026-08-13T15:35:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "submodule changed",
            "manual approve",
            "can be tested",
            "comp-promql"
          ],
          "author": "fallintoplace",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:35e332a13022995ee264",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114188",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114188",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Ignore redundant parentheses in stored table definitions",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/110833 Related: https://github.com/ClickHouse/ClickHouse/pull/92340 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `ATTACH PARTITION FROM`, `REPLACE PARTITION`, `MOVE PARTITION TO TABLE` and adding a `ReplicatedMergeTree` replica failing with `Tables have different ...`, `METADATA_MISMATCH` or `INCOMPATIBLE_COLUMNS` for tables whose definitions were written with redundant parentheses, such as `PARTITION BY (a)`, `ORDER BY (b)`, `INDEX ix (b * c) TYPE minmax`, `PROJECTION p (SELECT (b) ...)`, `CONSTRAINT c CHECK (a > 0)`, `TTL (d + INTERVAL 10 YEAR)` or `DEFAULT (a + 1)`. ### Description This is an alternative to https://github.com/ClickHouse/ClickHouse/pull/110833 that contains only the parentheses normalization, so that it is easy to backport. It does not touch `getTreeHash` and does not change how definitions are compared - they are still compared as text. The definition expressions of a table (keys, `TTL`, indices, projections, constraints, column defaults) are stored in ZooKeeper as text and compared as text. #92340 made the AST remember whether the user wrote redundant parentheses around an expression, so the same definition now has two spellings, and every comparison of a stored definition started to reject one against the other: ```sql CREATE TABLE dst (x UInt64, y UInt64) ENGINE = MergeTree PARTITION BY x ORDER BY y; CREATE TABLE src (x UInt64, y UInt64) ENGINE = MergeTree PARTITION BY (x) ORDER BY y; INSERT INTO src VALUES (1, 1); ALTER TABLE dst ATTACH PARTITION 1 FROM src; -- Code: 36. DB::Exception: Tables have different partition key. (BAD_ARGUMENTS) ``` An upgrade or a mixed-version cluster is not needed: both tables above are created by the same binary. Measured on released binaries, 25.8, 26.3 and 26.4 accept this and 26.5, 26.6 and 26.7 reject it, which matches the branches that contain `b38a892dfbfab6`. The fix is a single formatting mode. `IAST::FormatSettings::ignore_redundant_parentheses` suppresses the parentheses of the `parenthesized` flag - the one place that emits them, `decideParensEmission`, is skipped - and `IAST::formatIgnoringRedundantParentheses` is the one-line formatter built on it. Because the formatter visits everything it prints, this covers nested parentheses (`PARTITION BY (x + (1))`) and every kind of definition, not just the top level of a key. It is used where a stored definition is serialized or compared: - `ReplicatedMergeTreeTableMetadata::formattedAST` / `formattedASTNormalized` (the ZooKeeper `/metadata` payload and the replica-join comparison). This replaces the partial `stripArtificialParens` workaround added by #92340, which only cleared the flag on an expression list and its immediate children, and it no longer has to clone the AST. - `IndicesDescription::explicitToString` / `allToString`, `ProjectionsDescription::toString`, `ConstraintsDescription::toString` and `ColumnDescription::writeText` - the serializers of the `indices`, `projections`, `constraints` and `columns` parts of the replicated metadata. `ColumnsDescription::operator==` compares two sets of columns through the same serializer. - `MergeTreeData::checkStructureAndGetMergeTreeData` - the `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` structure gate (sorting/partition/primary keys, secondary indices, projections). - `StorageReplicatedMergeTree::alter` - the definitions an `ALTER` writes back into Keeper, so an `ALTER` never publishes a parenthesized definition either. - `AlterCommand::isTTLAlter` (whether restating a `TTL` schedules a `MATERIALIZE TTL` mutation) and the `MODIFY ORDER BY` no-op detection in `AlterCommands::prepare`. The text these produce is exactly what every server produced before #92340, so a new replica and an old one agree in both directions. Nothing else changes: the parser and the formatter are untouched, the table metadata keeps what the user wrote, and `SHOW CREATE` still prints `PARTITION BY (x + (1))`. Definitions that differ in more than the parentheses are still rejected (`a` vs `b`, a different index expression, a different column default). The `columns` payload of `StorageKeeperMap` and `ObjectStorageQueue` also contains these expressions, but both storages now compare it structurally, so a table created by any version keeps working. `StorageKeeperMap` already parses both sides and re-serializes them with the current serializer before comparing. `ObjectStorageQueueTableMetadata` compared the stored string verbatim; it now falls back to comparing the parsed `ColumnsDescription` when the strings differ, so a queue table created by 26.5-26.7 with redundant parentheses in a column `DEFAULT`, `CODEC` or `TTL` is accepted in both directions, and this also un-breaks such a table created by 25.8 and earlier (which is broken on 26.5-26.7 today). Retrying the creation of a replica that failed between the ZooKeeper transaction and saving the local metadata is also covered: `createReplicaAttempt` recognizes an already-created empty replica by comparing its `/metadata` and `/columns`, and now falls back to a structural comparison (`ReplicatedMergeTreeTableMetadata::checkEquals`, parsed `ColumnsDescription`) when the raw strings differ, so such a retry works across the old and new spellings instead of failing with `REPLICA_ALREADY_EXISTS`. Tests: `04836_parenthesized_definitions_attach_partition_from`, `04837_parenthesized_definitions_replicated_metadata` and `04850_parenthesized_definitions_replica_recovery`. Each case in them fails on master with the error it is named after. The unit test `ObjectStorageQueueTableMetadata.ColumnsComparisonIgnoresRedundantParentheses` covers the stored-JSON compatibility of the queue metadata. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1335` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114188",
          "createdAt": "2026-08-10T16:56:26Z",
          "updatedAt": "2026-08-13T15:35:57Z",
          "timestamp": "2026-08-13T15:35:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cbac0a4c7ce43db4c981",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114604",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114604",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix flaky 02999_scalar_subqueries_bug_2",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Failure report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=2485f1496f7dff85c6468cc59d1856cec0f92a9d&name_0=MasterCI&name_1=Stateless%20tests%20%28amd_tsan%2C%20parallel%29 Related: https://github.com/ClickHouse/ClickHouse/pull/111983 Related: https://github.com/ClickHouse/ClickHouse/pull/91744 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02999_scalar_subqueries_bug_2` checked that the scalar subquery in a materialized view definition is not executed at CREATE time by racing a 2 second `max_execution_time` against a 3 second `sleepEachRow`. That makes the test depend on machine load rather than on the property under test, and it failed on master in `Stateless tests (amd_tsan, parallel)`: ``` Code: 159. DB::Exception: Timeout exceeded: maximum: 2000 ms. (TIMEOUT_EXCEEDED) (query 10, line 17) ``` The diagnostics rerun passed 43/43 with the same randomized settings. The scalar was not executed. The message carries no `elapsed ... ms` clause, while the in-sleep deadline check (`sleep.cpp:168`) goes through `checkTimeLimit`, which always reports a non-zero elapsed; the statement was cancelled while still pending, since every query including DDL is registered with the `CancellationChecker` watchdog unconditionally (`ProcessList.cpp:381`). The engine behaved correctly and only the clock failed. `max_execution_time` is not randomized by the runner, so there is no setting to pin, and widening the bound was already tried on the structural twin (#91744) without holding. I removed the timing oracle and assert the property directly with `throwIf(1)`, following @ alexey-milovidov's fix for that twin in #111983. If the scalar is not executed `throwIf` never fires; if it ever is, the statement fails `FUNCTION_THROW_IF_VALUE_IS_NON_ZERO`. `throwIf::isSuitableForConstantFolding` returns `false` (`throwIf.cpp:94`), the same guard `sleepEachRow` relies on, so the mechanism the old oracle depended on is preserved, and the check is now stronger: it detects execution at all, not only execution slower than 2 seconds. I also cover `CREATE TABLE ... EMPTY AS SELECT`, which reaches a different analysis site, and added an executed-position arm that must fail so the check has teeth. The reference file stays empty. Validation: the new test is green 50/50 under `enable_analyzer=1` and 50/50 under `=0`, green with 8 concurrent copies, and two mutations of the new oracle each redden it. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1338` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114604",
          "createdAt": "2026-08-13T08:49:21Z",
          "updatedAt": "2026-08-13T15:35:52Z",
          "timestamp": "2026-08-13T15:35:52Z",
          "metrics": {
            "reactions": 2,
            "comments": 6
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e71ee7d0dd79ce50b46f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114508",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114508",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add ProbeDeck SQL client integration",
          "text": "Adds a community-maintained ProbeDeck integration guide for iOS and iPadOS. The guide covers ClickHouse Cloud and self-hosted setup, database authentication, mTLS, SSH bastions, monitoring sources, a reproducible query, limits, and troubleshooting. It also adds ProbeDeck to the SQL client navigation and overview. The screenshots use deterministic synthetic data and contain no production credentials or infrastructure. Related: https://github.com/ClickHouse/clickhouse-docs/pull/6625 ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114508",
          "createdAt": "2026-08-12T15:25:38Z",
          "updatedAt": "2026-08-13T15:35:48Z",
          "timestamp": "2026-08-13T15:35:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "manual approve",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "chamav",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ecb31bee44a4848f98b9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113937",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113937",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: update gui.mdx with CHOPs tool",
          "text": "Added CHOPs UI and Admin tool in the docs ### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113937",
          "createdAt": "2026-08-08T10:13:10Z",
          "updatedAt": "2026-08-13T15:35:43Z",
          "timestamp": "2026-08-13T15:35:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "rva-quantrail",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c58e7d1b6f2818316e3a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110892",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110892",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain analyze join stats",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/110668 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add information of internal state of joins to `EXPLAIN ANALYZE`",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110892",
          "createdAt": "2026-07-17T16:06:34Z",
          "updatedAt": "2026-08-13T15:35:42Z",
          "timestamp": "2026-08-13T15:35:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "Fgrtue",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:92a1888c1f6d2f0bc3dd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114546",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114546",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrapped `Time64` values from an overflowing scale conversion in `convertFieldToType`",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/94537 Related: https://github.com/ClickHouse/ClickHouse/pull/111534 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes incorrect, often sign-flipped, `Time64` literals and `IN`-list constants produced when rescaling a lower-scale `Decimal64` overflowed `Int64`. Such a conversion now reports `DECIMAL_OVERFLOW`, matching the `DateTime64` branch and explicit `CAST`. ### Description `convertFieldToTypeImpl` rescales a `Decimal64`-backed field into a `Time64` column at `src/Interpreters/convertFieldToType.cpp:473-476`. The scale-increasing arm multiplied without a range check, so `value * scale_multiplier_diff` could exceed `Int64` and wrap. The wrapped product was then handed to `decimalFromComponentsWithMultiplier<Time64>(value, 0, 1)`, whose own `mulOverflow` check is a no-op at multiplier `1`, so the corrupted value became the field. Observed on master (`bb1bd307`): `Values 'x Time64(6)' (253402207200000::Decimal64(0))` returned `-999:59:59.722624`, and `INSERT` persisted it. The wrapped `Int64` of `253402207200000 * 1000000` is `-4852209831933722624`, which renders as exactly that; the negated input wraps to `+4852209831933722624`, so a negative input returned a positive time. Explicit `CAST(... AS Time64(6))` already reported `DECIMAL_OVERFLOW` here, so the literal path disagreed with `CAST`. The file carried this same statement twice, for `DateTime64` and `Time64`, both unguarded. The related PR above guarded the `DateTime64` one and added test `03797`; its `Time64` twin was left as it was. This change mirrors that guard onto the twin. The operand also becomes `Int64`, which the guard requires: `mulOverflow` on an unsigned operand reports overflow for every negative value, which would reject in-range negative times. Reporting rather than returning Null matches the `DateTime64` twin, which `03797` asserts, and explicit `CAST`. The neighbouring `Date32` branches keep their Null contract and are untouched. Found by a UBSan signed-overflow report on this line; there is no open issue for it. It keeps reproducing on `master`, in both `asan_ubsan` stress jobs, as `signed integer overflow: 253402239600000 * 1000000 cannot be represented in type 'long'` at `src/Interpreters/convertFieldToType.cpp:475`: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=86363e819f9cb704a892063da9998e4c1281eb76&name_0=MasterCI&name_1=Stress%20test%20%28amd_asan_ubsan%29 - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=895af217db902c1458fecc608c6f6b3a78cc985e&name_0=MasterCI&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Only inputs that were producing wrong values change: `9223372036854` at scale 6, the largest whose rescale still fits, is still accepted and returns `999:59:59.000000` identically. Verified by building both arms and diffing: every overflow arm goes from a wrong value to `DECIMAL_OVERFLOW`, while 30 in-range controls across scales 0/3/6/9, both signs and both scale directions are byte-identical. `03797` and `04837` still pass. New test `04883` fails on master and is green here over 100 randomized runs.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114546",
          "createdAt": "2026-08-12T21:31:34Z",
          "updatedAt": "2026-08-13T15:34:31Z",
          "timestamp": "2026-08-13T15:34:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4fb04f1a7391883848bf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114216",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114216",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for the RemovePart part_log row in 02950 and 02491",
          "text": "### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02950_part_log_bytes_uncompressed` and `02491_part_log_has_table_uuid` read `system.part_log` immediately after a DDL that removes a part and assert the `RemovePart` row is already there. That row is written asynchronously, so the assertion is unsound and the tests fail with just that line missing. `RemovePart` has a single emitter, `MergeTreeData::removePartsFinally`, which on both DDL paths here is reached only through `grabOldParts()`. Both DDLs (`DROP PART`, `TRUNCATE`) do run cleanup in the query thread, so usually the row lands first. But that grab selects nothing if it loses `try_lock` on `grab_old_parts_mutex`, or if the part's `shared_ptr` is not unique, e.g. while a concurrent read holds a reference. The part then stays `Outdated` and a later cleanup pass writes the row: late, not lost, so the engine is correct and only the tests need fixing. `01686_event_time_microseconds_part_log.sh` already asserts the same shape behind a bounded poll. Both tests become `.sh`, since a bounded poll is not expressible in `.sql`, and wait for the row under a 60 s bound. Each iteration issues `SYSTEM START CLEANUP <table>` to schedule a pass, because a cleanup thread that found nothing to do backs off up to `max_cleanup_delay_period`. Assertion queries, tags and both `.reference` files are unchanged; no `CREATE TABLE` setting is added here (02491's `old_parts_lifetime` pin and 02950's tags are pre-existing, kept verbatim); no `no-parallel` and no blanket `no-random-*`. Validation: holding a reference across the DDL reproduces the non-unique-ownership path deterministically. With that lever and no poll both tests fail 8/8; with the poll they pass 10/10. Dropping only the `SYSTEM START CLEANUP` line fails at the deadline once the cleanup thread is backed off. Deleting the DDL makes the poll time out rather than pass, so it is not satisfied by a stale row. Then 150/150 green per test across default, `-j 8` and randomized-order batches.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114216",
          "createdAt": "2026-08-10T20:03:37Z",
          "updatedAt": "2026-08-13T15:34:30Z",
          "timestamp": "2026-08-13T15:34:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:27e9d28b2ef1be3e2b7e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114623",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114623",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "CI: Cache: restrict cross-branch reuse to pull_request workflows only",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112358 CI cache reuse in praktika keyed the reuse decision off the branch that produced a record. Branch, though, was only ever a proxy for how much a record can be trusted. This gates reuse on the producing **workflow event** instead and ignores the branch entirely in the reuse decision. Correctness is unaffected either way — a digest match already means identical inputs; the producing event only encodes how much we trust the result. Each cache record now carries the event that produced it. The policy: | Producing event → <br> Reusing event ↓ | `pull_request` | `push` / `schedule` / `dispatch` / `merge_queue` | |---|:---:|:---:| | `pull_request` | ✅ | ✅ | | `push` / `schedule` / `dispatch` / `merge_queue` | ❌ | ✅ | In words: a `pull_request` run reuses any record; every other (trusted) event reuses any record **except** one produced by a `pull_request`. The write side is the dual — a `pull_request` run only fills an empty `(job, digest)` slot (`if_not_exist=True`), while every trusted event overwrites it, so the shared slot always keeps a record reusable by every lane. This keeps the two properties the branch rule aimed at, more directly: - **Security.** A `pull_request` run executes untrusted (possibly fork) code, so a trusted lane must never reuse a record it produced. The trust boundary is stated as pull_request-vs-not rather than inferred from a branch name. - **Drift guard (#112358).** `merge_queue` never reuses a `pull_request` record, so a PR's green flaky-check result can't satisfy the queue's lookup and skip the re-run against the merge-group state — and this now holds independently of whether the PR and merge-queue digests happen to coincide. It also removes a clobbering hole: reuse no longer checks the branch, so a record produced by a release-branch push (or any trusted event) is reused whenever the digest matches, instead of being rejected on a branch mismatch. `CACHE_VERSION` is bumped to 2 so pre-existing records, which lack the event field, are not reused under the new rule. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114623",
          "createdAt": "2026-08-13T11:50:30Z",
          "updatedAt": "2026-08-13T15:34:11Z",
          "timestamp": "2026-08-13T15:34:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [
            "leshikus"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e9106583093f8b81eee5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114551",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114551",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support quoted identifiers in PromQL selectors",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support quoted metric and label names in PromQL selectors. ### Description Prometheus 3 supports UTF-8 metric and label names. Prometheus's own end-to-end test uses selectors such as: - `{\"http.requests\", \"service.name\"=\"api-server\", instance=\"0\", group=\"canary\"}` AWS CloudWatch also documents quoted OpenTelemetry selectors such as: - `{\"http.server.active_requests\", \"@resource.service.name\"=\"myservice\"}` These names are also used in OpenTelemetry metrics and user-facing PromQL products. For example, SigNoz documents queries such as: - `{\"system.cpu.utilization\", \"service.name\"=\"frontend\"}` - `sum by (\"k8s.pod.name\") (rate({\"container.cpu.utilization\", \"k8s.namespace.name\"=\"ns\"}[5m]))` ClickHouse currently accepts only identifier-style tokens for selector names. A quoted selector identifier is rejected by the grammar before query evaluation. This change: - accepts quoted metric and label names in selectors - treats a standalone quoted selector name as an `__name__` matcher - validates and unquotes the identifier - keeps non-legacy names quoted when serializing the query tree References: - [Prometheus UTF-8 guide](https://prometheus.io/docs/guides/utf8/) - [Prometheus end-to-end UTF-8 selector test](https://github.com/prometheus/prometheus/blob/main/promql/promqltest/test_test.go#L166-L191) - [AWS CloudWatch PromQL examples](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-PromQL-Querying.html) - [OpenTelemetry HTTP metric conventions](https://opentelemetry.io/docs/specs/semconv/http/http-metrics/) - [OpenTelemetry Prometheus compatibility survey](https://opentelemetry.io/blog/2024/prometheus-compatibility-survey/) - [SigNoz PromQL UTF-8 guide](https://signoz.io/docs/userguide/write-a-prom-query-with-new-format/) ### Tests - Regenerated the ANTLR parser artifacts. - Added `gtest_PromQLParser` coverage for quoted metric names, quoted label names, and invalid identifiers. - Ran focused syntax checks for the generated parser and touched PromQL sources. - Ran focused grammar checks for quoted metric and label selectors.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114551",
          "createdAt": "2026-08-12T21:52:00Z",
          "updatedAt": "2026-08-13T15:33:58Z",
          "timestamp": "2026-08-13T15:33:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "submodule changed",
            "can be tested",
            "comp-promql"
          ],
          "author": "fallintoplace",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f37ba72981d698376287",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104217",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104217",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix SQLite WHERE predicate pushdown for strings with special characters",
          "text": "`StorageSQLite::read` used `LiteralEscapingStyle::Regular`, which escapes single quotes as `\\'`. SQLite does not recognise backslash escapes; its only valid string escape is `''`. A pushed-down predicate like `WHERE col = 'it\\'s'` causes SQLite to parse `'it\\'` as a closed string and `s'` as a stray token — a SQL syntax error or injection vector. Switching to `LiteralEscapingStyle::PostgreSQL` would fix single quotes but still emit `\\n`, `\\r`, `\\t` as backslash sequences (which `writeAnyEscapedString` applies unconditionally). SQLite does not interpret those, so predicates on control-character strings would silently return no rows. This PR adds a dedicated `LiteralEscapingStyle::SQLite` backed by `writeQuotedStringSQLite`: only `'` → `''`; all other bytes (including `\\`, newline, tab) are embedded literally. NUL bytes cannot be embedded — SQLite's tokenizer loop in `sqlite3GetToken` terminates on `c==0` even inside a string literal, returning `TK_ILLEGAL` — so a predicate whose string literal (possibly nested in an `IN` tuple, array or map) contains a NUL byte is not pushed down at all: ClickHouse evaluates it locally, and with `external_table_strict_query = 1` the query is rejected instead of silently returning wrong rows. This is a follow-up to PR #74144 which fixed the DDL/PRAGMA and INSERT paths for SQLite but left the SELECT pushdown path using the wrong escaping style. During review the same class of bug was fixed on the PostgreSQL pushdown path as well: strings nested inside `Array` / `Tuple` / `Map` literals (e.g. the elements of a pushed-down `IN` list) now stay in the selected dialect all the way down instead of falling back to the regular ClickHouse escaping, and PostgreSQL string literals that contain backslashes or control characters are emitted as escape string constants (`E'...'`), so the server reads back exactly the original bytes regardless of `standard_conforming_strings` (a real tab used to be sent as the two characters `\\t`). Predicates whose string literals contain a NUL byte are not pushed down to PostgreSQL either, since a PostgreSQL string value cannot contain NUL. The same row-value restriction is applied on the normal `WHERE` pushdown path: a multi-column tuple is written as the row value `(a, b)`, which SQLite and MySQL accept only next to a comparison or `IN`, so a predicate such as `WHERE (id, val) IS NOT NULL` is no longer pushed down to them (ClickHouse evaluates it, and with `external_table_strict_query = 1` the query is rejected) instead of being sent as SQL the external database cannot parse (SQLite reports `row value misused`). For PostgreSQL, whose row constructors are ordinary value expressions, it is still pushed down. A tuple used as the whole condition is ClickHouse's list-of-predicates form and keeps being pushed down as a conjunction, `WHERE (\"a\" > 0) AND (\"column\" > 10)`, for every dialect - no external database accepts a row value as a condition. The user-provided `(SELECT ...)` table argument of `sqlite` / `postgresql` / `mysql`, which is re-serialized from the parsed AST and sent to the external database as is, no longer leaks ClickHouse-only syntax into that SQL: `Array` / `Map` literals and tuples with fewer than two elements (which could only be written back as `tuple(...)`) now throw `BAD_ARGUMENTS` instead of producing SQL the external database cannot parse, an explicit `tuple(a, b)` call is re-serialized as the parenthesized row value `(a, b)` - for SQLite and MySQL only in positions where those databases accept a row value (an operand of a comparison or `IN`); in any other position, such as the SELECT list, both the `tuple(...)` call and the equivalent tuple literal throw `BAD_ARGUMENTS`, because the parenthesized form is a syntax error there (SQLite reports `row value misused`). PostgreSQL row constructors are ordinary value expressions, valid in any expression position (`SELECT (a, b)`, `WHERE (a, b) IS NOT NULL`), so for PostgreSQL such tuples are sent through as row values everywhere instead of being rejected - everywhere except a boolean position, since no database accepts a record as a condition. A tuple in a boolean position - the `WHERE` / `HAVING` of the passed query, or an operand of `AND` / `OR` / `NOT` - is ClickHouse's list-of-predicates form, and is lowered to a conjunction for every dialect: `(SELECT ... WHERE (a > 0, b > 10))` reaches the external database as `WHERE (a > 0) AND (b > 10)`, the same rewrite the normal pushdown path applies; `PREWHERE`, which is ClickHouse-only syntax no external database can parse, is lowered into `WHERE` on that path as well (merging with an existing `WHERE` via `AND`), and the lowered filter gets the same boolean-position normalization. the equivalent tuple literal of constants is not a list of predicates the external database could evaluate and throws `BAD_ARGUMENTS` there instead. `array` / `map` calls on that path are rejected for all three databases. The internal `_CAST(literal, 'Type')` wrapper that the analyzer's `ConstantNode::toAST` puts around a tuple literal used as a plain expression operand (e.g. `WHERE (id, val) = (2, 'y')`) when it re-serializes the subquery argument from the query tree is unwrapped back to the literal, instead of leaking the ClickHouse-internal `_CAST` function into the SQL sent to the external database. A single-row multi-column `IN` set keeps its outer parentheses for both carriers - the fast-path literal `(a, b) IN ((1, 'x'))` and the explicit call `(a, b) IN (tuple(1, 'x'))` - so it reaches the external database as `IN ((1, 'x'))` instead of collapsing to the scalar list `IN (1, 'x')`. This normalization applies to MySQL as well: it shares the same re-serialization path, and although its `Regular` literal escaping style is correct for MySQL string literals (MySQL interprets backslash escapes like ClickHouse), the `tuple(...)` / `array(...)` / `map(...)` forms and `Array` / `Map` / single-element-tuple literals are not MySQL syntax either. The JDBC/ODBC (`StorageXDBC`) pushdown path is intentionally out of scope: the bridge protocol only reports the identifier quoting style, not the literal escaping dialect of the remote database, so plumbing a dialect-aware escaping style through it needs a bridge protocol extension. That path keeps the historical `Regular` escaping, and the limitation is now documented at the call site. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect SQL literal escaping in `StorageSQLite` and `sqlite()` table function when pushing `WHERE` predicates to SQLite: single quotes and control characters (`\\n`, `\\r`, `\\t`, `\\`) were escaped with backslashes, which SQLite does not interpret, causing syntax errors or wrong query results. Also fixed the escaping of string literals pushed down to PostgreSQL: strings nested inside `IN` lists kept ClickHouse escaping, and control characters were sent as backslash sequences that PostgreSQL reads back as different bytes. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104217",
          "createdAt": "2026-05-06T11:55:23Z",
          "updatedAt": "2026-08-13T15:33:55Z",
          "timestamp": "2026-08-13T15:33:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 18
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:553a3fff3a7abc1692dd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114401",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keeper: do not lose a session request when the Raft leader changes",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/78474 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a bug where a client connecting to ClickHouse Keeper during a Raft leader change could be held for the whole `session_timeout_ms` (30 seconds by default) before its connection was rejected, instead of being rejected as soon as the in-flight request was dropped. The connecting client can now reconnect to another replica sooner. ### Description A client's `Connect` makes Keeper submit an internal `SessionID` request. If the Raft append stream breaks while it is in flight, during a leader election say, that request is lost silently. **Root cause.** Such a request carries `session_id = -1` and no xid; its identity lives in `(server_id, internal_id)`. Two places used the wrong key: * `KeeperRequestDispatcher::onCommit` correlated a commit with its in-flight head by `(session_id, xid)`. Every `SessionID` request shares `(-1, 0)`, and `onCommit` runs on every node for every commit, so a `SessionID` committed for another server retired a still-uncommitted local request. * The error path queued the response for lookup by `session_id`, which `-1` has no callback for, so it was discarded and `getSessionID` timed out. The real waiter is a promise keyed by `internal_id`. **The change.** `onCommit` additionally requires `(server_id, internal_id)` to match for `OpNum::SessionID`. Error responses go through one `SessionID`-aware routing helper on `KeeperDispatcher`, wired into **both** dispatchers: `use_new_dispatcher` is a setting and the old one shares the defect. The request now fails at the in-flight drain bound rather than at `session_timeout_ms`. `KeeperTCPHandler` does not branch on the error code, so that earlier rejection is the whole user-visible gain; the accurate `ZCONNECTIONLOSS` only improves the server log. `test_keeper_force_recovery` also gets a retry around the connect after the election, since dropping in-flight appends is deliberate. The correlation fix is scoped to `OpNum::SessionID`. The garbage collectors also use negative session ids, but their `TryRemove` is idempotent and unwaited, so they are unaffected. **Validation.** New gtests cover both the routing decision and the production wiring behind it, each verified to go red under a mutation of the change it covers; the old dispatcher was exercised with `use_new_dispatcher = false`. An injected fault at the connect under test reddens the integration test without the retry and passes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114401",
          "createdAt": "2026-08-12T00:10:11Z",
          "updatedAt": "2026-08-13T15:33:51Z",
          "timestamp": "2026-08-13T15:33:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5afa82872bc35703e099",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114558",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114558",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support PromQL @ start() and @ end() modifiers",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): PromQL: support `@ start()` and `@ end()` timestamp modifiers. ### Summary PromQL allows `start()` and `end()` as special values for the `@` modifier. They mean the fixed start and end of a range query. For an instant query, both resolve to the evaluation time. Without this support, valid PromQL such as `http_requests_total @ start()` and `rate(http_requests_total[5m] @ end())` cannot be used against ClickHouse. Fixed-`@` subqueries keep their complete inner time grid and remain step-invariant across the outer range query. Standalone `start()` and `end()` functions are not included. ### Why this matters This is standard PromQL surface used by major Prometheus-compatible systems: - VictoriaMetrics has executable range-query tests for both `time() @ start()` and `time() @ end()`. - Grafana Mimir, Thanos, and Cortex replace these modifiers with fixed timestamps before splitting range queries, including subqueries. - Elasticsearch includes `@ start()`, `@ end()`, and subquery forms in its valid PromQL grammar fixtures. These are public implementation and compatibility references, not a claim about private customer query volume. Supporting the syntax keeps dashboards and recording/alerting queries portable across Prometheus-compatible backends. References: - [Prometheus Querying basics: `@` modifier](https://prometheus.io/docs/prometheus/latest/querying/basics/) - [VictoriaMetrics executable tests](https://github.com/VictoriaMetrics/VictoriaMetrics/blob/master/app/vmselect/promql/exec_test.go#L1207-L1227) - [Grafana Mimir query splitting](https://github.com/grafana/mimir/blob/main/pkg/frontend/querymiddleware/split_and_cache.go#L2838-L2948) - [Thanos query splitting](https://github.com/thanos-io/thanos/blob/main/internal/cortex/querier/queryrange/split_by_interval.go#L593-L690) - [Cortex query splitting](https://github.com/cortexproject/cortex/blob/master/pkg/querier/tripperware/queryrange/split_by_interval.go#L1061-L1157) - [Elasticsearch valid PromQL grammar fixtures](https://github.com/elastic/elasticsearch/blob/main/x-pack/plugin/esql/qa/testFixtures/src/main/resources/promql/grammar/queries-valid.promql#L1112-L1122) ### Changes - Add `start()` and `end()` to the PromQL timestamp grammar. - Preserve the symbolic modifier in the Prometheus query tree. - Resolve the symbols against the outer query boundaries during evaluation. - Keep offset and range-selector handling consistent with numeric `@` timestamps. - Preserve the complete inner grid for fixed-`@` subqueries. ### Tests - Added parser coverage for both symbols, modifier order, literals, and subqueries. - Added instant and range query coverage, including a range selector under `@ end()`. - Added `last_over_time()` coverage for symbolic and numeric `@` on subqueries. - Compared the new range-query cases with Prometheus through the existing integration helpers. - Regenerated the ANTLR parser artifacts.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114558",
          "createdAt": "2026-08-12T23:36:48Z",
          "updatedAt": "2026-08-13T15:33:49Z",
          "timestamp": "2026-08-13T15:33:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-improvement",
            "submodule changed",
            "comp-promql"
          ],
          "author": "fallintoplace",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4caae4d1dcb072da369c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104591",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104591",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add optimize_row_order_if_no_order_by (reopen #103919)",
          "text": "Reopen of https://github.com/ClickHouse/ClickHouse/pull/103919. This closes #103839. Adds a new `MergeTree` setting `optimize_row_order_if_no_order_by` (default `1`) that enables `optimize_row_order` automatically for tables without an explicit `ORDER BY`. **Behavior change and migration path.** With an empty sorting key (`ORDER BY ()` / `ORDER BY tuple()`) no query can rely on the physical row order, so the rows of every inserted block are reordered to improve compressibility. This makes such inserts slower (a 5M-row insert benchmark shows roughly `+180%`..`+230%` on the insert itself) in exchange for a smaller on-disk size and faster filters on low-cardinality columns; the CI `clickbench` and `tpch_adapted` runs report no significant change. Existing tables are affected on upgrade. To keep the old behavior: - per table: `SETTINGS optimize_row_order_if_no_order_by = 0` (or an explicit `optimize_row_order = 0`, which also opts out); - server-wide: set it in the `merge_tree` config section; - by version: it is recorded in the `MergeTree` settings changes history, so `compatibility` set to a version before `26.8` keeps it off. Note that `MergeTree` setting defaults are resolved once, when the server materializes its global `MergeTreeSettings`, so `compatibility` has to come from the default profile (`users.xml`) - a `SET compatibility` in an already-running session does not change them. The `MergeTree` documentation is updated accordingly. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `optimize_row_order_if_no_order_by` `MergeTree` setting. When enabled (default), row order optimization is applied to inserts into tables without an explicit `ORDER BY` clause. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes default insert-time behavior for `MergeTree` tables with `ORDER BY ()`, which can affect CPU cost and on-disk row ordering/compression. Also alters row-order optimization internals to swallow some `NOT_IMPLEMENTED` errors, which could mask type-specific issues if incorrect. > > **Overview** > Introduces a new `MergeTree` setting `optimize_row_order_if_no_order_by` (default **on**) and wires it into `MergeTreeDataWriter` so row-order optimization is automatically applied on insert for tables that *lack a sorting key* (`ORDER BY ()`), while preserving existing behavior for tables with an explicit `ORDER BY`. > > Hardens `RowOrderOptimizer` by catching `NOT_IMPLEMENTED` during cardinality estimation and falling back to an all-distinct upper bound instead of failing inserts, and records the setting in settings-change history. > > Updates many stateless/perf tests to explicitly disable the new default (`SETTINGS optimize_row_order_if_no_order_by = 0`) to keep deterministic baselines, adds new coverage for the setting’s default/override behavior and the regression on unsupported nested types, and extends `tests/performance/scripts/perf.py` to strip this setting from `CREATE TABLE` when running against older servers that don’t recognize it. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 02d68c628d9a3625824d33b91f87bd12d6fd3938. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104591",
          "createdAt": "2026-05-11T14:01:19Z",
          "updatedAt": "2026-08-13T15:33:39Z",
          "timestamp": "2026-08-13T15:33:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 44
          },
          "labels": [
            "pr-improvement",
            "pr-performance",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0046aaea672977e61269",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112816",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112816",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow running queries detached from client session",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New setting `run_query_in_background`. The server accepts the query, immediately returns an empty result, and runs it to completion regardless of what happens to the connection. The result is discarded. Track the query by its `query_id` in `system.processes` and `system.query_log`. Intended for long queries like `INSERT ... SELECT`, `CREATE TABLE … AS SELECT`, or `CREATE MATERIALIZED VIEW … POPULATE` that must not die with a dropped connection.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112816",
          "createdAt": "2026-07-31T20:49:40Z",
          "updatedAt": "2026-08-13T15:31:52Z",
          "timestamp": "2026-08-13T15:31:52Z",
          "metrics": {
            "reactions": 3,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b66de84ac8396dd712ff",
        "signalId": "github:ClickHouse/ClickHouse:issue:114657",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114657",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "The remote leg of a Distributed query re-serializes a lenient toDecimal64 constant into a strict CAST that throws: local and distributed results diverge",
          "text": "**Describe what's wrong** A query whose `WHERE` clause contains a constant expression producing an out-of-precision `Decimal` value — e.g. `toDecimal64(1000000000000000000, 0)`, a 19-digit value in `Decimal(18, 0)` — executes fine against a local table, but the same query through a `Distributed` table throws `Code: 69. DB::Exception: Too many digits (19 > 18) in decimal value` from the remote leg. The distributed layer re-serializes the folded constant as `_CAST('1000000000000000000', 'Decimal(18, 0)')`, and that string-literal form validates digits strictly while the original integer-argument `toDecimal64` call does not (the long-standing leniency described in #14124). So the rewritten query the initiator ships to the shards is not executable, although the user's original query is — the same statement returns a result locally and an exception through `Distributed`. **Does it reproduce on the most recent release?** Reproduced on master `26.8.1.1307` (official build), 20/20 through the `Distributed` table (exception) and 20/20 on the local table (correct result). **How to reproduce** ```sql CREATE TABLE t (a Decimal(18, 0)) ENGINE = MergeTree ORDER BY a; INSERT INTO t SELECT toDecimal64(number * 100000000000, 0) FROM numbers(100); CREATE TABLE dist_t AS t ENGINE = Distributed(test_shard_localhost, currentDatabase(), 't'); -- local: returns 0 SELECT count() FROM t WHERE intDiv(a, toDecimal64(1000000000000000000, 0)) NOT IN (0, 1); -- distributed: Code: 69. DB::Exception: Too many digits (19 > 18) in decimal value: -- In scope SELECT count() FROM ... WHERE intDiv(a, _CAST('1000000000000000000', 'Decimal(18, 0)')) NOT IN (0, 1) SELECT count() FROM dist_t WHERE intDiv(a, toDecimal64(1000000000000000000, 0)) NOT IN (0, 1) SETTINGS prefer_localhost_replica = 0; ``` `prefer_localhost_replica = 0` is needed only because the repro cluster is a localhost loopback; on a cluster with genuinely remote shards the remote leg always takes the serialization path and the exception fires at default settings. Simpler predicates hit the same thing: `WHERE a != toDecimal64(1000000000000000000, 0)` and `WHERE a < toDecimal64(1000000000000000000, 0)` both error 69 through `Distributed` and succeed locally. The constant itself evaluates leniently on the initiator: ```sql SELECT toDecimal64(1000000000000000000, 0); -- returns 1000000000000000000 SELECT CAST('1000000000000000000', 'Decimal(18, 0)'); -- Code: 69, Too many digits (19 > 18) ``` **Expected behavior** The same query should either succeed on both paths or fail on both. Whichever way the `toDecimal64` integer-form leniency is resolved, the SQL the distributed layer generates for the remote legs should be executable whenever the original query is — e.g. by serializing the folded constant in a form that round-trips (keeping the original function call, or a value form that does not re-validate into a stricter type). Related: https://github.com/ClickHouse/ClickHouse/issues/14124 Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114657",
          "createdAt": "2026-08-13T15:31:16Z",
          "updatedAt": "2026-08-13T15:31:16Z",
          "timestamp": "2026-08-13T15:31:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "potential bug"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:c258535201af10245673",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:94148",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:94148",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Introducing a new fuzzer check in the ClickHouse CI for my dear friend",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ---- New buzzhouse job <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds a new multi-node, docker-in-docker fuzzer job to CI and refactors shared fuzzer log analysis/parsing, which could affect how failures are detected and reported across fuzz/stress pipelines. Changes also touch integration test cluster process tracking (`exec_id` handling), so misclassification of failures or missed crashes is the main risk. > > **Overview** > **CI now runs a new fuzzer check, `La Casa Del Dolor`, in both `master` and `pull_request` workflows**, replacing the previous `BuzzHouse` job names/keys and wiring it into `finish_workflow` dependencies and job definitions. > > **Adds `ci/jobs/lacasadeldolor_job.py`** to execute the fuzzer via `tests/casa_del_dolor/dolor.py` in the integration-tests runner (docker-in-docker), collect multi-node logs/config artifacts, and reuse a new shared `analyze_job_logs()` routine. > > **Refactors failure detection/reporting**: `ast_fuzzer_job.py` extracts log analysis into `analyze_job_logs()` (including expanded accepted exit codes and improved OOM/sanitizer handling), and `FuzzerLogParser` is updated to search across multiple (including `.gz`) server/stderr logs and prefer the matched log when extracting stack traces/failed queries; `stress_job.py` and `clickhouse_proc.py` are updated to the new parser API. > > **Updates Casa del Dolor test harness and cluster helpers** to support the new CI mode: fixes imports under `tests.integration`, adjusts generator config/tempfile handling and exit-code validation, tightens disk/policy XML generation, and tracks ClickHouse container `exec_id` in `ClickHouseCluster/Instance` for more reliable shutdown/exit-code inspection. > > <sup>Written by [Cursor Bugbot](https://cursor.com/dashboard?tab=bugbot) for commit 82bd7ef20780c0d5b5bc648ace847418429573cb. This will update automatically on new commits. Configure [here](https://cursor.com/dashboard?tab=bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/94148",
          "createdAt": "2026-01-14T11:15:28Z",
          "updatedAt": "2026-08-13T15:29:54Z",
          "timestamp": "2026-08-13T15:29:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:846b0163844696acaa95",
        "signalId": "github:ClickHouse/ClickHouse:issue:114630",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114630",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "KeyCondition's pointInPolygon primary-key analysis skips is_valid validation (unlike pointInPolygon/spatial_bbox/GeoParquet pruning)",
          "text": "### Describe the unexpected behaviour `KeyCondition`'s primary-key range analysis for `pointInPolygon` (`analyze_point_in_polygon` in `src/Storages/MergeTree/KeyCondition.cpp:3583-3663`) builds the query polygon directly from the literal argument, calls `boost::geometry::correct` and `boost::geometry::envelope`, but never calls `boost::geometry::is_valid`. This is inconsistent with the other two places in the codebase that parse the same kind of literal: - `FunctionPointInPolygon::parseConstPolygon`/`parseConstMultiPolygon` (`src/Functions/pointInPolygon.cpp:842-861,894-913`), which validate the assembled polygon with `bg::is_valid` and throw `BAD_ARGUMENTS` when `validate_polygons` is enabled (the default). - The `spatial_bbox` `MergeTree` skip index and GeoParquet row-group pruning (`src/Common/GeoBbox.h`, added in #104437), which validate every constant geometry argument and fail closed (decline to prune) rather than derive a bbox from an invalid one. `analyze_point_in_polygon` also only recognizes the plain 2-argument `pointInPolygon(point, ring)` form (see the `Case1 no holes in polygon` comment at line 3715); a polygon-with-holes literal (3+ arguments) isn't analyzed for primary-key pruning at all. That part is safe (it just declines to prune, so `KeyCondition` falls back to scanning), it's the missing validity check on the single-ring case that's the actual gap. ### Practical impact Under the default `validate_polygons = 1`, this gap is effectively unreachable through SQL: ClickHouse evaluates `WHERE`-clause constant expressions once on a zero-row block before any index analysis or pruning runs, so an invalid constant polygon literal passed to `pointInPolygon` always raises `BAD_ARGUMENTS` immediately — `KeyCondition`'s analysis, which runs later during part/granule selection, never gets a chance to act on the unvalidated bbox. However, if a user explicitly sets `validate_polygons = 0` — a documented, supported way to bypass `pointInPolygon`'s own geometry validation — the dry-run exception no longer fires, and `KeyCondition`'s PK-range analysis still unconditionally computes a bbox/envelope from the same, now genuinely unvalidated, polygon and may use it to prune primary-key ranges. `boost::geometry`'s query algorithms (`within`, `intersects`, etc.) have undefined results for invalid geometries, so pruning decisions derived from such a polygon are not guaranteed to be sound. ### How to reproduce Not reproducible with a concrete wrong-result example yet — this is a code-review finding, not an observed bug. The scenario would require: a `MergeTree` table with a `pointInPolygon`-friendly primary key, `SETTINGS validate_polygons = 0`, and a self-intersecting/otherwise invalid constant polygon literal in the `WHERE` clause, compared against a full scan of the same query. ### Expected behavior `analyze_point_in_polygon` should either validate the polygon the same way `FunctionPointInPolygon` and `Common/GeoBbox.h` do (and fail closed / decline to prune when invalid, honoring `validate_polygons` the same way the function itself does), or the inconsistency should be a deliberate, documented decision. ### Additional context Found while auditing ClickHouse's geospatial pruning code for duplicated/drifted logic during #104437 (which unified the equivalent bbox-extraction-and-validation logic between the `spatial_bbox` skip index and GeoParquet row-group pruning in `src/Common/GeoBbox.h`). Filing this as a lower-priority follow-up for future reference rather than addressing it in that PR, since it's a different subsystem (primary-key range analysis) and not currently reachable under default settings. Related: https://github.com/ClickHouse/ClickHouse/pull/104437",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114630",
          "createdAt": "2026-08-13T12:47:16Z",
          "updatedAt": "2026-08-13T15:28:59Z",
          "timestamp": "2026-08-13T15:28:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "comp-geo"
          ],
          "author": "bacek",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:542a326b8880cd3e3e54",
        "signalId": "github:ClickHouse/ClickHouse:issue:112376",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112376",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "KILL QUERY requires SELECT on system.processes (and it should not)",
          "text": "### Company or project name ClickHouse Cloud support on behalf of a customer ### Describe what's wrong `KILL QUERY` requires the user to have `SELECT` privilege on `system.processes` table in the Cloud. In specific use cases (e.g. using a BI tool like Metabase), it can lead to a security hole where a user would be forced to open all records if that table to a used needing to kill a query -- not just their own. Creation of a ROW POLICY, however, is not possible (revoked by the Cloud for system.*). The ask is to make it possible to run KILL QUERY (by query_id) without having SELECT access on system.processes. ### Does it reproduce on the most recent release? Yes ### How to reproduce In the Cloud, create a user without SELECT access to system.processes As that user, execute KILL QUERY: ``` me@myLaptop ch_client % clickhouse client --host foo.bar.aws.clickhouse.cloud --secure --user me_test --password '*****' \\ --query \"KILL QUERY WHERE query_id='hello_kitty'\" Received exception from server (version 26.4.1): Code: 497. DB::Exception: Received from foo.bar.aws.clickhouse.cloud:9440. DB::Exception: me_test: Not enough privileges. To execute this query, it's necessary to have the grant SELECT ON system.processes. (ACCESS_DENIED) (query: KILL QUERY WHERE query_id='hello_kitty') ``` ### Expected behavior An existing query should be deleted if it belongs to the user. Otherwise (query doesn't exist, query doesn't belong to the user) a no-op. ### Error message and/or stacktrace see above ### Related issues and pull requests _No response_ ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112376",
          "createdAt": "2026-07-28T23:30:15Z",
          "updatedAt": "2026-08-13T15:28:08Z",
          "timestamp": "2026-08-13T15:28:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "comp-rbac",
            "potential bug"
          ],
          "author": "romanbukarev-clh",
          "state": "open",
          "assignees": [
            "pufit"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:eb08073a3a6a171b034b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:89945",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:89945",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a Chinese tokenizer (jieba) for the tokens function and text indexes",
          "text": "<!--- Disable AI PR formatting assistant: false --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a `chinese` tokenizer for the `tokens` function and `MergeTree` text indexes. It segments Chinese text into words using a dictionary and a Hidden Markov Model (the algorithm follows [jieba](https://github.com/fxsjy/jieba)), with `coarse_grained` (default) and `fine_grained` granularities. Continues #80174. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) ### Implementation notes - The tokenizer is a from-scratch C++ reimplementation following the fxsjy/jieba algorithm (dictionary-based maximum-probability segmentation with an HMM fallback). The embedded dictionary and HMM model are generated from a pinned [cppjieba](https://github.com/yanyiwu/cppjieba) commit (MIT) and verified by SHA-256; regenerator scripts are included. - The engine lives in a standalone repository, https://github.com/ClickHouse/jieba_cpp, consumed here as the `contrib/jieba_cpp` submodule with a `contrib/jieba_cpp-cmake` wrapper that builds it against ClickHouse's in-tree abseil/darts-clone/zstd. Built by default (`ENABLE_CHINESE_TOKENIZER`); the dictionary ships in little- and big-endian variants and is `#embed`-ed per host byte order, so it works on big-endian targets too. - `hasToken` now bypasses the text index for any non-`splitByNonAlpha` tokenizer (use `hasAnyTokens` / `hasAllTokens` for tokenizer-aware matching). This also fixes a pre-existing case where `hasToken` over `ngrams`/`array`/`splitByString`/`asciiCJK` indexes could return wrong rows.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/89945",
          "createdAt": "2025-11-12T16:27:31Z",
          "updatedAt": "2026-08-13T15:27:43Z",
          "timestamp": "2026-08-13T15:27:43Z",
          "metrics": {
            "reactions": 3,
            "comments": 36
          },
          "labels": [
            "pr-feature",
            "submodule changed",
            "can be tested",
            "hold"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "Ergus"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:44e2dee406e520b202a6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86353",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cascades cost-based optimizer for distributed query plans",
          "text": "A Cascades-style cost-based optimizer that chooses distribution strategies for the multi-stage distributed query plans of #106020. It explores alternatives in a memo (a shared store of equivalent plan fragments) with top-down, goal-directed search and picks the cheapest plan satisfying the required distribution and sorting properties, inserting exchange operators (plan steps that move rows between nodes) as needed. Implemented: - **Join strategies**: shuffle hash join, broadcast hash join (with `ReplicatedRead` — every worker repeats the same read of a small table instead of a network broadcast, assuming shared storage where all workers see the same data), replicated join (a small deterministic join is recomputed identically on every node over such reads, so its result never crosses the network; nested joins compose), local join. - **Aggregation strategies**: two-phase (partial + merge), shuffle by group keys, local; `distributed_aggregation_memory_efficient` and `distributed_plan_force_shuffle_aggregation` are honored. - **Top-N**: two-stage distributed top-N (per-node bounded sort, sorted-merge gather, coordinator limit); disabled under `exact_rows_before_limit`, which needs the full row count. - **Read strategies**: parallel N-way read, replicated read, local read. For `FINAL`, #108148 (already in master) taught the rule-based distributed plan to split a `FINAL` read into disjoint primary-key-range buckets where that is safe; Cascades now reuses that machinery, so `FINAL` no longer forces a serial read here either. The coordinator ships each bucket's marks in the `read_bucket` task parameters. - **Properties and enforcers**: distribution (node count, replication, partitioning columns with equivalence classes and the types the keys are cast to before hashing) and sorting; when a plan alternative lacks a required property, an enforcer inserts the step that provides it (`Gather`/`Shuffle`/`Broadcast`/`ScatterExchange`, `Sort`). - **Transformations**: join commutativity (only for semantics-preserving joins: `INNER ALL`, `CROSS`, `SEMI`/`ANY`/`ANTI`; never `ASOF`), two-phase aggregation split, two-stage top-N split. - **Cost model**: `work`, `network`, and `sequential` components, a fixed per-exchange overhead, broadcast costed per receiving node, and statistics clamped to join kind and strictness semantics; exchange costs use per-row byte widths measured from the parts' column sizes (followed through renames, not derived from types); all weights and calibration constants are overridable at query time. What this improves over the rule-based distributed planner, on TPC-H plans. The rule-based planner broadcasts a small table when its read is below `distributed_plan_max_rows_to_broadcast`, but it often cannot size the result of a join, so a small join result (`nation x region`, 5 rows after the region filter) is scattered across nodes, joined there, and shuffled again (repartitioned across nodes) by the next join key. It also often inserts a shuffle at join and aggregation boundaries even when the rows are already divided by the right key. Cascades estimates sizes through joins, knows which partitioning already holds, compares broadcast against shuffle by cost for each join, and recomputes a small deterministic join on every node when that is cheaper than moving its result. On TPC-H SF100 over 8 nodes (same binary, same run window; times are server-side averages of the hot runs, Cascades averaged over two runs that produced identical plans) the join-heavy queries improve: | Query | Rule-based -> Cascades | What changed in the plan | |---|---|---| | Q18 | 12.03 s -> 6.94 s | Fewer shuffles (8 -> 5): the aggregated `lineitem` subquery joins without a re-shuffle (`l_orderkey = o_orderkey`), and the final GROUP BY reuses the existing partitioning. | | Q21 | 8.02 s -> 5.54 s | 34 exchanges -> 14: all four `supplier x nation` joins are recomputed by every node over full local reads, so nothing is gathered or broadcast for them. | | Q09 | 4.19 s -> 2.25 s | 11 exchanges -> 6: `part` and `nation` are read in full by every node, so `lineitem` and `supplier` are not shuffled to meet them. | | Q08 | 2.32 s -> 1.00 s | 15 exchanges -> 7: `nation x region` (5 rows) is recomputed by every node; `part` is read in full per node, so `lineitem` is not shuffled to join it. | | Q02 | 2.24 s -> 0.93 s | 16 exchanges -> 5: the `supplier x nation x region` chain is recomputed by every node, so only `partsupp` and `supplier` are shuffled. The top-100 sort becomes two-stage, sending at most 100 rows per node. | | Q17 | 2.77 s -> 1.87 s | The small per-part average is broadcast to every node, so the outer 600M-row `lineitem` read is not shuffled. | | Q11 | 0.49 s -> 0.23 s | 5 exchanges -> 2: `supplier x nation` is recomputed by every node and `partsupp` joins it in place. | The remaining queries change less. Summed over all 22 queries, hot server time drops from 46.0 s to 32.1 s (about 30% lower). The largest regression is Q15 (0.22 s -> 0.39 s): the extra time is initiator-side planning — the memo search and statistics loading currently repeat for the view subqueries; fixing that is a follow-up. `EXPLAIN pretty = 1, estimates = 1` shows the chosen plan with a row estimate and the accumulated cost for each step. Also in this PR, two improvements to the shared bucketed-read machinery (they benefit the rule-based path too): the `FINAL` layer split no longer depends on the coordinator's core count, and a many-partition `FINAL` split groups its layers into the target task count instead of falling back to a serial read. Design, a worked example on TPC-H data (a simplified 3-table query traced through the memo), and current limitations are documented in `src/Processors/QueryPlan/Optimizations/Cascades/ARCHITECTURE.md`. Plan-shape tests cover the actual TPC-H queries (`03836_tpch_join_order_plans`). Disabled by default. Requires the analyzer; remote execution requires the stateless-worker configuration, while `distributed_plan_execute_locally = 1` runs the stages in-process without it: ```sql SET enable_cascades_optimizer = 1, make_distributed_plan = 1; ``` For tests, `param__internal_cascades_cluster_node_count` overrides the cluster size, `param__internal_cascades_cost_config` overrides the cost model configuration, `param__internal_join_table_stat_hints` injects table statistics, `param__internal_cascades_task_limit` lowers the task budget (it can never raise it). Related: https://github.com/ClickHouse/ClickHouse/pull/106020 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental Cascades cost-based optimizer for distributed query plans, enabled by `enable_cascades_optimizer = 1` together with `make_distributed_plan = 1`. It chooses between shuffle, broadcast, replicated, and local join strategies, two-phase, shuffle, and local aggregation, two-stage distributed top-N, and parallel and replicated reads by estimated cost, inserting exchange operators as needed. ### Documentation entry for user-facing changes - [ ] Documentation written in [/docs](https://github.com/ClickHouse/ClickHouse/tree/master/docs)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86353",
          "createdAt": "2025-08-28T11:27:09Z",
          "updatedAt": "2026-08-13T15:25:43Z",
          "timestamp": "2026-08-13T15:25:43Z",
          "metrics": {
            "reactions": 21,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:31b79a5e1d73a5519aa5",
        "signalId": "github:ClickHouse/ClickHouse:issue:114649",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114649",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Segmentation fault in `UniqExactSet::merge()` with `countDistinct(String)` and GROUP BY",
          "text": "### Company or project name _No response_ ### Describe what's wrong A SELECT query crashes `clickhouse-server` with SIGSEGV while merging aggregate states for: ```sql countDistinct(note) = 1 ```` ClickHouse version: ```text 26.3.17.110 (official build) git hash: 59141459d999fabdcf3d1dd88cdd5ed6f10136ee ``` The relevant part of the query can be simplified to: ```sql SELECT count() FROM ( SELECT countIf(event_id, note = '') AS cnt, min(event_time) AS first_event, max(event_time) AS last_event FROM events WHERE event_time >= ... AND event_time <= ... GROUP BY key1, key2, key3, key4, key5, (topic, partition) HAVING cnt > 1 AND countDistinct(note) = 1 ) ``` The server crashes with: ```text Received signal Segmentation fault (11) Address: NULL pointer. Access: read. ``` Relevant stack frames: ```text TwoLevelHashTable<...>::TwoLevelHashTable(...) DB::UniqExactSet<...>::merge(...) DB::AggregateFunctionUniqExactData<String, true>::mergeBatch(...) DB::Aggregator::mergeStreamsImplCase(...) DB::Aggregator::mergeBlocks(...) DB::MergingAggregatedBucketTransform::transform(...) ``` So the crash happens while merging `uniqExact` states, apparently around single-level / two-level hash table merging or conversion. ### Workaround / additional evidence In this particular query: ```sql countIf(event_id, note = '') > 1 AND countDistinct(note) = 1 ``` can be replaced with the equivalent condition: ```sql countIf(event_id, note = '') > 1 AND countIf(note != '') = 0 ``` After replacing `countDistinct(note) = 1` with: ```sql countIf(note != '') = 0 ``` the same query completes successfully and no longer crashes the server. This strongly suggests that the crash is specifically related to the `uniqExact` / `countDistinct` aggregation path rather than the rest of the query. ### Does it reproduce on the most recent release? Yes ### How to reproduce See the description section. We didn't found easy way to reproduce this bug. ### Expected behavior The query should either complete successfully or fail with a ClickHouse exception. A SELECT query should not crash the server process. ### Error message and/or stacktrace [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.565819 [ 136891 ] <Fatal> BaseDaemon: Address: NULL pointer. Access: read. Sent by the kernel. [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.565835 [ 136891 ] <Fatal> BaseDaemon: Stack trace: 0x0000000019771551 0x0000000019773138 0x0000000019769925 0x000000001c7844ca 0x000000001c67f1c5 0x000000001c704c60 0x000000001c70f634 0x000000001f0dca4b 0x000000001ed504cb 0x000000001ed7157d 0x000000001ed62584 0x000000001ed66983 0x0000000017cc8816 0x0000000017ccfa1d 0x0000000017cc597d 0x0000000017ccd15a 0x00007f2901cc11f5 0x00007f2901d418ec [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566004 [ 136891 ] <Fatal> BaseDaemon: 2. TwoLevelHashTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>, HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>>, 8ul>::TwoLevelHashTable<HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>>>(HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>> const&) @ 0x0000000019771551 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566061 [ 136891 ] <Fatal> BaseDaemon: 3. DB::UniqExactSet<HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>>, TwoLevelHashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>>>::merge(DB::UniqExactSet<HashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, HashTableGrower<3ul>, AllocatorWithStackMemory<Allocator<true, true>, 128ul, 1ul>>, TwoLevelHashSetTable<wide::integer<128ul, unsigned int>, HashTableCell<wide::integer<128ul, unsigned int>, UInt128TrivialHash, HashTableNoState>, UInt128TrivialHash, TwoLevelHashTableGrower<8ul>, Allocator<true, true>>> const&, ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>*, std::atomic<bool>*) @ 0x0000000019773138 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566118 [ 136891 ] <Fatal> BaseDaemon: 4. DB::IAggregateFunctionHelper<DB::AggregateFunctionUniq<String, DB::AggregateFunctionUniqExactData<String, true>>>::mergeBatch(unsigned long, unsigned long, char**, unsigned long, char* const*, ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>&, std::atomic<bool>&, DB::Arena*) const @ 0x0000000019769925 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566196 [ 136891 ] <Fatal> BaseDaemon: 5. void DB::Aggregator::mergeStreamsImplCase<DB::ColumnsHashing::HashMethodSerialized<PairNoInit<std::basic_string_view<char, std::char_traits<char>>, char*>, char*, true, true>, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>>(DB::Arena*, DB::ColumnsHashing::HashMethodSerialized<PairNoInit<std::basic_string_view<char, std::char_traits<char>>, char*>, char*, true, true>&, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>&, bool, char*, unsigned long, unsigned long, std::vector<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, std::allocator<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*>> const&, std::atomic<bool>&, DB::Arena*) const @ 0x000000001c7844ca [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566288 [ 136891 ] <Fatal> BaseDaemon: 6. void DB::Aggregator::mergeStreamsImpl<DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>>(DB::Arena*, DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>&, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>&, char*, DB::ColumnsHashing::LastElementCacheStats&, bool, unsigned long, unsigned long, std::vector<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*, std::allocator<DB::PODArray<char*, 4096ul, Allocator<false, false>, 63ul, 64ul> const*>> const&, std::vector<DB::IColumn const*, std::allocator<DB::IColumn const*>> const&, std::atomic<bool>&, DB::Arena*) const @ 0x000000001c67f1c5 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566373 [ 136891 ] <Fatal> BaseDaemon: 7. void DB::Aggregator::mergeStreamsImpl<DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>>(std::vector<COW<DB::IColumn>::immutable_ptr<DB::IColumn>, std::allocator<COW<DB::IColumn>::immutable_ptr<DB::IColumn>>> const&, unsigned long, DB::Arena*, DB::AggregationMethodSerialized<HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>, true, true>&, HashMapTable<std::basic_string_view<char, std::char_traits<char>>, HashMapCellWithSavedHash<std::basic_string_view<char, std::char_traits<char>>, char*, StringViewHash64, HashTableNoState>, StringViewHash64, HashTableGrowerWithPrecalculation<8ul>, Allocator<true, true>>&, char*, DB::ColumnsHashing::LastElementCacheStats&, bool, std::atomic<bool>&, DB::Arena*) const @ 0x000000001c704c60 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566410 [ 136891 ] <Fatal> BaseDaemon: 8. DB::Aggregator::mergeBlocks(std::list<DB::Aggregator::AggregatedChunk, std::allocator<DB::Aggregator::AggregatedChunk>>&, bool, std::atomic<bool>&) @ 0x000000001c70f634 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566480 [ 136891 ] <Fatal> BaseDaemon: 9. DB::MergingAggregatedBucketTransform::transform(DB::Chunk&) @ 0x000000001f0dca4b [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566528 [ 136891 ] <Fatal> BaseDaemon: 10. DB::ISimpleTransform::work() @ 0x000000001ed504cb [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566562 [ 136891 ] <Fatal> BaseDaemon: 11. DB::ExecutionThreadContext::executeTask() @ 0x000000001ed7157d [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566596 [ 136891 ] <Fatal> BaseDaemon: 12. DB::PipelineExecutor::executeStepImpl(unsigned long, DB::IAcquiredSlot*, std::atomic<bool>*) @ 0x000000001ed62584 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566669 [ 136891 ] <Fatal> BaseDaemon: 13. void std::__function::__policy_func<void ()>::__call_func[abi:fe210105]<DB::PipelineExecutor::spawnThreads(std::shared_ptr<DB::IAcquiredSlot>)::$_0>(std::__function::__policy_storage const*) @ 0x000000001ed66983 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566712 [ 136891 ] <Fatal> BaseDaemon: 14. ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x0000000017cc8816 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566774 [ 136891 ] <Fatal> BaseDaemon: 15. void std::__function::__policy_func<void ()>::__call_func[abi:fe210105]<ThreadFromGlobalPoolImpl<false, true>::ThreadFromGlobalPoolImpl<void (ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::*)(), ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool*>(void (ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool*&&)::'lambda'()>(std::__function::__policy_storage const*) @ 0x0000000017ccfa1d [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566801 [ 136891 ] <Fatal> BaseDaemon: 16. ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x0000000017cc597d [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566842 [ 136891 ] <Fatal> BaseDaemon: 17. void* std::__thread_proxy[abi:fe210105]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x0000000017ccd15a [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566914 [ 136891 ] <Fatal> BaseDaemon: 18. ? @ 0x00000000000891f5 [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:55.566934 [ 136891 ] <Fatal> BaseDaemon: 19. ? @ 0x00000000001098ec [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:56.281387 [ 136891 ] <Fatal> BaseDaemon: Integrity check of the executable successfully passed (checksum: 4079FE6D02AC0B0EFCF8B7F04A4EBE95) [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:56.281479 [ 136891 ] <Fatal> BaseDaemon: ClickHouse version 26.3.17.110 is old and should be upgraded to the latest version. [site1-telia-vm-qa-cas1-1] 2026.08.13 13:38:56.281657 [ 136891 ] <Fatal> BaseDaemon: Changed settings: max_query_size = 800000, connect_timeout_with_failover_ms = 1000, connect_timeout_with_failover_secure_ms = 2000, use_uncompressed_cache = false, distributed_foreground_insert = true, load_balancing = 'random', enable_positional_arguments = false, log_queries_cut_to_length = 250000, distributed_product_mode = 'global', insert_quorum = 0, select_sequential_consistency = 0, max_http_get_redirects = 1, any_join_distinct_right_table_keys = true, distributed_ddl_task_timeout = 30, max_bytes_before_external_group_by = 15000000000, max_bytes_before_external_sort = 15000000000, max_execution_time = 7200., max_ast_depth = 2000, max_ast_elements = 100000, max_expanded_ast_elements = 1000000, max_memory_usage = 15000000000, formatdatetime_parsedatetime_m_is_month_name = false, allow_drop_detached = true, deduplicate_blocks_in_dependent_materialized_views = false, async_insert = false, allow_experimental_analyzer = false, allow_experimental_window_functions = true, background_schedule_pool_size = 100, output_format_json_quote_64bit_integers = true, output_format_pretty_max_value_width = 250000 Error on processing query: Code: 32. DB::Exception: Attempt to read after eof: while receiving packet from 127.0.0.1:9017, local address: 127.0.0.1:14871. (ATTEMPT_TO_READ_AFTER_EOF) (version 26.3.17.110 (official build)) ### Related issues and pull requests This looks related to the previous `uniqExact` parallel merge issues: * #108912 * #108928 * #109389 However, the crashing build already contains the fixes from those changes, so this may be another issue in the same code path. ### Additional context _No response_",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114649",
          "createdAt": "2026-08-13T14:34:08Z",
          "updatedAt": "2026-08-13T15:25:32Z",
          "timestamp": "2026-08-13T15:25:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "crash"
          ],
          "author": "vasyaabr",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:ac3b5e3e0f8770cb7d4e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112973",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112973",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `sorted_merge` and `parallel_sorted_merge` join algorithms",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109005 Implements the two algorithms planned in [this discussion](https://github.com/ClickHouse/ClickHouse/pull/109005#discussion_r3694159098): `full_sorting_merge` and `parallel_full_sorting_merge` are always supported, so anything listed after them in `join_algorithm` is unreachable — listing them is an unconditional choice, not a preference. The new `sorted_merge` and `parallel_sorted_merge` algorithms execute the same merge join, but are **available only when both join inputs can be efficiently read in the order of the join keys** (e.g. MergeTree tables whose primary key starts with the join keys), so the pre-join sorts become cheap `FinishSorting` or disappear. When the tables' order cannot be exploited, the selection falls through to the next algorithm in the list. That makes them meaningful as a high-priority preference: `join_algorithm = 'sorted_merge,parallel_hash'` uses the streaming, low-memory merge join exactly when it is certainly beneficial, and a hash join otherwise. `sorted_merge` runs a single in-order merge join. `parallel_sorted_merge` additionally shards the join by ranges of the tables' common primary-key prefix into independent per-shard merge joins running in parallel — the same source-side sharding `query_plan_join_shard_by_pk_ranges` applies, enabled for this join by the algorithm itself; the in-order reads stay intact (no scatter, no re-sort). When the sharding cannot apply (e.g. an `ASOF` join), it degrades to a single `sorted_merge`. Implementation notes: - Eligibility is decided during plan physicalization, where the input subplans are visible: `JoinStepLogical::inputsCanBeReadInJoinKeyOrder` finds the `ReadFromMergeTree` below each input (mirroring `findReadingStep` of `optimizeReadInOrder`, without descending through nested joins) and probes the actual read-in-order matcher (`wouldReadInOrderBeUseful`, the same side-effect-free probe `topKThroughJoin` uses) with the join-key sort description. The predicted decision degrades gracefully in both directions: a false positive runs like `full_sorting_merge` (with a full sort), a false negative falls through to the next algorithm. - The same memoized predicate makes `tryAddJoinRuntimeFilter` keep its hands off an eligible join listed before the first hash-family algorithm. Without this, planting a runtime filter erases the merge algorithms from the list (a merge join reads both sides concurrently and cannot use a runtime filter), silently overriding the priority order — the defining feature of these algorithms. For non-eligible joins the runtime filter (and the erasure) stays, because those algorithms are not selectable anyway. When `applyParallelReplicas` later breaks the eligibility (a join input becomes a distributed read), the filter pass is re-run for exactly the joins it had skipped, so the `hash` fall-through gets its runtime filter back. - `FullSortingMergeJoin` now carries the selected algorithm (`getSelectedAlgorithm`) instead of an `is_parallel` flag; the hash-scatter rewrite (`optimizeParallelFullSortingMergeJoin`) stays exclusive to `parallel_full_sorting_merge`, and `optimizeJoinByShards` runs in a restricted mode (only `parallel_sorted_merge`-selected joins, with a cheap pre-scan bail-out) when `query_plan_join_shard_by_pk_ranges` is off. When the sharded stream counts diverge at pipeline-building time (e.g. a data-dependent `PREWHERE` prunes one side to a single empty stream), `JoinStep` merges each side's per-shard sorted streams back into one sorted stream and runs the single-stream merge join instead of failing (the same approach as #109393, applied at the sharding fallback). - The `CreateSetAndFilterOnTheFlyStep` pair (`max_rows_in_set_to_optimize_join`) is not added for sorted-merge joins: it sits between the read and the sort and would defeat the in-order read the algorithm was selected for. - The old analyzer has no query plan at selection time, so there the algorithms are never selected and the list falls through (documented). With only `sorted_merge` listed and no exploitable order, the query fails with `NOT_IMPLEMENTED`, like other unsupported single-algorithm configurations. - The known lower-priority-fallback side effects of listing merge algorithms (stricter `USING` key-type inference, `topKThroughJoin` deferral) extend to the new values and are documented in the `join_algorithm` setting description. Tests: `04669_sorted_merge_join_selection` pins the selection gating via `EXPLAIN PIPELINE` (selected on a primary-key join with no re-sort, falls through on non-key joins / disabled read-in-order / old analyzer, priority order respected, error when listed alone without exploitable order) and correctness against `hash` for `INNER`/`LEFT`/`RIGHT`/`FULL`/`ANY`/`join_use_nulls`. `04670_parallel_sorted_merge_join` pins the primary-key-range sharding (`Sharding:` marker with `query_plan_join_shard_by_pk_ranges = 0`, no `ScatterByPartitionTransform`, no `MergeSortingTransform`), the `ASOF` degradation, and correctness. `04824_sorted_merge_join_parallel_replicas_fallthrough` and `04894_sorted_merge_join_parallel_replicas_runtime_filter` pin the parallel-replicas edge (fall-through to `hash` with the runtime filter restored), `04893_parallel_sorted_merge_join_shard_stream_divergence` pins the diverged-shard degradation, and `04760`/`04850` pin the `join_use_nulls` and `query_plan_join_shard_by_pk_ranges` contracts. The PR #109005 regression tests and the join runtime filter tests pass unchanged. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added new `join_algorithm` values `sorted_merge` and `parallel_sorted_merge`: merge-join algorithms that are available only when both join inputs can be efficiently read in the order of the join keys (so the join benefits from the tables' order instead of sorting), and otherwise fall through to the next algorithm in the list. `parallel_sorted_merge` additionally shards the join by primary-key ranges into independent per-shard merge joins running in parallel. Listing them first, e.g. `join_algorithm = 'sorted_merge,parallel_hash'`, uses the streaming low-memory merge join exactly when it is certainly beneficial. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112973",
          "createdAt": "2026-08-02T03:03:03Z",
          "updatedAt": "2026-08-13T15:24:35Z",
          "timestamp": "2026-08-13T15:24:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2f8a06f15e37d188cc67",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114298",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114298",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113509 to 26.6: Add a dedicated thread pool for lightweight snapshot creation",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113509 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114240 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31479296882/job/93740282532) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.34` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114298",
          "createdAt": "2026-08-11T10:05:22Z",
          "updatedAt": "2026-08-13T15:24:31Z",
          "timestamp": "2026-08-13T15:24:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "closed",
          "assignees": [
            "Diskein",
            "hanfei1991"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:02e5b5436bd14d03132d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114625",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114625",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix the bit-sliced full adder in `groupNumericIndexedVector`",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/106208 Related: https://github.com/ClickHouse/ClickHouse/pull/110072 `addValue` set a bit when the computed sum bit was 1 but never cleared it when the sum bit was 0, so every carry left the lower bit set and the write path was not addition: ```sql SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), val)) FROM (SELECT arrayJoin([toInt64(10), toInt64(10)]) AS val); -- {5:30}, expected {5:20} SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), toInt64(1))) FROM numbers(8); -- {5:255}, expected {5:8} ``` Any repeated index whose addends share a set bit is affected, on every index type and in both the small and the promoted representation. `n` rows of value `1` at one index accumulate to `2^n - 1`. Additions whose bits are disjoint need no carry and were already correct (`10 + 5` gives `15`), which is why the existing tests and the documented examples — all of which use distinct indexes — did not catch it. `merge` and `numericIndexedVectorPointwiseAdd` share `pointwiseAddInplace`, which computes whole-bitmap XORs and assigns the result, so clearing is implicit there and those paths were already correct. Only the per-row path was wrong, which is why `numericIndexedVectorAllValueSum` disagreed with `sum(value)` over the same rows. `RoaringBitmapWithSmallSet` had no way to clear an element, so this adds a `remove`. `SmallSet` has no erase and the small set holds at most `small_set_size` elements, so that path rebuilds it without the removed value. `zero_indexes` is now maintained too. It holds the present indexes whose value is zero, so it has to gain an index when an update drives the value to zero and lose it when the value becomes non-zero — the same invariant the pointwise operations restore when they finish: ```cpp /// For any of the total_indexes, if it is not in the non-zero index of the result, the result is 0. total_indexes->rb_andnot(*getAllNonZeroIndex()); zero_indexes = total_indexes; ``` With that, adding `5` and then `-5` row by row produces `{5:0}`, matching what merging the two values already produced, and `numericIndexedVectorGetValue` and `numericIndexedVectorCardinality` agree with the map. `numericIndexedVectorBuild` also goes through `addValue`, but from a map whose keys are unique, so it never carried and is unaffected. The test asserts `numericIndexedVectorAllValueSum` equals `sum(value)` over repeated indexes with negative and fractional values, that the row-by-row path agrees with the pointwise path, and that an index driven to zero is present with value zero. Eight of its eleven assertions fail without the fix; the three that pass are controls — an addition with disjoint bits, the pointwise reference path, and adding zero to an index that already holds a value. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `groupNumericIndexedVector` producing wrong values when the same index appears in more than one row. The bit-sliced adder never cleared a bit when the computed sum bit was zero, so every carry left the lower bit set: eight rows of value `1` at one index accumulated to `255` instead of `8`, and `10 + 10` produced `30` instead of `20`. Values that share no set bits were unaffected. An index whose value is driven to zero is now reported as present with value zero, consistent with `numericIndexedVectorPointwiseAdd` and with merging aggregate states. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114625",
          "createdAt": "2026-08-13T12:12:29Z",
          "updatedAt": "2026-08-13T15:23:17Z",
          "timestamp": "2026-08-13T15:23:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7f4323b7561f41e82702",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114548",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114548",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support trailing commas in PromQL grouping labels",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Accept trailing commas in PromQL grouping label lists. ### Description The current Prometheus operators guide explicitly documents trailing commas in label lists. Both `(label1, label2)` and `(label1, label2,)` are valid syntax. Prometheus has supported this since 2.16.0. The upstream change was merged in [PromQL: Support trailing commas in grouping opts](https://github.com/prometheus/prometheus/pull/6480) and is listed in the [Prometheus changelog](https://github.com/prometheus/prometheus/blob/main/CHANGELOG.md). It followed an upstream request for consistency with trailing commas already supported in label matchers ([issue #6470](https://github.com/prometheus/prometheus/issues/6470)). This is useful for generated or templated queries where a label list can be assembled with a final comma. ClickHouse previously rejected these valid PromQL queries. The shared ANTLR `labelNameList` rule now accepts one optional trailing comma for `by`, `without`, `on`, `ignoring`, `group_left`, and `group_right`. Examples: - `sum by (job,) (up)` - `sum by (job, instance,) (up)` - `foo + on(job,) bar` - `foo + ignoring(instance,) bar` ### References - [Prometheus operators guide](https://prometheus.io/docs/prometheus/latest/querying/operators/) - [Prometheus implementation PR #6480](https://github.com/prometheus/prometheus/pull/6480) - [Prometheus issue #6470](https://github.com/prometheus/prometheus/issues/6470) - [Prometheus changelog](https://github.com/prometheus/prometheus/blob/main/CHANGELOG.md) ### Tests - Added `PromQLParser.TrailingCommasInGroupingLabelLists`. - Covered one-label and multi-label trailing commas. - Covered malformed empty and double-comma lists with rejection assertions. - Regenerated the ANTLR parser artifacts. - Ran parser smoke tests for all six forms and malformed empty/double-comma lists.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114548",
          "createdAt": "2026-08-12T21:35:23Z",
          "updatedAt": "2026-08-13T15:22:55Z",
          "timestamp": "2026-08-13T15:22:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "submodule changed",
            "can be tested",
            "comp-promql"
          ],
          "author": "fallintoplace",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5132b1bbf5f06769c21c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113895",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113895",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Strip the cosmetic parenthesized flag before comparing stored definitions",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/92340 Related: https://github.com/ClickHouse/ClickHouse/pull/110833 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed comparison of stored table definitions that were written with redundant parentheses (`PARTITION BY (a)`, `PRIMARY KEY (key)`). Since #92340 the formatter preserves those parentheses, so `ATTACH`/`REPLACE`/`MOVE PARTITION FROM` rejected two otherwise identical tables with `Tables have different partition key`, and a `KeeperMap` table created by 26.5 or 26.6 could not be opened by a server of another version. ### Description #92340 started preserving the parentheses a user writes around a definition expression. They are cosmetic, but stored table metadata is compared as text against a form that may have been written by another server version, so two identical definitions began to differ as strings. Two user-visible consequences: - **`ATTACH PARTITION FROM`.** A table declared `PARTITION BY (a)` no longer matches one declared `PARTITION BY a`. This requires neither an upgrade nor a mixed-version cluster: the comparison is between two in-memory ASTs of which only one carries the flag, so both tables created by the same binary already fail. Measured on released binaries, 26.3 and 26.4 accept the pair and 26.5 and later reject it. ```sql CREATE TABLE src (a UInt32, b UInt32) ENGINE=MergeTree PARTITION BY (a) ORDER BY b; CREATE TABLE dst (a UInt32, b UInt32) ENGINE=MergeTree PARTITION BY a ORDER BY b; INSERT INTO src VALUES (1, 1); ALTER TABLE dst ATTACH PARTITION 1 FROM src; -- 26.4: ok -- 26.6: Code: 36. DB::Exception: Tables have different partition key. (BAD_ARGUMENTS) ``` - **`KeeperMap`.** The primary key is serialized into Keeper and compared there against the text written by whichever version created the table. 26.5 and 26.6 write `primary key: (key)`, every other version writes `primary key: key`, so a server of the other version refuses to open the table: ``` Path ... is already used but the stored primary key definition doesn't match. Stored metadata: ... primary key: (key) local metadata: ... primary key: key ``` On a multi-replica setup the replicas that cannot apply the definition never finish startup. ### Implementation `ReplicatedMergeTreeTableMetadata` already stripped the flag, but only on the top level of an expression list, and the helper was private to that file, so the other two comparison sites never got it. This promotes it to `Parsers/stripArtificialParens.h`, makes it walk the whole tree, and applies it in the two places that were missed. The walk also reaches members that are not in `children` and would otherwise be skipped: the `GROUP BY` keys, the `GROUP BY` assignments and the recompression codec of a TTL element, and the `parameters` and `lambda` of a projection's `APPLY` transformer. `StorageKeeperMap` additionally accepts a stored primary key that still carries the parentheses, so tables already created by 26.5 or 26.6 keep working after this change. If that stored text cannot be parsed the comparison stays strict and the mismatch is reported, rather than being silently accepted. Only the comparison changes. There are no parser or formatter changes: what the user wrote is still what is stored and what `SHOW CREATE` and `system.tables` report, and genuinely different definitions still differ (covered by negative cases in both tests). This also fixes two transposed format arguments in the `KeeperMap` mismatch message, which rendered the path and the field name in the wrong order (`Path columns is already used but the stored /keeper_map_tables/... definition doesn't match`). ### Relationship to #110833 This is the compatibility part of #110833, extracted so it can be reviewed and backported on its own, as requested in https://github.com/ClickHouse/ClickHouse/pull/110833#issuecomment-5058455151. The `getTreeHash` work from that PR is a separate, larger change and is not included here. Credit for the original diagnosis and the wider fix goes to @groeneai. ### Tests - `04821_parenthesized_key_attach_partition_from` covers `ATTACH PARTITION FROM` across the two spellings, asserts that `system.tables` still reports the parentheses the user wrote, and that a genuinely different partition key is still rejected. - `04822_keeper_map_parenthesized_primary_key` asserts that the primary key stored in Keeper does not depend on the spelling, that a second table on the same path with the other spelling opens, and that a different primary key is still rejected.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113895",
          "createdAt": "2026-08-07T22:24:44Z",
          "updatedAt": "2026-08-13T15:22:29Z",
          "timestamp": "2026-08-13T15:22:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "v26.6-must-backport"
          ],
          "author": "fm4v",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5dd1ff0fd529c2004ada",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114316",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114316",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use the vector similarity index for integer reference vectors",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112233 Related: https://github.com/ClickHouse/ClickHouse/issues/114291 The reference vector of an ANN query is extracted only when its array type is `Float64`, `Float32` or `BFloat16` and every element is a `Float64` field. An integer literal such as `[1, 2]` is typed `Array(UInt8)`, so `tryUseVectorSearch` bails out and the query silently falls back to a brute-force scan over the whole table, although `[1, 2]` denotes the same point as `[1.0, 2.0]` and `L2Distance` accepts it. `EXPLAIN indexes = 1` shows no `vector_similarity` entry and gives no hint why, so a one-character difference in a literal becomes a sharp performance cliff on large tables. Native integer arrays are now accepted as reference vectors and their elements are converted to `Float64`, which is the type the reference vector is stored in anyway. Added `02354_vector_search_bug112233`, covering unsigned, signed, mixed integer/float, and not-exactly-representable reference vectors, plus an equality check between the integer and float spellings of the same query. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Vector search queries now use the `vector_similarity` index when the reference vector is written as an integer array literal, e.g. `ORDER BY L2Distance(vec, [1, 2])`. Previously such queries silently fell back to a brute-force scan.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114316",
          "createdAt": "2026-08-11T12:52:06Z",
          "updatedAt": "2026-08-13T15:22:05Z",
          "timestamp": "2026-08-13T15:22:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "hamidr",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ce25b6a7cfbfdbab9543",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114152",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114152",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "[WIP]Add incremental rmv core",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add incremental refreshable materialized views managed via refresh_incremental, which when used each refresh reads only the data committed to the single source table since the previous run and persists the advanced cursor in the RMV's Keeper CoordinationZnode for at-least-once resumption. depends on https://github.com/ClickHouse/ClickHouse/pull/111794 cc @alesapin @Michicosun",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114152",
          "createdAt": "2026-08-10T12:45:59Z",
          "updatedAt": "2026-08-13T15:21:13Z",
          "timestamp": "2026-08-13T15:21:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "SmitaRKulkarni",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:49aa1fa96f30a5d449b3",
        "signalId": "github:ClickHouse/ClickHouse:issue:114481",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114481",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Dropping a nested column group bypasses every ALTER protection",
          "text": "### Company or project name _No response_ ### Describe what's wrong when shared nested offsets are enabled a name in `DROP COLUMN <name>` that is not a column of the table denotes the whole group of columns `<name>.*` and the drop removes all of them The stages of an `ALTER` disagree about this meaning: validation and the mutation stage treat the name as the group, while the metadata update and every protective check (dependent views, unfinished mutations, key columns, dependencies) compare column names exactly and do not see the group. Each mismatch below is one observable consequence ### Does it reproduce on the most recent release? Yes ### How to reproduce ## 1. `DROP COLUMN IF EXISTS <group>` silently destroys the group's data `ALTER` reports OK, `DESCRIBE` still shows `n.a` and `n.b` and the `SELECT` returns `0 0 100`. The columns stay in the table while their values are gone ```sql CREATE TABLE t (`n.a` UInt64, `n.b` UInt64, x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t VALUES (1, 10, 100); ALTER TABLE t DROP COLUMN IF EXISTS n; SELECT * FROM t; -- 0 0 100 ``` Fiddle: https://fiddle.clickhouse.com/7de96186-e49d-45c1-b8af-a56803993cd0 The same split also swallows pending updates: ```sql CREATE TABLE t (`n.a` UInt64, x UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t VALUES (1, 1); ALTER TABLE t UPDATE `n.a` = 2 WHERE 1 SETTINGS mutations_sync = 0; ALTER TABLE t DROP COLUMN IF EXISTS n; SELECT * FROM t; -- 0 1 ``` ## 2. A dependent materialized view does not protect the group's columns ```sql CREATE TABLE src (`n.a` UInt64, `n.b` UInt64, x UInt64) ENGINE = MergeTree ORDER BY x; CREATE MATERIALIZED VIEW mv ENGINE = Null AS SELECT `n.a` FROM src; ALTER TABLE src DROP COLUMN n; INSERT INTO src (x) VALUES (1); ``` Drop succeeds, and from that point every insert into `src` fails with `UNKNOWN_IDENTIFIER`, because the view still selects `n.a`. With `IF EXISTS` the outcome is the data destruction from problem 1, with the view still subscribed Fiddle: https://fiddle.clickhouse.com/4dd388c7-af62-41d4-b097-bf4007e25041 ## 3. An unfinished mutation on a group's column does not block dropping the group ```sql CREATE TABLE t (`n.a` UInt64, x UInt64, c UInt64) ENGINE = MergeTree ORDER BY x; INSERT INTO t VALUES (1, 1, 1); ALTER TABLE t UPDATE c = `n.a` + 1 WHERE 1 SETTINGS mutations_sync = 0; ALTER TABLE t DROP COLUMN n; SELECT command, is_done, latest_fail_reason FROM system.mutations WHERE table = 't'; ``` When the background mutation has not finished yet, the drop passes, the group is removed, and the queued mutation now references a column that no longer exists - it stays in the queue forever until `KILL MUTATION`. Whether the statement is safe depends on a race with the background pool ### Expected behavior _No response_ ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context #114468 #114163",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114481",
          "createdAt": "2026-08-12T12:11:22Z",
          "updatedAt": "2026-08-13T15:20:53Z",
          "timestamp": "2026-08-13T15:20:53Z",
          "metrics": {
            "reactions": 1,
            "comments": 1
          },
          "labels": [
            "potential bug"
          ],
          "author": "m7kss1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:bf29cc084a0ea71f783b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113076",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113076",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Populate query_log views column for inserts through Alias",
          "text": "Share query access info with the forwarded target insert so materialized views triggered on the target are recorded in the `views` column of `system.query_log`. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `views` column in `query_log` for inserts to an `Alias` table whose target triggers materialized views, which was always empty.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113076",
          "createdAt": "2026-08-03T10:16:33Z",
          "updatedAt": "2026-08-13T15:20:13Z",
          "timestamp": "2026-08-13T15:20:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "eclbg",
          "state": "open",
          "assignees": [
            "scanhex12"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b7016bf13a1f0ab033dc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114626",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114626",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not merge-sort a distributed gather whose sort description is all-constant",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 Related: https://github.com/ClickHouse/ClickHouse/issues/106237 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description With `make_distributed_plan = 1`, a window `PARTITION BY <constant>` builds a `GatherSend` fragment whose sort description is entirely constant. `GatherSendStep::updatePipeline` adds an order-preserving `MergingSortedTransform` for any non-empty description, and that merge waits for *every* input stream to have data (`IMergingTransformBase::prepareInitializeInputs`). Hashing a constant key sends all rows to one bucket, so the other branches are never fed and stay `NeedData` while the loaded branch is `PortFull`. Neither side can move: the query deadlocks at zero CPU until `receive_timeout` and then raises `Pipeline stuck` (a logical error, so the server aborts in debug and sanitizer builds). An all-constant description orders nothing, so this change takes the `pipeline.resize(1)` branch that already exists in that function. A `ResizeProcessor` pairs any waiting output with any ready input and has no all-inputs barrier, which is why the code before [4a1ab1e](https://github.com/ClickHouse/ClickHouse/commit/4a1ab1e3d94fb0d) - which replaced an unconditional `resize(1)` with this conditional merge - could not wedge this way. Every row compares equal under an all-constant description, so any interleaving is validly sorted and the contract `GatherReceiveStep` relies on still holds. A description with at least one real column keeps the merge. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 (where this failure was reported) Related: https://github.com/ClickHouse/ClickHouse/issues/106237 (context only, different mechanism) Found by `AST fuzzer (amd_debug, targeted, old_compatibility)`: [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113558&sha=b9fe1650232abf4fefc16992f3efcce44222bebc&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The `BufferedShardByHashTransform` deadlock tracked in #106237 is a different mechanism: that transform does not appear in the failing pipeline at all. The new test `04888` fails on master with `Pipeline stuck` for a constant key, a `LowCardinality` constant and a `Nullable` constant, and passes with this change. Its controls - a real column key, a mixed constant-plus-column key, and the non-distributed plan - pass both before and after, so the change is narrow. `04837_distributed_plan_window_partition_shuffle`, which pins the full distributed plan for a column-keyed window, still passes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114626",
          "createdAt": "2026-08-13T12:23:35Z",
          "updatedAt": "2026-08-13T15:15:43Z",
          "timestamp": "2026-08-13T15:15:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "davenger"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6d6e5ca809d5bb48c1e1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114413",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114413",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix loss of primary-key pruning for `DateTime64` columns under a `toUnixTimestamp` sorting key",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114407 Related: https://github.com/ClickHouse/ClickHouse/issues/79977 Related: https://github.com/ClickHouse/ClickHouse/pull/101814 For a `MergeTree` table with `ORDER BY toUnixTimestamp(ts)` where `ts` is `DateTime64`, a plain range filter on `ts` (e.g. `WHERE ts >= '2026-06-15'`) stopped using the primary key in 26.7: `EXPLAIN indexes = 1` shows `PrimaryKey Condition: true` and every granule of the matched parts is read. The regression came from #101814, which correctly changed the monotonicity gate in `KeyCondition::extractMonotonicFunctionsChainFromKey` to ask `getMonotonicityForRange` over the function's argument type instead of its result type. That exposed a gap: `ToNumberMonotonicity` (the monotonicity implementation behind `toUnixTimestamp` and the `toInt*`/`toUInt*` family) did not support `DateTime64` arguments at all and answered \"unknown\", so the monotonic chain was rejected and the pruning was silently lost. Keys like `toDate(ts)` or `toStartOfHour(ts)` were unaffected because their monotonicity classes handle `DateTime64` explicitly, and `toUnixTimestamp` over a plain `DateTime` column was unaffected because `DateTime` is on the whitelist of `ToNumberMonotonicity`. This PR teaches `ToNumberMonotonicity` to answer for `DateTime64` arguments. The conversion takes the whole number of seconds, truncating the fractional part toward zero, and throws a `DECIMAL_OVERFLOW` exception instead of wrapping around (see `DecimalUtils::convertTo`), so it preserves order everywhere it is defined: - For a signed target of at least 64 bits (e.g. `toInt64`), the conversion is total — the whole number of seconds of any `DateTime64` fits — so it is reported as `is_always_monotonic`, for concrete ranges as well. As a bonus, the mirror-image filter `toInt64(ts) >= c` over `ORDER BY ts` now prunes too (the batched application over columns of index values can never throw), and stays consistent with the exact-ranges `count()` optimization. - For narrower or unsigned targets (e.g. `toUnixTimestamp`, whose result is `UInt32`), the conversion throws for a part of the domain, so it must not claim `is_always_monotonic`: `matchesExactContinuousRange` takes that claim as a promise that the per-range analysis can confirm every granule of a found range, while a part may hold out-of-range values, for which the per-range analysis has to answer \"unknown\". The first version of this PR made exactly that inconsistent claim, and the AST fuzzer found the debug assertion \"Inconsistent `KeyCondition` behavior\" (a logical error, so the stateless runs on the debug build also showed \"Server died\"). Instead, the conversion now reports a new flag `Monotonicity::is_always_monotonic_where_defined` — monotonic over the subset of the domain where the evaluation succeeds — which is consumed only by the constant-pushdown gate in `canConstantBeWrappedByMonotonicFunctions`. It is sound there: stored keys cannot correspond to out-of-range values, because computing the sorting key at insert time would have thrown, and an unrepresentable constant is rejected gracefully by the guards in `applyFunctionChainToColumn`. Per the review findings, those guards are also fixed: the pre-execution range checks are extended from `Date`/`DateTime`/`UInt32` to the whole native integer family, so an out-of-range constant pushed through e.g. `ORDER BY toUInt8(ts)` is rejected gracefully instead of throwing a `DECIMAL_OVERFLOW` exception during index analysis, and the negative-value fast reject is now unsigned-specific, so signed integer keys (e.g. `ORDER BY toInt64(ts)`) keep pruning for pre-`1970-01-01` filters. For partial conversions like `toUnixTimestamp`, the mirror-image case #79977 (`WHERE toUnixTimestamp(ts) >= c` over `ORDER BY ts`, a full scan since 23.1) remains out of scope: index analysis applies chain functions to whole columns of index values (see `applyFunction` in `KeyCondition.cpp`), and a part may contain out-of-range values next to the checked range, for which the batched conversion would throw an exception. For total conversions like `toInt64` it is fixed here. The test covers the restored pruning (with `force_primary_key`), sub-second bounds of the relaxed atom (positive and negative), constants outside the `UInt32` range, `toInt64` sorting keys including pre-`1970-01-01` filters, out-of-range constants over a narrow `toUInt8` key, the mixed-part case where index analysis must not throw an exception, and the exact-ranges `count()` optimization over a total conversion. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a regression in 26.7: for a `MergeTree` table ordered by `toUnixTimestamp` (or another integer conversion) of a `DateTime64` column, a plain range filter on that column no longer used the primary key and read all granules of the matched parts. Additionally, a filter like `toInt64(ts) >= c` over a table ordered by the raw `DateTime64` column now uses the primary key.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114413",
          "createdAt": "2026-08-12T02:15:56Z",
          "updatedAt": "2026-08-13T15:15:36Z",
          "timestamp": "2026-08-13T15:15:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fe68a23cf92490a27779",
        "signalId": "github:ClickHouse/ClickHouse:issue:114639",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114639",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Push dynamic TopN thresholds into MergeTree reads for ORDER BY ... LIMIT (2-3x on ClickBench Q24/Q26)",
          "text": "**Use case.** `ORDER BY ... LIMIT n` over a sorted-by-something-else table with a narrow projection, e.g. ClickBench Q24/Q26: ```sql SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10; SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime, SearchPhrase LIMIT 10; ``` The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 implemented a **dynamic TopN threshold pushdown** (`Push TopN threshold to MergeTreeSource`, commit `2e2f594308f4`, plus `Better query condition cache: make TopN dynamic filters deterministic and reusable`, gated by `max_limit_to_push_down_topn_predicate = 100`): while the partial-sort transform maintains the current top-`n`, the running n-th-best value of the `ORDER BY` key is pushed down into the `MergeTree` read as a dynamic threshold, so granules whose key range cannot beat the current top-`n` are skipped instead of read, decompressed, and sorted. This mechanism was dropped during the upstreaming of that PR (the scaffolding was removed from the branch; the surviving `RewriteOrderByLimitPass` is a different, row-offset-based approach — it is off by default and measures performance-neutral on ClickBench when enabled). Master's lazy materialization (`query_plan_optimize_lazy_materialization`) covers the wide-`SELECT *` case (Q23), but does not prune reads for narrow projections: every granule passing the `WHERE` is still fully processed. Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64), original bench-opt build (25.9.1.1) vs master `405e218ff`: Q24 0.009 s vs 0.013 s, Q26 0.008 s vs 0.013 s (1.2–1.3x with times this small). The gap is much larger on the official `c7a.metal-48xl` numbers: Q24 0.014 s vs 0.046 s, Q26 0.014 s vs 0.044 s (**2–3x**). A related consideration from the original design: the dynamic filter interacts with the query condition cache, so the thresholds need to be deterministic/reusable (or excluded from the cache key) — the original branch had a follow-up commit specifically making the TopN dynamic filters deterministic for that reason. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114639",
          "createdAt": "2026-08-13T12:58:53Z",
          "updatedAt": "2026-08-13T15:14:49Z",
          "timestamp": "2026-08-13T15:14:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6c03474de9bcbb828e53",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112309",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112309",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add hierarchicalKMeans and assignCentroid",
          "text": "Adds functions for computing cluster centroids and assigning new vectors to clusters. Ref : https://github.com/ClickHouse/ClickHouse/issues/112578 ### Changelog category - Experimental Feature ### Changelog entry - Added` hierarchicalKMeans()` and `assignCentroid()` functions.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112309",
          "createdAt": "2026-07-28T16:11:44Z",
          "updatedAt": "2026-08-13T15:14:28Z",
          "timestamp": "2026-08-13T15:14:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "shankar-iyer",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:cb98ec26b6e24bcc2e79",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96130",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96130",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Randomize tests with DETACH/ATTACH table before query execution",
          "text": "Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that enable randomly detaching and reattaching tables used in a query before its execution. This is a testing-only feature designed to find bugs related to table reattachment. Before executing a query, the system collects all tables referenced in the AST, and for each eligible table (stores data on disk, supports detaching, has no action locks or dependencies), it performs a `DETACH` followed by `ATTACH`. Changes: - Add `supportsDetachingTables` virtual method to `IDatabase` (overridden to `false` for engines that do not support non-permanent `DETACH TABLE`: `DatabaseDictionary`, `DatabaseReplicated`, `DatabaseSQLite`, `DatabaseBackup`, `DatabaseFilesystem`, `DatabaseHDFS`, `DatabaseS3`, `DatabaseURL`, `DatabaseRemote`, `DatabaseDataLake`, `DatabaseMaterializedPostgreSQL`) - Add `has`/`hasAny` methods to `ActionLocksManager` for checking existing locks (skipping expired `weak_ptr` entries) - Add table collection visitor and reattach logic in `executeQuery` (runs after AST validations, process list admission, and external tables initialization; skips `EXPLAIN`, transactions, internal/non-initial queries, and CTE name collisions) - Fix off-by-one in `MergeTreeDeduplicationLog::dropOutdatedLogs` (don't drop the active log) and add `sync` call in shutdown - Add `no-random-detach` tag to tests incompatible with this feature - Add `--no-random-detach` and `--reattach-tables-probability` options to `clickhouse-test` - Add `02461_reattach_tables` test Continuation of #55943. Continuation of #42336 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that randomly `DETACH` and `ATTACH` tables used in a query before its execution. This is a testing-only feature that helps find reattachment-related bugs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Introduces new pre-execution mutations (internal `DETACH`/`ATTACH`) in `executeQuery`, which can affect table availability and concurrency behavior if enabled; guarded by new experimental settings but touches core query execution paths. > > **Overview** > Adds experimental settings `reattach_tables_before_query_execution` and `..._probability` to optionally **DETACH and ATTACH back** eligible tables referenced by a query immediately before execution, including AST table discovery that accounts for CTE scoping, privilege checks, dependency/lock checks, and safety skips (e.g. `system`, non-disk storages, dynamic-structure columns, transactions, `EXPLAIN`, internal/non-initial queries). > > Extends `IDatabase` with `supportsDetachingTables()` and marks multiple database engines as not supporting non-permanent detach; adds `ActionLocksManager::has/hasAny` helpers to avoid detaching tables with active action locks. Updates the test runner and stress tooling to randomize this behavior (with `--no-random-detach` and probability control), adds a new `02461_reattach_tables` test, and tags many existing tests to opt out where DETACH/ATTACH would add flakiness/overhead. Also fixes `MergeTreeDeduplicationLog` cleanup to avoid dropping the active log and ensures writer `sync()` on shutdown. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit a369371ff61cb1934815ea1e8caf7debd83c98c4. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96130",
          "createdAt": "2026-02-05T23:27:47Z",
          "updatedAt": "2026-08-13T15:13:59Z",
          "timestamp": "2026-08-13T15:13:59Z",
          "metrics": {
            "reactions": 2,
            "comments": 91
          },
          "labels": [
            "pr-build"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4fd4140cc8e95ff3ad24",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112667",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112667",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Silk integration",
          "text": "Splits the silk runtime integration out of https://github.com/ClickHouse/ClickHouse/pull/111275, so that it can be reviewed on its own. This adds the plumbing that lets ClickHouse run work on [silk](https://github.com/ClickHouse/silk) fibers, without yet putting any subsystem on them. - `Silk::initializeFiberScheduler` / `Silk::destroyFiberScheduler`, called by the server when the `enable_silk_runtime` server setting is enabled. The fiber stack size is configurable through the `silk.fiber_stack_size` configuration key (320 KiB by default, which leaves enough room for OpenSSL handshakes). - `FiberLocal` - fiber-local storage. A fiber can migrate between operating-system threads, so it must not observe another fiber's `thread_local` state; the values of the registered slots are swapped in and out on every fiber switch instead. `current_thread` (`ThreadStatus`), the OpenTelemetry tracing context, and the memory-tracker and exception blockers are moved to it. - The silk thread-local-storage sanitizer: an LLVM pass in `utils/silk-thread-local-storage-sanitizer` that instruments every `thread_local` access and aborts when a fiber touches raw thread-local storage. Without it, a variable that was not migrated to `FiberLocal` produces silent corruption rather than a diagnostic. It is enabled in the debug and ASan CI builds. - `Silk::ConnectionPool` and `Silk::streamSocketFactory` - a `Connection` pool and a socket factory that suspend the calling fiber instead of blocking the operating-system thread. `PoolBase` and `ConnectionPool` are templated on the lock and the condition variable to make that possible, and `ConnectionPool` stays an alias of the `std::mutex` instantiation, so the existing call sites are unchanged. - Memory that the runtime maps outside the C++ heap - fiber stacks and `io_uring` rings - is charged to `total_memory_tracker` through silk's mmap accounting hooks. - The low-level silk runtime counters are exported to `system.asynchronous_metrics` under a `Silk` prefix. - `Common/Fiber.h` and `Common/FiberStack.h` are renamed to `Common/StackfulCoroutine.h` and `Common/CoroutineStack.h`. They implement the boost-context coroutines used by `AsyncTaskExecutor`, which are unrelated to silk fibers, and having two different things called \"fiber\" in the same codebase is confusing. Related: https://github.com/ClickHouse/ClickHouse/pull/111275 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental support for the [silk](https://github.com/ClickHouse/silk) fiber runtime, enabled with the `enable_silk_runtime` server setting. When it is enabled, the server initializes the silk fiber scheduler at startup, so that subsystems supporting it can run their jobs on fibers instead of occupying an operating-system thread while waiting for I/O.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112667",
          "createdAt": "2026-07-30T21:32:29Z",
          "updatedAt": "2026-08-13T15:13:24Z",
          "timestamp": "2026-08-13T15:13:24Z",
          "metrics": {
            "reactions": 1,
            "comments": 1
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "CheSema"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ced0e537adb081c1ecaa",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113651",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113651",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for container boot before installing packages in `Install packages`",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Install packages (amd_release)` intermittently fails its `Install server rpm` substep with one line and nothing else: ``` + yum localinstall '--disablerepo=*' --allowerasing -y /packages/clickhouse-server-...rpm ... [Errno 2] No such file or directory: '/var/cache/dnf/metadata_lock.pid' ``` Example: [PR #109299](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109299&sha=3c7aad076ea307634a39a3089615b6a73e208621&name_0=PR&name_1=Install%20packages%20%28amd_release%29). **Root cause.** The image boots systemd, and the Dockerfile deliberately keeps `systemd-tmpfiles-setup.service`, which runs `systemd-tmpfiles --create --remove --boot`. centos:8 ships `/usr/lib/tmpfiles.d/dnf.conf`, whose entire content is remove directives for the dnf lock files, `metadata_lock.pid` among them. `test_install` starts the container `--detach` and `docker exec`s `install.sh` immediately, so `yum` and that lock wipe run concurrently. dnf takes the metadata lock in `Base.fill_sack` before it even opens the rpm files and does not guard against it vanishing, so the substep aborts before any package work: install, start and the smoke test never run. **Change.** Wait for the boot transaction before the first `docker exec`, in the shared `test_install` helper, so every substep of both images is covered. Only the rpm image is affected in practice: ubuntu:22.04 ships no `dnf.conf`, and the dpkg and apt locks survive the same run. Waiting for D-Bus first is the load-bearing half: in the first milliseconds `systemctl` cannot reach systemd at all, so a gate built on it alone silently does nothing in the window it guards. A bare `systemctl start` returned `Failed to connect to bus` in 5 of 8 tries; with the bus wait in front, `rc=0` in 8 of 8. **Validation.** Built the real image locally and amplified the race with 24 concurrent containers, no fault injection: **6 of 144** runs reproduced the exact CI line without the wait, **0 of 144** with it, gate `rc=0` in 144/144. Cost is ~0.17 s per container, so ~3.1 s over the 18 containers a job starts, against a job whose 30-day median is ~240 s. No related open issue found.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113651",
          "createdAt": "2026-08-06T11:22:40Z",
          "updatedAt": "2026-08-13T15:12:31Z",
          "timestamp": "2026-08-13T15:12:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:50726120e9d05087b917",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113383",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113383",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Push tuple element predicates into Parquet and ORC subcolumn reads",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/112575 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Filter pushdown now works for `Tuple` subcolumns in Parquet and ORC files. A predicate such as `WHERE tup.1 = 555555` over `file()`, `s3()` or `url()` now prunes row groups and row index strides using the tuple element's own statistics instead of reading the whole file. ### Description Closes: #112575 Three independent defects, all needed for the reported query to prune. Parquet's reader itself was fine. **The analyzer never produced a subcolumn.** `StorageFile`/`StorageURL`/`StorageObjectStorage` return `false` from `supportsOptimizationToSubcolumns`, so `tupleElement(tup, 1)` was never rewritten to `tup.1`; the pass also accepted only `TableNode`, skipping table functions. That blanket `false` keeps #106147 fixed (`NOT_FOUND_COLUMN_IN_BLOCK` on `.null` in PREWHERE), so rather than flipping it this adds a narrow `supportsOptimizationToTupleElementSubcolumns` virtual defaulting to the existing one, with a `{Tuple, tupleElement}`-only allow-list. `04303_object_storage_prewhere_isnotnull_subcolumn` passes unmodified. Source identity also moves to the table-expression node: every `file()` resolves to the same `_table_function.file` ID, so two such sources shared a key. Accepting table functions generally also makes `format()`, `values()` and `view()` eligible for the other transformers; those storages already answer `supportsSubcolumns()`, so the default covers them. **ORC's search argument builder resolved top-level names only,** so any dotted name emitted `YES_NO_NULL` while the read path in the same file resolved them recursively. Resolving recursively also reaches the flattened-`Nested` descent, which rewrites the type it is given, so the builder keeps the key's own type: an array-typed predicate over a flattened `Nested` leaf is not pushed, since scalar element statistics cannot decide it. **ORC built its KeyCondition from the reader header,** which carries only the parent column, so the predicate degraded to `unknown`. ORC now passes `initKeyConditionOnce` a local copy extended with the tuple element paths the filter references; `FormatFilterInfo`, the Parquet call site and the reader header are untouched. Admission requires a named tuple at every level (unwrapping `Nullable`/`LowCardinality`/`Array`), which refuses Map `.keys`/`.values`: they use `SubstreamType::TupleElement` but have no per-element statistics. <details> <summary>Measurements (100k rows, one row group / stride per 10k)</summary> | Arm | master | this PR | |---|---|---| | ORC `WHERE tup.1 = 55555` | 200000 rows read | **20000** | | ORC `WHERE id = 55555` (control) | 20000 | 20000 | | Parquet `WHERE tup.1 = 55555` | 7 row groups / 0 pruned | **1 / 6** | | Parquet `WHERE id = 55555` (control) | 1 / 6 | 1 / 6 | Results identical in every arm. Refusal arms (Map `.keys`/`.values`, `.null`, `.size0`, unnamed tuple, `Array(Tuple)`, type-mismatch structure hint) return correct results with pruning off. Multi-level `tup.2.1` stays unpruned: correct, and a separate optimization. Regression sweep over `*functions_to_subcolumns*`, `*tuple_element*`, `*_parquet_*`, `*_orc_*` and the named pushdown tests, run on this build and on an unmodified master build for attribution: no regression attributable to this change. New tests are 50/50 green under randomized settings. </details> Also noticed, not touched here: reading a standalone dotted ORC tuple element with its inferred type returns column defaults, because `Nested::flatten` does not descend a `Nullable(Tuple)`. #109741 (open) rewrites that helper for the `Arrow` spelling.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113383",
          "createdAt": "2026-08-04T20:17:22Z",
          "updatedAt": "2026-08-13T15:12:26Z",
          "timestamp": "2026-08-13T15:12:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 18
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4ecea6a6931ed3b03aed",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113357",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113357",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cache the finalized cardinality in uniq statistics",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113038 --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Speed up query planning when column statistics are used. `uniq` and `uniq_v2` statistics now cache the estimated number of distinct values instead of recomputing the sketch on every request, which join order optimization issues many times per query. This is most visible on aarch64. Closes #113038. ### Description `IStatistics::estimateCardinality` is a pure function of the aggregate state, but both implementations recomputed it on every call, and plan optimization calls it many times per query: once per column with statistics in `ConditionSelectivityEstimator::estimateRelationProfileImpl`, and once or twice per equality atom. For `uniq_v2` that recomputation is expensive. `uniq_v2` is `uniqCombined64` with `K = 12`, which selects the fully functional `Denominator` specialization whose `get` (`src/Common/HyperLogLogCounter.h:180-189`) walks all 54 rank buckets in `long double`. On aarch64 that is software binary128, so every step compiles to `__multf3` / `__floatunsitf` / `__addtf3` libcalls; on x86-64 it is native x87 arithmetic. That is why the regression is aarch64 only. Since 26.7 the default `auto_statistics_types` includes `uniq_v2`. The fix memoizes the finalized value in the object owning the state, as `cardinality + 1` so `0` means \"not computed yet\" without reserving a representable cardinality as a sentinel (0 distinct values is legal for an all NULL column). The member is a `mutable std::atomic<UInt64>` because the estimator is shared between concurrent queries; relaxed ordering suffices as all racers compute the identical value. No invalidation is needed: every caller finishes building and merging before any cardinality is read, which I measured across 10387 mutator entries without a single live memo. Both implementations are fixed: `findUniqStats` prefers `Uniq` over `UniqV2`, so `StatisticsUniq` is a live carrier too. On this PR's own arm_release Performance Comparison, `JoinOptimizeMicroseconds` drops 73% on TPC-DS Q14 and 63% on TPC-H Q20, the two queries in the report, with `server_time` -31% and -51%. Estimates and the chosen plan are unchanged (1206 `system.parts_columns` estimate rows and the `EXPLAIN indexes=1` output of a 6 way join are byte identical). The analysis, patch direction and aarch64 measurements are @ egor-click's. This differs from the patch in the issue in replacing the `std::numeric_limits<UInt64>::max()` sentinel with the `+1` encoding, and in fixing `StatisticsUniq` too.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113357",
          "createdAt": "2026-08-04T16:52:54Z",
          "updatedAt": "2026-08-13T15:11:13Z",
          "timestamp": "2026-08-13T15:11:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "hanfei1991"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:3efa0a02995f34c2ee36",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114655",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114655",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #114188 to 26.7: Ignore redundant parentheses in stored table definitions",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114655",
          "createdAt": "2026-08-13T15:09:46Z",
          "updatedAt": "2026-08-13T15:09:54Z",
          "timestamp": "2026-08-13T15:09:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:6f2ac5abbb741fb37e1e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110344",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110344",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong primary-key pruning for toStartOfDay and relative-number functions on out-of-range DateTime64",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/90461 Related: https://github.com/ClickHouse/ClickHouse/pull/108018 --> Related: #90461 Related: #108018 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong `count()` results and dropped rows when a `DateTime64` primary-key column is filtered through `toStartOfDay`, `toRelativeSecondNum`, `toRelativeMinuteNum`, `toRelativeHourNum`, `toRelativeDayNum`, `toRelativeWeekNum`, `toMonthNumSinceEpoch` or `toYearNumSinceEpoch` with values outside the `UInt32`-seconds range (before 1970 or beyond 2106). These functions claim to be always monotonic to the primary index but their standard-precision results wrapped for out-of-range `DateTime64`, breaking primary-key pruning. ### Description The `DateTime64` sibling of #108018 (which fixed the same class for `Date32`). `toStartOfDay` and the relative-number transforms have `FactorTransform = ZeroTransform`, so `IFunctionDateOrDateTime::getMonotonicityForRange` reports them as always monotonic. Their standard-precision `DateTime64` code paths narrowed the result to `UInt32`/`UInt16` without saturating, so for arguments outside that range the value wrapped and the function stopped being monotonic. This makes primary-key range analysis produce exact ranges that extend before the selected mark range, which: - in release builds: silently drops granules holding matching rows, returning a wrong `count()`; - in debug/sanitizer builds: trips `chassert(exact_ranges[i].begin >= range.begin)` in the trivial-count projection optimization (`optimizeUseAggregateProjection.cpp`). Found by the AST fuzzer (amd_msan) on a query that mutated a `Date32` key column to `DateTime64(5)`: `SELECT count() FROM t WHERE toStartOfDay(d) >= toDateTime('2000-01-01 00:00:00','UTC') SETTINGS force_primary_key = 1`. Report: https://s3.amazonaws.com/clickhouse-test-reports/PRs/110310/e242401ecdb1e933c8646546bc2f905d4ed106cd/ast_fuzzer_amd_msan/fatal.log The fix saturates the `DateTime64` `execute` overloads to `[0, result-type max]`, matching the `Date32` fix in #108018, keeping each function monotonic over the whole `DateTime64` domain. Adds `04538_datetime64_zerotransform_monotonicity_pruning`; updates the references of `01768_extended_range`, `04408_datediff_datetime64_overflow` and `02403_enable_extended_results_for_datetime_functions`, which asserted the previous wrapped standard-precision results.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110344",
          "createdAt": "2026-07-14T07:56:02Z",
          "updatedAt": "2026-08-13T15:09:28Z",
          "timestamp": "2026-08-13T15:09:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:c4fc2a266a258825c9eb",
        "signalId": "github:ClickHouse/ClickHouse:issue:114581",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114581",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "IN (SELECT ...) inside a higher-order-function lambda always evaluates to 0 when the query is a derived table or the lambda is in WHERE",
          "text": "**Describe what's wrong** An `IN (SELECT ...)` predicate inside a higher-order-function lambda (`arrayExists`, `arrayFilter`, `arrayMap`, ...) always evaluates to `0` when the enclosing SELECT is used as a derived table (or when the lambda sits in an outer `WHERE`). The identical expression at the top level returns the correct result. **Does it reproduce on the most recent release?** Reproduces on current master, `26.8.1.1068` and `26.8.1.1194` (official builds). Deterministic: 20/20 runs. **How to reproduce** No tables needed: ```sql SELECT arrayExists(x -> x IN (SELECT 2), [2]); -- 1 (correct) SELECT * FROM (SELECT arrayExists(x -> x IN (SELECT 2), [2])); -- 0 (wrong: the same expression, only wrapped in a derived table) ``` More shapes of the same mechanism: ```sql SELECT arrayFilter(x -> x IN (SELECT 2), [1, 2, 3]); -- [2] (correct) SELECT * FROM (SELECT arrayFilter(x -> x IN (SELECT 2), [1, 2, 3])); -- [] (wrong) SELECT arrayMap(x -> x IN (SELECT '2'), [2, 3]); -- [1,0] (correct) SELECT * FROM (SELECT arrayMap(x -> x IN (SELECT '2'), [2, 3])); -- [0,0] (wrong) -- WHERE context loses rows: SELECT count() FROM (SELECT 1 AS k) WHERE arrayExists(x -> x IN (SELECT 1), [k]); -- 0 (wrong: expected 1) WITH tm1 AS (SELECT arrayExists(x -> x IN (SELECT 2), [2])) SELECT * FROM tm1; -- 0 (wrong) ``` All at default settings; `query_plan_enable_optimizations = 0` does not cure it, so it looks like the set for the lambda-captured `IN` is not built/bound when the expression is resolved inside a subquery scope, rather than a plan-optimization issue. Possibly related observation: with `enable_analyzer = 0` the derived-table form fails outright with an exception `Code: 47` `UNKNOWN_IDENTIFIER`, where the required column is spelled `... in(x, _subquery1) ...` but the available column is `... in(x, _subquery2) ...` — the same set-identity confusion visible in the old analyzer. **Expected behavior** Wrapping a SELECT in a derived table (or moving the expression into `WHERE`) must not change the value of `IN (SELECT ...)` inside a lambda: all the wrapped forms above should return the top-level results (`1`, `[2]`, `[1,0]`, `1`). Found by an automatic optimizer-testing framework (differential testing of optimizer settings, query plans, and equivalent rewrites).",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114581",
          "createdAt": "2026-08-13T03:36:46Z",
          "updatedAt": "2026-08-13T15:09:25Z",
          "timestamp": "2026-08-13T15:09:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "potential bug"
          ],
          "author": "zlareb1",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9520254ab997850063f8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114654",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114654",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #114188 to 26.6: Ignore redundant parentheses in stored table definitions",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114654",
          "createdAt": "2026-08-13T15:09:14Z",
          "updatedAt": "2026-08-13T15:09:22Z",
          "timestamp": "2026-08-13T15:09:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:e8f2cc40451160d0bae4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114653",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114653",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #114188 to 26.5: Ignore redundant parentheses in stored table definitions",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114653",
          "createdAt": "2026-08-13T15:08:36Z",
          "updatedAt": "2026-08-13T15:08:44Z",
          "timestamp": "2026-08-13T15:08:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:1781c82317b73e96bdb9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114652",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114652",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #114188 to 26.3: Ignore redundant parentheses in stored table definitions",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114652",
          "createdAt": "2026-08-13T15:07:56Z",
          "updatedAt": "2026-08-13T15:08:04Z",
          "timestamp": "2026-08-13T15:08:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:cce1ee8039176576388f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113833",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113833",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add JOIN observability columns to system.query_log",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111352 Related: https://github.com/ClickHouse/ClickHouse/issues/111748 ## Motivation `system.query_log` says almost nothing about what the JOINs in a query actually did. Answering questions like \"which queries ran a `CROSS` join we didn't expect\", \"which joins fell back to `grace_hash` and spilled to disk\", or \"which algorithm was really chosen when `join_algorithm = 'auto'`\" currently requires re-running `EXPLAIN PIPELINE` (impossible post-mortem) or scraping `ProfileEvents` query by query. ## Changes Four columns are added to `system.query_log`, with the matching fields in `QueryLogElement`: - `used_number_of_joins` (`UInt64`) — the number of physical joins executed by the query. It is collected from the query pipeline, so it reflects the joins that really ran after all optimizations, not the number of `JOIN` clauses in the query text. - `used_join_algorithms` (`Array(LowCardinality(String))`) — the algorithms that were actually used, e.g. `hash`, `parallel_hash`, `grace_hash`, `direct`, `full_sorting_merge`, `partial_merge`. The `join_algorithm` setting only lists the allowed algorithms; the choice among them happens at runtime, and an algorithm can even be replaced mid-execution (a `hash` join switching to `grace_hash` under memory pressure). - `used_join_kinds` (`Array(LowCardinality(String))`) — Kind of the joins present in the query. - `used_join_strictness` (`Array(LowCardinality(String))`) — Strictness of the joins present in the query. - `join_spilled_to_disk` (`UInt8`) — whether any of the joins spilled to disk. This PR currently adds the schema only. The fields are declared but nothing populates them yet, so the columns read as `0` and `[]`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added four columns to `system.query_log` describing the JOINs a query executed: `used_number_of_joins` (the number of physical joins in the executed pipeline), `used_join_algorithms` (the algorithms actually used at runtime, which can differ from the `join_algorithm` setting), `used_join_kinds` (`INNER`, `LEFT`, `CROSS`, `ASOF` and so on), and `join_spilled_to_disk` (whether any join wrote temporary data to disk). This makes it possible to find problematic JOIN patterns across a fleet without reproducing each query.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113833",
          "createdAt": "2026-08-07T14:11:36Z",
          "updatedAt": "2026-08-13T15:07:49Z",
          "timestamp": "2026-08-13T15:07:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "Manerone",
          "state": "open",
          "assignees": [
            "Fgrtue"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:93f46c676b89ac1a2e38",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114651",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114651",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cherry pick #114188 to 25.8: Ignore redundant parentheses in stored table definitions",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/114188 ## Do not merge this PR manually This pull-request is a first step of an automated backporting. It contains changes similar to calling `git cherry-pick` locally. If you intend to continue backporting the changes, then resolve all conflicts if any. Otherwise, if you do not want to backport them, then just close this pull-request. The check results does not matter at this step - you can safely ignore them. ### Troubleshooting #### If the conflicts were resolved in a wrong way If this cherry-pick PR is completely screwed by a wrong conflicts resolution, and you want to recreate it: - delete the `pr-cherrypick` label from the PR - delete this branch from the repository You also need to check the **Original pull-request** for `pr-backports-created` label, and delete if it's presented there ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31711952805/job/94486903456)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114651",
          "createdAt": "2026-08-13T15:07:10Z",
          "updatedAt": "2026-08-13T15:07:19Z",
          "timestamp": "2026-08-13T15:07:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "do not test",
            "pr-bugfix",
            "pr-cherrypick"
          ],
          "author": "robot-ch-test-poll",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:626b7a389e9568c4d7c0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114650",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114650",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the projection's intermediates out of the WITH FILL header",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114404 Caused by: https://github.com/ClickHouse/ClickHouse/pull/107700 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NUMBER_OF_COLUMNS_DOESNT_MATCH` for a distributed query combining an `ALIAS` column whose body is an expression with `ORDER BY ... WITH FILL ... INTERPOLATE`. ### Description `analyzeSort` passed every column *available* to the `Before INTERPOLATE` step through as an output. A step's available columns are all the nodes of the previous step's `ActionsDAG` — the intermediates of a computed expression included, since a later step is allowed to *reference* any of them. Turning \"referenceable\" into \"must be in the stream\" pinned those intermediates into the header of the `Filling` step. That made the header depend on something that should not matter. Two query trees that differ only by whether an `ALIAS` column was inlined into its defining expression produced different headers, because `v * 2` contributes its `2` and the un-inlined `a_v` does not: ``` un-inlined inlined __table1.k __table1.k __table1.a_v multiply(__table1.v, 2_UInt8) __table1.v __table1.v 2_UInt8 <- extra materialize(__table1.k) materialize(__table1.k) a_v a_v ``` A distributed query is planned from the un-inlined tree on the initiator and executed from the inlined one on the shard, so the two headers meet. `buildShardCollapseFanOut` only handles a *smaller* shard header, and the positional `makeConvertingActions` then throws `NUMBER_OF_COLUMNS_DOESNT_MATCH`. The fix holds back the projection's own intermediates and passes everything else through as before, including the sort columns materialized just above, which is what `Filling` fills by. They all remain *inputs* either way, so the step still asks the previous one to produce them. One exception has to be carved out: a column that an `INTERPOLATE` expression names. Those expressions become actions later, in `Planner`, against a dag built from this step's *header*, so whatever they reference must survive as a column even when the query does not select it — `INTERPOLATE (inter AS inter2 + inter)` where `inter2` is not in the `SELECT` list. Those names are collected by visiting the expressions into a throwaway dag. ### Scope Everything here is inside `if (query_node.hasInterpolate())`, so only queries with `INTERPOLATE` change. The shape was already broken over a `Distributed` table before #107700, since that path has always inlined `ALIAS` columns; #107700 extended the inlining to the parallel-replicas paths and so exposed it there too. Both are fixed. It also stops the reconciliation failure from masking a query error: `INTERPOLATE (k AS k)` on an `ORDER BY` column reports `INVALID_WITH_FILL_EXPRESSION` again instead of `NUMBER_OF_COLUMNS_DOESNT_MATCH`. ### Validation Built and run against a three-replica localhost cluster and a two-shard `Distributed` table, compared with the CI binaries of the commit before #107700 (`75b17ad`) and of its merge (`dd01d270e`): | | before #107700 | after #107700 | this | |---|---|---|---| | repro, parallel replicas shipping a plan | pass | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | repro, parallel replicas shipping SQL | pass | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | repro over a 2-shard `Distributed` table | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | `INTERPOLATE (k AS k)` under parallel replicas | `INVALID_WITH_FILL_EXPRESSION` | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | `INVALID_WITH_FILL_EXPRESSION` | The new test `04891_with_fill_interpolate_alias_column_header` reports four exceptions on the commit before #107700 and seven on the merge commit, and passes here. Every self-contained stateless test mentioning `WITH FILL` or `INTERPOLATE`, plus the `ALIAS`-shipping tests from #107700, passes: 94 of 94. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-dd669bf9-3d72-4334-961f-0871feaf9f98?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-dd669bf9-3d72-4334-961f-0871feaf9f98&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114650",
          "createdAt": "2026-08-13T14:39:41Z",
          "updatedAt": "2026-08-13T15:04:51Z",
          "timestamp": "2026-08-13T15:04:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:765977979fd57e402d85",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110477",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110477",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "EXPLAIN SYNTAX: return the pretty-printed query as a single multi-line record",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/80410 Related: https://github.com/ClickHouse/ClickHouse/pull/107925 --> Closes: #80410 Related: #107925 (closed, superseded by this PR) ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `EXPLAIN SYNTAX` now returns the pretty-printed (multi-line) query as a single result record instead of one record per line. The `oneline` option (`EXPLAIN SYNTAX oneline = 1`) still collapses the output to a single physical line. Queries that consumed the previous per-line output should treat the result as one row whose value contains embedded newlines. ### Description `EXPLAIN SYNTAX <query>` formatted the reformatted query and split it into one result row per physical line. This PR emits the whole formatted, copy-pasteable query as a single record with newlines preserved (issue #80410). Only the `AnalyzedSyntax` code path is changed; `PLAN`, `PIPELINE`, `AST` and the `oneline` option are unchanged. The previous attempt (#107925) was closed because it flipped the `oneline` default to `1`, collapsing the query to a single physical line rather than keeping the multi-line pretty form. This PR keeps `oneline = false` by default and only changes the result shape from N rows to one multi-line record. Reference files for tests consuming EXPLAIN SYNTAX output (including `.oldanalyzer.reference` variants) were regenerated. A regression test (`04545_explain_syntax_single_record`) asserts the single-record multi-line default and the `oneline = 1` collapse.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110477",
          "createdAt": "2026-07-15T00:34:56Z",
          "updatedAt": "2026-08-13T15:04:44Z",
          "timestamp": "2026-08-13T15:04:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:619328ffbfb102bb7587",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114323",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114323",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: require canonical internal links",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114230 This is a one-off cleanup of repository-authored documentation links that use legacy redirect aliases. It updates the current English documentation and source-embedded reference documentation to use routes relative to the docs root, so the automated translation PR can parse and localize them without producing missing locale routes. This PR intentionally adds no permanent CI checks or ongoing enforcement. Its scope is limited to the current link corrections needed to get the automated translation PR parsing successfully. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114230&sha=e599855a281a4dd841be20794036908a58acbf63&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114323&sha=86d2d7e194d7fbf75cc84616a8fa0da1ed84802e&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114323",
          "createdAt": "2026-08-11T13:26:47Z",
          "updatedAt": "2026-08-13T15:04:37Z",
          "timestamp": "2026-08-13T15:04:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci",
            "pr-autogenerated-docs"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e47277889545bc2435b0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111794",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111794",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add unordered stream modifier",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM UNORDERED modifier: skip the per-snapshot commit-order sort depends on https://github.com/ClickHouse/ClickHouse/pull/110653 (not for functional reason, only test) cc @alesapin @Michicosun",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111794",
          "createdAt": "2026-07-24T13:31:48Z",
          "updatedAt": "2026-08-13T15:04:28Z",
          "timestamp": "2026-08-13T15:04:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "SmitaRKulkarni",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1341bf28a205aa22568b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114642",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114642",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: document filtering for `system.query_log` initial queries",
          "text": "Document the recommended filters for analyzing `system.query_log`. The page now explains that `is_initial_query = 1` selects top-level client queries, while `initial_query_id` correlates the full cascade of a distributed query across nodes. The existing basic example also filters for initial queries so child executions are not treated as separate client queries. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114642",
          "createdAt": "2026-08-13T13:10:53Z",
          "updatedAt": "2026-08-13T15:03:27Z",
          "timestamp": "2026-08-13T15:03:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:65a9cc50db942b7f3aac",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113573",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add `NOSIGN` to public `s3` documentation queries",
          "text": "Update anonymous public S3 documentation examples to pass `NOSIGN` explicitly following the ClickHouse 26.7 server credential behavior change. This prevents public reads from attempting to use server-managed credentials. The audit also found and repairs two stale public paths: the LAION guide now uses the surviving 10-million-row shard, and the S3 brace-expansion example references the four files that currently exist. Authenticated, requester-pays, write, placeholder, and Foursquare examples are intentionally outside this PR. Related: https://linear.app/clickhouse/issue/DOC-945/update-public-s3-docs-examples-to-use-nosign ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Updated public S3 documentation examples to use `NOSIGN` and repaired stale public dataset paths.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113573",
          "createdAt": "2026-08-05T20:49:36Z",
          "updatedAt": "2026-08-13T15:01:42Z",
          "timestamp": "2026-08-13T15:01:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "pr-autogenerated-docs"
          ],
          "author": "dhtclk",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:19119329280c72c531e9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113059",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113059",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Collect SQL stacktraces on the hung-check and server-died abort paths",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/42701 Related: https://github.com/ClickHouse/ClickHouse/pull/112265 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ### Description Related: #42701, #112265 On an ASan build, a stateless run aborting on the hung check records no stack of the hung server: ``` Hung check failed: server is not responding Cannot collect C stacktraces under ASan: debugger attach is disabled. ``` Two things combine. `print_c_stacktraces` declines to attach lldb on ASan builds, because ptrace disables LeakSanitizer; that refusal is correct and stays. And `print_sql_stacktraces`, needing no debugger, was unreachable: its pre-check `check_server_liveness` probes **HTTP** (`http_port`, default 8123), while it collects over native **TCP** (`args.client --port=`, default 9000). Different listeners, different ports, so this signature (alive, not answering HTTP, TCP still serving) failed the gate. The hung-check abort site did not call it at all. This drops the mismatched pre-check and lets the collector be its own liveness test: it is already bounded (`timeout=30`) and reports failure as one trimmed line, so a dead socket costs at most 30 s and cannot re-emit the `Code: 210` tracebacks that motivated the pre-check. The dump is added to the three abort sites that had only the C path: hung check, server died, and the startup check. The stateless job keeps attaching the dump to its result and additionally clears any left by a previous job in the same workspace, so an aborted run cannot upload a stale dump as its own. #114143 has since added the same attachment upstream; this replaces it with the equivalent helper rather than attaching twice. It fixes no hang and does not restore C++ stacks on ASan. A server alive but not answering HTTP now yields the full `system.stack_trace` view with per-thread `query_id`, identifying the wedged query; one dead on both transports records \"tried, got nothing\" instead of silence. Validated against a live server: with HTTP dead and TCP live, master skips and writes nothing, while this branch writes a `sql_stacktraces.log` carrying `thread_name` and `query_id`. Green runs are unaffected. [Prompting report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=1ce152efafc5b00bf31eb7a0a64757fecbbc6e4a&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20distributed%20plan%2C%20parallel%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113059",
          "createdAt": "2026-08-03T06:03:46Z",
          "updatedAt": "2026-08-13T14:59:40Z",
          "timestamp": "2026-08-13T14:59:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:87ba9dc0024eaa435c0f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114087",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114087",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check column structure against the declared type in `collectOffsetsColumns`",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113925 Related: https://github.com/ClickHouse/ClickHouse/pull/113225 Related: https://github.com/ClickHouse/ClickHouse/issues/113891 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a column's declared type and its data diverge during a `MergeTree` read (mixed type provenance, e.g. after an `ALTER TABLE ... MODIFY COLUMN` whose mutation has not finished), the server now reports a clear exception naming the column and both structures, instead of an unchecked cast: debug and sanitizer builds previously aborted with a bare `Bad cast from type A to B` naming no column, and release builds walked mismatched memory silently. ### Description The type-directed `enumerateStreams` walk in `collectOffsetsColumns` (`src/Interpreters/inplaceBlockConversions.cpp`) pairs each available column's declared type with its data. When an entry's type was resolved from the table's metadata while the column was read from a data part with an older type - the mixed type provenance behind #113925 - the walk `assert_cast`s the column to the wrong class at the first diverging wrapper. In debug/sanitizer builds that is an abort whose message names two column classes and no column (the crash-classifier issue #113891, STID 4256-3fc1, took a cross-file hunt to attribute); in release builds `assert_cast` does not check at all, so the walk misbehaves on memory-unsafe reinterpretation. The new `columnMatchesTypeStructure` check runs right before the walk, only on the missing-columns path where the walk already happens. It descends exactly the levels whose serializations pair the declared type's structure with the column's: - `Array`, `Nullable`, `Tuple`, `Map` - the wrappers whose serializations `assert_cast` the column; - `Variant` - `SerializationVariant::enumerateStreams` walks every alternative declared by the type and pairs it with the column's variant of the same global discriminator, so the alternative list has to match in length and element-wise; - typed paths of `Object` - `SerializationObject::enumerateStreams` walks every typed path declared by the type and looks it up in the column, so the typed paths have to match by name and structure. Dynamic paths and shared data are taken from the column itself and need no check; - `ColumnReplicated` is unwrapped. `Dynamic` is checked by class only, because its `enumerateStreams` takes both the type and the column of its variant from the column itself (`column_dynamic->getVariantInfo().variant_type`), so the two cannot diverge. Every leaf the check does not know is accepted - leaf divergence, such as a part storing `UInt32` for a column widened to `UInt64`, is legitimate. Because the checked structure mirrors exactly what the `enumerateStreams` implementations themselves assert, the check cannot fire on any pairing that debug CI does not already abort on (or, for `Object`, fail with a raw typed-path lookup) - it only converts that failure into a diagnosable exception and closes the release-build hole. With the #113925 reproducer on a `RelWithDebInfo` build of master before that fix, the witness query fails with: ``` Code: 49. DB::Exception: Column `arr.n` is listed with type Array(Nullable(String)) among available columns, but its data has incompatible structure Array(size = 1, UInt64(size = 1), String(size = 2)). It is likely that a type resolved from the table's metadata was combined with a column read from a data part with an older type: (while reading from part .../all_1_1_0/ ...) ``` instead of silent unchecked-cast behavior. #113925 (which fixes the type selection) has since merged and is included here, so this check now guards the invariant against future regressions of the same family. **Validation.** Witness reproduced as above on a local `RelWithDebInfo` build. False-positive sweep: all 381 stateless tests matching `nested`/`subcolumn` run against that binary - 294 passed, 87 failed for documented bare-server environmental reasons (no Keeper, no clusters, no `protoc`, no `/var/lib/clickhouse`), and the check's message appears zero times in the whole run. After the `Variant`/`Object` extension, all 826 stateless `.sql` tests matching `variant`/`dynamic`/`json`/`nested`/`subcolumn`/`object` were run again - the check's message and `Bad cast` both appear zero times - plus a targeted check that reads old parts through newly added `Variant`, `JSON`, `Dynamic`, `Tuple`, `Map` and `Nested` columns and through unfinished `ALTER TABLE ... MODIFY COLUMN` mutations of `Variant` and `JSON`. The `Memory`-engine caller of `fillMissingColumns` always casts columns to the requested types first (`tryGetColumnFromBlock`), so it cannot trip the check either. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114087",
          "createdAt": "2026-08-10T00:15:02Z",
          "updatedAt": "2026-08-13T14:59:16Z",
          "timestamp": "2026-08-13T14:59:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5d979c3d015813c9dc6e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:99495",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:99495",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `GradualResizeProcessor` to limit effective parallelism for GROUP BY on small data volumes",
          "text": "When ClickHouse processes GROUP BY, it often overestimates the number of threads needed. With `max_threads = 64` but only a few thousand rows, all 64 `AggregatingTransform` instances get data, produce 64 partial hash tables, and the merge phase has to combine all of them — most nearly empty. This wastes time on merging overhead, which is especially noticeable for heavy aggregate states such as `uniq`, `uniqExact`, `groupArray`, etc. The new `GradualResizeProcessor` starts by pushing data to a single output port (or one port per split group when `min_outstreams_per_resize_after_split` applies), and activates all aggregation streams at once as soon as the configured row or byte threshold is crossed. For small datasets, only one aggregating thread receives data (or one per split group); for large datasets, all threads are used as before. New settings: - `min_rows_per_stream_for_gradual_resize` (default: `1000`) - `min_bytes_per_stream_for_gradual_resize` (default: `0`) When either threshold is non-zero, the pre-aggregation `StrictResize` is replaced with `GradualResize` in the pipeline. The optimization is enabled by default; set both `min_rows_per_stream_for_gradual_resize = 0` and `min_bytes_per_stream_for_gradual_resize = 0` to opt out. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve performance of GROUP BY on small data volumes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/99495",
          "createdAt": "2026-03-14T07:35:33Z",
          "updatedAt": "2026-08-13T14:58:19Z",
          "timestamp": "2026-08-13T14:58:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 30
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2cb55a873efa8f80f9ea",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112327",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112327",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use-after-free on a sparse join key in a direct dictionary join",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109225 --> Related: https://github.com/ClickHouse/ClickHouse/pull/109225 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-after-free when a `JOIN` with `join_algorithm = 'direct'` onto a dictionary was given a join key that is stored with sparse serialization. `getColumnVectorData` returned a reference to a temporary column, so the dictionary lookup read freed memory: release builds could return wrong results and debug or sanitizer builds aborted. ### Description `getColumnVectorData` (`src/Dictionaries/DictionaryHelpers.h`) materializes its key column into a function-local `ColumnPtr` and then returns a `PaddedPODArray` reference **into that local**. It copied the data into the caller's `backup_storage` only when the input was `Const`. For a dense column that was still safe, because every conversion is a no-op returning `getPtr()` and the caller's own `ColumnPtr` keeps the buffer alive. It is not safe for a column that has to be materialized: `ColumnSparse::convertToFullColumnIfSparse` and `ColumnReplicated::convertToFullColumnIfReplicated` each allocate a **new** column that nothing else owns, so the returned reference dangles as soon as the function returns. The fix takes the copy whenever a conversion actually produced a different column (`full_column.get() != column.get()`). This is pointer identity rather than a type test, so it covers `Const`, `Sparse` and `ColumnReplicated` with one predicate, and it is fail-closed for any representation added later. The previous `Const` check is strictly subsumed: `ColumnConst::convertToFullColumn` returns either the inner column or a `replicate` result, never the `ColumnConst` itself, so `Const` behaviour is unchanged. The dense path is unaffected and adds no copy, which I verified by instrumenting both live call sites: a dense key reports zero copies, a sparse key reports one at each site. A sparse key is the case that is a use-after-free today, and it already paid for a full materialization inside `removeSpecialRepresentations`, so the extra `memcpy` of that same buffer is negligible. Reaching the bug requires a path that hands a non-materialized key to the dictionary. `dictGet`, `dictHas` and the hierarchy functions cannot: `IFunction::useDefaultImplementationForSparseColumns()` and `...ForReplicatedColumns()` both default to true and no dictionary function overrides them, so a dense, caller-owned column arrives. `IDictionary::getByKeys` (the direct join) is the reaching path, because its own `removeSpecialRepresentations` call sits inside a Nullable-only branch and a non-Nullable sparse key passes through untouched. That is also why `04627_direct_join_dictionary_nullable_key`, which does exercise a sparse key, never caught this: its key is Nullable, so it gets materialized. All 14 `getColumnVectorData` call sites are fixed by this single change. Two of them are reachable today, both in `FlatDictionary` and both on the same `getByKeys` call: `hasKeys` (the site in the reports below) and `getColumn`, reached through `getColumns`. The remaining 12 are hierarchy-only and reachable solely through the pre-converting function path. `Hashed` and `HashedArray` never reach the helper on the `getByKeys` path at all, because `DictionaryKeysExtractor` holds its converted column by value and therefore owns it. I also swept every other `convertToFullColumnIf*` / `recursiveRemove*` / `removeSpecialRepresentations` call site under `src/` for the same \"derived data escapes the owning local\" shape and found no second instance, so no sibling fix is needed. The bug dates to 2021 (`b5b624f3d7e9bf`, which introduced the conversion here) and was widened in 2025 by `2b6cb36d1dc936`, which added `ColumnReplicated` as a second carrier. The same code is present on 26.7, 26.6, 26.5, 26.4 and 26.3. Found while triaging CI on #109225 and reproduced on unmodified master. It is latent in CI only because no existing test combined a sparse-serialized left key with a direct join over a `FLAT()` dictionary; CI randomizes `ratio_of_defaults_for_sparse_serialization`, so any test that does hit this combination fails roughly 40% of the time. Reports on `31b4a2961ef4c5183f7d15dda7f755541a77a98c`: - `AddressSanitizer: heap-use-after-free`, allocated by `ColumnSparse::convertToFullColumnIfSparse`, freed at the end of `getColumnVectorData`, read by `FlatDictionary::hasKeys`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20flaky%20check%29 - `Logical error: '(n >= (static_cast<ssize_t>(pad_left_) ? -1 : 0)) && (n <= static_cast<ssize_t>(this->size()))'` from the same read: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20flaky%20check%29 The new test `04652_direct_join_dictionary_sparse_key` covers both live call sites, the aggregation shape from the report, and a key carried through `ARRAY JOIN` over a sparse base column, which is a third shape where the key has to be materialized. It pins its results against a dense table and against `join_algorithm = 'hash'` instead of hand-written constants, and asserts both that the key really is sparse and that `DirectKeyValueJoin` is still chosen, so it cannot pass vacuously. On master it aborts; with the fix it passes 50/50 with and without randomized settings. Reverting only the new predicate makes it abort again. A follow-up cleanup worth doing separately: the helper carries a `/// TODO: Remove` and would be better returning the owning `ColumnPtr` alongside the data, which removes the need for `backup_storage` entirely. That touches all 14 call sites and four dictionary classes, so it does not belong here.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112327",
          "createdAt": "2026-07-28T17:38:47Z",
          "updatedAt": "2026-08-13T14:58:09Z",
          "timestamp": "2026-08-13T14:58:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:df739c9c98d1748dcd95",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114648",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114648",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: move ODBC driver documentation to clickhouse-odbc",
          "text": "## Summary - move the ODBC driver guide into a group-backed folder under `concepts/features/interfaces` - remove the duplicate ODBC connector page and its navigation entry - redirect the retired connector URL and its legacy alias directly to the ODBC table-engine reference - update remaining English documentation links to bypass the redirect ## Why The retired connector page duplicated the ODBC table-engine reference. The actual driver guide belongs with `ClickHouse/clickhouse-odbc`, where it can evolve alongside the driver and later be split into focused pages. ## Coordination Paired draft PR: https://github.com/ClickHouse/clickhouse-odbc/pull/581 The paired PR vendors the driver guide and adds verification and one-way documentation sync workflows. This PR prepares the corresponding target folder and navigation reference in `ClickHouse/ClickHouse`. ## Validation - scoped Mintlify validation passed for `concepts/features/interfaces/odbc` - snippet and component import checks passed - internal link validation passed with zero errors - redirect validation passed with zero errors and no redirect chains ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Move the ODBC driver guide to the driver repository and retire the duplicate connector page.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114648",
          "createdAt": "2026-08-13T14:23:22Z",
          "updatedAt": "2026-08-13T14:56:55Z",
          "timestamp": "2026-08-13T14:56:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:56ffc3f496d0158a9ce8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114620",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114620",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix serialization of Map-valued settings in access entities",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114591 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a bug where an access entity carrying a `Map`-valued setting, such as a settings profile with `http_response_headers` or `additional_table_filters`, was stored in a form that ClickHouse could not read back, so the entity became permanently unloadable after a restart. Closes #114591. ### Description Closes #114591. `CREATE SETTINGS PROFILE p SETTINGS http_response_headers = '{...}'` succeeds, but the entity is stored as `http_response_headers = [('k', 'v')]`, which nothing can parse back. No error appears at `CREATE` time, so the entity is **permanently unloadable** after a restart: ``` stored: ATTACH SETTINGS PROFILE `p` SETTINGS http_response_headers = [('a', 'b')] CONST; after restart: Code: 62. Syntax error: failed at position 65 ([) ... Could not parse <path>/access/<uuid>.sql ``` The list rebuild drops it silently, reading it back throws, so `SELECT` from `system.settings_profile_elements` fails while it is present, as does `RESTORE` of a backup holding it. Root cause: the value is cast to the setting's native type, so a Map setting holds a `Map` Field, which `FieldVisitorToString` renders as an array of tuples; but `ParserSettingsProfileElement` reads values with a scalar-only `ParserLiteral` and cannot open a `[`. That spelling is rejected everywhere, `SET http_response_headers = [('a','b')]` included, so the write side is wrong. Fix: when a profile element's value, MIN or MAX is a `Map` and the setting is builtin, emit the setting's canonical text as a quoted string. Write side only, no grammar change. Custom settings are excluded because `castValueUtil` returns their value unchanged, so a string would come back a `String` rather than a `Map`. Covers `CREATE USER`/`ROLE`, `ALTER ... SETTINGS` and MIN/MAX, and transitively BACKUP/RESTORE and both storages. `SHOW CREATE` now prints a quoted string rather than `[('k', 'v')]`, intended since the new form is copy-pasteable. Entities already stored in the broken form are not repaired, as they were never parseable; recreate them. Downgrade is safe: a pre-fix binary reads the new form correctly, so no versioning is needed. New test `04902_access_entity_map_setting_round_trip`: 11 of its 15 arms fail on pristine master and pass here, covering empty, multi-key and hostile maps plus a two-process on-disk reload.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114620",
          "createdAt": "2026-08-13T11:13:29Z",
          "updatedAt": "2026-08-13T14:55:52Z",
          "timestamp": "2026-08-13T14:55:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6356b3e729a9c7c74110",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114457",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114457",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix header column order after the no-rescoring vector search rewrite",
          "text": "<!--- A technical description of your changes with a motivation --> The second-pass vector search optimization (`vector_search_with_rescoring = 0`) removed the distance function from the `ExpressionStep` outputs and re-appended the rewritten `_distance` alias at the end, changing the header column order. Steps created above the `Sorting` before that rewrite runs — the local top-N `Limit` and the exchange steps of a distributed plan — kept the original column order, and `makeDistributedPlan` failed to rebuild the plan fragments with a logical error: `Cannot add step Limit to QueryPlan because it has incompatible header with root step Sorting`. Non-distributed plans never re-validate step headers after optimization, which is why this only surfaced with `make_distributed_plan = 1`. The fix reinserts the rewritten output node at the original position of the distance column, so the header column order is preserved and all previously created parent steps stay consistent. Found by AST fuzzer on [#42701](https://github.com/ClickHouse/ClickHouse/pull/42701) ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=019244e0d46b4d9029fc261a549e1c5cd79b1bbb&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)); the same signature also hit the unrelated #98789 on 2026-08-08. The regression test asserts only the row count: the distributed plan currently returns an incorrect top-N for vector search queries regardless of the rescoring mode — a separate, pre-existing bug. Related: https://github.com/ClickHouse/ClickHouse/issues/114456 Related: https://github.com/ClickHouse/ClickHouse/pull/42701 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a logical error `Cannot add step Limit to QueryPlan because it has incompatible header with root step Sorting` when a vector search query with `vector_search_with_rescoring = 0` was executed with the experimental `make_distributed_plan` setting.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114457",
          "createdAt": "2026-08-12T09:43:36Z",
          "updatedAt": "2026-08-13T14:55:29Z",
          "timestamp": "2026-08-13T14:55:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bce55bf26e29399470ed",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113181",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113181",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: Split large function reference pages",
          "text": "Split the four regular-function families containing more than 100 functions into individual reference pages with compact searchable overview indexes. This draft tests a different information architecture for the largest Mintlify reference pages, reducing the amount of interactive code-block content rendered on a single page while preserving generated documentation and legacy fragment navigation. The generator now emits function pages, family navigation, manifests, and shared searchable index components for array, date and time, other, and type-conversion functions. It includes 546 individual function pages and a focused generator regression test. The result was validated with the regular-function and session-settings generator tests, generated-route and anchor coverage checks, `git diff --check`, and local Mintlify preview inspection. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): N/A",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113181",
          "createdAt": "2026-08-03T20:17:11Z",
          "updatedAt": "2026-08-13T14:52:47Z",
          "timestamp": "2026-08-13T14:52:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "dhtclk",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:1c6ac92b43d1a720a78d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110626",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110626",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain data/expected structure mismatches in INSERT parse errors",
          "text": "Improves the error message shown when parsing the data of an `INSERT` fails. Previously, inserting data whose structure does not match the destination produced a confusing low-level parse error (for example `Cannot parse input: expected '\\t' before ...`) with no hint about the real cause. In the linked issue, the destination's inferred schema had integer columns while the inserted `TSV` data had string columns, and the reported error pointed at a tab that was actually present. Now, on a parse failure, ClickHouse infers the structure of the data being inserted (only when the format has a data schema reader, and solely for diagnostics) and, if it does not correspond to the expected structure, appends an explanation listing both the inferred and the expected structure. For example: ``` Code: 27. DB::Exception: Cannot parse input: expected '\\t' before: 'page_view... ... The structure of the data being inserted does not match the structure expected by the query, which is likely the cause of the parsing error. Inferred structure of the input data (in format `TSV`): c1 Nullable(Int64) c2 Nullable(String) c3 Nullable(String) Expected structure: c1 Int64 c2 Int64 c3 Int64 ``` The check is wired through a lazy provider on `IInputFormat` that runs only on a genuine parse error, so there is no cost on the happy path. The comparison ignores the artificial `Nullable` wrapper that schema inference adds by default, so inserting valid data into non-nullable columns is not falsely flagged. It covers the synchronous local/server path, client-side parsing, and the asynchronous insert queue (the default path now that `async_insert` is enabled by default). Closes: https://github.com/ClickHouse/ClickHouse/issues/110622 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When parsing of the data being inserted by an `INSERT` fails, the error message now explains a likely structure mismatch by comparing the structure inferred from the data with the structure expected by the query.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110626",
          "createdAt": "2026-07-15T23:41:14Z",
          "updatedAt": "2026-08-13T14:52:22Z",
          "timestamp": "2026-08-13T14:52:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 25
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0942c179bde25aabc3b9",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:91062",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:91062",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add an experimental regex-free glob parser",
          "text": "Introduces `GlobAST`, a regex-free glob parser, and wires it into the file and object-storage listing paths. It is **off by default** (the experimental `use_glob_ast_parser` setting), so `master` behavior is unchanged — the goal is to land the parser and all the wiring now so that enabling it later is a one-setting switch. The legacy engine builds a regex (`makeRegexpPatternFromGlobs` + `re2::RE2`), which has noticeable downsides: - lists whole prefixes for enum globs (#73333) - blows up to megabyte regexes on numeric ranges (#43456) - has no formal grammar (#80950) - diverges from POSIX shell semantics in places (e.g. stripping single-element brace groups like `{a}`/`{-}`). `GlobAST` parses a pattern once into typed expressions (constant, `?`/`*`/`**`, `{M..N}` range, `{a,b,…}` enum) and matches/expands directly — ranges by numeric bounds checks, enum-only globs by expanding to concrete keys. A `GlobMatcher` front-end selects the new or legacy backend per the setting, so call sites are unchanged; the grammar is documented in `parseGlobs.h`. Tested by unit tests, a differential fuzzer against the legacy matcher (`GlobASTLegacyMatchFuzz`), and a stateless parity test. Related: https://github.com/ClickHouse/ClickHouse/issues/73333 Related: https://github.com/ClickHouse/ClickHouse/issues/43456 Related: https://github.com/ClickHouse/ClickHouse/issues/80950 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental setting `use_glob_ast_parser` (default `false`) that switches glob matching for the `file`/`s3`/object-storage listing paths to a new regex-free parser (`GlobAST`), avoiding regex blow-up on large numeric ranges and listing fewer keys for brace-enumeration globs.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/91062",
          "createdAt": "2025-11-27T22:07:46Z",
          "updatedAt": "2026-08-13T14:52:19Z",
          "timestamp": "2026-08-13T14:52:19Z",
          "metrics": {
            "reactions": 1,
            "comments": 14
          },
          "labels": [
            "pr-experimental",
            "hold"
          ],
          "author": "thevar1able",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:51a5dfd084f411f404d5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111287",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111287",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix double free when finalizing -State aggregates under looping combinators",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server crash (double free) that could happen when finalizing an aggregate function with the `-State` combinator nested under a looping combinator (`-Resample`, `-ForEach`, `-Map`), for example `groupArrayStateResample`, if a memory limit was reached during finalization. ### Description Reported on https://github.com/ClickHouse/ClickHouse/pull/110975 (unrelated to that PR). Found by the Stress test (amd_debug): a segfault in `Aggregator::prepareChunkAndFillWithoutKey`, reached from `ConvertingAggregatedToChunksTransform::initialize`. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110975&sha=838d0b61235b06939c8acd923ebd396f51cf5b10&name_0=PR&name_1=Stress%20test%20%28amd_debug%29 Root cause: the `-State` combinator transfers its result by aliasing the raw aggregate state pointer into a `ColumnAggregateFunction` (`AggregateFunctionState::insertResultInto` -> `getData().push_back(place)`); ownership passes to the column. `Aggregator::insertAggregatesIntoColumns` relies on this transfer being atomic per place: on an exception it destroys the whole place exactly once. A looping combinator nested over `-State` aliases many sub-states one at a time; the `push_back` into the column's pointer array can reallocate and, being memory-tracked, throw `MEMORY_LIMIT_EXCEEDED` mid-loop. The already-transferred sub-states are then freed once by the aggregator's full `destroy()` and again by `~ColumnAggregateFunction`, i.e. a double free. Reproducer (crashes without the fix, returns a memory-limit error with it): ```sql SELECT arrayMap(x -> finalizeAggregation(x), state) FROM (SELECT groupArrayStateResample(0, 1048576, 1)(number, number % 20) AS state FROM numbers(100000)) SETTINGS max_memory_usage = 150000000, max_rows_to_read = 0; ``` Fix: reserve the destination columns before the transfer loop so the aliasing `push_back`s cannot reallocate (and therefore cannot throw) once a transfer has started. `ColumnAggregateFunction` used the no-op `IColumn::reserve`, so a real `reserve()`/`capacity()` over its state-pointer array is added. For `-Map`, the (possibly variable-width) key inserts are moved into their own loop before the value transfer, keeping the throwing work out of the aliasing loop. Reserving happens before any aliasing, so a throw there is harmless. The transfer loop is now non-throwing at the point of aliasing, restoring the atomic-per-place contract; results are unchanged. The fix covers all three looping transfer combinators (`-Resample`, `-ForEach`, `-Map`), which share the aliasing path; non-looping combinators delegate a single call and are already atomic. The added stateless test reproduces the crash deterministically via `-Resample` (empty buckets keep memory low until the finalization transfer, so a memory limit reliably lands the throw mid-transfer). `-ForEach` and `-Map` build their sub-states eagerly during aggregation, so they are not deterministically reproducible under a memory limit, but are fixed as the same class via the shared transfer path.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111287",
          "createdAt": "2026-07-21T20:34:03Z",
          "updatedAt": "2026-08-13T14:52:03Z",
          "timestamp": "2026-08-13T14:52:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3618a2d13180f3e900d4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109891",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109891",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reintroduce borrowed threadgroup async uaf fix",
          "text": "Reintroduce #108988 Related: https://github.com/ClickHouse/ClickHouse/pull/107030 Related: https://github.com/ClickHouse/ClickHouse/pull/108577 CI: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105890&sha=3dc0e76362eb18e27f1fffcd3f61ca9f13725fe8&name_0=PR&name_1=Stateless%20tests%20%28amd_tsan%2C%20s3%20storage%2C%20sequential%2C%201%2F2%29 Fixes a use-after-free risk in asynchronous work scheduled while a borrowed `ThreadGroup` is current. Borrowed `ThreadGroup` objects used by materialized view and async insert flush paths point their `performance_counters` and `memory_tracker` to the parent query group, so they are valid only while that parent group is alive. Async callbacks could capture such a borrowed group and later attach it on a pool thread after the parent query group had finished. Instead of keeping the parent `ThreadGroup` alive with a `shared_ptr`, this change keeps borrowed accounting scoped. Borrowed groups are marked explicitly, async callback capture drops borrowed groups, and thread pool callback runners capture the normalized group at task enqueue time rather than when a potentially long-lived runner object is created. This preserves synchronous borrowed accounting, but async work started from a borrowed scope runs under normal thread/global accounting instead of writing into, or prolonging the lifetime of, an already finished query group. Full ASAN reports https://gist.github.com/filimonov/1ec59047c65e3a5367c5c83f6021cc27 Compared to #108988 - added one commit with code comments + fix of the test failure https://github.com/ClickHouse/ClickHouse/issues/109841 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a memory safety issue where asynchronous work scheduled from materialized view processing could keep using query-level accounting after the query had finished.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109891",
          "createdAt": "2026-07-09T12:34:20Z",
          "updatedAt": "2026-08-13T14:52:00Z",
          "timestamp": "2026-08-13T14:52:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 47
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "comp-query-execution"
          ],
          "author": "filimonov",
          "state": "open",
          "assignees": [
            "azat",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f6afda2e298a65efd800",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114247",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114247",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Clear plain LIMIT/OFFSET in the window-view backfill source query",
          "text": "Follow-up to https://github.com/ClickHouse/ClickHouse/pull/113759: the last AI review finding on that PR landed after the PR had been added to the merge queue, so the branch could no longer be updated. This addresses it. `StorageWindowView::getSourceTableSelectQuery` builds the raw-source backfill query for `CREATE WINDOW VIEW ... POPULATE`. The rows it produces are inserted into the window view, where `writeIntoWindowView` executes the mergeable view query over them, so the backfill query must deliver the raw source rows and leave all row transformations of the original query to the view query — otherwise the initialized state diverges from live behavior. The PR originally made the helper clear the leftovers of the original `SELECT` that violated this invariant one by one — `LIMIT`/`OFFSET` (and `WITH TIES`), plain `DISTINCT`, `ARRAY JOIN`, `WHERE`/`PREWHERE`, table-expression `SAMPLE`/`FINAL` — on top of the `JOIN`/`GROUP BY`/`ORDER BY`/`LIMIT BY`/`WINDOW`/`QUALIFY`/`INTERPOLATE` handling it already had. Review then found that this strip-the-leftovers approach misses wrapped sources: the same constructs inside a `FROM (SELECT ...)` subquery or a CTE definition survived the rewrite, and covering them would have required recursing the rewrite into every nested select. So the helper now builds the backfill query from scratch instead of stripping a clone of the view query. The contract makes this valid: `writeIntoWindowView` always receives raw source-table blocks (`getInputHeader` is the source table header no matter how the view query wraps or transforms the table — `PushingToWindowViewSink` is created with exactly that header), so the correct backfill query is exactly `SELECT <source columns> FROM <source table>`, plus `ORDER BY` on the timestamp column so the watermark is initialized from the earliest record. Wrapped sources, joins, and every row-shaping clause are covered by construction because the user query is no longer cloned at all. The helper shrinks by ~100 lines. Note: the divergence is currently unobservable because `CREATE WINDOW VIEW ... POPULATE` fails before writing any rows (https://github.com/ClickHouse/ClickHouse/issues/113493), so no regression test is possible yet; this keeps the rewrite invariant consistent for when `POPULATE` is fixed. Verified with `clickhouse-local` probes that `POPULATE` over wrapped-subquery/CTE/`JOIN`/`WHERE`/`FINAL` sources now analyzes the backfill query cleanly and proceeds to the pre-existing #113493 sink error. Related: https://github.com/ClickHouse/ClickHouse/pull/113759 Related: https://github.com/ClickHouse/ClickHouse/issues/113493 ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114247",
          "createdAt": "2026-08-10T23:40:02Z",
          "updatedAt": "2026-08-13T14:50:15Z",
          "timestamp": "2026-08-13T14:50:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5a7861df2cdff5ac1055",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114131",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114131",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use a continuous primary-key range for whole-metric PromQL selectors of TimeSeries tables",
          "text": "A PromQL selector over a `TimeSeries` table filters the samples table with `id IN (SELECT id FROM tags WHERE <matchers>)`. For a metric with tens of thousands of series, `KeyCondition` runs its single-threaded generic exclusion search with the whole set: 284 ms per selector on a 62-billion-row part with 1.9M marks (503 ms at 8.1M marks), and rule-style queries evaluate up to 5 selectors. With the two-component id layout `Tuple(UInt64, UUID)` the canonical id generator derives the first component from the metric name alone, so all series of one metric form one continuous primary-key range. When a selector matches a whole metric — verified by metadata checks plus one `LIMIT 1` probe on the tags table, which also detects out-of-range ids left by an earlier `ALTER ... MODIFY SETTING id_generator` — the generated WHERE additionally carries `id >= tuple(hash(name), min) AND id <= tuple(hash(name), max)` and the inner query sets `use_index_for_in_with_subqueries_max_values = 1`. Index analysis uses the range; the `IN` stays for exact row filtering, so both emissions return identical rows on any data. Any failed check emits today's SQL unchanged. The range can select a few extra boundary granules (+5 of 135,005 marks on a 30-day scan). ## Measured effect tsbench PromQL suite: 62.455B samples / 361,432 series, 1.9M-mark part; Ryzen 9950X (16C/32T); baseline = clean master 9b6a2d7346f. Cold medians of 3 interleaved rounds: | query | master | this PR | delta | |---|--:|--:|--:| | s07 (30m range) | 2.22 s | 1.23 s | −44.8% | | r03 (rule, 3 selectors) | 5.02 s | 3.07 s | −38.8% | | s11 (24h range) | 4.30 s | 2.98 s | −30.8% | | r02 (25.6k-series instant) | 3.04 s | 2.17 s | −28.7% | | s06 (24h instant) | 5.04 s | 3.61 s | −28.4% | | full 24-query suite, cold geomean | 950 ms | 847 ms | **−10.8%** | Selectors that do not match a whole metric fall back and are unaffected (r05, s05: ±0.3%). Probe cost on non-firing selectors: ~2–4 ms each (r07: 66 → 73 ms); single-component id layouts never reach the probe. The removed cost grows with mark count, so the effect is larger at `index_granularity_bytes = 262144`. ## Tests `04836_time_series_selector_whole_metric_pk_range`: fires for whole-metric selectors (plan carries the range, `IN` retained), falls back byte-identically for label-filtered, regex, custom-generator, and ALTERed-`id_generator` history cases; both tuple layouts; full `prometheusQuery`/`prometheusQueryRange` results compared. `tests/performance/promql_selector_pk_range.xml`: 20,000-series metric, range path plus fallback control. Related: #113768 (open) — removes no-op casts in the same generated SELECT; complementary, each stands alone. --- ### Changelog category (leave one): - Performance Improvement ### Changelog entry: PromQL selectors that match all series of one metric now filter the samples table of a `TimeSeries` table with a continuous primary-key range on `id` during index analysis instead of a large `id IN <set>` condition, when the id layout is a two-component tuple with the canonical id generator. Removes the dominant single-threaded index-analysis cost of selector-heavy PromQL queries: up to −45% cold latency on dashboard and rule query shapes, −11% cold geomean over the full suite on a 62-billion-sample table. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114131",
          "createdAt": "2026-08-10T10:07:19Z",
          "updatedAt": "2026-08-13T15:55:30Z",
          "timestamp": "2026-08-13T15:55:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-performance",
            "comp-promql"
          ],
          "author": "nikitamikhaylov",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e3f1b1355e4c12ba6a5b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114624",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114624",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Find a LowCardinality needle equal to the type's default value",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `has`, `indexOf`, `countEqual`, `mapContainsKey`, `mapContainsValue` and `Map` subscript returning \"not found\" for a constant `LowCardinality` needle equal to the element type's default value, such as an empty `String` or a zero number. ### Description **The problem.** A constant needle equal to the element type's default value is never found in an `Array(LowCardinality(T))` or a `Map` with a `LowCardinality` key. No exception, so a `WHERE` on such a predicate silently drops rows: ```sql CREATE TABLE t (a Array(LowCardinality(String))) ENGINE = Memory; INSERT INTO t VALUES (['', 'a']); SELECT has(a, '') FROM t; -- 0, expected 1 ``` Likewise for `0` over the numeric types, a NUL-padded `FixedString` and the `1970-01-01` `Date`, while `m['']` returns `''` instead of the stored value. Any non-default needle is correct. Reproduces on 26.5 to 26.7. **Root cause.** A `LowCardinality` dictionary reserves prefix slots for the default and NULL values, and `ReverseIndex` is built with `num_prefix_rows_to_skip`, so the reserved slot is never indexed. `ColumnUnique::uniqueInsertData` compensates for that on the write path; the read path had no counterpart. **The change.** `ColumnUnique::getOrFindValueIndex` now performs the same default-slot match as the write path. `ReverseIndex` is untouched, so no write behaviour moves. That slot becoming reachable brings two equality details. The lookup casts the constant into the element type without reporting loss, so `UInt64(256)` arrived as `UInt8(0)` and would match the default; that slot now answers only if the element type can represent the constant. And a dictionary can hold `-0.0` and `0.0` as separate entries while the shortcut has room for one index, so a zero constant over a float element type is left to the value-comparing path, as `indexOfAssumeSorted` already is. The new test covers both call sites. Lookup timing is unchanged. **Overlap with #112953.** That open PR declines this same shortcut for every float element type, subsuming the float-zero decline here, and the two conflict textually; whichever merges second should keep the broader decline and drop this one, collapsing the duplicated representability predicate to one copy. Non-float types are independent of it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114624",
          "createdAt": "2026-08-13T11:55:50Z",
          "updatedAt": "2026-08-13T14:47:54Z",
          "timestamp": "2026-08-13T14:47:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d989a911a43b61165b7c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113505",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113505",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "S3 tables engine",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): S3 tables engine catalog for datalakes. Same as https://github.com/ClickHouse/ClickHouse/pull/103220, but with working INSERT",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113505",
          "createdAt": "2026-08-05T14:34:19Z",
          "updatedAt": "2026-08-13T14:45:53Z",
          "timestamp": "2026-08-13T14:45:53Z",
          "metrics": {
            "reactions": 3,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ed30c82480d1225dcac5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109225",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109225",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong results with parallel_hash JOIN and read-in-order-through-join",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/issues/109216 Related: https://github.com/ClickHouse/ClickHouse/pull/110671 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong results (silently dropped or mis-grouped rows) when `join_algorithm = 'parallel_hash'` is combined with a sorted consumer such as `optimize_aggregation_in_order`, `optimize_distinct_in_order` or `LIMIT BY`. With several join slots and a single-level hash map, `ConcurrentHashJoin` scatters the left block across slots, so the read-in-order-through-join optimization must no longer advertise the left sort order in that case. ### Description The `partial_merge` / `prefer_partial_merge` half of this PR has since been fixed on master by #110671, which added `IJoin::preservesLeftBlockOrder()` (defaulting to `true`) plus the `MergeJoin` / `JoinSwitcher` overrides and the `findReadingStep` gate. What remains here is a different carrier, and it is still a live wrong result on master. **`parallel_hash` (`ConcurrentHashJoin`) with several slots and a single-level map.** It inherits the `true` default from master, but `chooseMethod` leaves a key that materializes to one or two bytes (`key8` / `key16`) single-level - wider keys, including string and fixed-string ones, get a two-level variant. For a single-level map `joinBlock` scatters the left block across slots and `ConcurrentHashJoinResult` emits slot 0, then slot 1, and so on, so equal left-key values stop being contiguous while `findReadingStep` still installs the ordered read. Measured on master (`54dee101`, which contains #110671) against a debug build of this branch: ```sql CREATE TABLE t3 (a UInt32, j UInt8) ENGINE = MergeTree ORDER BY (a, j); CREATE TABLE t4 (j UInt8, v UInt64) ENGINE = MergeTree ORDER BY j; INSERT INTO t3 SELECT intDiv(number, 8)::UInt32, (number % 8)::UInt8 FROM numbers(64); INSERT INTO t4 SELECT (number % 8)::UInt8, number FROM numbers(8); SET join_algorithm = 'parallel_hash', max_threads = 8, optimize_aggregation_in_order = 1, max_bytes_before_external_join = 0, max_bytes_ratio_before_external_join = 0; SELECT a, count() FROM t3 LEFT ALL JOIN t4 ON t3.j = t4.j GROUP BY a ORDER BY a; ``` Master returns `1, 1, 1, 1, 1, 1, 1, 57`; the correct answer is 8 per group, which this branch returns. Ground truth was confirmed three independent ways (`optimize_aggregation_in_order = 0`, `join_algorithm = 'hash'`, `query_plan_read_in_order_through_join = 0`). The fix flips the `IJoin::preservesLeftBlockOrder()` default from `true` (fail-open) to `false` (fail-closed) and makes each join that really does stream the left side through once opt in: `HashJoin`, `DirectKeyValueJoin`, `ConstantJoin`, `PasteJoin` unconditionally, and `ConcurrentHashJoin` only when it does not scatter (`slots == 1 || twoLevelMapIsUsed()`). Flipping the default is what makes the contract hold by property rather than by accident. On master `FullSortingMergeJoin` has no override, so it inherits `true` - it is safe today only because `JoinStepLogical` inserts a `Sorting (Sort Left before JOIN)` step that `findReadingStep` does not descend through. That is a property of the current plan shape, not of the join, so any future plan change would silently reintroduce a wrong result. Under the fail-closed default it is safe by property. Precision was verified in both directions, so the stricter default does not cost the optimization anywhere it was previously correct: a two-level `UInt64` key still reads in order, a single-slot (`max_threads = 1`) `parallel_hash` join still reads in order, and `hash` / `direct` are unchanged. `GraceHashJoin` and `SpillingHashJoin` remain excluded through `hasDelayedBlocks()` as before. `topKThroughJoin.cpp` keeps its explicit `FullSortingMergeJoin` type check for its own mode 2 (a pre-JOIN `Sort` on the preserved input); its comment is updated to say the `preservesLeftBlockOrder()` read already covers that join and the type check is now belt-and-braces. Tests: `04498_distinct_in_order_partial_merge_join` fails on current master on exactly the `parallel_hash` block and passes here, so it is a live regression test rather than a restatement of #110671. `04500_read_in_order_through_constant_join` covers the `ConstantJoin` and `DirectKeyValueJoin` opt-ins, and `04500_limit_by_in_order_partial_merge_join` guards the `LIMIT BY` consumer. Each assertion was verified by mutation: with the corresponding override removed the assertion flips. All of them pin the whole read-in-order trio (`optimize_read_in_order`, `query_plan_read_in_order`, `query_plan_read_in_order_through_join`), since the stateless runner randomizes all three and a drawn `0` would make the plan assertions blind. #109216 is downgraded to `Related:` because the shape it reports (`prefer_partial_merge` + `optimize_distinct_in_order`) no longer reproduces on master after #110671; its reproducer now returns the correct 6 rows over repeated runs. This PR covers the sibling `parallel_hash` carrier of the same class, so it should not auto-close that issue.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109225",
          "createdAt": "2026-07-02T20:34:33Z",
          "updatedAt": "2026-08-13T14:45:46Z",
          "timestamp": "2026-08-13T14:45:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 42
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1d4c1e8ae0ef24998b96",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113909",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113909",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check the query cancellation while filling `system.parts` and its siblings",
          "text": "The tables based on `StorageSystemPartsBase` (`system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`) build the whole result eagerly in `initializePipeline`, so a cancelled or timed out query kept building rows over every storage and part until the very end. In a stress test with ThreadFuzzer this took minutes and tripped the hung check: a `SELECT` over `system.parts_columns` (with a filter matching every active part on the server) stayed in the process list for 252 seconds with `is_cancelled = 1` and `max_execution_time = 10`. Check the query status per storage and per part, following the pattern of `system.zookeeper` and `system.remote_data_paths`. The check also covers the storage-discovery prepass in `StoragesInfoStream` (the eager enumeration of all databases and tables), the analogous prepass in `StoragesDroppedInfoStream` (so `system.dropped_tables_parts` is interruptible as well), the skip loop in `StoragesInfoStreamBase::next`, and the per-storage part enumeration itself: the `MergeTreeData` helpers (`getDataPartsVectorForInternalUsage`, `getAllDataPartsVector`, `getProjectionPartsVectorForInternalUsage`, `getAllProjectionPartsVector`) take an optional `need_stop` callback that is checked periodically while the parts snapshot is being built. The two column-oriented tables additionally check the query status inside their column-enumeration loops and their metadata prepass, so a very wide table does not create a long uninterruptible stretch inside a single storage. The return value of `checkTimeLimit` is honored, so with `timeout_overflow_mode = 'break'` the eager build stops at the soft deadline and returns the rows collected so far. The table-lock acquisition in `StoragesInfoStreamBase::tryLockTable` is interruptible as well: instead of a single wait inside `RWLockImpl::getLock` for the whole `lock_acquire_timeout`, the lock is acquired in 100 ms slices with a query-status poll between the attempts (the total timeout and the `DEADLOCK_AVOIDED` semantics are preserved), so a killed or soft-timed-out query does not sit in the lock wait while a concurrent DDL query holds the drop lock. The test `04869_system_parts_lock_wait_cancellation` pins this with a share lock held by a long `SELECT` and a `DROP TABLE` in an `Ordinary` database queued behind it. The test uses the new `slowdown_system_parts_enumeration` failpoint, which only affects the specially named test tables (so concurrently running tests are unaffected). It sleeps 500 ms on every enumerated part, so building the full result for a 20-part table takes at least 10 seconds, and it sleeps 1 second per `COLUMNS_CANCELLATION_CHECK_PERIOD` (128) enumerated columns of a part, so building the full `system.parts_columns` / `system.projection_parts_columns` result over a single part with 1301 columns (and a projection over all of them) also takes at least 10 seconds. The test asserts that queries with a 1 second deadline in the `break` mode finish well under that, which is only possible by stopping at the per-part and per-column cancellation checkpoints. Timed assertions are needed because a plain row-count assertion cannot distinguish a build with the fix from one without: in the `break` mode the executor drops the eagerly built result after the deadline in both cases. The test also asserts partial row counts under a pre-expired deadline for all five tables, including `system.dropped_tables_parts` over a deterministic dropped-table fixture. The pre-expired-deadline checks also run under the failpoint, so the fewer-rows assertion is deterministic even on a machine fast enough to build the whole result in under a millisecond. For tables with the `_snap` name marker, the failpoint additionally slows down the parts-snapshot walks inside `MergeTreeData` (500 ms per enumerated part) and makes them poll the stop callback on every element, and for tables with the `_meta` name marker it slows down the column-metadata prepass of the column-oriented tables (1 second per 128 enumerated metadata columns), so the timed checks also prove that the snapshot materialization and the prepass are interruptible: all six of these checks fail against a binary with the `need_stop` polls and the prepass checkpoints disabled. Caught by `Stress test (arm_asan_ubsan, s3)`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113722&sha=ceedb772f6aa36e221d25ce48f9af71f032a7b08&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Related: https://github.com/ClickHouse/ClickHouse/pull/113722 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Queries over `system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`, and `system.dropped_tables_parts` now react to cancellation and `max_execution_time` while the result is being built.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113909",
          "createdAt": "2026-08-08T01:18:22Z",
          "updatedAt": "2026-08-13T14:44:50Z",
          "timestamp": "2026-08-13T14:44:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7cfd41400ac8433dde31",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114613",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114613",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `jsonPathValues` tokenizer for JSON text indexes",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 Related: https://github.com/ClickHouse/ClickHouse/pull/110757 I first tried to fix the edge cases between `JSONAllValues` and the text index. I kept finding more cases where the index could return an incorrect result. I therefore needed to disable more useful index paths to keep queries correct. The main problem is that `JSONAllValues` stores plain text. It does not retain the path or type of values from `Dynamic` JSON columns. PR #113376 documents several examples. A tokenizer made specifically for `JSON` seemed like a better fit. The new `jsonPathValues` tokenizer stores the JSON path, type, and value in each token. This gives the index enough information for safe typed comparisons. ```text +-------------------+-------+---------------------+-------+--------+------------------+ | escaped JSON path | 00 00 | binary-encoded type | 00 00 | 1-byte | payload | | | | | | kind | | +-------------------+-------+---------------------+-------+--------+------------------+ Payload: complete value : | full value | truncated value : | value prefix | SipHash-2-4 (8 bytes, LE) | map entry : | escaped key | 00 00 | complete/truncated value | validation : | empty | Kinds: 1/2 = scalar, 3/4 = array element, 5/6 = map entry (complete/truncated), 7 = dynamic validation Escaping: 00 -> 00 01 Component terminator: 00 00 ``` The path-value format also enables direct reads. This part was inspired by [ClickStack's use of text-index direct reads for dynamic map attributes](https://clickhouse.com/blog/making-clickstack-5x-faster-clickhouse-observability). `jsonPathValues` brings this model directly to `JSON` columns. It does not require user-defined alias columns or query rewrites. Long values use a bounded prefix and a hash. ClickHouse validates candidate rows when a token is truncated or a runtime type is unsafe. This prevents false-negative results. The tokenizer supports declared and dynamic paths, arrays, scalar leaves in `Array(JSON)`, and declared `Map(String, String)` paths. It supports equality, `IN`, `has`, map lookups, prefix and suffix searches, `LIKE`, `ILIKE`, regular expressions, and multi-search functions. `startsWith` uses an ordered dictionary lookup, so it does not require a full dictionary scan. Substring searches still scan the dictionary. ## Benchmark We tested `jsonPathValues(1024)` with 99,999,984 distinct events from 100 JSONBench files. Both target tables had the same 192-part layout. Each indexed query returned the same result as the query without an index. | Metric | No index | `jsonPathValues(1024)` | |---|---:|---:| | Build time | 214 s | 863 s | | Total storage | 9.29 GiB | 23.48 GiB | | Exact `cid` lookup | 1,926 ms | 13 ms | | Exact long-value lookup | 637 ms | 8 ms | | `startsWith` | 29 ms | 16 ms | | Substring `LIKE` | 145 ms | 93 ms | The index used 14.19 GiB, or approximately 152 bytes per JSON object. It made the table 2.53 times larger and made index construction 4.03 times slower. The exact lookups were 80 to 148 times faster. `startsWith` was 1.81 times faster. Query times are medians from ten warm runs. This PR also adds a checked-in performance test with 500,000 generated JSON rows. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add the `jsonPathValues` tokenizer for bounded, type-safe text indexes on `JSON` paths, arrays, and maps.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114613",
          "createdAt": "2026-08-13T09:35:07Z",
          "updatedAt": "2026-08-13T14:42:44Z",
          "timestamp": "2026-08-13T14:42:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:eb40b73c99fe8fb1eaa6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105499",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105499",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Drop for detached tables",
          "text": "### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental `DROP DETACHED TABLE` support, gated by `allow_experimental_drop_detached_table`, to remove metadata and data for detached tables. Continues #62490",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105499",
          "createdAt": "2026-05-21T09:59:53Z",
          "updatedAt": "2026-08-13T14:42:08Z",
          "timestamp": "2026-08-13T14:42:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-experimental"
          ],
          "author": "UberDever",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b614dd75e8609f80778f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112945",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112945",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use `pread` when `preadv2` with `RWF_NOWAIT` cannot be used, and recognize `EPERM` from it",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/104634 Related: https://github.com/ClickHouse/ClickHouse/issues/49149 Related: https://github.com/ClickHouse/ClickHouse/issues/39753 The `pread_threadpool` read method hands every read off to a thread pool, unless the data is already in the page cache, which it checks with the `preadv2` system call and the `RWF_NOWAIT` flag. Two things can go wrong with that check, and both were handled badly. **`EPERM` was not recognized.** It is what a `seccomp` profile of a container runtime answers for a system call that is not in its allow list. `ThreadPoolReader::submit` handed the read off to the thread pool for `ENOSYS` and `EOPNOTSUPP`, but let `EPERM` through to the throw, failing the query with `CANNOT_READ_FROM_FILE_DESCRIPTOR`. This is the signature reported in #49149. **The check was never verified in advance.** `hasBugInPreadV2` only compared the kernel version, and nothing else was checked, so on a system where the check cannot work `pread_threadpool` kept paying for a thread pool hand-off on every read - including the reads that only had to copy the data from the page cache. In #104634, on Amazon Linux 2 (kernel 5.10), the profile shows `ThreadPoolReaderPageCacheMiss` 1,971,514 out of `LocalThreadPoolJobs` 1,972,484 - every read went to the pool - while the device only moved ~22 GB of the 258 GB the file descriptor delivered, i.e. ~90% of the data was in the page cache and still paid for the hand-off. The system call is now probed once, before it is used, by `preadNoWaitUnavailableReason`. The probe passes an invalid file descriptor on purpose: `seccomp` filters and the system call table are consulted before the descriptor is looked up, so an available system call answers `EBADF` without reading anything, while a blocked one answers `EPERM` or `ENOSYS`. When it says the check cannot be used, `applySettingsQuirks` switches the default value of `local_filesystem_read_method` from `pread_threadpool` to `pread` at start time, and says why in the server log. Nothing downstream has to know: the reader, the userspace page cache eligibility in `DiskLocal::prepareRead` and the prefetched read pool all see a plain `pread` setting. As with the other settings quirks, a value set explicitly - in the configuration, or with `SET` at runtime - is left alone. Such a session keeps `pread_threadpool` and keeps paying for the hand-off, which is what happens today. The switch is a property of the host, so the setting is left marked as unchanged: only the changed settings are serialized into the query the initiator sends to the remote shards, and a host that cannot use the system call must not impose `pread` on the shards that can. An explicitly requested value stays changed and is still sent. The per-read `errno` handling in `ThreadPoolReader::submit` is kept, with `EPERM` added to it: the probe answers for the system call, but a particular filesystem can still reject the flag (`tmpfs` answers `EOPNOTSUPP`, for example), and such a read is handed off to the thread pool instead of failing the query. ### How it was tested `preadv2` was rejected the way a container runtime does it, with a `seccomp` filter installed by a small wrapper (`SECCOMP_RET_ERRNO`), and an old kernel was simulated with `setarch --uname-2.6`. Reading a 2 million row `MergeTree` table with the default `local_filesystem_read_method`, before (the released 26.7.1 binary) and after: | | before | after | |---|---|---| | no filter | `ThreadPoolReaderPageCacheHit` 39, `LocalThreadPoolJobs` 90 | `ThreadPoolReaderPageCacheHit` 34, `ThreadPoolReaderPageCacheMiss` 3, `LocalThreadPoolJobs` 93 - the read method stays `pread_threadpool` | | `preadv2` → `EPERM` | `Code: 74 ... errno: 1, Operation not permitted (CANNOT_READ_FROM_FILE_DESCRIPTOR)`, already while attaching the table | the read method is `pread`, the query succeeds, no `ThreadPoolReader` events | | `preadv2` → `ENOSYS` | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 - every read to the pool | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 80 | | kernel reported as older than 5.11 | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 81 | In all three rejected cases the reason is in the log, for example: ``` <Warning> SettingsQuirks: The default value of local_filesystem_read_method has been switched from 'pread_threadpool' to 'pread' (you can explicitly set it back still), because the `preadv2` system call is not available (the probe with an invalid file descriptor answered errno: 1, strerror: Operation not permitted instead of `EBADF`); it is typically rejected by a `seccomp` profile of a container runtime, and can be allowed in the runtime configuration. ... ``` An explicitly requested `local_filesystem_read_method = 'pread_threadpool'` is kept, on every one of those systems, and this is where the per-read `EPERM` handling earns its place: under the `EPERM` filter the same query fails on the released binary and succeeds here, with `ThreadPoolReaderPageCacheMiss` 45 and `LocalThreadPoolJobs` 135 - every read handed off to the pool, which is the documented cost of asking for it there. Under the same filters, `system.settings` reports `local_filesystem_read_method = 'pread'` with `changed = 0`, so nothing is forwarded to the remote shards, while an explicitly requested `pread_threadpool` reports `changed = 1` and is still sent. Unit tests cover the `errno` classification, the probe's `EBADF` contract, and the quirk itself (the default is switched exactly when the probe says the system call is unusable, the switched value stays out of `Settings::changes()`, and an explicitly set value is never switched). An automated end-to-end test would need an instance with a restrictive `seccomp` profile, which the integration test framework cannot express today - it starts every instance with `seccomp:unconfined`. ### Documented behavior impact On a system where the page cache cannot be checked without waiting for the disk - Linux older than 5.11, a sandbox that rejects `preadv2` with an error code, and systems other than Linux, where `preadv2` does not exist and the method never checked the page cache in the first place - the default value of `local_filesystem_read_method` becomes `pread` instead of `pread_threadpool`, so local reads are performed in the calling thread. This includes the reads with `O_DIRECT`, which never look at the page cache and do not need the check; the read method is now resolved once, for the whole server, so they follow the same value. Nothing changes on a supported system, and nothing changes for an explicitly configured read method. The setting description in `src/Core/Settings.cpp`, from which the docs are generated, is updated in this PR. A `seccomp` profile that terminates the process instead of rejecting the system call with an error code cannot be detected: the startup probe is itself a `preadv2` call, so such a profile kills the server there. On `master` it kills it at the first read instead, under the same default read method. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `pread_threadpool` read method needs the `preadv2` system call with the `RWF_NOWAIT` flag to read the data that is already in the page cache without handing the read off to a thread pool. It is now checked at start time whether that system call can be used, and if it cannot - the Linux kernel is older than 5.11, or a `seccomp` profile of a container runtime rejects the system call - the default value of `local_filesystem_read_method` is switched to `pread`, and the reason is reported in the server log. Previously, every read paid for a thread pool hand-off on such systems, and a `seccomp` profile that answers `EPERM` made queries fail with `CANNOT_READ_FROM_FILE_DESCRIPTOR`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112945",
          "createdAt": "2026-08-01T20:09:06Z",
          "updatedAt": "2026-08-13T14:41:57Z",
          "timestamp": "2026-08-13T14:41:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 26
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5b0552047e468bcb82f8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114212",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114212",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain the required column order when a dictionary QUERY returns columns in the wrong order",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a dictionary with a PostgreSQL source and a custom `QUERY` returns its columns in an order that does not match the dictionary structure, the error now names the destination dictionary column and the column order the query must return, instead of only reporting that a value could not be parsed. ### Description A `RANGE_HASHED` dictionary over `SOURCE(POSTGRESQL(... QUERY '...'))` failed with `Cannot parse LocalDate: 97955`, where `97955` was an unrelated `id` column's value: the DDL was valid, only the column order was wrong, and the message pointed at the data. A dictionary's expected source structure is keys-first: key(s), `RANGE MIN`, `RANGE MAX`, then the attributes. ClickHouse aliases every column when it builds the `SELECT` itself, but a user `QUERY` passes through verbatim and the result is read strictly by position, so `id` was deserialized into the `Date` attribute `contract_time`. On a conversion failure `PostgreSQLSource` now names the destination column and its position, and for a custom `QUERY` states the required order, derived from the dictionary's own structure so it suits every layout (the `RANGE` wording appears only for range dictionaries). Successful loads are untouched, and `BAD_ARGUMENTS` is preserved because `MaterializedPostgreSQL` relies on it. <details><summary>The message for the reported case</summary> ``` Cannot parse PostgreSQL value '97955' as Date: Cannot parse LocalDate: 97955: while reading column 2 of the result into `contract_time`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `meter_no`, `contract_time`, `end_date`, `id` (the key column(s) first, then the RANGE MIN and RANGE MAX columns, then the remaining attributes) ``` For a `FLAT` dictionary the same failure omits the `RANGE` clause: ``` Cannot parse PostgreSQL value 'abc' as Int64: Could not convert string to l: 'abc': while reading column 2 of the result into `num`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `id`, `num`, `name` ``` </details> A new integration test covers the range case, the documented order, and a `FLAT` dictionary asserting no `RANGE` wording appears; it fails on master. `04401_composite_key_dict_non_leading_key` gains the range shape over the ClickHouse source, pinning its by-name reordering against regression. Scope, since the mechanism is wider than the fix. The MySQL dictionary source (and likely XDBC and Cassandra) maps positionally too, but its parse errors do not funnel through one place, so folding it in grows the diff; say the word and I will extend it. Two cases stay untouched: an unnamed-column query still matches positionally and can silently return a wrong value, and a short row silently defaults trailing columns. Matching by name is possible, not ruled out: a `LIMIT 0` probe, or projecting the expected names inside the streamed query. I avoided it because it silently changes behaviour for every dictionary relying on positional matching, and arbitrary queries can return duplicate or unnamed columns.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114212",
          "createdAt": "2026-08-10T19:51:47Z",
          "updatedAt": "2026-08-13T14:40:07Z",
          "timestamp": "2026-08-13T14:40:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7cfd5cbca82a76da1380",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110183",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110183",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use-of-uninitialized-value in WITH FILL suffix over a merge",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107074 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-of-uninitialized-value in `ORDER BY ... WITH FILL ... STALENESS` when the fill suffix generates no rows (for example when the staleness window leaves nothing to fill) and the result is read through a merge of sorted streams (such as a two-shard `Distributed` table). ### Description Found by the AST fuzzer under MSan on top of #107074. `FillingTransform`, when all input chunks are processed, may run the suffix path. If the fill constraints are satisfied but no fill rows are produced (e.g. `WITH FILL ... STALENESS` with an exhausted staleness window), `generateSuffixIfNeeded` returns `true` while the result columns are freshly `cloneEmpty()`'d and carry no data. Previously the transform still emitted a 0-row chunk built from those empty columns. A downstream `MergingSortedTransform` (as used when reading from a two-shard `Distributed` table) then built a sort cursor over that empty chunk and compared row 0, reading past the end of the empty column. Fix: do not emit the suffix chunk when it has no rows. Minimal reproducer (needs a two-shard merge on the initiator): ```sql CREATE TABLE m (key Int) ENGINE = Memory; INSERT INTO m VALUES (100); CREATE TABLE d2 AS m ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), m); SELECT _shard_num FROM d2 ORDER BY _shard_num ASC WITH FILL TO 46 STALENESS 1; ``` CI finding: `AST fuzzer (amd_msan)`, report https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107074&sha=3ec54c88d4eb535a5d644fc1ab91af31d717f9a5&name_0=PR&name_1=AST%20fuzzer%20%28amd_msan%29 MSan use-of-uninitialized-value in `ColumnVector<UInt32>::doCompareAt` (`MergingSortedAlgorithm::consume`), origin `FillingTransform::initColumns` (`cloneEmpty`) via the suffix path. Verified on a local `amd_msan` build: reproduces before the fix (server aborts), clean after; the new regression test returns `1\\n2`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110183",
          "createdAt": "2026-07-12T21:10:52Z",
          "updatedAt": "2026-08-13T14:40:04Z",
          "timestamp": "2026-08-13T14:40:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b7104341fbdc59fe1780",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107305",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107305",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make ALTER MODIFY COLUMN on named Tuple metadata-only when adding subfields",
          "text": "`ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` that only adds new subfields is now metadata-only (no mutation). Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1` (default `false`). Subfield additions through `Array`/`Map`/nested `Tuple` wrappers are also handled. `Nullable(Tuple(...))` is blocked (null map incompatibility). Removing/renaming subfields or changing types still triggers a mutation. Key/index/projection guards reject the metadata-only path when `primary.idx` or skip-index bytes would become invalid (whole tuple in key, or subcolumn whose type changes). Subcolumn references with unchanged types (e.g. `ORDER BY t.a` when only `t.c` is added) are allowed. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` is now metadata-only when only adding subfields, matching the speed of top-level `ADD COLUMN`. Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1`. ### Documentation entry for user-facing changes - [x] Documentation is not required (behavioral improvement; semantics unchanged)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107305",
          "createdAt": "2026-06-12T07:42:58Z",
          "updatedAt": "2026-08-13T15:55:26Z",
          "timestamp": "2026-08-13T15:55:26Z",
          "metrics": {
            "reactions": 1,
            "comments": 8
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:58df92d47633cdf03c39",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114001",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114001",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Normalize open range bounds only for integer key types",
          "text": "Closes https://github.com/ClickHouse/ClickHouse/issues/113993 `Range` previously tightened every `Int64` or `UInt64` endpoint without knowing the key type it would be compared against. Move that normalization into `KeyCondition` after checking the key type is represented by an integer, preserving integer pruning while keeping fractional and future non-integer domains open. Related: https://github.com/ClickHouse/ClickHouse/issues/113993 Caused by: https://github.com/ClickHouse/ClickHouse/pull/98410 <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed incorrect MergeTree index pruning for strict comparisons between fractional types and integer constants.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114001",
          "createdAt": "2026-08-09T01:50:42Z",
          "updatedAt": "2026-08-13T14:39:25Z",
          "timestamp": "2026-08-13T14:39:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "v26.4-must-backport"
          ],
          "author": "EmeraldShift",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:f58248d88a841f0aa55b",
        "signalId": "github:ClickHouse/ClickHouse:issue:72380",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:72380",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "`RENAME DATABASE` query doesn't work for materialized views",
          "text": "**Company or project name** Prefer not to specify **Describe what's wrong** `RENAME DATABASE` query doesn't seem to be working correctly with materialized view statements. After renaming a database with `RENAME DATABASE old_name TO new_name` query, the materialized views would still reference tables with the `old_name` as could be seen with `SHOW CREATE new_name.materialized_view` statement in the repro. That leads to a bunch of issues with database permissions and `SELECT` queries. And while `INSERT` queries into the main table seem to be working, they are not triggering any relative materialized views. Repro: https://fiddle.clickhouse.com/797f730d-d037-4b64-8d17-ce47911fccdd **Does it reproduce on the most recent release?** Yes **Enable crash reporting** Doesn't crash **How to reproduce** I first got this on `24.8.1.10452` in ClickHouse Cloud, but I'm also able to reproduce it on current latest. ``` CREATE DATABASE test; CREATE TABLE test.sample ( id UUID DEFAULT generateUUIDv4(), data TEXT DEFAULT '' ) ENGINE = MergeTree PRIMARY KEY (id) ORDER BY (id); -- Create materialized view without explicitly specifying target table CREATE MATERIALIZED VIEW test.inline_mat_view ( uuid UUID, data TEXT ) ENGINE MergeTree ORDER BY (uuid) AS SELECT id as uuid, data FROM test.sample; -- Create table for materialized view explicitly CREATE TABLE test.explicit_table ( uuid UUID, data TEXT ) ENGINE MergeTree ORDER BY (uuid); -- Create materialized view pointing to explicitly created table CREATE MATERIALIZED VIEW test.explicit_mat_view TO test.explicit_table AS SELECT id as uuid, data FROM test.sample; -- Output original CREATE statements for both materialized views SHOW CREATE test.inline_mat_view; SHOW CREATE test.explicit_mat_view; RENAME DATABASE test TO dev; -- Inserting data into main table doesn't trigger any error, but materialized views are not updated INSERT INTO dev.sample (data) VALUES ('test1'), ('test2'), ('test3'), ('test4'), ('test5'); -- CREATE statements for materialized views after renaming the database still points to the old database name SHOW CREATE dev.inline_mat_view; SHOW CREATE dev.explicit_mat_view; -- Exception due to missing 'test' database SELECT * FROM dev.inline_mat_view; SELECT * FROM dev.explicit_mat_view; ``` **Expected behavior** When renaming the database I would expect materialized view to be updated accordingly and to point to the correct table in the renamed database.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/72380",
          "createdAt": "2024-11-25T10:53:38Z",
          "updatedAt": "2026-08-13T14:39:02Z",
          "timestamp": "2026-08-13T14:39:02Z",
          "metrics": {
            "reactions": 1,
            "comments": 5
          },
          "labels": [
            "potential bug"
          ],
          "author": "biased-badger",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:60d4eee1bbc8b7527e73",
        "signalId": "github:ClickHouse/ClickHouse:issue:113993",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113993",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Minmax index incorrectly prunes fractional DateTime64 values for integer bounds",
          "text": "### Company or project name _No response_ ### Describe what's wrong A minmax data-skipping index can change the result of a strict comparison between DateTime64 and an integer. Without skip-index pruning, `time > 0` correctly matches a value of 0.01 seconds. With a minmax index, the same row is pruned. ### Does it reproduce on the most recent release? Yes ### How to reproduce https://fiddle.clickhouse.com/66f7a197-cd87-4248-8e8c-6c44a4b85e4b ```sql CREATE TABLE t ( time DateTime64(9), INDEX idx_time time TYPE minmax GRANULARITY 1 ) ENGINE = MergeTree ORDER BY tuple() SETTINGS index_granularity = 1; INSERT INTO t VALUES (0.01::Decimal(9, 2)::DateTime64(9)); SELECT count() FROM t WHERE time > 0 SETTINGS use_skip_indexes = 0; -- 1 SELECT count() FROM t WHERE time > 0 SETTINGS force_data_skipping_indices = 'idx_time'; -- 0 EXPLAIN indexes = 1 SELECT * FROM t WHERE time > 0; ``` `EXPLAIN` shows the minmax condition as `time in [1, +Inf)`. ### Expected behavior Both queries should return 1. A data-skipping index must not change the query result. ### Error message and/or stacktrace _No response_ ### Related issues and pull requests _No response_ ### Additional context `Range::shrinkToIncludedIfPossible` changes the open UInt64 bound `> 0` into the closed bound `>= 1`. That transformation is valid for an integer value domain, but not for DateTime64(9), which can contain values between 0 and 1. KeyCondition treats native integers and DateTime64 as directly comparable, so the integer bound reaches this normalization without first being converted into the DateTime64 domain.",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113993",
          "createdAt": "2026-08-08T23:59:27Z",
          "updatedAt": "2026-08-13T14:36:45Z",
          "timestamp": "2026-08-13T14:36:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "clickgap-analyzed",
            "culprit-pr-pinned"
          ],
          "author": "EmeraldShift",
          "state": "open",
          "assignees": [
            "amosbird"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:786274be2fcf03d92cbb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113681",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113681",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Replace the per-bucket hash map in `timeSeries*ToGrid` with a sorted-append sample array",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replaced the per-bucket hash map inside the `timeSeries*ToGrid` aggregate functions with a flat sorted array of samples: sample ingestion becomes an O(1) append for in-order inputs (the overwhelmingly common case) and the per-bucket copy-and-sort at finalization is gone. ### Description The `timeSeries*ToGrid` functions kept each bucket's samples in an `absl::flat_hash_map<timestamp, value>`: every `add()` paid a hash-map emplace, and the order-dependent functions (`rate`, `increase`, `delta`, `changes`, `resets`) copied and sorted every bucket at finalization. Yet the input is almost perfectly ordered — samples come from MergeTree tables sorted by `(id, timestamp)`; an instrumented probe on a 32-thread read of a 62.5-billion-sample table counted **1 out-of-order add in 1,474,559,998**. The bucket is now a flat, memory-tracked vector of `(timestamp, value)` pairs: O(1) append while timestamps ascend, in-place max on an equal timestamp, and a rare out-of-order add just clears a `sorted` flag — normalization (sort + max-dedup) runs lazily, only for buckets that actually saw disorder. `merge()` is a linear merge of sorted runs with an append fast path for disjoint time ranges. `forEachSample` now guarantees ascending order, so the copy-and-sort buffers are deleted from the rate/delta/changes aggregators. The wire format and `FORMAT_VERSION`s are unchanged; `deserialize()` assumes no order of incoming pairs (old peers send hash-map iteration order), so mixed-version clusters interoperate — verified in both directions. Duplicate timestamps keep the larger value with the old `std::max` argument order. Measured on the 62.5-billion-sample table: a 30-day `sum by(...)(rate(...))` over 25,600 series drops ~6% of total query CPU (1106 s -> 1037 s); wall time and peak memory move little (the scan dominates the critical path, and raw sample storage dominates the state either way). The structural point is what this enables: the sorted buffer is the prerequisite for O(1) per-segment summaries in the rate family (follow-up), which is where the ~44 GiB state peaks of such queries actually go away. All query results are fingerprint-identical. Tests: a new stateless test (shuffled and duplicate-timestamp inputs in both orders, NaN at duplicated timestamps, `-Merge` of unsorted in-memory states, interleaved parts, two-level merges, serialized-state merges through `remote('127.0.0.{1,2}', ...)`, an `AggregatingMergeTree` roundtrip, a fixed state literal in old-peer wire order); all 40 existing timeseries/PromQL stateless tests pass byte-identically; the perf test gains an ingestion-heavy scenario (50M rows, 10k series). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113681",
          "createdAt": "2026-08-06T13:46:38Z",
          "updatedAt": "2026-08-13T15:36:36Z",
          "timestamp": "2026-08-13T15:36:36Z",
          "metrics": {
            "reactions": 1,
            "comments": 8
          },
          "labels": [
            "pr-performance",
            "comp-promql"
          ],
          "author": "nikitamikhaylov",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5980139a10a49f19a127",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114479",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114479",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enable reading in reverse order with FINAL for ReplacingMergeTree",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/58035 Related: https://github.com/ClickHouse/ClickHouse/pull/58361 Related: https://github.com/ClickHouse/ClickHouse/pull/111609 When a query with `FINAL` sorts in reverse order of the sorting key, for example `ORDER BY key DESC LIMIT n`, the read-in-order optimization was disabled and the query read the whole table. It now applies for `ReplacingMergeTree`. `ReplacingSortedAlgorithm` learns a `read_in_reverse` mode. A row with a strictly higher version always replaces the selected one; among rows with equal (or absent) versions, the previously selected row is kept unless the current row comes from a newer data part. That mirrors the \"last written row wins\" rule of the direct reading order, because in the reverse reading order rows within one part arrive backwards while the parts still arrive from the oldest to the newest one. This is what makes the reverse read select the same row of a duplicate key group as a direct read, which is the correctness concern that stopped https://github.com/ClickHouse/ClickHouse/pull/58361. The other engines keep the previous behavior, since their merging algorithms depend on the direct order of rows: the sequence of sign rows in `CollapsingMergeTree`, the order of rows fed to order-dependent aggregate functions in `AggregatingMergeTree`, and so on. On a 110 million row `ReplacingMergeTree` table with two overlapping parts, `SELECT x FROM t FINAL ORDER BY x DESC LIMIT 1` (measured on a `RelWithDebInfo` build): | | before | after | |---|---|---| | Rows read | 110.1 million | 1.71 million | | Elapsed | 4.5 s | 0.22 s | | Peak memory | 41 MB | 23 MB | Trade-off: as with the already existing direct-order in-order reads with `FINAL`, an in-order plan disables vertical `FINAL` and the splitting of parts ranges into intersecting and non-intersecting ones. A query that reads the full result with `ORDER BY key DESC` and no small `LIMIT` may therefore become slower on a wide or mostly merged table. The new setting `optimize_read_in_reverse_order_final` (default enabled) turns the optimization off, and `compatibility` with an earlier version restores the previous plans. Out of scope, to keep this change reviewable: `Merge` tables over `ReplacingMergeTree`, and the interaction with `topKThroughJoin`, which keeps its own optimization for `... FINAL LEFT JOIN ... ORDER BY key DESC LIMIT n`. Both can follow separately. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Enable the read-in-order optimization for queries with the `FINAL` modifier that sort in reverse order of the sorting key on `ReplacingMergeTree` tables, so that queries such as `SELECT ... FROM t FINAL ORDER BY key DESC LIMIT n` read only the relevant tail of the data instead of the whole table. Can be disabled with the new setting `optimize_read_in_reverse_order_final`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114479",
          "createdAt": "2026-08-12T12:07:10Z",
          "updatedAt": "2026-08-13T14:33:14Z",
          "timestamp": "2026-08-13T14:33:14Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "cwurm",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ce3f74e31b3639302831",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114178",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114178",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"NATS: add inline credentials setting\"",
          "text": "Reverts https://github.com/ClickHouse/ClickHouse/pull/110733 (merge commit `48ebef16af656664bab51cd5bb605a65e9b8ec9e`), which added the `nats_credentials` setting to the `NATS` table engine. Removed by this pull request: * the `nats_credentials` setting and its use in `NATSConnection` (`natsOptions_SetUserCredentialsFromMemory`), so the engine again takes credentials only as a path through `nats_credential_file`; * the mutual-exclusion check between `nats_credential_file` and `nats_credentials`, including the named-collection provenance handling in `registerStorageNATS`; * the runtime handling of `nats_credentials` only — the logging-only masking of the old spelling is deliberately kept. `nats_credentials` stays in both mask lists (`NATS::SETTINGS_TO_HIDE` and `FunctionSecretArgumentsFinder::nats_secret_keys`), because a query is formatted for logging before the storage settings are validated, so the old spelling must not leak the raw JWT/seed into `query_log` even though the server then rejects it. The masking branch (`findNATSTableEngineSecretArguments`) is kept as well: it also masks the still-supported secret keys (`nats_password`, `nats_token`, `nats_credential_file`) and the `nats_url` userinfo password in the table-engine argument form `ENGINE = NATS(collection, key = ...)`. The `ParserCreateQuery.MaskNATS*` unit tests are kept, and a new `MaskNATSTableEngineRemovedCredentialsSetting` test covers both spellings of the removed setting; * the test `04665_nats_credentials_named_collection`; * the `nats_credentials` lines from the `Documentation` block of `registerStorageNATS`. Two notes on the mechanics of the revert: * `src/Parsers/FunctionSecretArgumentsFinder.{h,cpp}` are resolved by hand: `findNATSTableEngineSecretArguments` and the full `nats_secret_keys` list (including `nats_credentials`, for masking only) are kept. Every other file comes out either byte-identical to its pre-`#110733` state or equal to it plus unrelated later changes (the message-broker schedule pool now returns a `shared_ptr`). * `docs/reference/engines/table-engines/integrations/nats.mdx` still mentions `nats_credentials`. That page is generated from the `Documentation` block this pull request updates, and direct edits of an `{/*AUTOGENERATED_START*/}` region are rejected by the docs check, so the page is left to the nightly documentation autogeneration. The setting was merged into `master` for the unreleased `26.8` and is not part of any release branch (the newest is `26.7`), so no changelog entry is needed. Related: https://github.com/ClickHouse/ClickHouse/pull/110733 Related: https://github.com/ClickHouse/ClickHouse/issues/85213 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1332` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114178",
          "createdAt": "2026-08-10T15:32:57Z",
          "updatedAt": "2026-08-13T14:32:50Z",
          "timestamp": "2026-08-13T14:32:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d849688b5b3053d0242c",
        "signalId": "github:ClickHouse/ClickHouse:issue:112586",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:112586",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Range condition in JOIN ON against a single-row side is not used for index analysis (no part pruning) since the logical join step became default",
          "text": "### Company or project name ClickHouse Inc. Found while comparing two server versions on an internal test cluster, on a reporting workload that derives its date window from a joined single-row subquery. ### Describe the situation When a range predicate on a `MergeTree` primary/partition key column is expressed in the `ON` clause of a `JOIN` against a single-row constant side, the condition is no longer turned into a `KeyCondition`. Index analysis reports `Condition: true` and every part is read, instead of pruning. The logical join step turns the inequality `ON` into a `cross` join plus a `Filter (Post Join Actions)` sitting above the join, and that filter never reaches index analysis. The same predicate written as a plain `WHERE` (or as scalar subqueries in `WHERE`) prunes correctly, and results are identical either way, so this is purely lost part/granule pruning. This is a common shape for \"current week / current month\" reporting queries, where the bounds are computed once in a CTE and joined in: ```sql WITH bounds AS (SELECT toStartOfWeek(today(), 1) AS lo, today() AS hi) SELECT ... FROM big_table AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi ``` On a real workload (a ~365M row table partitioned by month) this was the difference between reading 0 parts and reading all 302 parts, and between **1.4s and 35-68s (~25x)** for a query whose correct answer was an empty result set. ### Which ClickHouse versions are affected? The pruning is lost whenever the logical join step is used, i.e. whenever `query_plan_use_new_logical_join_step = 1`. **The regression is 26.5.** That is where https://github.com/ClickHouse/ClickHouse/pull/104017 made the setting `Obsolete`, hardcoding it to `true`. Up to and including 26.4 it was a normal `Production` setting, so anyone hitting this could simply set it to `0` and get pruning back. From 26.5 on there is **no setting-level mitigation at all** and the only fix is to rewrite the query. For completeness on the earlier history: the declared default went `false` -> `true` in 25.2 (https://github.com/ClickHouse/ClickHouse/pull/74909), so a 25.2..26.4 user who never touched the setting also saw no pruning. But that was recoverable, and deployments that pinned `query_plan_use_new_logical_join_step = 0` (ClickHouse Cloud among them, via its `compatibility` profile) kept correct pruning all the way through 26.4 and only lost it on upgrade past 26.5. Verified with official builds: | Version | Default behavior | With `query_plan_use_new_logical_join_step = 0` | |---|---|---| | 26.3.1.876 | `Parts: 14/14` (no pruning) | `Parts: 1/14` (prunes) | | 26.7.1.2043 | `Parts: 14/14` | setting is `Obsolete`, no effect | | 26.8.1.48 (master, `d38cb514a590e60af07777d8f48c34afdf20aa1f`) | `Parts: 14/14` | setting is `Obsolete`, no effect | `SETTINGS compatibility = '26.4'` does **not** restore the pruning on master either. ### How to reproduce Works in `clickhouse local`, no special settings needed. ```sql CREATE TABLE repro_join_prune (d Date, v UInt32) ENGINE = MergeTree PARTITION BY toYYYYMM(d) ORDER BY d; INSERT INTO repro_join_prune SELECT toDate('2025-01-01') + number, number FROM numbers(400); -- (1) bounds arrive via JOIN ON: no pruning on 25.2+ EXPLAIN indexes = 1 WITH bounds AS (SELECT toDate('2025-06-01') AS lo, toDate('2025-06-10') AS hi) SELECT count() FROM repro_join_prune AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi; -- (2) same predicate as a plain WHERE: prunes correctly on every version EXPLAIN indexes = 1 SELECT count() FROM repro_join_prune AS t WHERE t.d >= toDate('2025-06-01') AND t.d <= toDate('2025-06-10'); -- (3) same predicate as scalar subqueries: also prunes correctly on every version EXPLAIN indexes = 1 SELECT count() FROM repro_join_prune AS t WHERE t.d >= (SELECT toDate('2025-06-01')) AND t.d <= (SELECT toDate('2025-06-10')); ``` All three queries return `10`, so results are correct in every case. Plan for query (1) on master. Note that the range predicate is present as a `Filter (Post Join Actions)` but index analysis still gets `Condition: true`: ``` Aggregating │ Keys: │ Aggregates: count() │ Skip merging: 0 └──Filter (Post Join Actions) │ Filter column: d >= '2025-06-01' AND d <= '2025-06-10' └──Join (JOIN FillRightFirst) │ t[400] ⋈ system.one[1] │ Type: cross | Strictness: all | Algorithm: HashJoin │ Result rows: 400 │ Output: │ Left: __join_result_dummy, d │ Right: hi, lo ├──ReadFromMergeTree (default.repro_join_prune) │ Read type: Default │ Parts: 14 | Granules: 14 │ Output: d │ Indexes: │ Min-Max │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ Partition │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ PrimaryKey │ Condition: true │ Parts: 14/14 │ Granules: 14/14 │ Ranges: 14 └──ReadFromSystemOne ``` Plan for query (2) on master, for comparison: ``` Aggregating │ Keys: │ Aggregates: count() │ Skip merging: 0 └──Filter ((WHERE + Change column names to column identifiers)) │ Filter column: d >= '2025-06-01' AND d <= '2025-06-10' └──ReadFromMergeTree (default.repro_join_prune) Read type: Default Parts: 1 | Granules: 1 Output: d Indexes: Min-Max Keys: d Condition: and((d in (-Inf, 20249]), (d in [20240, +Inf))) Parts: 1/14 Granules: 1/14 Partition Keys: toYYYYMM(d) Condition: and((toYYYYMM(d) in (-Inf, 202506]), (toYYYYMM(d) in [202506, +Inf))) Parts: 1/1 Granules: 1/1 PrimaryKey Keys: d Condition: and((d in (-Inf, 20249]), (d in [20240, +Inf))) Parts: 1/1 Granules: 1/1 Search Algorithm: binary search Ranges: 1 ``` To see the \"good\" plan for query (1), run it on 26.4 or earlier with `SETTINGS query_plan_use_new_logical_join_step = 0`. ### Expected performance Query (1) should prune the same way queries (2) and (3) do, since the join is against a single-row constant side and the `ON` conditions are a plain range over the sorting/partition key. It did prune before the logical join step became the default, and it still prunes on 26.4 and earlier with `query_plan_use_new_logical_join_step = 0`. Note that a `CROSS JOIN` with the same predicate moved into `WHERE` does **not** prune on any version tested, including 26.4 with the old join step. That looks like a separate, pre-existing limitation rather than part of this regression, but it may share a root cause. ### Related issues and pull requests Caused by: https://github.com/ClickHouse/ClickHouse/pull/104017 Related: https://github.com/ClickHouse/ClickHouse/pull/74909 ### Additional context The following settings were tried on 26.6+ and none of them restore the pruning: - `query_plan_merge_filter_into_join_condition = 1` - `query_plan_convert_join_to_in = 1` - `allow_general_join_planning = 1` - `query_plan_filter_push_down = 1` - `query_plan_convert_outer_join_to_inner_join = 1` - `query_plan_optimize_join_order_limit = 10` - `use_join_disjunctions_push_down = 1` - `compatibility = '26.4'` The only workaround we found is a query rewrite: duplicate the bounds as a plain `WHERE` on the scanned table. The `JOIN` can stay in place, so the rewrite is semantics-preserving, and it restored the real workload from 35-68s to 1.3s. ```sql WITH bounds AS (SELECT toDate('2025-06-01') AS lo, toDate('2025-06-10') AS hi) SELECT count() FROM repro_join_prune AS t JOIN bounds ON t.d >= bounds.lo AND t.d <= bounds.hi WHERE t.d >= toDate('2025-06-01') AND t.d <= toDate('2025-06-10'); ``` Because the escape-hatch setting is now obsolete, users upgrading from 26.4 (or from any version where they had pinned `query_plan_use_new_logical_join_step = 0`) to 26.5+ hit this as a hard performance regression with no setting-level mitigation. <!-- ch-version-info:start --> ### Version info - Resolved by: #113484 - Backported to: `26.7.4.24`, `26.6.3.33` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/112586",
          "createdAt": "2026-07-30T12:41:56Z",
          "updatedAt": "2026-08-13T14:32:19Z",
          "timestamp": "2026-08-13T14:32:19Z",
          "metrics": {
            "reactions": 2,
            "comments": 6
          },
          "labels": [
            "performance"
          ],
          "author": "fm4v",
          "state": "closed",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ccc87bab77dbbc5b4549",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:63383",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:63383",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve the performance of `MODIFY TTL`",
          "text": "`ALTER TABLE ... MODIFY TTL` currently rewrites every part of the table, which on a large table means reading and writing all of its data just to change when rows expire. Very often the new TTL is the old one shifted in time - the retention period is extended or shortened, e.g. `create_time + INTERVAL 300 DAY` becomes `create_time + INTERVAL 10 DAY`. In that case every row's expiry time moves by the same constant number of seconds, so the parts do not have to be rewritten at all: it is enough to shift the expiry timestamps ClickHouse already stores per part. This pull request adds that fast path, and the result for the user is that such a `MODIFY TTL` completes almost instantly instead of taking minutes or hours: ```sql CREATE TABLE test_fast_ttl (`id` UInt32, `name` String, `create_time` DateTime) ENGINE = MergeTree ORDER BY id TTL create_time + toIntervalDay(300); INSERT INTO test_fast_ttl SELECT number, 'AAA', date_sub(day, 100, now()) from numbers(100000000); -- Before ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 25.564 sec. -- After ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 0.046 sec. ``` There is nothing to enable and no new syntax: the optimization is applied automatically inside the `MATERIALIZE TTL` mutation that `MODIFY TTL` already produces, and a plain `ALTER TABLE ... MATERIALIZE TTL` benefits from it as well. The observable result is exactly the same as before - the same rows expire and the parts end up with the same TTL bounds - only the work is avoided. Per part, the mutation now does one of the following: - the part is fully expired under the new TTL - it is replaced with an empty part; - no row of the part is expired yet - the part is cloned (its data files hardlinked) and only its stored TTL bounds are shifted; - otherwise - the part is rewritten exactly as before. The fast path is only taken when it is provably equivalent to the rewrite. It requires that the unconditional rows TTL (`TTL <expr>`) is the only TTL of the table, and that the old and the new TTL are the same date/time column shifted by constant fixed-length intervals, so that `new_ttl(row) - old_ttl(row)` is one constant for every row. Calendar `MONTH`/`YEAR` intervals, `DAY`/`WEEK` intervals in a time zone with daylight saving time, and row-dependent expressions are all rejected. The proof is redone for each part against the TTL expression (and time zone) that the part's stored timestamps were actually computed under, so a part that lags the table metadata, or was written by an older server, falls back to the regular rewrite rather than being shifted unsoundly. The same applies to the boundary cases of the stored timestamps themselves: a part containing a row whose TTL timestamp is exactly `1970-01-01 00:00:00` UTC (which ClickHouse treats as \"no TTL\"), a shift that would move some timestamp onto that value, and a part whose stored TTL is already fully expired (which the regular rewrite drops wholesale, even when the new TTL is longer) all take the regular rewrite. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `ALTER TABLE ... MODIFY TTL` no longer rewrites the table's data when the new TTL is the old one shifted by a constant amount of time (the same date/time column plus fixed-length intervals), which is the common case of extending or shortening the retention period. Fully expired parts are replaced with empty ones and the rest are cloned with only their stored TTL metadata shifted, which makes such an `ALTER` nearly instant. Cases where the shift is not provably constant - calendar month/year intervals, day/week intervals in a time zone with daylight saving time, or row-dependent TTL expressions - fall back to the regular rewrite. `ALTER TABLE ... MATERIALIZE TTL` benefits from the same optimization. <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> > Information about CI checks: https://clickhouse.com/docs/en/development/continuous-integration/ <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration Tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASAN - [ ] <!---ci_include_tsan--> All with TSAN - [ ] <!---ci_include_analyzer--> All with Analyzer - [ ] <!---ci_include_azure --> All with Azure - [ ] <!---ci_include_KEYWORD--> Add your option here #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration Tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASAN - [ ] <!---ci_exclude_tsan--> All with TSAN - [ ] <!---ci_exclude_msan--> All with MSAN - [ ] <!---ci_exclude_ubsan--> All with UBSAN - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 - [ ] <!---ci_exclude_KEYWORD--> Add your option here #### Extra options: - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 <details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/63383",
          "createdAt": "2024-05-05T15:40:47Z",
          "updatedAt": "2026-08-13T14:31:57Z",
          "timestamp": "2026-08-13T14:31:57Z",
          "metrics": {
            "reactions": 1,
            "comments": 33
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "zhongyuankai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fd707cfab59d566769f5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114570",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114570",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix presentation URL in `clickhouse-git-import` help",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Documentation entry for user-facing changes The help message of `clickhouse-git-import` referenced the presentation at `https://presentations.clickhouse.com/matemarketing_2020/`. Replace it with `https://presentations.clickhouse.com/2020-matemarketing/`. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1328` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114570",
          "createdAt": "2026-08-13T02:12:26Z",
          "updatedAt": "2026-08-13T14:31:52Z",
          "timestamp": "2026-08-13T14:31:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e9b6d9b7afb612f55447",
        "signalId": "github:ClickHouse/ClickHouse:issue:109678",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:109678",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "SIGSEGV: data race on `AsynchronousBoundedReadBuffer::prefetch_future` (ParquetV3 prefetcher)",
          "text": "### Describe the bug Having difficulty reproducing it. # The crash `26.7.1.569`, thread from the Parquet prefetcher fast pool: ``` <Fatal> BaseDaemon: Address: 0x28. Access: read. Address not mapped to object. 3. pthread_mutex_lock 5. std::mutex::lock() 6. std::__assoc_state<DB::IAsynchronousReader::Result>::move() future:636 (future::get) 7. DB::AsynchronousBoundedReadBuffer::readBigAt() AsynchronousBoundedReadBuffer.cpp:486 8. DB::Parquet::Prefetcher::readSync() Prefetcher.cpp:102 9. DB::Parquet::Prefetcher::runTask() Prefetcher.cpp:523 ... DB::ThreadPoolCallbackRunnerFast::threadFunction() (query: SELECT * FROM `d4`.`t93` LIMIT 10000 INTO OUTFILE '/tmp/file.data' TRUNCATE FORMAT Null;) ``` `d4.t93` is an S3-backed data-lake table (`engine: iceberg`) read through `ParquetV3`. ## Root cause — confirmed by code (a real ClickHouse bug, unreported) `AsynchronousBoundedReadBuffer::readBigAt()` is a positioned read intended to be safe for concurrent use, but it also *consumes an in-flight sequential prefetch* with no synchronization (`AsynchronousBoundedReadBuffer.cpp:480-489`): ```cpp if (prefetch_future.valid()) { ... result = prefetch_future.get(); // line 486 prefetch_future = {}; // line 488 ... } ``` `Prefetcher::readSync()` calls `reader->readBigAt(...)` from multiple `runTask` threads on the fast pool **without a lock in `ReadMode::RandomRead`** (`Prefetcher.cpp:102`) — note the `SeekAndRead` branch *does* take `read_mutex`. So two prefetcher threads race on the shared `prefetch_future`: ``` A: prefetch_future.valid() -> true B: prefetch_future.get(); prefetch_future = {} A: prefetch_future.get() on the moved-from future -> __state_ == nullptr -> __assoc_state::move() locks &__mut_ at offset 0x28 of nullptr -> SIGSEGV at 0x28 ``` The `0x28` fault address is exactly the offset of the `std::mutex` inside a null `__assoc_state`. Enablers in the fuzzer run: `local_filesystem_read_method='pread_fake_async'` (routes even local reads through the async buffer), `input_format_parquet_use_offset_index=1` (RandomRead → `readBigAt`), and `ThreadFuzzer` active (frame 4), which widened the tiny `valid()`-vs-`get()` window. ## Suggested fix Make the prefetch consumption in `readBigAt()` thread-safe, or take `read_mutex` in the `ReadMode::RandomRead` branch of `Prefetcher::readSync()`. Cleanest: guard the `if (prefetch_future.valid())` block in `readBigAt()` (e.g. atomically claim the future under a mutex / `std::exchange`), since `readBigAt()` is contractually concurrent. ### How to reproduce No success in reproducing it a second time. `d4`.`t93` is an IcebergS3 table. ### Error message and/or stacktrace Stack trace: ``` [node0] 2026.07.07 14:53:56.692776 [ 2113 ] <Fatal> BaseDaemon: ######################################## [node0] 2026.07.07 14:53:56.692837 [ 2113 ] <Fatal> BaseDaemon: (version 26.7.1.569 (official build), build id: 5654F8064F99BC8E4D1CB60CB74EBEE5400C905D, git hash: 20e1dcae5f06d30435d11bfc051996307d1972e7) (from thread 3980) (query_id: 2093fbc7-1d33-4441-ba60-ae33b42628cf) (query: SELECT * FROM `d4`.`t93` LIMIT 10000 INTO OUTFILE '/tmp/file.data' TRUNCATE FORMAT Null;) Received signal Segmentation fault (11) [node0] 2026.07.07 14:53:56.692853 [ 2113 ] <Fatal> BaseDaemon: Address: 0x28. Access: read. Address not mapped to object. [node0] 2026.07.07 14:53:56.692865 [ 2113 ] <Fatal> BaseDaemon: Stack trace: 0x000071b47cb33ef5 0x000058f615feec6f 0x000058f62ebed889 0x000058f619d12721 0x000058f619e564fb 0x000058f6230c1852 0x000058f6230c4082 0x000058f6230c7e33 0x000058f6177ce96d 0x000058f6160a3d89 0x000058f6160acd8b 0x000058f6160a0fbc 0x000058f6160aa14e 0x000071b47cb30ac3 0x000071b47cbc28d0 [node0] 2026.07.07 14:53:56.692930 [ 2113 ] <Fatal> BaseDaemon: 3. pthread_mutex_lock@@GLIBC_2.2.5 @ 0x0000000000097ef5 [node0] 2026.07.07 14:53:56.866200 [ 2113 ] <Fatal> BaseDaemon: 4. src/Common/ThreadFuzzer.cpp:447:1: pthread_mutex_lock @ 0x00000000155c8c6f [node0] 2026.07.07 14:53:56.928475 [ 2113 ] <Fatal> BaseDaemon: 5.0. inlined from contrib/llvm-project/libcxx/include/__thread/support/pthread.h:95: std::__libcpp_mutex_lock[abi:sqe220101](pthread_mutex_t*) [node0] 2026.07.07 14:53:56.928498 [ 2113 ] <Fatal> BaseDaemon: 5. contrib/llvm-project/libcxx/src/mutex.cpp:30:10: std::mutex::lock() @ 0x000000002e1c7889 [node0] 2026.07.07 14:53:57.044961 [ 2113 ] <Fatal> BaseDaemon: 6.0. inlined from contrib/llvm-project/libcxx/include/__mutex/unique_lock.h:44: unique_lock [node0] 2026.07.07 14:53:57.044988 [ 2113 ] <Fatal> BaseDaemon: 6. contrib/llvm-project/libcxx/include/future:636:11: std::__assoc_state<DB::IAsynchronousReader::Result>::move() @ 0x00000000192ec721 [node0] 2026.07.07 14:53:57.132500 [ 2113 ] <Fatal> BaseDaemon: 7.0. inlined from contrib/llvm-project/libcxx/include/future:986: std::future<DB::IAsynchronousReader::Result>::get() [node0] 2026.07.07 14:53:57.132527 [ 2113 ] <Fatal> BaseDaemon: 7. src/Disks/IO/AsynchronousBoundedReadBuffer.cpp:486:15: DB::AsynchronousBoundedReadBuffer::readBigAt(char*, unsigned long, unsigned long, std::function<bool (unsigned long)> const&) const @ 0x00000000194304fb [node0] 2026.07.07 14:53:57.332448 [ 2113 ] <Fatal> BaseDaemon: 8. src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:102:29: DB::Parquet::Prefetcher::readSync(char*, unsigned long, unsigned long) @ 0x000000002269b852 [node0] 2026.07.07 14:53:57.359378 [ 2113 ] <Fatal> BaseDaemon: 9. src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:523:13: DB::Parquet::Prefetcher::runTask(DB::Parquet::Prefetcher::Task*) @ 0x000000002269e082 [node0] 2026.07.07 14:53:57.382614 [ 2113 ] <Fatal> BaseDaemon: 10.0. inlined from src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:424: operator() [node0] 2026.07.07 14:53:57.382649 [ 2113 ] <Fatal> BaseDaemon: 10.1. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>::type std::__invoke[abi:sqe220101]<DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>(DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&) [node0] 2026.07.07 14:53:57.382681 [ 2113 ] <Fatal> BaseDaemon: 10.2. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>(DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&) [node0] 2026.07.07 14:53:57.382696 [ 2113 ] <Fatal> BaseDaemon: 10.3. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&>(DB::Parquet::Prefetcher::scheduleTask(DB::Parquet::Prefetcher::Task*)::$_0&) [node0] 2026.07.07 14:53:57.382707 [ 2113 ] <Fatal> BaseDaemon: 10. contrib/llvm-project/libcxx/include/__functional/function.h:443:17: ? @ 0x00000000226a1e33 [node0] 2026.07.07 14:53:57.506449 [ 2113 ] <Fatal> BaseDaemon: 11.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.506480 [ 2113 ] <Fatal> BaseDaemon: 11.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.506489 [ 2113 ] <Fatal> BaseDaemon: 11. src/Common/threadPoolCallbackRunner.cpp:225:12: DB::ThreadPoolCallbackRunnerFast::threadFunction() @ 0x0000000016da896d [node0] 2026.07.07 14:53:57.592515 [ 2113 ] <Fatal> BaseDaemon: 12.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.592537 [ 2113 ] <Fatal> BaseDaemon: 12.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.592548 [ 2113 ] <Fatal> BaseDaemon: 12. src/Common/ThreadPool.cpp:1103:12: ThreadPoolImpl<ThreadFromGlobalPoolImpl<false, true>>::ThreadFromThreadPool::worker() @ 0x000000001567dd89 [node0] 2026.07.07 14:53:57.627238 [ 2113 ] <Fatal> BaseDaemon: 13.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.627266 [ 2113 ] <Fatal> BaseDaemon: 13.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.627274 [ 2113 ] <Fatal> BaseDaemon: 13.2. inlined from src/Common/ThreadPool.cpp:1293: operator() [node0] 2026.07.07 14:53:57.627294 [ 2113 ] <Fatal> BaseDaemon: 13.3. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:90: std::__invoke_result_impl<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>::type std::__invoke[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) [node0] 2026.07.07 14:53:57.627310 [ 2113 ] <Fatal> BaseDaemon: 13.4. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:350: void std::__invoke_void_return_wrapper<void, true>::__call[abi:sqe220101]<startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) [node0] 2026.07.07 14:53:57.627321 [ 2113 ] <Fatal> BaseDaemon: 13.5. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:356: void std::__invoke_r[abi:sqe220101]<void, startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&>(startThreadFromGlobalPool(std::shared_ptr<ThreadFromGlobalPoolState>, std::function<void ()>, unsigned long, unsigned long, bool, bool)::$_0&) [node0] 2026.07.07 14:53:57.627329 [ 2113 ] <Fatal> BaseDaemon: 13. contrib/llvm-project/libcxx/include/__functional/function.h:443:12: ? @ 0x0000000015686d8b [node0] 2026.07.07 14:53:57.638085 [ 2113 ] <Fatal> BaseDaemon: 14.0. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:502: ? [node0] 2026.07.07 14:53:57.638103 [ 2113 ] <Fatal> BaseDaemon: 14.1. inlined from contrib/llvm-project/libcxx/include/__functional/function.h:754: ? [node0] 2026.07.07 14:53:57.638130 [ 2113 ] <Fatal> BaseDaemon: 14. src/Common/ThreadPool.cpp:1113:12: ThreadPoolImpl<std::thread>::ThreadFromThreadPool::worker() @ 0x000000001567afbc [node0] 2026.07.07 14:53:57.658980 [ 2113 ] <Fatal> BaseDaemon: 15.0. inlined from contrib/llvm-project/libcxx/include/__type_traits/invoke.h:0: std::__invoke_result_impl<void, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>::type std::__invoke[abi:sqe220101]<void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>(void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*&&)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*&&) [node0] 2026.07.07 14:53:57.659015 [ 2113 ] <Fatal> BaseDaemon: 15.1. inlined from contrib/llvm-project/libcxx/include/__thread/thread.h:161: void std::__thread_execute[abi:sqe220101]<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*, 0ul, 1ul>(std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>&, std::__integer_sequence<unsigned long, 0ul, 1ul>) [node0] 2026.07.07 14:53:57.659030 [ 2113 ] <Fatal> BaseDaemon: 15. contrib/llvm-project/libcxx/include/__thread/thread.h:169: void* std::__thread_proxy[abi:sqe220101]<std::tuple<std::unique_ptr<std::__thread_struct, std::default_delete<std::__thread_struct>>, void (ThreadPoolImpl<std::thread>::ThreadFromThreadPool::*)(), ThreadPoolImpl<std::thread>::ThreadFromThreadPool*>>(void*) @ 0x000000001568414e [node0] 2026.07.07 14:53:57.659079 [ 2113 ] <Fatal> BaseDaemon: 16. start_thread @ 0x0000000000094ac3 [node0] 2026.07.07 14:53:57.659105 [ 2113 ] <Fatal> BaseDaemon: 17. __GI___clone3 @ 0x00000000001268d0 [node0] 2026.07.07 14:53:58.543009 [ 2113 ] <Fatal> BaseDaemon: Integrity check of the executable successfully passed (checksum: C14EA45234712E0579DCDA202C426E57) [node0] 2026.07.07 14:53:58.704024 [ 2113 ] <Fatal> BaseDaemon: ClickHouse version 26.7.1.569 is old and should be upgraded to the latest version. [node0] 2026.07.07 14:53:58.704285 [ 2113 ] <Fatal> BaseDaemon: Changed settings: min_compress_block_size = 32, use_strict_insert_block_limits = false, max_insert_block_size_bytes = 1024, min_insert_block_size_rows = 1824, min_insert_block_size_bytes = 6127879, min_insert_block_size_rows_for_materialized_views = 4096, min_insert_block_size_bytes_for_materialized_views = 0, min_external_table_block_size_bytes = 32, max_joined_block_size_rows = 0, min_joined_block_size_rows = 1, min_joined_block_size_bytes = 6511566, joined_block_split_single_row = false, parallel_non_joined_rows_processing = false, max_final_threads = 2, max_threads_for_indexes = 1, max_threads = 32, max_parsing_threads = 2, max_read_buffer_size = 2048, max_read_buffer_size_remote_fs = 4, use_hedged_requests = true, s3_truncate_on_insert = true, azure_truncate_on_insert = true, s3_create_new_file_on_insert = false, s3_skip_empty_files = false, s3_allow_parallel_part_upload = true, s3queue_enable_logging_to_s3queue_log = true, hdfs_create_new_file_on_insert = true, dictionary_validate_primary_key_type = false, distributed_background_insert_batch = true, optimize_move_to_prewhere = true, enable_multiple_prewhere_read_steps = true, move_primary_key_columns_to_end_of_prewhere = false, allow_reorder_prewhere_conditions = false, load_balancing = 'hostname_levenshtein_distance', allow_suspicious_low_cardinality_types = true, allow_suspicious_fixed_string_types = true, allow_suspicious_indices = true, allow_suspicious_ttl_expressions = true, allow_suspicious_variant_types = true, allow_suspicious_primary_key = true, allow_suspicious_types_in_group_by = true, allow_suspicious_types_in_order_by = true, variant_throw_on_type_mismatch = false, dynamic_throw_on_type_mismatch = true, compile_expressions = false, min_count_to_compile_expression = 1, min_count_to_compile_sort_description = 0, group_by_two_level_threshold_bytes = 1024, distributed_aggregation_memory_efficient = false, aggregation_memory_efficient_merge_threads = 32, enable_positional_arguments = true, enable_positional_arguments_for_projections = false, allow_nonconst_timezone_arguments = true, enable_time_time64_type = true, function_locate_has_mysql_compatible_argument_order = false, parallel_distributed_insert_select = 1, optimize_distributed_group_by_sharding_key = true, optimize_skip_unused_shards = false, allow_nondeterministic_optimize_skip_unused_shards = true, min_chunk_bytes_for_parallel_parsing = 498990, merge_tree_min_rows_for_concurrent_read = 16384, merge_tree_min_bytes_for_concurrent_read = 4, merge_tree_min_rows_for_seek = 1, merge_tree_max_bytes_to_use_cache = 4096, enable_automatic_decision_for_merging_across_partitions_for_final = true, split_parts_ranges_into_intersecting_and_non_intersecting_final = true, split_intersecting_parts_ranges_into_layers_final = false, defer_partition_pruning_after_final = false, optimize_min_equality_disjunction_chain_length = 0, optimize_min_inequality_conjunction_chain_length = 0, min_bytes_to_use_direct_io = 4, min_bytes_to_use_mmap_io = 1, use_lightweight_primary_key_index_analysis = true, use_partition_pruning = false, use_skip_indexes_if_final_exact_mode = false, use_skip_indexes_on_data_read = true, use_statistics_for_part_pruning = false, use_top_k_dynamic_filtering_for_variable_length_types = true, query_plan_max_limit_for_top_k_optimization = 0, per_part_index_stats = false, secondary_indices_enable_bulk_filtering = true, max_streams_to_max_threads_ratio = 0.74030601978302, log_queries = false, distributed_product_mode = 'allow', deduplicate_insert_select = 'enable_when_possible', insert_quorum_parallel = false, select_sequential_consistency = 1, update_sequential_consistency = true, update_parallel_mode = 'async', table_function_remote_max_addresses = 62, enable_http_compression = false, count_distinct_implementation = 'uniqExact', send_profile_events = true, http_wait_end_of_query = false, join_output_by_rowlist_perkey_rows_threshold = 16, query_plan_join_swap_table = false, query_plan_optimize_join_order_limit = 64, query_plan_optimize_join_order_randomize = 0, enable_join_transitive_predicates = true, preferred_block_size_bytes = 16384, parts_to_throw_insert = 32, number_of_mutations_to_delay = 32, number_of_mutations_to_throw = 2, ignore_on_cluster_for_replicated_udf_queries = true, insert_allow_materialized_columns = true, http_max_fields = 512, http_make_head_request = false, use_index_for_in_with_subqueries = false, analyze_index_with_space_filling_curves = false, allow_key_condition_coalesce_rewrite = true, empty_result_for_aggregation_by_constant_keys_on_empty_set = false, allow_distributed_ddl = true, allow_suspicious_codecs = true, opentelemetry_start_keeper_trace_probability = 0.9900000095367432, max_bytes_before_external_group_by = 8, max_bytes_ratio_before_external_group_by = 0.99, prefer_external_sort_block_bytes = 0, max_bytes_ratio_before_external_sort = 0.1, max_bytes_before_remerge_sort = 10485760, remerge_sort_lowered_memory_bytes_ratio = 0.10000000149011612, max_execution_time = 60., allow_fuzz_query_functions = true, rows_before_aggregation = true, cross_to_inner_join_rewrite = 2, cross_join_min_rows_to_compress = 8, default_max_bytes_in_join = 4096, temporary_files_codec = 'none', max_rows_to_transfer = 16384, max_reverse_dictionary_lookup_cache_size_bytes = 16, log_profile_events = true, log_query_views = true, enable_optimize_predicate_expression_to_final_subquery = true, allow_push_predicate_when_subquery_contains_with = true, allow_custom_error_code_in_throwif = true, prefer_localhost_replica = true, allow_ddl = true, parallel_view_processing = false, enable_unaligned_array_join = false, read_in_order_use_virtual_row_per_block = true, optimize_aggregation_in_order = false, optimize_aggregation_in_order_limit = false, read_in_order_use_buffering = true, cancel_http_readonly_queries_on_client_close = false, allow_hyperscan = false, reject_expensive_hyperscan_regexps = true, allow_introspection_functions = true, allow_execute_multiif_columnar = true, parsedatetime_e_requires_space_padding = true, functions_h3_default_if_invalid = true, check_query_single_value_result = true, allow_drop_detached = true, allow_replace_partition_from_empty_source = true, dynamic_disk_allow_from_zk = false, max_parts_to_move = 0, max_partition_size_to_drop = 1, glob_expansion_max_elements = 8, show_table_uuid_in_table_create_query_if_not_nil = true, enable_scalar_subquery_optimization = true, optimize_trivial_count_query = false, optimize_trivial_approximate_count_query = true, optimize_trivial_group_by_limit_query = false, optimize_count_from_files = true, use_cache_for_count_from_files = true, optimize_respect_aliases = true, enable_lightweight_delete = true, lightweight_delete_mode = 'lightweight_update', optimize_normalize_count_variants = false, optimize_injective_functions_inside_uniq = false, optimize_arithmetic_operations_in_aggregate_functions = true, optimize_redundant_functions_in_order_by = true, optimize_if_transform_strings_to_enum = true, optimize_substitute_columns = false, normalize_function_names = true, allow_materialized_view_with_bad_select = true, materialized_views_squash_parallel_inserts = true, use_compact_format_in_distributed_parts_names = false, validate_polygons = true, recursive_cte_max_steps_in_type_inference = 10, allow_settings_after_format_in_insert = true, allow_nondeterministic_mutations = true, cast_keep_nullable = true, allow_non_metadata_alters = true, enable_materialized_cte = true, flatten_nested = false, optimize_skip_merged_partitions = false, optimize_use_projections = false, optimize_use_implicit_projections = false, optimize_use_projection_filtering = false, insert_null_as_default = true, enable_lightweight_update = true, apply_patch_parts = true, apply_patch_parts_join_cache_buckets = 64, mutations_execute_subqueries_on_initiator = true, mutations_max_literal_size_to_replace = 0, delta_lake_log_metadata = true, delta_lake_reload_schema_for_consistency = false, iceberg_metadata_log_level = 'metadata', iceberg_data_file_size_upper_threshold_compaction = 4, iceberg_compaction_delay_bias = 2652., iceberg_compaction_data_cleanup = 4642., use_query_cache = true, enable_writes_to_query_cache = false, enable_reads_from_query_cache = false, query_cache_for_subqueries = true, query_cache_nondeterministic_function_handling = 'ignore', query_cache_system_table_handling = 'ignore', query_cache_max_size_in_bytes = 16, query_cache_max_entries = 1024, query_cache_min_query_runs = 5481, query_cache_compress_entries = false, query_cache_squash_partial_results = true, use_query_condition_cache = false, optimize_rewrite_aggregate_function_with_if = false, optimize_rewrite_has_to_in = true, optimize_dictget_tuple_element = true, use_hash_table_stats_for_join_reordering = true, allow_experimental_kafka_offsets_storage_in_keeper = true, enable_software_prefetch_in_aggregation = true, allow_aggregate_partitions_independently = false, force_aggregate_partitions_independently = true, allow_limit_by_partitions_independently = false, min_hit_rate_to_use_consecutive_keys_optimization = 0.5, engine_file_empty_if_not_exists = false, engine_file_truncate_on_insert = true, enable_url_encoding = true, database_replicated_allow_replicated_engine_arguments = 1, database_replicated_allow_heavy_create = true, cloud_mode_engine = 0, external_storage_max_read_rows = 16, external_storage_max_read_bytes = 2048, allow_experimental_correlated_subqueries = true, max_streams_for_union_step = 32, optimize_aggregators_of_group_by_keys = false, optimize_injective_functions_in_group_by = true, legacy_column_name_of_tuple_literal = true, query_plan_enable_optimizations = true, query_plan_lift_up_array_join = false, query_plan_push_down_limit = true, query_plan_top_k_through_join = false, query_plan_split_filter = true, query_plan_merge_expressions = true, query_plan_filter_push_down = false, query_plan_convert_outer_join_to_inner_join = false, query_plan_convert_any_join_to_semi_or_anti_join = true, query_plan_merge_filter_into_join_condition = true, optimize_prewhere_after_pushdown = false, query_plan_execute_functions_after_sorting = false, query_plan_reuse_storage_ordering_for_window_functions = true, query_plan_lift_up_union = true, query_plan_read_in_order = false, query_plan_read_in_order_through_join = false, query_plan_aggregation_in_order = false, query_plan_optimize_lazy_final = false, max_rows_for_lazy_final = 4, max_bytes_for_lazy_final = 9963842, min_filtered_ratio_for_lazy_final = 0.10000000149011612, enable_lazy_columns_replication = false, enable_software_prefetch_in_join = false, correlated_subqueries_substitute_equivalent_expressions = true, correlated_subqueries_use_in_memory_buffer = true, function_range_max_elements_in_block = 0, function_base58_max_input_size = 4, local_filesystem_read_method = 'pread_fake_async', local_filesystem_read_prefetch = true, remote_filesystem_read_prefetch = true, merge_tree_min_rows_for_concurrent_read_for_remote_filesystem = 8, merge_tree_min_bytes_for_concurrent_read_for_remote_filesystem = 2, remote_read_min_bytes_for_seek = 0, merge_tree_min_bytes_per_task_for_remote_reading = 1, merge_tree_determine_task_size_by_prewhere_columns = false, merge_tree_min_read_task_size = 8192, async_insert_max_query_number = 2, cluster_function_process_archive_on_multiple_nodes = false, max_streams_for_files_processing_in_cluster_functions = 32, enable_filesystem_cache = false, filesystem_cache_name = 'fcache0', enable_filesystem_cache_on_write_operations = false, filesystem_cache_max_download_size = 3724327, throw_on_error_from_cache_on_write_operations = false, filesystem_cache_segments_batch_size = 50, filesystem_cache_boundary_alignment = 4, use_page_cache_with_distributed_cache = true, use_page_cache_for_object_storage = true, read_from_page_cache_if_exists_otherwise_bypass_cache = false, page_cache_inject_eviction = true, page_cache_block_size = 2048, page_cache_max_coalesced_bytes = 4096, allow_prefetched_read_pool_for_local_filesystem = true, prefetch_buffer_size = 0, filesystem_prefetch_step_bytes = 10485760, allow_calculating_subcolumns_sizes_for_merge_tree_reading = true, check_table_dependencies = true, check_named_collection_dependencies = false, allow_unrestricted_reads_from_keeper = true, allow_rank_dense_rank_arguments = false, schema_inference_use_cache_for_file = false, schema_inference_cache_require_modification_time_for_url = false, read_through_distributed_cache = false, distributed_cache_throw_on_error = true, distributed_cache_read_request_max_tries = 65536, distributed_cache_alignment = 8688762, distributed_cache_min_bytes_for_seek = 10485760, write_through_distributed_cache_buffer_size = 0, table_engine_read_through_distributed_cache = true, read_from_distributed_cache_if_exists_otherwise_bypass_cache = false, distributed_cache_registry_show_certificate_and_signature = false, filesystem_cache_enable_background_download_during_fetch = false, parallelize_output_from_storages = false, multiple_joins_try_to_keep_original_names = true, keeper_max_retries = 15, optimize_uniq_to_count = true, enable_order_by_all = true, allow_dynamic_type_in_join_keys = true, cast_string_to_variant_use_inference = false, enable_blob_storage_log_for_read_operations = true, allow_create_index_without_type = true, allow_named_collection_override_by_default = true, use_async_executor_for_materialized_views = true, short_circuit_function_evaluation_for_nulls_threshold = 1., allow_experimental_geo_types_in_iceberg = true, show_data_lake_catalogs_in_system_tables = true, delta_lake_throw_on_engine_predicate_error = true, delta_lake_insert_max_bytes_in_data_file = 4262856, allow_experimental_delta_lake_writes = true, use_iceberg_partition_pruning = true, extract_key_value_pairs_max_pairs_per_row = 0, allow_experimental_parallel_reading_from_replicas = 0, automatic_parallel_replicas_mode = 2, cluster_for_parallel_replicas = 'cluster0', parallel_replicas_allow_in_with_subquery = true, parallel_replicas_for_non_replicated_merge_tree = false, parallel_replicas_prefer_local_join = false, parallel_replicas_index_analysis_only_on_coordinator = false, parallel_replicas_support_projection = true, parallel_replicas_only_with_analyzer = false, parallel_replicas_allow_materialized_views = true, parallel_replicas_allow_view_over_mergetree = true, distributed_index_analysis = false, allow_experimental_database_iceberg = true, allow_experimental_database_unity_catalog = true, allow_experimental_database_glue_catalog = true, allow_experimental_analyzer = true, max_limit_for_vector_search_queries = 7447, vector_search_with_rescoring = false, use_join_disjunctions_push_down = true, shared_merge_tree_sync_parts_on_partition_operations = false, allow_general_join_planning = true, cluster_table_function_buckets_batch_size = 32, validate_enum_literals_in_operators = true, use_hive_partitioning = false, s3_uri_style = 'virtual_hosted', iceberg_insert_max_partitions = 2, min_outstreams_per_resize_after_split = 7655, enable_add_distinct_to_in_subqueries = true, jemalloc_profile_text_collapsed_use_count = true, allow_experimental_nullable_tuple_type = true, archive_adaptive_buffer_max_size_bytes = 8, max_bytes_before_external_join = 4236297, enable_join_fixed_hash_table_conversion = true, query_plan_min_columns_for_join_lazy_indexing = 5, allow_experimental_materialized_postgresql_table = true, allow_experimental_funnel_functions = true, allow_experimental_nlp_functions = true, allow_experimental_hash_functions = true, allow_experimental_time_series_table = true, allow_experimental_unique_key = true, allow_experimental_codecs = true, wait_changes_become_visible_after_commit_mode = 'wait_unknown', grace_hash_join_max_buckets = 1024, join_to_sort_minimum_perkey_rows = 2, allow_experimental_join_right_table_sorting = true, allow_experimental_json_lazy_type_hints = true, allow_statistics_optimize = true, use_statistics = false, allow_statistics = true, enable_full_text_index = true, query_plan_text_index_add_hint = true, use_text_index_like_evaluation_by_dictionary_scan = false, text_index_like_max_postings_to_read = 286, use_text_index_header_cache = false, text_index_lazy_intersection_density_threshold = 1., allow_experimental_window_view = true, allow_experimental_database_materialized_postgresql = true, allow_nullable_tuple_in_extracted_subcolumns = true, allow_experimental_database_hms_catalog = true, allow_experimental_kusto_dialect = true, allow_experimental_prql_dialect = true, allow_experimental_polyglot_dialect = true, allow_experimental_delta_kernel_rs = true, allow_insert_into_iceberg = true, allow_experimental_iceberg_compaction = true, allow_iceberg_remove_orphan_files = true, allow_experimental_expire_snapshots = true, write_full_path_in_iceberg_metadata = true, iceberg_metadata_compression_method = 'deflate', distributed_plan_force_exchange_kind = 'Streaming', distributed_plan_prefer_replicas_over_workers = true, allow_experimental_ytsaurus_table_engine = true, allow_experimental_ytsaurus_table_function = true, allow_experimental_ytsaurus_dictionary_source = true, enable_join_runtime_filters = true, join_runtime_bloom_filter_bytes = 2, join_runtime_bloom_filter_hash_functions = 1, join_runtime_filter_blocks_to_skip_before_reenabling = 16384, rewrite_in_to_join = false, allow_experimental_time_series_aggregate_functions = true, allow_experimental_paimon_storage_engine = true, use_paimon_partition_pruning = true, allow_experimental_object_storage_queue_hive_partitioning = true, query_plan_optimize_join_order_algorithm = 'dpsize', allow_experimental_database_paimon_rest_catalog = true, allow_experimental_ai_functions = true, allow_experimental_query_deduplication = true, update_insert_deduplication_token_in_dependent_materialized_views = true, allow_experimental_alias_table_engine = true, partial_merge_join_optimizations = 0, allow_not_comparable_types_in_order_by = true, allow_not_comparable_types_in_comparison_functions = true, enable_zstd_qat_codec = true, enable_deflate_qpl_codec = true, throw_if_deduplication_in_dependent_materialized_views_enabled_with_async_insert = false, async_insert_threads = 2, distributed_cache_read_alignment = 4, output_format_parallel_formatting = false, input_format_null_as_default = false, input_format_parquet_preserve_order = false, input_format_parquet_filter_push_down = false, input_format_parquet_enable_json_parsing = true, input_format_parquet_memory_low_watermark = 1024, input_format_parquet_memory_high_watermark = 1024, input_format_parquet_use_offset_index = true, input_format_parquet_local_time_as_utc = false, input_format_allow_seeks = false, input_format_orc_use_fast_decoder = true, input_format_orc_filter_push_down = true, input_format_orc_dictionary_as_low_cardinality = false, input_format_parquet_local_file_min_bytes_for_seek = 5762911, input_format_parquet_enable_row_group_prefetch = false, input_format_csv_use_best_effort_in_schema_inference = true, input_format_parquet_prefer_block_bytes = 4096, input_format_capn_proto_skip_fields_with_unsupported_types_in_schema_inference = true, schema_inference_make_columns_nullable = 3, schema_inference_make_json_columns_nullable = true, input_format_json_read_bools_as_numbers = true, input_format_json_read_bools_as_strings = false, input_format_json_try_infer_numbers_from_strings = true, input_format_json_infer_incomplete_types_as_strings = false, input_format_json_named_tuples_as_objects = true, input_format_json_throw_on_bad_escape_sequence = true, type_json_skip_duplicated_paths = true, input_format_try_infer_datetimes = true, input_format_try_infer_datetimes_only_datetime64 = true, input_format_protobuf_flatten_google_wrappers = false, input_format_tsv_skip_trailing_empty_lines = false, output_format_native_use_flattened_dynamic_and_json_serialization = false, input_format_values_accurate_types_of_literals = false, input_format_avro_null_as_default = false, input_format_binary_read_json_as_string = false, output_format_binary_write_json_as_string = true, output_format_json_quote_64bit_floats = true, output_format_json_array_of_rows = false, output_format_pretty_max_rows = 500, output_format_pretty_glue_chunks = 1, output_format_parquet_row_group_size = 8, output_format_parquet_row_group_size_bytes = 2872120, output_format_parquet_batch_size = 1571, output_format_parquet_write_bloom_filter = false, output_format_avro_codec = 'null', output_format_avro_sync_interval = 10000000, input_format_geojson_unsupported_geometry_handling = 'throw', output_format_pretty_multiline_fields = true, output_format_pretty_named_tuples_as_json = false, output_format_arrow_use_signed_indexes_for_dictionary = false, output_format_arrow_use_64_bit_indexes_for_dictionary = true, format_capn_proto_use_autogenerated_schema = true, output_format_sql_insert_include_column_names = true, input_format_bson_skip_fields_with_unsupported_types_in_schema_inference = true, format_display_secrets_in_show_and_select = true, validate_experimental_and_suspicious_types_inside_nested_types = false, show_create_query_identifier_quoting_rule = 'user_display', input_format_parquet_allow_geoparquet_parser = true ``` <!-- ch-version-info:start --> ### Version info - Resolved by: #112573 - Merged into: `26.8.1.1240` (included in `26.8` and later) - Backported to: `26.7.4.27` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/109678",
          "createdAt": "2026-07-07T16:18:13Z",
          "updatedAt": "2026-08-13T14:31:48Z",
          "timestamp": "2026-08-13T14:31:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "crash",
            "fuzz",
            "comp-parquet-reader-v3"
          ],
          "author": "PedroTadim",
          "state": "closed",
          "assignees": [
            "grantholly-clickhouse"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:fbdd783ebcd5ba02c84e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114608",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114608",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make the 04780 index-analysis allocation oracle the minimum of several runs",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/104948 `04780_json_subcolumn_index_match_not_quadratic` compares `MemoryAllocatedWithoutCheckBytes` of a dotted-constant `EXPLAIN indexes = 1` against a no-dots control with a 150% threshold, measuring each arm with a single query. A single run is not a stable oracle: whichever query happens to be the first to touch a cache or spin up a thread pool absorbs a transient multi-megabyte allocation, and the dotted arm always runs first in the loop, so such a one-off lands on it and inflates the ratio arbitrarily. Locally the effect is easy to see: the first query after a server start reports 8–22 MB in this counter against a ~2.8 MB steady state for the identical query. The test failed this way on at least 6 unrelated PRs since 2026-08-12 (e.g. `Stateless tests (amd_asan_ubsan, distributed plan, parallel)` on https://github.com/ClickHouse/ClickHouse/pull/104948 at commit 81300c20: `longidx index analysis over a constant with 100000 dots allocated 4216780 bytes, more than 150% of the no-dots control (22224 bytes)` — reruns passed; CIDB shows the same failure on #106011, #108522, #114475, #96130, #114476). Measure each arm three times and take the minimum: a genuine quadratic regression is deterministic and shows up in every run, so the oracle keeps discriminating (the minimum can only remove one-sided transient noise), while a one-off transient can no longer fail the test. Verified against the current master binary: the modified test passes repeatedly. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1329` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114608",
          "createdAt": "2026-08-13T09:13:17Z",
          "updatedAt": "2026-08-13T14:33:03Z",
          "timestamp": "2026-08-13T14:33:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:84810c3708fef9ef1a5b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112573",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix async bounded read buffer readbigat race",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix AsynchronousBoundedReadBuffer's readBigAt data race. Closes https://github.com/ClickHouse/ClickHouse/issues/109678. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1240` (included in `26.8` and later) - Backported to: `26.7.4.27` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112573",
          "createdAt": "2026-07-30T10:54:24Z",
          "updatedAt": "2026-08-13T14:31:46Z",
          "timestamp": "2026-08-13T14:31:46Z",
          "metrics": {
            "reactions": 1,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v26.4-must-backport"
          ],
          "author": "kssenii",
          "state": "closed",
          "assignees": [
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b4e0c21cf8c938bacb05",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114278",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114278",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `XDG_CACHE_HOME` being read from `XDG_STATE_HOME`",
          "text": "A copy-paste slip in `XDGBaseDirectories`: `ENV_XDG_CACHE_HOME` was defined as `\"XDG_STATE_HOME\"`, so `XDGBaseDirectories::getCacheHome` ignored the `XDG_CACHE_HOME` environment variable and obeyed `XDG_STATE_HOME` instead. `getCacheHome` currently has no in-tree callers, so this does not change observable behavior yet; it fixes the helper before callers appear. Found while reviewing https://github.com/ClickHouse/ClickHouse/pull/112824. Related: https://github.com/ClickHouse/ClickHouse/pull/112824 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114278",
          "createdAt": "2026-08-11T06:33:24Z",
          "updatedAt": "2026-08-13T14:31:43Z",
          "timestamp": "2026-08-13T14:31:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f3bb3b77c1e0ceeb3c93",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112152",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112152",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the function name of a stack frame attributed to a libc++ `__functional` header",
          "text": "Caused by: https://github.com/ClickHouse/ClickHouse/pull/57201 Related: https://github.com/ClickHouse/ClickHouse/pull/100419 `collapseDemangledNames` replaced a stack frame's function name with `?` whenever the frame's source location was a file in a directory ending in `functional`. The intent was to hide `std::function` plumbing frames, whose demangled names spell out the whole captured lambda type and say nothing the neighbouring frames do not already say. But the file of a frame is the source line the *instruction* maps to, which is not necessarily where the enclosing function is defined. An ordinary function can have individual instructions attributed to a libc++ `__functional` header - an inlined `std::function` operation, or compiler-generated code reported with line 0 - and then its name was dropped too, which loses the only useful part of the frame. This PR requires the symbol to name the `std::function` plumbing as well: the type-erasing wrappers (`std::__function::__func`, `__value_func`, `__policy_func`, ...) and the type erasure of `std::function` itself - its `operator()`, its copy, move and callable-taking constructors and assignment operators, and its destructor. The ordinary members of `std::function` (`swap`, `target_type`, `operator bool`, the `nullptr` reset `operator=(std::nullptr_t)`, the empty-constructing `function()` / `function(std::nullptr_t)`, ...) do work of their own, so they keep their names too. Those are still displayed as `?`, exactly as before; every other frame keeps its name - not only an ordinary `DB` function, but also a meaningful libc++ symbol that happens to live in a `__functional` header, such as `std::hash<String>::operator()` from `__functional/hash.h`, or the generic invocation helpers `std::invoke` / `std::__invoke` / `std::mem_fn`, which are not `std::function`-specific and whose frames can name the callable they dispatch to. This is much more likely in a build with ThinLTO enabled, where `std::function` calls are inlined across translation units, so it went unnoticed for a long time: `amd_cfi` is the only integration-test build with ThinLTO on, and its `test_crash_log` failure is what surfaced it. In that build, the frame that actually terminated the server was displayed as ``` 8. contrib/llvm-project/libcxx/include/__functional/function.h:0:7: ? @ 0x1c01f41f ``` both in the fatal log and in `trace_full` in `system.crash_log`, where `0x1c01f41f` is inside `DB::executeQuery` (the symbol resolves correctly - only the display suppressed it). Official release builds also use ThinLTO, so the same frames were being anonymised for users. `collapseDemangledNames` becomes a static member of `StackTrace` so that the heuristic is covered by a unit test in every build, rather than only by the weekly ThinLTO job. Fixes `test_crash_log/test.py::test_crash_log_extra_fields[terminate_with_exception-trace_full]` and `[terminate_with_std_exception-trace_full]`, seen in https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=b05161aa67d75ad84151a392f3e87156efd42842&name_0=WeeklyCFI&name_1=Integration%20tests%20%28amd_cfi%2C%202%2F4%29 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a stack frame being displayed as `?` instead of its function name in fatal log messages and in the `trace_full` column of `system.crash_log`. It affected frames of ordinary functions that have code attributed to a libc++ `__functional` header, which is common in release builds. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1334` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112152",
          "createdAt": "2026-07-27T17:43:33Z",
          "updatedAt": "2026-08-13T14:31:40Z",
          "timestamp": "2026-08-13T14:31:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 16
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f187252e5b5658ca1290",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114539",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114539",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enable the query condition cache for `ORDER BY ... LIMIT n` queries by default",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/111492 Related: https://github.com/ClickHouse/ClickHouse/pull/104478 Related: https://github.com/ClickHouse/ClickHouse/pull/110507 `use_query_condition_cache_for_top_k` was introduced by #111492 and defaulted to `false` as a precaution while the soundness of the query condition cache entries written by a TopK (`ORDER BY <column> LIMIT n`) read was being established. Such entries are partitioned by the TopK plan parameters and by a snapshot of the set of parts read, so only a query with the same plan over the same parts reuses them. The gate is no longer needed, so the default becomes `true` and `ORDER BY ... LIMIT n` queries use the query condition cache again. This effectively reverts #111492 by flipping its setting, rather than by removing it: the setting and every gating point it drives are kept, so the cache can still be kept out of TopK reads with `use_query_condition_cache_for_top_k = 0`. The tests that cover that configuration are kept too — they already pinned the setting explicitly rather than relying on the default — while the tests of the feature itself no longer have to enable it. `04628_query_condition_cache_topk_default_off` is renamed to `04628_query_condition_cache_topk_gate_off` since it no longer describes the default. The settings-history entry keeps `previous_value = false`, because the gate was backported to 26.7. `compatibility` with 26.7 or earlier therefore still turns the query condition cache off for TopK reads, while `compatibility = '26.8'` keeps it on; `04631_query_condition_cache_topk_compatibility` pins this. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The [query condition cache](/operations/query-condition-cache) is now enabled by default for queries that use the `ORDER BY <column> LIMIT n` (TopK) optimization. It can be turned off again with the setting `use_query_condition_cache_for_top_k`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114539",
          "createdAt": "2026-08-12T20:09:36Z",
          "updatedAt": "2026-08-13T14:31:39Z",
          "timestamp": "2026-08-13T14:31:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-performance",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f13f75d504f267590991",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114090",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114090",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a regression test for duplicate parallel replicas announcements from a self-matching merge() child",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Add a regression test for the parallel replicas coordinator guard trips (`Initiator received more initial requests than there are replicas: replica_num=1`, `Duplicate announcement received for replica number N`) that fired when a `merge()` table function's regex matched the table the outer query reads. #112849 fixed this in code — the `enable_parallel_replicas` clear in `ReadFromMerge::createChildrenPlans` — but its test `04665_merge_table_parallel_replicas_child_plan` never reaches the announcement guards on the pre-fix code: its fixture aborts earlier with `Cannot serialize FutureSetFromSubquery with no query plan`, and it uses the default index granularity, at which the initiator's replica claims every mark range and the followers are cancelled before they announce. So if the `enable_parallel_replicas` clear were ever narrowed (the way the `make_distributed_plan` clear was narrowed to `queryHasSubquerySets`), the signature would regress with `04665` still green. The fixture is the one @groeneai verified against three binaries (see https://github.com/ClickHouse/ClickHouse/pull/102192#issuecomment-5232296945): the outer read and the `merge()` child read of the same table derive the identical `stream_id`, so without the clear one follower builds several read pools under one `replica_num` and announces more than once into the same coordinator. A small `index_granularity` gives the followers enough marks to survive to the announcement; the committed fixture uses 16384 rows with `index_granularity = 16` (~1024 marks). Heavier shapes (500000/128 and 62500/16, both ~3900 marks) exceeded the 180s per-test limit of the flaky check, which runs 50 copies concurrently on a debug build: `system.query_log` from the failed run shows the `parallel_replicas_local_plan = 0` SELECT stalling up to 333s wall at 12.5s CPU with 834 thread-seconds in `NetworkReceiveElapsedMicroseconds` and only 9 coordinator round-trips, followers idle after finishing their physical read in seconds — the round-trips of that mode degrade sharply on an oversubscribed server, so the mark count is kept as low as the reproduction allows (at ~512 marks the local-plan mode stops reproducing, so ~1024 keeps a 2x margin). Verified locally on a 3-replica server: reddens in both `parallel_replicas_local_plan` modes 3/3 runs with the two `enable_parallel_replicas` clears in `StorageMerge.cpp` disabled, passes with them in place, also under 8 concurrent runs. The test asserts the successful result rather than a guard message, since the pre-fix failure surfaces as either of the two adjacent guard messages depending on follower interleaving. To make sure it cannot silently stop exercising the announcement path, it runs with `allow_experimental_parallel_reading_from_replicas = 2` (an unsupported-shape fallback to a plain local read is an error, not a silent success) and additionally asserts `ProfileEvents['ParallelReplicasHandleRequestMicroseconds'] > 0` on the initiator's `system.query_log` entry, like `04545_parallel_replicas_projection_short_circuit_unknown_stream.sql` — a follower must survive past the announcement into the coordinator's request path for it to fire, so the initiator claiming every range and cancelling the followers early fails the test instead of passing it (a coordinator-creation log line alone could not distinguish that). The regex is anchored to the table name so concurrent tests cannot leak into the `merge()`, and the table lives in `default` because a single-argument `merge()` resolves against the default database on each hop. Because `default` is shared by every concurrently running test, the test is a `.sh` so the table name can carry `$CLICKHOUSE_DATABASE`: the first revision used a fixed literal name, and two runs of the test then raced on `CREATE`/`DROP` of the same table (`UNKNOWN_TABLE` / `TABLE_ALREADY_EXISTS`) - the flaky check runs each new test 50 times with `--jobs nproc-1`, and the ordinary parallel jobs repeat newly modified tests. Related: https://github.com/ClickHouse/ClickHouse/pull/112849 Related: https://github.com/ClickHouse/ClickHouse/pull/110972",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114090",
          "createdAt": "2026-08-10T01:00:49Z",
          "updatedAt": "2026-08-13T14:31:28Z",
          "timestamp": "2026-08-13T14:31:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fab010de3ab0fe3456e5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114619",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114619",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112573 to 26.7: Fix async bounded read buffer readbigat race",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112573 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31690189837/job/94415478202) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.27` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114619",
          "createdAt": "2026-08-13T10:34:20Z",
          "updatedAt": "2026-08-13T14:31:07Z",
          "timestamp": "2026-08-13T14:31:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll",
          "state": "closed",
          "assignees": [
            "kssenii",
            "arsenmuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:71cb182a321af9e6d9cc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114492",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114492",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113484 to 26.6: Push down plan level constants from joins",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113484 Cherry-pick pull-request https://github.com/ClickHouse/ClickHouse/pull/114488 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31605886263/job/94144581099) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.33` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114492",
          "createdAt": "2026-08-12T14:35:23Z",
          "updatedAt": "2026-08-13T14:31:05Z",
          "timestamp": "2026-08-13T14:31:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse",
          "state": "closed",
          "assignees": [
            "diegomestre2",
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7ee550c13da5219bb82e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114262",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114262",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Lazy materialization for reading local Parquet files (`file` / `File`)",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/110970 Follow-up to the lazy materialization for object storage (#110970), as discussed in https://github.com/ClickHouse/ClickHouse/pull/110970#issuecomment-5247858385: implement it for plain local Parquet files read through `StorageFile` — the `file` table function and the `File` table engine. For `ORDER BY ... LIMIT n` queries, the columns that are not needed for sorting and filtering are read only for the `n` rows that survive the `LIMIT`. The format side (row-selective Parquet reads via `FormatFilterInfo::rows_to_read`) and the plan split (`JoinLazyColumnsStep`, `LazyMaterializingTransform`) from #110970 are reused as-is; this PR adds the `StorageFile` counterpart of the two branches: - The main pass appends a `__global_row_index` column (file index in a per-query `LazyFileRegistry` + physical row numbers from `ChunkInfoRowNumbers`) in `StorageFileSource::generate`. - The lazy branch (`LazilyReadFromFile` → `LazyReadFromFileSource` → `StorageFileLazyRowsSource`) reopens only the surviving files (at most `LIMIT n` of them) with a per-file set of rows to read and `parquet.preserve_order`. - The storage-agnostic column-split logic of `ReadFromObjectStorageStep::keepOnlyRequiredColumnsAndCreateLazyReadStep` (which columns can be deferred: `DEFAULT` expression dependencies, PREWHERE inputs, hive partition and virtual columns) is extracted into the shared `splitLazilyReadColumnsFromFormatInfo` and reused by both storages. **Generation safety.** POSIX has no conditional read, so the reread cannot be pinned the way `If-Match` pins it on S3. Instead the reread fails close with the new `FILE_CHANGED_DURING_READ` error when the file's generation token — sub-second mtime + inode + size, reusing `computeFileCacheVersionToken` from the query condition cache integration — no longer matches the one captured when the main pass opened the file. The token is validated at file registration time in the main pass, and both before and right after the reopen in the lazy pass. Replace-by-rename (the common atomic-update pattern) is always caught since it changes the inode; an in-place rewrite is caught up to the filesystem timestamp tick (and already tears a single-pass read today). **Gates** (`ReadFromFile::canUseLazyMaterialization`): Parquet format only, no file descriptor reads (stdin cannot be reopened), no archive entries, no `distributed_processing`, no `--rename_files_after_processing`, uncompressed files only. Controlled by the new setting `query_plan_optimize_lazy_materialization_for_file` (enabled by default), on top of `query_plan_optimize_lazy_materialization`. **Bug fix for #110970 shared via the helper:** deferring a requested subcolumn (e.g. of a `JSON` column) left the lazy branch's format header empty and the query failed with `Not found column or subcolumn ... in block`, because the format header contains the parent column of a requested subcolumn while the split filtered it by subcolumn names. The split now maps requested columns to their storage-level names — the parent of a deferred subcolumn moves to the lazy branch, and stays in the main branch as well when a sort key still needs another subcolumn of it. Covered for both storages by the new tests. On a 1 GB local Parquet file (2M rows × 105 columns equivalent shape: two 300-byte strings), `SELECT * ... ORDER BY k DESC LIMIT 5` runs 3× faster (0.15 s vs 0.45 s, page-cache warm), with results verified identical with the optimization on and off. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Lazy materialization for `ORDER BY ... LIMIT n` queries (#110970) now also applies to local Parquet files read with the `file` table function and the `File` table engine: the columns that are not needed for sorting and filtering are read only for the `n` rows that survive the `LIMIT`. The second read of a surviving file fails close with the new `FILE_CHANGED_DURING_READ` error if the file was modified between the two passes. Controlled by the new setting `query_plan_optimize_lazy_materialization_for_file` (enabled by default). Also fixes lazy materialization for object storage failing with `Not found column or subcolumn ... in block` when a requested subcolumn (e.g. of a `JSON` column) is deferred.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114262",
          "createdAt": "2026-08-11T03:20:36Z",
          "updatedAt": "2026-08-13T14:30:55Z",
          "timestamp": "2026-08-13T14:30:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:29efa3876c5212caf716",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113912",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113912",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix UB in avg over Date/Time types at the Int64 boundary",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fix undefined behavior and wrong results in `avg` over `Date`/`DateTime`/`DateTime64`/`Time`/`Time64`: the average is now computed exactly in integer space, fixing both the `Int64`-boundary overflow (UB, wrong result on x86) and `Float64` precision loss above 2^53 (visible at nanosecond scale). `avgResultToValue` cast the rounded `Float64` average straight to the result's native integer type. The exact average always lies within the range of the inputs, but the `Float64` computation is inexact: for ticks near the bounds of `Int64` it can land on 2^63 exactly, and the cast is undefined behavior. UBSan caught it in a stress run as `9.22337e+18 is outside the range of representable values of type 'long'` at `AggregateFunctionAvg.h:39` (`avg` over `DateTime64`). On x86 the cast wraps to `INT64_MIN`, turning the average of values near the upper bound of `DateTime64` into `1677-09-21 00:12:43.145224192`; on AArch64 `fcvtzs` saturates silently, hiding the problem. A saturating cast alone is not enough (review finding): `Float64` cannot distinguish the last 1024 ticks of `Int64`, so clamping off the rounded `Float64` would still corrupt valid values just inside the boundary, and more generally the `Float64` division loses precision for any tick count above 2^53. Instead, for integer-backed Date/Time result types the exact accumulated numerator is now divided by the denominator in integer space, rounding half to even (matching the `nearbyint` semantics of the previous path), with saturation only when the accumulated sum itself has overflowed. The `Float64` path keeps a saturating cast for non-exact numerators. This also fixes a pre-existing precision artifact: `avg` of `2020-01-01 00:00:00.000000000` and `2020-01-01 00:00:00.000000002` at scale 9 now returns `.000000001` instead of `.000000000` (reference of `03799_avg_date_time_types` updated). CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109453&sha=2218c5bc60fe31decb1e21d9545bd9c9f9fc7a56&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%29 Related: https://github.com/ClickHouse/ClickHouse/pull/109453 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1198` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113912",
          "createdAt": "2026-08-08T02:35:56Z",
          "updatedAt": "2026-08-13T14:24:56Z",
          "timestamp": "2026-08-13T14:24:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "pr-synced-to-cloud"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:3045cfd85f3dcaa97717",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:98498",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:98498",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert dangerous change in CreateUniqueArrayJoinAliasesVisitor",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Revert the dangerous change from https://github.com/ClickHouse/ClickHouse/pull/98376, which breaks the invariant of the query tree. ColumnNode is supposed to always have a valid source. cc @alexey-milovidov ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/98498",
          "createdAt": "2026-03-02T14:17:32Z",
          "updatedAt": "2026-08-13T14:23:26Z",
          "timestamp": "2026-08-13T14:23:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "close in a month if not active"
          ],
          "author": "novikd",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:8a897a090d7f12d808aa",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114629",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114629",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix DPsub join reordering silently dropping single-table ON-clause filters",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111898 The `dpsub` join-order algorithm silently dropped single-table filter and constant predicates that live in a `JOIN ... ON` clause (e.g. `t1.value = 'x'`), returning extra rows. `greedy` uses a different placement path and was unaffected. The predicates are placed in `collectJoinEdgesMask`. Two placement conditions there each silently dropped such a predicate: - `two_relations` required the whole join step to be exactly two relations, so the predicate was dropped whenever its relation was introduced against an already-multi-relation subplan (e.g. `t1` at the top of `t1 JOIN (t2 JOIN t3)`). - `fromLeft() || fromRight() || fromNone()` dropped the predicate for any single-table filter on a relation whose id is >= 2, because `fromLeft`/`fromRight` test relation ids 0 and 1 specifically (they describe the two inputs of a binary join step, not \"references a single relation\"). The predicate is now attached at the join that introduces its relation (the split whose one side is exactly that relation), and pure constants at the earliest two-relation join. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fixed `dpsub` join-order optimization (`query_plan_optimize_join_order_algorithm = 'dpsub'`) silently dropping single-table filter conditions from a `JOIN ... ON` clause, which could return extra rows.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114629",
          "createdAt": "2026-08-13T12:38:37Z",
          "updatedAt": "2026-08-13T14:22:26Z",
          "timestamp": "2026-08-13T14:22:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "fkastrati",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:04642e75d36ee3f8edd4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114647",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114647",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add LIKE/NOT LIKE/ILIKE filtering to SHOW access entity statements",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111692 Implement `[NOT] [I]LIKE 'pattern'` for all `SHOW` statements that list access entities: - `SHOW USERS` - `SHOW ROLES` / `SHOW CURRENT ROLES` / `SHOW ENABLED ROLES` - `SHOW SETTINGS PROFILES` - `SHOW QUOTAS` - `SHOW ROW POLICIES` - `SHOW MASKING POLICIES` This allows filtering access entities by name pattern, matching the existing behavior of `SHOW TABLES LIKE` and `SHOW DATABASES LIKE`: ```sql SHOW USERS LIKE '%admin%' SHOW ROLES NOT LIKE '%-internal' SHOW USERS ILIKE '%ALICE%' ``` The implementation mirrors the existing `SHOW TABLES` approach: three fields on the AST (`like`, `not_like`, `case_insensitive_like`), parsing in the shared access entity parser, and query rewriting in the interpreter to append a `WHERE name [NOT] [I]LIKE '...'` clause to the `system.*` table query. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `[NOT] [I]LIKE 'pattern'` filtering to `SHOW USERS`, `SHOW ROLES`, `SHOW SETTINGS PROFILES`, `SHOW QUOTAS`, `SHOW ROW POLICIES`, and `SHOW MASKING POLICIES` statements. Made with [Cursor](https://cursor.com)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114647",
          "createdAt": "2026-08-13T14:20:37Z",
          "updatedAt": "2026-08-13T14:20:37Z",
          "timestamp": "2026-08-13T14:20:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "anandheritage",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:24dcb66c5c9604273f3f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:106673",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:106673",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add LossyQuantile codec for lossy compression of floating-point numbers",
          "text": "Resolves #106219 Implement a new experimental compression codec `LossyQuantile` that encodes floating-point values as K-bit indices into per-group empirical quantile codebooks. **Syntax:** `CODEC(LossyQuantile(bits [, group_size [, stripe_size]]))` The codec accepts parameters: - `bits` (1-8): number of bits per encoded value, controls compression ratio vs. accuracy tradeoff - `group_size` (default 1048576): number of values per quantile group - `stripe_size` (default 1): interleaving stride for multi-dimensional data (e.g., fixed-size embedding arrays) **Compression:** For each group, sort the values, compute 2^K interpolated quantiles at levels `(q+1)/(2^K+1)`, then encode each value as the index of its nearest centroid. Bit-pack all K-bit indices. **Decompression:** Read centroids from the header, unpack indices, look up centroid values. **Special value handling:** NaN is mapped to the lowest centroid, Inf/-Inf are mapped to extreme centroids. These are excluded from quantile calculation. **Use case:** Brute-force vector search on slow media (S3) where I/O bandwidth dominates compute cost. Achieves 4-32x compression depending on K, with bounded distribution-adaptive reconstruction error. Requires `SET allow_experimental_codecs = 1` (same as ALP). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental `LossyQuantile` codec for lossy compression of floating-point columns using quantile-based scalar quantization. Encodes each value as a K-bit index (1-8 bits) into per-group empirical quantile codebooks, achieving 4-32x compression with bounded reconstruction error. Supports stripe mode for multi-dimensional array data. Made with [Cursor](https://cursor.com)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/106673",
          "createdAt": "2026-06-07T11:09:19Z",
          "updatedAt": "2026-08-13T14:17:34Z",
          "timestamp": "2026-08-13T14:17:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "can be tested",
            "pr-experimental"
          ],
          "author": "anandheritage",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:3304a0a1f0dd5c1c32d5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113401",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix paimon timestamp precision",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> fix: https://github.com/ClickHouse/ClickHouse/issues/112768 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables partitioned by a `TIMESTAMP` or `TIMESTAMP WITH LOCAL TIME ZONE` column of precision above 3. Such tables previously failed with `scale 6 is not supported, only support scale <= 3` before returning any row, which affected every timestamp-partitioned table written by Spark, since Spark maps both `TIMESTAMP` and `TIMESTAMP_NTZ` to Paimon `TIMESTAMP(6)`. Partition pruning on such a column now uses the full precision as well.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113401",
          "createdAt": "2026-08-05T01:12:47Z",
          "updatedAt": "2026-08-13T14:16:10Z",
          "timestamp": "2026-08-13T14:16:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "JiaQiTang98",
          "state": "open",
          "assignees": [
            "hanfei1991"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bfbbefc39724fe594387",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109602",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109602",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix the streaming-insert block wait not expiring under the query profiler",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109592 Related: https://jira.mariadb.org/browse/CONC-834 **Problem.** The file-descriptor poll timeout can silently never expire on a query thread. `ReadBufferFromFileDescriptor::poll` restarts the interrupted `poll` with the full original timeout after `EINTR`, resetting the deadline on every signal. Query threads receive periodic sampling-profiler timer signals (`query_profiler_real_time_period_ns`, default 1 s; the handler's `SA_RESTART` does not apply — per `signal(7)`, `poll` is never auto-restarted), so whenever the signal period does not exceed the timeout, the wait becomes unbounded while the fd stays silent. The user-visible path is the streaming-insert block wait: `IRowInputFormat::read` calls `poll` with the remaining `input_format_max_block_wait_ms` budget over a `StorageFile` fd/file or stdin buffer. With the profiler active, a stalled input source delays the partial-block flush indefinitely instead of flushing when the wait limit is reached. (All socket paths use `ReadBufferFromPocoSocket*::poll`, which is already deadline-aware, and are not affected.) **Root cause.** Same defect class as the MySQL `connect_timeout` fix in #109592 (mariadb-connector-c, upstream [CONC-834](https://jira.mariadb.org/browse/CONC-834)): an `EINTR` retry loop that passes the original timeout instead of the remainder. This was the last such site in `src/` — a sweep of all timed waits (`poll`/`epoll_wait`/`select`/`nanosleep`/`sigtimedwait`/io_uring/timerfd) found every other one deadline-aware (Poco sockets, `Epoll`, `KeeperTCPHandler`, `ShellCommandSource`, `base/sleep`). **Fix.** Re-poll with the remaining time computed **in microseconds from a single monotonic anchor**, reporting a timeout once the budget is exhausted. Per-retry whole-millisecond accounting (as in some existing call sites) would truncate a sub-millisecond retry interval to zero and make no progress under a sub-millisecond signal cadence, so the remainder is derived from the untouched anchor instead. The unit test (`gtest_fd_read_buffer_poll_under_signals.cpp`) waits on a pipe with no writer while a per-thread kernel timer (`timer_create` + `SIGEV_THREAD_ID`, the query profiler's own mechanism, so delivery deterministically targets the polling thread) interrupts the poll at two cadences: 10 ms (the classic deadline reset) and 0.5 ms (pins the microsecond-precision accounting). A watchdog thread disarms the timer after 3 s so a regressed build fails the elapsed assertion at ~3.2 s instead of hanging; the fixed poll returns in ~200 ms. Verified in both directions: the naive loop fails both cadences, a millisecond-accounting variant fails exactly the sub-millisecond one. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the block-wait timeout of streaming inserts (`input_format_max_block_wait_ms`) potentially never expiring while the query profiler is active: the file-descriptor poll restarted with the full timeout after every profiler signal, so a stalled input source could delay the partial-block flush indefinitely instead of flushing when the wait limit is reached.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109602",
          "createdAt": "2026-07-07T07:53:58Z",
          "updatedAt": "2026-08-13T14:15:36Z",
          "timestamp": "2026-08-13T14:15:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "tiandiwonder",
          "state": "open",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:228b082de4a141ff07ef",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108371",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108371",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Array subscript operator supports array of integers as index.",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/108095 The array subscript operator now accepts an array of integers as the index, so `arr[indexes]` gathers the elements at all of those positions at once. It is equivalent to `arrayMap(i -> arr[i], indexes)`, including the result type, the handling of negative indexes and the value produced for an out-of-range position, but it has its own implementation: a constant source array is not materialized per row, and numeric element types are gathered through the `PODArray` directly. ```sql SELECT [10, 20, 30, 40][[2, 4, 1]]; -- [20,40,10] SELECT [10, 20, 30][[1, 5, -1]]; -- [10,0,30] SELECT arrayElementOrNull([10, 20, 30], [1, 5]); -- [10,NULL] SELECT [10, 20, 30][[1, NULL, 2]]; -- [10,NULL,20] ``` The positions may be nullable, and a `NULL` position produces `NULL`, just as for a scalar index. The main use case is a lookup table: a constant dictionary array indexed by a per-row array of positions. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The array subscript operator supports an array of integers as the index: `arr[indexes]` returns the elements at all of the given positions, equivalently to `arrayMap(i -> arr[i], indexes)`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108371",
          "createdAt": "2026-06-24T12:15:02Z",
          "updatedAt": "2026-08-13T14:13:55Z",
          "timestamp": "2026-08-13T14:13:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "ucasfl",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6e7988884643d826e076",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104948",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104948",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Declarative function signatures, continuation of #3775",
          "text": "This is an experimental continuation of [#3775 (2018)](https://github.com/ClickHouse/ClickHouse/pull/3775) by @alexey-milovidov, which proposed a declarative way to describe function signatures so that argument validation and return-type inference can be driven by a small DSL instead of hand-written `getReturnTypeImpl` logic per function. ### What's in this branch A `getSignatureString()` method on `IFunction` and `IFunctionOverloadResolver`. When a function exposes a non-empty signature, the base `getReturnTypeImpl(ColumnsWithTypeAndName)` parses and applies it via a small grammar of type **matchers** (`UInt`, `Number`, `Array(T)`, `MaybeNullable(T)`, `Function((args), R)`, `T : Any` capture, …) and type **functions** (`leastSupertype`, `nativeNumber`, `aggregateFunctionReturnType(AggregateFunction(name, …))`, `subcolumnTypeOf`, `typeFromString`, `DateTime64(scale, tz)`, …). The DSL supports variadic positions (`…`), ellipsis grouping for repeated argument units (`T1, V1, …` repeats the pair), `OR` between alternatives, optional positions `[T]`, const-value capture (`const name String`), lambdas, etc. The grammar / parser / type matchers / type functions live under `src/DataTypes/FunctionSignature.h`, `FunctionSignature.cpp`, `TypeMatchers.cpp`, `TypeFunctions.cpp`. ### Coverage ~191 commits on top of master, each is a small per-family round titled `Function signatures: round N — …`. Current state on this branch: ``` SELECT count(*) AS total, countIf(signature != '') AS with_sig, round(countIf(signature != '') * 100.0 / count(*), 1) AS pct FROM system.functions WHERE NOT is_aggregate AND alias_to = '' AND origin = 'System'; 1416 1395 98.5 ``` (With \\`allow_experimental_nlp_functions = 1\\`; the few remaining unset ones are setting / config-gated functions like \\`aiClassify\\`, \\`region*\\`, \\`synonyms\\`, whose \\`create\\` throws unless the relevant config block / setting is present, so \\`system.functions.tryGet\\` returns null even though the signature is in source. With the right config + setting they all surface — verified in tmp/ test config.) Roughly: - **Authoritative** (DSL drives type-check and return type, the legacy `getReturnTypeImpl` is either gone or bypassed): higher-order array functions (`arrayMap`, `arrayFilter`, `arrayFirst*`, …), `toIntervalX`, comparison, `multiSearch*` / `multiMatch*`, vector L-norms / distances / dot product, `UUIDv7ToDateTime`, `arrayReduce` / `arrayReduceInRanges`, `reverse`, `mapKeys` / `mapValues`, `getSubcolumn`, the `least` / `greatest` resolver, paired-variadic `timeSeriesTagsToGroup` / `timeSeriesStoreTags`, … - **Documentation-only** (signature surfaced via `system.functions` but the legacy `getReturnTypeImpl` still runs because the result type uses promotion / widening / setting-dependent dispatch the current DSL can't express): arithmetic (`plus`, `minus`, `multiply`, `divide`, modulo / intDiv family), array widening (`arraySum` / `arrayCumSum*` / `arrayDifference`), `mapContains*Like`, `dateTrunc`, `toStartOfWeek`, `parseDateTime*`, `transform`, `range`, `toX` conversions, etc. There is a `signature_documentation` opt-in alongside `signature` in the binary-arithmetic, unary-arithmetic, FunctionArrayMapped, and FunctionMapToArrayAdapter families specifically so a function can advertise a signature without the DSL accidentally hijacking the legacy widening logic. ### Status Draft / experimental. - This is an experiment — I'm not asking for it to be merged. There's significant disruption (191 commits touching every function family), and the gain is mostly documentation: today only ~30% of the converted functions are actually DSL-authoritative; the rest are decorative. The arithmetic promotion matrix and the per-Op widening rules in particular would need a richer type-function vocabulary (or new matchers) before they can be expressed declaratively. - The branch has been kept rebased on master throughout and builds cleanly. I've run the stateless test suite on each round. The pre-existing-on-master test failures I hit are listed in commit messages of the rounds where I encountered them (none introduced by this work). - The original PR has been open since 2018 — this branch is intended as a concrete data point on \\\"what fraction of ClickHouse functions can be reasonably described by a declarative signature, and what would the DSL need to grow to cover the rest.\\\" That's the question I'd love feedback on. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): ... ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104948",
          "createdAt": "2026-05-14T14:35:16Z",
          "updatedAt": "2026-08-13T14:11:52Z",
          "timestamp": "2026-08-13T14:11:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 38
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8731633c027d19083fdc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114475",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114475",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113289 to 26.5: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114475",
          "createdAt": "2026-08-12T12:06:09Z",
          "updatedAt": "2026-08-13T14:11:33Z",
          "timestamp": "2026-08-13T14:11:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll3",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:50eba92b4df6194b6239",
        "signalId": "github:ClickHouse/ClickHouse:issue:114641",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114641",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Text index: support arbitrary `LIKE` patterns with `tokenizer = 'array'`",
          "text": "### Company or project name ClickHouse Inc. ### Use case A text index with `tokenizer = 'array'` stores the whole column value as a single token, so its dictionary is the set of distinct values of the column, and a `LIKE` filter over that dictionary is exactly the query predicate. Today only patterns of the shape `%needle%`, with an alphanumeric needle, are evaluated using the index. Any other pattern reads the whole column, although the index already contains everything needed to answer it. ```sql CREATE TABLE t ( name String, INDEX idx name TYPE text(tokenizer = 'array') ) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO t VALUES ('alpha-service-prod'), ('beta-service-prod'), ('gamma-svc-4999-dev'); ``` | Query | Uses the index today | Why not | |---|---|---| | `SELECT * FROM t WHERE name LIKE '%4999%'` | yes | | | `SELECT * FROM t WHERE name LIKE '%svc-4999%'` | no | punctuation in the needle | | `SELECT * FROM t WHERE name LIKE 'alpha%'` | no | anchored at the start | | `SELECT * FROM t WHERE name LIKE '%prod'` | no | anchored at the end | | `SELECT * FROM t WHERE name LIKE '%alpha%prod%'` | no | more than one needle | | `SELECT * FROM t WHERE name LIKE 'alpha_service%'` | no | `_` wildcard | ### Describe the solution you'd like For `tokenizer = 'array'`, evaluate arbitrary `LIKE` and `ILIKE` patterns using the text index: anchors, punctuation, `_` wildcards, several `%`-separated needles, escaped metacharacters. The existing cost guards should keep working — the minimum required literal length and the limit on how much of the index may be read, falling back to reading the column when the limit is exceeded. ### Describe alternatives you've considered An additional `ngrams(3)` index on the same column answers these patterns, but doubles index storage and write amplification for data the `array` dictionary already describes exactly. ### Additional context Actually, any predicate over the column with an `array` tokenizer can be supported, but it requires non-trivial development. The `LIKE` improvement comes almost for free, so let's start with this. Related: https://github.com/ClickHouse/ClickHouse/pull/98149 Related: https://github.com/ClickHouse/ClickHouse/issues/97723",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114641",
          "createdAt": "2026-08-13T13:03:00Z",
          "updatedAt": "2026-08-13T14:11:21Z",
          "timestamp": "2026-08-13T14:11:21Z",
          "metrics": {
            "reactions": 1,
            "comments": 0
          },
          "labels": [
            "feature",
            "comp-text-index"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [
            "ahmadov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:834d34ada14e1262e529",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114627",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114627",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112217 to 26.7: Measure cancellation server-side in test_cancel_backup.py",
          "text": "Backport of https://github.com/ClickHouse/ClickHouse/pull/112217 to `26.7`. Related: https://github.com/ClickHouse/ClickHouse/pull/112217 Related: https://github.com/ClickHouse/ClickHouse/pull/114478 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Description `test_backup_restore_on_cluster/test_cancel_backup.py::test_cancel_backup` is flaky on the `26.7` branch with the same signature that #112217 fixed on `master`: the client-side `time_to_cancel` stopwatch measures the integration harness (several `docker exec` clickhouse-client launches plus `wait_status` poll quantum) rather than the cancellation itself, and the resulting body exception is masked in reports as `NoTrashChecker.__exit__` asserting `'QUERY_WAS_CANCELLED' in []`. It just failed this way on the backport PR https://github.com/ClickHouse/ClickHouse/pull/114478 (report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114478&sha=f5364950f96d60206363d2b2c1bff278a6976ee8&name_0=BackportPR&name_1=Integration%20tests%20%28amd_asan_ubsan%2C%20db%20disk%2C%20old%20analyzer%2C%201%2F6%29), and #112217 itself documents an occurrence on the `26.6` release branch, so release branches keep hitting it. This is a clean cherry-pick of the test-only fix (measure cancellation server-side via `system.backups` timestamps).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114627",
          "createdAt": "2026-08-13T12:29:22Z",
          "updatedAt": "2026-08-13T14:10:31Z",
          "timestamp": "2026-08-13T14:10:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e704bd1e788bfda5cd28",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112950",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112950",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support the Vortex file format",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/87327 Related: https://github.com/ClickHouse/rust_vendor/pull/74 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added support for reading and writing the [Vortex](https://github.com/vortex-data/vortex) columnar file format (the `Vortex` input and output format). This closes [#87327](https://github.com/ClickHouse/ClickHouse/issues/87327). ### Documentation entry for user-facing changes The implementation uses the Rust `vortex` crate (v0.83.0) through a new C FFI crate `rust/workspace/vortex` (`_ch_rust_vortex`), following the same pattern as `prql` and `polyglot`. Data crosses the FFI boundary through the Arrow C Data Interface and is converted with the same `ArrowColumnToCHColumn`/`CHColumnToArrowColumn` code as the `Arrow` format. IO is delegated back to ClickHouse through callbacks: reads go through ClickHouse's own read buffers (range reads for seekable inputs, whole-file buffering otherwise), and the produced file is streamed into the output buffer. All work is driven by a single-threaded runtime on the calling thread — the library spawns no threads, and Rust panics are caught at the FFI boundary and turned into exceptions. Features: - Reading with projection pushdown: only the columns used by the query are read from the file. - Schema inference and `count()`-only queries answered from file metadata without reading data. - Writing with the library's default adaptive compression (BtrBlocks-style cascading encodings + zstd), including a valid empty file for empty results. - Graceful errors on malformed and truncated files (fuzzer-friendly: no aborts, Rust panics become exceptions). Limitations (documented in `docs/reference/formats/Vortex.mdx`): - `Map`, `Int128`/`UInt128`/`Int256`/`UInt256`, `IPv6`, and `Interval` columns cannot be written (no corresponding Vortex type). - `String` and `FixedString` are written as Vortex `Binary` (ClickHouse strings are arbitrary bytes, while Vortex requires `Utf8` to be valid UTF-8). - The format is disabled in MSan builds: the MSan-instrumented library (with origin tracking) is so large that linking `unit_tests_dbms` overflows the 2 GiB `R_X86_64_PC32` relocation range (same approach as `wasmtime` and `delta-kernel-rs`). - Reading and writing are single-threaded in this first version: the whole scan (I/O, decompression, and decoding) runs on one thread, so on ClickBench reads are significantly slower than `Parquet`, which ClickHouse decodes with multiple threads (see the [benchmark results](https://github.com/ClickHouse/ClickHouse/pull/112950#issuecomment-5274283519) and [the explanation](https://github.com/ClickHouse/ClickHouse/pull/112950#issuecomment-5274483899)). Filter pushdown (added in https://github.com/ClickHouse/ClickHouse/pull/114373, `input_format_vortex_filter_push_down`, on by default) reduces the amount of data decoded by selective queries, but does not parallelize the scan. Writes are also slower than `Parquet` (the adaptive compressor samples many encodings per column) — parallelism can be added later. The 127 new vendored Rust crates are added in https://github.com/ClickHouse/rust_vendor/pull/74 (the `contrib/rust_vendor` submodule is bumped to that branch).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112950",
          "createdAt": "2026-08-01T21:39:00Z",
          "updatedAt": "2026-08-13T14:10:30Z",
          "timestamp": "2026-08-13T14:10:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 19
          },
          "labels": [
            "pr-feature",
            "submodule changed",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f434920d24b107731fec",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114523",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114523",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong results of the partial aggregation strategy in distributed query plans",
          "text": "The partial+merge aggregation strategy of `make_distributed_plan` rewrites a final aggregation into a partial `AggregatingStep` plus a memory-efficient `MergingAggregatedStep`. The memory-efficient merge consumes each input as a stream of two-level buckets in ascending order, but the rewrite kept `should_produce_results_in_order_of_bucket_number = false` on the partial step, so its multi-stream output was combined into the single exchange stream in arbitrary order. When a bucket arrived after the merge had already emitted it, the merge emitted it a second time: duplicated `GROUP BY` keys with split aggregate states (for ClickBench Q08, a wrong `COUNT(DISTINCT UserID)`). The fix, one commit each: - Make the row count estimation look through `LogicalExchangeStep`. The `#if CLICKHOUSE_CLOUD` guard around it was needed when the exchange steps existed only in the private repo; since they are in the public repo, the guard only made every aggregation over a distributed read fall back to the Shuffle strategy in non-cloud builds (and masked this bug there). - Validate the bucket delivery order in `GroupingAggregatedTransform`: an unannounced late bucket now fails with a `LOGICAL_ERROR` exception instead of silently duplicating groups. - Wait for delayed buckets in `GroupingAggregatedTransform` when the consumer needs the output in bucket order, and release the announcements of finished inputs. Covered by a gtest. - Build the partial step with `should_produce_results_in_order_of_bucket_number = true` when the merge is memory-efficient (`AggregatingStep::cloneAsPartial`). The partial aggregation then produces a single bucket-ordered stream per worker, the same contract the classic distributed path establishes on shards. Covered by a stateless test that reproduces the duplicated keys on the code without the fix (about 80% of single runs, 8 repetitions, 50k rows, ~0.5s). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix duplicated `GROUP BY` keys in the partial aggregation strategy of `make_distributed_plan`, and allow this strategy for aggregations over a distributed read. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114523",
          "createdAt": "2026-08-12T16:51:05Z",
          "updatedAt": "2026-08-13T14:08:11Z",
          "timestamp": "2026-08-13T14:08:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [
            "nickitat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:768b0dbe10912f34c1fe",
        "signalId": "github:ClickHouse/ClickHouse:issue:114612",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114612",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Heap-use-after-free: Parquet v3 prefetcher reads and writes through a `ReadBuffer` freed by `IInputFormat::onFinish`",
          "text": "🕵️ ## Describe what's wrong `ParquetV3BlockInputFormat` does not override `resetReadBuffer()`, so `IInputFormat::onFinish()` frees the format's owned `ReadBuffer` while the Parquet `Prefetcher`'s IO tasks are still reading through it on the prefetch thread pool. The tasks then dereference freed memory. ASan on a plain `clickhouse local` run of the reproducer below (master, 26.8.1.1310): ``` ==ERROR: AddressSanitizer: heap-use-after-free on address 0x7112bd001800 WRITE of size 1048576 at 0x7112bd001800 thread T9 (ParquetPrefetch) #3 in DB::ReadBuffer::next() src/IO/ReadBuffer.cpp:113:15 #5 in DB::ReadBuffer::read(char*, unsigned long) src/IO/ReadBuffer.h:169:37 #6 in DB::Parquet::Prefetcher::readSync(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:113:29 #7 in DB::Parquet::Prefetcher::runTask(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:540:13 #8 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:435 #15 in DB::ThreadPoolCallbackRunnerFast::threadFunction() src/Common/threadPoolCallbackRunner.cpp:225:13 ``` It is a **1 MiB write into freed heap**, so this is memory corruption, not only a bad read. The same defect was hit independently by the `La Casa Del Dolor (arm_asan_ubsan)` CI job, where the report is a read and carries the full free/allocation stacks: ``` ERROR: AddressSanitizer: heap-use-after-free on address 0xfcdd2aec39c0 READ of size 8 at 0xfcdd2aec39c0 thread T423 (ThreadPool) #0 in DB::Parquet::Prefetcher::readSync(...) Prefetcher.cpp:110:21 <- reader->setReadUntilEnd() #1 in DB::Parquet::Prefetcher::runTask(...) Prefetcher.cpp:523:13 #2 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() Prefetcher.cpp:424:17 #9 in DB::ThreadPoolCallbackRunnerFast::threadFunction() threadPoolCallbackRunner.cpp:225:13 0xfcdd2aec39c0 is located 0 bytes inside of 272-byte region freed by thread T468 (ThreadPool) here: #1 in std::default_delete<DB::ReadBuffer>::operator()(DB::ReadBuffer*) #7 in std::vector<std::unique_ptr<DB::ReadBuffer>>::~vector() #8 in DB::ISource::work() src/Processors/ISource.cpp:140:13 ... #18 in DB::IPolygonDictionary::loadData() src/Dictionaries/PolygonDictionary.cpp:337:8 previously allocated by thread T468 (ThreadPool) here: #2 in DB::FileDictionarySource::loadAll() src/Dictionaries/FileDictionarySource.cpp:60:19 ``` `ISource.cpp:140` is the `onFinish()` call inside the `catch (...)` handler, and the freed 272-byte region is the `ReadBufferFromFile` that `FileDictionarySource::loadAll` handed to the format via `addBuffer`. ## How to reproduce Reproduces on master (26.8.1.1310) with the official `build_amd_asan_ubsan` binary, **8 out of 10 runs, with no ThreadFuzzer and no unusual settings**. Under ThreadFuzzer it hit on the first run. Run Fiddle: https://fiddle.clickhouse.com/06f26a34-d286-40a1-85f5-7d3e952d47b8 On a release build this prints only the load error and exits 53: `Code: 53. CAST AS Array can only be performed between same-dimensional Array ... While executing ParquetV3BlockInputFormat. (TYPE_MISMATCH)` — the 1 MiB write into freed memory is silent. On an ASan build the process dies with the report above instead (5/5 runs of exactly the commands above; 8/10 in an earlier variant, so it is a race, but a very wide one). The polygon dictionary is only a convenient way to make the pipeline throw mid-read while owning its `ReadBuffer` — the defect is in the format, not in the dictionary. ## Root cause `Prefetcher` protects only *its own* lifetime. `scheduleTask` captures the shutdown handle and each task takes `std::shared_lock(*_shutdown, std::try_to_lock)`; `~Prefetcher` calls `shutdown->shutdown()` (`ShutdownHelper`, `src/Common/threadPoolCallbackRunner.h:507`), which blocks until every in-flight task has released. As its own comment says, \"`this` is safe to access as long as `shutdown_lock` is held\". But `Prefetcher::reader` points at a `ReadBuffer` owned by `IInputFormat::owned_buffers`, an unrelated lifetime that the handshake does not cover: - `IInputFormat::onFinish()` -> `resetReadBuffer()` -> `resetOwnedBuffers()` -> `owned_buffers.clear()` - `ParquetV3BlockInputFormat` overrides `resetParser()` and `onCancel()`, but **not `resetReadBuffer()`**. `resetParser()` happens to be safe only because it destroys the reader *before* delegating to the base. `resetReadBuffer()` has no such ordering, so the buffer dies while the `ReadManager` -> `Reader` -> `Prefetcher` chain is still alive with tasks running. Three routes reach it: `onFinish()` from the `catch (...)` in `ISource::work()` (the one above), `onFinish()` on the normal completion path (`ISource.cpp:133`) with speculative prefetches still outstanding, and `IInputFormat::setReadBuffer` when a format is reused across files. ## Suggested fix Mirror what `resetParser()` already does, so `~Prefetcher` drains the IO tasks before the base frees the buffer: ```cpp void ParquetV3BlockInputFormat::resetReadBuffer() { /// ~Prefetcher waits for in-flight IO tasks, which read through the buffer that /// IInputFormat::resetReadBuffer() is about to free. { std::lock_guard lock(reader_mutex); reader.reset(); } IInputFormat::resetReadBuffer(); } ``` ## Additional context Distinct from #109678 / #112573. That was a data race on `AsynchronousBoundedReadBuffer::prefetch_future` consumed by concurrent `readBigAt` (the `RandomRead` branch, `Prefetcher.cpp:102`). This one is a lifetime bug in the `SeekAndRead` branch, which already holds `read_mutex` — locking cannot help once the object is freed — and the freed buffer here is a plain `ReadBufferFromFile`, so `AsynchronousBoundedReadBuffer` is not involved at all. The binary used above already contains #112573 (merged into 26.8.1.1240). Found by this Dolor run: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=94148&sha=a2ef409981fc77a168eaf8cc68d7ddf67c9fadec&name_0=PR&name_1=La+Casa+Del+Dolor+%28arm_asan_ubsan%29",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114612",
          "createdAt": "2026-08-13T09:21:28Z",
          "updatedAt": "2026-08-13T14:07:54Z",
          "timestamp": "2026-08-13T14:07:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "bug",
            "blocker",
            "comp-parquet-reader-v3"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6e38d7172a6321f05efc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110958",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110958",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix toTime key-expression type mismatch under use_legacy_to_time",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107951 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server abort (`Bad cast from type ColumnVector<UInt32> to ColumnVector<Int32>`, a `LOGICAL_ERROR`) when inserting into a `MergeTree` table whose `PRIMARY KEY` or `ORDER BY` uses `toTime(...)` from a session whose `use_legacy_to_time` value differs from the one under which the key metadata was built. The `toTime` legacy/new resolution now follows the context that builds the expression, so a persisted key expression's type no longer depends on the writing session. ### Description `toTime()` resolves to two different functions depending on the `use_legacy_to_time` setting: the new `toTime` returns `Time` (Int32-backed), the legacy variant returns `DateTime` (UInt32-backed). The selection in `FunctionFactory::tryGetImpl` read the thread-local query context and ignored the `context` argument the caller passed. A `MergeTree` table with `PRIMARY KEY (toTime(c1))` / `ORDER BY toTime(c1)` persists only the expression AST. Its primary-index on-disk type comes from `metadata_snapshot->getPrimaryKey().data_types`, derived by rebuilding the key expression with the storage's global context (server-default `use_legacy_to_time`). The part-writer serializes the index with that persisted type. When an `INSERT` runs in a session with a different `use_legacy_to_time`, the write-path key expression re-resolved `toTime` to the other variant, producing a column whose physical type mismatched the persisted serialization, so `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow` hit `assert_cast<ColumnVector<Int32>>(ColumnVector<UInt32>)` and aborted the server (a handled exception in release builds, an abort under debug/sanitizers). Fix: resolve the `toTime` legacy swap from the caller-provided `context` (falling back to the thread-local query context only when no context is supplied). Stored key/sorting expressions are rebuilt with the storage global context, so they now resolve `toTime` consistently with the type persisted in the table metadata, while normal query resolution still honors the session setting. Found by the BuzzHouse fuzzer. - CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109351&sha=b0cca3209ae52b188fa092a31eb1e754db40608d&name_0=PR&name_1=BuzzHouse%20%28amd_msan%29 - Check: `BuzzHouse (amd_msan)`; assertion `Bad cast from type DB::ColumnVector<unsigned int> to DB::ColumnVector<int>` at `SerializationNumber<int>::serializeBinary` <- `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow`. Reproducer: ```sql CREATE TABLE t (c0 Int32, c1 DateTime64 MATERIALIZED nowInBlock64()) ENGINE = MergeTree() PRIMARY KEY (toTime(c1)); INSERT INTO t (c0) SETTINGS use_legacy_to_time = 1 SELECT number FROM numbers(1000); ``` ### DDL behaviour change carried by the fix Previously, `CREATE TABLE` in a session whose `use_legacy_to_time` differed from the server-wide default stamped the stored key type with the session's resolution (e.g. `DateTime` when the session set `use_legacy_to_time = 1` on a server defaulting to `0`). With this fix, the stored key type always resolves under the server-wide default, so the session setting at `CREATE` time no longer affects the persisted key type. This is observable via `DESCRIBE mergeTreeIndex(...)` and is pinned by the test. Upgrade note for that narrow window (table created while the session setting differed from the server global, on a pre-fix binary): after the upgrade the same table resolves its key as `Time`, so parts written before and after store different raw key values for the same timestamp (e.g. `90000` vs `3600` for `01:00:00`); reads, inserts and merges succeed, but a key-range predicate may miss rows from old parts. Where session and global agreed (the overwhelmingly common case), nothing changes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110958",
          "createdAt": "2026-07-18T23:24:10Z",
          "updatedAt": "2026-08-13T14:05:11Z",
          "timestamp": "2026-08-13T14:05:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 21
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a7c90144c4df665ba5d2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113691",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113691",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `theilsU` window state returning noise when the frame's first argument is constant",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/80373 Related: https://github.com/ClickHouse/ClickHouse/pull/93384 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `theilsU` over a window frame returning an arbitrary value instead of 0 when the first argument is constant within the frame. ### Description `TheilsUWindowData::getResult` (the window-optimized state introduced in https://github.com/ClickHouse/ClickHouse/pull/93384) computes the entropy `H(A)` from cached incremental `Σ n·log n` sums. When the first argument is constant within the frame, the true `H(A)` is zero, and the computed value is pure rounding noise from the incremental updates. The code compared it against exact zero, so a tiny positive noise value passed the check, and `1 - H(A|B) / H(A)` then divided noise by noise: in debug builds this tripped the sanity check as the exception `Logical error: 'res < 1.0 + 1e-4'`, and in release builds the function could return an arbitrary value in $[0, 1]$ instead of 0. The exact (non-window) code path recomputes the entropies from the count maps, where a constant column gives `log(1) = 0` exactly, so it is not affected. The fix compares `H(A)` against an error bound proportional to `N · ε · log N` instead of exact zero, and widens the sanity-check tolerance by the same relative amount so that near-threshold frames do not trip it either. Found by the AST fuzzer on an unrelated PR (it hit https://github.com/ClickHouse/ClickHouse/pull/80373 and https://github.com/ClickHouse/ClickHouse/pull/107667): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=80373&sha=6c271049214aa5a94bd9a5f12fb27a9ffa75648f&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113691",
          "createdAt": "2026-08-06T15:21:02Z",
          "updatedAt": "2026-08-13T14:04:48Z",
          "timestamp": "2026-08-13T14:04:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6ebd5a6d3ce45b47184b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110321",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110321",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support INSERT ... VALUES in the polyglot SQL dialect",
          "text": "Enable `INSERT ... VALUES` with inline data in the polyglot SQL dialect (`dialect = 'polyglot'`). Previously, running e.g. `INSERT INTO t VALUES (1), (2), (3)` with `dialect = 'polyglot'` failed with `Multi-statement queries are not supported in polyglot dialect mode`. The underlying problem is that transpiling inside the parser cannot deliver inline data to the executor: the transpiled buffer is transient, and the executor overwrites `ASTInsertQuery::tail` with the external input stream, so the inline-data pointers (`data`/`end`) must reference a live query buffer. This transpiles the query up front instead of inside the parser: - The server (`executeQuery`) transpiles a foreign-dialect query to ClickHouse SQL before parsing, keeps the transpiled text alive on the query context, and parses it with the standard parser. Inline INSERT data then points into a live buffer and is processed by the normal machinery. `SET` queries are still parsed as-is so `dialect`/`polyglot_dialect` can always be changed back. - The client (`clickhouse-client`/`clickhouse-local`) parses a non-ClickHouse-dialect query into an AST — which, for a foreign dialect, means transpiling it locally only to drive client-side handling (statement classification, output format, INSERT detection) — but then sends the *original* query text verbatim, without splitting off inline data. The server performs the authoritative transpilation whose result is actually executed, so inline INSERT data lives in a server-owned buffer and survives parsing. The client-side transpilation is throwaway; note this means the transpiler must also be available on the client (a client built without `USE_POLYGLOT` fails locally with `SUPPORT_IS_DISABLED`), and the client and server transpilers are assumed to agree — acceptable for this experimental dialect. Every parse-time setting the query was parsed under (`dialect`, `allow_experimental_polyglot_dialect`, `polyglot_dialect`, `allow_settings_after_format_in_insert`, `implicit_select`, and the parse limits `max_query_size`, `max_parser_depth`, `max_parser_backtracks`) is pinned in the per-query settings sent along with the verbatim text, so the query's own `SETTINGS` clause cannot change how the server reparses that same text (it still applies to the query's execution, and a `SET` still takes effect for subsequent queries). All changes are gated on the dialect, so ordinary ClickHouse INSERTs are unaffected. Validated over the HTTP interface, the native client, and `clickhouse-local` (multi-row and single-row `VALUES`, `INSERT ... SELECT`, and PostgreSQL literal transpilation such as `true`/`false`); `SET` passthrough and multi-statement rejection are preserved. External insert data combined with a foreign-dialect `INSERT` is rejected with `NOT_IMPLEMENTED` instead of being silently dropped, on both surfaces: the client rejects piped stdin and `INFILE` (it sends the query verbatim and cannot forward a data tail), and the server rejects a non-empty HTTP request body appended to a streaming `INSERT` (`POST /?query=INSERT ... &dialect=polyglot` with a body). A foreign-dialect `INSERT` is transpiled as a whole, so the body would go through neither the transpiler nor the `max_query_size` guard, mixing two parsing rules in one `INSERT`. An empty body still works, which is the normal way to run a polyglot `INSERT` over HTTP. Limitations (scoped, experimental): because a foreign-dialect query is transpiled as a whole (the transpiler rewrites the inline data too and cannot know where the SQL header ends without parsing the dialect), the inline `INSERT ... VALUES` data counts towards `max_query_size` — unlike a native ClickHouse `INSERT`, whose inline data is streamed and is not bounded by `max_query_size`. An oversized payload fails-close with a dedicated, actionable error rather than silently changing the `INSERT` size contract; increase `max_query_size` to submit larger inline payloads. Only `INSERT ... VALUES` inline data is transpilable by the bundled dialects. `INSERT ... FORMAT ...` is not: `FORMAT` is a ClickHouse-only extension, so a foreign-dialect parser rejects the query at the inline data that follows (empirically, `postgresql`/`mysql`/`sqlite`/`duckdb`/`snowflake`/`bigquery` all fail at the first data row after `FORMAT`; a hypothetical identity transpiler even drops the raw `FORMAT` payload rather than re-emitting it). A foreign-dialect `INSERT ... FORMAT` therefore fails cleanly with a syntax error and inserts nothing — like `EXPLAIN INSERT ... VALUES`, which is also not transpilable by the bundled dialects (rejected at the `VALUES` token). The server-owned transpiled buffer that carries the inline data is itself format-agnostic and would handle `FORMAT` data if a transpiler ever produced such a query; the parser also defensively clears the inline-data pointers of an `EXPLAIN`-wrapped `INSERT` — the same way the client unwraps it — so both forms are safe if a future transpiler supports them. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `INSERT ... VALUES` with inline data when using the experimental `polyglot` SQL dialect. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110321",
          "createdAt": "2026-07-13T21:39:09Z",
          "updatedAt": "2026-08-13T14:00:19Z",
          "timestamp": "2026-08-13T14:00:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d6bd0b1f8d3e6393921b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114536",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114536",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add ClickStack guide for isolating read and write workloads",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/113284 ### Changelog category (leave one): - Documentation (changelog entry is not required) --- ClickStack's docs claim the ability to \"independently isolate read and write workloads with Warehouses\" in five places without documenting how. The only existing guidance (in `deployment/managed`, added in ClickHouse/clickhouse-docs#5452) covers one half of it: that the ClickStack UI binds to whichever Cloud service it is launched from. This adds a guide for the full topology and links it from the places that previously mentioned the capability without explaining it. **New page** — `clickstack/managing/isolating-read-write.mdx`: - Why isolate: read-only services run no background merges, so their compute is dedicated to queries; ingestion is insulated from expensive queries; each side scales independently against the ingest and query figures from the sizing model. - Recommended topology: one read-write service for ingestion, one read-only service for ClickStack, plus the planning constraints — the first service is always read-write, type is fixed at creation, and merge assignment crosses read-write services. - Setup steps: prepare the read-write service and ingestion user, add the read-only service, point the collector at the read-write endpoint, point the UI at read-only compute (managed and self-hosted), then verify the split with `system.query_log` — including the `all_groups.default` note, since `system` tables are per-service. - Where DDL and materialized views execute, and how alert evaluation follows the connection of the source it is attached to. - Advanced: separating merges from ingestion onto a dedicated merge service, flagged as requiring a support request, with the caveats that come with it (mutation tracking, TTL deletion, auto-idling, keeping queries off both read-write services). **Cross-links**: `managing/overview` (admin guides table), `managing/production` (new subsection), `managing/estimating-resources` (ties the ingest and query vCPU split to separate services), `deployment/managed` (extends the existing read-only-compute section), and the previously unlinked bullets in `overview` and `architecture`. `getting-started/managed.mdx` carries the same bullet but is deliberately left untouched — #112118 rewrites that section and already links the concept, so editing it here would only create a conflict.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114536",
          "createdAt": "2026-08-12T19:46:18Z",
          "updatedAt": "2026-08-13T13:59:25Z",
          "timestamp": "2026-08-13T13:59:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "andremm",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:de8593ffb4a98438599b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109450",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109450",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix split-dependent hash of JSON/Object columns",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109428 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `JSON`/`Object`/`Dynamic` columns hashing logically equal values differently depending on their physical layout: how `JSON`/`Object` paths were split between dynamic subcolumns and shared data, or whether a `Dynamic` value sat in a typed or the shared variant. Hash-based joins on such a key silently missed matches, and a `grace_hash` spill could raise a `LOGICAL_ERROR` (\"Invalid state transition\"). A `JSON` key needed no non-default setting; `Dynamic` also required `allow_dynamic_type_in_join_keys`. ### Description **The implementation changed since the earlier approval here, so please treat it as needing a fresh look.** The approved version hashed canonical `serializeValueIntoArena` bytes and cost up to +568% CPU; this one produces none. Both measured below. `computeHashInto` is the per-row weak hash behind the in-memory scatter paths (sharded aggregation, `grace_hash` bucketing, window `PARTITION BY`, join scatter). For `JSON`/`Object` and `Dynamic` it hashed the physical layout: `ColumnDynamic` forwarded to `ColumnVariant`, which hashes a shared-variant value by its serialized blob but a typed one by the column's representation; `ColumnObject` chained sub-columns in section order, so a path's contribution moved with it. Both splits follow insertion history and can change across a temp-file round-trip, while `compareAt` already treats the representations as equal. `JSON` hits this at default settings, since `hasDynamicType` misses its inner `Dynamic`. `ColumnDynamic::computeHashInto` forwards to `ColumnVariant` as master does, then overwrites only shared-discriminator rows with the leaf hash the value would have when typed. `ColumnObject::computeHashInto` folds `(path, value)` over the sorted union of the dynamic paths and `shared_data`, one cursor per row. `updateHashFast` and `updateHashWithValueRange` stay layout-dependent. `04505_json_object_hash_split_invariance` checks that equal values with different layouts match under the three hash joins, collapse under `GROUP BY`/`DISTINCT`/`uniqExact`, and that a spill flushes real payload rather than only constructing buckets: six assertions fail on master, all pass here. Three `ComputeHashInto` gtests cover the per-row hash and scratch state SQL cannot see. `ShuffleSendStep` shuffles rows to remote workers with this hash and marks no basis in its serialized step, so mixed versions could disagree on a shuffle key. Capability gate, or unsupported anyway? <details><summary>Provenance and performance</summary> Found by the AST fuzzer on #109428 as `Invalid state transition, expected WRITING_BLOCKS, got JOINING_BLOCKS` in `GraceHashJoin::FileBucket`, STID 3913-4579 ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109428&sha=15a6faf088a1b097b3ed56c69ec1fb4d5d602d17&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29)): a key re-hashing into an earlier bucket mid-spill violates that state machine. Release-shaped build, three arms from one revision (master with these files reverted; the approved implementation; this one), `UserTimeMicroseconds` min over 11 alternating reps; all arms agreed on all 12 workloads first. | workload | vs master | vs appr. | |---|---|---| | JSON `grace_hash`, no shared | +2.8% | -61.2% | | JSON `grace_hash`, mixed | +8.2% | -33.8% | | JSON `grace_hash`, all shared | +11.7% | +6.5% | | window JSON, no shared | +12.8% | -88.3% | | window JSON, mixed | **-10.6%** | -1.7% | | window JSON, 4 shared | +51.5% | -46.8% | | `Dynamic` `grace_hash` | +44.3% | -34.3% | | `Dynamic`, 100% shared | +70.3% | -1.2% | Non-JSON controls are flat and `ArenaAlloc*` matches master. Making a byte-producing basis cheap failed three times: even hoisting the buffer so the block allocated almost nothing still cost +1253%. The residual tracks the shared representation only: each shared value pays one decode, one leaf hash, one fresh scratch column. Of the worst arm's +70.3%, about +30.4% is that scratch reset (+24.9%, +18.5% on the next two). It uses `cloneEmpty()` rather than `popBack` because `popBack` is a row operation: `ColumnLowCardinality::popBack` drops indexes but not the dictionary, which `computeHashInto` re-hashes in full per row, and `ColumnVariant` nests the shape. Conditioning it on a measured property of the scratch column did not survive review, so it stays unconditional. The cliff is bounded per `computeHashInto` call, not by the table; a 5k-40k ladder is linear. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109450",
          "createdAt": "2026-07-05T18:43:19Z",
          "updatedAt": "2026-08-13T13:56:00Z",
          "timestamp": "2026-08-13T13:56:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "harikrishnan94"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:5fe53ad76011bedfc090",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114606",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114606",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "PostgreSQL: allow an empty TLS contents override when the collection stores no credential",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113947 Related: https://github.com/ClickHouse/ClickHouse/pull/110615 Port of the #113947 narrowing (MySQL) to `StoragePostgreSQL::getSSLParams`. An empty `ssl*_pem` query override is rejected only when it would actually drop a TLS credential the named collection carries — a path (a query cannot override path keys, so the value read is the collection's own) or contents, read in the pre-override form via `NamedCollection::getValueBeforeQueryOverride`. On a collection with no TLS keys at all, the empty override stays the no-op it is for the direct arguments, instead of throwing `BAD_ARGUMENTS`. The rejection cases are unchanged and remain covered by `test_path_overrides_are_rejected` and `test_tls_credentials_in_sql_named_collection`; the new `test_empty_override_without_stored_credential_is_noop` covers the no-op case (it fails without the code change). ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix of the unreleased #110615: an empty `ssl*_pem` override on a PostgreSQL named collection without TLS credentials is a no-op again instead of an error.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114606",
          "createdAt": "2026-08-13T09:11:09Z",
          "updatedAt": "2026-08-13T13:52:08Z",
          "timestamp": "2026-08-13T13:52:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b0bf8f591482cfb65095",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114607",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114607",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix an infinite uncancellable loop in `hop` and `windowID` on an interval whose span wraps modulo 2^32",
          "text": "<!-- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed an infinite, uncancellable loop in functions `hop` and `windowID` when the span of an interval argument in seconds is a multiple of 2^32 (for example, `toIntervalDay(2147483648)`): the wrapped subtraction dodged the time-overflow check, and with constant arguments the loop ran at analysis time, where the query could not even be killed. An interval like `toIntervalDay(2147483648)` spans `2^31 * 86400 = 43200 * 2^32` seconds, so subtracting it in the wrapping `UInt32` arithmetic of `AddTime` is a no-op: `wstart == wend`, which dodges the `wstart > wend` time-overflow guard added by https://github.com/ClickHouse/ClickHouse/pull/61523. The window-searching loop of `executeHop` then decrements `wend` by one hop at a time past zero, where it wraps back to ~2^32 while staying congruent to its starting point modulo `gcd(86400, 2^32) = 128`, never hits a value `<= time`, and spins forever. With constant arguments the loop runs during constant folding in `QueryAnalyzer::resolveFunction`, at analysis time, where nothing checks the cancellation. The same loop shape exists in `executeHopSlice` (`windowID`). The fix makes the guard `wstart >= wend` - subtracting a whole positive interval must change the time, and equality only happens on a wrap - and adds a wrap check to the loop itself (the new `wend` coming out greater than the old one throws the same `Time overflow` error), which terminates every wrapping loop regardless of how the pre-loop values were corrupted. The new test `04891_time_window_functions_wrapped_interval_overflow` covers both functions with the fuzzed arguments (both hung before the fix) and checks that a sane `hop`/`windowID` is unaffected. Found by the AST fuzzer in the stress test of a CI run of https://github.com/ClickHouse/ClickHouse/pull/112930 (hung check): [Stress test (arm_debug) report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=112930&sha=1c5bf0b613146bd6666b6d2bd2be55f539f6c0d8&name_0=PR&name_1=Stress%20test%20%28arm_debug%29). Closes: https://github.com/ClickHouse/ClickHouse/issues/114605 Related: https://github.com/ClickHouse/ClickHouse/issues/61521 Related: https://github.com/ClickHouse/ClickHouse/pull/61523",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114607",
          "createdAt": "2026-08-13T09:11:49Z",
          "updatedAt": "2026-08-13T13:51:50Z",
          "timestamp": "2026-08-13T13:51:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e9bded3e6aee0a286c79",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113903",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113903",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add file IO and integrate new keeper storage",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Keeper can store data on disk now, in a custom LSM tree. It has similar performance to the previous storage. Use coordnation setting `use_new_storage = true` to enable, `storage_memory_only = false` to store files on disk, add `data_storage_path` or `data_storage_disk` to config to specify where to put the files. To write to S3, point `data_storage_disk` to a disk of type `s3_plain`. --- https://github.com/ClickHouse/ClickHouse/pull/107261 added new storage with memory-only mode. This PR * Adds ability to store data in files. The storage is not actually persistent, the directory is wiped on startup and re-created from snapshot, just like with previous rocksdb storage. So the files don't have things like headers and version numbers yet, we can add that later if we want faster startup. * Integrates the new storage into keeper. (This PR replaces https://github.com/ClickHouse/ClickHouse/pull/112378 )",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113903",
          "createdAt": "2026-08-08T00:05:18Z",
          "updatedAt": "2026-08-13T13:51:35Z",
          "timestamp": "2026-08-13T13:51:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "al13n321",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:7c8ff501a1acd5a8099e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114219",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114219",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.5: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114219",
          "createdAt": "2026-08-10T20:06:41Z",
          "updatedAt": "2026-08-13T13:49:39Z",
          "timestamp": "2026-08-13T13:49:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:924c47131006fce96d54",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111394",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111394",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fsync backup files and directories when writing a backup to local disk",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111320 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `BACKUP ... TO File(...)` / `Disk(...)` now fsyncs the backup data files, the `.backup` manifest and the containing directories before reporting `BACKUP_CREATED`, so an acknowledged backup to local storage survives power loss. Controlled by the new backup setting `fsync_backup_files` (default `true`). Object-storage destinations (`S3`/`Azure`) are unaffected. ### Description Fixes #111320. `BACKUP ... TO File()/Disk()` returned `BACKUP_CREATED` without issuing any `fsync`/`fdatasync` at the destination: not the data files, not the `.backup` manifest, and not the destination directories (there was no `fsync` anywhere in `src/Backups/`). On power loss after the acknowledgement the backup could be lost entirely or left torn, even though `BACKUP_CREATED` is exactly what an operator relies on before dropping the source data. Object-storage destinations were already durable (a completed upload is persisted server-side); only local `File()`/`Disk()` were affected. Report URL: https://github.com/ClickHouse/ClickHouse/issues/111320 (reproduced 3/3 with a `dm-flakey` power-loss simulation). Fix, gated on the new backup setting `fsync_backup_files` (default `true`), following the durability audit family (#68958 -> #111346, #111269 -> #111335): - Two writer hooks with a no-op default on `IBackupWriter`, overridden only by the local `File`/`Disk` writers (`S3`/`Azure`/`Memory`/`Null` inherit the no-op): `syncFileToDisk(file_name)` (fdatasync a written file, covering both the buffered and the native `fs::copy`/`IDisk::copyFile` paths) and `syncDirectoriesToDisk()` (fdatasync every directory the backup created, deepest-first, plus the backup root's parent, via `LocalDirectorySyncGuard` / `IDisk::getDirectorySyncGuard`). - Each data file is synced right after it is written in `BackupImpl::writeFile` (safe under the concurrent write path: each call fsyncs its own file). - In `BackupImpl::finalizeWriting` the `.backup` manifest (or, for archives, the archive file) is synced last, after all data files, so a persisted manifest never precedes its payload. Directory syncing runs for every writer, including the internal writers of `BACKUP ON CLUSTER` which write their own data files. Verified locally with ProfileEvents: `fsync_backup_files=1` issues `FileSync`/`DirectorySync` for the whole backup (data files + manifest + every nested directory); `fsync_backup_files=0` issues none (matching the previous behavior); the backup still restores correctly. Regression test `tests/queries/0_stateless/04412_backup_to_file_fsync.sh` asserts, via the `FileSync`/`DirectorySync` ProfileEvents of the `BACKUP` query in `system.query_log`, that the fsyncs are issued when `fsync_backup_files=1` and are absent when `fsync_backup_files=0`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111394",
          "createdAt": "2026-07-22T13:30:01Z",
          "updatedAt": "2026-08-13T15:03:54Z",
          "timestamp": "2026-08-13T15:03:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "jkartseva"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4152c82958704aed6fbe",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112605",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112605",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-apply LIMIT BY on the initiator for custom-key parallel replicas",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/111555 Follow-up to https://github.com/ClickHouse/ClickHouse/pull/111919, which fixed the `WITH FILL` half of #111555. This fixes the remaining `LIMIT BY` half. ### Problem Under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`), `LIMIT n BY` was applied per replica and the initiator never re-applied it, so it returned up to `n * replicas` rows per group: ```sql CREATE TABLE wf_min (g UInt16, k UInt32) ENGINE = MergeTree ORDER BY k; INSERT INTO wf_min SELECT number % 3, number FROM numbers(30); SELECT g FROM wf_min ORDER BY g LIMIT 2 BY g SETTINGS enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_range', parallel_replicas_custom_key = 'k', parallel_replicas_custom_key_range_upper = 30; -- returned 18 rows (6 per group); correct is 6 (0,0,1,1,2,2) ``` ### Root cause Custom-key parallel replicas splits rows across replicas by an arbitrary key and forces the initiator's input `from_stage` to `WithMergeableStateAfterAggregation(AndLimit)` (`PlannerJoinTree`), which tells the planner \"shards already finalized aggregation-stage processing\". The initiator therefore skips `LIMIT BY` (`Planner.cpp`, guarded by `!isFromAggregationState()`). That is correct for genuine sharding-key-aligned pushdown (each group lives on one shard), but the custom key does not align with the `LIMIT BY` key, so the per-replica `LIMIT BY` is not final. This is the same class of \"no initiator-side finalization over custom-key replica streams\" as the `WITH FILL` half fixed in #111919. ### Fix Re-apply `LIMIT BY` on the finalizing initiator (`isFinalizingStage()`) for custom-key parallel replicas, in addition to the normal `!isFromAggregationState()` case. On a replica that still emits a mergeable state, `LIMIT BY` runs as a preliminary that keeps `offset + length` rows and defers `OFFSET` to the initiator, so `LIMIT n OFFSET m BY` stays correct too. Custom-key parallel replicas is detected via a new `JoinTreeQueryPlan::is_parallel_replicas_custom_key` flag set only where the custom-key path is actually built — the replica custom-key filter, and the `Distributed` / MergeTree initiator dispatch in `PlannerJoinTree` — rather than the ambient `canUseParallelReplicasCustomKey()` setting (which is true whenever the profile enables a `custom_key_*` mode, even when the query used genuine sharding-key pushdown and never built a custom-key plan). This keeps `optimize_distributed_group_by_sharding_key` pushdown (e.g. `01244_optimize_distributed_group_by_sharding_key`) unaffected. The change is analyzer-only; the deprecated legacy interpreter's custom-key path is left unchanged (its custom-key handling is inconsistent across modes and a blanket re-application there regressed `custom_key_sampling` + `OFFSET`). ### Verification Built ClickHouse from an earlier revision of this branch and checked against a running 3-replica custom-key cluster: - `LIMIT 2 BY g` returns 6 rows (was 18) for both `custom_key_range` and `custom_key_sampling`. - `LIMIT 1 OFFSET 1 BY g` returns 3 rows (`0,1,2`), matching the non-distributed result — `OFFSET` applied once. - No regression: normal distributed `LIMIT 2 BY g` over `remote(...)` still returns the correct 6 rows. (The later review fix — switching from the ambient setting to the plan-tied flag — was validated by review/CI; the sandbox VM was recycled so it was not re-run locally. CI runs both `01244_optimize_distributed_group_by_sharding_key` and the new `04657_parallel_replicas_custom_key_limit_by`.) ### Test Added `tests/queries/0_stateless/04657_parallel_replicas_custom_key_limit_by.sql` (tagged `no-old-analyzer`), covering both custom-key modes and the `OFFSET` case. Confirmed it returns 18 (buggy) on unpatched and 6 on patched. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `LIMIT n BY` returning up to `n * replicas` rows per group under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`). `LIMIT BY` is now re-applied on the initiator over the merged replica streams. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112605",
          "createdAt": "2026-07-30T14:04:23Z",
          "updatedAt": "2026-08-13T14:32:19Z",
          "timestamp": "2026-08-13T14:32:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e88d6130c56ab44f688d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114622",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114622",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not apply DROP fault injection to a refreshable view's cleanup DROP",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/114420 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a refreshable materialized view leaking its rotated-out target table when `ignore_drop_queries_probability` is enabled. The `DROP` a refresh issues to clean up the previous target is a step of the refresh, not a `DROP` the user asked for, so the fault injection no longer applies to it. ### Description Follow-up to #114420, requested by @ tiandiwonder in https://github.com/ClickHouse/ClickHouse/pull/114420#discussion_r3772879738. `ignore_drop_queries_probability` makes a `DROP TABLE` silently do nothing (or become a `TRUNCATE`) so the stress suite exercises \"the table you dropped is still there\". It must only affect `DROP`s the user issued. **Root cause.** After a refresh swaps in a new target, `StorageMaterializedView::dropTempTable` drops the rotated-out one via `InterpreterDropQuery(drop_query, refresh_context)`. `createRefreshContext` never marks that context internal, so the gate treated the cleanup `DROP` as a user `DROP`. Its existing refreshable-view exemption does not cover this: that test inspects the table *being dropped*, which here is the inner target, not the view. Both refresh exits are affected. When it fires, the old target survives as `.tmp.inner_id.<uuid>` still holding a full copy of the view's data, outside the view's metadata and surviving restart. **The change.** `InterpreterDropQuery` gains an `internal` member mirroring the one `InterpreterCreateQuery` already has, `dropTempTable` sets it, and the shared `refresh_context` is untouched. Marking that context instead would break the refresh outright in a `Replicated` database: the publishing `RENAME` runs on it, and `DatabaseReplicated` rejects a non-initial query unless the interpreter also passes `flags.internal`, which neither the Rename nor the Drop interpreter did. That is why #114420's one-liner was safe there and is not here. `QueryFlags{ .internal = internal }` at the replicated enqueue is behaviour-neutral for every other `DROP`, since all its members default to false. The interpreter's other policy branches (`ON CLUSTER` dispatch, access and dependency checks) deliberately keep applying. **Validation.** New test `04887_refreshable_mv_cleanup_drop_not_ignored`: three positive arms (success path, failure path, `Replicated` database) leak on master and are clean with the fix; a fourth asserts a genuine user `DROP` is still skipped. 50 randomized runs plus 50 with `ignore_drop_queries_probability=0.2` were green, as were `04247`, `04796`, `04218` and `04327`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114622",
          "createdAt": "2026-08-13T11:43:28Z",
          "updatedAt": "2026-08-13T14:13:03Z",
          "timestamp": "2026-08-13T14:13:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3d0bfd71c833ebe4c176",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114671",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114671",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make release creation steps self-gating and drop fail-closed flag defaults",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/114472 --> Makes the release-creation steps self-gating so an idempotent operation decides for itself whether there is work to do, instead of the orchestrator gating on `is_recovery` / `is_late_recovery` flags with fail-closed `True` defaults. - `ci/jobs/scripts/create_release.py`: `push_release_tag` self-skips when `Git.tag_exists(release_tag)` — a recovery finds the tag already published and does nothing. `update_version_and_contributors_list` self-skips a superseded (late) recovery for a patch release, so it never rewrites the branch version backwards. - `ci/jobs/release_job.py`: the tag-push step now runs unconditionally and the patch bump runs for every patch release. The `is_recovery = True` / `is_late_recovery = True` defaults are removed; `is_recovery` is read only under `ok` (remaining reads short-circuit on it), and `is_late_recovery` is no longer needed at the top level (only the docker floating-tag block reads it, from its own release info). Behavior-preserving: the tag-exists skip derives from the same local `Git.tag_exists` that `prepare` used to compute `is_recovery`, and the late-recovery skip mirrors the old `not is_late_recovery` bump gate. Addresses the review that an idempotent step should not need an `is_recovery` gate. Stacked in content on #114472; until it merges the diff shows its commits too. The only new commit here is `Make release creation steps self-gating and drop fail-closed flag defaults`. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114671",
          "createdAt": "2026-08-13T17:33:22Z",
          "updatedAt": "2026-08-13T17:43:01Z",
          "timestamp": "2026-08-13T17:43:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "leshikus",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:56216898b6fec02c8f54",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101791",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101791",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "In case of trivial views, push whole outer query to shards.",
          "text": "### Changelog category: - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): In case of trivial views over distributed table push whole outer query to shards. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Description / Proposed Solution When a VIEW is defined over a Distributed table, ClickHouse traditionally executes it on the shards without enclosing outer query. This means filters and expressions declared in the outer query are evaluated on the coordinator after pulling raw data from shards. For views whose body is a plain SELECT (column references, *, or arbitrary expressions — but no aggregation, grouping, ordering, joins, window functions, or scalar subqueries) over a single Distributed table, we can do better: inline the view body as a subquery and hand the whole thing to StorageDistributed. Each shard then receives the full outer query with the view body inlined, evaluates it against its local table, and only ships the result back. A view qualifies as \"trivial\" if its inner query: - Has a single SELECT (no UNION) - Selects only column references, *, or expressions — but no window functions (require the full dataset) and no scalar subqueries in the SELECT list - Has no WITH, PREWHERE, GROUP BY, HAVING, QUALIFY, ORDER BY, LIMIT, LIMIT BY, DISTINCT, or ARRAY JOIN - Has no subqueries in the WHERE clause - Reads from exactly one table with no joins, no table functions, no FINAL, no SAMPLE - Is not a parameterized view and does not use SQL SECURITY DEFINER **The optimization can be disbaled by setting (enabled by default):** ``` SET optimize_trivial_view_pushdown_to_distributed = 0; ``` ### Example: Env setup: ``` create table x engine = MergeTree ORDER BY tuple() AS SELECT intDiv(number,100000) as a, number as b FROM numbers(1000000000); SET prefer_localhost_replica = 0; CREATE TABLE x_dist AS x ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), x); CREATE VIEW v_computed AS SELECT a + 1 AS x, b AS y FROM x_dist WHERE a != 0; ``` Performance: ``` :) SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: 6baedc55-c8e7-4b2b-9946-b0d828abaf25 ┌─plus(a, 1)─┬──────────sum(b)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └────────────┴─────────────────┘ 1 row in set. Elapsed: 17.199 sec. Processed 2.00 billion rows, 32.00 GB (116.29 million rows/s., 1.86 GB/s.) Peak memory usage: 38.58 MiB. :) SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 7c4f2854-3de2-4d40-b960-efcc44a7b26d ┌─────x─┬──────────sum(y)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └───────┴─────────────────┘ 1 row in set. Elapsed: 16.497 sec. Processed 2.00 billion rows, 32.00 GB (121.24 million rows/s., 1.94 GB/s.) Peak memory usage: 38.82 MiB. ``` Plan: ``` :) explain SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; EXPLAIN SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: e2d32d5f-7024-4e6a-b686-45c7c0d5f2ef ┌─explain──────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression (Before GROUP BY) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression (Before GROUP BY) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └──────────────────────────────────────────────────────────────────────────────────┘ :) explain SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; EXPLAIN SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 88213305-46a5-493e-862c-a80c675c9452 ┌─explain───────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes query planning/execution for a subset of views over `Distributed` tables and touches access checks/row policy enforcement and SQL SECURITY semantics, which can affect correctness and security-sensitive behavior. > > **Overview** > Adds a new default-on setting `optimize_trivial_view_pushdown_to_distributed` to inline *trivial* views over `Distributed` tables and push the full outer query down to shards, reducing coordinator-side filtering/processing and network transfer. > > Implements planner rewrites to swap the view table expression with an analyzed subquery, merge `FINAL`/`SAMPLE` modifiers, and preserve semantics by suppressing pushdown when the outer query contains non-deterministic functions, while also explicitly handling SQL SECURITY modes, row-policy injection/logging, and column-pruned privilege checks. > > Extends integration/stateless tests to cover modifier propagation, non-determinism suppression, row-policy enforcement, SQL SECURITY behavior, and interactions with `max_rows_to_read_leaf`. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 583e1e6c0e8e25081391d7a07af086c6f9888c6f. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101791",
          "createdAt": "2026-04-04T19:32:01Z",
          "updatedAt": "2026-08-13T17:42:24Z",
          "timestamp": "2026-08-13T17:42:24Z",
          "metrics": {
            "reactions": 3,
            "comments": 15
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "simonmichal",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:badc1d7065dd986c22c7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109881",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109881",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Read Keeper changelogs in parallel at startup",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Speed up Keeper startup by reading multiple changelog files concurrently instead of serially, controlled by new settings `log_startup_read_max_streams` and `log_startup_read_buffer_size`. ### Documentation entry: - [ ] Documentation is written",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109881",
          "createdAt": "2026-07-09T10:41:36Z",
          "updatedAt": "2026-08-13T17:41:41Z",
          "timestamp": "2026-08-13T17:41:41Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-improvement",
            "submodule changed",
            "comp-keeper"
          ],
          "author": "antonio2368",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:a115c1258e5d53829ccf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114466",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114466",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: internationalize master",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114466",
          "createdAt": "2026-08-12T10:49:59Z",
          "updatedAt": "2026-08-13T17:41:19Z",
          "timestamp": "2026-08-13T17:41:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 47
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "locadex-agent[bot]",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:21d09d97b1903ad78916",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111973",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111973",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Let read-in-order propagate through SpillingHashJoin",
          "text": "`SpillingHashJoin::hasDelayedBlocks` was hardcoded to `true`, even in the `IN_MEMORY_JOIN` state where nothing is ever delayed. That is the flag gating the read-in-order-through-join and top-k-through-join optimizations, so wrapping a hash join for auto-spilling silently disabled both — and `max_bytes_ratio_before_external_join` defaults to `0.5`, which wraps every hash join. The in-tree comment in `topKThroughJoin.cpp` already describes this as the steady state. `IJoin` documents that \"SpillingHashJoin overrides `keepLeftPipelineInOrder` to forbid switching to GraceHashJoin at runtime\", but no such override existed, so the escape hatch the comment describes was never implemented. This implements it. `keepLeftPipelineInOrder` now pins the join to its in-memory algorithm, and `hasDelayedBlocks` reports `false` from that point on. Pinning is required for **correctness**, not just speed: once the plan drops a sort because the join preserves the left order, a later switch to `GraceHashJoin` would scatter rows by hash and silently return them in the wrong order. The optimizer asks before it commits (`findReadingStep` checks feasibility, and `keepLeftPipelineInOrder` is only called later, if reading in order actually turns out to be possible). So a new `IJoin::canKeepLeftPipelineInOrder` carries the question \"would you preserve the order if I asked?\", defaulting to `!hasDelayedBlocks()` so every other join keeps its current answer. `SpillingHashJoin` answers yes while still reporting delayed blocks, and only stops reporting them once actually pinned — if the optimization turns out not to apply, nothing is pinned and the delayed-block transforms are still built. **Trade-off:** a pinned join can no longer spill, so it holds the whole right side in memory and the memory tracker enforces the limit, exactly as it would with no auto-spill threshold configured. Since dropping the sort and then spilling would be a wrong-results bug, the only alternative is the conservative status quo of never propagating read-in-order through these joins. Both are now reachable: the new setting `query_plan_read_in_order_through_spilling_join` (default `1`) turns the optimization off again, and it is registered in `SettingsChangesHistory` with `previous_value = false`, so `compatibility` set to a version before 26.8 restores the old behavior. `ConcurrentHashJoin` does not spill on its own (its threshold argument only bounds preallocation), so `switchToGraceHashJoin` is the single place the invariant has to hold. The same gate applies to the first-pass `topKThroughJoin`: an `ORDER BY left_key LIMIT n` over a spill-capable `LEFT JOIN` now steps aside for the second-pass read-in-order plan instead of materializing a pushed-down `Sort` and `Limit`, and goes back to pushing them down when the setting is `0`. The handoff keeps the deferral's pre-existing conditions — most notably, `query_plan_join_swap_table` must be explicitly `false`, because under the default `auto` a later optimization may swap the join sides and invalidate the read-in-order plan; with `auto`, such queries keep the pushed-down `Sort` and `Limit`. Making this the default path exposed a pre-existing problem in read-in-order through a `JOIN`, which #110283 pinned down with `04516_join_order_estimation_pruned_parts` a few hours before this branch entered the merge queue. `max_rows_to_read` with `read_overflow_mode = 'throw'` is not checked against the rows a query reads, but against the rows the reading steps announce up front: `ReadProgressCallback::onProgress` substitutes `progress.total_rows_to_read` for `progress.read_rows` whenever the estimate is the larger of the two, and `ReadFromMergeTree` announces `min(part rows, InputOrderInfo::limit)`. `buildSortingDAG` dropped the limit at every `JoinStep`, so a plan reading in order through a `JOIN` always announced whole parts, and `ORDER BY left_key LIMIT 1` over a `LEFT JOIN` that reads 20 rows announced 1010 and failed `max_rows_to_read = 100`. Dropping the limit is right for a join that can filter the left stream, but a `LEFT ALL`/`LEFT ANY` join emits at least one row for every left row, so `n` output rows need at most `n` left rows - duplication only makes fewer left rows necessary. The limit now survives those joins and is dropped everywhere else (`INNER`, `SEMI`, `ANTI`, ...). `InputOrderInfo::limit` never truncates a read - it feeds the announcement, the read-pool task size, the `take_full_part` heuristic and `use_buffering` - so this cannot change results. The behavior was already reachable on `master` with `max_bytes_ratio_before_external_join = 0`, which makes `topKThroughJoin` defer to the second pass exactly as it now does by default. Related: https://github.com/ClickHouse/ClickHouse/pull/111248 Related: https://github.com/ClickHouse/ClickHouse/pull/111972 Related: https://github.com/ClickHouse/ClickHouse/pull/110283 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reading in order through a JOIN now also works when a hash join has an automatic spill-to-disk threshold configured (the default since 26.5). When this optimization applies, the join stays in memory so it preserves the left-side order; set `query_plan_read_in_order_through_spilling_join = 0` to restore the previous behavior, where such joins are not used for read-in-order and remain free to spill. As part of this, an `ORDER BY ... LIMIT` over a `LEFT JOIN` read in order no longer reports the whole table as the number of rows it intends to read, so it no longer trips `max_rows_to_read` with `read_overflow_mode = 'throw'` for a read that stops early.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111973",
          "createdAt": "2026-07-26T19:07:03Z",
          "updatedAt": "2026-08-13T17:41:16Z",
          "timestamp": "2026-08-13T17:41:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:7252fbda259b80152de7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114656",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114656",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "docs: convert merge loop code block to Steps component",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replace the numbered code block describing the background merge loop on the AWS performance page with a `<Steps>` component for clearer rendering. <!-- mintlify-agent-attribution --> --- Generated by Mintlify Agent. Requested by: shaun.struwig@clickhouse.com via Slack Mintlify session: slack_1785790244.316179_D0B0U0V3Z88",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114656",
          "createdAt": "2026-08-13T15:16:40Z",
          "updatedAt": "2026-08-13T17:41:13Z",
          "timestamp": "2026-08-13T17:41:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "mintlify[bot]",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f8a020bcbd6e5ad5c03e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114543",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114543",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix spelling of 'prefetches' and 'prefetched' in documentation",
          "text": "Corrected spelling of 'prefetches' and 'prefetched' in multiple sections. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114543",
          "createdAt": "2026-08-12T21:18:03Z",
          "updatedAt": "2026-08-13T17:41:12Z",
          "timestamp": "2026-08-13T17:41:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "linhgiang24",
          "state": "closed",
          "assignees": [
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d1b92b100af88650ea46",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114300",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114300",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "TimeSeries: store all tags in the `tags` column",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): TimeSeries: store all tags in the `tags` column The `tags` column of the tags target table now contains all the tags, including the `__name__` tag with the metric name and the tags with dedicated columns from the `tags_to_columns` setting, so the map alone fully identifies a time series. The default id generator now hashes just `tags`, i.e. now it looks like `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(tags)))` (instead of `tuple(sipHash64(metric_name), reinterpretAsUUID(sipHash128(metric_name, all_tags)))` )",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114300",
          "createdAt": "2026-08-11T10:29:30Z",
          "updatedAt": "2026-08-13T17:41:11Z",
          "timestamp": "2026-08-13T17:41:11Z",
          "metrics": {
            "reactions": 1,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-promql"
          ],
          "author": "vitlibar",
          "state": "closed",
          "assignees": [
            "nikitamikhaylov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:47cf851fafe0fc8595eb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114670",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114670",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: expand the Managed Postgres autoscaling documentation",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ## Summary - Expand the one-line Autoscaling section in the Managed Postgres scaling docs with the behavior sourced from the Ubicloud codebase: the 85% storage notification, the 90% automatic scale-up, and the 95% maintenance-window bypass - Document what autoscaling means for the cutover and client connections: same process as a manual instance change, connections dropped and in-flight transactions rolled back, DNS repointed to the new primary under the same hostname - Document read-only mode: free-space trigger and recovery thresholds per disk size, the error writes receive, and automatic recovery after the scale-up - Add a worked example scenario for a 1024 GB instance scaling to 2048 GB",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114670",
          "createdAt": "2026-08-13T16:59:03Z",
          "updatedAt": "2026-08-13T17:41:05Z",
          "timestamp": "2026-08-13T17:41:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "amogiska",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:a7dff0fa1411813149a7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114273",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114273",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _This is a test-only PR — no source code changes. Please review: test quality, whether the claimed coverage gaps are real, and whether test output makes sense._ Adds test coverage for 1 untested code path, found during automated review of [PR #109744](https://github.com/ClickHouse/ClickHouse/pull/109744). That PR (1) The PR changes `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:403-421): the whole-column `input_format_null_as_default` short-circuit (`SerializationNullable::deserializeNullAsDefaultOrNestedTextCSV`) is now skipped for a bare, non-empty `Tuple` whose `tuple_ **1. A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested** `src/Processors/Formats/Impl/CSVRowInputFormat.cpp:413`, `src/DataTypes/Serializations/SerializationTuple.cpp:683` **Risk:** `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:409-413) now skips the whole-column `null_as_default` short-circuit for a bare `Tuple`, so the leading field is handed to `SerializationTuple::deserializeTextCSV`, whose per-element arm at src/DataTypes/Serializations/SerializationTuple.cpp:683 applies `null_as_default` to the FIRST ELEMENT — and when that element is itself a `Tuple`, one field stands for the whole nested element. This is exactly the sentence … **What's unique vs PR tests:** The PR's 04405_csv_tuple_leading_null_null_as_default.sql covers a leading `\\N` for a SCALAR first element and, for nested tuples, only `Tuple(Nullable(Int32), Tuple(Int32, Int32))` with `\\N,2,3` where the nested tuple's own fields are present; its comment states `a null inside a nested tuple is read as that whole nested element and is not covered`. This test covers precisely that: the `\\N` field is the field of a nested `Tuple` element (and of a `Nullable(Tuple)` element), the resulting one- … [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/e4ec131d-5685-4047-bd72-b61c347ba13b) cc @groeneai (author of #109744), @Avogar (merged/approved #109744) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114273",
          "createdAt": "2026-08-11T05:31:50Z",
          "updatedAt": "2026-08-13T17:40:23Z",
          "timestamp": "2026-08-13T17:40:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f7a256b9377a16acd25e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111394",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111394",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fsync backup files and directories when writing a backup to local disk",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111320 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `BACKUP ... TO File(...)` / `Disk(...)` now fsyncs the backup data files, the `.backup` manifest and the containing directories before reporting `BACKUP_CREATED`, so an acknowledged backup to local storage survives power loss. Controlled by the new backup setting `fsync_backup_files` (default `true`). Object-storage destinations (`S3`/`Azure`) are unaffected. ### Description Fixes #111320. `BACKUP ... TO File()/Disk()` returned `BACKUP_CREATED` without issuing any `fsync`/`fdatasync` at the destination: not the data files, not the `.backup` manifest, and not the destination directories (there was no `fsync` anywhere in `src/Backups/`). On power loss after the acknowledgement the backup could be lost entirely or left torn, even though `BACKUP_CREATED` is exactly what an operator relies on before dropping the source data. Object-storage destinations were already durable (a completed upload is persisted server-side); only local `File()`/`Disk()` were affected. Report URL: https://github.com/ClickHouse/ClickHouse/issues/111320 (reproduced 3/3 with a `dm-flakey` power-loss simulation). Fix, gated on the new backup setting `fsync_backup_files` (default `true`), following the durability audit family (#68958 -> #111346, #111269 -> #111335): - Two writer hooks with a no-op default on `IBackupWriter`, overridden only by the local `File`/`Disk` writers (`S3`/`Azure`/`Memory`/`Null` inherit the no-op): `syncFileToDisk(file_name)` (fdatasync a written file, covering both the buffered and the native `fs::copy`/`IDisk::copyFile` paths) and `syncDirectoriesToDisk()` (fdatasync every directory the backup created, deepest-first, plus the backup root's parent, via `LocalDirectorySyncGuard` / `IDisk::getDirectorySyncGuard`). - Each data file is synced right after it is written in `BackupImpl::writeFile` (safe under the concurrent write path: each call fsyncs its own file). - In `BackupImpl::finalizeWriting` the `.backup` manifest (or, for archives, the archive file) is synced last, after all data files, so a persisted manifest never precedes its payload. Directory syncing runs for every writer, including the internal writers of `BACKUP ON CLUSTER` which write their own data files. Verified locally with ProfileEvents: `fsync_backup_files=1` issues `FileSync`/`DirectorySync` for the whole backup (data files + manifest + every nested directory); `fsync_backup_files=0` issues none (matching the previous behavior); the backup still restores correctly. Regression test `tests/queries/0_stateless/04412_backup_to_file_fsync.sh` asserts, via the `FileSync`/`DirectorySync` ProfileEvents of the `BACKUP` query in `system.query_log`, that the fsyncs are issued when `fsync_backup_files=1` and are absent when `fsync_backup_files=0`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111394",
          "createdAt": "2026-07-22T13:30:01Z",
          "updatedAt": "2026-08-13T17:40:07Z",
          "timestamp": "2026-08-13T17:40:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "jkartseva"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c0b3d731a4c9ccc0cf20",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114212",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114212",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain the required column order when a dictionary QUERY returns columns in the wrong order",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 --> Closes: https://github.com/ClickHouse/ClickHouse/issues/113935 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): When a dictionary with a PostgreSQL source and a custom `QUERY` returns its columns in an order that does not match the dictionary structure, the error now names the destination dictionary column and the column order the query must return, instead of only reporting that a value could not be parsed. ### Description A `RANGE_HASHED` dictionary over `SOURCE(POSTGRESQL(... QUERY '...'))` failed with `Cannot parse LocalDate: 97955`, where `97955` was an unrelated `id` column's value: the DDL was valid, only the column order was wrong, and the message pointed at the data. A dictionary's expected source structure is keys-first: key(s), `RANGE MIN`, `RANGE MAX`, then the attributes. ClickHouse aliases every column when it builds the `SELECT` itself, but a user `QUERY` passes through verbatim and the result is read strictly by position, so `id` was deserialized into the `Date` attribute `contract_time`. On a conversion failure `PostgreSQLSource` now names the destination column and its position, and for a custom `QUERY` states the required order, derived from the dictionary's own structure so it suits every layout (the `RANGE` wording appears only for range dictionaries). Successful loads are untouched, and `BAD_ARGUMENTS` is preserved because `MaterializedPostgreSQL` relies on it. <details><summary>The message for the reported case</summary> ``` Cannot parse PostgreSQL value '97955' as Date: Cannot parse LocalDate: 97955: while reading column 2 of the result into `contract_time`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `meter_no`, `contract_time`, `end_date`, `id` (the key column(s) first, then the RANGE MIN and RANGE MAX columns, then the remaining attributes) ``` For a `FLAT` dictionary the same failure omits the `RANGE` clause: ``` Cannot parse PostgreSQL value 'abc' as Int64: Could not convert string to l: 'abc': while reading column 2 of the result into `num`: the columns of a dictionary QUERY are taken by position, so it must return them in this order: `id`, `num`, `name` ``` </details> A new integration test covers the range case, the documented order, and a `FLAT` dictionary asserting no `RANGE` wording appears; it fails on master. `04401_composite_key_dict_non_leading_key` gains the range shape over the ClickHouse source, pinning its by-name reordering against regression. Scope, since the mechanism is wider than the fix. The MySQL dictionary source (and likely XDBC and Cassandra) maps positionally too, but its parse errors do not funnel through one place, so folding it in grows the diff; say the word and I will extend it. Two cases stay untouched: an unnamed-column query still matches positionally and can silently return a wrong value, and a short row silently defaults trailing columns. Matching by name is possible, not ruled out: a `LIMIT 0` probe, or projecting the expected names inside the streamed query. I avoided it because it silently changes behaviour for every dictionary relying on positional matching, and arbitrary queries can return duplicate or unnamed columns.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114212",
          "createdAt": "2026-08-10T19:51:47Z",
          "updatedAt": "2026-08-13T17:39:53Z",
          "timestamp": "2026-08-13T17:39:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8f634d04aba43109d782",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114624",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114624",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Find a LowCardinality needle equal to the type's default value",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes `has`, `indexOf`, `countEqual`, `mapContainsKey`, `mapContainsValue` and `Map` subscript returning \"not found\" for a constant `LowCardinality` needle equal to the element type's default value, such as an empty `String` or a zero number. ### Description **The problem.** A constant needle equal to the element type's default value is never found in an `Array(LowCardinality(T))` or a `Map` with a `LowCardinality` key. No exception, so a `WHERE` on such a predicate silently drops rows: ```sql CREATE TABLE t (a Array(LowCardinality(String))) ENGINE = Memory; INSERT INTO t VALUES (['', 'a']); SELECT has(a, '') FROM t; -- 0, expected 1 ``` Likewise for `0` over the numeric types, a NUL-padded `FixedString` and the `1970-01-01` `Date`, while `m['']` returns `''` instead of the stored value. Any non-default needle is correct. Reproduces on 26.5 to 26.7. **Root cause.** A `LowCardinality` dictionary reserves prefix slots for the default and NULL values, and `ReverseIndex` is built with `num_prefix_rows_to_skip`, so the reserved slot is never indexed. `ColumnUnique::uniqueInsertData` compensates for that on the write path; the read path had no counterpart. **The change.** `ColumnUnique::getOrFindValueIndex` now performs the same default-slot match as the write path. `ReverseIndex` is untouched, so no write behaviour moves. That slot becoming reachable brings two equality details. The lookup casts the constant into the element type without reporting loss, so `UInt64(256)` arrived as `UInt8(0)` and would match the default; that slot now answers only if the element type can represent the constant. And a dictionary can hold `-0.0` and `0.0` as separate entries while the shortcut has room for one index, so a zero constant over a float element type is left to the value-comparing path, as `indexOfAssumeSorted` already is. The new test covers both call sites. Lookup timing is unchanged. **Overlap with #112953.** That open PR declines this same shortcut for every float element type, subsuming the float-zero decline here, and the two conflict textually; whichever merges second should keep the broader decline and drop this one, collapsing the duplicated representability predicate to one copy. Non-float types are independent of it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114624",
          "createdAt": "2026-08-13T11:55:50Z",
          "updatedAt": "2026-08-13T17:39:52Z",
          "timestamp": "2026-08-13T17:39:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7a0dd57ccf025719bbbd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112605",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112605",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-apply LIMIT BY on the initiator for custom-key parallel replicas",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/111555 Follow-up to https://github.com/ClickHouse/ClickHouse/pull/111919, which fixed the `WITH FILL` half of #111555. This fixes the remaining `LIMIT BY` half. ### Problem Under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`), `LIMIT n BY` was applied per replica and the initiator never re-applied it, so it returned up to `n * replicas` rows per group: ```sql CREATE TABLE wf_min (g UInt16, k UInt32) ENGINE = MergeTree ORDER BY k; INSERT INTO wf_min SELECT number % 3, number FROM numbers(30); SELECT g FROM wf_min ORDER BY g LIMIT 2 BY g SETTINGS enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_range', parallel_replicas_custom_key = 'k', parallel_replicas_custom_key_range_upper = 30; -- returned 18 rows (6 per group); correct is 6 (0,0,1,1,2,2) ``` ### Root cause Custom-key parallel replicas splits rows across replicas by an arbitrary key and forces the initiator's input `from_stage` to `WithMergeableStateAfterAggregation(AndLimit)` (`PlannerJoinTree`), which tells the planner \"shards already finalized aggregation-stage processing\". The initiator therefore skips `LIMIT BY` (`Planner.cpp`, guarded by `!isFromAggregationState()`). That is correct for genuine sharding-key-aligned pushdown (each group lives on one shard), but the custom key does not align with the `LIMIT BY` key, so the per-replica `LIMIT BY` is not final. This is the same class of \"no initiator-side finalization over custom-key replica streams\" as the `WITH FILL` half fixed in #111919. ### Fix Re-apply `LIMIT BY` on the finalizing initiator (`isFinalizingStage()`) for custom-key parallel replicas, in addition to the normal `!isFromAggregationState()` case. On a replica that still emits a mergeable state, `LIMIT BY` runs as a preliminary that keeps `offset + length` rows and defers `OFFSET` to the initiator, so `LIMIT n OFFSET m BY` stays correct too. Custom-key parallel replicas is detected via a new `JoinTreeQueryPlan::is_parallel_replicas_custom_key` flag set only where the custom-key path is actually built — the replica custom-key filter, and the `Distributed` / MergeTree initiator dispatch in `PlannerJoinTree` — rather than the ambient `canUseParallelReplicasCustomKey()` setting (which is true whenever the profile enables a `custom_key_*` mode, even when the query used genuine sharding-key pushdown and never built a custom-key plan). This keeps `optimize_distributed_group_by_sharding_key` pushdown (e.g. `01244_optimize_distributed_group_by_sharding_key`) unaffected. The change is analyzer-only; the deprecated legacy interpreter's custom-key path is left unchanged (its custom-key handling is inconsistent across modes and a blanket re-application there regressed `custom_key_sampling` + `OFFSET`). ### Verification Built ClickHouse from an earlier revision of this branch and checked against a running 3-replica custom-key cluster: - `LIMIT 2 BY g` returns 6 rows (was 18) for both `custom_key_range` and `custom_key_sampling`. - `LIMIT 1 OFFSET 1 BY g` returns 3 rows (`0,1,2`), matching the non-distributed result — `OFFSET` applied once. - No regression: normal distributed `LIMIT 2 BY g` over `remote(...)` still returns the correct 6 rows. (The later review fix — switching from the ambient setting to the plan-tied flag — was validated by review/CI; the sandbox VM was recycled so it was not re-run locally. CI runs both `01244_optimize_distributed_group_by_sharding_key` and the new `04657_parallel_replicas_custom_key_limit_by`.) ### Test Added `tests/queries/0_stateless/04657_parallel_replicas_custom_key_limit_by.sql` (tagged `no-old-analyzer`), covering both custom-key modes and the `OFFSET` case. Confirmed it returns 18 (buggy) on unpatched and 6 on patched. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `LIMIT n BY` returning up to `n * replicas` rows per group under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_range'` / `'custom_key_sampling'`). `LIMIT BY` is now re-applied on the initiator over the merged replica streams. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-3e3c9bb8-0278-4d60-aba9-69fdb160110a\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112605",
          "createdAt": "2026-07-30T14:04:23Z",
          "updatedAt": "2026-08-13T17:39:00Z",
          "timestamp": "2026-08-13T17:39:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:add241b8c50fa6a21c4f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114628",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114628",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for DETACH DATABASE to release tables in three integration tests",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/93064 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/issues/93064 Three integration tests intermittently fail with `Code: 219 ... Database <db> cannot be detached, because some tables are still in use. Retry later.` thrown from `DatabaseAtomic::assertCanBeDetached` (`src/Databases/DatabaseAtomic.cpp:521`). Root cause: non-SYNC `DETACH DATABASE` is best-effort by contract, and these tests assume it is atomic. `assertCanBeDetached` throws if any table of the database still has a live refcount. The engine already has the deterministic wait, `waitDetachedTableNotInUse`, but `executeToDatabaseImpl` calls it only under `query.sync` (`src/Interpreters/InterpreterDropQuery.cpp:779-790`). Stateless CI installs `tests/config/users.d/database_atomic_drop_detach_sync.xml` (`database_atomic_wait_for_drop_and_detach_synchronously=1`); the integration helpers install no such profile and run at the server default of 0. Hence the failures are confined to the integration suite, and 16 of the 17 integration occurrences in 180 days are sanitizer builds, where the holder's window is wider. Code 219 is intended behaviour of the asynchronous form and is pinned in-tree: `01107_atomic_db_detach_attach.sh:19` sets the setting to 0 to provoke it and asserts it. So this is a test-side defect, and the change asks for the wait the engine already implements rather than altering it. I added `SYNC` to the five `DETACH DATABASE` statements with observed CI failures: `test_drop_replica` (11 hits/180d), `test_replicated_table_structure_alter` (4), and `test_drop_database_replica:192` (2, both stacks confirm that line is the thrower). All 21 non-SYNC sites under `tests/integration/` were enumerated; the other 16 are excluded because they have zero CIDB hits in 180 days, already pass `database_atomic_wait_for_drop_and_detach_synchronously`, or sit inside the existing `detach_database_with_retry` helper, whose docstring documents a different holder. Validation: with a holder injected the way `01107` does it, the DETACH returns 219 without `SYNC` and succeeds with it, 10/10 on `Atomic` and on `Replicated`. The three tests pass 42/42 runs. Each `DETACH` in `test_drop_replica` measures ~0.115 s over 25 measurements, so the wait costs nothing when no table is held, and it is cancellable and shutdown-aware. CI report for the master failure: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=91b700711adf0a07172e1046c8d8f013f1e0b82c&name_0=MasterCI&name_1=Integration%20tests%20%28amd_tsan%2C%206%2F6%29",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114628",
          "createdAt": "2026-08-13T12:36:13Z",
          "updatedAt": "2026-08-13T17:38:59Z",
          "timestamp": "2026-08-13T17:38:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:12fedb96a970a177a6a4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114645",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114645",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Speed up `IN (subquery)` set building by pre-deduplicating each `MergeTree` partition independently",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/108326 Related: https://github.com/ClickHouse/ClickHouse/pull/105126 The set for `IN (subquery)` is built by a single `CreatingSetsTransform`: all streams of the subquery are merged into one and every row is hashed serially, no matter how many threads read the data. If the partition expression of the subquery's table is a function of the subquery's output columns (the set is keyed on all of them), the reading will now emit each partition through a single port and each stream is deduplicated independently before the filling transform. Because a key then lives in exactly one stream, per-stream deduplication is complete, and the single filling transform only hashes unique rows — the serial part of the build shrinks from all rows to distinct rows, and the deduplication itself runs in parallel. ```sql CREATE TABLE t (a UInt64) ENGINE = MergeTree ORDER BY tuple() PARTITION BY a % 8; INSERT INTO t SELECT number % 1000000 FROM numbers(100000000); OPTIMIZE TABLE t FINAL; EXPLAIN PIPELINE SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t) SETTINGS allow_creating_set_partitions_independently = 1, max_threads = 8; ``` ```response (CreatingSets) DelayedPorts 9 → 8 (Expression) ExpressionTransform × 8 (Aggregating) Resize 1 → 8 AggregatingTransform (Expression) ExpressionTransform (Filter) FilterTransform (ReadFromSystemNumbers) NumbersRange 0 → 1 (CreatingSet) CreatingSetsTransform <- the single filling transform now hashes ~1M unique rows instead of 100M Resize 8 → 1 DistinctTransform × 8 <- new: parallel pre-deduplication on partition-disjoint streams (Expression) ExpressionTransform × 8 (ReadFromMergeTree) MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) × 8 0 → 1 <- per-partition reading (8 partitions → 8 streams) ``` On the table above (100M rows, 1M distinct keys, 8 balanced partitions; 64-core machine, average of 3 runs after 2 warm-ups): | query | off | on | speedup | |------------------------------------------------------------|--------|--------|---------| | `SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t)` | 0.643s | 0.113s | 5.7× | ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Speed up set building for `IN (subquery)` on partitioned `MergeTree` tables by keeping each partition's rows within a single stream and deduplicating each stream independently, so the single set-filling transform — previously hashing every row serially — only sees unique rows. This applies when the partition expression is a deterministic function of the subquery's output columns. The optimization is not applied when the largest partition holds more than twice the rows of the average partition; the new setting `force_creating_set_partitions_independently` (disabled by default) bypasses this check. Controlled by the new setting `allow_creating_set_partitions_independently` (enabled by default).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114645",
          "createdAt": "2026-08-13T14:04:03Z",
          "updatedAt": "2026-08-13T17:38:43Z",
          "timestamp": "2026-08-13T17:38:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "nihalzp",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3a136cd2eb3163b4e087",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114644",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114644",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Revert \"NATS: add inline credentials setting\"\"",
          "text": "Reverts https://github.com/ClickHouse/ClickHouse/pull/114178, restoring the `nats_credentials` setting of the `NATS` table engine (originally added in https://github.com/ClickHouse/ClickHouse/pull/110733), and fixes the reason of the original revert: the possibility of referencing arbitrary server paths through `nats_credential_file`. `nats_credential_file` is a path on the server filesystem: the server opens it with its own privileges, and during authentication the credentials are sent to `nats_url`, which comes from the same query. So a path taken from SQL lets anyone who can define a `NATS` source probe the local filesystem (the error text distinguishes a missing file, a permission error, and a file without a seed), and exfiltrate files the server can read to a NATS server they control (the part of the file before the seed line is sent verbatim in the `CONNECT` frame's `jwt` field). See the analysis in https://github.com/ClickHouse/ClickHouse/pull/110733#discussion_r3750203599. This hole predates #110733 (`nats_credential_file` exists since 24.2), but with the inline `nats_credentials` setting restored, the path form is no longer needed in SQL at all. The restriction, modelled on `StorageMySQL::getSSLParams` (`e700bbec4c84c585`): `nats_credential_file` is accepted only from a named collection defined in the server configuration file, or as `nats.credential_file` in the server configuration itself. Every SQL spelling — the `SETTINGS` clause, an engine-argument override `NATS(collection, nats_credential_file = ...)`, or a named collection created by `CREATE NAMED COLLECTION` — throws `BAD_ARGUMENTS` with a message directing to `nats_credentials`. Replacing a configured path with inline `nats_credentials` from the query remains allowed (the path itself is not used then). Loading from previously-validated metadata (server startup, force-restore, and short-syntax `ATTACH`) is exempt via `isLoadingFromExistingMetadata`, so existing tables keep working after an upgrade; a user-issued full `ATTACH TABLE` query is still checked. `ALTER TABLE ... MODIFY SETTING` is not a bypass: the `NATS` engine does not support settings alter. Tests: a new stateless test `04891_nats_credential_file_path_restriction` covers all rejected SQL spellings and the accepted configuration-file sources, using a `NATS` named collection added to the stateless-test server configuration (`tests/config/config.d/named_collection.xml`, with `nats_url = '127.0.0.1:1'`, so passing the validation surfaces as `CANNOT_CONNECT_NATS`); `04665_nats_credentials_named_collection` is updated — the `nats_credential_file` query-override direction is now rejected. Related: https://github.com/ClickHouse/ClickHouse/pull/110733 Related: https://github.com/ClickHouse/ClickHouse/pull/114178 Related: https://github.com/ClickHouse/ClickHouse/issues/85213 ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `NATS` table engine accepts credentials inline in the new `nats_credentials` setting (the same payload as a `.creds` file), and no longer accepts `nats_credential_file` from SQL: the path is a reference to a file on the server filesystem, which the server opens with its own privileges, so it can only be specified in a named collection defined in the server configuration file, or as `nats.credential_file` in the server configuration itself. Tables created before this restriction keep working after an upgrade.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114644",
          "createdAt": "2026-08-13T14:03:24Z",
          "updatedAt": "2026-08-13T17:38:08Z",
          "timestamp": "2026-08-13T17:38:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:563e1ccebb7cd43a4d52",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111830",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111830",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject truncated/incomplete AI text function responses",
          "text": "The AI text functions (`aiGenerate`, `aiClassify`, `aiExtract`, `aiTranslate`) could silently return a truncated answer. When a provider stops early — most commonly by hitting the `max_tokens` limit — it reports this in the response, but the shared base class `FunctionBaseAI` never inspected the signal and returned whatever partial text came back. For example, `aiGenerate('Write three sentences about the ocean.', map(..., 'max_tokens', '5'))` returned the fragment `The ocean covers more than` with no error. This change normalizes each provider's native stop reason (OpenAI `finish_reason`, Anthropic `stop_reason`) into a canonical `FinishReason` enum, so the base class can make a single completeness decision without knowing the provider dialect: - `Truncated` (token/context limit) → throws `AI_PROVIDER_RESPONSE_TRUNCATED`. - `ContentFilter` (OpenAI `content_filter`, Anthropic `refusal`) and `ToolCall` → throws `AI_PROVIDER_RESPONSE_INCOMPLETE`. - `Complete` (natural end, or a caller stop sequence such as Anthropic `stop_sequence`) and `Unknown` (unrecognized reason) are accepted, so benign non-`stop` reasons are not misclassified as truncation. The rejection is thrown inside the existing per-row `try`, so it is non-retriable (retrying would hit the same limit) and honors `ai_function_throw_on_error`: with `1` the exception propagates; with `0` the row becomes the column default. `aiEmbed` is unaffected (embeddings have no finish reason). Integration tests in `test_ai_functions` cover truncation (throw + graceful), content-filter, an accepted unknown reason, and the Anthropic `stop_sequence` (must not throw) and `max_tokens` (must throw) cases. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): AI text functions (`aiGenerate`, `aiClassify`, `aiExtract`, `aiTranslate`) now reject truncated or otherwise incomplete provider responses (e.g. when the model hits the `max_tokens` limit) instead of silently returning partial output. Behavior follows the `ai_function_throw_on_error` setting.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111830",
          "createdAt": "2026-07-24T17:24:29Z",
          "updatedAt": "2026-08-13T17:37:56Z",
          "timestamp": "2026-08-13T17:37:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "george-larionov",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:777f547aa7d9540709fb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114642",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114642",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: document filtering for `system.query_log` initial queries",
          "text": "Document the recommended filters for analyzing `system.query_log`. The page now explains that `is_initial_query = 1` selects top-level client queries, while `initial_query_id` correlates the full cascade of a distributed query across nodes. The existing basic example also filters for initial queries so child executions are not treated as separate client queries. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114642",
          "createdAt": "2026-08-13T13:10:53Z",
          "updatedAt": "2026-08-13T17:37:19Z",
          "timestamp": "2026-08-13T17:37:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation",
            "pr-autogenerated-docs"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7c3dbf548e3f763c8716",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112816",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112816",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow running queries detached from client session",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New setting `run_query_in_background`. The server accepts the query, immediately returns an empty result, and runs it to completion regardless of what happens to the connection. The result is discarded. Track the query by its `query_id` in `system.processes` and `system.query_log`. Intended for long queries like `INSERT ... SELECT`, `CREATE TABLE … AS SELECT`, or `CREATE MATERIALIZED VIEW … POPULATE` that must not die with a dropped connection.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112816",
          "createdAt": "2026-07-31T20:49:40Z",
          "updatedAt": "2026-08-13T17:37:14Z",
          "timestamp": "2026-08-13T17:37:14Z",
          "metrics": {
            "reactions": 4,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:bf3157f9e24ccd3efc8e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114531",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114531",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject a lossy codec on columns backing keys and indexes",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114406 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A lossy codec such as `SZ3` is now rejected at DDL time on any column that backs the sorting key, primary key, partition key, a secondary index or the unique key. Because such a codec does not return the value that was written, a merged part could be stored out of physical order and index analysis would skip rows matching the query. Column statistics are no longer built for, or used to prune parts by, a lossily compressed column, which returned too few rows on default settings. Existing tables stay loadable. ### Description `ORDER BY` makes a physical promise: rows are stored inside a part in sorting-key order and the primary index samples those stored values. A lossy codec breaks `read(write(v)) == v`, so the merge sorts pre-compression values while the part stores post-compression ones. `SZ3` is not monotonic, so the stored sequence is not sorted. The same applies to anything else computed from the pre-write block: skip-index granules, the unique-key index, partition values. The issue's reproducer gives `min(i - prev) = -0.2436889648437699` after `OPTIMIZE TABLE t FINAL` and 0 after one `INSERT`: the disorder appears only at the merge, where a debug build aborts in `CheckSortedTransform`. Lossiness comes from the existing `ICompressionCodec::isLossyCompression()`, not a codec-name list. `CompressionCodecMultiple` did not override it, so a stacked `CODEC(SZ3(...), LZ4)` reported itself lossless; it now ORs over its children. Classification is per serialized substream, mirroring `MergeTreeDataPartWriterWide::addStreams`, so `ORDER BY arr.size0` stays allowed while `arraySum(arr)` is not. The check sits at the two user-facing entry points: `registerStorageMergeTree` for CREATE and full-definition ATTACH, `checkAlterIsPossible` for ALTER on the initiating execution, following `4a29ef847411256`. A replica replaying a durable DDL entry is not re-checked, which would wedge its DDL worker. Column statistics, found during review, needed a different remedy: they are built pre-write too, but `auto_statistics_types` attaches `basic` to every numeric column, so a DDL rejection would ban the codec on any table that does not opt out. Instead none are built for such a column, and the pruner ignores any an earlier version wrote. With no setting changed, a predicate matching 2 rows returned 0. `system.columns` still lists them, as the two settings' descriptions now note. Seen in CI on one AST-fuzzer run, on #113575: [report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113575&sha=84df30db17165b65e3c78bc11513ed26534187d7&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%29). <details> <summary>Related cases deliberately left out of this PR</summary> - A **projection** over a lossily compressed column answers differently than the base part: `sum(i)` is `249957173.79` read from the base part and `249998750` via the projection. The mechanism is a separate post-write consumer, which writes the base part lossily and then computes the projection from the unchanged pre-compression block, so it follows in its own PR rather than being folded in here. - A legacy table can still gain an implicit minmax index over such a column through server configuration on the load path. Reachable, but a boundary sweep over 300 values produced no wrong result. - A part written by an earlier version keeps its statistics if the codec is then replaced by a lossless one without any merge or mutation rewriting that part. Deciding this needs per-part codec provenance, which parts do not record. `ALTER TABLE ... MATERIALIZE STATISTICS` or `OPTIMIZE TABLE ... FINAL` clears it. - `ORDER BY length(arr)` is now rejected although its value is exact. A key expression records which columns it needs and not which of their streams, so at this layer `length(arr)` and `arraySum(arr)` are indistinguishable, and `arraySum` is a genuine wrong-results carrier. The syntactic form `ORDER BY arr.size0` stays allowed because it names the stream. - Alias-dependent index rebuilds look incorrect independently of codecs: an alias-only `MODIFY COLUMN` requires no mutation while the index expression is rebuilt, so existing parts keep a same-named index built from the old expression. No claim is made about it here. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114531",
          "createdAt": "2026-08-12T18:16:56Z",
          "updatedAt": "2026-08-13T17:37:12Z",
          "timestamp": "2026-08-13T17:37:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5c946118a42a90d193f4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:94148",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:94148",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Introducing a new fuzzer check in the ClickHouse CI for my dear friend",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required) ---- New buzzhouse job <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Adds a new multi-node, docker-in-docker fuzzer job to CI and refactors shared fuzzer log analysis/parsing, which could affect how failures are detected and reported across fuzz/stress pipelines. Changes also touch integration test cluster process tracking (`exec_id` handling), so misclassification of failures or missed crashes is the main risk. > > **Overview** > **CI now runs a new fuzzer check, `La Casa Del Dolor`, in both `master` and `pull_request` workflows**, replacing the previous `BuzzHouse` job names/keys and wiring it into `finish_workflow` dependencies and job definitions. > > **Adds `ci/jobs/lacasadeldolor_job.py`** to execute the fuzzer via `tests/casa_del_dolor/dolor.py` in the integration-tests runner (docker-in-docker), collect multi-node logs/config artifacts, and reuse a new shared `analyze_job_logs()` routine. > > **Refactors failure detection/reporting**: `ast_fuzzer_job.py` extracts log analysis into `analyze_job_logs()` (including expanded accepted exit codes and improved OOM/sanitizer handling), and `FuzzerLogParser` is updated to search across multiple (including `.gz`) server/stderr logs and prefer the matched log when extracting stack traces/failed queries; `stress_job.py` and `clickhouse_proc.py` are updated to the new parser API. > > **Updates Casa del Dolor test harness and cluster helpers** to support the new CI mode: fixes imports under `tests.integration`, adjusts generator config/tempfile handling and exit-code validation, tightens disk/policy XML generation, and tracks ClickHouse container `exec_id` in `ClickHouseCluster/Instance` for more reliable shutdown/exit-code inspection. > > <sup>Written by [Cursor Bugbot](https://cursor.com/dashboard?tab=bugbot) for commit 82bd7ef20780c0d5b5bc648ace847418429573cb. This will update automatically on new commits. Configure [here](https://cursor.com/dashboard?tab=bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/94148",
          "createdAt": "2026-01-14T11:15:28Z",
          "updatedAt": "2026-08-13T17:37:11Z",
          "timestamp": "2026-08-13T17:37:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:09241fb681644aab1bfb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110029",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110029",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Read-through filesystem cache for the experimental ReaderExecutor",
          "text": "First of two PRs adding a read-through cache to the experimental `use_reader_executor` read path (default off), split out of the `ReaderExecutor` series (#103706), after decryption (#109702). This PR adds the cache-provider interface and the filesystem cache tier; the page cache and the richer, coordinated driver follow in a second PR. Adds `ICacheProvider` and the FileCache-backed `DiskCacheProvider`, consulted and populated per read window. The interface and provider are adopted from #103706's redesigned head: a whole-range `resolve(object, offset, range)` returns the window's residency as an ordered list of `Resolution`s in one cache transaction — each hit carrying a `CacheReader`, each populating miss carrying its open `CacheWriter` (a read-only/bypass tier returns writer-less misses). The executor drives it with the simplest possible per-window loop: `resolve` the window start, serve a cache hit straight from the tier's buffer (zero-copy), or on a miss claim the covering cell(s), fetch from the source, populate, and serve one block. The miss read goes through the long connection, so a cold sequential scan streams from one held connection; the window is returned as a `ChainedBuffers` (block-chunked) and decrypted per node on an encrypted disk. Concurrent readers of the same cold cell elect a single downloader via a claim taken before the fetch, so only one populates each cell; a cell another reader is already downloading is fetched through from the source (its populate lands zero bytes) rather than waited on. Coordinated waiting arrives with the page cache PR. Scope of this PR: - The filesystem cache serves known-size sources only — an FS cache is never attached to an unknown-size object, so the cache path is never entered for one. The executor's handling of unknown-size sources is otherwise unchanged from `master`. - No page cache, no cross-window plan, no prefetch, and none of #103706's plan machinery (`ResidencyIterator`, `CoverageMap`, `MemoryPressureMonitor`). - `ChainedBuffers` is functionally unchanged from `master` (only a couple of over-long comments trimmed). - Everything is gated behind `use_reader_executor`; the executor also falls back for the distributed cache and async prefetch, which it does not implement. New settings: `reader_executor_window_size` (serve window, 4 MiB) and `reader_executor_block_size` (buffer chunk, 1 MiB), each at least 4 KiB (rejected at settings load otherwise). Tests: `04511_reader_executor_disk_cache` (an `s3_cache` MergeTree) asserts the executor engages and consults the filesystem cache; `04604_reader_executor_min_size` asserts the sub-4-KiB window/block rejection; IO gtests cover the executor, the provider, and the offset map. Related: https://github.com/ClickHouse/ClickHouse/pull/103706 Related: https://github.com/ClickHouse/ClickHouse/pull/109702 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a read-through filesystem cache to the experimental `ReaderExecutor` read path (`use_reader_executor`, disabled by default). ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110029",
          "createdAt": "2026-07-10T16:50:28Z",
          "updatedAt": "2026-08-13T17:36:45Z",
          "timestamp": "2026-08-13T17:36:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f685bc865849e83ce8db",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114669",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114669",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: Clarify `PREWHERE` behavior with `JOIN`",
          "text": "This clarifies how `PREWHERE` behaves in queries with `JOIN`. It removes the misleading claim that `SELECT` clauses follow a single execution order, clearly states the single-table restriction, and adds a runnable minimal example with expected output. ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required for a documentation change.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114669",
          "createdAt": "2026-08-13T16:57:35Z",
          "updatedAt": "2026-08-13T17:36:29Z",
          "timestamp": "2026-08-13T17:36:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "dhtclk",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:0307cdb5cc5c19777958",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113833",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113833",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add JOIN observability columns to system.query_log",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111352 Related: https://github.com/ClickHouse/ClickHouse/issues/111748 ## Motivation `system.query_log` says almost nothing about what the JOINs in a query actually did. Answering questions like \"which queries ran a `CROSS` join we didn't expect\", \"which joins fell back to `grace_hash` and spilled to disk\", or \"which algorithm was really chosen when `join_algorithm = 'auto'`\" currently requires re-running `EXPLAIN PIPELINE` (impossible post-mortem) or scraping `ProfileEvents` query by query. ## Changes Four columns are added to `system.query_log`, with the matching fields in `QueryLogElement`: - `used_number_of_joins` (`UInt64`) — the number of physical joins executed by the query. It is collected from the query pipeline, so it reflects the joins that really ran after all optimizations, not the number of `JOIN` clauses in the query text. - `used_join_algorithms` (`Array(LowCardinality(String))`) — the algorithms that were actually used, e.g. `hash`, `parallel_hash`, `grace_hash`, `direct`, `full_sorting_merge`, `partial_merge`. The `join_algorithm` setting only lists the allowed algorithms; the choice among them happens at runtime, and an algorithm can even be replaced mid-execution (a `hash` join switching to `grace_hash` under memory pressure). - `used_join_kinds` (`Array(LowCardinality(String))`) — Kind of the joins present in the query. - `used_join_strictness` (`Array(LowCardinality(String))`) — Strictness of the joins present in the query. - `join_spilled_to_disk` (`UInt8`) — whether any of the joins spilled to disk. This PR currently adds the schema only. The fields are declared but nothing populates them yet, so the columns read as `0` and `[]`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added four columns to `system.query_log` describing the JOINs a query executed: `used_number_of_joins` (the number of physical joins in the executed pipeline), `used_join_algorithms` (the algorithms actually used at runtime, which can differ from the `join_algorithm` setting), `used_join_kinds` (`INNER`, `LEFT`, `CROSS`, `ASOF` and so on), and `join_spilled_to_disk` (whether any join wrote temporary data to disk). This makes it possible to find problematic JOIN patterns across a fleet without reproducing each query.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113833",
          "createdAt": "2026-08-07T14:11:36Z",
          "updatedAt": "2026-08-13T17:36:09Z",
          "timestamp": "2026-08-13T17:36:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "Manerone",
          "state": "open",
          "assignees": [
            "Fgrtue"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fba672da1a25603a5daf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86353",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cascades cost-based optimizer for distributed query plans",
          "text": "A Cascades-style cost-based optimizer that chooses distribution strategies for the multi-stage distributed query plans of #106020. It explores alternatives in a memo (a shared store of equivalent plan fragments) with top-down, goal-directed search and picks the cheapest plan satisfying the required distribution and sorting properties, inserting exchange operators (plan steps that move rows between nodes) as needed. Implemented: - **Join strategies**: shuffle hash join, broadcast hash join (with `ReplicatedRead` — every worker repeats the same read of a small table instead of a network broadcast, assuming shared storage where all workers see the same data), replicated join (a small deterministic join is recomputed identically on every node over such reads, so its result never crosses the network; nested joins compose; `ANY` joins are excluded because the kept row depends on the build order), local join. - **Aggregation strategies**: two-phase (partial + merge), shuffle by group keys, local; `distributed_aggregation_memory_efficient` and `distributed_plan_force_shuffle_aggregation` are honored. - **Top-N**: two-stage distributed top-N (per-node bounded sort, sorted-merge gather, coordinator limit); disabled under `exact_rows_before_limit`, which needs the full row count. - **Read strategies**: parallel N-way read, replicated read, local read. For `FINAL`, #108148 (already in master) taught the rule-based distributed plan to split a `FINAL` read into disjoint primary-key-range buckets where that is safe; Cascades now reuses that machinery, so `FINAL` no longer forces a serial read here either. The coordinator ships each bucket's marks in the `read_bucket` task parameters. - **`IN (subquery)`**: follows the `rewrite_in_to_join` setting like the rest of the planner (the forced join form is removed). In the default set form the set-building subqueries are planned separately and distributed like any other query. - **Properties and enforcers**: distribution (node count, replication, partitioning columns with equivalence classes and the types the keys are cast to before hashing) and sorting; when a plan alternative lacks a required property, an enforcer inserts the step that provides it (`Gather`/`Shuffle`/`Broadcast`/`ScatterExchange`, `Sort`). - **Transformations**: join commutativity (only for semantics-preserving joins: `INNER ALL`, `CROSS`, `SEMI`/`ANY`/`ANTI`; never `ASOF`, and never `ANY` under `join_any_take_last_row`), two-phase aggregation split, two-stage top-N split. - **Cost model**: `work`, `network`, and `sequential` components, each priced as wall-clock per node: a shuffle moves 1/N of the data per node, a broadcast payload is ingested once by every receiver in parallel, and a gather funnels every row through one endpoint, so its transfer stays undivided and pays a per-row cost. A hash-table build counts as parallel work (`parallel_hash` shards it across threads); a fixed per-exchange overhead keeps small inputs local. A table read is priced on its scan volume - the rows the primary key keeps - not on its output estimate, so a filter off the sorting key cannot make a replicated re-read look free. Standalone filters (e.g. `HAVING`) are estimated from column NDVs with join-key equivalence classes; join estimates are clamped to join kind and strictness semantics; exchange costs use per-row byte widths measured from the parts' column sizes (followed through renames, not derived from types). All weights and calibration constants are overridable at query time. What this improves over the rule-based distributed planner, on TPC-H plans. The rule-based planner broadcasts a small table when its read is below `distributed_plan_max_rows_to_broadcast`, but it often cannot size the result of a join, so a small join result (`nation x region`, 5 rows after the region filter) is scattered across nodes, joined there, and shuffled again (repartitioned across nodes) by the next join key. It also often inserts a shuffle at join and aggregation boundaries even when the rows are already divided by the right key. Cascades estimates sizes through joins, knows which partitioning already holds, compares broadcast against shuffle by cost for each join, and recomputes a small deterministic join on every node when that is cheaper than moving its result. On TPC-H SF100 over 8 nodes (same binary, same run window; times are server-side means of the hot runs) the join-heavy queries improve: | Query | Rule-based -> Cascades | What changed in the plan | |---|---|---| | Q21 | 7.66 s -> 5.34 s | The four `supplier x nation` joins compute their small results once and broadcast them, so `lineitem` is not shuffled to meet them. | | Q09 | 4.07 s -> 2.47 s | `part` and `nation` are read in full by every node, so `lineitem` and `supplier` are not shuffled to meet them. | | Q08 | 2.39 s -> 0.99 s | `nation x region` (5 rows) is recomputed by every node; `part` is read in full per node, so `lineitem` is not shuffled to join it. | | Q02 | 2.14 s -> 0.84 s | The `supplier x nation x region` chain is recomputed by every node, so only `partsupp` and `supplier` are shuffled. The top-100 sort becomes two-stage, sending at most 100 rows per node. | | Q05 | 2.05 s -> 1.12 s | The whole dimension side (`orders x customer x nation x region`) is recomputed by every node over full local reads; `lineitem` joins it in place with no shuffle at all. | | Q17 | 2.72 s -> 1.78 s | The small per-part average is broadcast to every node, so the outer 600M-row `lineitem` read is not shuffled. | | Q11 | 0.49 s -> 0.26 s | `supplier x nation` is computed once and broadcast; `partsupp` joins it in place. | | Q12 | 0.76 s -> 0.59 s | The filtered `lineitem` rows (~30K of 600M) are gathered once and broadcast; nothing is shuffled. | The remaining queries change less. Summed over all 22 queries, hot server time drops from 36.8 s to 28.8 s (about 22% lower). The remaining regressions are the `IN (subquery)` queries Q18 (2.43 s -> 2.85 s) and Q20 (2.24 s -> 2.99 s): with the forced join rewrite removed, both run their `IN`s as sets, and the set-form plan Cascades picks is slower than the rule-based one; a cost-informed choice of the `IN` form is a follow-up. `EXPLAIN pretty = 1, estimates = 1` shows the chosen plan with a row estimate and the accumulated cost for each step. Also in this PR, two improvements to the shared bucketed-read machinery (they benefit the rule-based path too): the `FINAL` layer split no longer depends on the coordinator's core count, and a many-partition `FINAL` split groups its layers into the target task count instead of falling back to a serial read. Design, a worked example on TPC-H data (a simplified 3-table query traced through the memo), and current limitations are documented in `src/Processors/QueryPlan/Optimizations/Cascades/ARCHITECTURE.md`. Plan-shape tests cover the actual TPC-H queries (`03836_tpch_join_order_plans`), and focused tests pin the cost-model contracts (e.g. `04869_cascades_read_cost_granule_volume`, `04838_cascades_filter_selectivity`). Disabled by default. Requires the analyzer; remote execution requires the stateless-worker configuration, while `distributed_plan_execute_locally = 1` runs the stages in-process without it: ```sql SET enable_cascades_optimizer = 1, make_distributed_plan = 1; ``` For tests, `param__internal_cascades_cluster_node_count` overrides the cluster size, `param__internal_cascades_cost_config` overrides the cost model configuration, `param__internal_join_table_stat_hints` injects table statistics, `param__internal_cascades_task_limit` lowers the task budget (it can never raise it). Related: https://github.com/ClickHouse/ClickHouse/pull/106020 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental Cascades cost-based optimizer for distributed query plans, enabled by `enable_cascades_optimizer = 1` together with `make_distributed_plan = 1`. It chooses between shuffle, broadcast, replicated, and local join strategies, two-phase, shuffle, and local aggregation, two-stage distributed top-N, and parallel and replicated reads by estimated cost, inserting exchange operators as needed. ### Documentation entry for user-facing changes - [ ] Documentation written in [/docs](https://github.com/ClickHouse/ClickHouse/tree/master/docs)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86353",
          "createdAt": "2025-08-28T11:27:09Z",
          "updatedAt": "2026-08-13T17:35:58Z",
          "timestamp": "2026-08-13T17:35:58Z",
          "metrics": {
            "reactions": 21,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3fb2def04f7cd3f30406",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113022",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113022",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a type-aware Bloom filter index for JSON",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113376 The original design used `JSONAllValues` as the input to a Bloom filter. `JSONAllValues` serializes each value as text. It does not preserve the runtime type. This loss of type information is important for `Dynamic` values. ClickHouse can compare JSON values with different runtime types. Some type pairs can match after conversion. Other type pairs can return an exception. A Bloom filter that stores only text cannot safely model these rules. It can skip a granule that contains a match. It can also hide an exception. `JSONAllValues` remains useful for text search, but it is not a safe base for this index. This PR replaces that design with `jsonbf_v1`. The new index creates tokens from these components: - JSON path - Container role - Runtime type - Binary value The container role separates scalar values, array elements, and map values. The index also processes nested JSON objects and named tuple fields. For `Dynamic` values, the index stores type-presence and complex-value presence tokens. Query analysis uses exact value tokens only when the comparison is safe. If analysis cannot prove safety, ClickHouse reads the granule. Container roles are preserved through nested tuples, and nested casts are handled conservatively. The index supports: - Equality and typed `IN` conditions - `has`, `hasAny`, and `hasAll` for arrays - Typed map values by key - Nested JSON objects and tuples The index does not optimize range conditions or whole-container equality. It also rejects unsafe comparison paths, such as Decimal-to-Float comparisons. An unsupported `Dynamic` runtime type disables skipping for its granule. In a one-million-row JSONBench test, the index reduced reads from 123 granules to 12–15 granules. Selected string equality queries were 1.6–2.3 times faster. Indexed inserts were approximately 2.9 times slower in the checked-in performance test. A local performance test produced these median results: | Operation | Without index | With `jsonbf_v1` | Difference | |---|---:|---:|---:| | Numeric equality, matching value | 13.05 ms | 10.99 ms | 1.19 times faster | | Numeric equality, missing value | 13.25 ms | 10.80 ms | 1.23 times faster | | String equality | 24.49 ms | 12.90 ms | 1.90 times faster | | Array `has` | 24.71 ms | 12.78 ms | 1.93 times faster | | Insert | 20.51 s | 40.87 s | 1.99 times slower | The small performance test shows limited benefit for numeric scalar equality and larger improvements for string equality and array membership. The larger JSONBench data set benefits more because the index skips more granules. Token generation and Bloom-filter construction increase insert time. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Adds the `jsonbf_v1` data-skipping index for type-aware equality and array membership on `JSON` values.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113022",
          "createdAt": "2026-08-02T18:31:45Z",
          "updatedAt": "2026-08-13T17:35:49Z",
          "timestamp": "2026-08-13T17:35:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "rorylshanks",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:59543768731ff370dc04",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114663",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114663",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113450 to 26.6: Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113450 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31717307521/job/94505163853)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114663",
          "createdAt": "2026-08-13T16:08:31Z",
          "updatedAt": "2026-08-13T17:32:57Z",
          "timestamp": "2026-08-13T17:32:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-1",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "JiaQiTang98",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3e2f8e639c29010a932f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112945",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112945",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use `pread` when `preadv2` with `RWF_NOWAIT` cannot be used, and recognize `EPERM` from it",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/104634 Related: https://github.com/ClickHouse/ClickHouse/issues/49149 Related: https://github.com/ClickHouse/ClickHouse/issues/39753 The `pread_threadpool` read method hands every read off to a thread pool, unless the data is already in the page cache, which it checks with the `preadv2` system call and the `RWF_NOWAIT` flag. Two things can go wrong with that check, and both were handled badly. **`EPERM` was not recognized.** It is what a `seccomp` profile of a container runtime answers for a system call that is not in its allow list. `ThreadPoolReader::submit` handed the read off to the thread pool for `ENOSYS` and `EOPNOTSUPP`, but let `EPERM` through to the throw, failing the query with `CANNOT_READ_FROM_FILE_DESCRIPTOR`. This is the signature reported in #49149. **The check was never verified in advance.** `hasBugInPreadV2` only compared the kernel version, and nothing else was checked, so on a system where the check cannot work `pread_threadpool` kept paying for a thread pool hand-off on every read - including the reads that only had to copy the data from the page cache. In #104634, on Amazon Linux 2 (kernel 5.10), the profile shows `ThreadPoolReaderPageCacheMiss` 1,971,514 out of `LocalThreadPoolJobs` 1,972,484 - every read went to the pool - while the device only moved ~22 GB of the 258 GB the file descriptor delivered, i.e. ~90% of the data was in the page cache and still paid for the hand-off. The system call is now probed once, before it is used, by `preadNoWaitUnavailableReason`. The probe passes an invalid file descriptor on purpose: `seccomp` filters and the system call table are consulted before the descriptor is looked up, so an available system call answers `EBADF` without reading anything, while a blocked one answers `EPERM` or `ENOSYS`. When it says the check cannot be used, `applySettingsQuirks` switches the default value of `local_filesystem_read_method` from `pread_threadpool` to `pread` at start time, and says why in the server log. Nothing downstream has to know: the reader, the userspace page cache eligibility in `DiskLocal::prepareRead` and the prefetched read pool all see a plain `pread` setting. As with the other settings quirks, a value set explicitly - in the configuration, or with `SET` at runtime - is left alone. Such a session keeps `pread_threadpool` and keeps paying for the hand-off, which is what happens today. The switch is a property of the host, so the setting is left marked as unchanged: only the changed settings are serialized into the query the initiator sends to the remote shards, and a host that cannot use the system call must not impose `pread` on the shards that can. An explicitly requested value stays changed and is still sent. The per-read `errno` handling in `ThreadPoolReader::submit` is kept, with `EPERM` added to it: the probe answers for the system call, but a particular filesystem can still reject the flag (`tmpfs` answers `EOPNOTSUPP`, for example), and such a read is handed off to the thread pool instead of failing the query. ### How it was tested `preadv2` was rejected the way a container runtime does it, with a `seccomp` filter installed by a small wrapper (`SECCOMP_RET_ERRNO`), and an old kernel was simulated with `setarch --uname-2.6`. Reading a 2 million row `MergeTree` table with the default `local_filesystem_read_method`, before (the released 26.7.1 binary) and after: | | before | after | |---|---|---| | no filter | `ThreadPoolReaderPageCacheHit` 39, `LocalThreadPoolJobs` 90 | `ThreadPoolReaderPageCacheHit` 34, `ThreadPoolReaderPageCacheMiss` 3, `LocalThreadPoolJobs` 93 - the read method stays `pread_threadpool` | | `preadv2` → `EPERM` | `Code: 74 ... errno: 1, Operation not permitted (CANNOT_READ_FROM_FILE_DESCRIPTOR)`, already while attaching the table | the read method is `pread`, the query succeeds, no `ThreadPoolReader` events | | `preadv2` → `ENOSYS` | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 - every read to the pool | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 80 | | kernel reported as older than 5.11 | `ThreadPoolReaderPageCacheMiss` 45, `LocalThreadPoolJobs` 135 | the read method is `pread`, no `ThreadPoolReader` events, `LocalThreadPoolJobs` 81 | In all three rejected cases the reason is in the log, for example: ``` <Warning> SettingsQuirks: The default value of local_filesystem_read_method has been switched from 'pread_threadpool' to 'pread' (you can explicitly set it back still), because the `preadv2` system call is not available (the probe with an invalid file descriptor answered errno: 1, strerror: Operation not permitted instead of `EBADF`); it is typically rejected by a `seccomp` profile of a container runtime, and can be allowed in the runtime configuration. ... ``` An explicitly requested `local_filesystem_read_method = 'pread_threadpool'` is kept, on every one of those systems, and this is where the per-read `EPERM` handling earns its place: under the `EPERM` filter the same query fails on the released binary and succeeds here, with `ThreadPoolReaderPageCacheMiss` 45 and `LocalThreadPoolJobs` 135 - every read handed off to the pool, which is the documented cost of asking for it there. Under the same filters, `system.settings` reports `local_filesystem_read_method = 'pread'` with `changed = 0`, so nothing is forwarded to the remote shards, while an explicitly requested `pread_threadpool` reports `changed = 1` and is still sent. Unit tests cover the `errno` classification, the probe's `EBADF` contract, and the quirk itself (the default is switched exactly when the probe says the system call is unusable, the switched value stays out of `Settings::changes()`, and an explicitly set value is never switched). An automated end-to-end test would need an instance with a restrictive `seccomp` profile, which the integration test framework cannot express today - it starts every instance with `seccomp:unconfined`. ### Documented behavior impact On a system where the page cache cannot be checked without waiting for the disk - Linux older than 5.11, a sandbox that rejects `preadv2` with an error code, and systems other than Linux, where `preadv2` does not exist and the method never checked the page cache in the first place - the default value of `local_filesystem_read_method` becomes `pread` instead of `pread_threadpool`, so local reads are performed in the calling thread. This includes the reads with `O_DIRECT`, which never look at the page cache and do not need the check; the read method is now resolved once, for the whole server, so they follow the same value. Nothing changes on a supported system, and nothing changes for an explicitly configured read method. The setting description in `src/Core/Settings.cpp`, from which the docs are generated, is updated in this PR. A `seccomp` profile that terminates the process instead of rejecting the system call with an error code cannot be detected: the startup probe is itself a `preadv2` call, so such a profile kills the server there. On `master` it kills it at the first read instead, under the same default read method. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): The `pread_threadpool` read method needs the `preadv2` system call with the `RWF_NOWAIT` flag to read the data that is already in the page cache without handing the read off to a thread pool. It is now checked at start time whether that system call can be used, and if it cannot - the Linux kernel is older than 5.11, or a `seccomp` profile of a container runtime rejects the system call - the default value of `local_filesystem_read_method` is switched to `pread`, and the reason is reported in the server log. Previously, every read paid for a thread pool hand-off on such systems, and a `seccomp` profile that answers `EPERM` made queries fail with `CANNOT_READ_FROM_FILE_DESCRIPTOR`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112945",
          "createdAt": "2026-08-01T20:09:06Z",
          "updatedAt": "2026-08-13T17:32:00Z",
          "timestamp": "2026-08-13T17:32:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 26
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2f60bddfa2d463abd876",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114247",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114247",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Clear plain LIMIT/OFFSET in the window-view backfill source query",
          "text": "Follow-up to https://github.com/ClickHouse/ClickHouse/pull/113759: the last AI review finding on that PR landed after the PR had been added to the merge queue, so the branch could no longer be updated. This addresses it. `StorageWindowView::getSourceTableSelectQuery` builds the raw-source backfill query for `CREATE WINDOW VIEW ... POPULATE`. The rows it produces are inserted into the window view, where `writeIntoWindowView` executes the mergeable view query over them, so the backfill query must deliver the raw source rows and leave all row transformations of the original query to the view query — otherwise the initialized state diverges from live behavior. The PR originally made the helper clear the leftovers of the original `SELECT` that violated this invariant one by one — `LIMIT`/`OFFSET` (and `WITH TIES`), plain `DISTINCT`, `ARRAY JOIN`, `WHERE`/`PREWHERE`, table-expression `SAMPLE`/`FINAL` — on top of the `JOIN`/`GROUP BY`/`ORDER BY`/`LIMIT BY`/`WINDOW`/`QUALIFY`/`INTERPOLATE` handling it already had. Review then found that this strip-the-leftovers approach misses wrapped sources: the same constructs inside a `FROM (SELECT ...)` subquery or a CTE definition survived the rewrite, and covering them would have required recursing the rewrite into every nested select. So the helper now builds the backfill query from scratch instead of stripping a clone of the view query. The contract makes this valid: `writeIntoWindowView` always receives raw source-table blocks (`getInputHeader` is the source table header no matter how the view query wraps or transforms the table — `PushingToWindowViewSink` is created with exactly that header), so the correct backfill query is exactly `SELECT <source columns> FROM <source table>`, plus `ORDER BY` on the timestamp column so the watermark is initialized from the earliest record. Wrapped sources, joins, and every row-shaping clause are covered by construction because the user query is no longer cloned at all. The helper shrinks by ~100 lines. Note: the divergence is currently unobservable because `CREATE WINDOW VIEW ... POPULATE` fails before writing any rows (https://github.com/ClickHouse/ClickHouse/issues/113493), so no regression test is possible yet; this keeps the rewrite invariant consistent for when `POPULATE` is fixed. Verified with `clickhouse-local` probes that `POPULATE` over wrapped-subquery/CTE/`JOIN`/`WHERE`/`FINAL` sources now analyzes the backfill query cleanly and proceeds to the pre-existing #113493 sink error. Related: https://github.com/ClickHouse/ClickHouse/pull/113759 Related: https://github.com/ClickHouse/ClickHouse/issues/113493 ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114247",
          "createdAt": "2026-08-10T23:40:02Z",
          "updatedAt": "2026-08-13T17:30:18Z",
          "timestamp": "2026-08-13T17:30:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d67eb717a8001bb66513",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114658",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114658",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix inconsistent AST formatting for a subquery argument of the view table function",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114004 (auto-closes the issue when this PR is merged into the default branch) --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114004 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed formatting of a subquery argument of the `view` and `viewIfPermitted` table functions when the enclosing query has a trailing `SETTINGS` clause. Such a query was formatted as `view((SELECT ...))`, which cannot be parsed back, so the query failed the internal format-parse-format check and raised `Inconsistent AST formatting`. ### Description `ASTQueryWithOutput::formatImpl` sets `parent_has_trailing_settings` so an inner `ASTSelectWithUnionQuery` parenthesizes its individual SELECTs, keeping the re-parser from consuming the trailing `SETTINGS` into the last SELECT. The flag is inherited down the format frame, so it also reached the argument of `view` / `viewIfPermitted`. There the parentheses are both redundant and rejected: the closing paren of the table function already terminates the select, and `ViewLayer::parse` bails out on a parenthesized lone select (`ExpressionListParsers.cpp:2981-2991`). The formatted text therefore did not parse back. The trailing `SETTINGS` clause is the trigger. Without it nothing sets the flag and the output is already correct. This clears the flag at the two places that cross the `view` argument boundary, which is what three other boundary owners already do: `ASTSubquery.cpp:98`, `ASTCreateQuery.cpp:1113` and `ASTAlterQuery.cpp:1032`. `ViewLayer` is the only producer of a bare-select function argument and serves exactly these two functions, so the two sites are the complete set. The issue frames parser-versus-formatter as the fork. This takes the formatter side, on the grounds that the parentheses carry no meaning in this position and that the same reset is the established pattern; the parser side would widen the accepted grammar to fix an output bug. The call is his to close. One output change beyond the broken shape: a multi-select argument such as `view(SELECT 1 UNION ALL SELECT 2)` also loses its branch parentheses. That form parsed before and parses now. Validation: `DESCRIBE TABLE view(SELECT 1) SETTINGS input_format_orc_use_fast_decoder = 0` aborts a debug server with `Logical error: 'Inconsistent AST formatting'` before the change and returns `1 UInt8` after it. The new test fails on the unpatched binary and passes 50/50 under randomized settings on the patched one. The `formatQuery` / round-trip / `1941` stateless family was run on both binaries and the failure sets are byte-identical.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114658",
          "createdAt": "2026-08-13T15:46:03Z",
          "updatedAt": "2026-08-13T17:30:00Z",
          "timestamp": "2026-08-13T17:30:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:6fce33c98954674a5590",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114659",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114659",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Record TopK-filtered granules in the query condition cache",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Granules fully filtered by the dynamic threshold of an `ORDER BY ... LIMIT n` (TopK) read are now recorded in the query condition cache, so repeat runs of such queries skip them at the mark-selection stage instead of re-reading and re-filtering most of the table. On the ClickBench Q24/Q26 shapes (`SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10`), a warm run drops from ~95M rows / ~12000 granules read to ~1.3M rows / ~160 granules, and from ~15–30 ms to ~8 ms on a 96-core machine. The `ORDER BY ... LIMIT n` (TopK) optimization pushes a dynamic `__topKFilter` threshold into the `MergeTree` read as a PREWHERE, dropping rows that cannot beat the running top-N. But the granules it emptied were never recorded in the query condition cache: the PREWHERE write path in `MergeTreeSelectProcessor::read` rejects non-deterministic conditions, and `__topKFilter` is one. Only the downstream WHERE `FilterTransform` wrote entries, at whole-chunk granularity, which learns almost nothing (a single surviving row voids the attribution of the whole chunk) — measured on hits, warm runs still selected 11766 of 12348 granules and re-read ~95M of 100M rows. The original out-of-tree build of this optimization (PR #81944, commit `2e2f594308f4` plus the follow-up making the TopN dynamic filters deterministic and reusable for the cache) recorded them and skipped ~98% of granules on warm runs — that is the 2–3x Q24/Q26 gap measured in issue #114639. Recording these granules is sound: for a fixed plan and data, the running threshold only tightens, so a granule none of whose rows survive the filter contains no row that could reach the final top-N, regardless of the threshold trajectory. The entries are keyed with the TopK plan salt (`TopKFilterInfo::condition_hash`: sort column, type, `LIMIT`, direction, `NULLS` direction, collation locale, number of sort columns, and the part-set snapshot), mirroring what the WHERE write path already does since #104478/#110507 — only the same TopK plan over the same part set ever reuses them. Changes: - `isDeterministicAllowingTopKFilter` (two identical static copies in `updateQueryConditionCache.cpp` and `ReadFromMergeTree.cpp`) moved into `VirtualColumnUtils`. - The TopK salt is plumbed to the reader via `MergeTreeReaderSettings::query_condition_cache_top_k_salt`. - The PREWHERE write path accepts a `__topKFilter`-bearing condition when the salt is present, folding the salt into the cache key. Any other non-deterministic condition is still never cached. - The PREWHERE consult path (`filterPartsByQueryConditionCache`) applies the salt exactly when the PREWHERE contains `__topKFilter`, so the keys match the write side; deterministic user PREWHERE conditions keep using plain, unsalted entries shared with non-TopK queries. - `04217_query_condition_cache_topk` and `04242_query_condition_cache_topk_collate` pin exact cache entry counts, which double (each TopK plan now writes a WHERE entry and a PREWHERE entry per part). - New test `04891_query_condition_cache_topk_prewhere_granules` isolates the PREWHERE write path with a WHERE-less TopK query (on current master such a query writes no cache entries at all), asserts that the ASC and DESC plans do not share granule decisions, and that a warm run selects fewer marks. - New performance test `topk_query_condition_cache.xml` with the ClickBench Q24/Q26 shapes over `hits_100m_single`. Measured on the 100M-row single-part ClickBench `hits` (96-core aarch64, hot, interleaved runs): warm Q24/Q26 at 0.007–0.009 s, on par with the original bench-opt build (0.009–0.011 s), against 0.015–0.03 s for current master; warm runs read 159 of 12348 granules against master's 11766. A first run with an empty cache shows no measurable overhead (0.016–0.019 s with the mechanism on and off). Closes: https://github.com/ClickHouse/ClickHouse/issues/114639 Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/110507",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114659",
          "createdAt": "2026-08-13T15:46:11Z",
          "updatedAt": "2026-08-13T17:29:49Z",
          "timestamp": "2026-08-13T17:29:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:92ea6900d782407fbc08",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:100371",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:100371",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `borrow_from_cache` object storage and `memory` metadata types",
          "text": "Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache using ephemeral `FileSegment`s. Each stored object is backed by a cache segment held alive via `FileSegmentsHolder`; when released, the cache reclaims the space. Add a new metadata type `memory` that keeps all file-to-blob mappings and directory structure entirely in memory with no persistence. On server restart all data is lost, which is acceptable for the intended use case of temporary tables. Configuration example: ``` disk(type=object_storage, object_storage_type='borrow_from_cache', metadata_type='memory', cache_name='some_cache') ``` ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A new object storage type and metadata type suitable for temporary tables. Add a new object storage type `borrow_from_cache` that allocates space in a named filesystem cache and holds it from eviction. Add a new metadata storage type `memory` that keeps the mapping in memory. For example, this could allow creating temporary tables with custom engines in the Cloud without using S3.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/100371",
          "createdAt": "2026-03-22T15:18:44Z",
          "updatedAt": "2026-08-13T17:29:44Z",
          "timestamp": "2026-08-13T17:29:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 32
          },
          "labels": [
            "pr-feature"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:59b1e031f376b8db3cc3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114401",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keeper: do not lose a session request when the Raft leader changes",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/78474 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a bug where a client connecting to ClickHouse Keeper during a Raft leader change could be held for the whole `session_timeout_ms` (30 seconds by default) before its connection was rejected, instead of being rejected as soon as the in-flight request was dropped. The connecting client can now reconnect to another replica sooner. ### Description A client's `Connect` makes Keeper submit an internal `SessionID` request. If the Raft append stream breaks while it is in flight, during a leader election say, that request is lost silently. **Root cause.** Such a request carries `session_id = -1` and no xid; its identity lives in `(server_id, internal_id)`. Two places used the wrong key: * `KeeperRequestDispatcher::onCommit` correlated a commit with its in-flight head by `(session_id, xid)`. Every `SessionID` request shares `(-1, 0)`, and `onCommit` runs on every node for every commit, so a `SessionID` committed for another server retired a still-uncommitted local request. * The error path queued the response for lookup by `session_id`, which `-1` has no callback for, so it was discarded and `getSessionID` timed out. The real waiter is a promise keyed by `internal_id`. **The change.** `onCommit` additionally requires `(server_id, internal_id)` to match for `OpNum::SessionID`. Error responses go through one `SessionID`-aware routing helper on `KeeperDispatcher`, wired into **both** dispatchers: `use_new_dispatcher` is a setting and the old one shares the defect. The request now fails at the in-flight drain bound rather than at `session_timeout_ms`. `KeeperTCPHandler` does not branch on the error code, so that earlier rejection is the whole user-visible gain; the accurate `ZCONNECTIONLOSS` only improves the server log. `test_keeper_force_recovery` also gets a retry around the connect after the election, since dropping in-flight appends is deliberate. The correlation fix is scoped to `OpNum::SessionID`. The garbage collectors also use negative session ids, but their `TryRemove` is idempotent and unwaited, so they are unaffected. **Validation.** New gtests cover both the routing decision and the production wiring behind it, each verified to go red under a mutation of the change it covers; the old dispatcher was exercised with `use_new_dispatcher = false`. An injected fault at the connect under test reddens the integration test without the retry and passes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114401",
          "createdAt": "2026-08-12T00:10:11Z",
          "updatedAt": "2026-08-13T17:29:07Z",
          "timestamp": "2026-08-13T17:29:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d6bca644a49e4a4a8220",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105710",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105710",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject divergent same-named parts in parallel replicas coordinator",
          "text": "The parallel replicas coordinator deduplicated parts purely by `MergeTreePartInfo` (name + version), so when two replicas announced a same-named part holding genuinely different data the second announcement was silently merged into the first. The coordinator then dispatched ranges from the first replica's snapshot to the second replica, whose local part was different, producing an exception `Trying to get non existing mark N, while size is M` inside `MergeTreeIndexGranularityConstant::getMarkRows` — or, when the layouts happened to line up, potentially incorrect results. This is reachable with `parallel_replicas_for_non_replicated_merge_tree = 1` against a cluster whose members each have independent local `MergeTree` data: block numbers of a non-replicated `MergeTree` come from a node-local `SimpleIncrement`, so each member's local first part is named `all_1_1_0` but can store different data. The fix makes the coordinator validate the identity of same-named parts across announcements: - `RangesInDataPartDescription` now carries the underlying part's total mark count, a content fingerprint (`checksums.getTotalChecksumUInt128`), and a `part_name_identity` tri-state saying whether a part name identifies the same content on every cluster member. The fields are gated on a new parallel replicas protocol version (`DBMS_PARALLEL_REPLICAS_PROTOCOL_VERSION` bumped to 10). - `part_name_identity` is `ClusterWide` when the engine coordinates block numbers through Keeper (`ReplicatedMergeTree` and descendants) *or* when all of the table's disks keep their metadata in storage shared by every cluster member (`MetadataStorageType::Plain`, `PlainRewritable`, `StaticWeb`, `WebIndex`, `Keeper`) — there every member enumerates literally the same parts. It is `NodeLocal` for a plain `MergeTree` on ordinary local or per-node remote disks. - `InOrderCoordinator` and `DefaultCoordinator` snapshot these fields from the first announcement of each part and compare later announcements against the snapshot. A fingerprint mismatch raises `BAD_ARGUMENTS` naming the diverging part and pointing the user at `ReplicatedMergeTree` or at disabling `parallel_replicas_for_non_replicated_merge_tree`. - When the fingerprint is unavailable on either side (a replica running an older server version, or a part whose checksums are not loaded) and either side reports `NodeLocal` part names, the coordinator fails closed with `BAD_ARGUMENTS` instead of weakening the identity check — such same-named parts cannot be verified, and merging them blindly could return incorrect results. For `ClusterWide` part names the part name implies identical content, so a mark-count fallback keeps mixed-version clusters working during rolling upgrades. Only analyzed-view fields (`rows`, `ranges`) are deliberately NOT compared: per-replica PK and skip-index analysis can legitimately select different mark subsets from the same underlying part, and `ranges` is consumed in place as the coordinator dispatches work. The protocol bump also required pinning one existing consumer of the serializer. The `read_bucket` task parameter of a distributed read plan (`make_distributed_plan`) embeds a `RangesInDataPartsDescription` blob but travels as an opaque query parameter with no handshake to negotiate a version on, so producer and consumer both used the build's own `DBMS_PARALLEL_REPLICAS_PROTOCOL_VERSION` and would disagree about the field layout across a rolling upgrade. The blob is now pinned to `DBMS_PARALLEL_REPLICAS_DISTRIBUTED_READ_BUCKET_VERSION = 8`, the layout in effect when it was introduced. Surfaced by the AST fuzzer (STID `4920-51f2`) on PR #105706. The 5-replica `parallel_replicas` cluster used by `tests/queries/0_stateless/02275_full_sort_join_long.sql.j2` exposes the divergence whenever local `t2` parts differ in size across replicas. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105706&sha=84790f83ec78aee26b08ab6e0bc6712f6e4f1745&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29 Related: https://github.com/ClickHouse/ClickHouse/pull/105706 Coverage is via `gtest_parallel_replicas_coordinator`, which exercises both coordination modes and both announcement orderings: rejection on divergent mark counts and divergent fingerprints, acceptance of identical announcements and of divergent analyzed views over the same part, the fail-closed path for `NodeLocal` part names without a fingerprint, the mark-count fallback for `ClusterWide` part names in mixed-version clusters, and a serializer round trip across protocol versions 8, 9 and 10 that requires the reader to consume the whole buffer. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed an exception `Trying to get non existing mark N, while size is M` and potential incorrect parallel-replica range assignment when `parallel_replicas_for_non_replicated_merge_tree = 1` is used on divergent local `MergeTree` data: the parallel replicas coordinator now validates same-named parts by a content fingerprint of the underlying part instead of merging announcements blindly, and fails closed when the fingerprint is unavailable and part names are not guaranteed to identify the same content on every cluster member. ### Documentation entry for user-facing changes - [x] Documentation is unchanged",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105710",
          "createdAt": "2026-05-23T23:10:40Z",
          "updatedAt": "2026-08-13T17:29:05Z",
          "timestamp": "2026-08-13T17:29:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 28
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:101e8f1a0a3923c060a4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114477",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114477",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow to create table without providing credentials if catalog specified",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow to create table without providing credentials if catalog specified. Earlier the query to create table looked like `create table catalog.table enigine = IcebergS3(url, key1, key2)`, but now we can create it via `create table catalog.table enigine = IcebergS3`",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114477",
          "createdAt": "2026-08-12T12:06:45Z",
          "updatedAt": "2026-08-13T17:28:10Z",
          "timestamp": "2026-08-13T17:28:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:d107e33b8fc30befbd9c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114522",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114522",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "QueryRunner follow-up: do not occupy threads eagerly",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): QueryRunner tables now start worker threads on demand and release them once idle, instead of occupying threads for the table's whole lifetime.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114522",
          "createdAt": "2026-08-12T16:42:37Z",
          "updatedAt": "2026-08-13T17:27:23Z",
          "timestamp": "2026-08-13T17:27:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:862b9bdae01c60ce5a9f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111457",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111457",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Compare read-in-order virtual row on its covered sort-key prefix",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/106740 Closes: https://github.com/ClickHouse/ClickHouse/issues/106630 Related: https://github.com/ClickHouse/ClickHouse/pull/110725 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a wrong result (mis-ordered merge) for read-in-order queries with a virtual row when `distinct-in-order` or `LIMIT BY` widens the read to a longer sort-key prefix than `ORDER BY` set it up for, and when a key column fixed by the filter is skipped by `ORDER BY` (e.g. `WHERE b = 1 ORDER BY a, c` on key `(a, b, c)`). The virtual row announced a wrong merge boundary: in release builds the merge could be silently mis-ordered, in debug builds the boundary assertion fired, and a `Nullable` key column after the skipped one threw the `Virtual row has different type` exception. The virtual row optimization now stays enabled in these cases. ### Description Alternative to https://github.com/ClickHouse/ClickHouse/pull/110725: instead of dropping the virtual row conversion when the in-order read prefix changes (and disabling it for skipped key columns), keep the optimization enabled and compare the virtual row only on the sort-key prefix it validly covers. **Root cause.** The read-in-order virtual row announced a wrong merge boundary in two ways: 1. *Widened prefix.* `optimizeReadInOrder` builds the virtual row conversion for the prefix `ORDER BY` needs (e.g. `CounterID`). A later optimization (`optimizeDistinctInOrder`, `optimizeLimitByInOrder`) re-requests the read with a longer prefix (`CounterID, EventDate`), but the `pk_block` width was derived from the conversion's input count, so the extra sort column was default-filled with `0` in `setVirtualRow`. In reverse order `0` understates the real values, so a real row exceeded the announced boundary: `Virtual row boundary violated in MergingSortedAlgorithm ... the virtual row announced UInt64_0 but the source then produced UInt64_1` in debug builds, a silently mis-ordered merge in release builds. 2. *Skipped fixed key.* For key `(a, b, c)` and `WHERE b = 1 ORDER BY a, c`, the fixed key `b` is skipped without an `ORDER BY` counterpart, but the conversion DAG indexed key columns densely, mapping `c` onto key column `b` (visible in `EXPLAIN actions=1`: input `b` aliased to `__table1.c`). The wrong value tripped the boundary check; a wrong type (`Nullable` key) threw the `Virtual row has different type` logical error even in release builds. Moreover, index values of the columns after the skipped key are semantically unusable: the index describes pre-filter data, so the entry `(5, 0, 9)` does not bound the filtered row `(5, 1, 3)` projected to `(a, c)`. **Fix.** - The virtual row conversion outputs only the sort-description prefix it can announce exactly: index values while the key prefix is contiguous, plus constants for fixed columns from `ORDER BY`. A skipped fixed key column ends the index-backed part: the index entry at a mark boundary may hold a filtered-out value for it, so its later components bound nothing in the filtered stream. A column fixed by the filter that stays in `ORDER BY` keeps disabling the virtual row, as before this fix. - The merge compares a virtual row only on the covered prefix and places it first on a covered-prefix tie (equivalent to treating the uncovered columns as minus infinity in the merge order, without materializing any values). The covered prefix is derived from the pk block column names in `MergingSortedAlgorithm` and carried per cursor in `SortCursorImpl::sort_prefix_limit`, honored by the generic `SortCursor::greaterAt`. A truncated virtual row can only occur with a multi-column sort description (its coverage is at least the first column), which always uses the generic cursor, so the single-column specialized queues are unaffected; the JIT comparator is bypassed when a truncated cursor participates. - `ReadFromMergeTree::readInOrder` reads index values for the whole used sorting-key prefix instead of only the conversion inputs. The in-order merges inside the read step sort by the full prefix, and index values are exact bounds for it even after filtering (a filter only removes rows), so these merges always see fully covered virtual rows. - `ReadFromMergeTree::requestReadingInOrder` drops the conversion only when a re-request makes it unsound: a prefix narrower than the one it was built for (the conversion could lose its inputs), or one not fully backed by the primary index. - `setVirtualRowConversions` builds the conversion with `project_inputs` so a raw index column cannot shadow a same-named conversion output when the merge looks sort columns up by name (matters with the old analyzer). This handles all key types uniformly — e.g. a descending `String` column after a skipped key keeps the optimization even though the type has no greatest value to pad with. Verified on release and debug builds (the boundary assertion is compiled only into debug builds): the previously aborting repros now return correct results, and `EXPLAIN` keeps `Virtual row conversions` for the widened and skipped-key reads. The test is based on the one from https://github.com/ClickHouse/ClickHouse/pull/110725, extended with checks that the optimization stays enabled, a descending skipped-key case, a `String` case in both directions, and a fixed key kept in `ORDER BY` (still disabled, as before).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111457",
          "createdAt": "2026-07-22T18:29:12Z",
          "updatedAt": "2026-08-13T17:27:14Z",
          "timestamp": "2026-08-13T17:27:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c8185479ff1c9bb423d2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114622",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114622",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not apply DROP fault injection to a refreshable view's cleanup DROP",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/114420 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a refreshable materialized view leaking its rotated-out target table when `ignore_drop_queries_probability` is enabled. The `DROP` a refresh issues to clean up the previous target is a step of the refresh, not a `DROP` the user asked for, so the fault injection no longer applies to it. ### Description Follow-up to #114420, requested by @ tiandiwonder in https://github.com/ClickHouse/ClickHouse/pull/114420#discussion_r3772879738. `ignore_drop_queries_probability` makes a `DROP TABLE` silently do nothing (or become a `TRUNCATE`) so the stress suite exercises \"the table you dropped is still there\". It must only affect `DROP`s the user issued. **Root cause.** After a refresh swaps in a new target, `StorageMaterializedView::dropTempTable` drops the rotated-out one via `InterpreterDropQuery(drop_query, refresh_context)`. `createRefreshContext` never marks that context internal, so the gate treated the cleanup `DROP` as a user `DROP`. Its existing refreshable-view exemption does not cover this: that test inspects the table *being dropped*, which here is the inner target, not the view. Both refresh exits are affected. When it fires, the old target survives as `.tmp.inner_id.<uuid>` still holding a full copy of the view's data, outside the view's metadata and surviving restart. **The change.** `InterpreterDropQuery` gains an `internal` member mirroring the one `InterpreterCreateQuery` already has, `dropTempTable` sets it, and the shared `refresh_context` is untouched. Marking that context instead would break the refresh outright in a `Replicated` database: the publishing `RENAME` runs on it, and `DatabaseReplicated` rejects a non-initial query unless the interpreter also passes `flags.internal`, which neither the Rename nor the Drop interpreter did. That is why #114420's one-liner was safe there and is not here. `QueryFlags{ .internal = internal }` at the replicated enqueue is behaviour-neutral for every other `DROP`, since all its members default to false. The interpreter's other policy branches (`ON CLUSTER` dispatch, access and dependency checks) deliberately keep applying. **Validation.** New test `04887_refreshable_mv_cleanup_drop_not_ignored`: three positive arms (success path, failure path, `Replicated` database) leak on master and are clean with the fix; a fourth asserts a genuine user `DROP` is still skipped. 50 randomized runs plus 50 with `ignore_drop_queries_probability=0.2` were green, as were `04247`, `04796`, `04218` and `04327`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114622",
          "createdAt": "2026-08-13T11:43:28Z",
          "updatedAt": "2026-08-13T17:26:44Z",
          "timestamp": "2026-08-13T17:26:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f3a994de6fc93fae5b42",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110183",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110183",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use-of-uninitialized-value in WITH FILL suffix over a merge",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107074 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-of-uninitialized-value in `ORDER BY ... WITH FILL ... STALENESS` when the fill suffix generates no rows (for example when the staleness window leaves nothing to fill) and the result is read through a merge of sorted streams (such as a two-shard `Distributed` table). ### Description Found by the AST fuzzer under MSan on top of #107074. `FillingTransform`, when all input chunks are processed, may run the suffix path. If the fill constraints are satisfied but no fill rows are produced (e.g. `WITH FILL ... STALENESS` with an exhausted staleness window), `generateSuffixIfNeeded` returns `true` while the result columns are freshly `cloneEmpty()`'d and carry no data. Previously the transform still emitted a 0-row chunk built from those empty columns. A downstream `MergingSortedTransform` (as used when reading from a two-shard `Distributed` table) then built a sort cursor over that empty chunk and compared row 0, reading past the end of the empty column. Fix: do not emit the suffix chunk when it has no rows. Minimal reproducer (needs a two-shard merge on the initiator): ```sql CREATE TABLE m (key Int) ENGINE = Memory; INSERT INTO m VALUES (100); CREATE TABLE d2 AS m ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), m); SELECT _shard_num FROM d2 ORDER BY _shard_num ASC WITH FILL TO 46 STALENESS 1; ``` CI finding: `AST fuzzer (amd_msan)`, report https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=107074&sha=3ec54c88d4eb535a5d644fc1ab91af31d717f9a5&name_0=PR&name_1=AST%20fuzzer%20%28amd_msan%29 MSan use-of-uninitialized-value in `ColumnVector<UInt32>::doCompareAt` (`MergingSortedAlgorithm::consume`), origin `FillingTransform::initColumns` (`cloneEmpty`) via the suffix path. Verified on a local `amd_msan` build: reproduces before the fix (server aborts), clean after; the new regression test returns `1\\n2`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110183",
          "createdAt": "2026-07-12T21:10:52Z",
          "updatedAt": "2026-08-13T17:26:43Z",
          "timestamp": "2026-08-13T17:26:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e8d63cfb42bfe3f9fa81",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96130",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96130",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Randomize tests with DETACH/ATTACH table before query execution",
          "text": "Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that enable randomly detaching and reattaching tables used in a query before its execution. This is a testing-only feature designed to find bugs related to table reattachment. Before executing a query, the system collects all tables referenced in the AST, and for each eligible table (stores data on disk, supports detaching, has no action locks or dependencies), it performs a `DETACH` followed by `ATTACH`. Changes: - Add `supportsDetachingTables` virtual method to `IDatabase` (overridden to `false` for engines that do not support non-permanent `DETACH TABLE`: `DatabaseDictionary`, `DatabaseReplicated`, `DatabaseSQLite`, `DatabaseBackup`, `DatabaseFilesystem`, `DatabaseHDFS`, `DatabaseS3`, `DatabaseURL`, `DatabaseRemote`, `DatabaseDataLake`, `DatabaseMaterializedPostgreSQL`) - Add `has`/`hasAny` methods to `ActionLocksManager` for checking existing locks (skipping expired `weak_ptr` entries) - Add table collection visitor and reattach logic in `executeQuery` (runs after AST validations, process list admission, and external tables initialization; skips `EXPLAIN`, transactions, internal/non-initial queries, and CTE name collisions) - Fix off-by-one in `MergeTreeDeduplicationLog::dropOutdatedLogs` (don't drop the active log) and add `sync` call in shutdown - Add `no-random-detach` tag to tests incompatible with this feature - Add `--no-random-detach` and `--reattach-tables-probability` options to `clickhouse-test` - Add `02461_reattach_tables` test Continuation of #55943. Continuation of #42336 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that randomly `DETACH` and `ATTACH` tables used in a query before its execution. This is a testing-only feature that helps find reattachment-related bugs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Introduces new pre-execution mutations (internal `DETACH`/`ATTACH`) in `executeQuery`, which can affect table availability and concurrency behavior if enabled; guarded by new experimental settings but touches core query execution paths. > > **Overview** > Adds experimental settings `reattach_tables_before_query_execution` and `..._probability` to optionally **DETACH and ATTACH back** eligible tables referenced by a query immediately before execution, including AST table discovery that accounts for CTE scoping, privilege checks, dependency/lock checks, and safety skips (e.g. `system`, non-disk storages, dynamic-structure columns, transactions, `EXPLAIN`, internal/non-initial queries). > > Extends `IDatabase` with `supportsDetachingTables()` and marks multiple database engines as not supporting non-permanent detach; adds `ActionLocksManager::has/hasAny` helpers to avoid detaching tables with active action locks. Updates the test runner and stress tooling to randomize this behavior (with `--no-random-detach` and probability control), adds a new `02461_reattach_tables` test, and tags many existing tests to opt out where DETACH/ATTACH would add flakiness/overhead. Also fixes `MergeTreeDeduplicationLog` cleanup to avoid dropping the active log and ensures writer `sync()` on shutdown. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit a369371ff61cb1934815ea1e8caf7debd83c98c4. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96130",
          "createdAt": "2026-02-05T23:27:47Z",
          "updatedAt": "2026-08-13T17:26:40Z",
          "timestamp": "2026-08-13T17:26:40Z",
          "metrics": {
            "reactions": 2,
            "comments": 91
          },
          "labels": [
            "pr-build"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a78a024e1941c2e0808a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114285",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114285",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
          "text": "Files written first turn the namespace into a plain directory, which a catalog sharing the storage view (SeaweedFS) rejects with HTTP 500; the swallowed error left an orphaned metadata file that broke retries. Ensure the namespace before the first write; propagate failures except 404-then-create and 409 (REST) / AlreadyExists (Glue). Example: https://pastila.nl/?01f91a25/620869a81af815ba6927860efbc72af4#NhE2Nbzh3AHhUHhA5Fdkzw==GCM ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114285",
          "createdAt": "2026-08-11T08:20:50Z",
          "updatedAt": "2026-08-13T17:26:05Z",
          "timestamp": "2026-08-13T17:26:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "azat",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ede453beb8d5fbf6c1a6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113983",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113983",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allowlist the expected FileLog bad-path reattach error in the upgrade check",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113781 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Upgrade check (amd_release)` intermittently fails its `Error message in clickhouse-server.log` sub-test on one benign line: ``` <Error> StorageFileLog (test_1.filelog_bad_path_attach): The absolute data path should be inside `user_files_path`(/var/lib/clickhouse/user_files/) ``` No product defect: the server starts, nothing crashes, no data is affected. Root cause. `04202_filelog_attach_path_outside_user_files` ATTACHes a FileLog table whose path is outside `user_files_path`. `ATTACH` is `LoadingStrictnessLevel::ATTACH` (2), which is `>= SECONDARY_CREATE` (1), so the constructor takes the relaxed branch at `src/Storages/FileLog/StorageFileLog.cpp:195-198`: it logs at `<Error>` and returns instead of throwing `BAD_ARGUMENTS`. That branch is deliberate and is what the test covers, since refusing to load at reattach time would break server startup. The table then outlives the test: stress threads run with a fixed `--database=test_N` (`ci/jobs/scripts/stress/stress.py`), and `clickhouse-test` skips its per-test teardown whenever `--database` is set (`need_cleanup = not args.database`), so that shared database is never dropped. The upgrade restart re-attaches the table, the relaxed branch fires again, and the line lands in the scanned log, where the post-restart scrub in `tests/docker_scripts/upgrade_runner.sh` had no entry for it. Hence the intermittency: `04202` must land on a fixed-database thread. Change. One `grep -av` entry in that scrub's existing secondary pipe, plus a short rationale comment next to the sibling entries. The pattern requires the fixture table name and the message together, and (bare parens are literals in BRE) the `StorageFileLog (db.table):` prefix shape. No source change, no test change. Validation. The scan pipeline, extracted verbatim from the runner, was run under GNU grep 3.11 against the failing run's own 19.9 MB `clickhouse-server.upgrade.log`. With the entry the artifact is empty; with it deleted the output is byte-identical to the 189-byte `upgrade_error_messages.txt` CI produced, so the sub-test flips `FAIL` to `OK`. Eight negative controls still surface, including a table whose name merely ends with the fixture name (`prod.other_filelog_bad_path_attach`), which the required `.` separator keeps visible.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113983",
          "createdAt": "2026-08-08T21:17:46Z",
          "updatedAt": "2026-08-13T17:25:44Z",
          "timestamp": "2026-08-13T17:25:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d5330b5f81b2e4c72de8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114668",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114668",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Stop Parquet background reads before releasing the format's read buffer",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114612 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed heap memory corruption when reading Parquet through an input format that owns its read buffer, for example a dictionary with `SOURCE(FILE(... format 'Parquet'))`. Background prefetch and decode tasks could still read and write through the buffer after the pipeline released it, which could abort the server. ### Description Closes #114612, reported by @ PedroTadim, who asked me to go ahead in https://github.com/ClickHouse/ClickHouse/issues/114612#issuecomment-5278451991. A `ReadBuffer` handed to an input format via `addBuffer()` lives in the format's `owned_buffers`. `ISource::work()` calls `onFinish()` on both the clean and the exception path, reaching `IInputFormat::resetReadBuffer()`, which clears `owned_buffers` and destroys the buffer. `ParquetV3BlockInputFormat` did not override that hook, and `Parquet::Prefetcher` holds a non-owning `SeekableReadBuffer *` to the same buffer, so background tasks kept using a destroyed object. It is not only a bad read: the report on the issue is a 1 MiB **write** into freed heap through `ReadBuffer::next()`, silent on a release build. `~Prefetcher()` does the right handshake, but only at format destruction, later than `onFinish()`; that window is the bug. The fix overrides `resetReadBuffer()` to drain background tasks before the base class releases the buffers, mirroring `ParallelParsingInputFormat::onFinish()`. Three details are forced by the surrounding code: the hook is `resetReadBuffer()`, not `onFinish()`, since it frees the buffers and has a second caller in `StreamingFormatExecutor`; the reader is drained but kept alive, because `getMatchedBuckets()` reads row group metadata after exhaustion; and `ReadManager` is drained before `Prefetcher`, because decode tasks re-enter `readSync` inline via `getRangeData()`. Sibling teardown paths: `resetParser()` already destroys the reader before delegating, `onCancel()` cancels it, and the reuse route through `StreamingFormatExecutor` calls `resetParser()` on every exit, so a drained reader is never reused. No fuzzer needed: `ReadBufferFromFile` leaves `use_pread` false, so a local file takes `SeekAndRead`, and any `CREATE DICTIONARY ... SOURCE(FILE(... format 'Parquet'))` whose load throws reaches it. Validated on ASAN: the added test fails 7/7 unfixed, passes 11/11 fixed; a standalone reproducer 31/33 unfixed, 0/35 fixed; a negative control with only the override reverted reddens at the unfixed rate. Parquet and polygon-dictionary suites gave an identical failing set on both binaries (187 each); 50 randomized-settings repeats passed clean.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114668",
          "createdAt": "2026-08-13T16:50:49Z",
          "updatedAt": "2026-08-13T17:24:33Z",
          "timestamp": "2026-08-13T17:24:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:3bd2d1887601282e732e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114533",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114533",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong results for non-boolean conditions taken out of `and`",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112236 `and` implicitly converts its arguments to booleans, so any non-zero value is true. When a query plan optimization takes a part of a conjunction away and a single conjunct is left as the new predicate, that conjunct was converted with a cast to the type of the original predicate. A cast is not a boolean conversion: it maps values like 256 or 0.1 to 0, so rows whose condition value is a non-zero multiple of 256 silently disappeared. ```sql CREATE TABLE t (id UInt32) ENGINE = MergeTree ORDER BY id; INSERT INTO t SELECT number FROM numbers(600); SELECT count() FROM t AS l LEFT JOIN t AS r ON l.id = r.id WHERE r.id AND l.id = r.id; ``` returned 597 instead of 599, the rows with `id = 256` and `id = 512` were dropped. `mergeFilterIntoJoinCondition` moves `l.id = r.id` into the JOIN and leaves `CAST(r.id, 'UInt8')` as the filter: ``` Filter column: CAST(id AS UInt8) ``` The same happens in `ActionsDAG::removeUnusedConjunctions` when a conjunct is pushed down and the filter column is still needed in the result. That one is reachable without a JOIN, and the value of the condition was wrong there as well (the raw value instead of a boolean): ```sql SELECT count(), sum(f) FROM ( SELECT id, (id != 1000 AND s) AS f FROM (SELECT id, sum(id) AS s FROM t GROUP BY id) WHERE id != 1000 AND s ); ``` returned `597 597` instead of `599 599`. It only converted floating point types, now every non-boolean type is converted. Both places now wrap the remaining conjunct into `and(x, true)`, the same way `toBoolIfNeeded` does it in `JoinStepLogical.cpp`. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong results when a non-boolean condition, such as a bare integer column in `WHERE`, is left alone after the other conditions are merged into the JOIN condition or pushed down. Rows whose condition value was a non-zero multiple of 256 were skipped. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114533",
          "createdAt": "2026-08-12T18:40:57Z",
          "updatedAt": "2026-08-13T17:24:11Z",
          "timestamp": "2026-08-13T17:24:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:424d4b3fd606ae8b50bd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107637",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107637",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Merge filters into join during join reordering",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Added a new setting `query_plan_merge_filters_into_join` that allows merging `Filter` steps into the `JOIN` step during join reordering, so `WHERE` predicates participate in join order optimization",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107637",
          "createdAt": "2026-06-16T15:25:08Z",
          "updatedAt": "2026-08-13T17:24:01Z",
          "timestamp": "2026-08-13T17:24:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:5195855fd20c73775eb2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:71028",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:71028",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Randomize parallel_replicas_min_number_of_rows_per_replica",
          "text": "<!--- Disable AI PR formatting assistant: true --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Details The setting enables code execution, which can trigger hidden bugs, in particular in GLOBAL JOINs with parallel replicas. Discovered one while doing https://github.com/ClickHouse/ClickHouse/pull/70658 within `02967_parallel_replicas_joins_and_analyzer` test",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/71028",
          "createdAt": "2024-10-24T14:24:19Z",
          "updatedAt": "2026-08-13T17:23:21Z",
          "timestamp": "2026-08-13T17:23:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 24
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "devcrafter",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:ec1d6790ec301d2d9ad7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107865",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107865",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add per-user filesystem cache disk usage metrics",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/105020 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added opt-in Prometheus gauges for current filesystem cache usage (`filesystem_cache_size_bytes` and `filesystem_cache_elements`), labeled by `cache_name` and `user_id`. The gauges are exposed through `system.dimensional_metrics` and the Prometheus endpoint and controlled by the `expose_prometheus_cache_usage_metrics_per_user` cache setting, which is disabled by default. --- ### Description Adds two gauges exposed through `system.dimensional_metrics` and the Prometheus endpoint: | Name | Labels | |---|---| | `filesystem_cache_size_bytes` | `cache_name`, `user_id` | | `filesystem_cache_elements` | `cache_name`, `user_id` | **Motivation.** Operators can identify which users currently occupy filesystem cache space, measured both in bytes and in file segments, for each cache. **Mechanism.** Each enabled cache owns a `FileCacheUsageTracker` containing shared per-user atomic counters. Main cache-priority entries retain the corresponding counters and update them together with the existing cache size and element accounting. `ServerAsynchronousMetrics` periodically obtains a per-user snapshot through `getUsageStatPerClient` and updates the dimensional gauges. This keeps `DimensionalMetrics` updates out of cache mutation paths. Composite LRU, SLRU, and split-cache priorities share the same tracker, including during SLRU queue transitions. Counters use shared ownership so inactive users can be reclaimed safely. When a user no longer has cache entries and its counters are zero, the next snapshot removes it from the tracker. Stale dimensional metric label combinations are also removed. **Sampling.** The metrics reflect the most recent asynchronous metrics update. **Cardinality.** In-process cardinality is bounded by users currently retained by each cache, plus labels awaiting the next asynchronous metrics update. Stale `(cache_name, user_id)` label combinations are removed. The feature remains disabled by default because enabling per-user metrics can still create significant Prometheus time-series cardinality. ### Documentation entry: - [x] Documentation is updated.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107865",
          "createdAt": "2026-06-18T14:22:58Z",
          "updatedAt": "2026-08-13T17:22:35Z",
          "timestamp": "2026-08-13T17:22:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "sacheendra",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:d1d1661bb76a4bdc5110",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107305",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107305",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make ALTER MODIFY COLUMN on named Tuple metadata-only when adding subfields",
          "text": "`ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` that only adds new subfields is now metadata-only (no mutation). Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1` (default `false`). Subfield additions through `Array`/`Map`/nested `Tuple` wrappers are also handled. `Nullable(Tuple(...))` is blocked (null map incompatibility). Removing/renaming subfields or changing types still triggers a mutation. Key/index/projection guards reject the metadata-only path when `primary.idx` or skip-index bytes would become invalid (whole tuple in key, or subcolumn whose type changes). Subcolumn references with unchanged types (e.g. `ORDER BY t.a` when only `t.c` is added) are allowed. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` is now metadata-only when only adding subfields, matching the speed of top-level `ADD COLUMN`. Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1`. ### Documentation entry for user-facing changes - [x] Documentation is not required (behavioral improvement; semantics unchanged)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107305",
          "createdAt": "2026-06-12T07:42:58Z",
          "updatedAt": "2026-08-13T17:22:17Z",
          "timestamp": "2026-08-13T17:22:17Z",
          "metrics": {
            "reactions": 1,
            "comments": 8
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ff5b5f3af5fae23b4170",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114328",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114328",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enhance MetadataStorageFromMemory",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Pure refactoring change, doesn't affect any working part of code.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114328",
          "createdAt": "2026-08-11T13:44:39Z",
          "updatedAt": "2026-08-13T17:21:58Z",
          "timestamp": "2026-08-13T17:21:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-object-storage-disks"
          ],
          "author": "alesapin",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5fa33ff9f40141682a5d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112717",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "title",
          "text",
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112717",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Rebuild skip indices, projections, TTL and MATERIALIZED columns left stale by ALTER MODIFY / UPDATE / CLEAR COLUMN",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/112302 Related: https://github.com/ClickHouse/ClickHouse/pull/85985 --> Related: https://github.com/ClickHouse/ClickHouse/pull/112302 Related: https://github.com/ClickHouse/ClickHouse/pull/85985 A mutation decides what to rebuild by comparing the mutated column against the columns that an index / projection / TTL expression requires. When the expression reads a **subcolumn** (`t.a`, `json.a`), those required columns carry the subcolumn name while the mutation tracks the whole stored column (`t`), so the comparison never matched. On **wide parts** the dependent files (`skp_idx_*`, projection parts, TTL info) were then hardlinked into the mutated part, leaving them describing data that no longer exists. Compact parts are unaffected, because a mutation rewrites the whole part. Every such site now resolves an expression's required columns to their top-level columns through one shared helper, `getRequiredColumnsWithSubcolumnsReplaced`, which rewrites subcolumns to `getSubcolumn` — the same rewrite used for `MATERIALIZED` expressions in #85985 — and re-analyzes the expression. The rewrite is applied where the value is recomputed as well, so it is derived from the *updated* parent column instead of a subcolumn pre-extracted from the source part. What is fixed: - **Skip index on a subcolumn** — rebuilt on `ALTER MODIFY COLUMN`, `ALTER UPDATE`, `ALTER CLEAR COLUMN` and patch updates of the parent. The recompute in the partial-rewrite path (`MutateSomePartColumns`) now extracts the subcolumn from the updated parent instead of failing with `NOT_FOUND_COLUMN_IN_BLOCK` or reading stale data. - **Projection** — rebuilt when the altered column feeds its sort key (a subcolumn sort key was missed) or its `WHERE` clause. Both decide what the projection stores: the sort key is persisted as raw bytes in the projection's own primary index, and the `WHERE` fixes the stored rows and aggregate states. A column used only in the projection's `SELECT` needs no rebuild, because that payload is converted on read. - **`alter_column_secondary_index_mode`** — the `throw`/`compatibility` guard in `checkAlterIsPossible` now also fires for an explicit index defined on a subcolumn of the altered column, instead of accepting the ALTER and then failing (or silently dropping the index) during the mutation. - **TTL on a subcolumn** — a `TTL t.a` is recalculated when the parent column is mutated, so rows and columns whose TTL moved into the past actually expire. `SHOW CREATE TABLE` still shows the original expression. - **`MATERIALIZED` column** — recalculated when `ALTER MODIFY COLUMN` changes the type of a column it reads (whole column or subcolumn), including chains of `MATERIALIZED` columns, and everything depending on them is rebuilt too. Previously only `ALTER UPDATE` recalculated them, so a value-changing conversion left them holding values computed from data the mutation had just rewritten. A `MATERIALIZED` column that is part of a key cannot be recalculated in an existing part — its sort order and partition id are fixed when the part is written, which is also why `MATERIALIZE COLUMN` refuses key columns — so such an ALTER is now rejected unless the conversion preserves values. Reproductions, all silent wrong results before this PR: ```sql -- skip index on a subcolumn CREATE TABLE t (id UInt32, t Tuple(a Int32, b String), INDEX idx t.a TYPE minmax GRANULARITY 1) ENGINE = MergeTree ORDER BY id SETTINGS index_granularity = 4, min_bytes_for_wide_part = 0, min_rows_for_wide_part = 0; INSERT INTO t SELECT number, (number, 'x') FROM numbers(16); ALTER TABLE t MODIFY COLUMN t Tuple(a Float32, b String); SELECT count() FROM t WHERE t.a >= 1; -- 0 before, 15 now SELECT count() FROM t WHERE t.a >= 1 SETTINGS use_skip_indexes=0; -- 15, ground truth -- projection filtered on a column whose type changes CREATE TABLE t2 (id UInt64, x Int64, PROJECTION p (SELECT sum(id) WHERE x < 0)) ENGINE = MergeTree ORDER BY id SETTINGS min_bytes_for_wide_part = 0; INSERT INTO t2 SELECT number, toInt64(3000000000) + number FROM numbers(100); ALTER TABLE t2 MODIFY COLUMN x Int32; -- every value wraps negative, so WHERE x < 0 now matches all rows SELECT sum(id) FROM t2 WHERE x < 0; -- 0 before, 4950 now -- MATERIALIZED column computed from a column whose type changes CREATE TABLE t3 (x Int64, m Int64 MATERIALIZED x) ENGINE = MergeTree ORDER BY tuple() SETTINGS min_bytes_for_wide_part = 0; INSERT INTO t3 VALUES (5000000000); ALTER TABLE t3 MODIFY COLUMN x Int32; SELECT x, m FROM t3; -- 705032704, 5000000000 before; 705032704, 705032704 now ``` Tests: `04617_stale_subcolumn_skip_index_after_mutation` and `04840_recalculate_materialized_column_on_source_type_change`. Known gaps left for follow-up pull requests: a subcolumn in a projection `WHERE` is accepted at `CREATE` but fails at `INSERT` (`Not found column t.a in block`), and the aggregation assignments of a GROUP BY TTL are not covered by the rewrite. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix wrong query results caused by a mutation leaving data derived from the mutated column stale on wide parts. A skip index or a TTL expression defined on a subcolumn (for example a `Tuple`, `Nested`, `Map` or `JSON` element) is now rebuilt or recalculated when `ALTER MODIFY COLUMN`, `ALTER UPDATE` or `ALTER CLEAR COLUMN` changes the parent column; a projection is rebuilt when the altered column feeds its sort key or its `WHERE` clause; and a `MATERIALIZED` column is recalculated when `ALTER MODIFY COLUMN` changes the type of a column it is computed from. An `ALTER MODIFY COLUMN` that would change the values of a `MATERIALIZED` column used in the sorting or partition key is now rejected, because such a column cannot be recalculated in an existing part.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112717",
          "createdAt": "2026-07-31T09:40:47Z",
          "updatedAt": "2026-08-13T17:20:11Z",
          "timestamp": "2026-08-13T17:20:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "Avogar",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5ecea989f0abd7f83f9c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110464",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110464",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `aiRedact` function for PII detection and redaction",
          "text": "`aiRedact` detects and redacts PII in text using an LLM provider. For example: SELECT aiRedact('Contact John Doe at john@doe.org', ['name', 'email']) -- 'Contact [REDACTED] at [REDACTED]' Pass an empty `categories` array to fall back to a default set of common categories (name, email, phone number, address, credit card, IP address). The redaction token defaults to `[REDACTED]` and can be changed with `replacement`. `aiRedact` is best-effort: detection and redaction are performed by an LLM, so the output is not reliable and may still contain PII depending on the model, prompt, and input. It must not be treated as a sufficient anonymization mechanism on its own, always review the output before relying on it. On error it behaves like the other AI functions: it throws by default, or returns an empty string when `ai_function_throw_on_error = 0`. Closes: https://github.com/ClickHouse/ClickHouse/issues/110362 Changelog category (leave one): - New Feature Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): New function `aiRedact` that detects and redacts personally identifiable information (PII) in text using an LLM provider. Specify the categories to redact (e.g. `['email', 'name']`) or pass an empty array to use a default set of PII categories. Matched values are replaced with a token (`[REDACTED]` by default). Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110464",
          "createdAt": "2026-07-14T23:07:06Z",
          "updatedAt": "2026-08-13T17:17:41Z",
          "timestamp": "2026-08-13T17:17:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested"
          ],
          "author": "davidmenggx",
          "state": "open",
          "assignees": [
            "george-larionov"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:0c90eef152a86b93d424",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114650",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114650",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the projection's intermediates out of the WITH FILL header",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/114404 Caused by: https://github.com/ClickHouse/ClickHouse/pull/107700 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `NUMBER_OF_COLUMNS_DOESNT_MATCH` for a distributed query combining an `ALIAS` column whose body is an expression with `ORDER BY ... WITH FILL ... INTERPOLATE`. ### Description `analyzeSort` passed every column *available* to the `Before INTERPOLATE` step through as an output. A step's available columns are all the nodes of the previous step's `ActionsDAG` — the intermediates of a computed expression included, since a later step is allowed to *reference* any of them. Turning \"referenceable\" into \"must be in the stream\" pinned those intermediates into the header of the `Filling` step. That made the header depend on something that should not matter. Two query trees that differ only by whether an `ALIAS` column was inlined into its defining expression produced different headers, because `v * 2` contributes its `2` and the un-inlined `a_v` does not: ``` un-inlined inlined __table1.k __table1.k __table1.a_v multiply(__table1.v, 2_UInt8) __table1.v __table1.v 2_UInt8 <- extra materialize(__table1.k) materialize(__table1.k) a_v a_v ``` A distributed query is planned from the un-inlined tree on the initiator and executed from the inlined one on the shard, so the two headers meet. `buildShardCollapseFanOut` only handles a *smaller* shard header, and the positional `makeConvertingActions` then throws `NUMBER_OF_COLUMNS_DOESNT_MATCH`. The fix holds back the projection's own intermediates and passes everything else through as before, including the sort columns materialized just above, which is what `Filling` fills by. They all remain *inputs* either way, so the step still asks the previous one to produce them. One exception has to be carved out: a column that an `INTERPOLATE` expression names. Those expressions become actions later, in `Planner`, against a dag built from this step's *header*, so whatever they reference must survive as a column even when the query does not select it — `INTERPOLATE (inter AS inter2 + inter)` where `inter2` is not in the `SELECT` list. Those names are collected by visiting the expressions into a throwaway dag. ### Scope Everything here is inside `if (query_node.hasInterpolate())`, so only queries with `INTERPOLATE` change. The shape was already broken over a `Distributed` table before #107700, since that path has always inlined `ALIAS` columns; #107700 extended the inlining to the parallel-replicas paths and so exposed it there too. Both are fixed. It also stops the reconciliation failure from masking a query error: `INTERPOLATE (k AS k)` on an `ORDER BY` column reports `INVALID_WITH_FILL_EXPRESSION` again instead of `NUMBER_OF_COLUMNS_DOESNT_MATCH`. ### Validation Built and run against a three-replica localhost cluster and a two-shard `Distributed` table, compared with the CI binaries of the commit before #107700 (`75b17ad`) and of its merge (`dd01d270e`): | | before #107700 | after #107700 | this | |---|---|---|---| | repro, parallel replicas shipping a plan | pass | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | repro, parallel replicas shipping SQL | pass | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | repro over a 2-shard `Distributed` table | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | pass | | `INTERPOLATE (k AS k)` under parallel replicas | `INVALID_WITH_FILL_EXPRESSION` | `NUMBER_OF_COLUMNS_DOESNT_MATCH` | `INVALID_WITH_FILL_EXPRESSION` | The new test `04891_with_fill_interpolate_alias_column_header` reports four exceptions on the commit before #107700 and seven on the merge commit, and passes here. Every self-contained stateless test mentioning `WITH FILL` or `INTERPOLATE`, plus the `ALIAS`-shipping tests from #107700, passes: 94 of 94. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-dd669bf9-3d72-4334-961f-0871feaf9f98?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-dd669bf9-3d72-4334-961f-0871feaf9f98&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114650",
          "createdAt": "2026-08-13T14:39:41Z",
          "updatedAt": "2026-08-13T17:17:23Z",
          "timestamp": "2026-08-13T17:17:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:08ca09886d8dd3571e88",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114615",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114615",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "diff-review skill second edition",
          "text": "Reworks the `diff-review` skill in `.claude/skills/diff-review` — the local in-browser diff review that Claude Code serves before a commit or a PR. Heads-up: I tailored this to how *I* like to review, so please treat it as a suggestion rather than a standard. Check the branch out, run it on one of your own diffs, and keep it only if you find it helpful. What changed: - **Multi-round review.** Sending a round no longer ends the review: the server stays up, the agent works on the comments it was handed while you keep reading, and each comment turns green in the page as it gets addressed. - **Comments are durable state.** They are written to the `--out` file as you type, so they survive a server restart or a page reload, and after the agent's fixes move the code they are relocated by their anchor line instead of pointing at the wrong place. Resolutions, replies and dismissals are part of that state. - **Explicit review range.** `--staged`, `--head <sha>` and `--committed` alongside `--base`, so a review of recorded work never picks up local edits; the header always states which range is on screen. - **Navigation.** Directory tree with per-file status, `+a −d` counts and open-comment badges, a path filter, an all-files page, whole-file mode with expandable folded regions, a Comments pane listing open and already-addressed comments, draggable sidebar and split, keyboard shortcuts with a `?` overlay, and light / dark / system themes. - **Tests.** `ui.html` is split into ES modules under `ui/`, covered by `ui_test.mjs`, and `persist_test.mjs` exercises the whole persistence round-trip against real servers on a throwaway repository. - Vendored `@pierre/diffs` bumped from 1.2.12 to 1.3.5. Local agent tooling only — nothing in the server, the build or the tests. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114615",
          "createdAt": "2026-08-13T09:54:07Z",
          "updatedAt": "2026-08-13T17:17:21Z",
          "timestamp": "2026-08-13T17:17:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4166fac3bb5a4bef31d0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114219",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114219",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.5: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114219",
          "createdAt": "2026-08-10T20:06:41Z",
          "updatedAt": "2026-08-13T17:15:21Z",
          "timestamp": "2026-08-13T17:15:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3058ac6d51b296576406",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113512",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113512",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "[WIP] Seal-gated reading: gate the probe side of a hash JOIN on the runtime filter and prune read ranges by it",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> <details> <summary>Claude</summary> ## Motivation JOIN runtime filters (`enable_join_runtime_filters`) filter probe-side rows only after they were read: the row-level `__applyFilter` conjunct discards non-matching rows, and the read-time index analysis from #109085 (`enable_join_runtime_filters_index_analysis`) skips granules inside already-created read tasks. In both cases the probe side starts reading concurrently with the build side, before the filter exists, so the early tasks are read unpruned. This PR adds a stronger, structural variant for the case when the probe-side join key is a prefix of the table's primary key: the probe side does not read anything until the build side completes, and the completed filter then prunes whole mark ranges by the primary key *before* read tasks are created. ## Approach The gating is expressed as an edge of the query pipeline, not as a waiting state: - `FillingRightJoinSideTransform` gets an optional \"seal\" output port. Exactly one of the concurrent filling transforms — the one that completes the build — emits a single seal chunk carrying the completed runtime filter in its chunk info; the rest just finish. - On the probe side, the sources of a gated `ReadFromMergeTree` are replaced by `SealGatedReadTransform`s: source-like processors whose only input is the seal. Until the seal arrives, the executor has nothing to schedule below them, so the read pools never cut a task. On the seal, the filter is handed to a `RuntimeFilterReadRangesRefiner` (installed on the read pool), which turns it into a primary key `KeyCondition` — an exact `IN`-set, or the `[min, max]` envelope when the exact set overflowed into a bloom filter — and drops non-matching mark ranges at task-cut time, reusing the refiner contract of the MergeTree read pools. - The plan-level pass `markSealGatedReading` finds hash joins whose pushed-down `__applyFilter` conjunct references a probe-side primary key column, and marks the join step and the reading step. `JoinStep::updatePipeline` then wires the seal to the pending seal inputs collected by the `Pipe`. - Everything fails open: a gated read whose seal cannot be wired (the build side of a join, `YShaped`/by-shards pipelines, cancellation) is fed from a `NullSource` and reads ungated with row-level filtering only. FINAL, parallel replicas, and joins sharded by primary key ranges are not gated. Both the default multi-threaded pool and the in-order reading paths are gated (including `max_threads = 1` and reads in primary key order). On a gated read, the read-time index analysis by the same runtime filter is skipped as redundant — the refiner is already granule-exact through the primary key; filters of other joins are kept. Enabled by the experimental setting `enable_join_seal_gated_reading` (default off) on top of `enable_join_runtime_filters`. This is also groundwork for epoch-based (punctuated) collocated joins, where the same reader will consume a stream of per-epoch seals. ## Results On a 90M-row probe table (`ORDER BY k`, warm cache; `tests/performance/join_seal_gated_reading.xml`, CI perf host numbers): - 100 sparse build keys (exact `IN`-set path): 172 ms → 12 ms - 1M build keys in a narrow band (bloom overflow, `[min, max]` envelope path): 645 ms → 54 ms Compared with the read-time granule pruning by the same runtime filter (`enable_join_runtime_filters_index_analysis` + `use_skip_indexes_on_data_read`), the single-query latency on local storage is on par (the probe side cannot run far ahead of the build even ungated: the join does not consume it until the hash table is ready, so port backpressure stalls it after about a chunk per stream). The structural difference of gating is that no read or prefetch is issued for pruned ranges at all, and that it is the seam for the per-epoch seals of collocated joins. The stateless test `04653_join_seal_gated_reading` asserts result equality with ungated execution, the pipeline structure, `ReadPoolRangeRefinerDroppedMarks`/`read_rows` contrast, the fail-open shapes, and the suppression of the redundant read-time analysis. </details> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental setting `enable_join_seal_gated_reading` (default off). When enabled together with `enable_join_runtime_filters` and the probe-side join key is a prefix of the table's primary key, the probe side of a hash JOIN starts reading only after the build side completes, and the completed runtime filter prunes whole mark ranges by the primary key before read tasks are created, instead of only filtering rows after they were read.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113512",
          "createdAt": "2026-08-05T15:13:35Z",
          "updatedAt": "2026-08-13T17:15:02Z",
          "timestamp": "2026-08-13T17:15:02Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-improvement",
            "pr-performance"
          ],
          "author": "KochetovNicolai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:17a41158345be33eb5cd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108090",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108090",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support DEFAULT expressions inside Tuple data types",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/issues/2797 `DEFAULT` expressions are now supported for named elements of `Tuple` data types, e.g.: ```sql CREATE TABLE t (id UInt8, c Tuple(a UInt8, s String DEFAULT 'Hello')) ENGINE = MergeTree ORDER BY id; ``` Such defaults exist only at the syntax level: they are pulled up to the column level (for `CREATE TABLE` as well as `ALTER TABLE ... ADD COLUMN` / `MODIFY COLUMN`), so the stored schema is normalized without any `DEFAULT` inside the data type. The column above is stored as type `Tuple(a UInt8, s String)` with a column-level `DEFAULT tuple(defaultValueOfTypeName('UInt8'), 'Hello')`. A default expression may reference other columns, but not other elements of the same tuple/nested; a reference that collides with an element name is rejected as ambiguous. Building an actual data type while a default is set throws. `DEFAULT` inside `Nested` (which is `Array(Tuple(...))`) or `Array` is not supported yet and is rejected with a clear `NOT_IMPLEMENTED` error, because a scalar element default cannot be represented as a static array column default. Note this is why issue #2797 (which asks specifically for `DEFAULT` on `Nested` elements) is linked as `Related` rather than `Closes`. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support `DEFAULT` expressions inside `Tuple` data types, e.g. `Tuple(a UInt8, s String DEFAULT 'Hello')`. The default is normalized away from the type and pulled up to the column level. Works for `CREATE TABLE` and `ALTER TABLE ... ADD`/`MODIFY COLUMN`. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108090",
          "createdAt": "2026-06-22T01:17:05Z",
          "updatedAt": "2026-08-13T17:13:35Z",
          "timestamp": "2026-08-13T17:13:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:41bf7f950bff24cd63e5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96487",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96487",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix LazilyReadFromMergeTree optimization with ALIAS columns (#96452)",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry: Fix missing lazy read and top-K read optimizations (skip-index top-K and dynamic top-K filtering) when selecting `ALIAS` columns with `WHERE ... ORDER BY ... LIMIT` on MergeTree tables. Closes: https://github.com/ClickHouse/ClickHouse/issues/96452 ### Documentation entry for user-facing changes N/A",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96487",
          "createdAt": "2026-02-09T21:21:08Z",
          "updatedAt": "2026-08-13T17:13:17Z",
          "timestamp": "2026-08-13T17:13:17Z",
          "metrics": {
            "reactions": 1,
            "comments": 31
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "jayvenn21",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d0e25c4b02363ee5762c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110892",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110892",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Explain analyze join stats",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/110668 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add information of internal state of joins to `EXPLAIN ANALYZE`",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110892",
          "createdAt": "2026-07-17T16:06:34Z",
          "updatedAt": "2026-08-13T17:13:11Z",
          "timestamp": "2026-08-13T17:13:11Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "Fgrtue",
          "state": "open",
          "assignees": [
            "vdimir"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:26a368e70926f7229090",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114660",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114660",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Stop SeaweedFS deleting live object folders in stateless CI",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/113828 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Since 2026-08-12, `s3 storage` stateless jobs intermittently fail an ordinary INSERT with `Code: 499 ... Immediately after upload: Object ... suddenly disappeared (S3_ERROR)`. The object really was destroyed: SeaweedFS's asynchronous empty-folder cleaner counts a folder's entries and then deletes the folder without excluding a PUT arriving in between, so a write landing in that window is lost. From a failing master run's own artifacts: ``` 16:18:17.675651 EmptyFolderCleaner: deleting empty folder /buckets/test/test/jys 16:18:17.676290 PUT of the object SUCCEEDS, 282 bytes 16:18:17.678805 verifying HEAD -> 404 16:18:17.681060 WriteBufferFromS3: Nothing to abort (so ClickHouse did not remove it) ``` This started with #113828, which replaced MinIO with SeaweedFS. ClickHouse keys objects as `<prefix>/<3 chars>/<random>`, so the bucket holds thousands of shallow folders that empty and refill continuously, and that churn arms the race: in one job 34,819 of 45,578 cleaner log lines were deletions, peaking at 3,238 per minute. CIDB has 0 hits in the preceding 180 days, then hits on 2026-08-12 only. I set the filer's empty-folder cleanup delay past any job's lifetime, so no folder becomes eligible for deletion while the suite runs. The cleaner keeps running and keeps queueing; only its eligibility window moves, and the folders persist in an instance destroyed at job end. No `src/` change: `s3_check_objects_after_upload` caught a genuinely destroyed write. Retrying or relaxing it would have hidden a real lost object. Validated against the shipped script, arms differing only by these lines: without the change the cleaner deletes 13 folders and 0 of 12 survive; with it, 0 deletions and 12 of 12 survive while the queue still holds its 12 items past 3m03s, where the baseline drained at 2m03s. The race is upstream's, which already carves `.uploads` out of this same cleaner for the identical reason (`empty_folder_cleaner.go:247`); this only keeps CI out of its way.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114660",
          "createdAt": "2026-08-13T15:52:45Z",
          "updatedAt": "2026-08-13T17:13:04Z",
          "timestamp": "2026-08-13T17:13:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:93b79aa9650faa403d00",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:93114",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:93114",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "WIP: Projection Index Text",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): TODO ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/93114",
          "createdAt": "2025-12-29T00:00:29Z",
          "updatedAt": "2026-08-13T17:12:47Z",
          "timestamp": "2026-08-13T17:12:47Z",
          "metrics": {
            "reactions": 7,
            "comments": 25
          },
          "labels": [
            "pr-improvement",
            "submodule changed"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:135ba05a2f2ce30f09a1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114472",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114472",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keep the patch release version bump increasing across recoveries",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113834 Related: https://github.com/ClickHouse/ClickHouse/pull/113528 --> Scheduled patch releases stopped incrementing the patch number — successive releases on a branch reused the same `vX.Y.P.*` line (e.g. `26.6.2.81`, `26.6.2.158`, `26.6.2.160`), because the post-release version bump was lost whenever a release was interrupted after the tag push, and every later recovery skipped it too. Prepare now classifies a run from the ref and the branch-tip version file into two flags: `is_recovery` (this run re-publishes an existing release rather than creating one) and `is_late_recovery` (the branch has already advanced to a newer release). The deferred bump is gated on `not is_late_recovery`, so a normal run and a current-release recovery complete the interrupted bump, while a superseded recovery never rewrites the version backwards. `stage_bump` refuses to write a version older than the branch tip, and asserts `is_recovery` on an empty bump — a fresh release must advance the version, a recovery may legitimately find it already done. Split out of #113834 — this is the version-bump half with a single push. Handling a non-fast-forward push (a backport moving the branch mid-release) is left to #113834. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md):",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114472",
          "createdAt": "2026-08-12T12:04:19Z",
          "updatedAt": "2026-08-13T17:12:19Z",
          "timestamp": "2026-08-13T17:12:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "leshikus",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:21dea433430556497e5a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113059",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113059",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Collect SQL stacktraces on the hung-check and server-died abort paths",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/42701 Related: https://github.com/ClickHouse/ClickHouse/pull/112265 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ### Description Related: #42701, #112265 On an ASan build, a stateless run aborting on the hung check records no stack of the hung server: ``` Hung check failed: server is not responding Cannot collect C stacktraces under ASan: debugger attach is disabled. ``` Two things combine. `print_c_stacktraces` declines to attach lldb on ASan builds, because ptrace disables LeakSanitizer; that refusal is correct and stays. And `print_sql_stacktraces`, needing no debugger, was unreachable: its pre-check `check_server_liveness` probes **HTTP** (`http_port`, default 8123), while it collects over native **TCP** (`args.client --port=`, default 9000). Different listeners, different ports, so this signature (alive, not answering HTTP, TCP still serving) failed the gate. The hung-check abort site did not call it at all. This drops the mismatched pre-check and lets the collector be its own liveness test: it is already bounded (`timeout=30`) and reports failure as one trimmed line, so a dead socket costs at most 30 s and cannot re-emit the `Code: 210` tracebacks that motivated the pre-check. The dump is added to the three abort sites that had only the C path: hung check, server died, and the startup check. The stateless job keeps attaching the dump to its result and additionally clears any left by a previous job in the same workspace, so an aborted run cannot upload a stale dump as its own. #114143 has since added the same attachment upstream; this replaces it with the equivalent helper rather than attaching twice. It fixes no hang and does not restore C++ stacks on ASan. A server alive but not answering HTTP now yields the full `system.stack_trace` view with per-thread `query_id`, identifying the wedged query; one dead on both transports records \"tried, got nothing\" instead of silence. Validated against a live server: with HTTP dead and TCP live, master skips and writes nothing, while this branch writes a `sql_stacktraces.log` carrying `thread_name` and `query_id`. Green runs are unaffected. [Prompting report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=42701&sha=1ce152efafc5b00bf31eb7a0a64757fecbbc6e4a&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20distributed%20plan%2C%20parallel%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113059",
          "createdAt": "2026-08-03T06:03:46Z",
          "updatedAt": "2026-08-13T17:12:07Z",
          "timestamp": "2026-08-13T17:12:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:32231b879571743cf6f2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114414",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114414",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add SHARED REGEXP path placement policy to JSON",
          "text": "Adds `SHARED REGEXP 'pattern'` to `JSON` type declarations so matching root-relative flattened paths are always stored in shared data instead of competing for dedicated dynamic-path subcolumns. This is useful for high-cardinality path families whose promotion would displace more useful paths. Matching is partial by default, like `SKIP REGEXP`. Set the persisted `JSON` type parameter `shared_regexp_use_partial_match=0` to require full-string matching. Typed paths take precedence over `SHARED REGEXP`; `SKIP` and `SKIP REGEXP` continue to discard matching data. Rules are evaluated against the complete path from the original `JSON` root, including derived sub-objects. The rules are compiled once into an immutable `RE2` matcher, with bounded rule count and pattern sizes. A single rule uses direct `RE2` matching and multiple rules use `RE2::Set`. `JSON` columns without rules keep a null matcher, their existing binary type encoding, and their existing row-placement path. Metadata snapshots are copied lazily only when retained placement provenance actually changes a type. Placement provenance is stored in part column metadata and retained by default through horizontal and vertical merges, wide and compact mutations, column renames, lightweight-update patch materialization, and `Array`/`Nullable`/`Tuple`/`Map` wrappers. Removing a rule therefore does not unexpectedly promote already-shared paths during a later rewrite. Set the `MergeTree` table setting `allow_json_shared_data_paths_repromotion=1` to opt into reconsidering those paths. Policy-only metadata changes inside `Variant` are documented as unsupported. The `Native` binary type encoding uses `JSON` encoding version 1 only when a policy is present; ordinary `JSON` types remain byte-identical version 0. The canonical `JSON` documentation and `Native` format specification are updated accordingly. Testing includes 18 focused unit tests and 9 stateless scenarios covering syntax, partial/full matching, root-relative sub-objects, all supported wrappers, flattened `Native` and `RowBinary` paths, horizontal/vertical merges, compact/wide mutations, renames, statistics, and both lightweight-patch application paths. An 88.6 MB real Elasticsearch `indices_stats` document produced 143,555 flattened paths; all 143,290 paths targeted by `^indices[.]` remained in shared data and none became dynamic. A one-run end-to-end smoke comparison was 1.66 s without the policy and 1.69 s with it, with a 0.055% max-RSS difference; these small differences are noise-level, not a statistical benchmark. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added `SHARED REGEXP` rules to the `JSON` data type to keep matching paths in shared data, with configurable full-string matching and explicit control over re-promoting paths after rules are removed.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114414",
          "createdAt": "2026-08-12T02:35:41Z",
          "updatedAt": "2026-08-13T17:11:36Z",
          "timestamp": "2026-08-13T17:11:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "valerypetrov",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:52a09f8105aaedb56f5c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:89945",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:89945",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add a Chinese tokenizer (jieba) for the tokens function and text indexes",
          "text": "<!--- Disable AI PR formatting assistant: false --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a `chinese` tokenizer for the `tokens` function and `MergeTree` text indexes. It segments Chinese text into words using a dictionary and a Hidden Markov Model (the algorithm follows [jieba](https://github.com/fxsjy/jieba)), with `coarse_grained` (default) and `fine_grained` granularities. Continues #80174. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) ### Implementation notes - The tokenizer is a from-scratch C++ reimplementation following the fxsjy/jieba algorithm (dictionary-based maximum-probability segmentation with an HMM fallback). The embedded dictionary and HMM model are generated from a pinned [cppjieba](https://github.com/yanyiwu/cppjieba) commit (MIT) and verified by SHA-256; regenerator scripts are included. - The engine lives in a standalone repository, https://github.com/ClickHouse/jieba_cpp, consumed here as the `contrib/jieba_cpp` submodule with a `contrib/jieba_cpp-cmake` wrapper that builds it against ClickHouse's in-tree abseil/darts-clone/zstd. Built by default (`ENABLE_CHINESE_TOKENIZER`); the dictionary ships in little- and big-endian variants and is `#embed`-ed per host byte order, so it works on big-endian targets too. - `hasToken` now bypasses the text index for any non-`splitByNonAlpha` tokenizer (use `hasAnyTokens` / `hasAllTokens` for tokenizer-aware matching). This also fixes a pre-existing case where `hasToken` over `ngrams`/`array`/`splitByString`/`asciiCJK` indexes could return wrong rows.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/89945",
          "createdAt": "2025-11-12T16:27:31Z",
          "updatedAt": "2026-08-13T17:11:31Z",
          "timestamp": "2026-08-13T17:11:31Z",
          "metrics": {
            "reactions": 3,
            "comments": 36
          },
          "labels": [
            "pr-feature",
            "submodule changed",
            "can be tested",
            "hold"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "Ergus"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:87e1049c537cb8e867c1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110477",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110477",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "EXPLAIN SYNTAX: return the pretty-printed query as a single multi-line record",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/80410 Related: https://github.com/ClickHouse/ClickHouse/pull/107925 --> Closes: #80410 Related: #107925 (closed, superseded by this PR) ### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `EXPLAIN SYNTAX` now returns the pretty-printed (multi-line) query as a single result record instead of one record per line. The `oneline` option (`EXPLAIN SYNTAX oneline = 1`) still collapses the output to a single physical line. Queries that consumed the previous per-line output should treat the result as one row whose value contains embedded newlines. ### Description `EXPLAIN SYNTAX <query>` formatted the reformatted query and split it into one result row per physical line. This PR emits the whole formatted, copy-pasteable query as a single record with newlines preserved (issue #80410). Only the `AnalyzedSyntax` code path is changed; `PLAN`, `PIPELINE`, `AST` and the `oneline` option are unchanged. The previous attempt (#107925) was closed because it flipped the `oneline` default to `1`, collapsing the query to a single physical line rather than keeping the multi-line pretty form. This PR keeps `oneline = false` by default and only changes the result shape from N rows to one multi-line record. Reference files for tests consuming EXPLAIN SYNTAX output (including `.oldanalyzer.reference` variants) were regenerated. A regression test (`04545_explain_syntax_single_record`) asserts the single-record multi-line default and the `oneline = 1` collapse.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110477",
          "createdAt": "2026-07-15T00:34:56Z",
          "updatedAt": "2026-08-13T17:11:09Z",
          "timestamp": "2026-08-13T17:11:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a61d2eec1cefb69c035a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114184",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114184",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Implementing generic block nested loop join",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Implemented generic block nested loop join",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114184",
          "createdAt": "2026-08-10T16:04:43Z",
          "updatedAt": "2026-08-13T17:10:41Z",
          "timestamp": "2026-08-13T17:10:41Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:aa1c0e614052f32190b4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113208",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113208",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix array membership over an erased element holding one concrete type",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112953 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `has`, `indexOf`, `indexOfAssumeSorted`, `countEqual`, `mapContainsKey` and `mapContainsValue` missing rows that `=` matches when the array element type is type-erasing (`Dynamic` or `Variant`, including nested in `Tuple`): `has([1::UInt64]::Array(Dynamic), 1::UInt8)` returned 0 while `1::UInt64::Dynamic = 1::UInt8` is 1. The fix applies to a row whose own elements resolve to one concrete alternative and `equals` over that peeled pair succeeds; other rows keep the previous behaviour. ### Description Related: https://github.com/ClickHouse/ClickHouse/pull/112953 At default settings these functions miss rows that scalar `=` matches, and `countEqual` breaks the `arrayCount(elem -> elem = x, arr)` equivalence its docs state. **Root cause.** `FunctionArrayIndex` tests equality with `IColumn::compareAt(...) == 0`, but for `ColumnDynamic`/`ColumnVariant` `compareAt` compares the variant type name (or discriminator) before the value: a sort order, not equality. Equal values under different variants never compare equal. **The change.** One dispatcher at both array entry points covers all six `FunctionArrayIndex` instantiations as constant, materialized and `LowCardinality`. For a type-erasing common type it peels an operand resolving to one concrete alternative and compares with the registered `equals`, folding each wrapper level's nullness at its own level so an outer NULL stays distinguishable from a nested one. `Tuple` recursion stops at `Array`/`Map`, also `compareAt`-based. **Scope, deliberately narrow, decided per ROW.** A block shares one flattened element column, so the decision is taken per row group: a row's answer follows from its own elements and the needle, not from which rows share its block. It declines, before any behaviour change, for a row mixing several concrete types, shared-variant rows, container alternatives, `LowCardinality` elements, and any pair `equals` rejects (the condition is `equals` succeeding, not the types, since comparability is partly value-dependent). Declined rows keep master's answer bit-for-bit. A `NULL` needle against a materialized erased array holding a NULL now matches: `has([NULL], NULL)` -> 1. No setting is added. **Validation.** New parallel-safe test `04706`: every cell asserted against an oracle in the same row, controls pinning each declined shape, and a group asserting three block partitions agree. A/B against pristine master over the 229 stateless tests reaching these functions: identical failure sets. 50/50: 100 OK. <details> <summary>Validation detail</summary> Every measurement pairs `SELECT lower(buildId())` against `readelf -n` on the serving binary. `#112953` touches this file but neither entry point. **Left for separate PRs**, each measured unchanged on both arms in a run where `has` itself moved 0 -> 1 on the same fixture: `hasAny`/`hasAll` (separate GatherUtils implementation); `arrayCompact`; container equality itself (`[1::UInt64::Dynamic] = [1::UInt8::Dynamic]` is 0 while the `Tuple` twin is 1, which is why this fix stops there); and the heterogeneous-row miss. **Mutations**, each rebuilt with the Build ID confirmed to move, the whole test re-run, then restored with it confirmed to return. Each reddens only what is named: | mutation | reddens | |---|---| | remove the new call | 15 lines; controls green | | guard on `Dynamic` only | the 2 `Variant` cells | | drop the NULL-vs-NULL match | the NULL cells | | bypass the dispatcher on the `Map` path | direct `has(map, key)` | | non-recursive, then one-level, `Tuple` recursion | the 3 `Tuple` cells; then `Tuple(Tuple(Dynamic))` | | decline constant arrays | the erased constant cell | | accept a row of several concrete types | the heterogeneous controls: the decline protects them | | never unwrap `Nullable` | the `Nullable(Tuple(Dynamic))` cells | | remove the `Map` cardinality normalisation | the 2 `Map` NULL-needle cells | | ignore a column's own nullness in the fold | the `LowCardinality`-needle cell | | decide per block, not per row group | the block-partition cells | **Performance** (debug, `max_threads=1`, median of 5): a 1000-element constant `Array(Dynamic)` over 2000 rows is unchanged at 0.023 s; a materialized 200k x 50 one goes 0.287 s -> **0.135 s**. **The updated reference** (`04338`, 2 lines) now enforces its own comment, \"`has(m, k)` must agree with `has(mapKeys(m), k)` for every row\": `has(map(NULL::Dynamic, 1), CAST(NULL, 'LowCardinality(Nullable(String))'))` was 0, the array path 1. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113208",
          "createdAt": "2026-08-04T00:05:50Z",
          "updatedAt": "2026-08-13T17:10:35Z",
          "timestamp": "2026-08-13T17:10:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4dcffa7447f5003bdff2",
        "signalId": "github:ClickHouse/ClickHouse:issue:114579",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114579",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Lazy FINAL with `optimize_aggregation_in_order` merges one group at a time (25000 `mergeBlocks` calls and 50000 log lines for a 35000-row table)",
          "text": "### Describe the situation When `query_plan_optimize_lazy_final` and `optimize_aggregation_in_order` are both on, the aggregation that the lazy `FINAL` replacement builds is merged **one group at a time**: `Aggregator::mergeBlocks` is called once per distinct key, and each call writes two log lines (a `Trace` \"Merging partially aggregated blocks\" and a `Debug` \"Merged partially aggregated blocks for bucket #-1\"). On a 35000-row table with 25000 distinct keys that is 25000 merges and 50000 log lines for a single query. A plain in-order `GROUP BY` of the same size does 7 merges, so this is specific to the lazy `FINAL` shape rather than to aggregation in order in general. The cost is dominated by the logging, so it scales with how verbose the server is configured to be. On a debug/sanitizer build with the log level CI uses, the query below took **525 s**, of which `LoggerElapsedNanoseconds` attributes **285 s** to the logger; the same query with `optimize_aggregation_in_order = 0` took ~3 s. This is not a correctness problem - the results match - and it is not new; it reproduces on builds well before the report below. ### How to reproduce Any recent `master`. With `clickhouse-local`: ```sql CREATE TABLE lf (k UInt64, version UInt64, is_deleted UInt8, v UInt64) ENGINE = ReplacingMergeTree(version, is_deleted) ORDER BY k; INSERT INTO lf SELECT number, 1, 0, number FROM numbers(20000); INSERT INTO lf SELECT number, 2, if(number % 10 = 0, 1, 0), number * 2 FROM numbers(10000, 15000); SELECT count(), sum(v) FROM lf FINAL WHERE k % 7 != 6 SETTINGS max_threads = 4, max_block_size = 8192, query_plan_optimize_lazy_final = 1, max_rows_for_lazy_final = 10000000, min_filtered_ratio_for_lazy_final = 0, optimize_aggregation_in_order = 1; ``` Run it with `--send_logs_level=trace` and count the merges: ``` optimize_aggregation_in_order = 1 -> 25000 \"Merging partially aggregated blocks\" lines optimize_aggregation_in_order = 0 -> 0 ``` The plan shows the replacement's own aggregation (`GROUP BY k` with `argMax` states) under `LazyReadReplacingFinal`; with aggregation in order it emits one chunk per group into the merge stage. ### Expected performance The merge stage should batch groups the way it does for an ordinary in-order `GROUP BY` (7 merges for 50000 groups), instead of one merge per group. ### Additional context Found while triaging a test timeout in https://github.com/ClickHouse/ClickHouse/pull/111459, where the randomized `optimize_aggregation_in_order = 1` setting made a small lazy `FINAL` test query take ~525 s in every flaky check. It is unrelated to that pull request - it reproduces with the feature under test switched off and on builds that predate it - and the test there now pins the setting, but the underlying pathology is worth fixing. Related: https://github.com/ClickHouse/ClickHouse/issues/113704",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114579",
          "createdAt": "2026-08-13T03:29:56Z",
          "updatedAt": "2026-08-13T17:10:18Z",
          "timestamp": "2026-08-13T17:10:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8fa512e07b4b0da83374",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113401",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix paimon timestamp precision",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> fix: https://github.com/ClickHouse/ClickHouse/issues/112768 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables partitioned by a `TIMESTAMP` or `TIMESTAMP WITH LOCAL TIME ZONE` column of precision above 3. Such tables previously failed with `scale 6 is not supported, only support scale <= 3` before returning any row, which affected every timestamp-partitioned table written by Spark, since Spark maps both `TIMESTAMP` and `TIMESTAMP_NTZ` to Paimon `TIMESTAMP(6)`. Partition pruning on such a column now uses the full precision as well.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113401",
          "createdAt": "2026-08-05T01:12:47Z",
          "updatedAt": "2026-08-13T17:09:40Z",
          "timestamp": "2026-08-13T17:09:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "JiaQiTang98",
          "state": "open",
          "assignees": [
            "hanfei1991"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:10f59240fc78e02763ae",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111852",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111852",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix(Silk): honor O_NONBLOCK in the fiber TLS BIO",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... --- This is a latent-interaction bug that appears only when **three** components are combined — none is wrong on its own: 1. **Silk fiber sockets** (#107680): the fiber-aware OpenSSL BIO `silkBioRead`/`silkBioWrite` submits io_uring I/O and parks the caller for the socket's send/receive timeout. That is correct for a blocking read — but the BIO never consults the fd's `O_NONBLOCK` flag, unlike OpenSSL's default socket BIO, which returns `EAGAIN` immediately when the flag is set. 2. **TLS** (`USE_SSL`): the affected path is the TLS BIO. The plain-socket staleness check uses a raw `recv(MSG_PEEK | MSG_DONTWAIT)` and never touches this code, so the problem is TLS-only. 3. **The connection pool's `SSL_peek` staleness probe** (#110402): before reusing a pooled connection, `DB::getSocketState` flips the fd non-blocking via `ScopedNonBlocking` (a raw `fcntl(F_SETFL, O_NONBLOCK)`, behind Poco's back) and calls `SSL_peek`, expecting an immediate `EAGAIN` — its comment reads *\"The socket is non-blocking, so this never blocks.\"* #110402 introduced this probe, replacing the previous `poll`-based keep-alive disconnect check. Only with all three present does that non-blocking `SSL_peek` route through the silk BIO, which ignores `O_NONBLOCK` and blocks for the receive timeout left on the socket by the previous request. Each component is individually correct; the fix lands on the silk BIO because it is the one whose behavior diverges from OpenSSL's socket-BIO contract (honoring `O_NONBLOCK`), while the TLS layer and the `#110402` probe are behaving as intended. On `master` today the silk BIO has no wired production consumer — it is infrastructure — so this three-way combination is not yet reachable in a shipped server; it was reproduced with downstream work that routes object-storage-disk connections through silk fiber sockets. **Impact.** Every borrow of a pooled TLS connection to an object-storage disk pays a timeout it should not. A server loading tables from an HTTPS object-store disk at startup does many such borrows and stalls — a deterministic ~13.5 s in a local reproduction, and an unbounded boot hang (never reaching \"Ready for connections\", no error logged) with production timeouts or a zero/unset receive timeout, where the wait becomes a deadline-less `future.wait()`. **Root-cause evidence** (local TLS-MinIO reproduction): a server-side request trace showed each request arriving only *after* its wait expired; the ~13.5 s decomposed exactly into the adaptive per-method receive timeouts paid in sequence (GET 500 ms + PUT 3000 ms + DELETE 10000 ms, `ConnectionTimeouts.cpp`); `ss` showed frozen `bytes_sent` through each stall; and a live backtrace was parked at `silkBioRead` ← `SSL_peek` ← `getSslSocketState` ← `isStale` ← `getConnection`. A control with silk sockets disabled does the same step in <15 ms. The plain-HTTP staleness probe uses a raw `recv(MSG_PEEK|MSG_DONTWAIT)` and is unaffected — the bug is TLS-specific. **Fix.** When the fd is non-blocking, `silkBioRead`/`silkBioWrite` do a direct `recv`/`send` with `MSG_DONTWAIT` and set the BIO retry flags (immediate `EAGAIN` → `SSL_ERROR_WANT_READ`), matching OpenSSL's default BIO. The fiber/io_uring path is unchanged for blocking sockets. The non-blocking state is read fresh from the fd on each call via `fcntl(F_GETFL)`, because `Poco::Net::SocketImpl::getBlocking()` is a cached flag the raw-`fcntl` probe never updates (and silk sockets reject `setBlocking(false)` outright). Adds a regression test (`SilkFiberSecureSocketTest.NonBlockingPeekDoesNotBlockOnIdleConnection`) that drives the real `getSocketState` path against an idle TLS connection with a 5 s receive timeout and asserts it returns in under 500 ms; without the fix it blocks the full timeout. Not for changelog: the silk fiber BIO is infrastructure with no in-tree production consumer yet, so no released user is affected. Related: https://github.com/ClickHouse/ClickHouse/pull/107680 Related: https://github.com/ClickHouse/ClickHouse/pull/110402",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111852",
          "createdAt": "2026-07-24T20:53:25Z",
          "updatedAt": "2026-08-13T17:09:37Z",
          "timestamp": "2026-08-13T17:09:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "mstetsyuk"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:525a85d5ab590b81f4ad",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114646",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114646",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fail-closed allowlist of hypothetical index types for WHATIF",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ...",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114646",
          "createdAt": "2026-08-13T14:06:56Z",
          "updatedAt": "2026-08-13T17:07:49Z",
          "timestamp": "2026-08-13T17:07:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "yariks5s",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:16cd56b3367fee37037e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112667",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112667",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Silk integration",
          "text": "Splits the silk runtime integration out of https://github.com/ClickHouse/ClickHouse/pull/111275, so that it can be reviewed on its own. This adds the plumbing that lets ClickHouse run work on [silk](https://github.com/ClickHouse/silk) fibers, without yet putting any subsystem on them. - `Silk::initializeFiberScheduler` / `Silk::destroyFiberScheduler`, called by the server when the `enable_silk_runtime` server setting is enabled. The fiber stack size is configurable through the `silk.fiber_stack_size` configuration key (320 KiB by default, which leaves enough room for OpenSSL handshakes). - `FiberLocal` - fiber-local storage. A fiber can migrate between operating-system threads, so it must not observe another fiber's `thread_local` state; the values of the registered slots are swapped in and out on every fiber switch instead. `current_thread` (`ThreadStatus`), the OpenTelemetry tracing context, and the memory-tracker and exception blockers are moved to it. - The silk thread-local-storage sanitizer: an LLVM pass in `utils/silk-thread-local-storage-sanitizer` that instruments every `thread_local` access and aborts when a fiber touches raw thread-local storage. Without it, a variable that was not migrated to `FiberLocal` produces silent corruption rather than a diagnostic. It is enabled in the debug and ASan CI builds. - `Silk::ConnectionPool` and `Silk::streamSocketFactory` - a `Connection` pool and a socket factory that suspend the calling fiber instead of blocking the operating-system thread. `PoolBase` and `ConnectionPool` are templated on the lock and the condition variable to make that possible, and `ConnectionPool` stays an alias of the `std::mutex` instantiation, so the existing call sites are unchanged. - Memory that the runtime maps outside the C++ heap - fiber stacks and `io_uring` rings - is charged to `total_memory_tracker` through silk's mmap accounting hooks. - The low-level silk runtime counters are exported to `system.asynchronous_metrics` under a `Silk` prefix. - `Common/Fiber.h` and `Common/FiberStack.h` are renamed to `Common/StackfulCoroutine.h` and `Common/CoroutineStack.h`. They implement the boost-context coroutines used by `AsyncTaskExecutor`, which are unrelated to silk fibers, and having two different things called \"fiber\" in the same codebase is confusing. Related: https://github.com/ClickHouse/ClickHouse/pull/111275 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental support for the [silk](https://github.com/ClickHouse/silk) fiber runtime, enabled with the `enable_silk_runtime` server setting. When it is enabled, the server initializes the silk fiber scheduler at startup, so that subsystems supporting it can run their jobs on fibers instead of occupying an operating-system thread while waiting for I/O.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112667",
          "createdAt": "2026-07-30T21:32:29Z",
          "updatedAt": "2026-08-13T17:07:43Z",
          "timestamp": "2026-08-13T17:07:43Z",
          "metrics": {
            "reactions": 1,
            "comments": 1
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "CheSema"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:806dd14ca36d78a7d66e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:80353",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:80353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Redis-wire protocol",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Add an opt-in Redis wire-protocol server backed by `Join` tables: a configured `redis.port` serves `GET`/`MGET` and `HGET`/`HMGET` point lookups (plus `AUTH`, `SELECT`, `PING`, `ECHO`) against ClickHouse tables mapped to Redis database numbers. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/80353",
          "createdAt": "2025-05-16T14:55:43Z",
          "updatedAt": "2026-08-13T17:03:46Z",
          "timestamp": "2026-08-13T17:03:46Z",
          "metrics": {
            "reactions": 4,
            "comments": 10
          },
          "labels": [
            "pr-feature",
            "comp-protocols",
            "manual approve",
            "can be tested"
          ],
          "author": "m4ttheux",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e19f2de52c0b4142b58e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:103706",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:103706",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "ReaderExecutor: pipeline-based read orchestration",
          "text": "Builds on top of #103234 (ReadPipeline). Gated by `SET use_reader_executor = 1` (experimental, default off). This commit replaces the matryoshka of `ReadBuffer` wrappers with a `ReaderExecutor` that owns offset mapping, cache decisions, prefetch and decryption in one place. Caches plug in via a uniform `ICacheProvider` / `ICacheHandle` API. ## What this brings **Zero-copy page-cache hits.** Reads served from `PageCache` reach the caller as a `shared_ptr` into the mmap'd cell — no `memcpy` between cache and working buffer (compare to `CachedInMemoryReadBufferFromFile`'s per-block copy). **Connection reuse across remote read calls.** Sequential reads against the same S3/Azure/HTTP object reuse an open buffer instead of issuing a new HTTP request per window. `SourceBufferLimit` caps live connections globally (`max_remote_read_connections`, default 1000); over-limit reads fall back to stateless open-read-close. Visible via `system.remote_read_connections`. **Coordinated cache + prefetch decisions.** The matryoshka design had each layer making independent choices (prefetch ahead in `AsynchronousBoundedReadBuffer`, decide what to cache in `CachedOnDiskReadBufferFromFile`, gather across objects in `ReadBufferFromRemoteFSGather`). The executor sees the whole window: - prefetches only what isn't already cache-hit - retains source over-read only when paired with a live connection (drops it otherwise) - knows which bytes are user-requested vs cache-fill, so it doesn't account cache-fill bytes against the user's read budget **Per-object cache identity, not per-pipeline.** `DiskCacheProvider` derives `FileCacheKey` / `FileCacheOriginInfo` per `StoredObject` (etag-keyed caching for `StorageObjectStorageSource`, segment-key-type classification for `Data` vs `System` queues). The old single-key-per-executor approach silently shared cache identity across all objects in a gather-mode read. **Foundation for new cache backends.** New caches (vector cache, distributed cache, etc.) implement one interface (`ICacheProvider::lookup`) and plug into the chain without touching read paths. ## Observability - `system.remote_read_connections` — open connections (path, query id, position, elapsed) - `system.reader_executor_log` — per-executor stats (cache hit/miss/populate bytes, prefetch hits/cancellations, source request count, decrypt time, …) - `ProfileEvents`: `ReaderExecutor*` counters and microsecond timers; `LiveSourceBuffer*` counters - `HistogramMetrics`: cache get/populate/source-read/prefetch-wait latencies ## Server settings - `reader_executor_prefetch_pool_size` (default 8) — shared prefetch thread pool size - `reader_executor_prefetch_queue_size` (default 0 = `pool_size * 10`) — pool queue depth - `max_remote_read_connections` (default 1000) — live-buffer slot count, reload-able ## Internals - `Rope` / `RopeNode` / `OwnedRopeBuffer` — refcounted buffer chains with mixed provenance (owned, page-cache-pinned), built-in cursor (`peek`/`advance`/`tryRewind`), sorted-on-insert nodes, disjoint-interval coverage tracking. - `ICacheProvider` / `ICacheHandle` — uniform cache API (lookup → status/get/put). Two implementations: `PageCacheProvider` (file-level, zero-copy), `DiskCacheProvider` (per-object, FileCache-backed). - `ISourceReader` — stateless range-read interface. `LocalSourceReader`, `ObjectStorageSourceReader`, `BufferSourceReader` (adapter for backup / `BufferCreator`). - `OffsetMap` — replaces `ReadBufferFromRemoteFSGather`'s gather logic with a logical-to-(object, object-offset) lookup; supports single-object unknown-size (S3 `HEAD` without `Content-Length` → streams to EOF). - `ReaderExecutor` — owns: position, offset map, cache chain, prefetch handle, live buffer, over-read tail, source-buffer slot. - `PipelineReadBuffer` — thin `ReadBufferFromFileBase` exposing the executor through `BufferBase::set/next/seek` (so legacy callers see no change). - `PrefetchThreadPool` — shared bounded pool, returns `nullptr` on overflow (sync fallback), task cancellation is race-free. ## Testing ~3,700 lines of tests (≈46% of the PR), 130+ gtest cases: - `gtest_rope` — 31 cases (append / peek / advance / tryRewind / slice / copyTo / coverage queries / shift). - `gtest_reader_executor` — 41+ cases (cache-chain combinations, per-object lookup, prefetch races, EOF release, slot leak fixes, unknown-size streaming, …). - `gtest_filecache` — 22 cases including over-read / bypass-mode / partial coverage scenarios. - `gtest_read_pipeline` — known/unknown size selection through the executor. - Functional: `04262_reader_executor_observability` plus opt-outs (`SET use_reader_executor = 0`) on tests that intentionally exercise the legacy path. ## Load tesing the simpliest cache API\" ``` ┌─────────┬────────────┬──────────┬───────┬────────────────────────┐ │ regime │ legacy QPS │ exec QPS │ QPS Δ │ per-query-type latency │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ cold │ 0.252 │ 0.320 │ +27% │ −12% (faster) ✅ │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ warm │ 0.407 │ 0.252 │ −38%¹ │ +27% (slower) ❌ │ ├─────────┼────────────┼──────────┼───────┼────────────────────────┤ │ partial │ 0.247 │ 0.251 │ +2% │ +37% (slower) ❌ │ └─────────┴────────────┴──────────┴───────┴────────────────────────┘ ``` changed cache API (stream aware): ``` Full clean verdict — 9e9abe9d (26.6.1.1) ┌─────────────┬─────────────────────────┬───────────────────────────────────────────────────┐ │ Regime │ Executor vs legacy │ Note │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ cold │ −15% (win) │ 0 slot fail; coalesced reads │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ warm │ −8% (win) │ 0 slot fail; cache reads healthy │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ⚠️ partial │ +7% │ S3-source over-read / churn on uncached half │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ❌ populate │ +17% (worst) │ 50% connection reuse on cache-fill path — churn │ ├─────────────┼─────────────────────────┼───────────────────────────────────────────────────┤ │ ✅ stress │ 27 ok / 1 OOM vs 4 / 11 │ controlled degradation = net win (bounded, alive) │ └─────────────┴─────────────────────────┴───────────────────────────────────────────────────┘ stress details: ┌───────────────────────┬───────────┬──────────┐ │ │ legacy │ executor │ ├───────────────────────┼───────────┼──────────┤ │ queries finished (ok) │ 4 │ 27 │ ├───────────────────────┼───────────┼──────────┤ │ failed │ 11 │ 1 │ ├───────────────────────┼───────────┼──────────┤ │ OOM │ 11 │ 1 │ ├───────────────────────┼───────────┼──────────┤ │ connection reuse │ 86.8% │ 98.8% │ ├───────────────────────┼───────────┼──────────┤ │ connection resets │ 233,618 │ 38,581 │ ├───────────────────────┼───────────┼──────────┤ │ peak TCP recv-buffer │ 4,152 MiB │ 813 MiB │ ├───────────────────────┼───────────┼──────────┤ │ TCP sockets │ 17,333 │ 1,548 │ ├───────────────────────┼───────────┼──────────┤ │ cgroup mem │ 24.4 GiB │ 24.6 GiB │ └───────────────────────┴───────────┴──────────┘ sha -- 494fcfe8825c ┌─────────┬──────────────┬──────────────┬──────────────┐ │ regime │ baseline QPS │ executor QPS │ exec/base │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ cold │ 0.207 │ 0.292 │ 1.41× (+41%) │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ warm │ 0.330 │ 0.307 │ 0.93× (−7%) │ ├─────────┼──────────────┼──────────────┼──────────────┤ │ partial │ 0.375 │ 0.420 │ 1.12× (+12%) │ └─────────┴──────────────┴──────────────┴──────────────┘ ``` That is good outcome. I need to optimize the work with caches. Need to stream data from available cache segments without any additional costs. sha -- dd4ae00fdaac (26.7.1.1) ``` ┌──────────┬──────────────┬──────────────┬──────────────┐ │ regime │ baseline QPS │ executor QPS │ exec/base │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ cold │ 0.230 │ 0.328 │ 1.43× (+43%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ warm │ 0.244 │ 0.302 │ 1.24× (+24%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ partial │ 0.332 │ 0.325 │ 0.98× (−2%) │ ├──────────┼──────────────┼──────────────┼──────────────┤ │ populate │ 0.348 │ 0.384 │ 1.10× (+10%) │ └──────────┴──────────────┴──────────────┴──────────────┘ ``` First build with every regime at parity or better. The warm coordination-CPU tax and the populate cache-write regression are both gone; populate over-read down to 14%. Run under production-default networking (`disk_connections_rcvbuf=204800`, `max_remote_read_connections=1000`), with `reader_executor_use_long_connections=1` set explicitly (its default flipped to 0 on this build). Long-connection hit rate 99.5–100% on all regimes. sha -- fcccec1c7e6c (26.7.1.1) — equal-workload methodology ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 959 s │ 705 s │ 1.36× (+36%) │ │ warm │ 665 s │ 450 s │ 1.48× (+48%) │ │ partial │ 592 s │ 540 s │ 1.10× (+10%) │ │ populate │ 552 s │ 516 s │ 1.07× (+7%) │ │ evict │ 601 s │ 565 s │ 1.06× (+6%) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Methodology fix: both arms now run the identical query multiset (`--iterations`, sequential — no `--randomize`), metric = total wall time. The earlier windowed-QPS numbers systematically penalized the faster arm (a free benchmark worker draws a new random query, so the faster arm attracts more heavy queries) — warm was reported −9…−14% but is actually **+48%**; `evict` = new regime with the FileCache as a transit buffer (continuous eviction, no cross-query hits). The executor is faster in all five cache regimes; the 128-thread stress arm now completes without stuck queries (previously required `KILL QUERY`), with 4× fewer sockets than legacy. No correctness errors or crashes across the campaign since `c089b7e5`. Known remaining costs (per-query analysis): (1) cold/partial S3 over-read 2.2–2.8× from prefetch speculation; (2) on mixed cache regimes small queries regress 2-3× — `cache_get` returns zero bytes for data that is present (`FileSegmentWait` on DOWNLOADING segments, then reads from source anyway) plus block-granular request amplification on narrow columns (~90 MiB requested for a 5 MiB column), paid at S3 first-byte latency with ~90% of prefetches cancelled. sha -- 018c0257567b (26.7.1.1) — plan-look-ahead window build ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 810 s │ 619 s │ 1.31× (+31%) │ │ warm │ 477 s │ 450 s │ 1.06× (+6%) │ │ partial │ 785 s │ 518 s │ 1.52× (+52%) │ │ populate │ 532 s │ 478 s │ 1.11× (+11%) │ │ evict │ 579 s │ 581 s │ 1.00× (par) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Same equal-workload methodology (identical 186-query multiset per arm). Since the workload is fixed, executor wall times are directly comparable across builds: vs `fcccec1c` the executor improved on cold (705→619 s, −12%), populate (516→478 s, −7%) and partial (540→518 s, −4%) — the plan-window changes helped. Baseline wall times swing between rounds (legacy path + master merge + day variance), so cross-build conclusions should use the executor-vs-executor columns, not the ratios. Also corrected with the fixed workload: the true equal-work S3 over-read is ~2.0× on cold and ~2.3× on partial (the earlier 2.8–2.9× figures were inflated by the windowed methodology — the faster arm simply ran more queries per window). Stress arm self-completed again (4th consecutive build). No `Code: 33`, no crashes. sha -- 2b43930ebbcf (26.8.1.1) — read-path log cut + settings-history dedup ``` ┌──────────┬───────────────┬───────────────┬──────────────┐ │ regime │ baseline wall │ executor wall │ speedup │ ├──────────┼───────────────┼───────────────┼──────────────┤ │ cold │ 496 s │ 315 s │ 1.57× (+57%) │ │ warm │ 223 s │ 226 s │ 0.99× (par) │ │ partial │ 326 s │ 254 s │ 1.28× (+28%) │ │ populate │ 242 s │ 240 s │ 1.01× (par) │ │ evict │ 281 s │ 304 s │ 0.92× (−8%) │ └──────────┴───────────────┴───────────────┴──────────────┘ ``` Two rounds this cycle. The morning round (head `a6465d03`) exposed a hot-read-path logging tax: `PipelineReadBuffer` logged one `Trace` message per window advance (6.3M messages per warm arm) plus ~1M per-buffer `Debug` \"Created\" messages, costing 795 s of thread time per arm (baseline: 6.7 s) — roughly 14% of the CPU budget on cache-served regimes, which showed as warm/populate 0.95×. The same-day fix (`da34b431`) is verified by this round: logger time dropped 795 → 170 s, allocation volume dropped 5.85 → 1.94 TB ≈ baseline's 1.84 TB (most of the long-standing \"executor allocates ~4× more\" observation was log-message formatting), and executor `UserTime` is now below baseline on warm. Populate flipped to parity-plus; warm is at 0.99× with the residual logger time still 34× baseline — a few more sites may be worth demoting. S3 over-read: populate 1.01× and evict 1.03× (the executor no longer over-fetches where the cache absorbs writes), cold 1.20× while making 7× fewer GETs at ~11 MiB each (this is what buys the 1.57× cold wall), partial 1.38× — the remaining cost item. The evict 0.92× turned out to be regime noise: a clean repeat pair on the settled cluster came back 323/279 s = 1.16× in the executor's favor, matching the morning round's 1.17× (evict has a history of single-pair swings — 0.78× → 1.18× between two runs of one build earlier). Verdict: executor ≥ baseline in all five regimes. Environment notes for cross-entry comparison: starting with this cycle the staging instance's default profile `compatibility` was bumped `25.12` → `26.8`, which changes effective defaults for both arms — all walls in this entry are ~2× faster than in previous entries for that reason; compare ratios, not walls, across entries. Relatedly, the settings-history dedup in this head means Cloud `compatibility` no longer resurrects the pre-release 8 MiB `reader_executor_plan_look_ahead_max_window` introduction default on private builds. Zero failed queries, no `Code: 33`, no crashes across all 20 arms of both rounds. ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add experimental `ReaderExecutor` for pipeline-based read orchestration with unified cache API (`PageCacheProvider`, `DiskCacheProvider`), connection reuse for remote reads, shared prefetch pool, and `system.remote_read_connections` / `system.reader_executor_log` observability tables. Enable with `SET use_reader_executor = 1`. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/103706",
          "createdAt": "2026-04-29T07:27:30Z",
          "updatedAt": "2026-08-13T17:03:28Z",
          "timestamp": "2026-08-13T17:03:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "CheSema",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b958c253ddb3200a3b22",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:63383",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:63383",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Improve the performance of `MODIFY TTL`",
          "text": "`ALTER TABLE ... MODIFY TTL` currently rewrites every part of the table, which on a large table means reading and writing all of its data just to change when rows expire. Very often the new TTL is the old one shifted in time - the retention period is extended or shortened, e.g. `create_time + INTERVAL 300 DAY` becomes `create_time + INTERVAL 10 DAY`. In that case every row's expiry time moves by the same constant number of seconds, so the parts do not have to be rewritten at all: it is enough to shift the expiry timestamps ClickHouse already stores per part. This pull request adds that fast path, and the result for the user is that such a `MODIFY TTL` completes almost instantly instead of taking minutes or hours: ```sql CREATE TABLE test_fast_ttl (`id` UInt32, `name` String, `create_time` DateTime) ENGINE = MergeTree ORDER BY id TTL create_time + toIntervalDay(300); INSERT INTO test_fast_ttl SELECT number, 'AAA', date_sub(day, 100, now()) from numbers(100000000); -- Before ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 25.564 sec. -- After ALTER TABLE test_fast_ttl MODIFY TTL create_time + INTERVAL 10 DAY; -- 0 rows in set. Elapsed: 0.046 sec. ``` There is nothing to enable and no new syntax: the optimization is applied automatically inside the `MATERIALIZE TTL` mutation that `MODIFY TTL` already produces, and a plain `ALTER TABLE ... MATERIALIZE TTL` benefits from it as well. The observable result is exactly the same as before - the same rows expire and the parts end up with the same TTL bounds - only the work is avoided. Per part, the mutation now does one of the following: - the part is fully expired under the new TTL - it is replaced with an empty part; - no row of the part is expired yet - the part is cloned (its data files hardlinked) and only its stored TTL bounds are shifted; - otherwise - the part is rewritten exactly as before. The fast path is only taken when it is provably equivalent to the rewrite. It requires that the unconditional rows TTL (`TTL <expr>`) is the only TTL of the table, and that the old and the new TTL are the same date/time column shifted by constant fixed-length intervals, so that `new_ttl(row) - old_ttl(row)` is one constant for every row. Calendar `MONTH`/`YEAR` intervals, `DAY`/`WEEK` intervals in a time zone with daylight saving time, and row-dependent expressions are all rejected. The proof is redone for each part against the TTL expression (and time zone) that the part's stored timestamps were actually computed under, so a part that lags the table metadata, or was written by an older server, falls back to the regular rewrite rather than being shifted unsoundly. The same applies to the boundary cases of the stored timestamps themselves: a part containing a row whose TTL timestamp is exactly `1970-01-01 00:00:00` UTC (which ClickHouse treats as \"no TTL\"), a shift that would move some timestamp onto that value, and a part whose stored TTL is already fully expired (which the regular rewrite drops wholesale, even when the new TTL is longer) all take the regular rewrite. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `ALTER TABLE ... MODIFY TTL` no longer rewrites the table's data when the new TTL is the old one shifted by a constant amount of time (the same date/time column plus fixed-length intervals), which is the common case of extending or shortening the retention period. Fully expired parts are replaced with empty ones and the rest are cloned with only their stored TTL metadata shifted, which makes such an `ALTER` nearly instant. Cases where the shift is not provably constant - calendar month/year intervals, day/week intervals in a time zone with daylight saving time, or row-dependent TTL expressions - fall back to the regular rewrite. `ALTER TABLE ... MATERIALIZE TTL` benefits from the same optimization. <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> > Information about CI checks: https://clickhouse.com/docs/en/development/continuous-integration/ <details> <summary>Modify your CI run</summary> **NOTE:** If your merge the PR with modified CI you **MUST KNOW** what you are doing **NOTE:** Checked options will be applied if set before CI RunConfig/PrepareRunConfig step #### Include tests (required builds will be added automatically): - [ ] <!---ci_include_fast--> Fast test - [ ] <!---ci_include_integration--> Integration Tests - [ ] <!---ci_include_stateless--> Stateless tests - [ ] <!---ci_include_stateful--> Stateful tests - [ ] <!---ci_include_unit--> Unit tests - [ ] <!---ci_include_performance--> Performance tests - [ ] <!---ci_include_asan--> All with ASAN - [ ] <!---ci_include_tsan--> All with TSAN - [ ] <!---ci_include_analyzer--> All with Analyzer - [ ] <!---ci_include_azure --> All with Azure - [ ] <!---ci_include_KEYWORD--> Add your option here #### Exclude tests: - [ ] <!---ci_exclude_fast--> Fast test - [ ] <!---ci_exclude_integration--> Integration Tests - [ ] <!---ci_exclude_stateless--> Stateless tests - [ ] <!---ci_exclude_stateful--> Stateful tests - [ ] <!---ci_exclude_performance--> Performance tests - [ ] <!---ci_exclude_asan--> All with ASAN - [ ] <!---ci_exclude_tsan--> All with TSAN - [ ] <!---ci_exclude_msan--> All with MSAN - [ ] <!---ci_exclude_ubsan--> All with UBSAN - [ ] <!---ci_exclude_coverage--> All with Coverage - [ ] <!---ci_exclude_aarch64--> All with Aarch64 - [ ] <!---ci_exclude_KEYWORD--> Add your option here #### Extra options: - [ ] <!---do_not_test--> do not test (only style check) - [ ] <!---no_merge_commit--> disable merge-commit (no merge from master before tests) - [ ] <!---no_ci_cache--> disable CI cache (job reuse) #### Only specified batches in multi-batch jobs: - [ ] <!---batch_0--> 1 - [ ] <!---batch_1--> 2 - [ ] <!---batch_2--> 3 - [ ] <!---batch_3--> 4 <details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/63383",
          "createdAt": "2024-05-05T15:40:47Z",
          "updatedAt": "2026-08-13T17:02:37Z",
          "timestamp": "2026-08-13T17:02:37Z",
          "metrics": {
            "reactions": 1,
            "comments": 33
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "zhongyuankai",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c6d5102d7791cd1f365d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104993",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104993",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Handle non-constant RHS for `IN`",
          "text": "Fix `IN` and `NOT IN` expressions with non-constant right-hand side operands that reference columns from the current row. # old analyzer Previously, the old analyzer tried to build a standalone `Set` for expressions such as `number % 2 IN (number % 3, number % 5)`, which made the right-hand side unable to resolve `number` and produced an `UNKNOWN_IDENTIFIER` exception. This change rewrites such expressions to row-wise `has` expressions instead. ```sql --- error was reported for the query below for old analyzer as mentioned in #58242 SET enable_analyzer = 0; SELECT number FROM numbers(10) WHERE number % 2 IN (number % 3, number % 5) ORDER BY number; ``` ### Out of scope: bare source-column RHS under the old analyzer Under the old analyzer (`enable_analyzer = 0`), a bare column of the `FROM` source as the right-hand side, such as `x IN (arr)` where `arr` is a column of the current row, still fails with `UNKNOWN_TABLE`: `MarkTableIdentifiersVisitor` rewrites `x IN ident` into `x IN (SELECT * FROM ident)` before source columns are collected, so the expression never reaches the row-wise rewrite. The new analyzer resolves the same query as a column and succeeds; tests `04234` and `04812` pin this divergence explicitly. Closing it needs either reordering that visitor after source columns are known, or falling back from a table to a column when the table does not exist, plus a compatibility decision for `x IN t` when a column shadows an existing table name - that is tracked in the review discussion and left out of this PR on purpose. # new analyzer The new analyzer already handled the basic non-constant right-hand side case, but some tuple and NULL cases still failed. This change fixes: * tuple-typed right-hand side expressions produced by functions other than tuple * tuple left-hand side membership checks that previously tried to create `Nullable(Tuple(...))` * `NULL` operands in non-constant tuple right-hand side operands, where the old cast target could become `Nullable(Nothing)` Examples: ```sql --- Before this fix, the new analyzer treated the tuple-typed if RHS as a single tuple value and failed with a type error; after this fix, it expands the tuple value one level for scalar IN, so the query returns 1 SELECT number IN (if(number >= 0, tuple(number, number + 1), tuple(0, 0))) FROM numbers(1); ``` ```sql --- tuple in tuple, user exepects some rows to match, however, error like `Cannot create column with type 'Nullable(Tuple(UInt8, UInt8))' because Nullable Tuple type is not allowed` will be reported before this fix SELECT number, (1, 1) IN ((number % 3, number % 2), (2, 2)) FROM numbers(6) ORDER BY number; ``` ```sql --- user expects `NULL` but error like `Conversion from UInt8 to Nothing is not supported` will be reported before this fix SELECT x IN (y, 1) FROM ( SELECT materialize(NULL) AS x, materialize(2) AS y ); ``` Issue: https://github.com/ClickHouse/ClickHouse/issues/58242 ### Changelog category (leave one): - Bug Fix ### Changelog entry: - Fix IN and NOT IN expressions with non-constant right-hand side operands referencing columns from the current row, and align new analyzer tuple right-hand side handling with existing ClickHouse IN semantics. Under the old analyzer, a bare source column as the right-hand side (`x IN (arr)`) still resolves as a table name and stays out of scope. This closes [#58242](https://github.com/ClickHouse/ClickHouse/issues/58242)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104993",
          "createdAt": "2026-05-15T04:11:24Z",
          "updatedAt": "2026-08-13T17:00:52Z",
          "timestamp": "2026-08-13T17:00:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 27
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "niyue",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3553ef014d4dfcc14e90",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108721",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108721",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Detect when tables behind a query have changed",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/108713 Adds a way to detect when the data behind a query has changed, and uses it to make the query cache consistent and to skip unnecessary refreshes of materialized views. A new virtual method `getModificationHash` on `IStorage` returns an optional `UInt128` that changes whenever the data behind the table changes (similar to an HTTP ETag). It is not a hash of the data and the way it is computed is engine-specific. It returns `NULL` when the engine cannot give a usable value, so callers fail closed. Implemented for engines that can expose a loop-free (no-ABA) value - one that never returns to an earlier value across a change-and-change-back (an `A -> B -> A` transition): - `MergeTree` family: the block-number ranges of the active data parts, their content checksums, the table structure/key metadata, a per-lifetime counter that advances on every active-part-set change (so a drop that restores an identical part set does not reproduce an earlier hash), and a per-lifetime metadata version that advances on every metadata change (so a metadata `ALTER` and back does not either). Changes on insert, merge, mutation, `ALTER`. - `Memory`: identity of the current set of blocks plus the row count and a monotonic version. - `Log`, `TinyLog`, `StripeLog`: total rows and bytes plus structure and a monotonic version. - `Merge` and `Distributed`: combine the hashes of the underlying tables (`Distributed` asks each shard through `system.tables`), folding each table's identity (database, name, UUID) so two different tables with the same hash are distinguished; they fail closed if any underlying table does. - `View` and `MaterializedView` are looked through: a view hashes the tables behind its stored `SELECT` (plus its own UUID, columns, and security metadata), a materialized view hashes its target table. They fail closed for parameterized views and in databases without table UUIDs (`Ordinary`), where incarnations of a re-created view cannot be told apart. `URL` and object storage (`S3`, ...) trust the resource's strong (non-weak) `ETag` when one is exposed, and fail closed (report `NULL`) otherwise - for glob/failover patterns, a weak or absent `ETag`, or an unreachable source. The `ETag` is loop-free for content (the same `ETag` denotes the same content), but unlike the engines above there is no monotonic version to fold, so an `A -> B -> A` rewrite back to byte-identical content within a single query's read window can in principle repeat it. That residual is narrow and both consumers below are opt-in, so we keep these engines in the feature rather than drop them. `File` fails closed: its only change signals are size and modification time, both weak. It is exposed as a lazily-computed `modification_hash` column in `system.tables`, and used for two features: - New setting `query_cache_use_only_when_data_was_not_changed`: when enabled, the combined modification hash of the tables referenced by a query is folded into the query cache key, so a cached result is reused only while none of those tables changed. If consistency cannot be guaranteed (e.g. a query calling a non-deterministic function, whose result can change while every referenced table is unchanged; a table function; a `File` table; a `URL`/object-storage table without a strong `ETag`; an object-storage read pruned by a `_path`/`_file`/Hive-partition filter, where the consumed subset cannot be compared with the full listing; a query inside a transaction, which reads the transaction's snapshot while the hash samples the live table state; or a referenced table with an active row policy for the current user, which changes what the user reads while every referenced table is unchanged - row policies existing only on remote shard servers of a `Distributed` table cannot be seen and remain part of the best-effort window), the query cache is bypassed for that query. - `REFRESH ... IF CHANGED` for refreshable materialized views: a scheduled refresh is skipped when none of the tables the view reads from changed since the last refresh that rebuilt the view (e.g. `REFRESH EVERY 1 MINUTE IF CHANGED`). It always rebuilds when a source table cannot prove it is unchanged. The `clickhouse` binary was built and all new tests pass locally, including a functional check that `Distributed.modification_hash` changes when a remote table changes and that `IF CHANGED` skips refreshes while the source is unchanged. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a `modification_hash` column to `system.tables` that changes whenever the data behind a table changes. Based on it, added a setting `query_cache_use_only_when_data_was_not_changed` to make the query cache consistent (a cached result is reused only while the referenced tables are unchanged) and a `REFRESH ... IF CHANGED` option for refreshable materialized views to skip refreshes when the source data did not change.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108721",
          "createdAt": "2026-06-28T00:01:15Z",
          "updatedAt": "2026-08-13T16:59:40Z",
          "timestamp": "2026-08-13T16:59:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 56
          },
          "labels": [
            "pr-feature",
            "pr-autogenerated-docs"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:e25a992b3f216b0bfe99",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114283",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114283",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add pre-hook to insert CI links into PR body",
          "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) -- Adds a `ci_links.py` pre-hook to the `PR` workflow that, on upstream `ClickHouse/ClickHouse` pull request runs, appends a `:ci_links:` block to the PR description with: - a link to the workflow report, and - a link to a GitHub search for the corresponding sync PR (`sync-upstream/pr/<number>`). The block is added only when it is not already present, so subsequent runs do not re-edit the PR body. Non-upstream / non-PR runs are skipped, and any failure is caught so it can never break the workflow. <!-- CI automatic block start :ci_links: --> --- Workflow [[PR](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114283&sha=latest&name_0=PR)] Sync PR [[sync-upstream/pr/114283](https://github.com/search?q=head%3Async-upstream%2Fpr%2F114283+org%3AClickHouse+type%3Apr&type=pullrequests)] <!-- CI automatic block end :ci_links: -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114283",
          "createdAt": "2026-08-11T07:48:23Z",
          "updatedAt": "2026-08-13T16:59:18Z",
          "timestamp": "2026-08-13T16:59:18Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:98b5ee2c92474858478f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113691",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113691",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `theilsU` window state returning noise when the frame's first argument is constant",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/80373 Related: https://github.com/ClickHouse/ClickHouse/pull/93384 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `theilsU` over a window frame returning an arbitrary value instead of 0 when the first argument is constant within the frame. ### Description `TheilsUWindowData::getResult` (the window-optimized state introduced in https://github.com/ClickHouse/ClickHouse/pull/93384) computes the entropy `H(A)` from cached incremental `Σ n·log n` sums. When the first argument is constant within the frame, the true `H(A)` is zero, and the computed value is pure rounding noise from the incremental updates. The code compared it against exact zero, so a tiny positive noise value passed the check, and `1 - H(A|B) / H(A)` then divided noise by noise: in debug builds this tripped the sanity check as the exception `Logical error: 'res < 1.0 + 1e-4'`, and in release builds the function could return an arbitrary value in $[0, 1]$ instead of 0. The exact (non-window) code path recomputes the entropies from the count maps, where a constant column gives `log(1) = 0` exactly, so it is not affected. The fix compares `H(A)` against an error bound proportional to `N · ε · log N` instead of exact zero, and widens the sanity-check tolerance by the same relative amount so that near-threshold frames do not trip it either. Found by the AST fuzzer on an unrelated PR (it hit https://github.com/ClickHouse/ClickHouse/pull/80373 and https://github.com/ClickHouse/ClickHouse/pull/107667): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=80373&sha=6c271049214aa5a94bd9a5f12fb27a9ffa75648f&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113691",
          "createdAt": "2026-08-06T15:21:02Z",
          "updatedAt": "2026-08-13T16:58:52Z",
          "timestamp": "2026-08-13T16:58:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d8c20e916b9eb4cfc170",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114629",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114629",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix DPsub join reordering silently dropping single-table ON-clause filters",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111898 The `dpsub` join-order algorithm silently dropped single-table filter and constant predicates that live in a `JOIN ... ON` clause (e.g. `t1.value = 'x'`), returning extra rows. `greedy` uses a different placement path and was unaffected. The predicates are placed in `collectJoinEdgesMask`. Two placement conditions there each silently dropped such a predicate: - `two_relations` required the whole join step to be exactly two relations, so the predicate was dropped whenever its relation was introduced against an already-multi-relation subplan (e.g. `t1` at the top of `t1 JOIN (t2 JOIN t3)`). - `fromLeft() || fromRight() || fromNone()` dropped the predicate for any single-table filter on a relation whose id is >= 2, because `fromLeft`/`fromRight` test relation ids 0 and 1 specifically (they describe the two inputs of a binary join step, not \"references a single relation\"). The predicate is now attached at the join that introduces its relation (the split whose one side is exactly that relation), and pure constants at the earliest two-relation join. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fixed `dpsub` join-order optimization (`query_plan_optimize_join_order_algorithm = 'dpsub'`) silently dropping single-table filter conditions from a `JOIN ... ON` clause, which could return extra rows.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114629",
          "createdAt": "2026-08-13T12:38:37Z",
          "updatedAt": "2026-08-13T16:57:48Z",
          "timestamp": "2026-08-13T16:57:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "fkastrati",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:32816f33c9eb83e8f880",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101039",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101039",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add leader election for non-replicated MergeTree on shared storage",
          "text": "Add leader election for non-replicated MergeTree tables on shared object storage (currently `S3`; `Azure` is implemented but not yet enabled, pending test coverage), enabling active/standby failover without external coordination (no Keeper). Uses conditional writes (`If-Match` / `If-None-Match`) on object storage to maintain a lease file with JSON content `{\"version\":1,\"leader_id\":\"...\",\"timestamp\":...}`. The leader renews its lease periodically; followers monitor and claim leadership when the lease expires. New MergeTree settings: - `leader_election` (Bool, default false) — enable leader election - `leader_election_heartbeat_interval` (Seconds, default 10) — lease renewal interval - `leader_election_session_timeout` (Seconds, default 30) — lease expiry threshold; must be at least 3x the heartbeat interval When not the leader, inserts, merges, mutations, and DDL are blocked; background data processing is skipped. Participating nodes should keep their clocks synchronized (NTP) to within `leader_election_session_timeout`; the conditional-write protocol always prevents split-brain, but excessive clock skew can cause leadership churn. Closes https://github.com/ClickHouse/ClickHouse/issues/91613 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `leader_election` setting for non-replicated MergeTree tables on shared object storage, enabling active/standby failover using conditional writes without external coordination. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) Three new MergeTree-level settings are added: - `leader_election` (Bool) — When enabled on a non-replicated MergeTree table stored on an `S3` object storage disk (`Azure` is implemented but rejected at table creation until it has test coverage), multiple ClickHouse instances sharing the same data path will elect a single leader. Only the leader performs writes, merges, and mutations. Followers act as read-only replicas and automatically claim leadership when the current leader's lease expires. - `leader_election_heartbeat_interval` (Seconds, default 10) — How often the leader renews its lease and followers check for an expired lease. - `leader_election_session_timeout` (Seconds, default 30) — How long a lease remains valid without renewal. Must be at least 3x `leader_election_heartbeat_interval`. Example: ```sql CREATE TABLE shared_table (x UInt64) ENGINE = MergeTree ORDER BY x SETTINGS leader_election = true, leader_election_heartbeat_interval = 10, leader_election_session_timeout = 30; ``` <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **High Risk** > High risk because it introduces new leader-election coordination and modifies core `MergeTree` write/DDL/drop and catalog cleanup paths; bugs could cause write unavailability or accidental deletion/incorrect visibility on shared storage. > > **Overview** > Adds a **Beta** `leader_election` mode for non-replicated `MergeTree` tables on shared S3/Azure, using a conditional-write lease file to ensure only one instance performs inserts/merges/mutations while followers remain read-only and can take over on failure. > > Introduces new MergeTree settings (`leader_election`, `leader_election_heartbeat_interval`, `leader_election_session_timeout`) with validation, new `MergeTreeLeaderElection*` metrics/events, and follower part-refresh + leader takeover sync to load new parts and advance block counters before enabling writes. > > Hardens destructive operations for shared-storage tables: blocks most `ALTER`/partition-mutation operations and `RENAME` under leader election, gates background processing on leadership, and updates drop/cleanup logic (`dropSkipsDataDirectoryCleanup`, `DatabaseCatalog`/`StorageTableProxy` fail-closed behavior) to avoid recursive deletion or hangs when tables are shared or cannot be materialized. Adds documentation and integration/stateless tests covering failover, metrics, validation, and rejection cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 78c1252915ba512bebf94a147932158e1bacecf0. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101039",
          "createdAt": "2026-03-28T21:56:21Z",
          "updatedAt": "2026-08-13T16:57:36Z",
          "timestamp": "2026-08-13T16:57:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 87
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8b071f117993bd8ff6ca",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114625",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114625",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix the bit-sliced full adder in `groupNumericIndexedVector`",
          "text": "<!-- CURSOR_AGENT_PR_BODY_BEGIN --> Closes: https://github.com/ClickHouse/ClickHouse/issues/106208 Related: https://github.com/ClickHouse/ClickHouse/pull/110072 `addValue` set a bit when the computed sum bit was 1 but never cleared it when the sum bit was 0, so every carry left the lower bit set and the write path was not addition: ```sql SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), val)) FROM (SELECT arrayJoin([toInt64(10), toInt64(10)]) AS val); -- {5:30}, expected {5:20} SELECT numericIndexedVectorToMap(groupNumericIndexedVectorState(toUInt8(5), toInt64(1))) FROM numbers(8); -- {5:255}, expected {5:8} ``` Any repeated index whose addends share a set bit is affected, on every index type and in both the small and the promoted representation. `n` rows of value `1` at one index accumulate to `2^n - 1`. Additions whose bits are disjoint need no carry and were already correct (`10 + 5` gives `15`), which is why the existing tests and the documented examples — all of which use distinct indexes — did not catch it. `merge` and `numericIndexedVectorPointwiseAdd` share `pointwiseAddInplace`, which computes whole-bitmap XORs and assigns the result, so clearing is implicit there and those paths were already correct. Only the per-row path was wrong, which is why `numericIndexedVectorAllValueSum` disagreed with `sum(value)` over the same rows. `RoaringBitmapWithSmallSet` had no way to clear an element, so this adds a `remove`. `SmallSet` has no erase and the small set holds at most `small_set_size` elements, so that path rebuilds it without the removed value. `zero_indexes` is now maintained too. It holds the present indexes whose value is zero, so it has to gain an index when an update drives the value to zero and lose it when the value becomes non-zero — the same invariant the pointwise operations restore when they finish: ```cpp /// For any of the total_indexes, if it is not in the non-zero index of the result, the result is 0. total_indexes->rb_andnot(*getAllNonZeroIndex()); zero_indexes = total_indexes; ``` With that, adding `5` and then `-5` row by row produces `{5:0}`, matching what merging the two values already produced, and `numericIndexedVectorGetValue` and `numericIndexedVectorCardinality` agree with the map. `numericIndexedVectorBuild` also goes through `addValue`, but from a map whose keys are unique, so it never carried and is unaffected. The test asserts `numericIndexedVectorAllValueSum` equals `sum(value)` over repeated indexes with negative and fractional values, that the row-by-row path agrees with the pointwise path, and that an index driven to zero is present with value zero. Eight of its eleven assertions fail without the fix; the three that pass are controls — an addition with disjoint bits, the pointwise reference path, and adding zero to an index that already holds a value. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `groupNumericIndexedVector` producing wrong values when the same index appears in more than one row. The bit-sliced adder never cleared a bit when the computed sum bit was zero, so every carry left the lower bit set: eight rows of value `1` at one index accumulated to `255` instead of `8`, and `10 + 10` produced `30` instead of `20`. Values that share no set bits were unaffected. An index whose value is driven to zero is now reported as present with value zero, consistent with `numericIndexedVectorPointwiseAdd` and with merging aggregate states. <!-- CURSOR_AGENT_PR_BODY_END --> <div><a href=\"https://cursor.com/agents/bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e?cursor_ref=pr_footer&cursor_cta=open_in_web\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-web-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-web-light.png\"><img alt=\"Open in Web\" width=\"114\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-web-dark.png\"></picture></a>&nbsp;<a href=\"https://cursor.com/background-agent?bcId=bc-457a9741-41fc-41c5-88d7-c8b0b5bb887e&cursor_ref=pr_footer&cursor_cta=open_in_cursor\"><picture><source media=\"(prefers-color-scheme: dark)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"><source media=\"(prefers-color-scheme: light)\" srcset=\"https://cursor.com/assets/images/open-in-cursor-light.png\"><img alt=\"Open in Cursor\" width=\"131\" height=\"28\" src=\"https://cursor.com/assets/images/open-in-cursor-dark.png\"></picture></a>&nbsp;</div>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114625",
          "createdAt": "2026-08-13T12:12:29Z",
          "updatedAt": "2026-08-13T16:57:26Z",
          "timestamp": "2026-08-13T16:57:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:dd36901c652dd9da7bb1",
        "signalId": "github:ClickHouse/ClickHouse:issue:114612",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114612",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Heap-use-after-free: Parquet v3 prefetcher reads and writes through a `ReadBuffer` freed by `IInputFormat::onFinish`",
          "text": "🕵️ ## Describe what's wrong `ParquetV3BlockInputFormat` does not override `resetReadBuffer()`, so `IInputFormat::onFinish()` frees the format's owned `ReadBuffer` while the Parquet `Prefetcher`'s IO tasks are still reading through it on the prefetch thread pool. The tasks then dereference freed memory. ASan on a plain `clickhouse local` run of the reproducer below (master, 26.8.1.1310): ``` ==ERROR: AddressSanitizer: heap-use-after-free on address 0x7112bd001800 WRITE of size 1048576 at 0x7112bd001800 thread T9 (ParquetPrefetch) #3 in DB::ReadBuffer::next() src/IO/ReadBuffer.cpp:113:15 #5 in DB::ReadBuffer::read(char*, unsigned long) src/IO/ReadBuffer.h:169:37 #6 in DB::Parquet::Prefetcher::readSync(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:113:29 #7 in DB::Parquet::Prefetcher::runTask(...) src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:540:13 #8 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() src/Processors/Formats/Impl/Parquet/Prefetcher.cpp:435 #15 in DB::ThreadPoolCallbackRunnerFast::threadFunction() src/Common/threadPoolCallbackRunner.cpp:225:13 ``` It is a **1 MiB write into freed heap**, so this is memory corruption, not only a bad read. The same defect was hit independently by the `La Casa Del Dolor (arm_asan_ubsan)` CI job, where the report is a read and carries the full free/allocation stacks: ``` ERROR: AddressSanitizer: heap-use-after-free on address 0xfcdd2aec39c0 READ of size 8 at 0xfcdd2aec39c0 thread T423 (ThreadPool) #0 in DB::Parquet::Prefetcher::readSync(...) Prefetcher.cpp:110:21 <- reader->setReadUntilEnd() #1 in DB::Parquet::Prefetcher::runTask(...) Prefetcher.cpp:523:13 #2 in DB::Parquet::Prefetcher::scheduleTask(...)::$_0::operator()() Prefetcher.cpp:424:17 #9 in DB::ThreadPoolCallbackRunnerFast::threadFunction() threadPoolCallbackRunner.cpp:225:13 0xfcdd2aec39c0 is located 0 bytes inside of 272-byte region freed by thread T468 (ThreadPool) here: #1 in std::default_delete<DB::ReadBuffer>::operator()(DB::ReadBuffer*) #7 in std::vector<std::unique_ptr<DB::ReadBuffer>>::~vector() #8 in DB::ISource::work() src/Processors/ISource.cpp:140:13 ... #18 in DB::IPolygonDictionary::loadData() src/Dictionaries/PolygonDictionary.cpp:337:8 previously allocated by thread T468 (ThreadPool) here: #2 in DB::FileDictionarySource::loadAll() src/Dictionaries/FileDictionarySource.cpp:60:19 ``` `ISource.cpp:140` is the `onFinish()` call inside the `catch (...)` handler, and the freed 272-byte region is the `ReadBufferFromFile` that `FileDictionarySource::loadAll` handed to the format via `addBuffer`. ## How to reproduce Reproduces on master (26.8.1.1310) with the official `build_amd_asan_ubsan` binary, **8 out of 10 runs, with no ThreadFuzzer and no unusual settings**. Under ThreadFuzzer it hit on the first run. Run Fiddle: https://fiddle.clickhouse.com/06f26a34-d286-40a1-85f5-7d3e952d47b8 On a release build this prints only the load error and exits 53: `Code: 53. CAST AS Array can only be performed between same-dimensional Array ... While executing ParquetV3BlockInputFormat. (TYPE_MISMATCH)` — the 1 MiB write into freed memory is silent. On an ASan build the process dies with the report above instead (5/5 runs of exactly the commands above; 8/10 in an earlier variant, so it is a race, but a very wide one). The polygon dictionary is only a convenient way to make the pipeline throw mid-read while owning its `ReadBuffer` — the defect is in the format, not in the dictionary. ## Root cause `Prefetcher` protects only *its own* lifetime. `scheduleTask` captures the shutdown handle and each task takes `std::shared_lock(*_shutdown, std::try_to_lock)`; `~Prefetcher` calls `shutdown->shutdown()` (`ShutdownHelper`, `src/Common/threadPoolCallbackRunner.h:507`), which blocks until every in-flight task has released. As its own comment says, \"`this` is safe to access as long as `shutdown_lock` is held\". But `Prefetcher::reader` points at a `ReadBuffer` owned by `IInputFormat::owned_buffers`, an unrelated lifetime that the handshake does not cover: - `IInputFormat::onFinish()` -> `resetReadBuffer()` -> `resetOwnedBuffers()` -> `owned_buffers.clear()` - `ParquetV3BlockInputFormat` overrides `resetParser()` and `onCancel()`, but **not `resetReadBuffer()`**. `resetParser()` happens to be safe only because it destroys the reader *before* delegating to the base. `resetReadBuffer()` has no such ordering, so the buffer dies while the `ReadManager` -> `Reader` -> `Prefetcher` chain is still alive with tasks running. Three routes reach it: `onFinish()` from the `catch (...)` in `ISource::work()` (the one above), `onFinish()` on the normal completion path (`ISource.cpp:133`) with speculative prefetches still outstanding, and `IInputFormat::setReadBuffer` when a format is reused across files. ## Suggested fix Mirror what `resetParser()` already does, so `~Prefetcher` drains the IO tasks before the base frees the buffer: ```cpp void ParquetV3BlockInputFormat::resetReadBuffer() { /// ~Prefetcher waits for in-flight IO tasks, which read through the buffer that /// IInputFormat::resetReadBuffer() is about to free. { std::lock_guard lock(reader_mutex); reader.reset(); } IInputFormat::resetReadBuffer(); } ``` ## Additional context Distinct from #109678 / #112573. That was a data race on `AsynchronousBoundedReadBuffer::prefetch_future` consumed by concurrent `readBigAt` (the `RandomRead` branch, `Prefetcher.cpp:102`). This one is a lifetime bug in the `SeekAndRead` branch, which already holds `read_mutex` — locking cannot help once the object is freed — and the freed buffer here is a plain `ReadBufferFromFile`, so `AsynchronousBoundedReadBuffer` is not involved at all. The binary used above already contains #112573 (merged into 26.8.1.1240). Found by this Dolor run: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=94148&sha=a2ef409981fc77a168eaf8cc68d7ddf67c9fadec&name_0=PR&name_1=La+Casa+Del+Dolor+%28arm_asan_ubsan%29",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114612",
          "createdAt": "2026-08-13T09:21:28Z",
          "updatedAt": "2026-08-13T16:56:05Z",
          "timestamp": "2026-08-13T16:56:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "bug",
            "blocker",
            "comp-parquet-reader-v3"
          ],
          "author": "PedroTadim",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3a9d11df08e282594046",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114623",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114623",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "CI: Cache: restrict cross-branch reuse to pull_request workflows only",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/112358 CI cache reuse in praktika keyed the reuse decision off the branch that produced a record. Branch, though, was only ever a proxy for how much a record can be trusted. This gates reuse on the producing **workflow event** instead and ignores the branch entirely in the reuse decision. Correctness is unaffected either way — a digest match already means identical inputs; the producing event only encodes how much we trust the result. Each cache record now carries the event that produced it. The policy: | Producing event → <br> Reusing event ↓ | `pull_request` | `push` / `schedule` / `dispatch` / `merge_queue` | |---|:---:|:---:| | `pull_request` | ✅ | ✅ | | `push` / `schedule` / `dispatch` / `merge_queue` | ❌ | ✅ | In words: a `pull_request` run reuses any record; every other (trusted) event reuses any record **except** one produced by a `pull_request`. The write side is the dual — a `pull_request` run only fills an empty `(job, digest)` slot (`if_not_exist=True`), while every trusted event overwrites it, so the shared slot always keeps a record reusable by every lane. This keeps the two properties the branch rule aimed at, more directly: - **Security.** A `pull_request` run executes untrusted (possibly fork) code, so a trusted lane must never reuse a record it produced. The trust boundary is stated as pull_request-vs-not rather than inferred from a branch name. - **Drift guard (#112358).** `merge_queue` never reuses a `pull_request` record, so a PR's green flaky-check result can't satisfy the queue's lookup and skip the re-run against the merge-group state — and this now holds independently of whether the PR and merge-queue digests happen to coincide. It also removes a clobbering hole: reuse no longer checks the branch, so a record produced by a release-branch push (or any trusted event) is reused whenever the digest matches, instead of being rejected on a branch mismatch. `CACHE_VERSION` is bumped to 2 so pre-existing records, which lack the event field, are not reused under the new rule. ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114623",
          "createdAt": "2026-08-13T11:50:30Z",
          "updatedAt": "2026-08-13T16:55:54Z",
          "timestamp": "2026-08-13T16:55:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [
            "leshikus"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fe146e6c4896f93df07e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111287",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111287",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix double free when finalizing -State aggregates under looping combinators",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server crash (double free) that could happen when finalizing an aggregate function with the `-State` combinator nested under a looping combinator (`-Resample`, `-ForEach`, `-Map`), for example `groupArrayStateResample`, if a memory limit was reached during finalization. ### Description Reported on https://github.com/ClickHouse/ClickHouse/pull/110975 (unrelated to that PR). Found by the Stress test (amd_debug): a segfault in `Aggregator::prepareChunkAndFillWithoutKey`, reached from `ConvertingAggregatedToChunksTransform::initialize`. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110975&sha=838d0b61235b06939c8acd923ebd396f51cf5b10&name_0=PR&name_1=Stress%20test%20%28amd_debug%29 Root cause: the `-State` combinator transfers its result by aliasing the raw aggregate state pointer into a `ColumnAggregateFunction` (`AggregateFunctionState::insertResultInto` -> `getData().push_back(place)`); ownership passes to the column. `Aggregator::insertAggregatesIntoColumns` relies on this transfer being atomic per place: on an exception it destroys the whole place exactly once. A looping combinator nested over `-State` aliases many sub-states one at a time; the `push_back` into the column's pointer array can reallocate and, being memory-tracked, throw `MEMORY_LIMIT_EXCEEDED` mid-loop. The already-transferred sub-states are then freed once by the aggregator's full `destroy()` and again by `~ColumnAggregateFunction`, i.e. a double free. Reproducer (crashes without the fix, returns a memory-limit error with it): ```sql SELECT arrayMap(x -> finalizeAggregation(x), state) FROM (SELECT groupArrayStateResample(0, 1048576, 1)(number, number % 20) AS state FROM numbers(100000)) SETTINGS max_memory_usage = 150000000, max_rows_to_read = 0; ``` Fix: reserve the destination columns before the transfer loop so the aliasing `push_back`s cannot reallocate (and therefore cannot throw) once a transfer has started. `ColumnAggregateFunction` used the no-op `IColumn::reserve`, so a real `reserve()`/`capacity()` over its state-pointer array is added. For `-Map`, the (possibly variable-width) key inserts are moved into their own loop before the value transfer, keeping the throwing work out of the aliasing loop. Reserving happens before any aliasing, so a throw there is harmless. The transfer loop is now non-throwing at the point of aliasing, restoring the atomic-per-place contract; results are unchanged. The fix covers all three looping transfer combinators (`-Resample`, `-ForEach`, `-Map`), which share the aliasing path; non-looping combinators delegate a single call and are already atomic. The added stateless test reproduces the crash deterministically via `-Resample` (empty buckets keep memory low until the finalization transfer, so a memory limit reliably lands the throw mid-transfer). `-ForEach` and `-Map` build their sub-states eagerly during aggregation, so they are not deterministically reproducible under a memory limit, but are fixed as the same class via the shared transfer path.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111287",
          "createdAt": "2026-07-21T20:34:03Z",
          "updatedAt": "2026-08-13T16:55:51Z",
          "timestamp": "2026-08-13T16:55:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4f5a423ed16667d59587",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114220",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114220",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.6: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114220",
          "createdAt": "2026-08-10T20:07:09Z",
          "updatedAt": "2026-08-13T16:55:43Z",
          "timestamp": "2026-08-13T16:55:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:41953d43556fecddbcfb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:104437",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:104437",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add spatial_bbox skip index for MergeTree geometry columns",
          "text": "Part of making ClickHouse fastest spatial analytical engine on Earth https://github.com/bacek/chgeos/blob/main/BENCHMARK.md ;) ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Adds `spatial_bbox` skip index for MergeTree geometry columns. The index stores a bounding box per granule and skips granules whose geometry cannot intersect the query geometry, reducing work for spatial predicates. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/104437",
          "createdAt": "2026-05-08T22:48:54Z",
          "updatedAt": "2026-08-13T16:55:01Z",
          "timestamp": "2026-08-13T16:55:01Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "bacek",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0d1a10c962c32fd270ce",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109368",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109368",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Require a join subquery alias only when it removes a real ambiguity",
          "text": "`joined_subquery_requires_alias = 1` (the default) rejected every unaliased subquery, table function or union used in a multi-table join, even when the missing alias could not cause any ambiguity. That is stricter than necessary: an alias only serves to qualify a column, so it is only needed when the unaliased table expression exposes a column whose name also occurs in another table expression of the same join. `validateJoinTableExpressionWithoutAlias` now throws `ALIAS_REQUIRED` only on such a name collision (computed with the existing `getColumnsFromTableExpression` helper over the sibling table expressions), and otherwise allows the missing alias. Genuine ambiguities involving non-sibling table expressions are still caught later by the normal `AMBIGUOUS_IDENTIFIER` resolution, exactly as they are for ordinary tables. This lets standard queries such as TPC-DS q14 (whose `cross_items` derived table has no correlation name) run without setting `joined_subquery_requires_alias = 0`: ```sql SELECT i_item_sk FROM item, (SELECT iss.i_brand_id AS brand_id FROM store_sales, item AS iss ...) WHERE i_brand_id = brand_id; -- no shared column name -> no alias needed ``` Notes: - Validation in `resolveJoin`/`resolveCrossJoin` is moved to run after all table expressions of the join are resolved, so sibling columns are known when the collision is checked. - The change is purely permissive: it never turns a previously-succeeding query into an error. When the columns of any side cannot be determined it falls back to the old strict behavior. - Only the analyzer is affected; the deprecated non-analyzer path keeps the stricter behavior. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): With `joined_subquery_requires_alias = 1` (the default), the analyzer now rejects an unaliased subquery or table function in a join only when the missing alias would make one of its columns unreachable, for example because the name collides with another joined table expression or is shadowed by an in-scope alias; otherwise the query is allowed. If the analyzer cannot determine the exposed names up front, it keeps the old strict behavior. Unambiguous queries, such as some standard TPC-DS queries, no longer require adding an alias. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109368",
          "createdAt": "2026-07-03T21:53:19Z",
          "updatedAt": "2026-08-13T16:54:37Z",
          "timestamp": "2026-08-13T16:54:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 26
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "novikd",
            "m-selmi"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:32421f50f32bbbc67137",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114667",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114667",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix File row policy / PREWHERE breaking DEFAULT columns missing from the data file",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114616 When a file-backed table (`File`, and the same path for object storage) has a `DEFAULT` column that is not present in the data file, a row policy or `PREWHERE` could prune the inputs of that default expression before `AddingDefaultsTransform` ran. That led to `UNKNOWN_IDENTIFIER` on current master, and on 26.7 to silently wrong row-policy results (type defaults instead of real values). This keeps columns required by `DEFAULT` expressions in the read set through prewhere pruning, and applies filters that reference defaulted columns after defaults are computed. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Fix queries over file-backed tables where a row policy or `PREWHERE` interacted with `DEFAULT` columns missing from the data file: they no longer fail with `UNKNOWN_IDENTIFIER`, and row policies are evaluated on real default values instead of type defaults.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114667",
          "createdAt": "2026-08-13T16:31:30Z",
          "updatedAt": "2026-08-13T16:52:46Z",
          "timestamp": "2026-08-13T16:52:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "Ria-K912",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:a4129a0bba2680d652d4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:99495",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:99495",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `GradualResizeProcessor` to limit effective parallelism for GROUP BY on small data volumes",
          "text": "When ClickHouse processes GROUP BY, it often overestimates the number of threads needed. With `max_threads = 64` but only a few thousand rows, all 64 `AggregatingTransform` instances get data, produce 64 partial hash tables, and the merge phase has to combine all of them — most nearly empty. This wastes time on merging overhead, which is especially noticeable for heavy aggregate states such as `uniq`, `uniqExact`, `groupArray`, etc. The new `GradualResizeProcessor` starts by pushing data to a single output port (or one port per split group when `min_outstreams_per_resize_after_split` applies), and activates all aggregation streams at once as soon as the configured row or byte threshold is crossed. For small datasets, only one aggregating thread receives data (or one per split group); for large datasets, all threads are used as before. New settings: - `min_rows_per_stream_for_gradual_resize` (default: `1000`) - `min_bytes_per_stream_for_gradual_resize` (default: `0`) When either threshold is non-zero, the pre-aggregation `StrictResize` is replaced with `GradualResize` in the pipeline. The optimization is enabled by default; set both `min_rows_per_stream_for_gradual_resize = 0` and `min_bytes_per_stream_for_gradual_resize = 0` to opt out. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve performance of GROUP BY on small data volumes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/99495",
          "createdAt": "2026-03-14T07:35:33Z",
          "updatedAt": "2026-08-13T16:52:41Z",
          "timestamp": "2026-08-13T16:52:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 30
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9ec2cf937bc4941f38ca",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114661",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114661",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Document that PREWHERE filters one join input before the JOIN\"",
          "text": "Reverts ClickHouse/ClickHouse#114484 - it is too low-quality, sorry. CC @PedroTadim @dhtclk <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1350` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114661",
          "createdAt": "2026-08-13T15:53:35Z",
          "updatedAt": "2026-08-13T16:52:06Z",
          "timestamp": "2026-08-13T16:52:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "rschu1ze",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ba8eb8ec1b057f4a43fd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113505",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113505",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "S3 tables engine",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): S3 tables engine catalog for datalakes. Same as https://github.com/ClickHouse/ClickHouse/pull/103220, but with working INSERT",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113505",
          "createdAt": "2026-08-05T14:34:19Z",
          "updatedAt": "2026-08-13T16:51:20Z",
          "timestamp": "2026-08-13T16:51:20Z",
          "metrics": {
            "reactions": 3,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fb934aca88b32e4807bc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112648",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112648",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix Iceberg Avro writer emitting optional complex fields as required",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/111775 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes an Iceberg table whose write format is `Avro` serializing a field declared `\"required\": false` whose type is a `list`, `map` or `struct` as required, with no `[\"null\", T]` union, so that the data file's own schema disagreed with the table metadata and other Iceberg readers saw a required field. Such a field is now written as the `[\"null\", T]` union the Iceberg spec uses for an optional one. ### Description `AvroSerializer::createSchemaWithSerializeFn` emits a `[\"null\", T]` union only for `Nullable` and `Variant`, so Avro nullability came purely from the type carrying a `Nullable` wrapper. The Iceberg read layer never puts that wrapper on a container: `IcebergSchemaProcessor::getFieldType` calls `makeNullable` only on the scalar branch, while the complex branch returns a bare `Array`/`Map`/`Tuple` regardless of the field's own `required` bit. An optional container's optionality is therefore unrecoverable from the type. This is the Avro half of #111775, which fixed the same defect for ORC and added the per-path metadata this consumes. The fix adds `createFieldSchemaWithSerializeFn`, used at the four sites that own a field (top-level column, tuple field, array element, map value). It consults that metadata and wraps the built schema in `[\"null\", T]`, unless the schema is already a union or null, so the union is added exactly once per path. `getIcebergType` publishes `required: true` for ClickHouse-authored containers, so the carrier is normally an externally authored schema written by ClickHouse with `'Avro'`. The one exception is `Nullable(Tuple)`, the only container ClickHouse DDL can declare optional. Field ids are unchanged and plain `FORMAT Avro` output is byte-identical. At the default `input_format_null_as_default = 1` such a file reads back unchanged. With the setting off the read throws `Cannot insert Avro Union(Null, T) into non-nullable type T`, because `getFieldType` derives a bare container for an optional complex field, the standing reader limitation Spark-written files already hit. For a `Nullable(Tuple)` column it is a change: master wrote a bare record, readable at either value, so with `enable_nullable_tuple_type = 1` (default off) and the setting persisted off at `CREATE`, a read that worked now errors. Optional element, value and field positions are unaffected, their targets being `Nullable`. The reader derivation is the root cause and is tracked separately. <details> <summary>Validation: per-field nullability in the written file's schema, master vs fix</summary> Fixture: hand-written `v1.metadata.json` with optional and required containers, `CREATE TABLE IF NOT EXISTS ... ENGINE = IcebergLocal(dir, 'Avro')`, one INSERT. Values read out of the written file's `avro.schema` header entry. | field (Iceberg) | master | with fix | |---|---|---| | `req_int` required int | required int | required int | | `opt_int` optional int | optional int | optional int | | **`opt_list` optional list** | **required array** | **optional array** | | **`opt_map` optional map** | **required map** | **optional map** | | **`opt_struct` optional struct** | **required record** | **optional record** | | `req_list` / `req_map` / `req_struct` required | required | required | | `list_opt_elem.element`, `map_opt_val.value`, `opt_struct.sy` optional | optional | optional | | **`list_opt_struct_elem.element` optional struct** | **required record** | **optional record** | | **`struct_opt_list_field.inner` optional list** | **required array** | **optional array** | | **`map_opt_struct_val.value` optional struct** | **required record** | **optional record** | Each of the four call sites is pinned separately: reverting one at a time to `createSchemaWithSerializeFn` moves a disjoint set of reference lines, the top-level site moving `opt_list`/`opt_map`/`opt_struct`, and the array-element, tuple-field and map-value sites moving one nested line each. All 13 Iceberg field ids are byte-identical between the arms, and the test also pins them at depth: not descending through the union drops four nested-id lines. 50/50 green, and 20/20 green under `compatibility='20.1'`, the regime whose pre-21.1 `input_format_null_as_default` default the read-back pin defends. </details> <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.487` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112648",
          "createdAt": "2026-07-30T18:47:15Z",
          "updatedAt": "2026-08-13T16:51:19Z",
          "timestamp": "2026-08-13T16:51:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:e03eda4848a1f72e76ca",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114620",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "assignees"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114620",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix serialization of Map-valued settings in access entities",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114591 (auto-closes the issue when this PR is merged into the default branch) --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a bug where an access entity carrying a `Map`-valued setting, such as a settings profile with `http_response_headers` or `additional_table_filters`, was stored in a form that ClickHouse could not read back, so the entity became permanently unloadable after a restart. Closes #114591. ### Description Closes #114591. `CREATE SETTINGS PROFILE p SETTINGS http_response_headers = '{...}'` succeeds, but the entity is stored as `http_response_headers = [('k', 'v')]`, which nothing can parse back. No error appears at `CREATE` time, so the entity is **permanently unloadable** after a restart: ``` stored: ATTACH SETTINGS PROFILE `p` SETTINGS http_response_headers = [('a', 'b')] CONST; after restart: Code: 62. Syntax error: failed at position 65 ([) ... Could not parse <path>/access/<uuid>.sql ``` The list rebuild drops it silently, reading it back throws, so `SELECT` from `system.settings_profile_elements` fails while it is present, as does `RESTORE` of a backup holding it. Root cause: the value is cast to the setting's native type, so a Map setting holds a `Map` Field, which `FieldVisitorToString` renders as an array of tuples; but `ParserSettingsProfileElement` reads values with a scalar-only `ParserLiteral` and cannot open a `[`. That spelling is rejected everywhere, `SET http_response_headers = [('a','b')]` included, so the write side is wrong. Fix: when a profile element's value, MIN or MAX is a `Map` and the setting is builtin, emit the setting's canonical text as a quoted string. Write side only, no grammar change. Custom settings are excluded because `castValueUtil` returns their value unchanged, so a string would come back a `String` rather than a `Map`. Covers `CREATE USER`/`ROLE`, `ALTER ... SETTINGS` and MIN/MAX, and transitively BACKUP/RESTORE and both storages. `SHOW CREATE` now prints a quoted string rather than `[('k', 'v')]`, intended since the new form is copy-pasteable. Entities already stored in the broken form are not repaired, as they were never parseable; recreate them. Downgrade is safe: a pre-fix binary reads the new form correctly, so no versioning is needed. New test `04902_access_entity_map_setting_round_trip`: 11 of its 15 arms fail on pristine master and pass here, covering empty, multi-key and hostile maps plus a two-process on-disk reload.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114620",
          "createdAt": "2026-08-13T11:13:29Z",
          "updatedAt": "2026-08-13T16:50:47Z",
          "timestamp": "2026-08-13T16:50:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "pufit"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:99067fc0e9684583366b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111867",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111867",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use Gaussian centroids for truncated QBit codes",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/109405 Related: https://github.com/ClickHouse/ClickHouse/pull/110911 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve reduced-precision `L2DistanceTransposedQuantized`, `cosineDistanceTransposedQuantized`, and `dotProductTransposedQuantized` by reconstructing `p < 8` `QBit(Int8)` codes with Gaussian conditional-mean prefix centroids. Full precision (`p = 8`) remains bit-exact; reduced-precision approximate distances may change, and recall/latency gains are workload-dependent. ### Motivation The existing reduced-precision path selects one middle fine Lloyd-Max reconstruction level for every truncated prefix. That value is not the conditional mean of the complete Gaussian interval represented by the prefix. This change uses the conditional mean `(phi(lo) - phi(hi)) / (Phi(hi) - Phi(lo))` which minimizes scalar MSE for that prefix under the standard-normal source model of the existing Lloyd-Max codec. The 127 positive values are evaluated from the existing `Float32` boundaries at high precision, rounded once to `Float32`, stored as hexadecimal literals, and mirrored exactly for negative prefixes. This keeps the hot distance loop as one LUT lookup and avoids platform-dependent libm work during LUT initialization. ### Validation - A new stateless regression fails on the exact official `26.7.1.1315` binary for all 20 reduced-precision centroid/sign checks and passes its `p = 8` control. The candidate passes all 20 checks and the control. - An independent all-raw probe covers every raw byte and every `p = 1..8`: expected level counts, finiteness, symmetry, prefix-block invariance, index-order monotonicity, independent Gaussian means (maximum 0 ULP), and bit-exact legacy `p = 8` reconstruction. - An engine probe covers all raw bytes and precisions, non-strided and `QBit(Int8, 16, 8)`, `used_dims` 8/16, dot/L2/cosine, and `optimize_qbit_distance_function_reads` 0/1. Partial-read modes match bitwise; bounded SimSIMD tolerances are used for L2/cosine. - Debug `programs/clickhouse` build passes. Updated `04504_transposed_distance_quantized` and new `04628_qbit_lloyd_max_prefix_centroid` match their references with empty stderr in clean `clickhouse local` paths. On 103,000 source-disjoint Nomic embedding vectors (768 dimensions, 200 queries, four randomized-Hadamard seeds), scalar coordinate MSE decreases at every changed precision. The transformed-space retrieval proxy is deliberately reported separately because it is not a production ClickHouse latency benchmark: | `p` | scalar MSE delta | recall@10 delta (pp) | hit@1 delta (pp) | |---:|---:|---:|---:| | 1 | -3.83% | 0.000 | 0.000 | | 2 | -3.49% | -0.250 | -0.375 | | 3 | -0.89% | +0.538 | +0.750 | | 4 | -15.17% | +1.738 | +1.625 | | 5 | -17.70% | +1.300 | +1.875 | | 6 | -24.08% | +0.800 | +0.125 | | 7 | -44.32% | +0.588 | +0.250 | | 8 | unchanged | 0.000 | 0.000 | There is no universal retrieval improvement claim: `p = 2` regresses slightly in this proxy. There is also no latency claim until a matched Release-build p50/p95 benchmark is available. The intentional compatibility boundary is numerical output at `p < 8`; function signatures, storage, and `p = 8` results are unchanged.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111867",
          "createdAt": "2026-07-25T01:52:18Z",
          "updatedAt": "2026-08-13T16:50:45Z",
          "timestamp": "2026-08-13T16:50:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 12
          },
          "labels": [
            "pr-improvement",
            "can be tested",
            "v26.7-must-backport"
          ],
          "author": "skuznetsov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9f8eedc6f9f5b5513fbd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114666",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114666",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Promote ConstantValue to Core with Field-free value accessors",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/113051 --> Continue removing `DB::Field` from constant handling by promoting the owning single-constant type to a shared primitive with `Field`-free value accessors. ### What - **Move `ConstantValue` from `Analyzer/` to `Core/`.** It bundles `{size-1 ColumnConst, DataTypePtr}`. It carries a `DataTypePtr`, and `DataTypes` depends on `Columns`, so `Core` is the correct layer (this also removes a lower layer having to reach up into `Analyzer` for the type). It stays deliberately distinct from `ColumnWithTypeAndName` (size-1 const invariant, no name, scalar accessors). - **Add value accessors** that read row 0 of the size-1 column without materializing a `Field`: `isNull`, `getUInt`, `getInt`, `getFloat64`, `getBool`, `getDataAt`. `ConstantNode` gains matching delegators, and `ConstantNode::getValue` now delegates to a single transitional `ConstantValue::getField`. - **`evaluateConstantExpressionAsColumn` now returns `ConstantValue`** instead of `std::pair<ColumnPtr, DataTypePtr>`. Callers updated: `numbers`/`primes`/`generateSeries`/`values` table functions, `ActionsVisitor`, `InterpreterSelectQuery` (LIMIT/OFFSET), prometheus/timeSeries selectors, and the gtest. The two `getStringConstArgument` helpers now read the value directly via `isNull`/`getDataAt`. ### Behavior No user-visible change: the value is the same size-1 const column with the same exact type, just bundled and readable without a `Field`. `ConstantValue`'s `Field` constructor and `getField` remain as the only `Field` entry/exit while the analyzer still folds constants into `Field`s; a later phase removes them. ### Verification `gtest_convert_column_to_type`, `gtest_evaluate_constant_expression`, and a `numbers`/`generate_series`/`values`/`LIMIT` (integer + fractional)/`IN` stateless spot-check. Related: https://github.com/ClickHouse/ClickHouse/pull/113051 ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not for changelog: internal refactor toward removing `DB::Field`; no user-visible behavior change.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114666",
          "createdAt": "2026-08-13T16:25:52Z",
          "updatedAt": "2026-08-13T16:49:57Z",
          "timestamp": "2026-08-13T16:49:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "yakov-olkhovskiy",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:e3f87f135d8a2524f102",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110838",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110838",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add introspection TCP port",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ClickHouse server now has an introspection port. This is a native protocol TCP listener that starts before the server begins attaching tables and stops only after the tables' detach completes. During these windows, an operator can connect to it with `clickhouse client` and run queries such as `SHOW PROCESSLIST`, `SELECT * FROM system.stack_trace`, or `SYSTEM INSTRUMENT ADD 'QueryMetricLog::startQuery' SLEEP ENTRY 0.5`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110838",
          "createdAt": "2026-07-17T11:07:23Z",
          "updatedAt": "2026-08-13T16:45:51Z",
          "timestamp": "2026-08-13T16:45:51Z",
          "metrics": {
            "reactions": 1,
            "comments": 16
          },
          "labels": [
            "pr-feature"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "evillique"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:556382e6d412fb7071a6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114580",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114580",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: Duplicate TLS argument rejection and positional-arity stripping untested",
          "text": "_Test-only PR. Review: are the gaps real, is the test right._ Adds test coverage for 1 untested code path, found during automated review of [PR #110615](https://github.com/ClickHouse/ClickHouse/pull/110615). That PR: (1) Adds TLS/SSL to every PostgreSQL integration: `sslmode` plus certificate/key either as server-local paths (`sslrootcert`/`sslcert`/`sslkey`, config-only) or as literal contents (`*_pem`, accepted from SQL, materialized into `TemporarySecretFile` and masked as secrets). New code: … **1. Duplicate TLS argument rejection and positional-arity stripping untested** `src/Storages/StoragePostgreSQL.cpp:726`, `src/Databases/PostgreSQL/DatabasePostgreSQL.cpp:573` **Risk:** `StoragePostgreSQL::extractSSLParamsFromArguments` strips trailing TLS `key = value` pairs and rejects repeats at `StoragePostgreSQL.cpp:726-727`; the stripped list then feeds the arity check at `DatabasePostgreSQL.cpp:573`. Risk if broken: a repeated `sslmode = 'require', sslmode = 'disable'` … **Unique vs PR tests:** 04820 formats queries with distinct TLS keys and checks masking plus path rejection; 04846 checks query-tree masking; test_postgresql_ssl exercises real handshakes through named collections. None repeats a TLS key (the `specified more than once` branch) and none combines the maximum positional … **Tags:** `-- Tags: no-fasttest` — `no-fasttest`: the PostgreSQL integration is not built in the fast test build; the PR's own 04820 uses the same tag cc @alexey-milovidov (author of #110615) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114580",
          "createdAt": "2026-08-13T03:32:59Z",
          "updatedAt": "2026-08-13T16:45:10Z",
          "timestamp": "2026-08-13T16:45:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d04cbf2eefe68ef6bf7c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111794",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111794",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add unordered stream modifier",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM UNORDERED modifier: skip the per-snapshot commit-order sort depends on https://github.com/ClickHouse/ClickHouse/pull/110653 (not for functional reason, only test) cc @alesapin @Michicosun",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111794",
          "createdAt": "2026-07-24T13:31:48Z",
          "updatedAt": "2026-08-13T16:43:58Z",
          "timestamp": "2026-08-13T16:43:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "SmitaRKulkarni",
          "state": "closed",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bb44ae0d63d07412220b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114634",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114634",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113742 to 26.6: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114634",
          "createdAt": "2026-08-13T12:49:55Z",
          "updatedAt": "2026-08-13T16:43:00Z",
          "timestamp": "2026-08-13T16:43:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8ecc877deb46c03e5114",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114316",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114316",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use the vector similarity index for integer reference vectors",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112233 Related: https://github.com/ClickHouse/ClickHouse/issues/114291 The reference vector of an ANN query is extracted only when its array type is `Float64`, `Float32` or `BFloat16` and every element is a `Float64` field. An integer literal such as `[1, 2]` is typed `Array(UInt8)`, so `tryUseVectorSearch` bails out and the query silently falls back to a brute-force scan over the whole table, although `[1, 2]` denotes the same point as `[1.0, 2.0]` and `L2Distance` accepts it. `EXPLAIN indexes = 1` shows no `vector_similarity` entry and gives no hint why, so a one-character difference in a literal becomes a sharp performance cliff on large tables. Native integer arrays are now accepted as reference vectors and their elements are converted to `Float64`, which is the type the reference vector is stored in anyway. Added `02354_vector_search_bug112233`, covering unsigned, signed, mixed integer/float, and not-exactly-representable reference vectors, plus an equality check between the integer and float spellings of the same query. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Vector search queries now use the `vector_similarity` index when the reference vector is written as an integer array literal, e.g. `ORDER BY L2Distance(vec, [1, 2])`. Previously such queries silently fell back to a brute-force scan.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114316",
          "createdAt": "2026-08-11T12:52:06Z",
          "updatedAt": "2026-08-13T16:41:55Z",
          "timestamp": "2026-08-13T16:41:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "hamidr",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bd7b307233e718590bb0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113909",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113909",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check the query cancellation while filling `system.parts` and its siblings",
          "text": "The tables based on `StorageSystemPartsBase` (`system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`) build the whole result eagerly in `initializePipeline`, so a cancelled or timed out query kept building rows over every storage and part until the very end. In a stress test with ThreadFuzzer this took minutes and tripped the hung check: a `SELECT` over `system.parts_columns` (with a filter matching every active part on the server) stayed in the process list for 252 seconds with `is_cancelled = 1` and `max_execution_time = 10`. Check the query status per storage and per part, following the pattern of `system.zookeeper` and `system.remote_data_paths`. The check also covers the storage-discovery prepass in `StoragesInfoStream` (the eager enumeration of all databases and tables), the analogous prepass in `StoragesDroppedInfoStream` (so `system.dropped_tables_parts` is interruptible as well), the skip loop in `StoragesInfoStreamBase::next`, and the per-storage part enumeration itself: the `MergeTreeData` helpers (`getDataPartsVectorForInternalUsage`, `getAllDataPartsVector`, `getProjectionPartsVectorForInternalUsage`, `getAllProjectionPartsVector`) take an optional `need_stop` callback that is checked periodically while the parts snapshot is being built. The two column-oriented tables additionally check the query status inside their column-enumeration loops and their metadata prepass, so a very wide table does not create a long uninterruptible stretch inside a single storage. The return value of `checkTimeLimit` is honored, so with `timeout_overflow_mode = 'break'` the eager build stops at the soft deadline and returns the rows collected so far. The table-lock acquisition in `StoragesInfoStreamBase::tryLockTable` is interruptible as well: instead of a single wait inside `RWLockImpl::getLock` for the whole `lock_acquire_timeout`, the lock is acquired in 100 ms slices with a query-status poll between the attempts (the total timeout and the `DEADLOCK_AVOIDED` semantics are preserved), so a killed or soft-timed-out query does not sit in the lock wait while a concurrent DDL query holds the drop lock. The test `04869_system_parts_lock_wait_cancellation` pins this with a share lock held by a long `SELECT` and a `DROP TABLE` in an `Ordinary` database queued behind it. The test uses the new `slowdown_system_parts_enumeration` failpoint, which only affects the specially named test tables (so concurrently running tests are unaffected). It sleeps 500 ms on every enumerated part, so building the full result for a 20-part table takes at least 10 seconds, and it sleeps 1 second per `COLUMNS_CANCELLATION_CHECK_PERIOD` (128) enumerated columns of a part, so building the full `system.parts_columns` / `system.projection_parts_columns` result over a single part with 1301 columns (and a projection over all of them) also takes at least 10 seconds. The test asserts that queries with a 1 second deadline in the `break` mode finish well under that, which is only possible by stopping at the per-part and per-column cancellation checkpoints. Timed assertions are needed because a plain row-count assertion cannot distinguish a build with the fix from one without: in the `break` mode the executor drops the eagerly built result after the deadline in both cases. The test also asserts partial row counts under a pre-expired deadline for all five tables, including `system.dropped_tables_parts` over a deterministic dropped-table fixture. The pre-expired-deadline checks also run under the failpoint, so the fewer-rows assertion is deterministic even on a machine fast enough to build the whole result in under a millisecond. For tables with the `_snap` name marker, the failpoint additionally slows down the parts-snapshot walks inside `MergeTreeData` (500 ms per enumerated part) and makes them poll the stop callback on every element, and for tables with the `_meta` name marker it slows down the column-metadata prepass of the column-oriented tables (1 second per 128 enumerated metadata columns), so the timed checks also prove that the snapshot materialization and the prepass are interruptible: all six of these checks fail against a binary with the `need_stop` polls and the prepass checkpoints disabled. Caught by `Stress test (arm_asan_ubsan, s3)`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113722&sha=ceedb772f6aa36e221d25ce48f9af71f032a7b08&name_0=PR&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Related: https://github.com/ClickHouse/ClickHouse/pull/113722 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Queries over `system.parts`, `system.parts_columns`, `system.projection_parts`, `system.projection_parts_columns`, and `system.dropped_tables_parts` now react to cancellation and `max_execution_time` while the result is being built.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113909",
          "createdAt": "2026-08-08T01:18:22Z",
          "updatedAt": "2026-08-13T16:41:37Z",
          "timestamp": "2026-08-13T16:41:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:89233633bf874c3eab28",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114216",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114216",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for the RemovePart part_log row in 02950 and 02491",
          "text": "### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `02950_part_log_bytes_uncompressed` and `02491_part_log_has_table_uuid` read `system.part_log` immediately after a DDL that removes a part and assert the `RemovePart` row is already there. That row is written asynchronously, so the assertion is unsound and the tests fail with just that line missing. `RemovePart` has a single emitter, `MergeTreeData::removePartsFinally`, which on both DDL paths here is reached only through `grabOldParts()`. Both DDLs (`DROP PART`, `TRUNCATE`) do run cleanup in the query thread, so usually the row lands first. But that grab selects nothing if it loses `try_lock` on `grab_old_parts_mutex`, or if the part's `shared_ptr` is not unique, e.g. while a concurrent read holds a reference. The part then stays `Outdated` and a later cleanup pass writes the row: late, not lost, so the engine is correct and only the tests need fixing. `01686_event_time_microseconds_part_log.sh` already asserts the same shape behind a bounded poll. Both tests become `.sh`, since a bounded poll is not expressible in `.sql`, and wait for the row under a 60 s bound. Each iteration issues `SYSTEM START CLEANUP <table>` to schedule a pass, because a cleanup thread that found nothing to do backs off up to `max_cleanup_delay_period`. Assertion queries, tags and both `.reference` files are unchanged; no `CREATE TABLE` setting is added here (02491's `old_parts_lifetime` pin and 02950's tags are pre-existing, kept verbatim); no `no-parallel` and no blanket `no-random-*`. Validation: holding a reference across the DDL reproduces the non-unique-ownership path deterministically. With that lever and no poll both tests fail 8/8; with the poll they pass 10/10. Dropping only the `SYSTEM START CLEANUP` line fails at the deadline once the cleanup thread is backed off. Deleting the DDL makes the poll time out rather than pass, so it is not satisfied by a stale row. Then 150/150 green per test across default, `-j 8` and randomized-order batches.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114216",
          "createdAt": "2026-08-10T20:03:37Z",
          "updatedAt": "2026-08-13T16:41:22Z",
          "timestamp": "2026-08-13T16:41:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "can be tested",
            "pr-synced-to-cloud",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c86c793e873294860156",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114546",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114546",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrapped `Time64` values from an overflowing scale conversion in `convertFieldToType`",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. --> Related: https://github.com/ClickHouse/ClickHouse/pull/94537 Related: https://github.com/ClickHouse/ClickHouse/pull/111534 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes incorrect, often sign-flipped, `Time64` literals and `IN`-list constants produced when rescaling a lower-scale `Decimal64` overflowed `Int64`. Such a conversion now reports `DECIMAL_OVERFLOW`, matching the `DateTime64` branch and explicit `CAST`. ### Description `convertFieldToTypeImpl` rescales a `Decimal64`-backed field into a `Time64` column at `src/Interpreters/convertFieldToType.cpp:473-476`. The scale-increasing arm multiplied without a range check, so `value * scale_multiplier_diff` could exceed `Int64` and wrap. The wrapped product was then handed to `decimalFromComponentsWithMultiplier<Time64>(value, 0, 1)`, whose own `mulOverflow` check is a no-op at multiplier `1`, so the corrupted value became the field. Observed on master (`bb1bd307`): `Values 'x Time64(6)' (253402207200000::Decimal64(0))` returned `-999:59:59.722624`, and `INSERT` persisted it. The wrapped `Int64` of `253402207200000 * 1000000` is `-4852209831933722624`, which renders as exactly that; the negated input wraps to `+4852209831933722624`, so a negative input returned a positive time. Explicit `CAST(... AS Time64(6))` already reported `DECIMAL_OVERFLOW` here, so the literal path disagreed with `CAST`. The file carried this same statement twice, for `DateTime64` and `Time64`, both unguarded. The related PR above guarded the `DateTime64` one and added test `03797`; its `Time64` twin was left as it was. This change mirrors that guard onto the twin. The operand also becomes `Int64`, which the guard requires: `mulOverflow` on an unsigned operand reports overflow for every negative value, which would reject in-range negative times. Reporting rather than returning Null matches the `DateTime64` twin, which `03797` asserts, and explicit `CAST`. The neighbouring `Date32` branches keep their Null contract and are untouched. Found by a UBSan signed-overflow report on this line; there is no open issue for it. It keeps reproducing on `master`, in both `asan_ubsan` stress jobs, as `signed integer overflow: 253402239600000 * 1000000 cannot be represented in type 'long'` at `src/Interpreters/convertFieldToType.cpp:475`: - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=86363e819f9cb704a892063da9998e4c1281eb76&name_0=MasterCI&name_1=Stress%20test%20%28amd_asan_ubsan%29 - https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=895af217db902c1458fecc608c6f6b3a78cc985e&name_0=MasterCI&name_1=Stress%20test%20%28arm_asan_ubsan%2C%20s3%29 Only inputs that were producing wrong values change: `9223372036854` at scale 6, the largest whose rescale still fits, is still accepted and returns `999:59:59.000000` identically. Verified by building both arms and diffing: every overflow arm goes from a wrong value to `DECIMAL_OVERFLOW`, while 30 in-range controls across scales 0/3/6/9, both signs and both scale directions are byte-identical. `03797` and `04837` still pass. New test `04883` fails on master and is green here over 100 randomized runs.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114546",
          "createdAt": "2026-08-12T21:31:34Z",
          "updatedAt": "2026-08-13T16:41:17Z",
          "timestamp": "2026-08-13T16:41:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2c5396398b3316c7dd07",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112828",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112828",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Recover a NATS JetStream subscription closed by the broker",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/96651 Related: https://github.com/ClickHouse/ClickHouse/pull/103557 Related: https://github.com/ClickHouse/ClickHouse/pull/112464 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a `NATS` table with `nats_stream` set silently consuming nothing after the NATS server is restarted. The `JetStream` subscription is now re-established automatically instead of requiring `DETACH TABLE` and `ATTACH TABLE`. ### Description A `NATS` engine table reading from a `JetStream` stream stops consuming permanently once the NATS server is restarted. Nothing is logged, the connection reports healthy, and only `DETACH TABLE` plus `ATTACH TABLE` or a server restart recovers it. It is also how the flaky `test_nats_restore_failed_connection_without_losses_on_write` fails on master. Root cause: an asynchronous pull subscription renews its pull request only when a message is delivered, and a reconnect resends the `SUB` line without the outstanding pull request, so with nothing in flight when the server goes away the chain never restarts. The server does report this, answering the outstanding request with `409 Server Shutdown`, and the client then closes the subscription. ClickHouse missed it because `isSubscribed` only tests whether the subscription vector is non-empty, so the existing re-subscribe path was gated on a predicate that cannot see a dead subscription. This adds a per-subscription liveness check and consults it in the streaming task, which drops the subscriptions so the existing re-subscribe runs in the same iteration. Only `JetStream` consumers opt in: core NATS subscriptions are already restored by the client, and recovery drops buffered messages core NATS never redelivers. Validated with five new integration tests. Three restart the broker and fail on master, 9 of 9 repeats, before passing after the change; a fourth asserts a healthy consumer never re-subscribes, so it passes either way and exists to bound the cost. The fifth covers a restart of a table reading two subjects, which nothing covered before. Not covered: a broker loss leaving the subscription with no status at all, such as a hard kill, a partition, or a loss coinciding with a re-subscribe. That needs a local fetch timeout, which would also periodically tear down healthy subscriptions. #103557 targets the same defect from a connection-level reconnect counter, but no longer applies to this code and has no integration test. Close whichever you prefer. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1343` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112828",
          "createdAt": "2026-07-31T22:25:07Z",
          "updatedAt": "2026-08-13T16:41:13Z",
          "timestamp": "2026-08-13T16:41:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 17
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "antaljanosbenjamin",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:53dbf12a697b41fb75b8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114484",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114484",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Document that PREWHERE filters one join input before the JOIN",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not required for a documentation change. ### Description `prewhere.mdx` did not mention `JOIN` at all. This is the documentation KochetovNicolai asked for when he closed issue 89097 as not-a-bug: \"We need to document this.\" One `<Note>`, next to the existing note that documents the same class of fact for `FINAL`. It states the rule and gives the equivalent explicit spelling as a filtered subquery. The `SELECT` clause list is annotated to agree with it. Measured on the example as it appears on the page: `PREWHERE b.y > 50` returns `[1,2,3,4]`, the filtered-subquery spelling the same, the `WHERE` spelling `[1,4]`. The `WHERE` result is unchanged whether `query_plan_filter_push_down` is on or off, which is why the note credits `WHERE` to the join result rather than to a fixed position, per PedroTadim's review. The note claims nothing about individual join kinds or strictnesses, since `FULL JOIN` and `ASOF INNER JOIN` both differ too, and it attributes the fill value to [join_use_nulls](https://clickhouse.com/docs/reference/settings/session-settings/join#join_use_nulls) rather than naming one. Related: https://github.com/ClickHouse/ClickHouse/issues/89097 Related: https://github.com/ClickHouse/ClickHouse/issues/114206 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1341` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114484",
          "createdAt": "2026-08-12T13:01:26Z",
          "updatedAt": "2026-08-13T16:41:08Z",
          "timestamp": "2026-08-13T16:41:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-documentation",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d27a9a1ccff965f29878",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110104",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110104",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Check for cancellation in AggregatingInOrderTransform",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107941 --> Related: https://github.com/ClickHouse/ClickHouse/issues/107941 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `optimize_aggregation_in_order` ignoring query cancellation. `AggregatingInOrderTransform` now checks for cancellation while aggregating a chunk, so a query stopped by `KILL QUERY` or by `max_execution_time` (in the default `timeout_overflow_mode = 'throw'`) stops promptly instead of running the whole chunk to completion. ### Description `AggregatingInOrderTransform::consume()` splits one input chunk into runs of equal keys in a loop. Query time and cancellation limits are only checked between pipeline steps (between `work()` calls), so a chunk with many distinct keys makes a single `consume()` call run for a long time (`O(distinct_keys)` iterations, each an `upper_bound` over the remaining rows) with no cancellation checkpoint. As a result a cancelled query (`KILL QUERY`, `max_execution_time`) using `optimize_aggregation_in_order` kept aggregating until the whole chunk was done; the connection thread then blocked in `PullingAsyncPipelineExecutor::cancel() -> ThreadFromGlobalPool::join()` waiting for that loop. The server-side AST fuzzer repeatedly hit this as `Hung check failed, possible deadlock found` (Stress test, all sanitizers), with `system.processes` showing `is_cancelled = 1` and `elapsed` far past the 90s hung-check window while the worker thread sat in `AggregatingInOrderTransform::consume -> Aggregator::executeImpl`. The loop now checks `isCancelled()` once per key interval (cheap) and returns early; the partial aggregation state is discarded because the pipeline is being torn down. This mirrors the existing per-loop cancellation checks in `WindowTransform` and `FillingTransform`. Scope: this covers cancellation that sets `is_cancelled` on the pipeline, i.e. `KILL QUERY` and `max_execution_time` in the default `timeout_overflow_mode = 'throw'`, which is what the reproduced hung check hit (`system.processes` showed `is_cancelled = 1`). The non-default `break` mode is a soft limit that returns a partial result and never sets `is_cancelled`; honoring it mid-chunk (as `FillingTransform` does via `process_list_element->checkTimeLimit()`) is a separate partial-result change, out of scope here. The regular `AggregatingTransform` behaves the same way. Regression test `04512_aggregation_in_order_cancellation` forces one long `consume()` over 40M distinct-key rows in a single chunk, `KILL QUERY ... SYNC` once every row is read: with the fix the KILL returns in a fraction of a second, without it it blocks for the several seconds the loop needs to finish. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1344` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110104",
          "createdAt": "2026-07-11T17:17:10Z",
          "updatedAt": "2026-08-13T16:41:03Z",
          "timestamp": "2026-08-13T16:41:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "yakov-olkhovskiy"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5dc37dbd1e4f5a322689",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114067",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114067",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "fix(Analyzer): skip rerunFunctionResolve for 'exists' nodes created by rewrite_in_to_join",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114026 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed `Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION)` thrown when a `PREWHERE` clause contains an `IN (subquery)` predicate and `rewrite_in_to_join = 1` (or `make_distributed_plan = 1`, which force-enables it) is set. The same query spelled with `WHERE` worked correctly. --- ### Problem With `rewrite_in_to_join = 1`, the analyzer rewrites `x IN (subquery)` into an `exists(...)` `FunctionNode` resolved via `FunctionExists`, a special function that is not registered in `FunctionFactory`. When `PREWHERE` is resolved, `ReplaceColumnsVisitor` calls `rerunFunctionResolve` on every `FunctionNode` in the predicate. `rerunFunctionResolve` (`src/Analyzer/Utils.cpp`) already special-cases `grouping` (also resolved outside the factory), but not `exists`, so it called `FunctionFactory::instance().get(\"exists\", context)` and threw `UNKNOWN_FUNCTION`. ### Fix Two parts: 1. `PREWHERE` is evaluated by the reading step and cannot execute a correlated subquery — the planner rejects one with `ILLEGAL_PREWHERE`. So the `rewrite_in_to_join` rewrite is now skipped while resolving a `PREWHERE` expression, and the plain `IN` is kept there. `PREWHERE x IN (subquery)` then returns the same result as its `WHERE` spelling, which is what the issue asks for. Subqueries nested inside `PREWHERE` still rewrite their own `IN` predicates. 2. `exists` is added to the special-case early return in `rerunFunctionResolve`, next to `grouping`. This matters for an explicitly written `PREWHERE EXISTS (correlated subquery)`, which is genuinely unsupported: it is now reported honestly as `ILLEGAL_PREWHERE` instead of `Unknown function exists`. ### Test `tests/queries/0_stateless/04820_rewrite_in_to_join_prewhere_exists.sql` covers `PREWHERE ... IN (subquery)` against the `WHERE` control arm, `NOT IN`, tuple `IN`, an `IN` nested in a subquery inside `PREWHERE`, and the explicit `PREWHERE EXISTS (...)` case asserting `ILLEGAL_PREWHERE`. ### Reproduction ```sql CREATE TABLE t (k UInt64, s String) ENGINE = MergeTree ORDER BY k; INSERT INTO t SELECT number, toString(number % 2) FROM numbers(1000); -- Threw: Code: 46. DB::Exception: Unknown function exists. (UNKNOWN_FUNCTION) SELECT count() FROM t PREWHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; -- Worked (control) SELECT count() FROM t WHERE s IN (SELECT '1') SETTINGS rewrite_in_to_join = 1, allow_experimental_correlated_subqueries = 1; ```",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114067",
          "createdAt": "2026-08-09T20:05:33Z",
          "updatedAt": "2026-08-13T16:40:58Z",
          "timestamp": "2026-08-13T16:40:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "RohithPariki",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c008160785c5f585861b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113899",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113899",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Avoid scans for constant sort keys",
          "text": "## What `isAlreadySorted` now returns early, right after the sort descriptors are resolved, when every sort key is a `ColumnConst`. Writing a `MergeTree` part whose sorting keys are all constant no longer walks the block doing adjacent-row comparisons to confirm an ordering that is constant by construction. Collation validation still runs before the early return, and any block that mixes constant and non-constant keys keeps the existing comparator path — only the all-constant case takes the new path. ## Why it helps When the sorting-key columns are constant across a block — a common shape when a leading `ORDER BY` column is fixed per part (time-ordered or batched ingestion, per-source or per-partition writes) — the sortedness check was doing a full comparison pass to reach a foregone conclusion. Returning as soon as the keys are known-constant removes that pass. Measured on `MergeTree inserts with constant and mixed sorting keys`, 64K/256K/1M rows (paired medians, co-measured on both trees): | Metric (1M rows) | Before | After | Δ | | --- | ---: | ---: | ---: | | Constant-key sortedness check, 1 key | 1106 µs | 2 µs | **−99.8%** | | Constant-key sortedness check, 4 keys | 4449 µs | 4 µs | **−99.9%** | | End-to-end insert latency, 4 keys | 24548 µs | 20201 µs | **−17.7%** | | Insert CPU, 4 keys | 20732 µs | 16301 µs | **−21.4%** | Smaller block sizes land in the same range (e.g. 1-key 64K: 66 µs → 2 µs). The sortedness check collapses to a near-constant cost, and that saving carries into a full insert as the end-to-end latency and CPU gains. Blocks that aren't all-constant take the unchanged comparator path. ## Testing Stateless coverage exercises single and multiple constant keys plus a non-constant suffix (the mixed case that must keep comparing). The declared correctness check passed. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Skip the redundant sortedness scan when writing MergeTree parts whose sorting keys are all constant. --- Contributed by [Perfloop](https://app.perfloop.ai): the numbers above were co-measured on both trees and independently re-verified before submission — the full public record is at [case_8ebwekrder](https://app.perfloop.ai/t/oss/case_8ebwekrder). Replies from this account are human-approved, and a human operator is accountable for this contribution. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1346` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113899",
          "createdAt": "2026-08-07T23:19:06Z",
          "updatedAt": "2026-08-13T16:40:53Z",
          "timestamp": "2026-08-13T16:40:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-performance",
            "can be tested",
            "pr-synced-to-cloud"
          ],
          "author": "perfloop-agent",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2d11ad469db1e4690b50",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113450",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113450",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix reading Paimon tables with a nullable ARRAY or MAP column",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/113337 Related: https://github.com/ClickHouse/ClickHouse/pull/113425 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed reading Paimon tables that contain a nullable `ARRAY` or `MAP` column. Such a table could not be read at all, because the schema mapper wrapped the composite type in `Nullable`, which ClickHouse forbids, so both `DESC` and `SELECT` failed with `Nested type Array(Nullable(Int32)) cannot be inside Nullable type`. A nullable composite column is now mapped to a non-`Nullable` composite type and a `NULL` value is read as an empty one. Closes #113337. ### Description This takes over #113425 by @zlareb1, who closed it and asked me to carry it forward. The diagnosis and the fixture are his; the source change uses the in-tree capability gate rather than deleting the wrap. **What breaks.** Paimon columns are nullable by default, so a plain `CREATE TABLE paimon.default.t (f ARRAY<INT>)` from Spark produces an unreadable table. `Paimon::DataType::parse` applied its `if (nullable)` wrap in the composite branch as well as the scalar one, and `DataTypeArray`/`DataTypeMap` return `canBeInsideNullable() == false`, so `DataTypeNullable`'s constructor threw. This happens while parsing the schema, so it takes out the whole table rather than one column. Affects `paimonS3`/`paimonLocal`/`paimonAzure`, their `*Cluster` variants, the `Paimon*` engines and the REST catalog. The engines need `allow_experimental_paimon_storage_engine`; the table functions do not. **The change.** The two inner wraps become a single `makeNullableSafe`, which wraps only when the type permits it. Neither the Iceberg nor the DeltaLake schema processor wraps a composite in `Nullable` (Iceberg gates on `canBeInsideNullable()`; DeltaLake keeps the wrap in its scalar branches only), so this aligns Paimon with them. The scalar wrap is untouched, so inner nullability survives: the fixture reads as `Array(Nullable(Int32))` and `Map(String, Nullable(Int32))`. The gate, rather than deletion, keeps the wrap available for a future `ROW`, whose `DataTypeTuple` does permit it. A `NULL` composite reads as an empty one, the Parquet reader's documented behaviour, so the two become indistinguishable. That is forced by the type system and matches Iceberg, DeltaLake, Arrow, ORC and Avro. **Validation.** New test `04757_paimon_nullable_composite_types` over @zlareb1's fixture fails on master with the error above and passes with the fix. The ten existing Paimon tests are identical on both binaries. Restoring either wrap individually reddens the new test on its own message. 50/50 randomized runs pass. `Types.h` is absent on 25.8, so 26.3 through 26.7 are affected; `must-backport` labels look appropriate.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113450",
          "createdAt": "2026-08-05T10:05:17Z",
          "updatedAt": "2026-08-13T16:40:49Z",
          "timestamp": "2026-08-13T16:40:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "pr-backports-created",
            "pr-synced-to-cloud",
            "pr-must-backport-synced",
            "v26.5-must-backport"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "JiaQiTang98"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:aea70155cf9f05ebc449",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114323",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114323",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: require canonical internal links",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114230 This is a one-off cleanup of repository-authored documentation links that use legacy redirect aliases. It updates the current English documentation and source-embedded reference documentation to use routes relative to the docs root, so the automated translation PR can parse and localize them without producing missing locale routes. This PR intentionally adds no permanent CI checks or ongoing enforcement. Its scope is limited to the current link corrections needed to get the automated translation PR parsing successfully. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114230&sha=e599855a281a4dd841be20794036908a58acbf63&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114323&sha=86d2d7e194d7fbf75cc84616a8fa0da1ed84802e&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114323",
          "createdAt": "2026-08-11T13:26:47Z",
          "updatedAt": "2026-08-13T16:39:30Z",
          "timestamp": "2026-08-13T16:39:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci",
            "pr-autogenerated-docs"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8433a15e5fafda78647f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110958",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110958",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix toTime key-expression type mismatch under use_legacy_to_time",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/107951 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server abort (`Bad cast from type ColumnVector<UInt32> to ColumnVector<Int32>`, a `LOGICAL_ERROR`) when inserting into a `MergeTree` table whose `PRIMARY KEY` or `ORDER BY` uses `toTime(...)` from a session whose `use_legacy_to_time` value differs from the one under which the key metadata was built. The `toTime` legacy/new resolution now follows the context that builds the expression, so a persisted key expression's type no longer depends on the writing session. ### Description `toTime()` resolves to two different functions depending on the `use_legacy_to_time` setting: the new `toTime` returns `Time` (Int32-backed), the legacy variant returns `DateTime` (UInt32-backed). The selection in `FunctionFactory::tryGetImpl` read the thread-local query context and ignored the `context` argument the caller passed. A `MergeTree` table with `PRIMARY KEY (toTime(c1))` / `ORDER BY toTime(c1)` persists only the expression AST. Its primary-index on-disk type comes from `metadata_snapshot->getPrimaryKey().data_types`, derived by rebuilding the key expression with the storage's global context (server-default `use_legacy_to_time`). The part-writer serializes the index with that persisted type. When an `INSERT` runs in a session with a different `use_legacy_to_time`, the write-path key expression re-resolved `toTime` to the other variant, producing a column whose physical type mismatched the persisted serialization, so `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow` hit `assert_cast<ColumnVector<Int32>>(ColumnVector<UInt32>)` and aborted the server (a handled exception in release builds, an abort under debug/sanitizers). Fix: resolve the `toTime` legacy swap from the caller-provided `context` (falling back to the thread-local query context only when no context is supplied). Stored key/sorting expressions are rebuilt with the storage global context, so they now resolve `toTime` consistently with the type persisted in the table metadata, while normal query resolution still honors the session setting. Found by the BuzzHouse fuzzer. - CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109351&sha=b0cca3209ae52b188fa092a31eb1e754db40608d&name_0=PR&name_1=BuzzHouse%20%28amd_msan%29 - Check: `BuzzHouse (amd_msan)`; assertion `Bad cast from type DB::ColumnVector<unsigned int> to DB::ColumnVector<int>` at `SerializationNumber<int>::serializeBinary` <- `MergeTreeDataPartWriterOnDisk::calculateAndSerializePrimaryIndexRow`. Reproducer: ```sql CREATE TABLE t (c0 Int32, c1 DateTime64 MATERIALIZED nowInBlock64()) ENGINE = MergeTree() PRIMARY KEY (toTime(c1)); INSERT INTO t (c0) SETTINGS use_legacy_to_time = 1 SELECT number FROM numbers(1000); ``` ### DDL behaviour change carried by the fix Previously, `CREATE TABLE` in a session whose `use_legacy_to_time` differed from the server-wide default stamped the stored key type with the session's resolution (e.g. `DateTime` when the session set `use_legacy_to_time = 1` on a server defaulting to `0`). With this fix, the stored key type always resolves under the server-wide default, so the session setting at `CREATE` time no longer affects the persisted key type. This is observable via `DESCRIBE mergeTreeIndex(...)` and is pinned by the test. Upgrade note for that narrow window (table created while the session setting differed from the server global, on a pre-fix binary): after the upgrade the same table resolves its key as `Time`, so parts written before and after store different raw key values for the same timestamp (e.g. `90000` vs `3600` for `01:00:00`); reads, inserts and merges succeed, but a key-range predicate may miss rows from old parts. Where session and global agreed (the overwhelmingly common case), nothing changes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110958",
          "createdAt": "2026-07-18T23:24:10Z",
          "updatedAt": "2026-08-13T16:39:25Z",
          "timestamp": "2026-08-13T16:39:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 22
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexey-milovidov",
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3de6e1423bf75ef2d7ca",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114643",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114643",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-land aggregate function `gini` in the `sum` family",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/113763 Related: https://github.com/ClickHouse/ClickHouse/pull/113868 Related: https://github.com/ClickHouse/ClickHouse/pull/112280 --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New aggregate function `gini`, which calculates the [Gini coefficient](https://en.wikipedia.org/wiki/Gini_coefficient) of a column of finite, non-negative numeric values. The result ranges from `0` (all values equal) towards `1` as inequality grows; for a sample of `n` values the maximum is `(n - 1) / n`. `NaN` values are skipped and infinite values are rejected. The function returns `Float64` and consumes `O(n)` memory. ### Description Re-lands the `gini` function that #113868 reverted, following the option-2 spec @ Manerone gave in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5217284116 and confirmed in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5279278069. It re-adds #112280's function with exactly three items removed, and touches no file under `src/AggregateFunctions/Combinators/`: - the `getArgumentsThatCanBeOnlyNull` override, - `.returns_default_when_only_null = true` on registration, - the `argument_type->onlyNull()` branch in the creator. The property was doing the work. It made `AggregateFunctionFactory::getImpl` skip its only-null guard and build a real `gini` instance over `Nullable(Nothing)`, so `gini(NULL)` returned `Float64` `nan` where every `sum`-family function folds to `Nullable(Nothing)`. Without it the fold happens in the `Null` combinator before any other combinator is applied, and the creator's only-null branch becomes unreachable, which is also why `createAggregateFunctionSum` has no equivalent. `gini` now registers exactly like `sum`. Runtime `Nullable` handling is a separate axis, via `getOwnNullAdapter`, and is unchanged: `gini` over `[1, NULL, 3]` still returns `0.25`. Validated on three binaries: post-revert master (`gini` absent), a build carrying #112280's version, and this branch. Every literal-`NULL` cell on this branch equals the `sum` value measured on the same binary, and all non-`NULL` results are unchanged from #112280. No documentation files are touched. The page body is generated from the `FunctionDocumentation` block in `AggregateFunctionGini.cpp`, which this PR restores, so the docs autogeneration workflow fills the page from that source. cc @Manerone",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114643",
          "createdAt": "2026-08-13T13:34:59Z",
          "updatedAt": "2026-08-13T16:38:13Z",
          "timestamp": "2026-08-13T16:38:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "Manerone"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1656e49d7484fdc2cdda",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112758",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112758",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Sync private settings",
          "text": "### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112758",
          "createdAt": "2026-07-31T14:55:06Z",
          "updatedAt": "2026-08-13T16:37:02Z",
          "timestamp": "2026-08-13T16:37:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:592dea450922be5085bf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114525",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114525",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Optimize merges of the text index",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved performance of merges of text indexes. ### Additional context A few optimizations: - The main one: reducing overhead on deserialization of embedded and small postings caused by the allocation of the bitmap - Removed unneeded conversion to roaring bitmap on build of the output posting list - Used specialized sort cursor and batch sorting strategy for merging of text index segments",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114525",
          "createdAt": "2026-08-12T17:15:12Z",
          "updatedAt": "2026-08-13T16:35:59Z",
          "timestamp": "2026-08-13T16:35:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:809be4a202f4592a672c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114602",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114602",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: regenerate reference documentation from source",
          "text": "This pull request is opened automatically by the nightly documentation autogeneration workflow. It regenerates settings, functions, table and database engines, data types, formats, table functions, window functions, system tables, and asynchronous metrics from the structured documentation embedded in the ClickHouse source and exposed through the corresponding `system.*` tables. The generator preserves page frontmatter and hand-written content outside the `{/*AUTOGENERATED_START*/}` / `{/*AUTOGENERATED_END*/}` regions. Some pages are fully generated below their frontmatter. Do not edit generated content by hand -- edit the structured documentation in the defining source code instead; the next nightly run regenerates the pages. ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1348` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114602",
          "createdAt": "2026-08-13T08:02:58Z",
          "updatedAt": "2026-08-13T16:35:49Z",
          "timestamp": "2026-08-13T16:35:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud",
            "pr-autogenerated-docs"
          ],
          "author": "clickhouse-gh[bot]",
          "state": "closed",
          "assignees": [
            "Blargian"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:3d452e5518f27a7d0205",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110344",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110344",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix wrong primary-key pruning for toStartOfDay and relative-number functions on out-of-range DateTime64",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/90461 Related: https://github.com/ClickHouse/ClickHouse/pull/108018 --> Related: #90461 Related: #108018 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed wrong `count()` results and dropped rows when a `DateTime64` primary-key column is filtered through `toStartOfDay`, `toRelativeSecondNum`, `toRelativeMinuteNum`, `toRelativeHourNum`, `toRelativeDayNum`, `toRelativeWeekNum`, `toMonthNumSinceEpoch` or `toYearNumSinceEpoch` with values outside the `UInt32`-seconds range (before 1970 or beyond 2106). These functions claim to be always monotonic to the primary index but their standard-precision results wrapped for out-of-range `DateTime64`, breaking primary-key pruning. ### Description The `DateTime64` sibling of #108018 (which fixed the same class for `Date32`). `toStartOfDay` and the relative-number transforms have `FactorTransform = ZeroTransform`, so `IFunctionDateOrDateTime::getMonotonicityForRange` reports them as always monotonic. Their standard-precision `DateTime64` code paths narrowed the result to `UInt32`/`UInt16` without saturating, so for arguments outside that range the value wrapped and the function stopped being monotonic. This makes primary-key range analysis produce exact ranges that extend before the selected mark range, which: - in release builds: silently drops granules holding matching rows, returning a wrong `count()`; - in debug/sanitizer builds: trips `chassert(exact_ranges[i].begin >= range.begin)` in the trivial-count projection optimization (`optimizeUseAggregateProjection.cpp`). Found by the AST fuzzer (amd_msan) on a query that mutated a `Date32` key column to `DateTime64(5)`: `SELECT count() FROM t WHERE toStartOfDay(d) >= toDateTime('2000-01-01 00:00:00','UTC') SETTINGS force_primary_key = 1`. Report: https://s3.amazonaws.com/clickhouse-test-reports/PRs/110310/e242401ecdb1e933c8646546bc2f905d4ed106cd/ast_fuzzer_amd_msan/fatal.log The fix saturates the `DateTime64` `execute` overloads to `[0, result-type max]`, matching the `Date32` fix in #108018, keeping each function monotonic over the whole `DateTime64` domain. Adds `04538_datetime64_zerotransform_monotonicity_pruning`; updates the references of `01768_extended_range`, `04408_datediff_datetime64_overflow` and `02403_enable_extended_results_for_datetime_functions`, which asserted the previous wrapped standard-precision results.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110344",
          "createdAt": "2026-07-14T07:56:02Z",
          "updatedAt": "2026-08-13T16:35:41Z",
          "timestamp": "2026-08-13T16:35:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:00808e3fe55d126b89ce",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114409",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114409",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Revert the PromQL topk/limitk streaming plan and its shared-subquery materialization\"",
          "text": "Reverts ClickHouse/ClickHouse#114326 Depends on https://github.com/ClickHouse/ClickHouse/pull/113397",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114409",
          "createdAt": "2026-08-12T01:15:49Z",
          "updatedAt": "2026-08-13T16:35:30Z",
          "timestamp": "2026-08-13T16:35:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:541aa93622f459081c93",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109252",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109252",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make JSONExtract honour cast_string_to_date_time_mode when parsing DateTime values",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): `JSONExtract` now honours `cast_string_to_date_time_mode` when converting string JSON values to `DateTime`/`DateTime64`, consistently with `CAST`. Closes #109126. --- Extracting a string JSON value into `DateTime`/`DateTime64` is a string-to-type cast, but `JSONExtract` keyed its parsing mode off `date_time_input_format` (an input-format parsing setting), while the equivalent `CAST` honours `cast_string_to_date_time_mode`. With `date_time_input_format = 'basic'` and `cast_string_to_date_time_mode = 'best_effort'` (reproduced on current master): ```sql SELECT JSONExtract('{\"date\":\"2020-01-01 00:00:00.123Z\"}', 'date', 'DateTime64(3)'); -- 1970-01-01 00:00:00.000 (silently returns default) SELECT toDateTime64('2020-01-01 00:00:00.123Z', 3); -- 2020-01-01 00:00:00.123",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109252",
          "createdAt": "2026-07-03T03:40:20Z",
          "updatedAt": "2026-08-13T16:34:17Z",
          "timestamp": "2026-08-13T16:34:17Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "Utkal059",
          "state": "open",
          "assignees": [
            "george-larionov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:174793135d1e999707e4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86768",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86768",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Feature: Enable overlay databases for server.",
          "text": "Enables server-side `Overlay` databases. An `Overlay` database is a read-only facade that exposes the union of the tables of several underlying databases, resolving each table name through the listed sources in order (the first source that has the table wins). DDL on the facade is rejected — it has no storage of its own — while `SELECT` and pass-through `INSERT` resolve to the underlying source table. Reading or writing through the facade requires the corresponding grant on both the facade database and the underlying source, and the facade's row policies are combined with the source's. Previously an `Overlay` database existed only as the implicit default database of `clickhouse-local`. Closes: https://github.com/ClickHouse/ClickHouse/issues/52764 <!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow creating `Overlay` databases. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86768",
          "createdAt": "2025-09-05T21:38:13Z",
          "updatedAt": "2026-08-13T16:33:47Z",
          "timestamp": "2026-08-13T16:33:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 108
          },
          "labels": [
            "pr-feature",
            "manual approve",
            "can be tested",
            "pr-autogenerated-docs"
          ],
          "author": "AlyHKafoury",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a423b68758a25374d6aa",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113996",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113996",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Observe the query time limit while collecting typo-correction hints",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/pull/86768 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed `max_execution_time` being ignored while a query was still being analyzed. Referencing an unresolved column of a deeply nested type sent typo correction into an enumeration of every subcolumn of every candidate column, exponential in the nesting depth and observing no time limit, so the query could keep running for minutes past its limit and after being cancelled. ### Description Reported on #86768, in the `Stress test (arm_tsan)` hung check ([report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=86768&sha=399ef4e99734b6a8c5cef4fd1bfbc2e724d33b8a&name_0=PR&name_1=Stress%20test%20%28arm_tsan%29)): a query with `is_cancelled: 1` had run 1379 s against `max_execution_time: 10`, stack top in `SerializationObjectPool::getOrCreate` under an `enumerateStreams` recursion, `read_rows: 0`. Root cause is in the analyzer, not the serialization pool. An unresolvable identifier takes the typo-correction path, where `TypoCorrection::collectCompoundExpressionValidIdentifiers` enumerates every subcolumn of the candidate column looking for a suggestion. That substream tree doubles per nesting level, the walk repeats per column, table expression and scope, and nothing on it observes the time limit. Two changes in `TypoCorrection.cpp`: - Return before entering the walk when it cannot contribute: the insert needs prefix plus subcolumn name to have exactly as many parts as the unresolved identifier, and a subcolumn name has at least one part, so once the prefix alone is that long nothing can come out. That holds at every call site, so the hint set is unchanged by construction. `collectScopeValidIdentifiers` already guards its walk this way. - Poll `QueryStatus::checkTimeLimit` in the surviving walk, and before the early return so the limit does not depend on entering it. It reads the query's own watch rather than waiting to be told, so the limit also holds under `timeout_overflow_mode = 'break'` and in `clickhouse-local`, neither of which ever marks the query cancelled. A `KILL QUERY` still reports its own cause; with no query attached it is a no-op. Reproduced without the fuzzer: `Array(Map(String, Tuple(...)))` nested N deep plus `SELECT alias.nosuchcol FROM t AS alias`. At depth 12 that goes from 44 s to 2.0 s, and under a 1 s limit reports `TIMEOUT_EXCEEDED` instead of running 116 s. Hints are unchanged over a 15-shape matrix. `IDataType` and `ISerialization` are untouched, so other `forEachSubcolumn` callers are unaffected. Only this carrier is closed, so the note is amended, not dropped. <details><summary>Measurements</summary> Subcolumns enumerated per candidate column: 50, 106, 218, 442 at nesting depth 3, 4, 5, 6. Debug build, `Array(Map(String, Tuple(a T, b T)))` nested N deep, `SELECT alias.nosuchcol FROM t AS alias`: | depth | 7 | 9 | 10 | 11 | 12 | |---|---|---|---|---|---| | before | 0.75 s | 3.21 s | 7.58 s | 18.09 s | 43.97 s | | after | 0.68 s | 0.50 s | 0.63 s | 1.09 s | 2.01 s | At depth 10 with `max_execution_time = 1`: before 116 s and `UNKNOWN_IDENTIFIER`, after `TIMEOUT_EXCEEDED` in 1.2 s. A one-part identifier, which no walk can ever answer, went from 44 s to 0.09 s; with a second table expression the same shape went from 35 s to 0.84 s. The two modes that never mark a query cancelled, depth 12 under a 0.001 s limit: `timeout_overflow_mode = 'break'` returned `UNKNOWN_IDENTIFIER` after 0.67 s with the limit ignored, and now returns `TIMEOUT_EXCEEDED` in 0.20 s; `clickhouse-local` went from 1.92 s with the limit ignored to `TIMEOUT_EXCEEDED` in 1.26 s. A `KILL QUERY` at depth 15 reports `QUERY_WAS_CANCELLED`, never a timeout, 8 of 8. Pre-fix the one-part shape took 348 s under a sanitizer build and 23 s on a release build, so the committed test's limits hold on every build type rather than only on debug. Hint set compared over 15 shapes (bare, alias-qualified, table-qualified and database-qualified prefixes; `Tuple`, `Nullable`, `LowCardinality`, `LowCardinality(Nullable)`, `Array(Tuple)`, `Map`, `Variant`, `Dynamic`, `JSON`, nested `Tuple`): byte-identical before and after, 12 of the 15 carrying a real hint. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113996",
          "createdAt": "2026-08-09T00:12:46Z",
          "updatedAt": "2026-08-13T16:33:43Z",
          "timestamp": "2026-08-13T16:33:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4e8981c31cbbd7a59f6d",
        "signalId": "github:ClickHouse/ClickHouse:issue:114639",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:114639",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Push dynamic TopN thresholds into MergeTree reads for ORDER BY ... LIMIT (2-3x on ClickBench Q24/Q26)",
          "text": "**Use case.** `ORDER BY ... LIMIT n` over a sorted-by-something-else table with a narrow projection, e.g. ClickBench Q24/Q26: ```sql SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime LIMIT 10; SELECT SearchPhrase FROM hits WHERE SearchPhrase <> '' ORDER BY EventTime, SearchPhrase LIMIT 10; ``` The original build of https://github.com/ClickHouse/ClickHouse/pull/81944 implemented a **dynamic TopN threshold pushdown** (`Push TopN threshold to MergeTreeSource`, commit `2e2f594308f4`, plus `Better query condition cache: make TopN dynamic filters deterministic and reusable`, gated by `max_limit_to_push_down_topn_predicate = 100`): while the partial-sort transform maintains the current top-`n`, the running n-th-best value of the `ORDER BY` key is pushed down into the `MergeTree` read as a dynamic threshold, so granules whose key range cannot beat the current top-`n` are skipped instead of read, decompressed, and sorted. This mechanism was dropped during the upstreaming of that PR (the scaffolding was removed from the branch; the surviving `RewriteOrderByLimitPass` is a different, row-offset-based approach — it is off by default and measures performance-neutral on ClickBench when enabled). Master's lazy materialization (`query_plan_optimize_lazy_materialization`) covers the wide-`SELECT *` case (Q23), but does not prune reads for narrow projections: every granule passing the `WHERE` is still fully processed. Measured on identical single-part ClickBench data (hot, interleaved runs, 96-core aarch64), original bench-opt build (25.9.1.1) vs master `405e218ff`: Q24 0.009 s vs 0.013 s, Q26 0.008 s vs 0.013 s (1.2–1.3x with times this small). The gap is much larger on the official `c7a.metal-48xl` numbers: Q24 0.014 s vs 0.046 s, Q26 0.014 s vs 0.044 s (**2–3x**). A related consideration from the original design: the dynamic filter interacts with the query condition cache, so the thresholds need to be deterministic/reusable (or excluded from the cache key) — the original branch had a follow-up commit specifically making the TopN dynamic filters deterministic for that reason. Related: https://github.com/ClickHouse/ClickHouse/pull/81944 Related: https://github.com/ClickHouse/ClickHouse/pull/81944#issuecomment-5280709772",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/114639",
          "createdAt": "2026-08-13T12:58:53Z",
          "updatedAt": "2026-08-13T16:33:01Z",
          "timestamp": "2026-08-13T16:33:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "shankar-iyer"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f25a5669111450025c9e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108820",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108820",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow Distributed and Remote tables in Replicated databases",
          "text": "`database_replicated_allow_only_replicated_engine` rejects table engines that keep their own unreplicated data on disk in a `Replicated` database. Before this change, `Distributed` was also rejected because its optional background `INSERT` queue makes `storesDataOnDisk` return true, even though the queue is a transient send buffer and the table's actual data belongs to its destination shards. This change distinguishes table-owned on-disk data from auxiliary delivery state. The restriction applies to non-`ATTACH` `CREATE` queries when `database_replicated_allow_only_replicated_engine = 1`. | Engine or group | Result | Reason | |---|---|---| | `ReplicatedMergeTree` family and `SharedMergeTree` | Allow | The engine provides a replication or shared-storage contract. | | Writable non-replicated `MergeTree` with a local or remote storage policy | Reject | It owns unreplicated on-disk table data; a remote disk alone does not establish replication. | | Static read-only non-replicated `MergeTree` | Allow | It cannot create new table-owned data. | | `Log`, `TinyLog`, `StripeLog`, `Set`, `Join`, `EmbeddedRocksDB`, and database `File` | Reject | They own unreplicated on-disk table data. | | `MaterializedPostgreSQL` | Reject | It owns a local nested table. | | `Distributed`, `Remote`, and `RemoteSecure` | Allow | Their optional local queue is auxiliary delivery state rather than data of the table itself. | | `Memory`, `Buffer`, and `Null` | Allow | They do not own on-disk table data. | | `Merge`, `Alias`, and `View` | Allow | They are metadata-only. | | Views with inner tables | Depends | Each generated inner table is created and checked separately. | | External-storage engines, data lakes, external databases, and external queues | Allow | Their data is managed outside ClickHouse-owned table storage. | | Lazy `StorageTableProxy` | Reject conservatively | The nested storage is unknown without loading it. | `ATTACH` remains outside the existing gate. The `Distributed` background `INSERT` queue is local to the node accepting an insert and is not replicated. Users requiring acknowledgement only after data reaches the destination shards should set `distributed_foreground_insert = 1`. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Allow `Distributed`, `Remote`, and `RemoteSecure` tables in `Replicated` databases when `database_replicated_allow_only_replicated_engine` is enabled, while continuing to reject writable non-replicated `MergeTree` tables using local or remote storage policies.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108820",
          "createdAt": "2026-06-29T15:29:14Z",
          "updatedAt": "2026-08-13T16:32:56Z",
          "timestamp": "2026-08-13T16:32:56Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "UberDever",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:48c1372d5cb9389569f4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114626",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114626",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not merge-sort a distributed gather whose sort description is all-constant",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 Related: https://github.com/ClickHouse/ClickHouse/issues/106237 --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description With `make_distributed_plan = 1`, a window `PARTITION BY <constant>` builds a `GatherSend` fragment whose sort description is entirely constant. `GatherSendStep::updatePipeline` adds an order-preserving `MergingSortedTransform` for any non-empty description, and that merge waits for *every* input stream to have data (`IMergingTransformBase::prepareInitializeInputs`). Hashing a constant key sends all rows to one bucket, so the other branches are never fed and stay `NeedData` while the loaded branch is `PortFull`. Neither side can move: the query deadlocks at zero CPU until `receive_timeout` and then raises `Pipeline stuck` (a logical error, so the server aborts in debug and sanitizer builds). An all-constant description orders nothing, so this change takes the `pipeline.resize(1)` branch that already exists in that function. A `ResizeProcessor` pairs any waiting output with any ready input and has no all-inputs barrier, which is why the code before [4a1ab1e](https://github.com/ClickHouse/ClickHouse/commit/4a1ab1e3d94fb0d) - which replaced an unconditional `resize(1)` with this conditional merge - could not wedge this way. Every row compares equal under an all-constant description, so any interleaving is validly sorted and the contract `GatherReceiveStep` relies on still holds. A description with at least one real column keeps the merge. Related: https://github.com/ClickHouse/ClickHouse/pull/113558 (where this failure was reported) Related: https://github.com/ClickHouse/ClickHouse/issues/106237 (context only, different mechanism) Found by `AST fuzzer (amd_debug, targeted, old_compatibility)`: [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113558&sha=b9fe1650232abf4fefc16992f3efcce44222bebc&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29). The `BufferedShardByHashTransform` deadlock tracked in #106237 is a different mechanism: that transform does not appear in the failing pipeline at all. The new test `04888` fails on master with `Pipeline stuck` for a constant key, a `LowCardinality` constant and a `Nullable` constant, and passes with this change. Its controls - a real column key, a mixed constant-plus-column key, and the non-distributed plan - pass both before and after, so the change is narrow. `04837_distributed_plan_window_partition_shuffle`, which pins the full distributed plan for a column-keyed window, still passes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114626",
          "createdAt": "2026-08-13T12:23:35Z",
          "updatedAt": "2026-08-13T16:32:28Z",
          "timestamp": "2026-08-13T16:32:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "davenger"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:36bff09f5b33a8ef4e4a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111494",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111494",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Text index: add trivial count optimization",
          "text": "Currently, the text index direct read optimization deserialize the sparse index, dictionary block and postings when there is token that exists in the index. Once the postings is read from disk, it fills the newly created boolean virtual column with postings data. With this optimization, we aim to reduce reading postings from disk and creating a virtual column. Instead we can answer queries using the token metadata from the dictionary block for specific query patterns as follows:. 1. `SELECT count() FROM table WHERE hasToken(column, 'foo');` 2. `SELECT count() FROM table WHERE hasAnyTokens(column, ['foo', 'bar']);` 3. `SELECT count() FROM table WHERE hasAllTokens(column, ['foo', 'bar']);` For the 1. case, we can avoid reading postings at all and use the cardinality metadata stored in the dictionary block to answer the query. For 2. and 3. cases, we would still read the postings but can avoid creating a virtual column. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Returns `COUNT()` queries directly from the text index cardinality metadata.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111494",
          "createdAt": "2026-07-22T22:33:56Z",
          "updatedAt": "2026-08-13T16:31:42Z",
          "timestamp": "2026-08-13T16:31:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-performance"
          ],
          "author": "ahmadov",
          "state": "open",
          "assignees": [
            "Ergus",
            "CurtizJ",
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bd3624e20d698b0c84bf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109367",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109367",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Scheduler: reclaimable memory tracking and dynamic spilling",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/109064 --> Scheduler-side support for reclaimable memory tracking and dynamic spilling, built on top of the memory reservations subsystem. The design is described in #109064. Queries can report the portion of an allocation that can be spilled or discarded on request (`IAllocationQueue::setReclaimable`), which is aggregated bottom-up as a new `reclaimable` field on every scheduler node. Each `AllocationLimit` gains a soft limit: when a workload's allocated memory exceeds it and the subtree has reclaimable memory, the scheduler asks a victim to reclaim memory (`ResourceAllocation::spillAllocation`, mirroring `killAllocation`) instead of waiting for the hard `max_memory` limit to force a kill. The request is replied: the query finishes it with `IAllocationQueue::finishSpill` after issuing the decreases for the freed memory (zero reclaimable doubling as a decline), the reply travels to the root as `Update::spilled` on the same propagation path as the state it describes, and at most one spill request is outstanding per subtree until it arrives; a victim that leaves the queue counts as having replied. Victim selection is deterministic and matches the kill order (largest usage, least precedence, largest allocation), descends a single root-to-leaf path via reclaimable-filtered ordered sets, and is fail-close: with nothing reclaimable, or with no soft limit configured, behavior is exactly as before. The soft limit is configured per workload via two new settings: `max_memory_before_spill` (absolute) and `max_memory_to_spill_ratio` (a fraction of the workload's own `max_memory`), smaller wins. The new state is observable in `system.scheduler`: `reclaimable`, `spills`, and the effective `soft_limit`. This is the scheduler side only. `MemoryReservation::spillAllocation` is currently a no-op; the query side that reports reclaimable memory and reacts to spill signals is a separate change. Documentation for the scheduler-side workload settings (`max_memory_before_spill`, `max_memory_to_spill_ratio`) and a `Spilling reclaimable memory` section are included here; the set of operators that can spill will be documented together with the query-side reaction. Invariants for the new machinery are documented on `ISpaceSharedNode` (I1-I8). Added unit tests cover fail-close, largest-first selection, fair descent skipping unreclaimable subtrees, the gate held until the victim's reply, the reply deferred to a pending decrease, re-signalling until under the soft limit, clamping reclaimable on a shrink, a declining victim, a victim removed mid-spill, the soft-limit enable transition via `CREATE OR REPLACE WORKLOAD`, the settings (absolute, ratio, and smaller-wins combination), a concurrency stress, and a parametrized throughput test. Verified under ThreadSanitizer (145 scheduler/workload gtests, 0 data races). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added workload settings `max_memory_before_spill` and `max_memory_to_spill_ratio` for the experimental memory reservation scheduling. They configure a soft memory limit above which a workload's queries are asked to spill reclaimable memory, before the hard `max_memory` limit forces an eviction. This is the scheduler-side foundation: queries do not yet report reclaimable memory or react to spill requests, so the settings have no effect until the query-side change lands. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109367",
          "createdAt": "2026-07-03T21:49:24Z",
          "updatedAt": "2026-08-13T16:29:58Z",
          "timestamp": "2026-08-13T16:29:58Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "serxa",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4d69871f20cae90be073",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:103182",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:103182",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Lightweight Updates v2",
          "text": "### Changelog category (leave one): - Backward Incompatible Change ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Lightweight `UPDATE` patch parts now use a new v2 on-disk format sorted by (`sorting_key..., _block_number, _block_offset`) and applied with a new merging algorithm. Peak memory is bounded by the largest equal-sort-key run instead of the full patch, and updates that cross merge boundaries no longer fall back to in-memory Join apply. Old-format patch parts remain readable. During a rolling upgrade from a version before 26.8, keep `patch_parts_version = 'v1'` or use the `compatibility` setting until all replicas are upgraded. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/103182",
          "createdAt": "2026-04-20T17:45:11Z",
          "updatedAt": "2026-08-13T16:29:39Z",
          "timestamp": "2026-08-13T16:29:39Z",
          "metrics": {
            "reactions": 2,
            "comments": 8
          },
          "labels": [
            "pr-backward-incompatible"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:0bc81f65791dd5149a5a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113573",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: add `NOSIGN` to public `s3` documentation queries",
          "text": "Update anonymous public S3 documentation examples to pass `NOSIGN` explicitly following the ClickHouse 26.7 server credential behavior change. This prevents public reads from attempting to use server-managed credentials. The audit also found and repairs two stale public paths: the LAION guide now uses the surviving 10-million-row shard, and the S3 brace-expansion example references the four files that currently exist. Authenticated, requester-pays, write, placeholder, and Foursquare examples are intentionally outside this PR. Related: https://linear.app/clickhouse/issue/DOC-945/update-public-s3-docs-examples-to-use-nosign ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Updated public S3 documentation examples to use `NOSIGN` and repaired stale public dataset paths.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113573",
          "createdAt": "2026-08-05T20:49:36Z",
          "updatedAt": "2026-08-13T16:23:29Z",
          "timestamp": "2026-08-13T16:23:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-documentation",
            "pr-autogenerated-docs"
          ],
          "author": "dhtclk",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a7593d82ae8e89f72f8b",
        "signalId": "github:ClickHouse/ClickHouse:issue:113741",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:issue:113741",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "issue",
          "title": "Custom-key parallel replicas over a Merge table with a Distributed child: children are offloaded to a finalized stage (CANNOT_CONVERT_TYPE, logical error in GroupingAggregatedTransform)",
          "text": "🕵 Reading a `Merge` table (or the `merge` table function) with custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_sampling'` or `'custom_key_range'`) fails with an exception when the common processing stage of the children is `WithMergeableState` — for example, when one of the underlying tables is a `Distributed` table. ```sql CREATE TABLE t_mrg_ck_1 (k UInt64) ENGINE = MergeTree ORDER BY k AS SELECT number FROM numbers(100000); CREATE TABLE t_mrg_ck_2 (k UInt64) ENGINE = MergeTree ORDER BY k AS SELECT number FROM numbers(100000); CREATE TABLE t_mrg_ck_3 (k UInt64) ENGINE = Distributed('test_shard_localhost', currentDatabase(), 't_mrg_ck_1'); SET enable_parallel_replicas = 1, max_parallel_replicas = 3, cluster_for_parallel_replicas = 'test_cluster_one_shard_three_replicas_localhost', parallel_replicas_for_non_replicated_merge_tree = 1, parallel_replicas_mode = 'custom_key_sampling', parallel_replicas_custom_key = 'k'; SELECT count() FROM merge(currentDatabase(), '^t_mrg_ck_'); ``` ``` Code: 70. DB::Exception: Conversion from UInt64 to AggregateFunction(count) is not supported: while converting source column `count()` to destination column `count()`: Child table: default.t_mrg_ck_1. (CANNOT_CONVERT_TYPE) ``` Depending on the aggregate function, it instead trips an assertion in the pipeline — an exception in debug and sanitizer builds (found by the AST fuzzer, STID `3970-479a`): ```sql SELECT 47, quantileExactInclusive(visibleWidth(['1', '2'])) IGNORE NULLS FROM merge(currentDatabase(), '^t_mrg_ck_') GROUP BY ALL LIMIT 973; ``` ``` Logical error: 'Chunk should have AggregatedChunkInfo/ChunkInfoWithAllocatedBytes in GroupingAggregatedTransform.'. ``` **Root cause.** The `Distributed` child reports `WithMergeableState` from its `getQueryProcessingStage`, so `StorageMerge::getQueryProcessingStage` sets the common stage of all children to `WithMergeableState`, and `ReadFromMerge::createPlanForTable` plans each `MergeTree` child through an interpreter with `SelectQueryOptions(WithMergeableState)`. Inside that child interpreter, the custom-key branch of `PlannerJoinTree` (`src/Planner/PlannerJoinTree.cpp`, the `canUseParallelReplicasCustomKey` block) offloads the child query to the replicas at the hard-coded stage `WithMergeableStateAfterAggregationAndLimit`, ignoring the requested `to_stage`. The child plan therefore produces finalized rows (post-aggregation, post-`LIMIT`) where the parent `ReadFromMerge` pipeline expects partial aggregation states: `convertAndFilterSourceStream` throws `CANNOT_CONVERT_TYPE` when the finalized type differs from the state type, and when the types coincide structurally, the chunks without `AggregatedChunkInfo` reach `GroupingAggregatedTransform` and trip the assertion. Only the analyzer path is affected (`enable_analyzer = 0` returns correct results). Reproduced on current master (verified on a binary with no unrelated changes); found by the targeted AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/110972 (which is unrelated: the failure reproduces without `parallel_replicas_allow_merge_tables`, a setting that does not exist on master), report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110972&sha=ac0a584ce316ace31f4dbc16b38e1262a2344751&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29 The custom-key offload should be skipped when the requested `to_stage` is below `WithMergeableStateAfterAggregationAndLimit`: a plan that must stop at a partial stage cannot accept a finalized remote read. Related: https://github.com/ClickHouse/ClickHouse/pull/110972 <!-- ch-version-info:start --> ### Version info - Resolved by: #113742 - Backported to: `26.7.4.29` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/issues/113741",
          "createdAt": "2026-08-06T23:14:04Z",
          "updatedAt": "2026-08-13T16:23:00Z",
          "timestamp": "2026-08-13T16:23:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "comp-distributed",
            "potential bug",
            "comp-parallel-replicas"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:41d050b42544e0c51e4c",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113742",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113742",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "<!--- A technical comment, you are free to remove or leave it as it is when PR is created The following categories are used in the next scripts, update them accordingly utils/changelog/changelog.py tests/ci/cancel_and_rerun_workflow_lambda/app.py --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `CANNOT_CONVERT_TYPE` and an exception in `GroupingAggregatedTransform` when reading a `Merge` table with one child being a `Distributed` table under custom-key parallel replicas (`parallel_replicas_mode = 'custom_key_sampling'` or `'custom_key_range'`). Closes [#113741](https://github.com/ClickHouse/ClickHouse/issues/113741). ### Documentation entry for user-facing changes The custom-key parallel replicas branch of the planner replaces the plan of a table expression with a remote read at the fixed stage `WithMergeableStateAfterAggregationAndLimit`, ignoring the stage the plan was requested up to. A `Merge` table over a `Distributed` child plans all of its children up to `WithMergeableState` through an interpreter, so a `MergeTree` child's plan produced finalized (post-aggregation, post-`LIMIT`) data where the parent `ReadFromMerge` expected partial aggregation states: `CANNOT_CONVERT_TYPE` for `count`, and the exception `Chunk should have AggregatedChunkInfo/ChunkInfoWithAllocatedBytes in GroupingAggregatedTransform` when the finalized type structurally coincides with the state type. Only the analyzer path is affected. The fix allows the custom-key read only when the requested stage is `Complete` or `WithMergeableStateAfterAggregationAndLimit` itself; a child planned to a partial stage now runs as a plain local read, as it does when parallel replicas are off. Found by the targeted AST fuzzer on https://github.com/ClickHouse/ClickHouse/pull/110972, where it is unrelated: the failure reproduces on master without `parallel_replicas_allow_merge_tables` (verified on a binary with that PR's changes swapped out to the merge base). Fuzzer report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110972&sha=ac0a584ce316ace31f4dbc16b38e1262a2344751&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%2C%20old_compatibility%29 (STID `3970-479a`). Closes: https://github.com/ClickHouse/ClickHouse/issues/113741 Related: https://github.com/ClickHouse/ClickHouse/pull/110972 <!-- ch-version-info:start --> ### Version info - Backported to: `26.7.4.29` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113742",
          "createdAt": "2026-08-06T23:22:57Z",
          "updatedAt": "2026-08-13T16:22:58Z",
          "timestamp": "2026-08-13T16:22:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e347bdc272fdcc464b8d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112498",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112498",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md): Fixed a crash when reading a Parquet file with inconsistent bloom filter metadata. Such files could also silently return fewer rows than they should. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ## Problem Reading a Parquet file whose bloom filter metadata is inconsistent segfaults the server. Seen in production: 25 crashes over 12 days on one instance, all with the same stack, while an hourly job read an Iceberg lake on S3. ``` (version 26.2.1.525 (official build), architecture: aarch64) Received signal 11 Signal description: Segmentation fault Address: 0xfffab740826f. Access: <not available>. Address not mapped to object. 2.0. inlined from base/base/../base/unaligned.h:12: unsigned int unalignedLoad<unsigned int>(void const*) 2. src/Processors/Formats/Impl/Parquet/Reader.cpp:906: DB::Parquet::Reader::BloomFilterLookup::findAnyHash(...) 3. src/Storages/MergeTree/KeyCondition.cpp:780: DB::mayExistOnBloomFilter(...) 4. src/Storages/MergeTree/KeyCondition.cpp:4127: DB::KeyCondition::checkInHyperrectangle(...) 5. src/Processors/Formats/Impl/Parquet/Reader.cpp:935: DB::Parquet::Reader::applyBloomAndDictionaryFilters(RowGroup&) 6. src/Processors/Formats/Impl/Parquet/ReadManager.cpp:117: DB::Parquet::ReadManager::finishRowGroupStage(...) ``` Consequences: - The server process dies, so every query on that instance fails, not just the Parquet one. The affected instance was single-replica, so each crash was a full outage until the pod restarted. - The crashing query retries and crashes again, once per retry. - When the out-of-range bloom filter block happens to land inside the read buffer instead of unmapped memory, there is no crash: the row group is pruned on unrelated bytes and rows go missing silently. `SELECT count() FROM file(...) WHERE s = '123456'` returned `0` for a value that is present. - Reproducible on `master`, and on every version since 25.11, when the v3 Parquet reader became the default (`input_format_parquet_bloom_filter_push_down` has defaulted to `1` since 25.5). ## Root cause `Reader::processBloomFilterHeader` takes the bloom filter bitset size from the file (`BloomFilterHeader.numBytes`, validated only for sign and 32-byte alignment) and derives the byte range of each 32-byte bloom filter block from it, at `bloom_filter_offset + header_size + block_idx * 32`. It never checks that `header_size + numBytes` fits inside the byte range it registered for the bloom filter — `ColumnMetaData.bloom_filter_length` when the file declares it, otherwise the \"next known offset\" upper bound computed in `initializePrefetches`. `Prefetcher::splitRange` was the only guard, and it missed the case in two independent ways: - `if (start < range.start || length > range.end - start)` **underflows**: when `start > range.end`, `range.end - start` wraps around, so the comparison is false and a subrange past the end of the range passes. - The check runs only while the parent range is still in state `HasRange`. The bloom filter header range (registered with `likely_to_be_used = true`) and the bloom filter data range start at the same file offset, and the data range is normally smaller than `min_bytes_for_seek`, so starting the header prefetch coalesces the data range into the same read task and flips it to `HasTask`. `splitRange`'s tail path then computes `req->task_offset = subranges[i].first - task->offset` with no validation at all. This is the path a real S3 read takes. `Prefetcher::getRangeData` guarded the resulting span with `chassert` only, which is compiled out in release builds, so it returned a `std::span` pointing outside `task->buf`, and `findAnyHash`'s `unalignedLoad<UInt32>` read unmapped memory. ## Fix Each of the three layers now fails closed: 1. `Reader::processBloomFilterHeader` rejects a header whose `header_size + numBytes` exceeds the bloom filter extent the file declared, with `INCORRECT_DATA` naming `input_format_parquet_bloom_filter_push_down=0` as the escape hatch — the same shape as the two bloom filter validation errors already there. The extent is remembered in the new `ColumnChunk::bloom_filter_data_bytes`, set at both `registerRange` call sites. 2. `Prefetcher::splitRange` makes the `HasRange` check underflow-safe and applies the equivalent check against the read task's byte range on the coalesced `HasTask` path, before touching refcount or `RequestState`s so throwing stays clean. 3. `Prefetcher::getRangeData` turns the buffer-bounds `chassert` into a real check against `task->length` (the invariant that holds for both the `buf` and the zero-copy `cached_region` paths), so a bookkeeping mistake anywhere surfaces as an error instead of an out-of-bounds read. Regression test: `04654_parquet_bloom_filter_bitset_out_of_bounds` reads a 1649-byte fixture whose `s` column `BloomFilterHeader` claims a 1 GiB bitset while its column metadata declares 272 bytes of bloom filter data; the read must report `INCORRECT_DATA`, and the same file still reads correctly with push-down off. The fixture is small on purpose, so its bloom filter stays under the default seek threshold and the test drives the coalesced path that production takes. Each of the three layers was also removed on its own and the test re-run, confirming none of them is dead code: without layer 1 the request is rejected by layer 2 (`Subrange out of bounds: [460964200, 460964232) not in read task [395, 1352)`), without layers 1 and 2 by layer 3, and with all three removed the reader aborts on the buffer-bounds assertion in a Debug build. Note that a file whose `bloom_filter_length` excludes the serialized header (`parquet.thrift` specifies that it includes it) is now rejected rather than read. Those reads were already either crashing or silently over-pruning; `input_format_parquet_bloom_filter_push_down=0` reads such files without their bloom filters. <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1303` (included in `26.8` and later) - Backported to: `26.7.4.25`, `26.6.3.35` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112498",
          "createdAt": "2026-07-30T00:31:49Z",
          "updatedAt": "2026-08-13T16:22:14Z",
          "timestamp": "2026-08-13T16:22:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested",
            "pr-synced-to-cloud",
            "pr-must-backport-synced"
          ],
          "author": "tiandiwonder",
          "state": "closed",
          "assignees": [
            "Algunenano"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:6aa7132194e1b9872679",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114635",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114635",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113742 to 26.7: Skip the custom-key parallel replicas read when the requested stage cannot absorb finalized data",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113742 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31700181405/job/94447176388) <!-- ch-version-info:start --> ### Version info - Merged into: `26.7.4.29` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114635",
          "createdAt": "2026-08-13T12:50:24Z",
          "updatedAt": "2026-08-13T16:21:35Z",
          "timestamp": "2026-08-13T16:21:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll4",
          "state": "closed",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b11091f8e5f3aadc92ce",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114575",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114575",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #112498 to 26.6: Fix segfault reading a Parquet file with an inconsistent bloom filter size",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/112498 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31660232512/job/94323336840) <!-- ch-version-info:start --> ### Version info - Merged into: `26.6.3.35` <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114575",
          "createdAt": "2026-08-13T02:35:48Z",
          "updatedAt": "2026-08-13T16:21:33Z",
          "timestamp": "2026-08-13T16:21:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll2",
          "state": "closed",
          "assignees": [
            "Algunenano",
            "tiandiwonder"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:deaefcba98fc48e1e809",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114665",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114665",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow overriding the HTTP method for SELECT through the url table function and URL engine",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/62352 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `http_method='POST'` — as a key-value argument of the `url` table function and the `URL` table engine, or through the `http_method`/`method` named collection keys — now applies to `SELECT` queries: reads use `POST` instead of the default `GET`, for servers that accept only `POST`. Schema inference follows the configured method. `PUT` keeps its write-only meaning: a `SELECT` through a configuration with `http_method='PUT'` still uses `GET`. ### Implementation notes - `IStorageURLBase::getReadMethod()` now returns `POST` when the configured `http_method` is `POST`. `PUT` still applies to writes only (pre-signed upload URLs, #44326), and `INSERT` behavior is unchanged: `POST` by default, `PUT` when configured. - The write path no longer mutates the shared `http_method` member when defaulting to `POST` — the storage instance is shared between queries, so persisting the write default would have flipped subsequent reads of the same table from `GET` to `POST` (and raced concurrent reads). - Schema inference (`getTableStructureAndFormatFromData` and the URL read-buffer iterator) uses the same effective read method instead of hardcoded `GET`, for both `url()`/`URL` and `urlCluster`. - The `http_method = '...'` (or `method = '...'`) key-value argument is parsed by `StorageURL::evalArgsAndCollectHeaders` alongside `headers(...)`, kept in the `CREATE` AST (so it survives `SHOW CREATE TABLE` and DETACH/ATTACH), validated to be `POST`/`PUT`, and rejected when the URL scheme dispatches to another backend (`file://`, `s3://`, ...), mirroring `headers(...)`. - In the analyzer, the argument's left-hand identifier is excluded from column resolution (`skipAnalysisForArguments`), like `headers(...)`. - Documentation for the `url` table function and the `URL` engine is updated. The new stateless test `04869_url_select_http_method` asserts the method actually used on the wire via `system.query_log.http_method`: default `GET`, `POST` override for both the data and the schema-inference requests, `PUT`-configured `SELECT` staying on `GET`, the `URL` engine path, rejection of unsupported methods, and persistence of the argument in the table DDL. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114665",
          "createdAt": "2026-08-13T16:17:38Z",
          "updatedAt": "2026-08-13T16:19:49Z",
          "timestamp": "2026-08-13T16:19:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "valerypetrov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f2193b79b7b3376d136b",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112327",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112327",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix use-after-free on a sparse join key in a direct dictionary join",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/109225 --> Related: https://github.com/ClickHouse/ClickHouse/pull/109225 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a use-after-free when a `JOIN` with `join_algorithm = 'direct'` onto a dictionary was given a join key that is stored with sparse serialization. `getColumnVectorData` returned a reference to a temporary column, so the dictionary lookup read freed memory: release builds could return wrong results and debug or sanitizer builds aborted. ### Description `getColumnVectorData` (`src/Dictionaries/DictionaryHelpers.h`) materializes its key column into a function-local `ColumnPtr` and then returns a `PaddedPODArray` reference **into that local**. It copied the data into the caller's `backup_storage` only when the input was `Const`. For a dense column that was still safe, because every conversion is a no-op returning `getPtr()` and the caller's own `ColumnPtr` keeps the buffer alive. It is not safe for a column that has to be materialized: `ColumnSparse::convertToFullColumnIfSparse` and `ColumnReplicated::convertToFullColumnIfReplicated` each allocate a **new** column that nothing else owns, so the returned reference dangles as soon as the function returns. The fix takes the copy whenever a conversion actually produced a different column (`full_column.get() != column.get()`). This is pointer identity rather than a type test, so it covers `Const`, `Sparse` and `ColumnReplicated` with one predicate, and it is fail-closed for any representation added later. The previous `Const` check is strictly subsumed: `ColumnConst::convertToFullColumn` returns either the inner column or a `replicate` result, never the `ColumnConst` itself, so `Const` behaviour is unchanged. The dense path is unaffected and adds no copy, which I verified by instrumenting both live call sites: a dense key reports zero copies, a sparse key reports one at each site. A sparse key is the case that is a use-after-free today, and it already paid for a full materialization inside `removeSpecialRepresentations`, so the extra `memcpy` of that same buffer is negligible. Reaching the bug requires a path that hands a non-materialized key to the dictionary. `dictGet`, `dictHas` and the hierarchy functions cannot: `IFunction::useDefaultImplementationForSparseColumns()` and `...ForReplicatedColumns()` both default to true and no dictionary function overrides them, so a dense, caller-owned column arrives. `IDictionary::getByKeys` (the direct join) is the reaching path, because its own `removeSpecialRepresentations` call sits inside a Nullable-only branch and a non-Nullable sparse key passes through untouched. That is also why `04627_direct_join_dictionary_nullable_key`, which does exercise a sparse key, never caught this: its key is Nullable, so it gets materialized. All 14 `getColumnVectorData` call sites are fixed by this single change. Two of them are reachable today, both in `FlatDictionary` and both on the same `getByKeys` call: `hasKeys` (the site in the reports below) and `getColumn`, reached through `getColumns`. The remaining 12 are hierarchy-only and reachable solely through the pre-converting function path. `Hashed` and `HashedArray` never reach the helper on the `getByKeys` path at all, because `DictionaryKeysExtractor` holds its converted column by value and therefore owns it. I also swept every other `convertToFullColumnIf*` / `recursiveRemove*` / `removeSpecialRepresentations` call site under `src/` for the same \"derived data escapes the owning local\" shape and found no second instance, so no sibling fix is needed. The bug dates to 2021 (`b5b624f3d7e9bf`, which introduced the conversion here) and was widened in 2025 by `2b6cb36d1dc936`, which added `ColumnReplicated` as a second carrier. The same code is present on 26.7, 26.6, 26.5, 26.4 and 26.3. Found while triaging CI on #109225 and reproduced on unmodified master. It is latent in CI only because no existing test combined a sparse-serialized left key with a direct join over a `FLAT()` dictionary; CI randomizes `ratio_of_defaults_for_sparse_serialization`, so any test that does hit this combination fails roughly 40% of the time. Reports on `31b4a2961ef4c5183f7d15dda7f755541a77a98c`: - `AddressSanitizer: heap-use-after-free`, allocated by `ColumnSparse::convertToFullColumnIfSparse`, freed at the end of `getColumnVectorData`, read by `FlatDictionary::hasKeys`: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_asan_ubsan%2C%20flaky%20check%29 - `Logical error: '(n >= (static_cast<ssize_t>(pad_left_) ? -1 : 0)) && (n <= static_cast<ssize_t>(this->size()))'` from the same read: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=109225&sha=31b4a2961ef4c5183f7d15dda7f755541a77a98c&name_0=PR&name_1=Stateless%20tests%20%28amd_debug%2C%20flaky%20check%29 The new test `04652_direct_join_dictionary_sparse_key` covers both live call sites, the aggregation shape from the report, and a key carried through `ARRAY JOIN` over a sparse base column, which is a third shape where the key has to be materialized. It pins its results against a dense table and against `join_algorithm = 'hash'` instead of hand-written constants, and asserts both that the key really is sparse and that `DirectKeyValueJoin` is still chosen, so it cannot pass vacuously. On master it aborts; with the fix it passes 50/50 with and without randomized settings. Reverting only the new predicate makes it abort again. A follow-up cleanup worth doing separately: the helper carries a `/// TODO: Remove` and would be better returning the owning `ColumnPtr` alongside the data, which removes the need for `backup_storage` entirely. That touches all 14 call sites and four dictionary classes, so it does not belong here.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112327",
          "createdAt": "2026-07-28T17:38:47Z",
          "updatedAt": "2026-08-13T16:39:38Z",
          "timestamp": "2026-08-13T16:39:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "alexbakharew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:fa1c0174affbd72fcea5",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114247",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114247",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Clear plain LIMIT/OFFSET in the window-view backfill source query",
          "text": "Follow-up to https://github.com/ClickHouse/ClickHouse/pull/113759: the last AI review finding on that PR landed after the PR had been added to the merge queue, so the branch could no longer be updated. This addresses it. `StorageWindowView::getSourceTableSelectQuery` builds the raw-source backfill query for `CREATE WINDOW VIEW ... POPULATE`. The rows it produces are inserted into the window view, where `writeIntoWindowView` executes the mergeable view query over them, so the backfill query must deliver the raw source rows and leave all row transformations of the original query to the view query — otherwise the initialized state diverges from live behavior. The PR originally made the helper clear the leftovers of the original `SELECT` that violated this invariant one by one — `LIMIT`/`OFFSET` (and `WITH TIES`), plain `DISTINCT`, `ARRAY JOIN`, `WHERE`/`PREWHERE`, table-expression `SAMPLE`/`FINAL` — on top of the `JOIN`/`GROUP BY`/`ORDER BY`/`LIMIT BY`/`WINDOW`/`QUALIFY`/`INTERPOLATE` handling it already had. Review then found that this strip-the-leftovers approach misses wrapped sources: the same constructs inside a `FROM (SELECT ...)` subquery or a CTE definition survived the rewrite, and covering them would have required recursing the rewrite into every nested select. So the helper now builds the backfill query from scratch instead of stripping a clone of the view query. The contract makes this valid: `writeIntoWindowView` always receives raw source-table blocks (`getInputHeader` is the source table header no matter how the view query wraps or transforms the table — `PushingToWindowViewSink` is created with exactly that header), so the correct backfill query is exactly `SELECT <source columns> FROM <source table>`, plus `ORDER BY` on the timestamp column so the watermark is initialized from the earliest record. Wrapped sources, joins, and every row-shaping clause are covered by construction because the user query is no longer cloned at all. The helper shrinks by ~100 lines. Note: the divergence is currently unobservable because `CREATE WINDOW VIEW ... POPULATE` fails before writing any rows (https://github.com/ClickHouse/ClickHouse/issues/113493), so no regression test is possible yet; this keeps the rewrite invariant consistent for when `POPULATE` is fixed. Verified with `clickhouse-local` probes that `POPULATE` over wrapped-subquery/CTE/`JOIN`/`WHERE`/`FINAL` sources now analyzes the backfill query cleanly and proceeds to the pre-existing #113493 sink error. Related: https://github.com/ClickHouse/ClickHouse/pull/113759 Related: https://github.com/ClickHouse/ClickHouse/issues/113493 ### Changelog category (leave one): - Not for changelog (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114247",
          "createdAt": "2026-08-10T23:40:02Z",
          "updatedAt": "2026-08-13T17:46:37Z",
          "timestamp": "2026-08-13T17:46:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:75bef5a15d565f91fbae",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:110105",
        "event": "discovered",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:110105",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add clickhouse-proxy application mode",
          "text": "Implements a new application mode in the `clickhouse` binary, named `proxy`. The proxy accepts connections over end-user ClickHouse protocols, finds the upstream backend based on configurable rules (hostname from TLS SNI or HTTP header, user name, database name, and — for HTTP — query type), and forwards the traffic to it. It is built on the `silk` fiber framework so that many concurrent connections are handled with low RAM usage. Start it with `clickhouse proxy --config-file proxy_config.xml`; a fully commented example config is in `programs/proxy/proxy_config.xml`, and the feature is documented at `docs/en/operations/clickhouse-proxy.md`. **What it does** - Protocol frontends for HTTP(S), native TCP, MySQL, PostgreSQL, transparent TLS-by-SNI, and opaque TCP streams. It handles only end-user protocols (not Keeper or inter-server replication). The user name and database are parsed from the first packets where the protocol allows it (HTTP headers/params/Basic auth, native `Hello`, PostgreSQL `StartupMessage`, HTTP query type). MySQL is server-speaks-first with in-band TLS, so it is forwarded transparently and routed by peer address or the default pool. - TLS routing options: terminate and re-encrypt (the proxy and each backend hold their own certificates), terminate only (unwrap: TLS to clients, plaintext to backends), and transparent (route by SNI without decrypting). Optional ACME certificate provisioning, as in `clickhouse-server`. - Multi-criteria routing rules matching host / user / database / query type / protocol, by exact value or regular expression; regexp captures can be substituted into a backend address template (e.g. route users `ch-<tenant>` to per-tenant backends). - Pools with pluggable load balancing (`random`, `round_robin`, `least_connections`, `lowest_latency`, `least_resources`) behind a common interface; a pool may list several backends for load balancing. - Session stickiness by `session_id` (from the HTTP URL) or by peer address, using a consistent (rendezvous) hash of the backends. - Backend health monitoring (connect latency, consecutive failures) and optional CPU/memory polling with per-backend credentials, feeding the `least_resources` strategy; a JSON status endpoint, `/ping`, and static pages are served by the proxy itself. - An abstract routing table exposing hooks (unknown route, no backends available, first time a user or database is seen) that run a shell command — for example to provision a backend on demand and wait for it to become available. Validated end to end against a live server: HTTP `/ping`, static pages, the status endpoint, GET/POST forwarding, the no-backend error paths, native protocol round-trips, and clean shutdown. ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added a new `proxy` mode to the `clickhouse` binary: a lightweight, protocol-aware proxy that routes end-user connections (HTTP, native, MySQL, PostgreSQL, TLS-by-SNI, and raw TCP) to backend servers based on a configurable routing table, with TLS termination/passthrough, load balancing, health checks, and session stickiness. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/110105",
          "createdAt": "2026-07-11T17:34:42Z",
          "updatedAt": "2026-08-13T17:46:26Z",
          "timestamp": "2026-08-13T17:46:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-feature",
            "submodule changed"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:e6521752a509c1fae00a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86353",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cascades cost-based optimizer for distributed query plans",
          "text": "A Cascades-style cost-based optimizer that chooses distribution strategies for the multi-stage distributed query plans of #106020. It explores alternatives in a memo (a shared store of equivalent plan fragments) with top-down, goal-directed search and picks the cheapest plan satisfying the required distribution and sorting properties, inserting exchange operators (plan steps that move rows between nodes) as needed. Implemented: - **Join strategies**: shuffle hash join, broadcast hash join (with `ReplicatedRead` — every worker repeats the same read of a small table instead of a network broadcast, assuming shared storage where all workers see the same data), replicated join (a small deterministic join is recomputed identically on every node over such reads, so its result never crosses the network; nested joins compose; `ANY` joins are excluded because the kept row depends on the build order), local join. - **Aggregation strategies**: two-phase (partial + merge), shuffle by group keys, local; `distributed_aggregation_memory_efficient` and `distributed_plan_force_shuffle_aggregation` are honored. - **Top-N**: two-stage distributed top-N (per-node bounded sort, sorted-merge gather, coordinator limit); disabled under `exact_rows_before_limit`, which needs the full row count. - **Read strategies**: parallel N-way read, replicated read, local read. For `FINAL`, #108148 (already in master) taught the rule-based distributed plan to split a `FINAL` read into disjoint primary-key-range buckets where that is safe; Cascades now reuses that machinery, so `FINAL` no longer forces a serial read here either. The coordinator ships each bucket's marks in the `read_bucket` task parameters. - **`IN (subquery)`**: follows the `rewrite_in_to_join` setting like the rest of the planner (the forced join form is removed). In the default set form the set-building subqueries are planned separately and distributed like any other query. - **Properties and enforcers**: distribution (node count, replication, partitioning columns with equivalence classes and the types the keys are cast to before hashing) and sorting; when a plan alternative lacks a required property, an enforcer inserts the step that provides it (`Gather`/`Shuffle`/`Broadcast`/`ScatterExchange`, `Sort`). - **Transformations**: join commutativity (only for semantics-preserving joins: `INNER ALL`, `CROSS`, `SEMI`/`ANY`/`ANTI`; never `ASOF`, and never `ANY` under `join_any_take_last_row`), two-phase aggregation split, two-stage top-N split. - **Cost model**: `work`, `network`, and `sequential` components, each priced as wall-clock per node: a shuffle moves 1/N of the data per node, a broadcast payload is ingested once by every receiver in parallel, and a gather funnels every row through one endpoint, so its transfer stays undivided and pays a per-row cost. A hash-table build counts as parallel work (`parallel_hash` shards it across threads); a fixed per-exchange overhead keeps small inputs local. A table read is priced on its scan volume - the rows the primary key keeps - not on its output estimate, so a filter off the sorting key cannot make a replicated re-read look free. Standalone filters (e.g. `HAVING`) are estimated from column NDVs with join-key equivalence classes; join estimates are clamped to join kind and strictness semantics; exchange costs use per-row byte widths measured from the parts' column sizes (followed through renames, not derived from types). All weights and calibration constants are overridable at query time. What this improves over the rule-based distributed planner, on TPC-H plans. The rule-based planner broadcasts a small table when its read is below `distributed_plan_max_rows_to_broadcast`, but it often cannot size the result of a join, so a small join result (`nation x region`, 5 rows after the region filter) is scattered across nodes, joined there, and shuffled again (repartitioned across nodes) by the next join key. It also often inserts a shuffle at join and aggregation boundaries even when the rows are already divided by the right key. Cascades estimates sizes through joins, knows which partitioning already holds, compares broadcast against shuffle by cost for each join, and recomputes a small deterministic join on every node when that is cheaper than moving its result. On TPC-H SF100 over 8 nodes (same binary, same run window; times are server-side means of the hot runs) the join-heavy queries improve: | Query | Rule-based -> Cascades | What changed in the plan | |---|---|---| | Q21 | 7.66 s -> 5.34 s | The four `supplier x nation` joins compute their small results once and broadcast them, so `lineitem` is not shuffled to meet them. | | Q09 | 4.07 s -> 2.47 s | `part` and `nation` are read in full by every node, so `lineitem` and `supplier` are not shuffled to meet them. | | Q08 | 2.39 s -> 0.99 s | `nation x region` (5 rows) is recomputed by every node; `part` is read in full per node, so `lineitem` is not shuffled to join it. | | Q02 | 2.14 s -> 0.84 s | The `supplier x nation x region` chain is recomputed by every node, so only `partsupp` and `supplier` are shuffled. The top-100 sort becomes two-stage, sending at most 100 rows per node. | | Q05 | 2.05 s -> 1.12 s | The whole dimension side (`orders x customer x nation x region`) is recomputed by every node over full local reads; `lineitem` joins it in place with no shuffle at all. | | Q17 | 2.72 s -> 1.78 s | The small per-part average is broadcast to every node, so the outer 600M-row `lineitem` read is not shuffled. | | Q11 | 0.49 s -> 0.26 s | `supplier x nation` is computed once and broadcast; `partsupp` joins it in place. | | Q12 | 0.76 s -> 0.59 s | The filtered `lineitem` rows (~30K of 600M) are gathered once and broadcast; nothing is shuffled. | The remaining queries change less. Summed over all 22 queries, hot server time drops from 36.8 s to 28.8 s (about 22% lower). The remaining regressions are the `IN (subquery)` queries Q18 (2.43 s -> 2.85 s) and Q20 (2.24 s -> 2.99 s). With the forced join rewrite removed, both run their `IN`s as sets in both modes, and the sets built are identical; the difference is in the plans around them. Cascades picks replicated-read shapes that rescan moderate tables on every node, which loses here for a structural reason: the main reads are filtered by `x IN <set>`, and the set contents do not exist at costing time, so the model cannot credit a shuffle-based plan for how few rows survive the set filter, while the replicated read's full scan is paid regardless. Two follow-ups: a cost-informed choice of the `IN` form, and set-filter selectivity from the subquery's output estimate. `EXPLAIN pretty = 1, estimates = 1` shows the chosen plan with a row estimate and the accumulated cost for each step. Also in this PR, two improvements to the shared bucketed-read machinery (they benefit the rule-based path too): the `FINAL` layer split no longer depends on the coordinator's core count, and a many-partition `FINAL` split groups its layers into the target task count instead of falling back to a serial read. Design, a worked example on TPC-H data (a simplified 3-table query traced through the memo), and current limitations are documented in `src/Processors/QueryPlan/Optimizations/Cascades/ARCHITECTURE.md`. Plan-shape tests cover the actual TPC-H queries (`03836_tpch_join_order_plans`), and focused tests pin the cost-model contracts (e.g. `04869_cascades_read_cost_granule_volume`, `04838_cascades_filter_selectivity`). Disabled by default. Requires the analyzer; remote execution requires the stateless-worker configuration, while `distributed_plan_execute_locally = 1` runs the stages in-process without it: ```sql SET enable_cascades_optimizer = 1, make_distributed_plan = 1; ``` For tests, `param__internal_cascades_cluster_node_count` overrides the cluster size, `param__internal_cascades_cost_config` overrides the cost model configuration, `param__internal_join_table_stat_hints` injects table statistics, `param__internal_cascades_task_limit` lowers the task budget (it can never raise it). Related: https://github.com/ClickHouse/ClickHouse/pull/106020 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental Cascades cost-based optimizer for distributed query plans, enabled by `enable_cascades_optimizer = 1` together with `make_distributed_plan = 1`. It chooses between shuffle, broadcast, replicated, and local join strategies, two-phase, shuffle, and local aggregation, two-stage distributed top-N, and parallel and replicated reads by estimated cost, inserting exchange operators as needed. ### Documentation entry for user-facing changes - [ ] Documentation written in [/docs](https://github.com/ClickHouse/ClickHouse/tree/master/docs)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86353",
          "createdAt": "2025-08-28T11:27:09Z",
          "updatedAt": "2026-08-13T17:45:48Z",
          "timestamp": "2026-08-13T17:45:48Z",
          "metrics": {
            "reactions": 21,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9b13e7bc31e631920ca8",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107305",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107305",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Make ALTER MODIFY COLUMN on named Tuple metadata-only when adding subfields",
          "text": "`ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` that only adds new subfields is now metadata-only (no mutation). Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1` (default `false`). Subfield additions through `Array`/`Map`/nested `Tuple` wrappers are also handled. `Nullable(Tuple(...))` is blocked (null map incompatibility). Removing/renaming subfields or changing types still triggers a mutation. Key/index/projection guards reject the metadata-only path when `primary.idx` or skip-index bytes would become invalid (whole tuple in key, or subcolumn whose type changes). Subcolumn references with unchanged types (e.g. `ORDER BY t.a` when only `t.c` is added) are allowed. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `ALTER TABLE ... MODIFY COLUMN <col> Tuple(...)` on a named `Tuple` is now metadata-only when only adding subfields, matching the speed of top-level `ADD COLUMN`. Gated behind `SET allow_experimental_metadata_only_named_tuple_alter = 1`. ### Documentation entry for user-facing changes - [x] Documentation is not required (behavioral improvement; semantics unchanged)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107305",
          "createdAt": "2026-06-12T07:42:58Z",
          "updatedAt": "2026-08-13T17:45:47Z",
          "timestamp": "2026-08-13T17:45:47Z",
          "metrics": {
            "reactions": 1,
            "comments": 8
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "amosbird",
          "state": "open",
          "assignees": [
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:40b854b7c12c8d932c1e",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96487",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96487",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix LazilyReadFromMergeTree optimization with ALIAS columns (#96452)",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry: Fix missing lazy read and top-K read optimizations (skip-index top-K and dynamic top-K filtering) when selecting `ALIAS` columns with `WHERE ... ORDER BY ... LIMIT` on MergeTree tables. Closes: https://github.com/ClickHouse/ClickHouse/issues/96452 ### Documentation entry for user-facing changes N/A",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96487",
          "createdAt": "2026-02-09T21:21:08Z",
          "updatedAt": "2026-08-13T17:45:34Z",
          "timestamp": "2026-08-13T17:45:34Z",
          "metrics": {
            "reactions": 1,
            "comments": 31
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "jayvenn21",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:58e9689b35778ee9d490",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114150",
        "event": "discovered",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114150",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Disable the map element to subcolumn rewrite by default and revert the PREWHERE grouping change",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/111879 Related: https://github.com/ClickHouse/ClickHouse/pull/99200 Related: https://github.com/ClickHouse/ClickHouse/pull/107988 Related: https://github.com/ClickHouse/ClickHouse/pull/111954 Related: https://github.com/ClickHouse/ClickHouse/issues/107912 ### Description In 26.3 we added bucketed serialization for the `Map` data type together with an analyzer rewrite of `m['key']` (internally `arrayElement(m, 'key')`) into the Map key subcolumn `m.key_<key>`, so that only the relevant bucket is read from disk. The bucketed serialization turned out to be imperfect and is planned to be reimplemented as a separate type. For regular (non-bucketed) `Map` serialization this rewrite can be harmful (for example, it complicates subcolumn size estimation) and gives no benefit for single-key lookups. **1. New setting.** `optimize_map_element_to_subcolumn` (default `false`) gates only the `{Map, arrayElement}` transformer in `FunctionToSubcolumnsPass`. All other `Map` subcolumn optimizations (`mapKeys`, `mapValues`, `length`, `empty`) keep following `optimize_functions_to_subcolumns`. The new setting takes effect only when `optimize_functions_to_subcolumns` is also enabled. This only disables the automatic rewrite; reading the `m.key_<key>` subcolumn explicitly still works, and query results are unchanged either way. **2. Revert of the PREWHERE grouping from #107988.** That change made `MergeTreeWhereOptimizer` and `tryBuildPrewhereSteps` group conditions by physical storage column instead of by the exact column set, so that several `m['kN']` subcolumn reads share one read step. Its only purpose was to compensate for the rewrite: with the rewrite disabled, all `m['kN']` conditions reference the same column `m` and the original exact-column-set grouping already places them into one step. The grouping also introduced a correctness regression, #111879: a guard condition and a throwing expression over subcolumns of the same column (for example `payload.longitude IS NOT NULL` and a `CAST(tuple(...), 'Point')` over `payload.latitude`) became adjacent and were merged into one read step. All conditions of a step are evaluated on the same unfiltered block, so the throwing expression ran on rows the guard had rejected. Measured with the repro from #111879 on pre-built binaries: | binary | result | |---|---| | master before #107988 (`cd3c529cd3a`) | `66666` | | master with #107988 | `CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN` | | this branch | `66666` | Reverting instead of adding a new heuristic keeps this change small enough to backport. Improving subcolumn read sharing in PREWHERE will be done separately on `master` and does not need backporting; #111954 is superseded by this change. The serialization part of #107988 (`ISerialization` cache keys, whole-map caching in `SerializationMapKeyValue`, the `SerializationSparse` refactor) is kept — it is unrelated to #111879 and still shares reads between subcolumns of the same column inside one read step. **Expected performance comparison results.** By default `m['key']` again reads the whole `Map` column, as before 26.3, so query shapes that read map elements are slower than on `master`, where the rewrite is enabled: `get_map_value` and the shapes in `tests/performance/map_subcolumns_prewhere.xml` (the setting opt-in was removed from that test so it measures the default path). This is the intended consequence of disabling the rewrite; enabling `optimize_map_element_to_subcolumn` restores the previous behavior. **Out of scope.** With `enable_multiple_prewhere_read_steps = 0` the whole PREWHERE is a single step, so no guard can protect a throwing expression and the repro still throws. This is pre-existing and unrelated to the rewrite: master before #107988 (`cd3c529cd3a`) already throws in that configuration, and it involves `JSON` subcolumns, not `Map`. The new test pins the setting to `1` because it is randomized in CI. **Tests.** `04815_prewhere_guard_throwing_expression_subcolumns` is a regression test for #111879. `04814_optimize_map_element_to_subcolumn_setting` covers the new setting and pins its value, because `clickhouse-test` randomizes it. The tests that assert the rewrite (`04000`, `04513`, `03989`, `04040`, `04207`) enable the setting explicitly. All 177 `prewhere`-named stateless tests were compared against a pre-revert binary and none changes behavior. ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Added a new setting `optimize_map_element_to_subcolumn` (disabled by default) that controls rewriting `m['key']` into the `Map` key subcolumn `m.key_<key>`. It is disabled by default because the bucketed `Map` serialization it relies on is being reworked. Also reverted the PREWHERE condition grouping that was introduced for that rewrite, which caused `CANNOT_INSERT_NULL_IN_ORDINARY_COLUMN` when a guard condition and a potentially throwing expression, such as a `CAST`, read subcolumns of the same column, for example a `JSON` or `Map` column.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114150",
          "createdAt": "2026-08-10T12:25:51Z",
          "updatedAt": "2026-08-13T17:45:22Z",
          "timestamp": "2026-08-13T17:45:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "Avogar",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:c69d773c99d7c0678028",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114670",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114670",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: expand the Managed Postgres autoscaling documentation",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ## Summary - Expand the one-line Autoscaling section in the Managed Postgres scaling docs with the behavior sourced from the Ubicloud codebase: the 85% storage notification, the 90% automatic scale-up, and the 95% maintenance-window bypass - Document what autoscaling means for the cutover and client connections: same process as a manual instance change, connections dropped and in-flight transactions rolled back, DNS repointed to the new primary under the same hostname - Document read-only mode: free-space trigger and recovery thresholds per disk size, the error writes receive, and automatic recovery after the scale-up - Add a worked example scenario for a 1024 GB instance scaling to 2048 GB",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114670",
          "createdAt": "2026-08-13T16:59:03Z",
          "updatedAt": "2026-08-13T17:45:05Z",
          "timestamp": "2026-08-13T17:45:05Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "amogiska",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4d6e36e7c2b4510fba98",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114466",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114466",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: internationalize master",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114466",
          "createdAt": "2026-08-12T10:49:59Z",
          "updatedAt": "2026-08-13T17:44:49Z",
          "timestamp": "2026-08-13T17:44:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 47
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "locadex-agent[bot]",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c92b03a5efbe3f5c5df3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109367",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109367",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Scheduler: reclaimable memory tracking and dynamic spilling",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/109064 --> Scheduler-side support for reclaimable memory tracking and dynamic spilling, built on top of the memory reservations subsystem. The design is described in #109064. Queries can report the portion of an allocation that can be spilled or discarded on request (`IAllocationQueue::setReclaimable`), which is aggregated bottom-up as a new `reclaimable` field on every scheduler node. Each `AllocationLimit` gains a soft limit: when a workload's allocated memory exceeds it and the subtree has reclaimable memory, the scheduler asks a victim to reclaim memory (`ResourceAllocation::spillAllocation`, mirroring `killAllocation`) instead of waiting for the hard `max_memory` limit to force a kill. The request is replied: the query finishes it with `IAllocationQueue::finishSpill` after issuing the decreases for the freed memory (zero reclaimable doubling as a decline), the reply travels to the root as `Update::spilled` on the same propagation path as the state it describes, and at most one spill request is outstanding per subtree until it arrives; a victim that leaves the queue counts as having replied. Victim selection is deterministic and matches the kill order (largest usage, least precedence, largest allocation), descends a single root-to-leaf path via reclaimable-filtered ordered sets, and is fail-close: with nothing reclaimable, or with no soft limit configured, behavior is exactly as before. The soft limit is configured per workload via two new settings: `max_memory_before_spill` (absolute) and `max_memory_to_spill_ratio` (a fraction of the workload's own `max_memory`), smaller wins. The new state is observable in `system.scheduler`: `reclaimable`, `spills`, and the effective `soft_limit`. This is the scheduler side only. `MemoryReservation::spillAllocation` is currently a no-op; the query side that reports reclaimable memory and reacts to spill signals is a separate change. Documentation for the scheduler-side workload settings (`max_memory_before_spill`, `max_memory_to_spill_ratio`) and a `Spilling reclaimable memory` section are included here; the set of operators that can spill will be documented together with the query-side reaction. Invariants for the new machinery are documented on `ISpaceSharedNode` (I1-I8). Added unit tests cover fail-close, largest-first selection, fair descent skipping unreclaimable subtrees, the gate held until the victim's reply, the reply deferred to a pending decrease, re-signalling until under the soft limit, clamping reclaimable on a shrink, a declining victim, a victim removed mid-spill, the soft-limit enable transition via `CREATE OR REPLACE WORKLOAD`, the settings (absolute, ratio, and smaller-wins combination), a concurrency stress, and a parametrized throughput test. Verified under ThreadSanitizer (145 scheduler/workload gtests, 0 data races). ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added workload settings `max_memory_before_spill` and `max_memory_to_spill_ratio` for the experimental memory reservation scheduling. They configure a soft memory limit above which a workload's queries are asked to spill reclaimable memory, before the hard `max_memory` limit forces an eviction. This is the scheduler-side foundation: queries do not yet report reclaimable memory or react to spill requests, so the settings have no effect until the query-side change lands. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109367",
          "createdAt": "2026-07-03T21:49:24Z",
          "updatedAt": "2026-08-13T17:44:40Z",
          "timestamp": "2026-08-13T17:44:40Z",
          "metrics": {
            "reactions": 1,
            "comments": 10
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "serxa",
          "state": "open",
          "assignees": [
            "azat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:de8e1cd2ced30592bc38",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114328",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114328",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Enhance MetadataStorageFromMemory",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Pure refactoring change, doesn't affect any working part of code.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114328",
          "createdAt": "2026-08-11T13:44:39Z",
          "updatedAt": "2026-08-13T17:44:07Z",
          "timestamp": "2026-08-13T17:44:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-not-for-changelog",
            "comp-object-storage-disks"
          ],
          "author": "alesapin",
          "state": "open",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c2950226042c25f0e2cf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108653",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108653",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support `GROUPS` frame mode for window functions",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> The query below applies the same `1 PRECEDING AND 1 FOLLOWING` bounds as a `ROWS`, a `RANGE`, and a `GROUPS` (this PR) frame. The `order` column contains duplicate and non-consecutive values, so the three modes cover different rows: ```sql CREATE TABLE wf_frame_groups (`order` UInt64, value UInt64) ENGINE = Memory; INSERT INTO wf_frame_groups FORMAT Values (10, 1), (10, 2), (20, 3), (30, 4), (30, 5); SELECT order, value, groupArray(value) OVER (ORDER BY order ROWS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS rows_frame, groupArray(value) OVER (ORDER BY order RANGE BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS range_frame, groupArray(value) OVER (ORDER BY order GROUPS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS groups_frame FROM wf_frame_groups ORDER BY order, value; ``` ```response ┌─order─┬─value─┬─rows_frame─┬─range_frame─┬─groups_frame─┐ │ 10 │ 1 │ [1,2] │ [1,2] │ [1,2,3] │ │ 10 │ 2 │ [1,2,3] │ [1,2] │ [1,2,3] │ │ 20 │ 3 │ [2,3,4] │ [3] │ [1,2,3,4,5] │ │ 30 │ 4 │ [3,4,5] │ [4,5] │ [3,4,5] │ │ 30 │ 5 │ [4,5] │ [4,5] │ [3,4,5] │ └───────┴───────┴────────────┴─────────────┴──────────────┘ ``` Each mode interprets the bounds differently: - `ROWS` counts physical rows, so the frame is at most three adjacent rows: the current row plus one on each side. - `RANGE` counts `order` values, so `1 PRECEDING` and `1 FOLLOWING` cover rows whose `order` is within 1 of the current row's. With gaps of 10, no neighbouring row qualifies, so the frame holds only the rows that share the current `order`. - `GROUPS` (added by this PR) counts peer groups, so `1 PRECEDING` and `1 FOLLOWING` always include the adjacent groups in full, whatever the gaps between `order` values. cc: @cwurm ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `GROUPS` frame mode for window functions (SQL:2011), e.g. `any(price) OVER (PARTITION BY symbol ORDER BY ts GROUPS BETWEEN CURRENT ROW AND 1 FOLLOWING)`. In a `GROUPS` frame the boundaries count whole peer groups — sets of rows that are equal on the `ORDER BY` key — so `N PRECEDING`/`N FOLLOWING` mean `N` peer groups before/after the current row's peer group, rather than physical rows (`ROWS`) or `ORDER BY` value distances (`RANGE`).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108653",
          "createdAt": "2026-06-26T20:45:31Z",
          "updatedAt": "2026-08-13T17:43:57Z",
          "timestamp": "2026-08-13T17:43:57Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "nihalzp",
          "state": "closed",
          "assignees": [
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:a0b0d90d4a4cb13aaa02",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113512",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113512",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "[WIP] Seal-gated reading: gate the probe side of a hash JOIN on the runtime filter and prune read ranges by it",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> <details> <summary>Claude</summary> ## Motivation JOIN runtime filters (`enable_join_runtime_filters`) filter probe-side rows only after they were read: the row-level `__applyFilter` conjunct discards non-matching rows, and the read-time index analysis from #109085 (`enable_join_runtime_filters_index_analysis`) skips granules inside already-created read tasks. In both cases the probe side starts reading concurrently with the build side, before the filter exists, so the early tasks are read unpruned. This PR adds a stronger, structural variant for the case when the probe-side join key is a prefix of the table's primary key: the probe side does not read anything until the build side completes, and the completed filter then prunes whole mark ranges by the primary key *before* read tasks are created. ## Approach The gating is expressed as an edge of the query pipeline, not as a waiting state: - `FillingRightJoinSideTransform` gets an optional \"seal\" output port. Exactly one of the concurrent filling transforms — the one that completes the build — emits a single seal chunk carrying the completed runtime filter in its chunk info; the rest just finish. - On the probe side, the sources of a gated `ReadFromMergeTree` are replaced by `SealGatedReadTransform`s: source-like processors whose only input is the seal. Until the seal arrives, the executor has nothing to schedule below them, so the read pools never cut a task. On the seal, the filter is handed to a `RuntimeFilterReadRangesRefiner` (installed on the read pool), which turns it into a primary key `KeyCondition` — an exact `IN`-set, or the `[min, max]` envelope when the exact set overflowed into a bloom filter — and drops non-matching mark ranges at task-cut time, reusing the refiner contract of the MergeTree read pools. - The plan-level pass `markSealGatedReading` finds hash joins whose pushed-down `__applyFilter` conjunct references a probe-side primary key column, and marks the join step and the reading step. `JoinStep::updatePipeline` then wires the seal to the pending seal inputs collected by the `Pipe`. - Everything fails open: a gated read whose seal cannot be wired (the build side of a join, `YShaped`/by-shards pipelines, cancellation) is fed from a `NullSource` and reads ungated with row-level filtering only. FINAL, parallel replicas, and joins sharded by primary key ranges are not gated. Both the default multi-threaded pool and the in-order reading paths are gated (including `max_threads = 1` and reads in primary key order). On a gated read, the read-time index analysis by the same runtime filter is skipped as redundant — the refiner is already granule-exact through the primary key; filters of other joins are kept. Enabled by the experimental setting `enable_join_seal_gated_reading` (default off) on top of `enable_join_runtime_filters`. This is also groundwork for epoch-based (punctuated) collocated joins, where the same reader will consume a stream of per-epoch seals. ## Results On a 90M-row probe table (`ORDER BY k`, warm cache; `tests/performance/join_seal_gated_reading.xml`, CI perf host numbers): - 100 sparse build keys (exact `IN`-set path): 172 ms → 12 ms - 1M build keys in a narrow band (bloom overflow, `[min, max]` envelope path): 645 ms → 54 ms Compared with the read-time granule pruning by the same runtime filter (`enable_join_runtime_filters_index_analysis` + `use_skip_indexes_on_data_read`), the single-query latency on local storage is on par (the probe side cannot run far ahead of the build even ungated: the join does not consume it until the hash table is ready, so port backpressure stalls it after about a chunk per stream). The structural difference of gating is that no read or prefetch is issued for pruned ranges at all, and that it is the seam for the per-epoch seals of collocated joins. The stateless test `04653_join_seal_gated_reading` asserts result equality with ungated execution, the pipeline structure, `ReadPoolRangeRefinerDroppedMarks`/`read_rows` contrast, the fail-open shapes, and the suppression of the redundant read-time analysis. </details> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental setting `enable_join_seal_gated_reading` (default off). When enabled together with `enable_join_runtime_filters` and the probe-side join key is a prefix of the table's primary key, the probe side of a hash JOIN starts reading only after the build side completes, and the completed runtime filter prunes whole mark ranges by the primary key before read tasks are created, instead of only filtering rows after they were read.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113512",
          "createdAt": "2026-08-05T15:13:35Z",
          "updatedAt": "2026-08-13T17:43:34Z",
          "timestamp": "2026-08-13T17:43:34Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-improvement",
            "pr-performance"
          ],
          "author": "KochetovNicolai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cddf71ea526da628e781",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114401",
        "event": "changed",
        "observedAt": "2026-08-13T17:47:07.884300Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114401",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Keeper: do not lose a session request when the Raft leader changes",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/78474 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a bug where a client connecting to ClickHouse Keeper during a Raft leader change could be held for the whole `session_timeout_ms` (30 seconds by default) before its connection was rejected, instead of being rejected as soon as the in-flight request was dropped. The connecting client can now reconnect to another replica sooner. ### Description A client's `Connect` makes Keeper submit an internal `SessionID` request. If the Raft append stream breaks while it is in flight, during a leader election say, that request is lost silently. **Root cause.** Such a request carries `session_id = -1` and no xid; its identity lives in `(server_id, internal_id)`. Two places used the wrong key: * `KeeperRequestDispatcher::onCommit` correlated a commit with its in-flight head by `(session_id, xid)`. Every `SessionID` request shares `(-1, 0)`, and `onCommit` runs on every node for every commit, so a `SessionID` committed for another server retired a still-uncommitted local request. * The error path queued the response for lookup by `session_id`, which `-1` has no callback for, so it was discarded and `getSessionID` timed out. The real waiter is a promise keyed by `internal_id`. **The change.** `onCommit` additionally requires `(server_id, internal_id)` to match for `OpNum::SessionID`. Error responses go through one `SessionID`-aware routing helper on `KeeperDispatcher`, wired into **both** dispatchers: `use_new_dispatcher` is a setting and the old one shares the defect. The request now fails at the in-flight drain bound rather than at `session_timeout_ms`. `KeeperTCPHandler` does not branch on the error code, so that earlier rejection is the whole user-visible gain; the accurate `ZCONNECTIONLOSS` only improves the server log. `test_keeper_force_recovery` also gets a retry around the connect after the election, since dropping in-flight appends is deliberate. The correlation fix is scoped to `OpNum::SessionID`. The garbage collectors also use negative session ids, but their `TryRemove` is idempotent and unwaited, so they are unaffected. **Validation.** New gtests cover both the routing decision and the production wiring behind it, each verified to go red under a mutation of the change it covers; the old dispatcher was exercised with `use_new_dispatcher = false`. An injected fault at the connect under test reddens the integration test without the retry and passes with it.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114401",
          "createdAt": "2026-08-12T00:10:11Z",
          "updatedAt": "2026-08-13T17:47:09Z",
          "timestamp": "2026-08-13T17:47:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f3439cdd231e2838ffe7",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107637",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107637",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Merge filters into join during join reordering",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Added a new setting `query_plan_merge_filters_into_join` that allows merging `Filter` steps into the `JOIN` step during join reordering, so `WHERE` predicates participate in join order optimization",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107637",
          "createdAt": "2026-06-16T15:25:08Z",
          "updatedAt": "2026-08-13T18:01:48Z",
          "timestamp": "2026-08-13T18:01:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:83b3900e3fa4ea27cadc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114673",
        "event": "discovered",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114673",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Read a subcolumn of an `ALIAS` parent through a `Merge` table instead of a default",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/112975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix reading a subcolumn (`.size0`, `.null`, a tuple element, `String.size`) through a `Merge` table when the underlying table declares the parent column as `ALIAS`. Such a read returned the type default, and when the parent was selected in the same query it could return the value of an unrelated column. ### Description Selecting `arr.size0` from a `Merge` table returns `0` when the child declares `arr` as an `ALIAS` column, while the same query against the child returns the correct `5`: ```sql CREATE TABLE ag (n UInt64, arr Array(UInt8) ALIAS [1,2,3,4,5]) ENGINE = MergeTree ORDER BY tuple(); INSERT INTO ag (n) VALUES (77); CREATE TABLE mg (arr Array(UInt8), n UInt64) ENGINE = Merge(currentDatabase(), '^ag$'); SELECT arr.size0 FROM mg; -- 0, should be 5 ``` Root cause: `ColumnsDescription::add` registers no subcolumns for an `ALIAS` column, because the value has to be extracted after the expression is evaluated. So `arr.size0` does not resolve against the child, but does against the `Merge` table, where `arr` is ordinary. `ReadFromMerge::getModifiedQueryInfo` reads that as \"the child does not have this column\": it substitutes a constant default, and its `with_aliases` loop skips the column, so the child is never asked for the data. With the parent also selected, alias expansion put the alias's dependency column in the child read list, and the misaligned read surfaced that value under the subcolumn's name. The analyzer already implements the extract-after-evaluation contract, turning such a name into `getSubcolumn` over the alias expression. This teaches both guards to recognise the case and route it to the existing alias branch with the full identifier, adding no new carrier. Also wrong for `Tuple.a`, `String.size` and `Nullable.null` (a NULL read back as non-NULL), under `optimize_functions_to_subcolumns` 0 and 1, and it mis-scoped a row policy. `Map` subcolumns raise `NO_SUCH_COLUMN_IN_TABLE` before and after. The old analyzer raises `UNKNOWN_IDENTIFIER` and is unchanged, hence the `no-old-analyzer` tag. `MATERIALIZED` and `DEFAULT` parents were already correct, and the default substitution for a child genuinely lacking a column is preserved; both have test arms. A mistyped child (`String ALIAS` under an `Array(UInt8)` Merge column) stays as it is, because `String` has no `size0`. That conversion case is #112975.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114673",
          "createdAt": "2026-08-13T17:59:07Z",
          "updatedAt": "2026-08-13T18:01:28Z",
          "timestamp": "2026-08-13T18:01:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:84ddb08fd199da3157e3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114184",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114184",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Implementing generic block nested loop join",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): * Implemented generic block nested loop join",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114184",
          "createdAt": "2026-08-10T16:04:43Z",
          "updatedAt": "2026-08-13T18:00:39Z",
          "timestamp": "2026-08-13T18:00:39Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4d38374f494cd8d249d3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114409",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114409",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Revert the PromQL topk/limitk streaming plan and its shared-subquery materialization\"",
          "text": "Reverts ClickHouse/ClickHouse#114326 Depends on https://github.com/ClickHouse/ClickHouse/pull/113397",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114409",
          "createdAt": "2026-08-12T01:15:49Z",
          "updatedAt": "2026-08-13T18:00:37Z",
          "timestamp": "2026-08-13T18:00:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-not-for-changelog"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:791fe34069cb16f9cb3d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111973",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111973",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Let read-in-order propagate through SpillingHashJoin",
          "text": "`SpillingHashJoin::hasDelayedBlocks` was hardcoded to `true`, even in the `IN_MEMORY_JOIN` state where nothing is ever delayed. That is the flag gating the read-in-order-through-join and top-k-through-join optimizations, so wrapping a hash join for auto-spilling silently disabled both — and `max_bytes_ratio_before_external_join` defaults to `0.5`, which wraps every hash join. The in-tree comment in `topKThroughJoin.cpp` already describes this as the steady state. `IJoin` documents that \"SpillingHashJoin overrides `keepLeftPipelineInOrder` to forbid switching to GraceHashJoin at runtime\", but no such override existed, so the escape hatch the comment describes was never implemented. This implements it. `keepLeftPipelineInOrder` now pins the join to its in-memory algorithm, and `hasDelayedBlocks` reports `false` from that point on. Pinning is required for **correctness**, not just speed: once the plan drops a sort because the join preserves the left order, a later switch to `GraceHashJoin` would scatter rows by hash and silently return them in the wrong order. The optimizer asks before it commits (`findReadingStep` checks feasibility, and `keepLeftPipelineInOrder` is only called later, if reading in order actually turns out to be possible). So a new `IJoin::canKeepLeftPipelineInOrder` carries the question \"would you preserve the order if I asked?\", defaulting to `!hasDelayedBlocks()` so every other join keeps its current answer. `SpillingHashJoin` answers yes while still reporting delayed blocks, and only stops reporting them once actually pinned — if the optimization turns out not to apply, nothing is pinned and the delayed-block transforms are still built. **Trade-off:** a pinned join can no longer spill, so it holds the whole right side in memory and the memory tracker enforces the limit, exactly as it would with no auto-spill threshold configured. Since dropping the sort and then spilling would be a wrong-results bug, the only alternative is the conservative status quo of never propagating read-in-order through these joins. Both are now reachable: the new setting `query_plan_read_in_order_through_spilling_join` (default `1`) turns the optimization off again, and it is registered in `SettingsChangesHistory` with `previous_value = false`, so `compatibility` set to a version before 26.8 restores the old behavior. `ConcurrentHashJoin` does not spill on its own (its threshold argument only bounds preallocation), so `switchToGraceHashJoin` is the single place the invariant has to hold. The same gate applies to the first-pass `topKThroughJoin`: an `ORDER BY left_key LIMIT n` over a spill-capable `LEFT JOIN` now steps aside for the second-pass read-in-order plan instead of materializing a pushed-down `Sort` and `Limit`, and goes back to pushing them down when the setting is `0`. The handoff keeps the deferral's pre-existing conditions — most notably, `query_plan_join_swap_table` must be explicitly `false`, because under the default `auto` a later optimization may swap the join sides and invalidate the read-in-order plan; with `auto`, such queries keep the pushed-down `Sort` and `Limit`. Making this the default path exposed a pre-existing problem in read-in-order through a `JOIN`, which #110283 pinned down with `04516_join_order_estimation_pruned_parts` a few hours before this branch entered the merge queue. `max_rows_to_read` with `read_overflow_mode = 'throw'` is not checked against the rows a query reads, but against the rows the reading steps announce up front: `ReadProgressCallback::onProgress` substitutes `progress.total_rows_to_read` for `progress.read_rows` whenever the estimate is the larger of the two, and `ReadFromMergeTree` announces `min(part rows, InputOrderInfo::limit)`. `buildSortingDAG` dropped the limit at every `JoinStep`, so a plan reading in order through a `JOIN` always announced whole parts, and `ORDER BY left_key LIMIT 1` over a `LEFT JOIN` that reads 20 rows announced 1010 and failed `max_rows_to_read = 100`. Dropping the limit is right for a join that can filter the left stream, but a `LEFT ALL`/`LEFT ANY` join emits at least one row for every left row, so `n` output rows need at most `n` left rows - duplication only makes fewer left rows necessary. The limit now survives those joins and is dropped everywhere else (`INNER`, `SEMI`, `ANTI`, ...). `InputOrderInfo::limit` never truncates a read - it feeds the announcement, the read-pool task size, the `take_full_part` heuristic and `use_buffering` - so this cannot change results. The behavior was already reachable on `master` with `max_bytes_ratio_before_external_join = 0`, which makes `topKThroughJoin` defer to the second pass exactly as it now does by default. Related: https://github.com/ClickHouse/ClickHouse/pull/111248 Related: https://github.com/ClickHouse/ClickHouse/pull/111972 Related: https://github.com/ClickHouse/ClickHouse/pull/110283 ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Reading in order through a JOIN now also works when a hash join has an automatic spill-to-disk threshold configured (the default since 26.5). When this optimization applies, the join stays in memory so it preserves the left-side order; set `query_plan_read_in_order_through_spilling_join = 0` to restore the previous behavior, where such joins are not used for read-in-order and remain free to spill. As part of this, an `ORDER BY ... LIMIT` over a `LEFT JOIN` read in order no longer reports the whole table as the number of rows it intends to read, so it no longer trips `max_rows_to_read` with `read_overflow_mode = 'throw'` for a read that stops early.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111973",
          "createdAt": "2026-07-26T19:07:03Z",
          "updatedAt": "2026-08-13T18:00:29Z",
          "timestamp": "2026-08-13T18:00:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:64c8451d19dc18f42841",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114665",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114665",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allow overriding the HTTP method for SELECT through the url table function and URL engine",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/62352 ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `http_method='POST'` — as a key-value argument of the `url` table function and the `URL` table engine, or through the `http_method`/`method` named collection keys — now applies to `SELECT` queries: reads use `POST` instead of the default `GET`, for servers that accept only `POST`. Schema inference follows the configured method. `PUT` keeps its write-only meaning: a `SELECT` through a configuration with `http_method='PUT'` still uses `GET`. ### Implementation notes - `IStorageURLBase::getReadMethod()` now returns `POST` when the configured `http_method` is `POST`. `PUT` still applies to writes only (pre-signed upload URLs, #44326), and `INSERT` behavior is unchanged: `POST` by default, `PUT` when configured. - The write path no longer mutates the shared `http_method` member when defaulting to `POST` — the storage instance is shared between queries, so persisting the write default would have flipped subsequent reads of the same table from `GET` to `POST` (and raced concurrent reads). - Schema inference (`getTableStructureAndFormatFromData` and the URL read-buffer iterator) uses the same effective read method instead of hardcoded `GET`, for both `url()`/`URL` and `urlCluster`. - The `http_method = '...'` (or `method = '...'`) key-value argument is parsed by `StorageURL::evalArgsAndCollectHeaders` alongside `headers(...)`, kept in the `CREATE` AST (so it survives `SHOW CREATE TABLE` and DETACH/ATTACH), validated to be `POST`/`PUT`, and rejected when the URL scheme dispatches to another backend (`file://`, `s3://`, ...), mirroring `headers(...)`. - In the analyzer, the argument's left-hand identifier is excluded from column resolution (`skipAnalysisForArguments`), like `headers(...)`. - Documentation for the `url` table function and the `URL` engine is updated. The new stateless test `04869_url_select_http_method` asserts the method actually used on the wire via `system.query_log.http_method`: default `GET`, `POST` override for both the data and the schema-inference requests, `PUT`-configured `SELECT` staying on `GET`, the `URL` engine path, rejection of unsupported methods, and persistence of the argument in the table DDL. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114665",
          "createdAt": "2026-08-13T16:17:38Z",
          "updatedAt": "2026-08-13T18:00:25Z",
          "timestamp": "2026-08-13T18:00:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "valerypetrov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b63ad20a3dfa01458dcf",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114273",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114273",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add test: A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested",
          "text": "_Found via ClickGap automated review. Please close or comment if this is incorrect or needs adjustment._ _This is a test-only PR — no source code changes. Please review: test quality, whether the claimed coverage gaps are real, and whether test output makes sense._ Adds test coverage for 1 untested code path, found during automated review of [PR #109744](https://github.com/ClickHouse/ClickHouse/pull/109744). That PR (1) The PR changes `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:403-421): the whole-column `input_format_null_as_default` short-circuit (`SerializationNullable::deserializeNullAsDefaultOrNestedTextCSV`) is now skipped for a bare, non-empty `Tuple` whose `tuple_ **1. A `\\N` CSV field belonging to a nested `Tuple` / `Nullable(Tuple)` element of a separate-columns `Tuple` is untested** `src/Processors/Formats/Impl/CSVRowInputFormat.cpp:413`, `src/DataTypes/Serializations/SerializationTuple.cpp:683` **Risk:** `CSVFormatReader::readFieldImpl` (src/Processors/Formats/Impl/CSVRowInputFormat.cpp:409-413) now skips the whole-column `null_as_default` short-circuit for a bare `Tuple`, so the leading field is handed to `SerializationTuple::deserializeTextCSV`, whose per-element arm at src/DataTypes/Serializations/SerializationTuple.cpp:683 applies `null_as_default` to the FIRST ELEMENT — and when that element is itself a `Tuple`, one field stands for the whole nested element. This is exactly the sentence … **What's unique vs PR tests:** The PR's 04405_csv_tuple_leading_null_null_as_default.sql covers a leading `\\N` for a SCALAR first element and, for nested tuples, only `Tuple(Nullable(Int32), Tuple(Int32, Int32))` with `\\N,2,3` where the nested tuple's own fields are present; its comment states `a null inside a nested tuple is read as that whole nested element and is not covered`. This test covers precisely that: the `\\N` field is the field of a nested `Tuple` element (and of a `Nullable(Tuple)` element), the resulting one- … [Try it on ClickHouse Fiddle](https://fiddle.clickhouse.com/e4ec131d-5685-4047-bd72-b61c347ba13b) cc @groeneai (author of #109744), @Avogar (merged/approved #109744) — could you take a look, and add the `can be tested` label if this looks good? ### Changelog category (leave one): - Not for changelog (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable — test-only change. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114273",
          "createdAt": "2026-08-11T05:31:50Z",
          "updatedAt": "2026-08-13T17:59:33Z",
          "timestamp": "2026-08-13T17:59:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-not-for-changelog",
            "can be tested"
          ],
          "author": "clickgapai",
          "state": "open",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:490deb8d212a09b1a837",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111287",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111287",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix double free when finalizing -State aggregates under looping combinators",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/110975 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed a server crash (double free) that could happen when finalizing an aggregate function with the `-State` combinator nested under a looping combinator (`-Resample`, `-ForEach`, `-Map`), for example `groupArrayStateResample`, if a memory limit was reached during finalization. ### Description Reported on https://github.com/ClickHouse/ClickHouse/pull/110975 (unrelated to that PR). Found by the Stress test (amd_debug): a segfault in `Aggregator::prepareChunkAndFillWithoutKey`, reached from `ConvertingAggregatedToChunksTransform::initialize`. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=110975&sha=838d0b61235b06939c8acd923ebd396f51cf5b10&name_0=PR&name_1=Stress%20test%20%28amd_debug%29 Root cause: the `-State` combinator transfers its result by aliasing the raw aggregate state pointer into a `ColumnAggregateFunction` (`AggregateFunctionState::insertResultInto` -> `getData().push_back(place)`); ownership passes to the column. `Aggregator::insertAggregatesIntoColumns` relies on this transfer being atomic per place: on an exception it destroys the whole place exactly once. A looping combinator nested over `-State` aliases many sub-states one at a time; the `push_back` into the column's pointer array can reallocate and, being memory-tracked, throw `MEMORY_LIMIT_EXCEEDED` mid-loop. The already-transferred sub-states are then freed once by the aggregator's full `destroy()` and again by `~ColumnAggregateFunction`, i.e. a double free. Reproducer (crashes without the fix, returns a memory-limit error with it): ```sql SELECT arrayMap(x -> finalizeAggregation(x), state) FROM (SELECT groupArrayStateResample(0, 1048576, 1)(number, number % 20) AS state FROM numbers(100000)) SETTINGS max_memory_usage = 150000000, max_rows_to_read = 0; ``` Fix: reserve the destination columns before the transfer loop so the aliasing `push_back`s cannot reallocate (and therefore cannot throw) once a transfer has started. `ColumnAggregateFunction` used the no-op `IColumn::reserve`, so a real `reserve()`/`capacity()` over its state-pointer array is added. For `-Map`, the (possibly variable-width) key inserts are moved into their own loop before the value transfer, keeping the throwing work out of the aliasing loop. Reserving happens before any aliasing, so a throw there is harmless. The transfer loop is now non-throwing at the point of aliasing, restoring the atomic-per-place contract; results are unchanged. The fix covers all three looping transfer combinators (`-Resample`, `-ForEach`, `-Map`), which share the aliasing path; non-looping combinators delegate a single call and are already atomic. The added stateless test reproduces the crash deterministically via `-Resample` (empty buckets keep memory low until the finalization transfer, so a memory limit reliably lands the throw mid-transfer). `-ForEach` and `-Map` build their sub-states eagerly during aggregation, so they are not deterministically reproducible under a memory limit, but are fixed as the same class via the shared transfer path.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111287",
          "createdAt": "2026-07-21T20:34:03Z",
          "updatedAt": "2026-08-13T17:59:29Z",
          "timestamp": "2026-08-13T17:59:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8f621dca81bb70c84103",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114475",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114475",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113289 to 26.5: Fix quadratic JSON subcolumn skip-index matching over a large dotted constant",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113289 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31593284251/job/94102850957)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114475",
          "createdAt": "2026-08-12T12:06:09Z",
          "updatedAt": "2026-08-13T17:59:14Z",
          "timestamp": "2026-08-13T17:59:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-ch-test-poll3",
          "state": "closed",
          "assignees": [
            "alexey-milovidov",
            "Avogar",
            "groeneai"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:f02f2f75c7c55ebcbf28",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:105429",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:105429",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Release pull request for branch 26.5",
          "text": "This PullRequest is a part of ClickHouse release cycle. It is used by CI system only. Do not perform any changes with it. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches security-sensitive internal/DDL execution gating by replacing `query_kind` checks with a new server-set flag, which could affect ON CLUSTER/replication/backup behavior if mis-propagated. Also tightens Arrow/Native input validation, which may reject previously-accepted malformed inputs and impact ingestion edge cases. > > **Overview** > Introduces a new `Context` flag `is_ddl_or_on_cluster_internal` (server-set and non-spoofable) and switches multiple code paths from `ClientInfo::QueryKind::SECONDARY_QUERY` to this flag for *security-sensitive* decisions, including internal backup/restore gating, Replicated DB DDL handling, `ON CLUSTER`/UUID-macro allowances, and context creation in DDL/replication/system operations. > > Hardens Arrow ingestion by reading geo metadata from the Arrow schema, validating BinaryArray offsets/lengths against buffer bounds (and handling absent/empty buffers), and avoiding `mutable_data()` usage in geo parsing; adds regression tests for corrupted Arrow offsets and geo metadata. > > Improves error classification for Variant deserialization under Native format (reports `INCORRECT_DATA` instead of `LOGICAL_ERROR`), adds an integration test for malformed Native Variant payloads, and updates planner behavior to disable/forbid parallel replicas when `additional_table_filters` is used without `serialize_query_plan` (with new stateless tests). > > Minor: CI style job skips test-number gap checks on release/backport branches, MySQL protocol test fixes dotnet working dir, and release metadata (version/contributors) is updated. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 1a682e2727943363cb077a81015b9242ab666852. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/105429",
          "createdAt": "2026-05-20T13:36:34Z",
          "updatedAt": "2026-08-13T17:59:14Z",
          "timestamp": "2026-08-13T17:59:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "release"
          ],
          "author": "robot-clickhouse",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fc3e5bf409366243bf8d",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114531",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114531",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reject a lossy codec on columns backing keys and indexes",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/114406 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): A lossy codec such as `SZ3` is now rejected at DDL time on any column that backs the sorting key, primary key, partition key, a secondary index or the unique key. Because such a codec does not return the value that was written, a merged part could be stored out of physical order and index analysis would skip rows matching the query. Column statistics are no longer built for, or used to prune parts by, a lossily compressed column, which returned too few rows on default settings. Existing tables stay loadable. ### Description `ORDER BY` makes a physical promise: rows are stored inside a part in sorting-key order and the primary index samples those stored values. A lossy codec breaks `read(write(v)) == v`, so the merge sorts pre-compression values while the part stores post-compression ones. `SZ3` is not monotonic, so the stored sequence is not sorted. The same applies to anything else computed from the pre-write block: skip-index granules, the unique-key index, partition values. The issue's reproducer gives `min(i - prev) = -0.2436889648437699` after `OPTIMIZE TABLE t FINAL` and 0 after one `INSERT`: the disorder appears only at the merge, where a debug build aborts in `CheckSortedTransform`. Lossiness comes from the existing `ICompressionCodec::isLossyCompression()`, not a codec-name list. `CompressionCodecMultiple` did not override it, so a stacked `CODEC(SZ3(...), LZ4)` reported itself lossless; it now ORs over its children. Classification is per serialized substream, mirroring `MergeTreeDataPartWriterWide::addStreams`, so `ORDER BY arr.size0` stays allowed while `arraySum(arr)` is not. The check sits at the two user-facing entry points: `registerStorageMergeTree` for CREATE and full-definition ATTACH, `checkAlterIsPossible` for ALTER on the initiating execution, following `4a29ef847411256`. A replica replaying a durable DDL entry is not re-checked, which would wedge its DDL worker. Column statistics, found during review, needed a different remedy: they are built pre-write too, but `auto_statistics_types` attaches `basic` to every numeric column, so a DDL rejection would ban the codec on any table that does not opt out. Instead none are built for such a column, and the pruner ignores any an earlier version wrote. With no setting changed, a predicate matching 2 rows returned 0. `system.columns` still lists them, as the two settings' descriptions now note. Seen in CI on one AST-fuzzer run, on #113575: [report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=113575&sha=84df30db17165b65e3c78bc11513ed26534187d7&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%2C%20targeted%29). <details> <summary>Related cases deliberately left out of this PR</summary> - A **projection** over a lossily compressed column answers differently than the base part: `sum(i)` is `249957173.79` read from the base part and `249998750` via the projection. The mechanism is a separate post-write consumer, which writes the base part lossily and then computes the projection from the unchanged pre-compression block, so it follows in its own PR rather than being folded in here. - A legacy table can still gain an implicit minmax index over such a column through server configuration on the load path. Reachable, but a boundary sweep over 300 values produced no wrong result. - A part written by an earlier version keeps its statistics if the codec is then replaced by a lossless one without any merge or mutation rewriting that part. Deciding this needs per-part codec provenance, which parts do not record. `ALTER TABLE ... MATERIALIZE STATISTICS` or `OPTIMIZE TABLE ... FINAL` clears it. - `ORDER BY length(arr)` is now rejected although its value is exact. A key expression records which columns it needs and not which of their streams, so at this layer `length(arr)` and `arraySum(arr)` are indistinguishable, and `arraySum` is a genuine wrong-results carrier. The syntactic form `ORDER BY arr.size0` stays allowed because it names the stream. - Alias-dependent index rebuilds look incorrect independently of codecs: an alias-only `MODIFY COLUMN` requires no mutation while the index expression is rebuilt, so existing parts keep a same-named index built from the old expression. No claim is made about it here. </details>",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114531",
          "createdAt": "2026-08-12T18:16:56Z",
          "updatedAt": "2026-08-13T17:58:38Z",
          "timestamp": "2026-08-13T17:58:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5dca3b01efb52460be48",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113681",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113681",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Replace the per-bucket hash map in `timeSeries*ToGrid` with a sorted-append sample array",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Replaced the per-bucket hash map inside the `timeSeries*ToGrid` aggregate functions with a flat sorted array of samples: sample ingestion becomes an O(1) append for in-order inputs (the overwhelmingly common case) and the per-bucket copy-and-sort at finalization is gone. ### Description The `timeSeries*ToGrid` functions kept each bucket's samples in an `absl::flat_hash_map<timestamp, value>`: every `add()` paid a hash-map emplace, and the order-dependent functions (`rate`, `increase`, `delta`, `changes`, `resets`) copied and sorted every bucket at finalization. Yet the input is almost perfectly ordered — samples come from MergeTree tables sorted by `(id, timestamp)`; an instrumented probe on a 32-thread read of a 62.5-billion-sample table counted **1 out-of-order add in 1,474,559,998**. The bucket is now a flat, memory-tracked vector of `(timestamp, value)` pairs: O(1) append while timestamps ascend, in-place max on an equal timestamp, and a rare out-of-order add just clears a `sorted` flag — normalization (sort + max-dedup) runs lazily, only for buckets that actually saw disorder. `merge()` is a linear merge of sorted runs with an append fast path for disjoint time ranges. `forEachSample` now guarantees ascending order, so the copy-and-sort buffers are deleted from the rate/delta/changes aggregators. The wire format and `FORMAT_VERSION`s are unchanged; `deserialize()` assumes no order of incoming pairs (old peers send hash-map iteration order), so mixed-version clusters interoperate — verified in both directions. Duplicate timestamps keep the larger value with the old `std::max` argument order. Measured on the 62.5-billion-sample table: a 30-day `sum by(...)(rate(...))` over 25,600 series drops ~6% of total query CPU (1106 s -> 1037 s); wall time and peak memory move little (the scan dominates the critical path, and raw sample storage dominates the state either way). The structural point is what this enables: the sorted buffer is the prerequisite for O(1) per-segment summaries in the rate family (follow-up), which is where the ~44 GiB state peaks of such queries actually go away. All query results are fingerprint-identical. Tests: a new stateless test (shuffled and duplicate-timestamp inputs in both orders, NaN at duplicated timestamps, `-Merge` of unsorted in-memory states, interleaved parts, two-level merges, serialized-state merges through `remote('127.0.0.{1,2}', ...)`, an `AggregatingMergeTree` roundtrip, a fixed state literal in old-peer wire order); all 40 existing timeseries/PromQL stateless tests pass byte-identically; the perf test gains an ingestion-heavy scenario (50M rows, 10k series). 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113681",
          "createdAt": "2026-08-06T13:46:38Z",
          "updatedAt": "2026-08-13T17:58:34Z",
          "timestamp": "2026-08-13T17:58:34Z",
          "metrics": {
            "reactions": 1,
            "comments": 9
          },
          "labels": [
            "pr-performance",
            "comp-promql"
          ],
          "author": "nikitamikhaylov",
          "state": "open",
          "assignees": [
            "vitlibar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:880066bbf18a4546893f",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114323",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114323",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: require canonical internal links",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/114230 This is a one-off cleanup of repository-authored documentation links that use legacy redirect aliases. It updates the current English documentation and source-embedded reference documentation to use routes relative to the docs root, so the automated translation PR can parse and localize them without producing missing locale routes. This PR intentionally adds no permanent CI checks or ongoing enforcement. Its scope is limited to the current link corrections needed to get the automated translation PR parsing successfully. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114230&sha=e599855a281a4dd841be20794036908a58acbf63&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114323&sha=86d2d7e194d7fbf75cc84616a8fa0da1ed84802e&name_0=PR&name_1=Docs%20check%20%28Mintlify%29 ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Not applicable.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114323",
          "createdAt": "2026-08-11T13:26:47Z",
          "updatedAt": "2026-08-13T17:58:20Z",
          "timestamp": "2026-08-13T17:58:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-ci",
            "pr-autogenerated-docs"
          ],
          "author": "Blargian",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d5097235c34cb984ee80",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114283",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114283",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add pre-hook to insert CI links into PR body",
          "text": "### Changelog category (leave one): - CI Fix or improvement (changelog entry is not required) -- Adds a `ci_links.py` pre-hook to the `PR` workflow that, on upstream `ClickHouse/ClickHouse` pull request runs, appends a `:ci_links:` block to the PR description with: - a link to the workflow report, and - a link to a GitHub search for the corresponding sync PR (`sync-upstream/pr/<number>`). The block is added only when it is not already present, so subsequent runs do not re-edit the PR body. Non-upstream / non-PR runs are skipped, and any failure is caught so it can never break the workflow. <!-- CI automatic block start :ci_links: --> --- Workflow [[PR](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=114283&sha=latest&name_0=PR)] Sync PR [[sync-upstream/pr/114283](https://github.com/search?q=head%3Async-upstream%2Fpr%2F114283+org%3AClickHouse+type%3Apr&type=pullrequests)] <!-- CI automatic block end :ci_links: -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114283",
          "createdAt": "2026-08-11T07:48:23Z",
          "updatedAt": "2026-08-13T17:58:00Z",
          "timestamp": "2026-08-13T17:58:00Z",
          "metrics": {
            "reactions": 1,
            "comments": 3
          },
          "labels": [
            "pr-ci"
          ],
          "author": "maxknv",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7452da7f96001aad4296",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:99495",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:99495",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `GradualResizeProcessor` to limit effective parallelism for GROUP BY on small data volumes",
          "text": "When ClickHouse processes GROUP BY, it often overestimates the number of threads needed. With `max_threads = 64` but only a few thousand rows, all 64 `AggregatingTransform` instances get data, produce 64 partial hash tables, and the merge phase has to combine all of them — most nearly empty. This wastes time on merging overhead, which is especially noticeable for heavy aggregate states such as `uniq`, `uniqExact`, `groupArray`, etc. The new `GradualResizeProcessor` starts by pushing data to a single output port (or one port per split group when `min_outstreams_per_resize_after_split` applies), and activates all aggregation streams at once as soon as the configured row or byte threshold is crossed. For small datasets, only one aggregating thread receives data (or one per split group); for large datasets, all threads are used as before. New settings: - `min_rows_per_stream_for_gradual_resize` (default: `1000`) - `min_bytes_per_stream_for_gradual_resize` (default: `0`) When either threshold is non-zero, the pre-aggregation `StrictResize` is replaced with `GradualResize` in the pipeline. The optimization is enabled by default; set both `min_rows_per_stream_for_gradual_resize = 0` and `min_bytes_per_stream_for_gradual_resize = 0` to opt out. ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improve performance of GROUP BY on small data volumes.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/99495",
          "createdAt": "2026-03-14T07:35:33Z",
          "updatedAt": "2026-08-13T17:57:39Z",
          "timestamp": "2026-08-13T17:57:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 30
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:70cd890cb6d261434e31",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114643",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114643",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Re-land aggregate function `gini` in the `sum` family",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/issues/113763 Related: https://github.com/ClickHouse/ClickHouse/pull/113868 Related: https://github.com/ClickHouse/ClickHouse/pull/112280 --> ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): New aggregate function `gini`, which calculates the [Gini coefficient](https://en.wikipedia.org/wiki/Gini_coefficient) of a column of finite, non-negative numeric values. The result ranges from `0` (all values equal) towards `1` as inequality grows; for a sample of `n` values the maximum is `(n - 1) / n`. `NaN` values are skipped and infinite values are rejected. The function returns `Float64` and consumes `O(n)` memory. ### Description Re-lands the `gini` function that #113868 reverted, following the option-2 spec @ Manerone gave in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5217284116 and confirmed in https://github.com/ClickHouse/ClickHouse/issues/113763#issuecomment-5279278069. It re-adds #112280's function with exactly three items removed, and touches no file under `src/AggregateFunctions/Combinators/`: - the `getArgumentsThatCanBeOnlyNull` override, - `.returns_default_when_only_null = true` on registration, - the `argument_type->onlyNull()` branch in the creator. The property was doing the work. It made `AggregateFunctionFactory::getImpl` skip its only-null guard and build a real `gini` instance over `Nullable(Nothing)`, so `gini(NULL)` returned `Float64` `nan` where every `sum`-family function folds to `Nullable(Nothing)`. Without it the fold happens in the `Null` combinator before any other combinator is applied, and the creator's only-null branch becomes unreachable, which is also why `createAggregateFunctionSum` has no equivalent. `gini` now registers exactly like `sum`. Runtime `Nullable` handling is a separate axis, via `getOwnNullAdapter`, and is unchanged: `gini` over `[1, NULL, 3]` still returns `0.25`. Validated on three binaries: post-revert master (`gini` absent), a build carrying #112280's version, and this branch. Every literal-`NULL` cell on this branch equals the `sum` value measured on the same binary, and all non-`NULL` results are unchanged from #112280. No documentation files are touched. The page body is generated from the `FunctionDocumentation` block in `AggregateFunctionGini.cpp`, which this PR restores, so the docs autogeneration workflow fills the page from that source. cc @Manerone",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114643",
          "createdAt": "2026-08-13T13:34:59Z",
          "updatedAt": "2026-08-13T17:57:20Z",
          "timestamp": "2026-08-13T17:57:20Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-feature",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "Manerone"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b87138cdbb8124591ad3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114285",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114285",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
          "text": "Files written first turn the namespace into a plain directory, which a catalog sharing the storage view (SeaweedFS) rejects with HTTP 500; the swallowed error left an orphaned metadata file that broke retries. Ensure the namespace before the first write; propagate failures except 404-then-create and 409 (REST) / AlreadyExists (Glue). Example: https://pastila.nl/?01f91a25/620869a81af815ba6927860efbc72af4#NhE2Nbzh3AHhUHhA5Fdkzw==GCM ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Register Iceberg namespace in the catalog before writing table files (needed for SeaweedFS)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114285",
          "createdAt": "2026-08-11T08:20:50Z",
          "updatedAt": "2026-08-13T17:56:38Z",
          "timestamp": "2026-08-13T17:56:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "azat",
          "state": "open",
          "assignees": [
            "alesapin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4380e05ce22f19cdb3a2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112805",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112805",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not drop a named collection that a detached table still uses",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/96181 Related: https://github.com/ClickHouse/ClickHouse/issues/77366 Related: https://github.com/ClickHouse/ClickHouse/pull/110529 A table detached with a plain `DETACH TABLE` keeps its metadata file, so the server attaches it again on the next start. It is gone from `DatabaseCatalog` though, so `isTableExist` returns false for it, and the `check_named_collection_dependencies` check (added in #96181) treated its dependency as a stale leftover of a failed `CREATE TABLE`: it removed the dependency and let `DROP NAMED COLLECTION` succeed. The `ATTACH` replayed at startup then threw `NAMED_COLLECTION_DOESNT_EXIST`, which aborts loading the metadata, and the server did not start at all. This is how the `Stress test (arm_tsan)` job fails on master with `Cannot start clickhouse-server`: the AST fuzzer makes `04320_url_engine_dispatch_partition_and_format` leave its `URL(named_collection)` table detached (the test's `ATTACH TABLE` never reaches the server), and the test then drops the named collection. CI report: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=8aad759007771032aa94d7f4eee3e18103dd63c3&name_0=MasterCI&name_1=Stress%20test%20%28arm_tsan%29 ``` Application: Caught exception while loading metadata: Code: 722. DB::Exception: Waited job failed: Code: 695. DB::Exception: Load job 'load table test_3.04320_..._n' failed: Code: 669. DB::Exception: There is no named collection `04320_..._nc`: Cannot attach table `test_3`.`04320_..._n` from metadata file store/e5b/.../04320_..._n.sql from query ATTACH TABLE ... ENGINE = URL(`04320_..._nc`, format = 'JSON'). (NAMED_COLLECTION_DOESNT_EXIST) ``` The same signature accounts for 6 of the 8 `Cannot start clickhouse-server` failures with a missing named collection in the last 60 days (per `play.clickhouse.com`); the other two come from `03822_named_collection_drop_dependency_check`, which drops the collection with `check_named_collection_dependencies = 0` on purpose, and are the case #110529 handles by tolerating the missing collection at startup. ### Changes Per review feedback, the implementation is a simple in-memory bookkeeping in `NamedCollectionFactory` (an earlier revision inspected the metadata the detached table would be attached from, which required probing database disks and sweeping metadata directories after renames): - `DETACH TABLE` (and `DETACH DATABASE`, for every table inside) moves the dependencies of the table into a list of (collection, database, table) entries. - `DROP NAMED COLLECTION` is refused with `NAMED_COLLECTION_IS_USED` while an entry for the collection exists. - `ATTACH` does not remove the entry by itself: the dependencies are registered while the engine arguments are resolved, and the attach can still fail after that (an unknown format name, a failure creating the storage), leaving the table detached — the entry must keep protecting it. Instead, the entry is removed by the events that prove the metadata under that name is gone or harmless: `DROP TABLE`, `DETACH TABLE ... PERMANENTLY`, and `RENAME` of the (necessarily re-attached) table; `DROP DATABASE` removes the entries of the database's detached tables, and `RENAME DATABASE` re-keys them. The `DROP NAMED COLLECTION` check itself removes nothing: the table's existence in `DatabaseCatalog` is racy against in-flight attaches and detaches (the table can exist while nothing in the drop query has validated its live dependency), so the drop path is read-only and every recorded entry refuses the drop. - `DETACH TABLE ... PERMANENTLY` does not record an entry: a permanently detached table is not loaded at startup, so dropping a collection it references cannot break the server start. A later explicit `ATTACH` of such a table fails cleanly with `NAMED_COLLECTION_DOESNT_EXIST` and is recoverable by recreating the collection. - The list lives in memory only, which is consistent across a restart: a plainly detached table is attached again at the next start, where regular dependency tracking picks it up, and a permanently detached one records no entry at all. The list is deliberately imprecise in one direction: a stale entry may keep refusing the drop for a while after the detached table itself is gone, or after the table was attached back (until the table is dropped or renamed). In exchange, the `DROP NAMED COLLECTION` path performs no disk access at all. - A `RENAME TABLE` that moves a table between an `Ordinary` and an `Atomic` database changes the identity the dependency is keyed by: the move into `Atomic` assigns a fresh UUID to the table and the move out of it drops the UUID, while a dependency is keyed by the UUID for tables of `Atomic` databases and by the name for tables of `Ordinary` ones. The rename interpreter only knows the names, so the entry used to keep the identity the table had before the move and nothing found it afterwards: the detach recorded no entry and, for the `Atomic -> Ordinary` direction, the drop check even classified the entry of the still attached table as a leftover of a failed `CREATE` and dropped the collection from under it. `DatabaseOnDisk::renameTable` now re-keys the entries of the moved table to its new `StorageID`, where both identities are known. `EXCHANGE` is unaffected: it is only supported between two `Atomic` databases, where the UUIDs do not change. - A table of a database with `lazy_load_tables = 1` is attached as a `StorageTableProxy` and its real storage is built only on the first access, so the engine arguments are not resolved at load time and the dependency on the named collection they name stayed unregistered. `DROP NAMED COLLECTION` was then allowed while such a table still referenced the collection - breaking it at the first access with `NAMED_COLLECTION_DOESNT_EXIST` - and a `DETACH` of it had no dependency to move to the list of the detached ones, so the protection above silently disappeared in that mode. `DatabaseOrdinary::loadTableLazy` now registers the dependency straight from the metadata, via the new `tryGetUsedNamedCollectionName` helper. An identifier first argument counts as a collection reference only for the engines that resolve their arguments through named collections (a new `StorageFactory::StorageFeatures::supports_named_collections` flag) - for other engines an identifier means something else, e.g. a cluster name for `Distributed` - and for them the signal is time-stable: whether the collection currently exists is deliberately not checked, so the dependency of a collection that is missing at load time (say, after a drop with `check_named_collection_dependencies = 0`) protects it when it is recreated later. The exception is `Remote`/`RemoteSecure`, where the same identifier is also a valid positional argument - a cluster name - when the named-collection lookup does not resolve (marked by the new `StorageFeatures::named_collection_argument_is_ambiguous`; every other flagged engine treats an unknown collection as an error, not as a fallback to a positional form). Syntax alone cannot prove that such a table uses a collection, so the helper replicates the decision the engine's own argument parsing would make at the same moment - which is exactly what a non-lazy load of the same metadata does: the named-collection branch is taken only when a collection with that name exists, and only `key = value` overrides may follow the collection name, which a positional argument list never looks like. `MongoDB` and `MaterializedPostgreSQL` were the only flagged engines whose eager argument resolution did not pass the dependent table to `tryGetNamedCollectionWithOverrides` (they registered the dependency at the lazy load but not when the storage is built); they now pass it, so both load modes register the same dependency. `addDependency` ignores an exact duplicate, because the same dependency is registered again when the proxy is materialized. Tables created with `CREATE TABLE ... AS f(...)` need no lazy branch: a database with `lazy_load_tables = 1` deliberately loads them eagerly as a `StorageTableFunctionProxy` (see `DatabaseOrdinary::shouldLazyLoad`), and that load registers the dependency via `ITableFunction::getUsedNamedCollectionName`. - The cleanup of stale *active* dependencies (leftovers of a failed `CREATE TABLE`, pre-existing from #96181) no longer treats the table's absence from `DatabaseCatalog` alone as a proof of staleness: the dependency of an in-flight `CREATE`/`ATTACH` is registered while the engine arguments are resolved, before the table is committed to the catalog, and a concurrent `DROP NAMED COLLECTION` could prune it and drop the collection while the create later succeeds — recreating the broken metadata this PR fixes. The creating query holds the `DDLGuard` of the table name for the whole window between the registration and the commit, so the drop re-checks the table's existence under that guard before pruning: once the guard is acquired, no create is in flight, and the table's absence proves the entry is stale. (Entries with an empty database name come from dictionaries defined in the configuration files, which are not created through DDL; they are pruned as before.) The pruning removes only the exact stale entry (the collection and the recorded database, table and UUID): `CREATE TABLE ... UUID` can reuse the UUID of a failed create under a different table name, which the guard of the recorded name does not synchronize with, and removing everything under the UUID would erase the live dependency of such an in-flight create — the collection it uses could then be dropped from under the committed table. A new `create_table_pause_before_commit` failpoint keeps a create inside the window for the tests. ### Documented behavior impact `DROP NAMED COLLECTION` now rejects a collection that a detached table or a table in a detached database references, where it previously succeeded (and left a server that could not start). This is what `check_named_collection_dependencies` already promises - \"Check that DROP NAMED COLLECTION will not break tables that depend on it\" - so the documented behavior of the setting does not change, and setting it to `0` still allows the drop. No documentation update is needed. ### Verified locally (release build) - Before: `CREATE NAMED COLLECTION` + `CREATE TABLE ... ENGINE = URL(nc)` + `DETACH TABLE` + `DROP NAMED COLLECTION` succeeds, and the server then fails to start with the exact error chain above (exit code 210). Same with `DETACH DATABASE`. - After: the drop is refused, the collection stays, the table attaches back, and a restart of a server with the detached table present succeeds. - `04660_drop_named_collection_detached_table`, `04698_drop_named_collection_detached_after_rename`, and `04823_drop_named_collection_broken_attach` (all new, covering plain `DETACH TABLE` (blocks the drop), `DETACH TABLE ... PERMANENTLY` (does not block; the later `ATTACH` fails cleanly), `DETACH DATABASE`, `Ordinary` databases, renames of the table and of the database before and after the detach, stale dependencies of failed `CREATE TABLE`, and an `ATTACH` that fails after the dependencies were registered — the drop stays refused), and `04836_drop_named_collection_inflight_create` (new, runs `DROP NAMED COLLECTION` against a `CREATE TABLE` and an `ATTACH TABLE` paused between the dependency registration and the commit to the catalog: the drop blocks on the `DDLGuard` and is refused), and `04840_drop_named_collection_cross_engine_rename` (new, moves a table between an `Ordinary` and an `Atomic` database in both directions and then detaches the table and its database), and `04848_drop_named_collection_reused_uuid` (new, prunes the stale entry of a failed `CREATE TABLE ... UUID` while a create of a different table reusing the UUID is paused inside the window: the drop of the old collection succeeds, and the drop of the collection the new table uses stays refused; verified to fail without the fix), plus `03822_named_collection_drop_dependency_check` and `04003_named_collection_drop_dependency_check_dict` pass. - `test_named_collections/test.py::test_drop_while_used_by_lazily_loaded_table` (new integration test: a table using a named collection in an `Atomic` database with `lazy_load_tables = 1`, a server restart so the table comes back as a never-accessed proxy, and the drop refused both while it is attached and after `DETACH TABLE`). It is an integration test because the hole is only reachable once the in-memory list is empty, i.e. after a restart: without one, the `DETACH DATABASE` that precedes `ATTACH DATABASE` leaves its own entry behind and that entry refuses the drop on its own. Verified that it fails without the fix (the drop succeeds) and passes with it. - `test_drop_collection_recreated_under_lazily_loaded_table` (new integration test: the collection is dropped with `check_named_collection_dependencies = 0`, the server restarts while it is missing, and the drop of the recreated collection is refused; verified to fail without the fix), and `test_drop_not_used_by_lazily_loaded_distributed_table` (new integration test: a collection named after the cluster of a lazily loaded `Distributed` table is droppable; passes before and after, pinning the behavior). - `test_drop_not_used_by_lazily_loaded_remote_table` (new integration test: a collection named after the cluster of a lazily loaded `ENGINE = Remote(cluster, system, one)` table, created after the table, is droppable after a restart; verified to fail without the fix - the drop was refused with `NAMED_COLLECTION_IS_USED` by the unrelated table), and `test_drop_while_used_by_lazily_loaded_table_function` (new integration test pinning that a `CREATE TABLE ... AS bigquery(collection)` table in a database with `lazy_load_tables = 1` keeps blocking the drop after a restart, before the first access and after a `DETACH TABLE`: such tables are loaded eagerly as a `StorageTableFunctionProxy`, which re-registers the dependency; passes without any code change, confirming no lazy table-function branch is needed). ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed the server failing to start after a named collection was dropped while a detached table (or a table in a detached database) still referenced it. `DROP NAMED COLLECTION` now counts detached tables as dependents and is refused with `NAMED_COLLECTION_IS_USED`, as `check_named_collection_dependencies` already does for attached tables.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112805",
          "createdAt": "2026-07-31T19:56:49Z",
          "updatedAt": "2026-08-13T17:56:31Z",
          "timestamp": "2026-08-13T17:56:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4ccdbc61e2fc85c524f0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109891",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109891",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Reintroduce borrowed threadgroup async uaf fix",
          "text": "Reintroduce #108988 Related: https://github.com/ClickHouse/ClickHouse/pull/107030 Related: https://github.com/ClickHouse/ClickHouse/pull/108577 CI: https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=105890&sha=3dc0e76362eb18e27f1fffcd3f61ca9f13725fe8&name_0=PR&name_1=Stateless%20tests%20%28amd_tsan%2C%20s3%20storage%2C%20sequential%2C%201%2F2%29 Fixes a use-after-free risk in asynchronous work scheduled while a borrowed `ThreadGroup` is current. Borrowed `ThreadGroup` objects used by materialized view and async insert flush paths point their `performance_counters` and `memory_tracker` to the parent query group, so they are valid only while that parent group is alive. Async callbacks could capture such a borrowed group and later attach it on a pool thread after the parent query group had finished. Instead of keeping the parent `ThreadGroup` alive with a `shared_ptr`, this change keeps borrowed accounting scoped. Borrowed groups are marked explicitly, async callback capture drops borrowed groups, and thread pool callback runners capture the normalized group at task enqueue time rather than when a potentially long-lived runner object is created. This preserves synchronous borrowed accounting, but async work started from a borrowed scope runs under normal thread/global accounting instead of writing into, or prolonging the lifetime of, an already finished query group. Full ASAN reports https://gist.github.com/filimonov/1ec59047c65e3a5367c5c83f6021cc27 Compared to #108988 - added one commit with code comments + fix of the test failure https://github.com/ClickHouse/ClickHouse/issues/109841 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixes a memory safety issue where asynchronous work scheduled from materialized view processing could keep using query-level accounting after the query had finished.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109891",
          "createdAt": "2026-07-09T12:34:20Z",
          "updatedAt": "2026-08-13T17:56:28Z",
          "timestamp": "2026-08-13T17:56:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 47
          },
          "labels": [
            "pr-bugfix",
            "can be tested",
            "comp-query-execution"
          ],
          "author": "filimonov",
          "state": "open",
          "assignees": [
            "azat",
            "alexey-milovidov"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ccd292da1bba0b8e1361",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111794",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111794",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add unordered stream modifier",
          "text": "### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add STREAM UNORDERED modifier: skip the per-snapshot commit-order sort depends on https://github.com/ClickHouse/ClickHouse/pull/110653 (not for functional reason, only test) cc @alesapin @Michicosun",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111794",
          "createdAt": "2026-07-24T13:31:48Z",
          "updatedAt": "2026-08-13T17:56:27Z",
          "timestamp": "2026-08-13T17:56:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement",
            "pr-synced-to-cloud"
          ],
          "author": "SmitaRKulkarni",
          "state": "closed",
          "assignees": [
            "Michicosun"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:de64bb4feb2c6787ccff",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114661",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114661",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Revert \"Document that PREWHERE filters one join input before the JOIN\"",
          "text": "Reverts ClickHouse/ClickHouse#114484 - it is too low-quality, sorry. CC @PedroTadim @dhtclk <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1350` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114661",
          "createdAt": "2026-08-13T15:53:35Z",
          "updatedAt": "2026-08-13T17:56:23Z",
          "timestamp": "2026-08-13T17:56:23Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-not-for-changelog",
            "pr-synced-to-cloud"
          ],
          "author": "rschu1ze",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:eee84e001366067015fc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107669",
        "event": "discovered",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107669",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add `input_format_json_max_object_size` setting to limit JSON object size on parsing",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `input_format_json_max_object_size` setting to limit JSON object size on parsing. Closes https://github.com/ClickHouse/ClickHouse/issues/106704",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107669",
          "createdAt": "2026-06-16T20:49:08Z",
          "updatedAt": "2026-08-13T17:55:50Z",
          "timestamp": "2026-08-13T17:55:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "Avogar",
          "state": "open",
          "assignees": [
            "Manerone"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:c88007cedc2117bd65ae",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:112667",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:112667",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Silk integration",
          "text": "Splits the silk runtime integration out of https://github.com/ClickHouse/ClickHouse/pull/111275, so that it can be reviewed on its own. This adds the plumbing that lets ClickHouse run work on [silk](https://github.com/ClickHouse/silk) fibers, without yet putting any subsystem on them. - `Silk::initializeFiberScheduler` / `Silk::destroyFiberScheduler`, called by the server when the `enable_silk_runtime` server setting is enabled. The fiber stack size is configurable through the `silk.fiber_stack_size` configuration key (320 KiB by default, which leaves enough room for OpenSSL handshakes). - `FiberLocal` - fiber-local storage. A fiber can migrate between operating-system threads, so it must not observe another fiber's `thread_local` state; the values of the registered slots are swapped in and out on every fiber switch instead. `current_thread` (`ThreadStatus`), the OpenTelemetry tracing context, and the memory-tracker and exception blockers are moved to it. - The silk thread-local-storage sanitizer: an LLVM pass in `utils/silk-thread-local-storage-sanitizer` that instruments every `thread_local` access and aborts when a fiber touches raw thread-local storage. Without it, a variable that was not migrated to `FiberLocal` produces silent corruption rather than a diagnostic. It is enabled in the debug and ASan CI builds. - `Silk::ConnectionPool` and `Silk::streamSocketFactory` - a `Connection` pool and a socket factory that suspend the calling fiber instead of blocking the operating-system thread. `PoolBase` and `ConnectionPool` are templated on the lock and the condition variable to make that possible, and `ConnectionPool` stays an alias of the `std::mutex` instantiation, so the existing call sites are unchanged. - Memory that the runtime maps outside the C++ heap - fiber stacks and `io_uring` rings - is charged to `total_memory_tracker` through silk's mmap accounting hooks. - The low-level silk runtime counters are exported to `system.asynchronous_metrics` under a `Silk` prefix. - `Common/Fiber.h` and `Common/FiberStack.h` are renamed to `Common/StackfulCoroutine.h` and `Common/CoroutineStack.h`. They implement the boost-context coroutines used by `AsyncTaskExecutor`, which are unrelated to silk fibers, and having two different things called \"fiber\" in the same codebase is confusing. Related: https://github.com/ClickHouse/ClickHouse/pull/111275 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added experimental support for the [silk](https://github.com/ClickHouse/silk) fiber runtime, enabled with the `enable_silk_runtime` server setting. When it is enabled, the server initializes the silk fiber scheduler at startup, so that subsystems supporting it can run their jobs on fibers instead of occupying an operating-system thread while waiting for I/O.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/112667",
          "createdAt": "2026-07-30T21:32:29Z",
          "updatedAt": "2026-08-13T17:55:03Z",
          "timestamp": "2026-08-13T17:55:03Z",
          "metrics": {
            "reactions": 1,
            "comments": 1
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "mstetsyuk",
          "state": "open",
          "assignees": [
            "CheSema"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4c5ff63a2e85657d1ec3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111457",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111457",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Compare read-in-order virtual row on its covered sort-key prefix",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/106740 Closes: https://github.com/ClickHouse/ClickHouse/issues/106630 Related: https://github.com/ClickHouse/ClickHouse/pull/110725 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix a wrong result (mis-ordered merge) for read-in-order queries with a virtual row when `distinct-in-order` or `LIMIT BY` widens the read to a longer sort-key prefix than `ORDER BY` set it up for, and when a key column fixed by the filter is skipped by `ORDER BY` (e.g. `WHERE b = 1 ORDER BY a, c` on key `(a, b, c)`). The virtual row announced a wrong merge boundary: in release builds the merge could be silently mis-ordered, in debug builds the boundary assertion fired, and a `Nullable` key column after the skipped one threw the `Virtual row has different type` exception. The virtual row optimization now stays enabled in these cases. ### Description Alternative to https://github.com/ClickHouse/ClickHouse/pull/110725: instead of dropping the virtual row conversion when the in-order read prefix changes (and disabling it for skipped key columns), keep the optimization enabled and compare the virtual row only on the sort-key prefix it validly covers. **Root cause.** The read-in-order virtual row announced a wrong merge boundary in two ways: 1. *Widened prefix.* `optimizeReadInOrder` builds the virtual row conversion for the prefix `ORDER BY` needs (e.g. `CounterID`). A later optimization (`optimizeDistinctInOrder`, `optimizeLimitByInOrder`) re-requests the read with a longer prefix (`CounterID, EventDate`), but the `pk_block` width was derived from the conversion's input count, so the extra sort column was default-filled with `0` in `setVirtualRow`. In reverse order `0` understates the real values, so a real row exceeded the announced boundary: `Virtual row boundary violated in MergingSortedAlgorithm ... the virtual row announced UInt64_0 but the source then produced UInt64_1` in debug builds, a silently mis-ordered merge in release builds. 2. *Skipped fixed key.* For key `(a, b, c)` and `WHERE b = 1 ORDER BY a, c`, the fixed key `b` is skipped without an `ORDER BY` counterpart, but the conversion DAG indexed key columns densely, mapping `c` onto key column `b` (visible in `EXPLAIN actions=1`: input `b` aliased to `__table1.c`). The wrong value tripped the boundary check; a wrong type (`Nullable` key) threw the `Virtual row has different type` logical error even in release builds. Moreover, index values of the columns after the skipped key are semantically unusable: the index describes pre-filter data, so the entry `(5, 0, 9)` does not bound the filtered row `(5, 1, 3)` projected to `(a, c)`. **Fix.** - The virtual row conversion outputs only the sort-description prefix it can announce exactly: index values while the key prefix is contiguous, plus constants for fixed columns from `ORDER BY`. A skipped fixed key column ends the index-backed part: the index entry at a mark boundary may hold a filtered-out value for it, so its later components bound nothing in the filtered stream. A column fixed by the filter that stays in `ORDER BY` keeps disabling the virtual row, as before this fix. - The merge compares a virtual row only on the covered prefix and places it first on a covered-prefix tie (equivalent to treating the uncovered columns as minus infinity in the merge order, without materializing any values). The covered prefix is derived from the pk block column names in `MergingSortedAlgorithm` and carried per cursor in `SortCursorImpl::sort_prefix_limit`, honored by the generic `SortCursor::greaterAt`. A truncated virtual row can only occur with a multi-column sort description (its coverage is at least the first column), which always uses the generic cursor, so the single-column specialized queues are unaffected; the JIT comparator is bypassed when a truncated cursor participates. - `ReadFromMergeTree::readInOrder` reads index values for the whole used sorting-key prefix instead of only the conversion inputs. The in-order merges inside the read step sort by the full prefix, and index values are exact bounds for it even after filtering (a filter only removes rows), so these merges always see fully covered virtual rows. - `ReadFromMergeTree::requestReadingInOrder` drops the conversion only when a re-request makes it unsound: a prefix narrower than the one it was built for (the conversion could lose its inputs), or one not fully backed by the primary index. - `setVirtualRowConversions` builds the conversion with `project_inputs` so a raw index column cannot shadow a same-named conversion output when the merge looks sort columns up by name (matters with the old analyzer). This handles all key types uniformly — e.g. a descending `String` column after a skipped key keeps the optimization even though the type has no greatest value to pad with. Verified on release and debug builds (the boundary assertion is compiled only into debug builds): the previously aborting repros now return correct results, and `EXPLAIN` keeps `Virtual row conversions` for the widened and skipped-key reads. The test is based on the one from https://github.com/ClickHouse/ClickHouse/pull/110725, extended with checks that the optimization stays enabled, a descending skipped-key case, a `String` case in both directions, and a fixed key kept in `ORDER BY` (still disabled, as before).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111457",
          "createdAt": "2026-07-22T18:29:12Z",
          "updatedAt": "2026-08-13T17:54:10Z",
          "timestamp": "2026-08-13T17:54:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "vdimir",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:298742597270acb706cb",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114220",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114220",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Backport #113291 to 26.6: Fix for virtual row is not being applied in some cases",
          "text": "Original pull-request https://github.com/ClickHouse/ClickHouse/pull/113291 This pull-request is a last step of an automated backporting. Treat it as a standard pull-request: look at the checks and resolve conflicts. Merge it only if you intend to backport changes to the target branch, otherwise just close it. ### The PR source The PR is created in the [CI job](https://github.com/ClickHouse/ClickHouse/actions/runs/31425800507/job/93577076250)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114220",
          "createdAt": "2026-08-10T20:07:09Z",
          "updatedAt": "2026-08-13T17:54:03Z",
          "timestamp": "2026-08-13T17:54:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "pr-bugfix",
            "pr-backport"
          ],
          "author": "robot-clickhouse-ci-2",
          "state": "open",
          "assignees": [
            "vdimir",
            "Avogar"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8d8ea6ecb304b8a5d275",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114672",
        "event": "discovered",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114672",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Do not hold the DDL guard during TRUNCATE",
          "text": "`TRUNCATE TABLE` held the table's DDL guard while removing data, and truncate can wait for running merges or for other replicas to process `DROP_RANGE`. Any DDL on that name — including background threads that need it — was blocked for that whole time. Release the guard before `IStorage::truncate` for databases with UUIDs, matching `ALTER TABLE ... DROP PARTITION`, which never took it. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `TRUNCATE TABLE` no longer blocks concurrent `DROP`, `RENAME` and other DDL queries on the same table while the data is being removed.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114672",
          "createdAt": "2026-08-13T17:52:48Z",
          "updatedAt": "2026-08-13T17:53:50Z",
          "timestamp": "2026-08-13T17:53:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-improvement"
          ],
          "author": "evillique",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:b943772fc032a3a731f6",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114668",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114668",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Stop Parquet background reads before releasing the format's read buffer",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/114612 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fixed heap memory corruption when reading Parquet through an input format that owns its read buffer, for example a dictionary with `SOURCE(FILE(... format 'Parquet'))`. Background prefetch and decode tasks could still read and write through the buffer after the pipeline released it, which could abort the server. ### Description Closes #114612, reported by @ PedroTadim, who asked me to go ahead in https://github.com/ClickHouse/ClickHouse/issues/114612#issuecomment-5278451991. A `ReadBuffer` handed to an input format via `addBuffer()` lives in the format's `owned_buffers`. `ISource::work()` calls `onFinish()` on both the clean and the exception path, reaching `IInputFormat::resetReadBuffer()`, which clears `owned_buffers` and destroys the buffer. `ParquetV3BlockInputFormat` did not override that hook, and `Parquet::Prefetcher` holds a non-owning `SeekableReadBuffer *` to the same buffer, so background tasks kept using a destroyed object. It is not only a bad read: the report on the issue is a 1 MiB **write** into freed heap through `ReadBuffer::next()`, silent on a release build. `~Prefetcher()` does the right handshake, but only at format destruction, later than `onFinish()`; that window is the bug. The fix overrides `resetReadBuffer()` to drain background tasks before the base class releases the buffers, mirroring `ParallelParsingInputFormat::onFinish()`. Three details are forced by the surrounding code: the hook is `resetReadBuffer()`, not `onFinish()`, since it frees the buffers and has a second caller in `StreamingFormatExecutor`; the reader is drained but kept alive, because `getMatchedBuckets()` reads row group metadata after exhaustion; and `ReadManager` is drained before `Prefetcher`, because decode tasks re-enter `readSync` inline via `getRangeData()`. Sibling teardown paths: `resetParser()` already destroys the reader before delegating, `onCancel()` cancels it, and the reuse route through `StreamingFormatExecutor` calls `resetParser()` on every exit, so a drained reader is never reused. No fuzzer needed: `ReadBufferFromFile` leaves `use_pread` false, so a local file takes `SeekAndRead`, and any `CREATE DICTIONARY ... SOURCE(FILE(... format 'Parquet'))` whose load throws reaches it. Validated on ASAN: the added test fails 7/7 unfixed, passes 11/11 fixed; a standalone reproducer 31/33 unfixed, 0/35 fixed; a negative control with only the override reverted reddens at the unfixed rate. Parquet and polygon-dictionary suites gave an identical failing set on both binaries (187 each); 50 randomized-settings repeats passed clean.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114668",
          "createdAt": "2026-08-13T16:50:49Z",
          "updatedAt": "2026-08-13T17:53:37Z",
          "timestamp": "2026-08-13T17:53:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "pr-bugfix",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ff420908aefb00196929",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:101791",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:101791",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "In case of trivial views, push whole outer query to shards.",
          "text": "### Changelog category: - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): In case of trivial views over distributed table push whole outer query to shards. ### Documentation entry for user-facing changes - [ ] Documentation is written (mandatory for new features) ### Description / Proposed Solution When a VIEW is defined over a Distributed table, ClickHouse traditionally executes it on the shards without enclosing outer query. This means filters and expressions declared in the outer query are evaluated on the coordinator after pulling raw data from shards. For views whose body is a plain SELECT (column references, *, or arbitrary expressions — but no aggregation, grouping, ordering, joins, window functions, or scalar subqueries) over a single Distributed table, we can do better: inline the view body as a subquery and hand the whole thing to StorageDistributed. Each shard then receives the full outer query with the view body inlined, evaluates it against its local table, and only ships the result back. A view qualifies as \"trivial\" if its inner query: - Has a single SELECT (no UNION) - Selects only column references, *, or expressions — but no window functions (require the full dataset) and no scalar subqueries in the SELECT list - Has no WITH, PREWHERE, GROUP BY, HAVING, QUALIFY, ORDER BY, LIMIT, LIMIT BY, DISTINCT, or ARRAY JOIN - Has no subqueries in the WHERE clause - Reads from exactly one table with no joins, no table functions, no FINAL, no SAMPLE - Is not a parameterized view and does not use SQL SECURITY DEFINER **The optimization can be disbaled by setting (enabled by default):** ``` SET optimize_trivial_view_pushdown_to_distributed = 0; ``` ### Example: Env setup: ``` create table x engine = MergeTree ORDER BY tuple() AS SELECT intDiv(number,100000) as a, number as b FROM numbers(1000000000); SET prefer_localhost_replica = 0; CREATE TABLE x_dist AS x ENGINE = Distributed(test_cluster_two_shards_localhost, currentDatabase(), x); CREATE VIEW v_computed AS SELECT a + 1 AS x, b AS y FROM x_dist WHERE a != 0; ``` Performance: ``` :) SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: 6baedc55-c8e7-4b2b-9946-b0d828abaf25 ┌─plus(a, 1)─┬──────────sum(b)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └────────────┴─────────────────┘ 1 row in set. Elapsed: 17.199 sec. Processed 2.00 billion rows, 32.00 GB (116.29 million rows/s., 1.86 GB/s.) Peak memory usage: 38.58 MiB. :) SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 7c4f2854-3de2-4d40-b960-efcc44a7b26d ┌─────x─┬──────────sum(y)─┐ 1. │ 10000 │ 199989999900000 │ -- 199.99 trillion └───────┴─────────────────┘ 1 row in set. Elapsed: 16.497 sec. Processed 2.00 billion rows, 32.00 GB (121.24 million rows/s., 1.94 GB/s.) Peak memory usage: 38.82 MiB. ``` Plan: ``` :) explain SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1; EXPLAIN SELECT a + 1, sum(b) FROM x_dist WHERE a != 0 GROUP BY a + 1 ORDER BY sum(b) DESC LIMIT 1 Query id: e2d32d5f-7024-4e6a-b686-45c7c0d5f2ef ┌─explain──────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression (Before GROUP BY) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression (Before GROUP BY) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └──────────────────────────────────────────────────────────────────────────────────┘ :) explain SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1; EXPLAIN SELECT x, sum(y) FROM v_computed GROUP BY x ORDER BY sum(y) DESC LIMIT 1 Query id: 88213305-46a5-493e-862c-a80c675c9452 ┌─explain───────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ 1. │ Expression (Project names) │ 2. │ Limit (preliminary LIMIT) │ 3. │ Sorting (Sorting for ORDER BY) │ 4. │ Expression ((Before ORDER BY + Projection)) │ 5. │ MergingAggregated │ 6. │ Union │ 7. │ Aggregating │ 8. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 9. │ Expression ((WHERE + Change column names to column identifiers)) │ 10. │ ReadFromMergeTree (default.x) │ 11. │ Aggregating │ 12. │ Expression ((Before GROUP BY + (Change column names to column identifiers + (Project names + Projection)))) │ 13. │ Expression ((WHERE + Change column names to column identifiers)) │ 14. │ ReadFromMergeTree (default.x) │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ``` <!--- Directly edit documentation source files in the \"docs\" folder with the same pull-request as code changes or Add a user-readable short description of the changes that should be added to docs.clickhouse.com below. At a minimum, the following information should be added (but add more as needed). - Motivation: Why is this function, table engine, etc. useful to ClickHouse users? - Parameters: If the feature being added takes arguments, options or is influenced by settings, please list them below with a brief explanation. - Example use: A query or command. --> <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes query planning/execution for a subset of views over `Distributed` tables and touches access checks/row policy enforcement and SQL SECURITY semantics, which can affect correctness and security-sensitive behavior. > > **Overview** > Adds a new default-on setting `optimize_trivial_view_pushdown_to_distributed` to inline *trivial* views over `Distributed` tables and push the full outer query down to shards, reducing coordinator-side filtering/processing and network transfer. > > Implements planner rewrites to swap the view table expression with an analyzed subquery, merge `FINAL`/`SAMPLE` modifiers, and preserve semantics by suppressing pushdown when the outer query contains non-deterministic functions, while also explicitly handling SQL SECURITY modes, row-policy injection/logging, and column-pruned privilege checks. > > Extends integration/stateless tests to cover modifier propagation, non-determinism suppression, row-policy enforcement, SQL SECURITY behavior, and interactions with `max_rows_to_read_leaf`. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 583e1e6c0e8e25081391d7a07af086c6f9888c6f. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/101791",
          "createdAt": "2026-04-04T19:32:01Z",
          "updatedAt": "2026-08-13T17:53:08Z",
          "timestamp": "2026-08-13T17:53:08Z",
          "metrics": {
            "reactions": 3,
            "comments": 16
          },
          "labels": [
            "pr-performance",
            "can be tested"
          ],
          "author": "simonmichal",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:771efbbd6adf4b613c05",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109946",
        "event": "discovered",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109946",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix propagation of settings in `accurateCastOrDefault`",
          "text": "### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix propagation of settings in `accurateCastOrDefault`. Closes https://github.com/ClickHouse/ClickHouse/issues/109943",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109946",
          "createdAt": "2026-07-09T23:21:47Z",
          "updatedAt": "2026-08-13T17:52:55Z",
          "timestamp": "2026-08-13T17:52:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-bugfix",
            "pr-must-backport"
          ],
          "author": "Avogar",
          "state": "open",
          "assignees": [
            "antonio2368"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:ac184e21aea788973a01",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113505",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113505",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "S3 tables engine",
          "text": "### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): S3 tables engine catalog for datalakes. Same as https://github.com/ClickHouse/ClickHouse/pull/103220, but with working INSERT",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113505",
          "createdAt": "2026-08-05T14:34:19Z",
          "updatedAt": "2026-08-13T17:52:39Z",
          "timestamp": "2026-08-13T17:52:39Z",
          "metrics": {
            "reactions": 3,
            "comments": 2
          },
          "labels": [
            "pr-feature"
          ],
          "author": "scanhex12",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:dedc7f1d5c587abff324",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114525",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114525",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Optimize merges of the text index",
          "text": "### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Improved performance of merges of text indexes. ### Additional context A few optimizations: - The main one: reducing overhead on deserialization of embedded and small postings caused by the allocation of the bitmap - Removed unneeded conversion to roaring bitmap on build of the output posting list - Used specialized sort cursor and batch sorting strategy for merging of text index segments",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114525",
          "createdAt": "2026-08-12T17:15:12Z",
          "updatedAt": "2026-08-13T17:52:32Z",
          "timestamp": "2026-08-13T17:52:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "CurtizJ",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4c159f38c9f800a04686",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114670",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114670",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: expand the Managed Postgres autoscaling documentation",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required) ## Summary - Expand the one-line Autoscaling section in the Managed Postgres scaling docs with the behavior sourced from the Ubicloud codebase: the 85% storage notification, the 90% automatic scale-up, and the 95% maintenance-window bypass - Document what autoscaling means for the cutover and client connections: same process as a manual instance change, connections dropped and in-flight transactions rolled back, DNS repointed to the new primary under the same hostname - Document read-only mode: free-space trigger and recovery thresholds per disk size, the error writes receive, and automatic recovery after the scale-up - Add a worked example scenario for a 1024 GB instance scaling to 2048 GB",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114670",
          "createdAt": "2026-08-13T16:59:03Z",
          "updatedAt": "2026-08-13T17:52:13Z",
          "timestamp": "2026-08-13T17:52:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "amogiska",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fad7f1f7c4583c692a02",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:109004",
        "event": "discovered",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:109004",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Assign merges for all partitions at once for OPTIMIZE FINAL",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/46770 For a non-replicated `MergeTree` table, `OPTIMIZE TABLE ... FINAL` without an explicit partition used to process partitions one by one: it selected and ran the merge for one partition, waited for it to finish, and only then moved on to the next one. On a table with many partitions this serialized all the work and only ever showed a single merge in `system.merges`. Now the per-partition merges are assigned and executed in parallel, so all partitions are merged at once. The degree of parallelism is bounded by the configured background merge concurrency (`background_pool_size` * `background_merges_mutations_concurrency_ratio`). This mirrors how `StorageReplicatedMergeTree::optimize` already assigns a merge per partition and then waits for all of them. Notes: - Tables inside an explicit transaction keep the sequential path, because parallel merges would otherwise share a single transaction object that is not made for concurrent use. - The wait for already-running merges inside `selectPartsToMerge` (the `OPTIMIZE FINAL` path) is now scoped to the current partition: merges in other partitions cannot prevent selecting all the parts of this one, and waiting for them would needlessly serialize the parallel assignment. Verified on a local build: on a table with 8 partitions, `OPTIMIZE TABLE ... FINAL` now runs 8 merges concurrently (observed via `system.merges`) instead of one at a time, and every partition is still correctly merged into a single part (checked for `MergeTree`, `OPTIMIZE ... FINAL DEDUPLICATE`, and `ReplacingMergeTree`, plus the `optimize_skip_merged_partitions` / `optimize_throw_if_noop` no-op paths and two concurrent `OPTIMIZE FINAL` queries). ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `OPTIMIZE TABLE ... FINAL` on a non-replicated `MergeTree` table now assigns and runs the merges for all partitions at once instead of processing them one partition at a time.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/109004",
          "createdAt": "2026-06-30T23:57:47Z",
          "updatedAt": "2026-08-13T17:51:12Z",
          "timestamp": "2026-08-13T17:51:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-performance"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:db1e55a07a3efbbd15a4",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114466",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114466",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Docs: internationalize master",
          "text": "### Changelog category (leave one): - Documentation (changelog entry is not required)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114466",
          "createdAt": "2026-08-12T10:49:59Z",
          "updatedAt": "2026-08-13T17:50:49Z",
          "timestamp": "2026-08-13T17:50:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 47
          },
          "labels": [
            "pr-documentation"
          ],
          "author": "locadex-agent[bot]",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:831adbebe8109f8d2ed1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114645",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114645",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Speed up `IN (subquery)` set building by pre-deduplicating each `MergeTree` partition independently",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> Related: https://github.com/ClickHouse/ClickHouse/pull/108326 Related: https://github.com/ClickHouse/ClickHouse/pull/105126 The set for `IN (subquery)` is built by a single `CreatingSetsTransform`: all streams of the subquery are merged into one and every row is hashed serially, no matter how many threads read the data. If the partition expression of the subquery's table is a function of the subquery's output columns (the set is keyed on all of them), the reading will now emit each partition through a single port and each stream is deduplicated independently before the filling transform. Because a key then lives in exactly one stream, per-stream deduplication is complete, and the single filling transform only hashes unique rows — the serial part of the build shrinks from all rows to distinct rows, and the deduplication itself runs in parallel. ```sql CREATE TABLE t (a UInt64) ENGINE = MergeTree ORDER BY tuple() PARTITION BY a % 8; INSERT INTO t SELECT number % 1000000 FROM numbers(100000000); OPTIMIZE TABLE t FINAL; EXPLAIN PIPELINE SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t) SETTINGS allow_creating_set_partitions_independently = 1, max_threads = 8; ``` ```response (CreatingSets) DelayedPorts 9 → 8 (Expression) ExpressionTransform × 8 (Aggregating) Resize 1 → 8 AggregatingTransform (Expression) ExpressionTransform (Filter) FilterTransform (ReadFromSystemNumbers) NumbersRange 0 → 1 (CreatingSet) CreatingSetsTransform <- the single filling transform now hashes ~1M unique rows instead of 100M Resize 8 → 1 DistinctTransform × 8 <- new: parallel pre-deduplication on partition-disjoint streams (Expression) ExpressionTransform × 8 (ReadFromMergeTree) MergeTreeSelect(pool: ReadPoolInOrder, algorithm: InOrder) × 8 0 → 1 <- per-partition reading (8 partitions → 8 streams) ``` On the table above (100M rows, 1M distinct keys, 8 balanced partitions; 64-core machine, average of 3 runs after 2 warm-ups): | query | off | on | speedup | |------------------------------------------------------------|--------|--------|---------| | `SELECT count() FROM numbers(10) WHERE number IN (SELECT a FROM t)` | 0.643s | 0.113s | 5.7× | ### Changelog category (leave one): - Performance Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Speed up set building for `IN (subquery)` on partitioned `MergeTree` tables by keeping each partition's rows within a single stream and deduplicating each stream independently, so the single set-filling transform — previously hashing every row serially — only sees unique rows. This applies when the partition expression is a deterministic function of the subquery's output columns. The optimization is not applied when the largest partition holds more than twice the rows of the average partition; the new setting `force_creating_set_partitions_independently` (disabled by default) bypasses this check. Controlled by the new setting `allow_creating_set_partitions_independently` (enabled by default).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114645",
          "createdAt": "2026-08-13T14:04:03Z",
          "updatedAt": "2026-08-13T17:50:48Z",
          "timestamp": "2026-08-13T17:50:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "pr-performance"
          ],
          "author": "nihalzp",
          "state": "open",
          "assignees": [
            "yariks5s"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4129e22267f421ad45a2",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113983",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113983",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Allowlist the expected FileLog bad-path reattach error in the upgrade check",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Related: https://github.com/ClickHouse/ClickHouse/pull/113781 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description `Upgrade check (amd_release)` intermittently fails its `Error message in clickhouse-server.log` sub-test on one benign line: ``` <Error> StorageFileLog (test_1.filelog_bad_path_attach): The absolute data path should be inside `user_files_path`(/var/lib/clickhouse/user_files/) ``` No product defect: the server starts, nothing crashes, no data is affected. Root cause. `04202_filelog_attach_path_outside_user_files` ATTACHes a FileLog table whose path is outside `user_files_path`. `ATTACH` is `LoadingStrictnessLevel::ATTACH` (2), which is `>= SECONDARY_CREATE` (1), so the constructor takes the relaxed branch at `src/Storages/FileLog/StorageFileLog.cpp:195-198`: it logs at `<Error>` and returns instead of throwing `BAD_ARGUMENTS`. That branch is deliberate and is what the test covers, since refusing to load at reattach time would break server startup. The table then outlives the test: stress threads run with a fixed `--database=test_N` (`ci/jobs/scripts/stress/stress.py`), and `clickhouse-test` skips its per-test teardown whenever `--database` is set (`need_cleanup = not args.database`), so that shared database is never dropped. The upgrade restart re-attaches the table, the relaxed branch fires again, and the line lands in the scanned log, where the post-restart scrub in `tests/docker_scripts/upgrade_runner.sh` had no entry for it. Hence the intermittency: `04202` must land on a fixed-database thread. Change. One `grep -av` entry in that scrub's existing secondary pipe, plus a short rationale comment next to the sibling entries. The pattern requires the fixture table name and the message together, and (bare parens are literals in BRE) the `StorageFileLog (db.table):` prefix shape. No source change, no test change. Validation. The scan pipeline, extracted verbatim from the runner, was run under GNU grep 3.11 against the failing run's own 19.9 MB `clickhouse-server.upgrade.log`. With the entry the artifact is empty; with it deleted the output is byte-identical to the 189-byte `upgrade_error_messages.txt` CI produced, so the sub-test flips `FAIL` to `OK`. Eight negative controls still surface, including a table whose name merely ends with the fixture name (`prod.other_filelog_bad_path_attach`), which the required `.` separator keeps visible.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113983",
          "createdAt": "2026-08-08T21:17:46Z",
          "updatedAt": "2026-08-13T17:50:47Z",
          "timestamp": "2026-08-13T17:50:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "manual approve",
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:aa138eeca86cb5a13abd",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:96130",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:96130",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Randomize tests with DETACH/ATTACH table before query execution",
          "text": "Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that enable randomly detaching and reattaching tables used in a query before its execution. This is a testing-only feature designed to find bugs related to table reattachment. Before executing a query, the system collects all tables referenced in the AST, and for each eligible table (stores data on disk, supports detaching, has no action locks or dependencies), it performs a `DETACH` followed by `ATTACH`. Changes: - Add `supportsDetachingTables` virtual method to `IDatabase` (overridden to `false` for engines that do not support non-permanent `DETACH TABLE`: `DatabaseDictionary`, `DatabaseReplicated`, `DatabaseSQLite`, `DatabaseBackup`, `DatabaseFilesystem`, `DatabaseHDFS`, `DatabaseS3`, `DatabaseURL`, `DatabaseRemote`, `DatabaseDataLake`, `DatabaseMaterializedPostgreSQL`) - Add `has`/`hasAny` methods to `ActionLocksManager` for checking existing locks (skipping expired `weak_ptr` entries) - Add table collection visitor and reattach logic in `executeQuery` (runs after AST validations, process list admission, and external tables initialization; skips `EXPLAIN`, transactions, internal/non-initial queries, and CTE name collisions) - Fix off-by-one in `MergeTreeDeduplicationLog::dropOutdatedLogs` (don't drop the active log) and add `sync` call in shutdown - Add `no-random-detach` tag to tests incompatible with this feature - Add `--no-random-detach` and `--reattach-tables-probability` options to `clickhouse-test` - Add `02461_reattach_tables` test Continuation of #55943. Continuation of #42336 ### Changelog category (leave one): - Build/Testing/Packaging Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Add `reattach_tables_before_query_execution` and `reattach_tables_before_query_execution_probability` settings that randomly `DETACH` and `ATTACH` tables used in a query before its execution. This is a testing-only feature that helps find reattachment-related bugs. ### Documentation entry for user-facing changes - [x] Documentation is written (mandatory for new features) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Introduces new pre-execution mutations (internal `DETACH`/`ATTACH`) in `executeQuery`, which can affect table availability and concurrency behavior if enabled; guarded by new experimental settings but touches core query execution paths. > > **Overview** > Adds experimental settings `reattach_tables_before_query_execution` and `..._probability` to optionally **DETACH and ATTACH back** eligible tables referenced by a query immediately before execution, including AST table discovery that accounts for CTE scoping, privilege checks, dependency/lock checks, and safety skips (e.g. `system`, non-disk storages, dynamic-structure columns, transactions, `EXPLAIN`, internal/non-initial queries). > > Extends `IDatabase` with `supportsDetachingTables()` and marks multiple database engines as not supporting non-permanent detach; adds `ActionLocksManager::has/hasAny` helpers to avoid detaching tables with active action locks. Updates the test runner and stress tooling to randomize this behavior (with `--no-random-detach` and probability control), adds a new `02461_reattach_tables` test, and tags many existing tests to opt out where DETACH/ATTACH would add flakiness/overhead. Also fixes `MergeTreeDeduplicationLog` cleanup to avoid dropping the active log and ensures writer `sync()` on shutdown. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit a369371ff61cb1934815ea1e8caf7debd83c98c4. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/96130",
          "createdAt": "2026-02-05T23:27:47Z",
          "updatedAt": "2026-08-13T17:49:48Z",
          "timestamp": "2026-08-13T17:49:48Z",
          "metrics": {
            "reactions": 2,
            "comments": 91
          },
          "labels": [
            "pr-build"
          ],
          "author": "alexey-milovidov",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4cf2c87402d7699d73f0",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:107865",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:107865",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Add per-user filesystem cache disk usage metrics",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/105020 ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added opt-in Prometheus gauges for current filesystem cache usage (`filesystem_cache_size_bytes` and `filesystem_cache_elements`), labeled by `cache_name` and `user_id`. The gauges are exposed through `system.dimensional_metrics` and the Prometheus endpoint and controlled by the `expose_prometheus_cache_usage_metrics_per_user` cache setting, which is disabled by default. --- ### Description Adds two gauges exposed through `system.dimensional_metrics` and the Prometheus endpoint: | Name | Labels | |---|---| | `filesystem_cache_size_bytes` | `cache_name`, `user_id` | | `filesystem_cache_elements` | `cache_name`, `user_id` | **Motivation.** Operators can identify which users currently occupy filesystem cache space, measured both in bytes and in file segments, for each cache. **Mechanism.** Each enabled cache owns a `FileCacheUsageTracker` containing shared per-user atomic counters. Main cache-priority entries retain the corresponding counters and update them together with the existing cache size and element accounting. `ServerAsynchronousMetrics` periodically obtains a per-user snapshot through `getUsageStatPerClient` and updates the dimensional gauges. This keeps `DimensionalMetrics` updates out of cache mutation paths. Composite LRU, SLRU, and split-cache priorities share the same tracker, including during SLRU queue transitions. Counters use shared ownership so inactive users can be reclaimed safely. When a user no longer has cache entries and its counters are zero, the next snapshot removes it from the tracker. Stale dimensional metric label combinations are also removed. **Sampling.** The metrics reflect the most recent asynchronous metrics update. **Cardinality.** In-process cardinality is bounded by users currently retained by each cache, plus labels awaiting the next asynchronous metrics update. Stale `(cache_name, user_id)` label combinations are removed. The feature remains disabled by default because enabling per-user metrics can still create significant Prometheus time-series cardinality. ### Documentation entry: - [x] Documentation is updated.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/107865",
          "createdAt": "2026-06-18T14:22:58Z",
          "updatedAt": "2026-08-13T17:48:49Z",
          "timestamp": "2026-08-13T17:48:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "sacheendra",
          "state": "open",
          "assignees": [
            "kssenii"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:669c9c40abdc3fd8ab22",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:111394",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:111394",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fsync backup files and directories when writing a backup to local disk",
          "text": "<!-- Closes: https://github.com/ClickHouse/ClickHouse/issues/111320 --> ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): `BACKUP ... TO File(...)` / `Disk(...)` now fsyncs the backup data files, the `.backup` manifest and the containing directories before reporting `BACKUP_CREATED`, so an acknowledged backup to local storage survives power loss. Controlled by the new backup setting `fsync_backup_files` (default `true`). Object-storage destinations (`S3`/`Azure`) are unaffected. ### Description Fixes #111320. `BACKUP ... TO File()/Disk()` returned `BACKUP_CREATED` without issuing any `fsync`/`fdatasync` at the destination: not the data files, not the `.backup` manifest, and not the destination directories (there was no `fsync` anywhere in `src/Backups/`). On power loss after the acknowledgement the backup could be lost entirely or left torn, even though `BACKUP_CREATED` is exactly what an operator relies on before dropping the source data. Object-storage destinations were already durable (a completed upload is persisted server-side); only local `File()`/`Disk()` were affected. Report URL: https://github.com/ClickHouse/ClickHouse/issues/111320 (reproduced 3/3 with a `dm-flakey` power-loss simulation). Fix, gated on the new backup setting `fsync_backup_files` (default `true`), following the durability audit family (#68958 -> #111346, #111269 -> #111335): - Two writer hooks with a no-op default on `IBackupWriter`, overridden only by the local `File`/`Disk` writers (`S3`/`Azure`/`Memory`/`Null` inherit the no-op): `syncFileToDisk(file_name)` (fdatasync a written file, covering both the buffered and the native `fs::copy`/`IDisk::copyFile` paths) and `syncDirectoriesToDisk()` (fdatasync every directory the backup created, deepest-first, plus the backup root's parent, via `LocalDirectorySyncGuard` / `IDisk::getDirectorySyncGuard`). - Each data file is synced right after it is written in `BackupImpl::writeFile` (safe under the concurrent write path: each call fsyncs its own file). - In `BackupImpl::finalizeWriting` the `.backup` manifest (or, for archives, the archive file) is synced last, after all data files, so a persisted manifest never precedes its payload. Directory syncing runs for every writer, including the internal writers of `BACKUP ON CLUSTER` which write their own data files. Verified locally with ProfileEvents: `fsync_backup_files=1` issues `FileSync`/`DirectorySync` for the whole backup (data files + manifest + every nested directory); `fsync_backup_files=0` issues none (matching the previous behavior); the backup still restores correctly. Regression test `tests/queries/0_stateless/04412_backup_to_file_fsync.sh` asserts, via the `FileSync`/`DirectorySync` ProfileEvents of the `BACKUP` query in `system.query_log`, that the fsyncs are issued when `fsync_backup_files=1` and are absent when `fsync_backup_files=0`.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/111394",
          "createdAt": "2026-07-22T13:30:01Z",
          "updatedAt": "2026-08-13T17:48:43Z",
          "timestamp": "2026-08-13T17:48:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 13
          },
          "labels": [
            "pr-bugfix",
            "manual approve",
            "can be tested"
          ],
          "author": "groeneai",
          "state": "open",
          "assignees": [
            "jkartseva"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:5e39108eb5adb8c7291a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:113691",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:113691",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix `theilsU` window state returning noise when the frame's first argument is constant",
          "text": "Related: https://github.com/ClickHouse/ClickHouse/pull/80373 Related: https://github.com/ClickHouse/ClickHouse/pull/93384 ### Changelog category (leave one): - Bug Fix (user-visible misbehavior in an official stable release) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Fix `theilsU` over a window frame returning an arbitrary value instead of 0 when the first argument is constant within the frame. ### Description `TheilsUWindowData::getResult` (the window-optimized state introduced in https://github.com/ClickHouse/ClickHouse/pull/93384) computes the entropy `H(A)` from cached incremental `Σ n·log n` sums. When the first argument is constant within the frame, the true `H(A)` is zero, and the computed value is pure rounding noise from the incremental updates. The code compared it against exact zero, so a tiny positive noise value passed the check, and `1 - H(A|B) / H(A)` then divided noise by noise: in debug builds this tripped the sanity check as the exception `Logical error: 'res < 1.0 + 1e-4'`, and in release builds the function could return an arbitrary value in $[0, 1]$ instead of 0. The exact (non-window) code path recomputes the entropies from the count maps, where a constant column gives `log(1) = 0` exactly, so it is not affected. The fix compares `H(A)` against an error bound proportional to `N · ε · log N` instead of exact zero, and widens the sanity-check tolerance by the same relative amount so that near-threshold frames do not trip it either. Found by the AST fuzzer on an unrelated PR (it hit https://github.com/ClickHouse/ClickHouse/pull/80373 and https://github.com/ClickHouse/ClickHouse/pull/107667): [CI report](https://s3.amazonaws.com/clickhouse-test-reports/json.html?PR=80373&sha=6c271049214aa5a94bd9a5f12fb27a9ffa75648f&name_0=PR&name_1=AST%20fuzzer%20%28amd_debug%29).",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/113691",
          "createdAt": "2026-08-06T15:21:02Z",
          "updatedAt": "2026-08-13T17:48:39Z",
          "timestamp": "2026-08-13T17:48:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [
            "pr-bugfix"
          ],
          "author": "alexey-milovidov",
          "state": "closed",
          "assignees": [
            "nihalzp"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4db37ea745b1a0d973c3",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114628",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "text",
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114628",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Wait for DETACH DATABASE to release tables in three integration tests",
          "text": "<!-- Related: https://github.com/ClickHouse/ClickHouse/issues/93064 --> ### Changelog category (leave one): - CI Fix or Improvement (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... ### Description Related: https://github.com/ClickHouse/ClickHouse/issues/93064 Three integration tests intermittently fail with `Code: 219 ... Database <db> cannot be detached, because some tables are still in use. Retry later.` thrown from `DatabaseAtomic::assertCanBeDetached` (`src/Databases/DatabaseAtomic.cpp:521`). Root cause: non-SYNC `DETACH DATABASE` is best-effort by contract, and these tests assume it is atomic. `assertCanBeDetached` throws if any table of the database still has a live refcount. The engine already has the deterministic wait, `waitDetachedTableNotInUse`, but `executeToDatabaseImpl` calls it only under `query.sync` (`src/Interpreters/InterpreterDropQuery.cpp:779-790`). Stateless CI installs `tests/config/users.d/database_atomic_drop_detach_sync.xml` (`database_atomic_wait_for_drop_and_detach_synchronously=1`); the integration helpers install no such profile and run at the server default of 0. Hence the failures are confined to the integration suite, and 16 of the 17 integration occurrences in 180 days are sanitizer builds, where the holder's window is wider. Code 219 is intended behaviour of the asynchronous form and is pinned in-tree: `01107_atomic_db_detach_attach.sh:19` sets the setting to 0 to provoke it and asserts it. So this is a test-side defect, and the change asks for the wait the engine already implements rather than altering it. I added `SYNC` to the five `DETACH DATABASE` statements with observed CI failures: `test_drop_replica` (11 hits/180d), `test_replicated_table_structure_alter` (4), and `test_drop_database_replica:192` (2, both stacks confirm that line is the thrower). All 21 non-SYNC sites under `tests/integration/` were enumerated; the other 16 are excluded because they have zero CIDB hits in 180 days, already pass `database_atomic_wait_for_drop_and_detach_synchronously`, or sit inside the existing `detach_database_with_retry` helper, whose docstring documents a different holder. Validation: with a holder injected the way `01107` does it, the DETACH returns 219 without `SYNC` and succeeds with it, 10/10 on `Atomic` and on `Replicated`. The three tests pass 42/42 runs. Each `DETACH` in `test_drop_replica` measures ~0.115 s over 25 measurements, so the wait costs nothing when no table is held, and it is cancellable and shutdown-aware. CI report for the master failure: https://s3.amazonaws.com/clickhouse-test-reports/json.html?REF=master&sha=91b700711adf0a07172e1046c8d8f013f1e0b82c&name_0=MasterCI&name_1=Integration%20tests%20%28amd_tsan%2C%206%2F6%29 <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1356` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114628",
          "createdAt": "2026-08-13T12:36:13Z",
          "updatedAt": "2026-08-13T17:55:30Z",
          "timestamp": "2026-08-13T17:55:30Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "can be tested",
            "pr-ci"
          ],
          "author": "groeneai",
          "state": "closed",
          "assignees": [
            "PedroTadim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1b04b5ac7549cc8c86cc",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114316",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114316",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Use the vector similarity index for integer reference vectors",
          "text": "Closes: https://github.com/ClickHouse/ClickHouse/issues/112233 Related: https://github.com/ClickHouse/ClickHouse/issues/114291 The reference vector of an ANN query is extracted only when its array type is `Float64`, `Float32` or `BFloat16` and every element is a `Float64` field. An integer literal such as `[1, 2]` is typed `Array(UInt8)`, so `tryUseVectorSearch` bails out and the query silently falls back to a brute-force scan over the whole table, although `[1, 2]` denotes the same point as `[1.0, 2.0]` and `L2Distance` accepts it. `EXPLAIN indexes = 1` shows no `vector_similarity` entry and gives no hint why, so a one-character difference in a literal becomes a sharp performance cliff on large tables. Native integer arrays are now accepted as reference vectors and their elements are converted to `Float64`, which is the type the reference vector is stored in anyway. Added `02354_vector_search_bug112233`, covering unsigned, signed, mixed integer/float, and not-exactly-representable reference vectors, plus an equality check between the integer and float spellings of the same query. ### Changelog category (leave one): - Improvement ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Vector search queries now use the `vector_similarity` index when the reference vector is written as an integer array literal, e.g. `ORDER BY L2Distance(vec, [1, 2])`. Previously such queries silently fell back to a brute-force scan.",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114316",
          "createdAt": "2026-08-11T12:52:06Z",
          "updatedAt": "2026-08-13T17:48:13Z",
          "timestamp": "2026-08-13T17:48:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "pr-improvement",
            "can be tested"
          ],
          "author": "hamidr",
          "state": "open",
          "assignees": [
            "rschu1ze"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:087065e4f0b547432bb1",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:86353",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "metrics"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:86353",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Cascades cost-based optimizer for distributed query plans",
          "text": "A Cascades-style cost-based optimizer that chooses distribution strategies for the multi-stage distributed query plans of #106020. It explores alternatives in a memo (a shared store of equivalent plan fragments) with top-down, goal-directed search and picks the cheapest plan satisfying the required distribution and sorting properties, inserting exchange operators (plan steps that move rows between nodes) as needed. Implemented: - **Join strategies**: shuffle hash join, broadcast hash join (with `ReplicatedRead` — every worker repeats the same read of a small table instead of a network broadcast, assuming shared storage where all workers see the same data), replicated join (a small deterministic join is recomputed identically on every node over such reads, so its result never crosses the network; nested joins compose; `ANY` joins are excluded because the kept row depends on the build order), local join. - **Aggregation strategies**: two-phase (partial + merge), shuffle by group keys, local; `distributed_aggregation_memory_efficient` and `distributed_plan_force_shuffle_aggregation` are honored. - **Top-N**: two-stage distributed top-N (per-node bounded sort, sorted-merge gather, coordinator limit); disabled under `exact_rows_before_limit`, which needs the full row count. - **Read strategies**: parallel N-way read, replicated read, local read. For `FINAL`, #108148 (already in master) taught the rule-based distributed plan to split a `FINAL` read into disjoint primary-key-range buckets where that is safe; Cascades now reuses that machinery, so `FINAL` no longer forces a serial read here either. The coordinator ships each bucket's marks in the `read_bucket` task parameters. - **`IN (subquery)`**: follows the `rewrite_in_to_join` setting like the rest of the planner (the forced join form is removed). In the default set form the set-building subqueries are planned separately and distributed like any other query. - **Properties and enforcers**: distribution (node count, replication, partitioning columns with equivalence classes and the types the keys are cast to before hashing) and sorting; when a plan alternative lacks a required property, an enforcer inserts the step that provides it (`Gather`/`Shuffle`/`Broadcast`/`ScatterExchange`, `Sort`). - **Transformations**: join commutativity (only for semantics-preserving joins: `INNER ALL`, `CROSS`, `SEMI`/`ANY`/`ANTI`; never `ASOF`, and never `ANY` under `join_any_take_last_row`), two-phase aggregation split, two-stage top-N split. - **Cost model**: `work`, `network`, and `sequential` components, each priced as wall-clock per node: a shuffle moves 1/N of the data per node, a broadcast payload is ingested once by every receiver in parallel, and a gather funnels every row through one endpoint, so its transfer stays undivided and pays a per-row cost. A hash-table build counts as parallel work (`parallel_hash` shards it across threads); a fixed per-exchange overhead keeps small inputs local. A table read is priced on its scan volume - the rows the primary key keeps - not on its output estimate, so a filter off the sorting key cannot make a replicated re-read look free. Standalone filters (e.g. `HAVING`) are estimated from column NDVs with join-key equivalence classes; join estimates are clamped to join kind and strictness semantics; exchange costs use per-row byte widths measured from the parts' column sizes (followed through renames, not derived from types). All weights and calibration constants are overridable at query time. What this improves over the rule-based distributed planner, on TPC-H plans. The rule-based planner broadcasts a small table when its read is below `distributed_plan_max_rows_to_broadcast`, but it often cannot size the result of a join, so a small join result (`nation x region`, 5 rows after the region filter) is scattered across nodes, joined there, and shuffled again (repartitioned across nodes) by the next join key. It also often inserts a shuffle at join and aggregation boundaries even when the rows are already divided by the right key. Cascades estimates sizes through joins, knows which partitioning already holds, compares broadcast against shuffle by cost for each join, and recomputes a small deterministic join on every node when that is cheaper than moving its result. On TPC-H SF100 over 8 nodes (same binary, same run window; times are server-side means of the hot runs) the join-heavy queries improve: | Query | Rule-based -> Cascades | What changed in the plan | |---|---|---| | Q21 | 7.66 s -> 5.34 s | The four `supplier x nation` joins compute their small results once and broadcast them, so `lineitem` is not shuffled to meet them. | | Q09 | 4.07 s -> 2.47 s | `part` and `nation` are read in full by every node, so `lineitem` and `supplier` are not shuffled to meet them. | | Q08 | 2.39 s -> 0.99 s | `nation x region` (5 rows) is recomputed by every node; `part` is read in full per node, so `lineitem` is not shuffled to join it. | | Q02 | 2.14 s -> 0.84 s | The `supplier x nation x region` chain is recomputed by every node, so only `partsupp` and `supplier` are shuffled. The top-100 sort becomes two-stage, sending at most 100 rows per node. | | Q05 | 2.05 s -> 1.12 s | The whole dimension side (`orders x customer x nation x region`) is recomputed by every node over full local reads; `lineitem` joins it in place with no shuffle at all. | | Q17 | 2.72 s -> 1.78 s | The small per-part average is broadcast to every node, so the outer 600M-row `lineitem` read is not shuffled. | | Q11 | 0.49 s -> 0.26 s | `supplier x nation` is computed once and broadcast; `partsupp` joins it in place. | | Q12 | 0.76 s -> 0.59 s | The filtered `lineitem` rows (~30K of 600M) are gathered once and broadcast; nothing is shuffled. | The remaining queries change less. Summed over all 22 queries, hot server time drops from 36.8 s to 28.8 s (about 22% lower). The remaining regressions are the `IN (subquery)` queries Q18 (2.43 s -> 2.85 s) and Q20 (2.24 s -> 2.99 s). With the forced join rewrite removed, both run their `IN`s as sets in both modes, and the sets built are identical; the difference is in the plans around them. Cascades picks replicated-read shapes that rescan moderate tables on every node, which loses here for a structural reason: the main reads are filtered by `x IN <set>`, and the set contents do not exist at costing time, so the model cannot credit a shuffle-based plan for how few rows survive the set filter, while the replicated read's full scan is paid regardless. Two follow-ups: a cost-informed choice of the `IN` form, and set-filter selectivity from the subquery's output estimate. `EXPLAIN pretty = 1, estimates = 1` shows the chosen plan with a row estimate and the accumulated cost for each step. Also in this PR, two improvements to the shared bucketed-read machinery (they benefit the rule-based path too): the `FINAL` layer split no longer depends on the coordinator's core count, and a many-partition `FINAL` split groups its layers into the target task count instead of falling back to a serial read. Design, a worked example on TPC-H data (a simplified 3-table query traced through the memo), and current limitations are documented in `src/Processors/QueryPlan/Optimizations/Cascades/ARCHITECTURE.md`. Plan-shape tests cover the actual TPC-H queries (`03836_tpch_join_order_plans`), and focused tests pin the cost-model contracts (e.g. `04869_cascades_read_cost_granule_volume`, `04838_cascades_filter_selectivity`). Disabled by default. Requires the analyzer; remote execution requires the stateless-worker configuration, while `distributed_plan_execute_locally = 1` runs the stages in-process without it: ```sql SET enable_cascades_optimizer = 1, make_distributed_plan = 1; ``` For tests, `param__internal_cascades_cluster_node_count` overrides the cluster size, `param__internal_cascades_cost_config` overrides the cost model configuration, `param__internal_join_table_stat_hints` injects table statistics, `param__internal_cascades_task_limit` lowers the task budget (it can never raise it). Related: https://github.com/ClickHouse/ClickHouse/pull/106020 ### Changelog category (leave one): - Experimental Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Added an experimental Cascades cost-based optimizer for distributed query plans, enabled by `enable_cascades_optimizer = 1` together with `make_distributed_plan = 1`. It chooses between shuffle, broadcast, replicated, and local join strategies, two-phase, shuffle, and local aggregation, two-stage distributed top-N, and parallel and replicated reads by estimated cost, inserting exchange operators as needed. ### Documentation entry for user-facing changes - [ ] Documentation written in [/docs](https://github.com/ClickHouse/ClickHouse/tree/master/docs)",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/86353",
          "createdAt": "2025-08-28T11:27:09Z",
          "updatedAt": "2026-08-13T17:45:48Z",
          "timestamp": "2026-08-13T17:45:48Z",
          "metrics": {
            "reactions": 22,
            "comments": 7
          },
          "labels": [
            "pr-experimental"
          ],
          "author": "davenger",
          "state": "open",
          "assignees": [
            "novikd"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d0a4ed11705392ceaced",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:108653",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:108653",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Support `GROUPS` frame mode for window functions",
          "text": "<!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> The query below applies the same `1 PRECEDING AND 1 FOLLOWING` bounds as a `ROWS`, a `RANGE`, and a `GROUPS` (this PR) frame. The `order` column contains duplicate and non-consecutive values, so the three modes cover different rows: ```sql CREATE TABLE wf_frame_groups (`order` UInt64, value UInt64) ENGINE = Memory; INSERT INTO wf_frame_groups FORMAT Values (10, 1), (10, 2), (20, 3), (30, 4), (30, 5); SELECT order, value, groupArray(value) OVER (ORDER BY order ROWS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS rows_frame, groupArray(value) OVER (ORDER BY order RANGE BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS range_frame, groupArray(value) OVER (ORDER BY order GROUPS BETWEEN 1 PRECEDING AND 1 FOLLOWING) AS groups_frame FROM wf_frame_groups ORDER BY order, value; ``` ```response ┌─order─┬─value─┬─rows_frame─┬─range_frame─┬─groups_frame─┐ │ 10 │ 1 │ [1,2] │ [1,2] │ [1,2,3] │ │ 10 │ 2 │ [1,2,3] │ [1,2] │ [1,2,3] │ │ 20 │ 3 │ [2,3,4] │ [3] │ [1,2,3,4,5] │ │ 30 │ 4 │ [3,4,5] │ [4,5] │ [3,4,5] │ │ 30 │ 5 │ [4,5] │ [4,5] │ [3,4,5] │ └───────┴───────┴────────────┴─────────────┴──────────────┘ ``` Each mode interprets the bounds differently: - `ROWS` counts physical rows, so the frame is at most three adjacent rows: the current row plus one on each side. - `RANGE` counts `order` values, so `1 PRECEDING` and `1 FOLLOWING` cover rows whose `order` is within 1 of the current row's. With gaps of 10, no neighbouring row qualifies, so the frame holds only the rows that share the current `order`. - `GROUPS` (added by this PR) counts peer groups, so `1 PRECEDING` and `1 FOLLOWING` always include the adjacent groups in full, whatever the gaps between `order` values. cc: @cwurm ### Changelog category (leave one): - New Feature ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): Support the `GROUPS` frame mode for window functions (SQL:2011), e.g. `any(price) OVER (PARTITION BY symbol ORDER BY ts GROUPS BETWEEN CURRENT ROW AND 1 FOLLOWING)`. In a `GROUPS` frame the boundaries count whole peer groups — sets of rows that are equal on the `ORDER BY` key — so `N PRECEDING`/`N FOLLOWING` mean `N` peer groups before/after the current row's peer group, rather than physical rows (`ROWS`) or `ORDER BY` value distances (`RANGE`). <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1354` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/108653",
          "createdAt": "2026-06-26T20:45:31Z",
          "updatedAt": "2026-08-13T17:53:45Z",
          "timestamp": "2026-08-13T17:53:45Z",
          "metrics": {
            "reactions": 2,
            "comments": 3
          },
          "labels": [
            "pr-feature"
          ],
          "author": "nihalzp",
          "state": "closed",
          "assignees": [
            "antaljanosbenjamin"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:8304f48510246736e15a",
        "signalId": "github:ClickHouse/ClickHouse:pull_request:114543",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:ClickHouse/ClickHouse:pull_request:114543",
          "source": "github",
          "group": "data-infrastructure",
          "project": "ClickHouse/ClickHouse",
          "kind": "pull_request",
          "title": "Fix spelling of 'prefetches' and 'prefetched' in documentation",
          "text": "Corrected spelling of 'prefetches' and 'prefetched' in multiple sections. <!-- Linked issues and pull requests. Use full GitHub URLs, one relationship per line; delete the lines you don't need. Closes: https://github.com/ClickHouse/ClickHouse/issues/NNNNN (auto-closes the issue when this PR is merged into the default branch) Related: https://github.com/ClickHouse/ClickHouse/pull/NNNNN --> ### Changelog category (leave one): - Documentation (changelog entry is not required) ### Changelog entry (a [user-readable short description](https://github.com/ClickHouse/ClickHouse/blob/master/docs/changelog_entry_guidelines.md) of the changes that goes into CHANGELOG.md): ... <!-- ch-version-info:start --> ### Version info - Merged into: `26.8.1.1353` (included in `26.8` and later) <!-- ch-version-info:end -->",
          "url": "https://github.com/ClickHouse/ClickHouse/pull/114543",
          "createdAt": "2026-08-12T21:18:03Z",
          "updatedAt": "2026-08-13T17:55:27Z",
          "timestamp": "2026-08-13T17:55:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "pr-documentation",
            "can be tested"
          ],
          "author": "linhgiang24",
          "state": "closed",
          "assignees": [
            "tiandiwonder"
          ],
          "change": "updated"
        }
      }
    ]
  }
}
