{
  "schemaVersion": 3,
  "dataset": {
    "version": 3,
    "date": "2026-08-13",
    "group": {
      "id": "data-infrastructure",
      "name": "Data / Messaging / Storage Infrastructure"
    },
    "repository": {
      "id": "redpanda",
      "repo": "redpanda-data/redpanda",
      "name": "Redpanda",
      "keywords": [
        "Redpanda"
      ]
    },
    "context": {
      "repository": "redpanda-data/redpanda",
      "url": "https://github.com/redpanda-data/redpanda",
      "description": "Redpanda is a streaming data platform for developers. Kafka API compatible. 10x faster. No ZooKeeper. No JVM!",
      "homepage": "https://redpanda.com",
      "language": "C++",
      "topics": [
        "containers",
        "cpp",
        "event-driven",
        "go",
        "kafka",
        "kubernetes",
        "microservices",
        "realtime",
        "redpanda",
        "seastar",
        "storage-engine",
        "stream-processing",
        "streaming"
      ],
      "defaultBranch": "dev",
      "stars": 12438,
      "forks": 784,
      "openIssues": 611,
      "archived": false,
      "collectedAt": "2026-08-13T18:02:02.916137+00:00"
    },
    "news": {
      "repository": "redpanda-data/redpanda",
      "collectedAt": "2026-08-13T18:02:02.916137+00:00",
      "latestRelease": {
        "repository": "redpanda-data/redpanda",
        "tag": "v26.1.15",
        "title": "v26.1.15",
        "url": "https://github.com/redpanda-data/redpanda/releases/tag/v26.1.15",
        "publishedAt": "2026-08-07T03:37:58Z",
        "notes": "## Bug Fixes\r\n* Fix the registered config name for  `leader_balancer_node_mute_timeout`. by @WillemKauf in [#31360](https://github.com/redpanda-data/redpanda/pull/31360)\r\n* Fixes a bug in which L0 batches in a cloud topic forgot to preserve `last_offset_delta` in their header, leading to an under-declared last offset which can stall consumers, skip records, or halt exact-offset replication. by @WillemKauf in [#31363](https://github.com/redpanda-data/redpanda/pull/31363)\r\n* Fixes a bug in which topics with `min.compaction.lag.ms` left unconfigured with produced batches holding timestamps in the future would be considered ineligible for compaction by @WillemKauf in [#31459](https://github.com/redpanda-data/redpanda/pull/31459)\r\n* Fixes a bug in which transient `TOPIC_AUTHORIZATION_FAILED` errors and SASL authentication failures were possible during application of a controller snapshot. by @WillemKauf in [#31437](https://github.com/redpanda-data/redpanda/pull/31437)\r\n* Fixes a bug where corrupted storage would not yield a bad CRC in returned record batches. by @andrwng in [#31388](https://github.com/redpanda-data/redpanda/pull/31388)\r\n* Fixes a bug where having a cloud topic read replica on a given cluster would prevent L0 objects on that cluster from being garbage collected. by @andrwng in [#31391](https://github.com/redpanda-data/redpanda/pull/31391)\r\n* Updating a Shadow Link that uses PLAIN authentication no longer fails when the password is omitted; the stored password is preserved. by @r-vasquez in [#31421](https://github.com/redpanda-data/redpanda/pull/31421)\r\n* [#31448](https://github.com/redpanda-data/redpanda/issues/31448) `rpk connect install --connect-version` no longer rejects versions with a\r\nsegment of three or more digits, which had blocked pinning any Redpanda Connect\r\nrelease since 4.100.0. Malformed versions with trailing characters are now\r\nrejected during validation rather than failing at download. by @prakhargarg105 in [#31449](https://github.com/redpanda-data/redpanda/pull/31449)\r\n* `rpk security secrets list` no longer truncates its output at 100 secrets. by @simon0191 in [#31436](https://github.com/redpanda-data/redpanda/pull/31436)\r\n* `rpk shadow create` no longer fails secret-reference validation on clusters\r\nwith more than one page of `REDPANDA_CLUSTER`-scoped secrets. by @simon0191 in [#31436](https://github.com/redpanda-data/redpanda/pull/31436)\r\n* `rpk shadow update` in editor mode now replaces the entire Shadow Link configuration instead of merging changed fields, so list-valued fields (e.g. topic filters) can shrink or be cleared. by @r-vasquez in [#31421](https://github.com/redpanda-data/redpanda/pull/31421)\r\n\r\n## Improvements\r\n* [#31335](https://github.com/redpanda-data/redpanda/issues/31335) Fixes an issue where `/v1/usage` responses could cause oversized allocations for clusters with a large number of Iceberg-enabled topics. by @WillemKauf in [#31337](https://github.com/redpanda-data/redpanda/pull/31337)\r\n* `rpk cluster health` will now display any nodes that may be in maintenance mode. by @alextreichler in [#31352](https://github.com/redpanda-data/redpanda/pull/31352)\r\n\r\n**Full Changelog**: https://github.com/redpanda-data/redpanda/compare/v26.1.14...v26.1.15",
        "highlights": [
          "Bug Fixes",
          "Fix the registered config name for leaderbalancernodemutetimeout. by @WillemKauf in #31360",
          "Fixes a bug in which L0 batches in a cloud topic forgot to preserve lastoffsetdelta in their header, leading to an under-declared last offset which can stall consumers, skip records, or halt exact-offset replication. by @WillemKauf in #3136",
          "Fixes a bug in which topics with min.compaction.lag.ms left unconfigured with produced batches holding timestamps in the future would be considered ineligible for compaction by @WillemKauf in #31459",
          "Fixes a bug in which transient TOPICAUTHORIZATIONFAILED errors and SASL authentication failures were possible during application of a controller snapshot. by @WillemKauf in #31437",
          "Fixes a bug where corrupted storage would not yield a bad CRC in returned record batches. by @andrwng in #31388"
        ],
        "prerelease": false
      },
      "upcoming": [
        {
          "repository": "redpanda-data/redpanda",
          "kind": "milestone",
          "title": "v26.2.x-next",
          "url": "https://github.com/redpanda-data/redpanda/milestone/329",
          "progress": 88,
          "openIssues": 8,
          "closedIssues": 58
        },
        {
          "repository": "redpanda-data/redpanda",
          "kind": "milestone",
          "title": "v26.1.16",
          "url": "https://github.com/redpanda-data/redpanda/milestone/332",
          "progress": 100,
          "openIssues": 0,
          "closedIssues": 11
        }
      ],
      "communityDiscussions": []
    },
    "runs": [
      {
        "collectedAt": "2026-08-13T12:26:38.318Z",
        "since": "2026-08-12T12:26:38.318Z",
        "observedCount": 54,
        "changedCount": 54
      },
      {
        "collectedAt": "2026-08-13T13:48:00.446149Z",
        "since": "2026-08-12T13:48:00.446149Z",
        "observedCount": 56,
        "changedCount": 56
      },
      {
        "collectedAt": "2026-08-13T16:19:22.035158Z",
        "since": "2026-08-12T16:19:22.035158Z",
        "observedCount": 45,
        "changedCount": 12
      },
      {
        "collectedAt": "2026-08-13T17:43:20.785491Z",
        "since": "2026-08-12T17:43:20.785491Z",
        "observedCount": 43,
        "changedCount": 11
      },
      {
        "collectedAt": "2026-08-13T17:47:07.884300Z",
        "since": "2026-08-12T17:47:07.884300Z",
        "observedCount": 42,
        "changedCount": 0
      },
      {
        "collectedAt": "2026-08-13T18:01:55.420671Z",
        "since": "2026-08-12T18:01:55.420671Z",
        "observedCount": 41,
        "changedCount": 3
      }
    ],
    "signals": [
      {
        "id": "github:redpanda-data/redpanda:issue:2749",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "issue",
        "title": "Add healthcheck for Docker",
        "text": "Hi, I'm not sure where the Dockerfile lives and this might be the wrong repo but I'll throw the issue here in the hope that it's the best place. :) ## The goal (aka what we're trying to do): We are trying to set up a [docker-compose example with redpanda and tremor](https://github.com/tremor-rs/tremor-redpanda/) to make it easy for users to test out this integration. Where we're trying to get is that a user can clone the repo, run `docker compose up` (or `docker-compose up`) and have a working set of systems to try, then extend and play with. ## The issue (aka what we observe): At times when the startup of the different nodes is out of sync slightly, the topic is reported as an unknown topic or partition. ``` tremor_out_1 | 2021-10-22T09:05:02.208061694+00:00 ERROR tremor_runtime::source::kafka - [Source::tremor://localhost/onramp/redpanda-in/filter/out] Subscription failed: UnknownTopicOrPartition (Broker: Unknown topic or partition). Stopping. [11:08 AM] ``` ## The expectations (aka what we would have expected): The redpanda instance is started with `auto_create_topics_enabled=true` so we would have expected for unknown topics to be created and not throw an error. ## The guesswork (aka what we think happens): The `docker-compose.yaml` specifies the dependencies of the different containers so they should start in order. But looking at the [Dockerfile](https://hub.docker.com/layers/vectorized/redpanda/dev/images/sha256-ca167ad57167c52db1e16faa7e426f2510b434f8b63f86ca3cc85d8d10eab1ae?context=explore) for redpanda (let's hope it's the right one) there is no [`HEALTHCHECK`](https://docs.docker.com/engine/reference/builder/#healthcheck) specified in here. So our guess is that there is a race condition where the container is booted and docker thinks it's up but the redpanda system isn't fully bootstrapped yet. ## The dream (aka what we hope would fix this): Adding a [`HEALTHCHECK`](https://docs.docker.com/engine/reference/builder/#healthcheck) to redpandas `Dockerfile` - that said we're somewhat unsure what would be the best way to check this so would love someone with more knowledge of the internal workings to get some input. @mfelsche suggested `http://localhost:9644/v1/status/ready` so would this work: ```docker HEALTHCHECK --interval=30s --timeout=1s --start-period=5s --retries=3 CMD curl -f http://localhost:9644/v1/status/ready || exit 1 ``` ## Reproduction 1. Clone https://github.com/tremor-rs/tremor-redpanda/ 2. run drocker compose up JIRA Link: [CORE-770](https://redpandadata.atlassian.net/browse/CORE-770) [CORE-770]: https://redpandadata.atlassian.net/browse/CORE-770?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/issues/2749",
        "createdAt": "2021-10-22T09:35:30Z",
        "updatedAt": "2026-08-12T16:43:22Z",
        "timestamp": "2026-08-12T16:43:22Z",
        "metrics": {
          "reactions": 1,
          "comments": 7
        },
        "labels": [
          "kind/enhance",
          "area/build",
          "community"
        ],
        "author": "Licenser",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:issue:31545",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "issue",
        "title": "rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
        "text": "## What happened `rpk connect upgrade` fails for every currently installed Redpanda Connect version, because it parses the *current* version through a regex that caps every segment at two digits before comparing it to latest. Connect's minor version passed 100 in 2026, so upgrade is broken for every install since ~mid-2026. ``` $ rpk connect --version Version: 4.103.1 Date: 2026-07-31T15:40:36Z $ rpk connect upgrade unable to determine current version of Redpanda Connect: unable to determine Redpanda version from \"4.103.1\": unable to get the redpanda version from \"4.103.1\" ``` Reproduced on rpk **v26.1.15** (darwin arm64) — the release that already contains the #31438 fix — so this is not the same code path as #31432/CON-529, just the same bug class in a sibling function that fix didn't touch. Also present on **v26.2.1** and on **dev** as of 2026-08-12. ## Root cause `src/go/rpk/pkg/redpanda/version.go`, `VersionFromString` (called by `connect/upgrade.go`'s `connectVersion` helper): ```go vMatch := regexp.MustCompile(`^v?(\\d{1,2})\\.(\\d{1,2})\\.(\\d{1,2})(?:\\s|-rc\\d{1,2}|-dev|-nightly|$)`).FindStringSubmatch(s) ``` Same defect as #31432: `\\d{1,2}` caps every segment at two digits, so parsing the installed version string (e.g. `4.103.1`) fails outright. `upgrade.go` calls this on the *currently installed* version before it can even fetch/compare against latest, so the command dies before doing anything else. ## Suggested fix Apply the same fix #31438 used for `install.go`'s regex: ```go regexp.MustCompile(`^v?(\\d+)\\.(\\d+)\\.(\\d+)(?:\\s|-rc\\d+|-dev|-nightly|$)`) ``` `VersionFromString` is exported from `pkg/redpanda` — worth grepping for other call sites that may hit the same wall (it's used for more than just Connect version parsing). ## Impact - `rpk connect upgrade` cannot run against any Connect install >= 4.100.0 — the primary way most users are told to update (\"If you want to upgrade to the latest version, please run 'rpk connect upgrade'\" is `install`'s own suggested next step). - Anything that shells out to `rpk connect upgrade` to ensure a plugin is current before reading its version breaks the same way; docs automation worked around it by calling `rpk connect install --force` instead (`install`'s default `\"latest\"` path never reaches this regex). Found while reproducing and fixing CON-529/#31432 for docs automation (DOC-2417) — installing a pinned >= 4.100.0 version to unblock the first bug immediately exposed this second one on the very next command.",
        "url": "https://github.com/redpanda-data/redpanda/issues/31545",
        "createdAt": "2026-08-12T10:28:09Z",
        "updatedAt": "2026-08-12T15:49:26Z",
        "timestamp": "2026-08-12T15:49:26Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "JakeSCahill",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:issue:31548",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "issue",
        "title": "[v26.2.x] rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
        "text": "Backport https://github.com/redpanda-data/redpanda/issues/31545 to branch v26.2.x. Requested by PR https://github.com/redpanda-data/redpanda/pull/31546",
        "url": "https://github.com/redpanda-data/redpanda/issues/31548",
        "createdAt": "2026-08-12T15:51:46Z",
        "updatedAt": "2026-08-12T18:12:43Z",
        "timestamp": "2026-08-12T18:12:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:issue:31550",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "issue",
        "title": "[v26.1.x] rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
        "text": "Backport https://github.com/redpanda-data/redpanda/issues/31545 to branch v26.1.x. Requested by PR https://github.com/redpanda-data/redpanda/pull/31546",
        "url": "https://github.com/redpanda-data/redpanda/issues/31550",
        "createdAt": "2026-08-12T15:52:53Z",
        "updatedAt": "2026-08-12T18:12:57Z",
        "timestamp": "2026-08-12T18:12:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:issue:31552",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "issue",
        "title": "[v25.3.x] rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
        "text": "Backport https://github.com/redpanda-data/redpanda/issues/31545 to branch v25.3.x. Requested by PR https://github.com/redpanda-data/redpanda/pull/31546",
        "url": "https://github.com/redpanda-data/redpanda/issues/31552",
        "createdAt": "2026-08-12T15:55:59Z",
        "updatedAt": "2026-08-12T18:33:24Z",
        "timestamp": "2026-08-12T18:33:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:issue:4191",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "issue",
        "title": "Add rpk standalone installscript for OSX,Linux",
        "text": "### Who is this for and what problem do they have today? - Developer wants to interact with redpanda. The developer is not involved in Redpanda's hosting or installation. - Developer uses Linux distributions other than rhel/centos/fedora/ubuntu. ### What are the success criteria? - rpk can be installed with a shell one-liner on \"all\" Linux Distros and OSX - without installing redpanda server. ### Why is solving this problem impactful? - Developer Experience for the affected person is impacted. Shell one-liner is quicker than downloading the artifact; especially if the user does not use a currently supported Linux distro. ### Additional notes - This is only an improvement. The binaries are available for download from GH releases. - Binaries of rpk are already distributed in GitHub releases, so the heavy lifting is already done. Either some custom shell script, or maybe https://goreleaser.com/ can be used to provide an install script. JIRA Link: [CORE-879](https://redpandadata.atlassian.net/browse/CORE-879) [CORE-879]: https://redpandadata.atlassian.net/browse/CORE-879?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/issues/4191",
        "createdAt": "2022-04-05T07:09:39Z",
        "updatedAt": "2026-08-12T16:48:12Z",
        "timestamp": "2026-08-12T16:48:12Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [
          "kind/enhance",
          "area/rpk"
        ],
        "author": "birdayz",
        "state": "closed",
        "assignees": [
          "birdayz"
        ]
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:29258",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] Kgo Verifier Producer set linger to 0",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/29234",
        "url": "https://github.com/redpanda-data/redpanda/pull/29258",
        "createdAt": "2026-01-14T16:29:13Z",
        "updatedAt": "2026-08-12T15:49:42Z",
        "timestamp": "2026-08-12T15:49:42Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:29280",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] Add comparison operators to iobuf fuzz test (and other enchancements)",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/29200",
        "url": "https://github.com/redpanda-data/redpanda/pull/29280",
        "createdAt": "2026-01-15T21:28:39Z",
        "updatedAt": "2026-08-12T15:42:49Z",
        "timestamp": "2026-08-12T15:42:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:29795",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] Implement cross-segment prefetching for small segments in cloud storage reads",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/29496",
        "url": "https://github.com/redpanda-data/redpanda/pull/29795",
        "createdAt": "2026-03-11T10:34:42Z",
        "updatedAt": "2026-08-12T17:57:36Z",
        "timestamp": "2026-08-12T17:57:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30314",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] kafka/server: cap fetch memory allocation at max message size limit",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30023",
        "url": "https://github.com/redpanda-data/redpanda/pull/30314",
        "createdAt": "2026-04-28T02:54:07Z",
        "updatedAt": "2026-08-12T17:45:48Z",
        "timestamp": "2026-08-12T17:45:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30531",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] config: refresh iceberg_enabled docstring",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30529",
        "url": "https://github.com/redpanda-data/redpanda/pull/30531",
        "createdAt": "2026-05-19T15:12:12Z",
        "updatedAt": "2026-08-12T15:41:49Z",
        "timestamp": "2026-08-12T15:41:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30673",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] `compaction`: avoid recompression of unchanged batches",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30669 - Command: git cherry-pick -x c2abd24ceb 57e62ea1a7 - Commits backported: 2 - Conflicts resolved: 2 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30669-v25.3.x-1780368343 ## Conflict details - c2abd24ceb (src/v/compaction/types.cc): target branch still uses free `operator<<` instead of dev's `format_to` member, and lacks `expired_tombstones_discarded`; added only `compressed_batches_reused` to the existing format string and arg list. - c2abd24ceb (src/v/compaction/types.h): added `compressed_batches_reused` field; skipped the unrelated `expired_tombstones_discarded` field that exists only on dev. - c2abd24ceb (src/v/storage/compaction_reducers.cc): kept v25.3.x's narrower placeholder-install condition (only when last batch in segment — `_stm_mgr` / idempotent-window logic does not exist on this branch) while wrapping the return in the new `filtered_batch{ rebuilt, ... }`; in the second hunk switched the index entry to use `to_append` per the cherry-pick's intent. - c2abd24ceb (src/v/storage/tests/segment_deduplication_test.cc): merged the new include set with the v25.3.x-only `gmock`, `random/generators`, `storage/chunk_cache` headers; dropped the `disk_log.stm_hookset()` argument from the new test's `copy_data_segment_reducer` constructor call since v25.3.x's constructor doesn't take an stm hookset. - 57e62ea1a7 (src/v/compaction/filter.cc): merged the `operator()` decompress path with the cherry-pick's `decompress` + reuse logic; adapted `filter_and_rewrite_with_sink` to v25.3.x's two-arg sink — when the batch is unchanged and a compressed original is available, hand the original to the sink with `compression::none` so it isn't re-compressed; otherwise pass the (possibly rebuilt) batch with the original compression so the sink compresses it as before. - 57e62ea1a7 (src/v/compaction/tests/reducer_test.cc): merged the two include sets (`bytes/bytes.h` from HEAD plus the new `compaction/key.h` and `compaction/key_offset_map.h` from the cherry-pick).",
        "url": "https://github.com/redpanda-data/redpanda/pull/30673",
        "createdAt": "2026-06-02T02:52:55Z",
        "updatedAt": "2026-08-12T15:49:39Z",
        "timestamp": "2026-08-12T15:49:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30716",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] build/deps: upgrade krb5 to 1.22.2",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30628 - Command: git cherry-pick -x 5fffb24bdc 5acac50623 ba231dd72d de492c91a7 - Commits backported: 4 - Conflicts resolved: 1 - Commits skipped (already on target): 1 - Backport branch: ai-backport-pr-30628-v25.3.x-1780530761 ## Conflict details - 5acac50623 (MODULE.bazel.lock): generated lockfile conflicted; accepted the incoming version from the PR (`git checkout --theirs`). Needs regeneration on the target branch (see below). - de492c91a7 (merge commit \"Merge branch 'dev' into fix/upgrade-kerberos-1.22\"): skipped. This is a sync merge of `dev` into the feature branch (first parent ba231dd72d was already applied, second parent is the `dev` tip). It carries no PR-specific delta, so cherry-picking it would pull unrelated `dev` changes onto the release branch. ## ⚠️ Generated files The following files were cherry-picked and may need regeneration: - MODULE.bazel.lock These files were accepted as-is from the source branch. Before merging, regenerate them on the target branch to ensure they're correct. For example: - MODULE.bazel.lock: run `bazel mod deps --lockfile_mode=update`",
        "url": "https://github.com/redpanda-data/redpanda/pull/30716",
        "createdAt": "2026-06-03T23:54:13Z",
        "updatedAt": "2026-08-12T15:49:37Z",
        "timestamp": "2026-08-12T15:49:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30776",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] iceberg: Push Parquest column stats to Iceberg manifests",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30704 - Command: git cherry-pick -x 7f833530e9 b4e8232b4e - Commits backported: 2 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30704-v25.3.x-1781219157 ## Conflict details - 7f833530e9 (src/v/datalake/base_types.h): include block conflicted only on context; accepted the incoming includes (`base/format_to.h`, `bytes/bytes.h`, `container/chunked_vector.h`, `serde/envelope.h`, `serde/rw/bytes.h`) needed by the new `per_column_stats` struct. - 7f833530e9 (src/v/datalake/BUILD): accepted the incoming `base_types` library deps (`//src/v/base`, `//src/v/bytes`, `//src/v/container:chunked_vector`, `//src/v/serde`, `//src/v/serde:bytes`). - 7f833530e9 (src/v/datalake/coordinator/BUILD): kept the target's `//src/v/base` dep and added the incoming `//src/v/bytes` dep for `iceberg_file_committer`. - 7f833530e9 (src/v/serde/parquet/writer.cc): kept the two new file-level accumulators (`file_value_count`, `file_column_size_bytes`); the surrounding bloom-filter block in the incoming context does not exist on v25.3.x and was omitted (bloom filters are not present on the target branch). - 7f833530e9 (src/v/serde/parquet/column_writer.cc): the target branch lacks the stats-truncation feature, so the commit's `build_statistics` helper was adapted to the target's non-truncated form (`encode_for_stats(...)`, `is_exact=true`). Replaced per-page `record_value()` calls with the commit's `_flushed_stats.merge(_current_page_stats)`, added the new `_file_stats` collector and `file_column_stats()` accumulation, and omitted the `_bloom_filter` member (bloom filters absent on target). The second commit (b4e8232b4e, ducktape test) applied cleanly.",
        "url": "https://github.com/redpanda-data/redpanda/pull/30776",
        "createdAt": "2026-06-11T23:11:14Z",
        "updatedAt": "2026-08-12T15:49:40Z",
        "timestamp": "2026-08-12T15:49:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30777",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] iceberg: Push Parquest column stats to Iceberg manifests",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30704 - Command: git cherry-pick -x 7f833530e9 b4e8232b4e - Commits backported: 2 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30704-v26.1.x-1781219179 ## Conflict details - 7f833530e9 (src/v/serde/parquet/writer.cc): only the two new file-level stat accumulation lines (`file_value_count`/`file_column_size_bytes`) were applied; the incoming bloom-filter block is a separate feature not present on v26.1.x and was omitted. - 7f833530e9 (src/v/serde/parquet/column_writer.cc): the commit was built atop a stats-truncation feature (`truncate_max`/`truncate_min`, `max_stats_truncate_length`, `is_utf8_string`) that does not exist on v26.1.x. Adapted `build_statistics()` and `flush_page()` to the target's non-truncating bound logic, dropped the redundant `_flushed_stats.record_value()` calls in favor of `merge()`, added the `_file_stats` collector, and omitted the unrelated `_bloom_filter` member. - 7f833530e9 (src/v/datalake/base_types.h): merged the new includes needed by `per_column_stats` (bytes, chunked_vector, serde envelope/rw); omitted `base/format_to.h` since v26.1.x's `local_file_metadata` still uses `operator<<` rather than `format_to`. - 7f833530e9 (src/v/datalake/BUILD): merged the new `base_types` deps (base, bytes, container:chunked_vector, serde, serde:bytes) into the target's deps list. - 7f833530e9 (src/v/datalake/coordinator/BUILD): kept both the target's `//src/v/base` dep and the incoming `//src/v/bytes` dep in `iceberg_file_committer`'s implementation_deps.",
        "url": "https://github.com/redpanda-data/redpanda/pull/30777",
        "createdAt": "2026-06-11T23:12:24Z",
        "updatedAt": "2026-08-12T15:43:24Z",
        "timestamp": "2026-08-12T15:43:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:30869",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] k/s/group_manager: fix inflated consumer group lag after log truncation",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30822 - Command: git cherry-pick -x a7da3be13f c3281ec0af b603df2d74 - Commits backported: 3 - Conflicts resolved: 2 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30822-v25.3.x-1782198716 ## Conflict details - a7da3be13f (src/v/cluster/BUILD): srcs/hdrs lists diverged around the insertion point; kept the target branch's `partition_balancer_types`/`partition_leaders_table` entries and added the new `partition_kafka_offsets.{cc,h}` in alphabetical order. - a7da3be13f (src/v/kafka/data/replicated_partition.cc): method ordering differs between branches (the commit deletes `partition_kafka_start_offset()` and `kafka_start_offset_with_override()`, which sit before the `stm_exact_offset_replicator`/`make_exact_offset_replicator` block on the target branch); removed the two now-extracted methods and kept the exact-offset-replicator block. The delegating edits to `start_offset()`/`high_watermark()`/`sync_effective_start()` applied cleanly. - c3281ec0af (src/v/cluster/health_monitor_backend.cc): include block diverged; the target branch had already dropped the unused `cluster/node_status_table.h` include, so only added `cluster/partition_kafka_offsets.h`. - c3281ec0af (src/v/cluster/health_monitor_types.cc): the target branch still prints `partition_status` via `operator<<` rather than the source branch's `format_to` member; applied the commit's intent by adding `log_start_offset` to the format string and argument list, keeping the `operator<<` form and `ps.` member prefixes.",
        "url": "https://github.com/redpanda-data/redpanda/pull/30869",
        "createdAt": "2026-06-23T07:16:58Z",
        "updatedAt": "2026-08-12T15:49:35Z",
        "timestamp": "2026-08-12T15:49:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31139",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "bazel/seastar: enable task queue shuffling in debug builds",
        "text": "Seastar's CMake build defines `SEASTAR_SHUFFLE_TASK_QUEUE` in the Debug and Sanitize configurations, randomizing task queue execution order to surface latent task ordering assumptions. The Bazel debug build was missing this define; add it to restore the old CMake behavior. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.1.x - [ ] v25.3.x - [ ] v25.2.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31139",
        "createdAt": "2026-07-16T21:10:55Z",
        "updatedAt": "2026-08-13T12:06:00Z",
        "timestamp": "2026-08-13T12:06:00Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31156",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] `storage`: make `_schemas` deletion exempt",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31130",
        "url": "https://github.com/redpanda-data/redpanda/pull/31156",
        "createdAt": "2026-07-17T13:30:48Z",
        "updatedAt": "2026-08-12T15:49:34Z",
        "timestamp": "2026-08-12T15:49:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31186",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] storage: mark snapshot writer/reader closed regardless of close outcome",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31180",
        "url": "https://github.com/redpanda-data/redpanda/pull/31186",
        "createdAt": "2026-07-20T18:05:16Z",
        "updatedAt": "2026-08-12T15:41:46Z",
        "timestamp": "2026-08-12T15:41:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31192",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] kafka/client: authenticate under the reconnect mutex",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31152 - Command: git cherry-pick -x b08c929724 - Commits backported: 1 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-31152-v26.1.x-1784570815 ## Conflict details - b08c929724 (src/v/kafka/client/test/BUILD): the commit added the granular `//src/v/cluster:*` sub-targets (`client_quota_backend`, `cluster_utils`, `controller_log_limiter`, `data_migration_table`, `feature_backend`, `leader_balancer_strategy`, `node_status_backend`, `self_test`, `topic_metrics_watcher`) plus `//src/v/config` and `//src/v/kafka/client:api_types` to the `cluster_test` deps. On v26.1.x `//src/v/cluster` is still a single monolithic target and those fine-grained sub-targets do not exist, so the granular deps were collapsed to the single `//src/v/cluster` target (the pattern already used by the other test targets in this file), keeping `//src/v/config` and `//src/v/kafka/client:api_types`. This satisfies every include added to cluster_test.cc (`cluster/security_frontend.h` is provided by `//src/v/cluster`).",
        "url": "https://github.com/redpanda-data/redpanda/pull/31192",
        "createdAt": "2026-07-20T18:11:03Z",
        "updatedAt": "2026-08-12T15:41:44Z",
        "timestamp": "2026-08-12T15:41:44Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31301",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "metrics: remove one-shot IMDS connect probe",
        "text": "## What Removes the one-shot IMDS connect probe from the cloud instance-type detector (`instance_type_detector.cc`). ## Why The detector probed the EC2 IMDS endpoint with a single bounded `base_transport` connect before issuing any request. That was a workaround (added alongside the capacity-metrics feature in #30742) for `http::client::get_connected` retrying `connect()` in a tight **no-backoff loop** that spun hot whenever the link-local `169.254.169.254` address was unreachable, i.e. whenever we're not on EC2. The spin flooded the log and burned a shard's CPU for the whole timeout window (this was CORE-16451, and it caused the CI blowup in build #85550). #30796 (\"http: dial all resolved addresses in a single bounded attempt\") replaced that retry loop with a single bounded dial pass, so the request path no longer spins on an unreachable endpoint. The probe is now redundant: off EC2, issuing the request directly fails fast via `query_ec2_instance_type`'s normal failure path, which already returns `std::nullopt` on failure. JIRA: [CORE-16451](https://redpandadata.atlassian.net/browse/CORE-16451) ## Backports Required - [x] none - not a bug fix (removes a now-unnecessary workaround) ## Release Notes * none [CORE-16451]: https://redpandadata.atlassian.net/browse/CORE-16451?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/pull/31301",
        "createdAt": "2026-07-27T15:45:26Z",
        "updatedAt": "2026-08-13T13:20:46Z",
        "timestamp": "2026-08-13T13:20:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "travisdowns",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31365",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "rptest: add OOM crash self-test; allow-list memory diagnostics",
        "text": "Revives #8570 (rebased onto dev, see #8565 for the original motivation): put the seastar memory diagnostics dump, which is emitted at ERROR level when a node runs out of memory, on the default log allow list so an OOM surfaces as the \"redpanda crashed\" diagnostic rather than a BadLogLines about the dump. Also clarifies in the raise_on_bad_logs docstrings that allow-list regexes are unanchored. The main value over the old PR is the meta-test: a new debug-only `oom` type for the `/v1/debug/trigger_crash` admin API that allocates until the seastar allocator gives out, plus a ducktape self-test (alongside the existing segfault/assert ones) that runs a node out of memory and checks the harness reports it as a NodeCrash. With the system allocator (sanitized builds) there is no seastar memory limit and unbounded allocation would exhaust host memory instead, so there the API refuses the request, and the self-test verifies the refusal and that the node survives. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.1.x - [ ] v25.3.x - [ ] v25.2.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31365",
        "createdAt": "2026-07-30T14:45:49Z",
        "updatedAt": "2026-08-13T17:12:15Z",
        "timestamp": "2026-08-13T17:12:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "travisdowns",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31366",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[CORE-16604] cluster_link: incremental Schema Registry sync by tailing _schemas",
        "text": "Schema Registry API sync only propagates changes on its periodic full sync, so a registration can wait a whole interval and every propagation costs a whole-registry scan of both sides. This adds a change feed: the link consumes the source's _schemas topic over the Kafka API and re-reads only the targets whose records it saw, reconciling them by the same per-subject path a full sync uses - so replay is harmless and semantics are identical. A CONTEXT record instead re-lists the source's contexts, since deletion is decided against the whole set. Full syncs still run on their interval as the backstop. A source that does not expose _schemas over Kafka - Confluent Cloud, Redpanda Serverless, or one too old for ListOffsets v4 - never arms. One that exposes it but refuses the fetch, e.g. a missing READ ACL, arms and then stops on its first poll. Either way the link keeps replicating on its full-sync interval alone. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Schema Registry API sync now propagates source changes by tailing the source's `_schemas` topic rather than waiting for the next full sync. Requires the link's Kafka credentials to have READ on `_schemas`; without it the link falls back to full syncs alone.",
        "url": "https://github.com/redpanda-data/redpanda/pull/31366",
        "createdAt": "2026-07-30T16:23:55Z",
        "updatedAt": "2026-08-13T14:13:02Z",
        "timestamp": "2026-08-13T14:13:02Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "mnajda-redpanda",
        "state": "open",
        "assignees": [
          "mnajda-redpanda"
        ]
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31411",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[CORE-15822] security/audit: make audit initialization controller-leader independent",
        "text": "## Problem On the RPC audit sink (the default since v25.3), `create_internal_topic()` is the **entire** configure step of `audit_client::initialize()`, and it unconditionally sends `CreateTopics` to the controller — so every audit enable/startup requires a reachable controller leader, even when the audit topic already exists. `topics_frontend::autocreate_topics` returns `no_leader_controller` immediately when no controller leader is elected (it does not wait for one), so initialization loops for as long as the leaderless state lasts. Transient churn (ordinary elections, rolling restarts with quorum) is absorbed by retries and the per-shard queues; but when the leaderless state outlasts that absorption (~90 s at default queue sizes — prolonged states of this kind are observed in production, e.g. CORE-15946), the drain timer is never armed, the queues fill and stay full, and with `audit_failure_policy=reject` authentication and the admin API are refused cluster-wide until a leader returns. This failure mode reproduces on HEAD. The Kafka-client sink (all pre-25.3 deployments, and the pre-25.3 → 25.3+ mixed-version upgrade window) has the same controller dependency twice over — its ACL write and its `CreateTopics` both need an elected leader — plus a credential-propagation design that this PR also fixes (below). Production history of this failure class: INC-2362 / CORE-15822 and CORE-12785 (full RCA with log evidence in [CORE-15822]). In both, the audit topic already existed and the produce path was viable the whole time. Those incidents ran the legacy Kafka-client sink, where initialization had an additional independent blocker (the client's all-or-nothing metadata bootstrap, fixed separately by the client rewrite in v25.2). ## Change ### RPC sink Check the local topic table (controller-log-backed, replicated to every node) before creating: if `_redpanda.audit_log` is already known, skip `CreateTopics` entirely. With the topic present — i.e. every enable/startup except the first in a cluster's life — RPC-sink initialization becomes dependency-free. ### Kafka-client sink The first version of this PR left this sink out, because its `CreateTopics` doubled as the ephemeral-credential bootstrap: the create's `sasl_authentication_failed` was the only trigger for `inform()`, which propagates the `__auditing` credential to brokers. Skipping the create without replacing that mechanism breaks authentication outright (an earlier skip-only revision failed 48 of 51 failing cases in a full audit-suite run with exactly that signature). Three changes make the sink leader-independent safely: 1. **Proactive `inform`-all** in `do_configure()`, right after minting the credential and before any broker contact. Credential propagation no longer rides on a create failure. This also fixes a standalone wedge reproduced while testing this PR: a broker that mints a fresh credential while its peers are unreachable (e.g. restarting into a leaderless window) keeps failing SASL against the returning peers (`Invalid credentials`); the kafka client aggregates those per-broker failures into `broker_error({node: -1}, broker_not_available)`, which is not `sasl_authentication_failed`, so `mitigate_error()` rethrows and `inform()` never runs — initialization loops forever (observed for minutes at the 75 s backoff cap) until a broker restart, which is exactly how the production incidents were remediated. Informs to unreachable peers are best-effort (logged, not thrown) and covered on demand by the existing mitigation path. 2. **ACL exists-skip**: `set_auditing_permissions()` returns early when both `__auditing` bindings (create + write on the audit topic) are readable in the local authorizer, eliding a raft-0 write that needs an elected leader. A stale-negative read is benign — `create_acls` is idempotent. The `create_acls` results are now also checked: its errors come back in-band (e.g. `no_leader_controller`) and were previously discarded, which the skip would have turned into a real hole — initialization could succeed with the topic present but no write ACL installed. Failing loudly keeps `initialize()` retrying until the ACL write lands. 3. **Topic exists-skip**, same as the RPC sink. Safe now that (1) owns credential propagation; and without a leader the request could not even be dispatched — `CreateTopics` must be routed to the controller broker. ### New error code handled in `update_status`: `topic_authorization_failed` Added to both the `_misconfigured` entry condition and `failed_codes`, alongside `illegal_sasl_state`. Why: if the `__auditing` ACL bindings are ever missing while auditing runs — the runtime deletion case, or the ACL exists-skip reading a stale positive — audit produces start failing authorization. Previously that state was **silent**: batches were dropped with only a WARN and the errors counter. Now it flips `_misconfigured`, so the enqueue gate degrades loudly (`Audit message rejected due to misconfigured authorization`, failure policy applied) and recovers automatically once the bindings are back. The exposure is narrow — `__auditing` is an `ephemeral_user` principal, a type the Kafka ACL API cannot express, so the bindings cannot be deleted by a principal-targeted filter — but a principal-less (wildcard) `DeleteAcls` on the topic does match them, and a silent-drop failure mode is the wrong default either way. Semantics notes: - `CreateTopics` never reconciled or altered an existing topic (`topic_already_exists` was a no-op), so skipping it cannot regress any re-creation path; a deleted topic is absent from the table and re-enable recreates it as before; - stale-negative table read (lagging, topic exists): falls through to the old path, controller answers `topic_already_exists` — degradation to today's behavior; - stale-positive read (topic deleted mid-enable): same exposure as today's create-then-delete race; recoverable by toggling `audit_enabled`; - stale-positive ACL read: degrades loudly via `topic_authorization_failed` above, recoverable by toggling `audit_enabled` (the skip then sees the bindings are gone and recreates them). ## The flow this fixes **Before** — a broker (re)starts, or auditing re-initializes, while the controller group has no elected leader: ``` init -> create_internal_topic -> needs controller-group leader -> no_leader_controller (immediately, no waiting) -> initialize() loops -> drain timer never armed -> per-shard queues fill (~90 s) -> reject refuses authn + admin API cluster-wide until a leader returns ``` **After** — same conditions, both sinks: ``` init -> exists-skips (local topic-table + authorizer reads, no network) -> configure() complete: no controller-group dependency -> drain armed -> produce goes to the audit partitions' own leaders (independent raft groups; commit = quorum of THEIR replicas) -> auditing fully functional with no controller leader at any step ``` The two paths use different leaders from different elections: `CreateTopics` needs the leader of the controller group (raft group 0), while writing audit records needs only the leaders of the `_redpanda.audit_log` partitions — which is why in both analyzed incidents the write path was healthy the entire time the init path was stuck. After this change the audit subsystem touches the controller group only for one-off events (the very first topic creation in a cluster's life, config changes), never for steady-state operation. Fine print: the very first enable still requires a controller leader (the topic and ACLs genuinely must be created), and flipping `audit_enabled` itself is a metadata mutation; the audit partitions must have elected leaders, but their elections are per-partition and do not involve the controller group. ## Testing All ducktape, all PASS on this branch: - `AuditLogLeaderlessControllerTest` (matrix over both transports): 5 brokers, the single audit partition pinned to the first three; stopping 3 of 5 leaves raft group 0 leaderless while the audit partition keeps 2/3 replicas. A surviving broker restarts into that window and must initialize through the exists-skips (no CreateTopics, no failure/retry loop, fibers start; the Kafka-client case additionally pins the ACL skip), and an admin request issued **during the window** must be consumed back from the audit topic **while the controller is still leaderless** — the end-to-end proof that a leaderless controller does not block auditing. After quorum returns, a controller-dependent event flows too, with no extra restart (the incidents required one). - `AuditLogTopicRecreateTest` (matrix over both transports): false-positive guard — with the audit topic genuinely deleted (auditing disabled, topic de-listed from `kafka_nodelete_topics`, then deleted), the skip must NOT fire; re-enable goes through `CreateTopics` and recreates it, and events flow. - `AuditLogTopicExistsTest` (matrix over both transports): re-enable with the topic present initializes without `CreateTopics` and still delivers; the Kafka-client case additionally pins the ACL skip and the proactive inform-all (`Informed:` logs). ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x v25.2.x / v25.1.x are not in the standard list, but if demand materializes they are now feasible as-is: this PR contains the Kafka-client variant (proactive inform + both skips), which those release lines run. ## Release Notes ### Improvements * Audit logging no longer requires controller availability on startup/enable when the audit topic already exists — on both the RPC transport (default since v25.3) and the Kafka-client transport. * A missing-ACL state for the audit principal now surfaces as rejected/permitted-per-policy audit events (misconfigured-authorization) instead of silently dropping records. 🤖 Generated with [Claude Code](https://claude.com/claude-code) [CORE-15822]: https://redpandadata.atlassian.net/browse/CORE-15822?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/pull/31411",
        "createdAt": "2026-08-04T15:30:22Z",
        "updatedAt": "2026-08-13T14:54:39Z",
        "timestamp": "2026-08-13T14:54:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 9
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "bartoszpiekny-redpanda",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31461",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
        "text": "`NodeDecommissionWaiter.wait_for_removal` asserts that a decommission has \"stopped making progress\" if `replicas_left`/bytes-left-to-move hasn't improved within `progress_timeout`. In tests that run `FailureInjectorBackgroundThread` alongside node operations (e.g. `ShadowLinkingRandomOpsTest.test_node_operations`), this assertion can trip even though the system is behaving correctly: the partition balancer legitimately stalls while nodes are being restarted, because each restart resets the ~30s disk-report freshness wait the balancer needs before it can compute new moves. `FailureInjectorBackgroundThread`'s default cadence (20-40s between injections) is faster than that recovery window, so a static `progress_timeout` can't distinguish \"still recovering from ongoing chaos\" from a genuine stall. This PR tracks failure-injection timestamps on `FailureInjectorBackgroundThread` and uses them in `NodeDecommissionWaiter` to hold off the \"stopped making progress\" check while chaos has been active in the last `chaos_settle_sec` (120s default), with an `absolute_max_wait_sec` (900s default) backstop so a genuine deadlock still fails in bounded time. <!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes * none <!-- If the changes in this PR do not need to be mentioned in the release notes, then don't add a sub-section and simply list `none`, e.g. * none Otherwise, adding a sub-section or `none` is REQUIRED if the PR is not a backport PR. If this is a backport PR, adding contents to this section will override the release notes section inherited from the original PR to dev. Add one or more of the sub-sections with a short description bullet point of the change, e.g. ### Bug Fixes * Short description of the bug fix if this is a PR to `dev` branch. ### Features * Short description of the feature. Explain how to configure. ### Improvements * Short description of how this PR improves existing behavior. -->",
        "url": "https://github.com/redpanda-data/redpanda/pull/31461",
        "createdAt": "2026-08-06T10:03:21Z",
        "updatedAt": "2026-08-13T07:18:28Z",
        "timestamp": "2026-08-13T07:18:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [],
        "author": "bartoszpiekny-redpanda",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31483",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[CORE-16998, CORE-17000, CORE-17001] - Consumer Groups (Classic): Fix some bugs where the offset map could diverge from the log",
        "text": "CORE-16998, CORE-17000, CORE-17001 Three bugs in the classic group's offset accounting, all the same defect: live state and a replay of the log resolve the same key differently. The group's committed offset then changes on a leader change, a restart or compaction, and no client asked for it. None is a KIP-848 regression, and each commit carries its own unit test. - **CORE-16998 (High)** the retention sweep can tombstone an offset that an open transaction has staged, and when the tombstone lands last, recovery finds no offset for the partition. Fix: `filter_expired_offsets` consults the staged partitions. - **CORE-17000 (High)** a transactional commit ordered by its staging offset loses to a plain commit that raced it, so live state and the log disagree. Fix: stamp the staged offsets with the commit marker's offset. - **CORE-17001 (Medium)** a request that names one partition twice commits the first value in memory and the last in the log. Fix: commit only the last entry for a partition. **One divergence from Kafka, in 17000.** Kafka keeps the plain value for that race. It writes the transactional record at TxnOffsetCommit time, so the plain commit is the later record in its log. Both coordinators keep whatever compaction leaves, so the outcomes differ only in record placement. To match Kafka we would write our transactional records at staging time, which is a separate change. **Backports.** The release branches predate the `offset_store` extraction, so all three land in `group.cc` as `group::` methods there. We adapt each cherry-pick by hand. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes ### Bug Fixes * Consumer groups: Fix a bug where retention could expire a committed offset while an open transaction holds a new value for it. * Consumer groups: Fix a bug where a plain offset commit landing while a transaction is open would temporarily overrule the eventual transaction offset commit. * Consumer groups: Fix a bug where an OffsetCommit naming the same partition twice would lead to the in-memory store and the log to disagree on the final value.",
        "url": "https://github.com/redpanda-data/redpanda/pull/31483",
        "createdAt": "2026-08-07T07:48:46Z",
        "updatedAt": "2026-08-13T06:05:48Z",
        "timestamp": "2026-08-13T06:05:48Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/redpanda",
          "claude-review"
        ],
        "author": "oleiman",
        "state": "open",
        "assignees": [
          "oleiman"
        ]
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31502",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "raft/tests: reproduce term span without configuration batch",
        "text": "The multi-term segments design assumes raft_configuration batches demarcate term boundaries. They do not. A candidate that wins an election but fails the local append of its initial configuration batch stays leader (`vote_stm.cc`: \"even if failed to replicate, don't step down\"). Nothing retries the append. Data of the new term then commits with no configuration batch for that term anywhere in the log. The test reproduces this and asserts the only durable record of the boundary is the file name of the segment rolled on term change — the safety net multi-term segments would remove. The first commit adds append failure injection to the fixture log. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31502",
        "createdAt": "2026-08-10T12:35:35Z",
        "updatedAt": "2026-08-13T10:45:49Z",
        "timestamp": "2026-08-13T10:45:49Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31512",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "k/s/tests: deflake fetch_memory_units cross-shard test",
        "text": "Units released back to their originating shard are returned via an unawaited cross-shard call, so they are not guaranteed to be visible to the immediately following read; under `--@seastar//:shuffle_task_queue=true` the read can run first (reproduced in 12/30 runs). Poll for the release instead. Test-only change. ## Backports Required - [ ] none - papercut/not impactful enough to backport - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [x] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31512",
        "createdAt": "2026-08-10T16:00:37Z",
        "updatedAt": "2026-08-13T00:33:19Z",
        "timestamp": "2026-08-13T00:33:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31517",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "hashing: back crc32c with abseil instead of google/crc32c",
        "text": "Backs `crc::crc32c` with abseil's CRC32C instead of google/crc32c. The wrapper becomes `basic_crc32c<Backend>` with backends for both libraries, and `crc::crc32c` is an alias for the abseil instantiation, so production switches over with no call site changes. CRC32C values are identical between the backends, which is what makes the switch safe: they are persisted in storage indices, snapshots and sst footers, and are part of the Kafka record batch wire format. The wrapper previously had no test coverage at all, so this adds a gtest suite typed over both backends that pins the absolute values to the RFC 3720 appendix B.4 vectors and the catalogued CRC-32/ISCSI check value — external references, not whatever the backing library computes — plus the properties callers rely on (incremental extends match one-shot at split points straddling both libraries' internal thresholds, the integral overload matches the equivalent bytes, `crc_extend_iobuf` is fragmentation independent). The final commit adds a benchmark that measures both backends through the wrapper, exactly as production calls them. It is the record of this decision and the way to revisit it; each fixture cross-checks that the backends agree before measuring. ## Benchmark `bazel run --config=release //src/v/hashing/tests:crc_bench_rpbench`. Times are the cost of one whole checksum, not one byte: | case | google/crc32c | abseil | abseil speedup | |---|--:|--:|--:| | batch header, 12 integer extends | 36.8ns | 2.19ns | **16.8x** | | 8B | 2.69ns | 1.34ns | 2.0x | | 57B | 4.29ns | 2.72ns | 1.6x | | 128B | 4.91ns | 5.13ns | 0.96x | | 512B | 27.0ns | 15.0ns | 1.8x | | 4032B | 113.6ns | 99.3ns | 1.14x | | 4096B | 99.4ns | 102.4ns | 0.97x | | 16256B | 447.0ns | 384.3ns | 1.16x | | 16384B | 382.9ns | 384.7ns | 1.0x | | 64KiB contiguous | 1.55µs | 1.52µs | 1.02x | | 64KiB iobuf, 512B fragments | 4.55µs | 2.46µs | **1.85x** | | 64KiB iobuf, 16KiB fragments | 1.56µs | 1.60µs | 0.98x | | 512B memcpy+crc, two-pass vs fused | 30.1ns | 20.7ns | 1.45x | | 64KiB memcpy+crc, two-pass vs fused | 2.36µs | 2.15µs | 1.10x | ## Why abseil is faster **Small extends are inlined behind a compile-time CPU check.** `absl::ExtendCrc32c` is defined in the header, and for buffers of 64 bytes or less expands in place to a straight-line sequence of hardware `crc32b/w/l/q` instructions (`crc_internal::ExtendCrc32cInline`), enabled at compile time by the SSE4.2 target baseline. `crc32c::Extend` is always an out-of-line library call that branches on a function-local `static bool can_use_sse42` and then calls the SSE4.2 kernel — for an 8-byte extend that is two calls and a dispatch check to reach a single `crc32q`, so per-call overhead dominates the actual CRC work. This is redpanda's hottest checksum shape: `model::internal_header_only_crc` feeds twelve 1–8 byte integers through the checksum for every replicated batch and pays the overhead twelve times. With abseil the entire sequence compiles to inlined crc32 instructions: 36.8ns → 2.2ns. The 128B row shows the flip side of the cutoff — just past 64B both libraries go out of line and google is marginally ahead. **Mid-size buffers get instruction-level parallelism sooner.** The `crc32` instruction has ~3-cycle latency but 1/cycle throughput, so a single dependency chain runs at a third of the hardware rate; full speed requires computing several independent streams and folding them together. google/crc32c engages its 3-way streams only in fixed block tiers — the smallest needs 3×336 = 1008 contiguous bytes, then 3×1360 and 3×5440 — and runs a serial 8-bytes-per-instruction loop below that and for whatever is left over. Abseil folds 3 streams starting at 256 bytes, and above 2KiB divides the whole buffer evenly across its streams. At 512 bytes google is fully serial while abseil is already 3-way: 27.0ns vs 15.0ns. Checksumming fragmented iobufs hits this shape repeatedly, hence 64KiB in 512B fragments going 4.55µs → 2.46µs. **Bulk throughput is a wash, but google sawtooths across its tiers.** Above ~4KiB both libraries run multi-stream at similar speed. Because google's tiers are fixed-size blocks, a length that just misses a tier drops its remainder to smaller tiers and finally the serial loop: 4032B (just under the 3×1360 tier) is ~14% slower per byte than 4096B (just over), and likewise 16256 vs 16384. Abseil's even split has no cliffs. This is also why the benchmark pins sizes on both sides of each tier boundary — a single size can rank the two libraries either way. Two footnotes recorded in the benchmark source: abseil's runtime CPU table does not recognize the benchmark host and falls back to its no-PCLMULQDQ streams for the out-of-line path, so abseil's mid/large numbers above are a floor rather than its best; and `absl::MemcpyCrc32c` (checksum fused with the copy in a single pass, non-temporal stores at large sizes) has no google/crc32c equivalent, so it is benchmarked raw against a memcpy-then-checksum two-pass. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Reduced the CPU cost of CRC32C checksumming (computed for every record batch header and for storage indices, snapshots and internal RPCs) by backing it with abseil's implementation. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/redpanda-data/redpanda/pull/31517",
        "createdAt": "2026-08-10T22:42:04Z",
        "updatedAt": "2026-08-12T20:45:32Z",
        "timestamp": "2026-08-12T20:45:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "dotnwat",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31521",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "rpk: fix grammar and wording defects in help text",
        "text": "## Summary Copyedits surfaced by auditing the docs team's rpk description overrides (`docs-data/rpk-overrides.json` in redpanda-data/docs) against rpk's source text — these are cases where the docs override exists only to paper over a source defect, so fixing the source lets the docs drop the override: - `profile clear`: \"unset **an prod** cluster profile\" → \"a production\" - `registry mode get`/`set`, `registry compatibility-level`: fix the \"global mode**, alternatively** you can…\" comma splice (same pattern in all three) - `security acl delete --cluster`: ACLs are removed *for*, not *to*, the cluster - `security acl list --registry-global`/`--registry-subject`: the list command **matches** ACLs, it doesn't **grant** them (text was copied from `acl create`, where \"grant\" is correct and unchanged) - `container start --retries`, `topic consume --fetch-max-bytes`/`--fetch-max-partition-bytes`: \"amount of\" → \"number of\" for countable nouns Help-text strings only; no behavior changes. Affected package tests pass. Related: #31520 (single-sources the `-X` flag docs from the same audit). 🤖 Generated with [Claude Code](https://claude.com/claude-code) ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31521",
        "createdAt": "2026-08-11T09:30:20Z",
        "updatedAt": "2026-08-13T16:41:50Z",
        "timestamp": "2026-08-13T16:41:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [
          "area/rpk"
        ],
        "author": "JakeSCahill",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31533",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[CORE-16631] tests/rptest: wait for target HWM to settle after failover",
        "text": "## Summary - `test_producer_ids_failover` occasionally hit kgo-verifier's \"possible idempotency bug\" check right after `failover_link()` — failover completion is a health-report heuristic (`link_status_reconciler::try_finish_failover`), not a guarantee that every previously-mirrored write has landed, so a fresh producer's one-shot `list-offsets(-1)` snapshot could race a trailing write and see a lower offset than what the broker later assigns the record. - Waits for the target partition's high watermark to stop moving before the second round of production, so the producer's offset baseline is taken after the partition has actually settled. ## Test plan - [x] `tools/type-checking/type-check.sh check rptest/tests/cluster_linking_e2e_test.py` — passes - [x] `tools/ruff format --diff tests/rptest/tests/cluster_linking_e2e_test.py` — clean - [x] Ran `test_producer_ids_failover[storage_mode=cloud]` 5x locally with the fix — 5/5 PASS, no added run-time overhead ## Caveats - Could not reproduce the original flake locally. CORE-16631's CI bot comments show ~48 occurrences over 2 months, all with the identical signature `sent=4 acked=2 bad_offsets=2 restarts=1`, but 0/125 local repeated runs (with or without this fix) hit it. An experiment widening `health_monitor_max_metadata_age` to 60s (to see if amplifying health-report staleness would force the race) also didn't reproduce it locally — CI data shows ~3x the rate on arm64 vs amd64 and is concentrated on shared `ci-docker-aws` runners, so the trigger likely needs real network/CPU jitter this local docker-compose setup doesn't have. - This fix closes a real, code-grounded race window, but is not proven to be the sole cause of every CORE-16631 occurrence — treat as a mitigation to validate against CI recurrence, not a confirmed root-cause fix. - Separately: `try_finish_failover` already has access to finer-grained per-partition watermark data (`source_partition_high_watermark`/`shadow_partition_high_watermark`) but doesn't use it in its completion decision — using it directly would be a more robust product-level fix than this test-side workaround. Worth a follow-up ticket if this mitigation doesn't fully close CORE-16631. Related to CORE-16631",
        "url": "https://github.com/redpanda-data/redpanda/pull/31533",
        "timestamp": "2026-08-12T13:30:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [],
        "author": "bartoszpiekny-redpanda",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31536",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "bazel: add a remote_download_minimal config to cut CI download volume",
        "text": "Adds `--remote_download_minimal` in bazelrc that stops CI materialising outputs no CI step reads. This reduces downloaded and materialized-on-disk bytes by about 70% and shaves maybe ~1 minute off of a cached build. jira: [CORE-16997](https://redpandadata.atlassian.net/browse/CORE-16997) ## Activated later It is inert until something passes `--config=ci-remote-cache`; the vtools change to do that in CI is a separate PR. Keeping the flags here means the whitelist lives next to the rest of the repo's build config rather than in CI config (a pattern we should follow for other cache related flags when it makes sense). ## Why Bazel's default is `--remote_download_outputs=toplevel`, which materialises every top-level artifact of `//...`. On a warm-cache `bazel test //...` that is ~56 GB, and ~53 (ish) GB of it is never opened by any CI step: | what | files | size | |---|---|---| | rpbench `_binary` targets | 63 | 15.86 GB | | intermediate `.a` / `.so` | 2,411 | 12.96 GB | | duplicate copies under `bazel/packaging` | 40 | 5.27 GB | | parquet test-case JSONs | 64 | 4.42 GB | | other | 1,636 | 3.40 GB | The rpbench `_binary` targets are the largest single category at ~16 GB, some individually 2.3 GB. They are `cc_binary`, not test rules, so Bazel's TEST-mode exclusion for test rules does not cover them, and the unit-test job never runs them. ## Measured before / after Two same-commit warm-cache CI runs. Spawn count, cache-hit count and total output bytes were **identical across both** (13,864 spawns / 13,860 cache hits / 95,855,439,177 output bytes), so the deltas are attributable to the flags alone. | config | materialised | files | result | |---|---|---|---| | default (`toplevel`) | **56.05 GB** | 15,475 | green | | **`--config=ci-remote-cache`** | **17.15 GB** (**-69%**) | 11,274 | green | Bytes actually on the wire are ~75% below the materialised figures because `--remote_cache_compression` is on: **17.37 GB -> 4.66 GB, -73%**. ### Timing | config | bazel wall time | critical path | CI job | |---|---|---|---| | default | 65.70 s | 25.86 s | 4.43 min | | **`--config=ci-remote-cache`** | **38.96 s (-41%)** | 13.17 s | 3.52 min (-21%) | Single sample per arm and the arms ran on different agents of the same type, so treat the timing as indicative; the byte figures are the solid part. ## What is still materialised under the new config Of the remaining 17.15 GB: | what | files | size | share | |---|---|---|---| | test logs / xml | 2,007 | 13.97 GB | 81.5% | | the 3 whitelisted packaging outputs | 25 | 3.00 GB | 17.5% | | everything else | 9,242 | 0.18 GB | 0.9% | So after this change the residue is essentially just test logs plus the three artifacts the packaging step needs. **Test logs cannot be removed by any download flag.** `test.log` is the test runner's stdout, and Bazel downloads stdout/stderr unconditionally for every cache-hit action regardless of `--remote_download_outputs`. They are the floor. Some are enormous: ``` 3586.7 MB src/v/storage/tests/storage_e2e_single_thread/test.log 1474.6 MB src/v/cluster/tests/partition_balancer_simulator_test/test.log 855.0 MB src/v/cloud_topics/tests/end_to_end_test/test.log 687.2 MB src/v/raft/tests/raft_reconfiguration_test/test.log 588.6 MB src/v/kafka/client/direct_consumer/tests/direct_consumer_test/test.log ``` We should try to address this as well, it would result in another similar relative improvement in downloaded bytes. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none [CORE-16997]: https://redpandadata.atlassian.net/browse/CORE-16997?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/pull/31536",
        "createdAt": "2026-08-11T15:30:09Z",
        "updatedAt": "2026-08-13T13:17:16Z",
        "timestamp": "2026-08-13T13:17:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [],
        "author": "travisdowns",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31538",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] bazel: define an empty ci-remote-cache config",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
        "url": "https://github.com/redpanda-data/redpanda/pull/31538",
        "createdAt": "2026-08-11T17:37:31Z",
        "updatedAt": "2026-08-12T17:09:10Z",
        "timestamp": "2026-08-12T17:09:10Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31539",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] bazel: define an empty ci-remote-cache config",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
        "url": "https://github.com/redpanda-data/redpanda/pull/31539",
        "createdAt": "2026-08-11T17:37:32Z",
        "updatedAt": "2026-08-12T15:42:46Z",
        "timestamp": "2026-08-12T15:42:46Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31540",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] bazel: define an empty ci-remote-cache config",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
        "url": "https://github.com/redpanda-data/redpanda/pull/31540",
        "createdAt": "2026-08-11T17:40:18Z",
        "updatedAt": "2026-08-13T16:42:32Z",
        "timestamp": "2026-08-13T16:42:32Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31542",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "Tq more changes",
        "text": "WIP ## Backports Required - [X] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31542",
        "createdAt": "2026-08-11T21:16:40Z",
        "updatedAt": "2026-08-12T18:50:54Z",
        "timestamp": "2026-08-12T18:50:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "WillemKauf",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31544",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[CORE-12930] - Storage: Some observability improvements for unrecoverable segments",
        "text": "<!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> A .cannotrecover file is a segment that failed startup recovery. The presence of these is a consistent source of confusion for support engineers. This PR adds some new metrics, naming, and logging around unrecoverable segments. - gauge for unrecoverable segments on disk, labelled by whether the segment was the tail of its log at recovery time (tail, mid_log, or unknown for files an older version left behind) - counted during the open_segments directory walk during recovery - try to surface a reason we couldn't recover the segment (not enough bytes, zero header, crc error, etc.) - put the reason and the position in the filename: `<segment>.<stop_reason>.<tail|mid_log>.cannotrecover`, so both are still there after the log line that reported them rotates away - in addition to the filename, log the size on disk, `log_replayer` checkpoint (incl the stop reason), and where the segment sat in the log - for non-tail segments, log at warn, otherwise keep logging at info as before. The idea here is to give a bit more signal to support or oncallers trying to determine whether cannotrecover files might indicate data loss, where to look, etc. Also adds some code to clean up previously orphaned index files from unrecoverable segments that could otherwise prevent partition removal from fully removing partition directories. Covers .base_index and .compaction_index, plus the compaction staging variant of each. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Improves observability and logging around unrecoverable segments. <!-- If the changes in this PR do not need to be mentioned in the release notes, then don't add a sub-section and simply list `none`, e.g. * none Otherwise, adding a sub-section or `none` is REQUIRED if the PR is not a backport PR. If this is a backport PR, adding contents to this section will override the release notes section inherited from the original PR to dev. Add one or more of the sub-sections with a short description bullet point of the change, e.g. ### Bug Fixes * Short description of the bug fix if this is a PR to `dev` branch. ### Features * Short description of the feature. Explain how to configure. ### Improvements * Short description of how this PR improves existing behavior. -->",
        "url": "https://github.com/redpanda-data/redpanda/pull/31544",
        "createdAt": "2026-08-12T00:24:39Z",
        "updatedAt": "2026-08-13T17:29:34Z",
        "timestamp": "2026-08-13T17:29:34Z",
        "metrics": {
          "reactions": 0,
          "comments": 2
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "claude-review"
        ],
        "author": "oleiman",
        "state": "open",
        "assignees": [
          "oleiman"
        ]
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31546",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "rpk/connect: don't cap version segments at two digits in VersionFromString",
        "text": "Fixes #31545 ## What this does `VersionFromString` (`pkg/redpanda/version.go`) parsed a version string through a regex that capped every segment at two digits: ```go regexp.MustCompile(`^v?(\\d{1,2})\\.(\\d{1,2})\\.(\\d{1,2})(?:\\s|-rc\\d{1,2}|-dev|-nightly|$)`) ``` It's called by `connect/upgrade.go`'s `connectVersion` helper to parse the *currently installed* Connect version before comparing it to latest. Connect's minor version crossed 100 in 2026, so `rpk connect upgrade` fails outright for any installed version >= 4.100.0 — i.e. every currently shipping version: ``` $ rpk connect --version Version: 4.103.1 Date: 2026-07-31T15:40:36Z $ rpk connect upgrade unable to determine current version of Redpanda Connect: unable to determine Redpanda version from \"4.103.1\": unable to get the redpanda version from \"4.103.1\" ``` This is the same bug class as #31432/CON-529 (fixed for `install.go`'s separate validator in #31438), in a sibling function that fix didn't touch. `VersionFromString` is exported from `pkg/redpanda`, so I widened `\\d{1,2}` to `\\d+` on all three segments plus the `-rc\\d{1,2}` suffix, matching #31438's fix. Filed as [CON-540](https://redpandadata.atlassian.net/browse/CON-540) / #31545. ## Test plan - Flipped the three existing test cases that asserted 3-digit segments *error* (`3 digits year/feature/patch`) to assert they now parse correctly — those cases were pinning the exact bug. - Added cases for the real-world trigger (`4.103.1`), a 3-digit segment with an `-rc` suffix, and a 3-digit segment with trailing build text. - `go test ./pkg/redpanda/...` and `go test ./pkg/cli/connect/...`: all pass. - Built `rpk` from this branch and reproduced the original failure end-to-end: `rpk connect upgrade` now correctly parses an installed `4.103.2` and upgrades to `4.104.0`, where it previously died before even fetching the manifest. ## Backports Required - [ ] none - not a bug fix - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x Same three branches #31438 was backported to — `VersionFromString` carries the defect on all of them, and this change touches only that one function, so it should cherry-pick cleanly. ## UX Changes ## Release Notes ### Bug Fixes * `rpk connect upgrade` no longer fails to determine the currently-installed Redpanda Connect version when that version has a segment of three or more digits, which had blocked upgrading any Connect install since 4.100.0. [CON-540]: https://redpandadata.atlassian.net/browse/CON-540?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/pull/31546",
        "createdAt": "2026-08-12T10:49:11Z",
        "updatedAt": "2026-08-12T15:51:36Z",
        "timestamp": "2026-08-12T15:51:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "area/rpk"
        ],
        "author": "JakeSCahill",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31547",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "raft: fix election timer starvation and busy reply term loss",
        "text": "Two raft liveness gaps let a follower storm elections against a live leader, the family of symptoms previously reported in #30815: * A follower in active recovery rejects lightweight heartbeats without refreshing its election timer, and the full heartbeat fallback after a lightweight heartbeat failure can be suppressed indefinitely by stalled in-flight appends. The follower's timer expires and it keeps starting elections that are doomed while the leader is healthy. * Full heartbeat replies fabricated as `follower_busy` carried no term and the leader dropped busy replies before its reply-term check, so a follower wedged one term above a live leader (a leadership transfer whose vote round was lost) could never depose it. Both livelocks are reproduced by new `raft_fixture` tests that fail without the fixes. A companion test pins the pre-existing invariant that rejected append entries from the current term leader still refresh the election timer. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [x] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Bug Fixes * Fixed raft livelocks where a follower could repeatedly start elections against a live leader: a recovering follower starved of election timer refreshes, and a follower wedged at a higher term whose busy heartbeat replies carried no term.",
        "url": "https://github.com/redpanda-data/redpanda/pull/31547",
        "createdAt": "2026-08-12T14:41:38Z",
        "updatedAt": "2026-08-13T10:45:43Z",
        "timestamp": "2026-08-13T10:45:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31549",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 Fixes: https://github.com/redpanda-data/redpanda/issues/31548,",
        "url": "https://github.com/redpanda-data/redpanda/pull/31549",
        "createdAt": "2026-08-12T15:51:49Z",
        "updatedAt": "2026-08-12T17:42:01Z",
        "timestamp": "2026-08-12T17:42:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/rpk",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31551",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 Fixes: https://github.com/redpanda-data/redpanda/issues/31550,",
        "url": "https://github.com/redpanda-data/redpanda/pull/31551",
        "createdAt": "2026-08-12T15:52:55Z",
        "updatedAt": "2026-08-13T16:42:31Z",
        "timestamp": "2026-08-13T16:42:31Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/rpk",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31553",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 - Command: git cherry-pick -x 8fc022e706 5f5ad1e384 - Commits backported: 2 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-31546-v25.3.x-1786550048 ## Conflict details - 8fc022e706 (src/go/rpk/pkg/redpanda/version.go): the `vMatch` regex line differed because `dev` also matches a `-nightly` suffix that v25.3.x does not have. Resolved by applying only this commit's intent — widening `\\d{1,2}` to `\\d+` for all three version segments and `-rc\\d{1,2}` to `-rc\\d+` — while keeping the target branch's suffix alternation (`\\s|-rc\\d+|-dev|$`), since `-nightly` predates this commit on `dev` and is not part of the change being backported. Fixes: https://github.com/redpanda-data/redpanda/issues/31552,",
        "url": "https://github.com/redpanda-data/redpanda/pull/31553",
        "createdAt": "2026-08-12T15:56:01Z",
        "updatedAt": "2026-08-12T18:30:33Z",
        "timestamp": "2026-08-12T18:30:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/rpk",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31554",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "rpk: add load-factor dashboard to generate grafana-dashboard",
        "text": "The Redpanda Load Factor dashboard was recently added to the [observability repo](https://github.com/redpanda-data/observability) in redpanda-data/observability#51. It shows utilization relative to capacity (load factor) for the key broker resources: CPU (reactor), IO scheduler, disk IOPS, memory, network bandwidth, and client connections, so users can identify which resource a cluster will exhaust first. It is built entirely on public metrics (`/public_metrics`). This PR exposes it through `rpk generate grafana-dashboard --dashboard load-factor`, following the same pattern as the existing dashboards: the JSON is fetched from the observability repo's `main` branch at runtime, with an embedded gzipped copy (added here, hash-verified in the unit test) as the offline fallback. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Features * `rpk generate grafana-dashboard` now offers a `load-factor` dashboard showing utilization relative to capacity for key broker resources (CPU, IO scheduler, disk IOPS, memory, network bandwidth, client connections).",
        "url": "https://github.com/redpanda-data/redpanda/pull/31554",
        "createdAt": "2026-08-12T16:49:58Z",
        "updatedAt": "2026-08-12T18:13:24Z",
        "timestamp": "2026-08-12T18:13:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [
          "area/rpk",
          "area/build"
        ],
        "author": "travisdowns",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31555",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] rpk: add load-factor dashboard to generate grafana-dashboard",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31554",
        "url": "https://github.com/redpanda-data/redpanda/pull/31555",
        "createdAt": "2026-08-12T18:15:04Z",
        "updatedAt": "2026-08-12T19:49:28Z",
        "timestamp": "2026-08-12T19:49:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/rpk",
          "area/build",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31556",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[UX-1427] chore: bump franz-go to 1.21.6",
        "text": "Particularly interested in replacing golang.org/x/crypto/pbkdf2 with the stdlib one. <!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31556",
        "createdAt": "2026-08-12T18:44:39Z",
        "updatedAt": "2026-08-13T17:14:35Z",
        "timestamp": "2026-08-13T17:14:35Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/rpk"
        ],
        "author": "r-vasquez",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31557",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "rpk: consolidate plugin version validation on pkg",
        "text": "Define the semver pattern once in pkg/redpanda: VersionFromString parses versions from command output, and a new ValidVersion strictly validates user-supplied flags. The connect, ai, check, and k8s install commands drop their private regexes and delegate to ValidVersion. The parser now accepts any prerelease or build suffix, so upgrade can parse any version that install accepts. The ai, check, and k8s validators previously capped segments at two digits and accepted trailing garbage; both bugs are gone with the shared pattern. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [X] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31557",
        "createdAt": "2026-08-12T22:12:08Z",
        "updatedAt": "2026-08-12T23:34:41Z",
        "timestamp": "2026-08-12T23:34:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/rpk",
          "area/build"
        ],
        "author": "r-vasquez",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31558",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "bazel/packaging: bake rpath at link time instead of patching",
        "text": "The packaging rules run two patchelf stages over every packaged binary, per consuming target: an rpath rewrite to `$ORIGIN/../lib` (OverrideBinaryRPath) and an interpreter rewrite (SetDynamicLoader). Because outputs are namespaced by the consuming target, a warm build of the packaging targets materializes nine patched copies of the 2.2 GB redpanda binary, each a distinct CAS blob that gets produced, uploaded to the remote cache, stored, and re-downloaded by consumers. This PR eliminates the rpath stage entirely by baking the rpath at link time instead: * `redpanda`, `rp_util` and `iotune` link with `-Wl,-rpath,$ORIGIN/../lib` via their `linkopts`. * The foreign_cc binaries get the same flag through their build systems: `hwloc` via an `LDFLAGS` env entry, `openssl` via a `Configure` command-line flag. The escaping in those two BUILD files is chosen so a literal `$ORIGIN` survives every evaluation layer (the foreign_cc script, configure, Makefile expansion, and the recipe shell); the resulting `DT_RUNPATH` is byte-identical to what patchelf used to write. The baked entry is inert in the build tree (`bin/../lib` does not exist next to the build output, so `bazel test` and dev workflows see no change), and the preexisting bazel `_solib` RUNPATH entries are inert in an install tree, so runtime behavior is identical in both contexts. With no binaries left to patch, the rpath patching machinery (`OverrideBinaryRPath`, the `rpath_override` attribute) is removed from the packaging rules; each packaged binary now produces exactly one patched output (the interpreter stage) instead of two chained ones. Savings, per warm build per configuration (sizes from the CI opt build; see the measurements in CORE-16997): * Patched redpanda copies drop from 9 (19.8 GB) to 4 (8.8 GB): the five rpath-only copies are no longer produced at all, about **11.0 GB less unique output** per build that is no longer written to disk, uploaded to and stored in the remote cache, or fetched by any cache consumer whose `--remote_download_regex` matches `bazel/packaging`. `iotune`, `hwloc-calc`, `hwloc-distrib` and `openssl` similarly lose their rpath copies (a further ~66 MB). * Five fewer patchelf passes over the 2.2 GB binary per build: roughly 11 GB fewer executor disk writes (about 22 GB less I/O counting the reads back), plus their wall clock off the packaging critical path. * On dev machines, the redpanda copies under `bazel-out` after a warm `bazel test //...` drop from 23.7 GB to about 12.7 GB. `oci_image_contents` only needed the rpath stage, so the container image now ships the binary byte-identical to the build output. The vtools release pipeline is unaffected: it takes the binaries from `bazel_out.tar.gz`, re-sets the interpreter itself, and relies on the `$ORIGIN/../lib` rpath, which is now baked in rather than patched in. Two behavioral notes: the shipped `libssl.so.3`/`libcrypto.so.3` gain a benign self-referential `$ORIGIN/../lib` runpath (their build flags changed, so expect a one-time openssl rebuild + relink), and redpanda/iotune keep their inert bazel `_solib` RUNPATH entries ahead of `$ORIGIN/../lib` where patchelf used to replace the whole list (verified those entries resolve nowhere in any install layout). The interpreter stage (SetDynamicLoader) is unchanged; deduplicating or eliminating the remaining four `redpanda_ld` copies is a follow-up. Verified locally: built all packaging targets and confirmed the rpath stage is gone from the action graph with only `_ld` outputs remaining; `readelf` shows the expected `DT_RUNPATH` and interpreter on every packaged binary (including the tuner deb's unpatched hwloc binaries); the in-tree binary runs normally; `redpanda --version` and `openssl version` run from the extracted vtools tarball tree via the shipped loader, with openssl resolving the bundled `libssl`/`libcrypto` purely via the baked rpath; an install-shaped tree without sysroot libs (the OCI configuration) resolves the bundled solibs via the baked rpath under the host loader. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31558",
        "createdAt": "2026-08-12T22:35:12Z",
        "updatedAt": "2026-08-12T23:23:40Z",
        "timestamp": "2026-08-12T23:23:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "travisdowns",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31559",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "sr: fix broker abort when a request fails before its deferred authz check",
        "text": "Schema Registry brokers abort when a request fails before its deferred authorization check. With `schema_registry_enable_authorization` set, a few endpoints only learn which resource to check against partway through handling, so the check runs inside the handler. Three individually reasonable things interact badly: 1. `request_auth_result` protects against handlers that never run an authorization check: if it is destroyed without being checked, it throws and fails the request with a 500 so that a reply which skipped authorization cannot reach the client. The throw is suppressed when another exception is already in flight (the request is already failing for some other reason). 2. The deferred-authorization handlers take a `request_auth_result` as a by-value coroutine parameter. 3. A coroutine destroys its parameters during frame teardown, which runs after the handler's own exception has already been captured into its future (meaning the handler's exception is no longer \"in flight\"), and inside the reactor's noexcept task entry point. So when a deferred-check handler fails before it can check the `request_auth_result`, the handler's exception gets caught by the coroutine machinery and stored in the future it returned. By the time the `request_auth_result` parameter is destroyed, no exception is in flight: the suppression in `request_auth_result`'s destructor does not apply, and the destructor throws a second exception in a noexcept context. `std::terminate` aborts the broker instead of the request returning an error. The fix stops the destructor from throwing (it becomes noexcept and log-only) and moves enforcement into the route wrapper, which now shares the auth result with the handler and inspects it once the handler completes: a failed handler keeps its original error, and a reply produced without the check is discarded and replaced with a 500. The admin server is unaffected because it authorizes eagerly at dispatch. Covered by unit tests and two ducktape tests. One of the ducktape tests exercises the described failure (an unreachable `_schemas` partition makes the handler throw before the auth result can be checked) and asserts the request fails with an error response while the broker keeps serving; on an unfixed build it fails with the broker abort described above. Fixes [CORE-17034](https://redpandadata.atlassian.net/browse/CORE-17034) ## Backports Required - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes ### Bug Fixes * Fixed Schema Registry aborting the broker when a request failed before its deferred authorization check with `schema_registry_enable_authorization` enabled. Such requests now return an error response. [CORE-17034]: https://redpandadata.atlassian.net/browse/CORE-17034?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
        "url": "https://github.com/redpanda-data/redpanda/pull/31559",
        "createdAt": "2026-08-12T23:01:13Z",
        "updatedAt": "2026-08-13T14:09:57Z",
        "timestamp": "2026-08-13T14:09:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "nguyen-andrew",
        "state": "open",
        "assignees": [
          "nguyen-andrew"
        ]
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31560",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "bazel: bump the toolchain sysroot to Ubuntu 24.04",
        "text": "The sysroot's `linux-libc-dev` decides which kernel uapi constants the build can see, no matter which kernel the broker actually runs on. Ours is built from Ubuntu 22.04, which pins that at 5.15 and costs us every addition since October 2021. Two of those we already reach for and do not find, and in both cases someone worked around it locally rather than the gap being caught by a build failure: - **`MADV_COLLAPSE` (6.1) is a live functional loss.** Seastar's memory prefaulter asks the kernel to synchronously collapse the freshly-populated seastar heap into transparent huge pages, guarded on `#ifdef MADV_COLLAPSE` in `src/core/smp.cc`. The constant is absent from the 22.04 sysroot, so that call is not compiled into anything we ship and the heap only gets whatever khugepaged does in the background. It is visible in the object code: `smp.pic.o` carries one `madvise` immediate (`$0x17`, `MADV_POPULATE_WRITE`) in every 22.04-built tree I have, and gains `$0x19` (25, `MADV_COLLAPSE`) with this change. Note `syschecks/hugepages.cc` already carries a local `#define MADV_COLLAPSE 25` for `promote_code_to_hugepages()`, so the gap was hit and patched in one translation unit while the seastar call site stayed dark. - **`STATX_DIOALIGN` (6.1) is no longer a functional loss, but its safety net is cut.** Seastar queries DIO alignment through a hand-rolled 256-byte `struct statx` stand-in with a local `0x2000`, so this works today (CORE-16295 / #30443). What does not work is the check on it: both `static_assert`s validating the stand-in's layout against the real struct sit inside the same `#ifdef`, so the hand-written padding has never been verified. Against real 6.8 headers they compile and pass, so there is no latent bug hiding behind them today, but nothing was keeping it that way. 24.04 carries 6.8 headers, which covers both. Two things worth stating explicitly because they look like risks and are not. **The glibc move is not one:** `//bazel/packaging` ships the sysroot's shared libraries and repoints the binaries' interpreter at the bundled loader, so the host's glibc is never consulted, and the loader's own minimum kernel is 3.2.0 on both 2.35 and 2.39. **io_uring is not gated by the sysroot** either, despite being the largest relevant slice of the header diff (74 new identifiers): seastar's backend goes through `@liburing`, which vendors its own `io_uring.h` already carrying `IORING_SETUP_SINGLE_ISSUER`, `DEFER_TASKRUN`, `REGISTER_RING_FDS` and `OP_SEND_ZC`. Bumping liburing is that lever, not this. 26.04 was measured too (glibc 2.43, 7.0 headers, same 3.2.0 floor) and deliberately skipped for now: it covers strictly more, but 24.04 clears both findings and is the smaller step for an input that invalidates every action in the build. The commits are (1) `Dockerfile.sysroot` onto `ubuntu:noble` with gcc's runtime dir moving 12 -> 13, plus README notes on the two floors a distro choice sets, and (2) the `repositories.bzl` / `MODULE.bazel.lock` pin. Verification, beyond the build passing: the tarball's file set outside `usr/include` is identical to the 22.04 sysroot on x86_64, and on aarch64 gains exactly gcc-13's `crtoffload{begin,end,table}.o` via the existing `*crt*.o` glob (additive, never linked unless offloading is used); no absolute symlinks are left for bazel's directory artifacts to reject; and `bazel build //src/v/base` builds all of seastar plus openssl and hwloc against it from scratch. **This is a draft because the tarballs are not published yet.** The third commit is marked DO NOT MERGE and carries them in-tree, handed to bazel through a distdir so the not-yet-existing `_SYSROOT_URL` is never fetched; it is here so CI can exercise the branch before anything is published. Reverting that one commit is the entire cleanup, because the pinned sha256s are already the final ones - a tarball's hash does not depend on where it is hosted. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31560",
        "createdAt": "2026-08-12T23:02:50Z",
        "updatedAt": "2026-08-13T17:29:33Z",
        "timestamp": "2026-08-13T17:29:33Z",
        "metrics": {
          "reactions": 0,
          "comments": 11
        },
        "labels": [
          "area/build"
        ],
        "author": "travisdowns",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31561",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] k/s/tests: deflake fetch_memory_units cross-shard test",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31512",
        "url": "https://github.com/redpanda-data/redpanda/pull/31561",
        "createdAt": "2026-08-13T00:34:49Z",
        "updatedAt": "2026-08-13T15:28:27Z",
        "timestamp": "2026-08-13T15:28:27Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31562",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] [ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31461",
        "url": "https://github.com/redpanda-data/redpanda/pull/31562",
        "createdAt": "2026-08-13T07:19:59Z",
        "updatedAt": "2026-08-13T07:20:01Z",
        "timestamp": "2026-08-13T07:20:01Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31563",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] [ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31461",
        "url": "https://github.com/redpanda-data/redpanda/pull/31563",
        "createdAt": "2026-08-13T07:20:01Z",
        "updatedAt": "2026-08-13T07:20:04Z",
        "timestamp": "2026-08-13T07:20:04Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31564",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] [ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31461",
        "url": "https://github.com/redpanda-data/redpanda/pull/31564",
        "createdAt": "2026-08-13T07:20:04Z",
        "updatedAt": "2026-08-13T07:20:06Z",
        "timestamp": "2026-08-13T07:20:06Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31565",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "bazel: Apply Wno-backend-plugin to abseil crc_x86_arm_combined.cc",
        "text": "We were seeing: ``` error: external/abseil-cpp+/absl/crc/internal/crc_x86_arm_combined.cc: Function _ZNK4absl12lts_2025081412crc_internal12_GLOBAL__N_145CRC32AcceleratedX86ARMCombinedMultipleStreamsILm1ELm2ELNS2_14CutoffStrategyE1EE6ExtendEPjPKvm is annotated as a hot function but the profile is cold [-Werror,-Wbackend-plugin] ``` `CRC32AcceleratedX86ARMCombinedMultipleStreams::Extend` is marked `__hot__`. However, it's a template with non-type template parameters which are chosen based on the host CPU type. In training naturally only ever one of those instantiations gets trained on. The others stay cold. LLVM emits a warning when a function marked `__hot__` has cold profile data. With our new -Werror in external files this caused a build failure. Silence the warning in that file to fix the build. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31565",
        "createdAt": "2026-08-13T09:06:26Z",
        "updatedAt": "2026-08-13T13:11:52Z",
        "timestamp": "2026-08-13T13:11:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "StephanDollberg",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31566",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "k/s/tests: deflake the offset_store producer lock test",
        "text": "`the_producer_lock_serializes_by_producer` asserted that two operations ran concurrently, which requires the first producer's operation to start before the second producer's runs. Both are scheduled as tasks in debug builds, and nothing ordered them: under `--@seastar//:shuffle_task_queue=true` the second can run to completion first (reproduced in 10/30 runs, 0/100 after the fix). Wait for the first operation to hold the lock instead. The mid-test checks also become non-fatal. `ASSERT_*_CORO` expands to `co_return`, which freed a coroutine frame the pending operations still referenced, so the assertion failure surfaced as a heap-use-after-free that masked it. Test-only change. ## Backports Required - [ ] none - papercut/not impactful enough to backport - [ ] none - not a bug fix - [ ] none - this is a backport - [x] none - issue does not exist in previous branches - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31566",
        "createdAt": "2026-08-13T13:39:22Z",
        "updatedAt": "2026-08-13T17:02:37Z",
        "timestamp": "2026-08-13T17:02:37Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31567",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "kafka/client/test: deflake data_queue blocking push test",
        "text": "`TestBlockingPushWhenFull` asserted the queue size between a `pop()` and the resumption of a push blocked on a full queue. `pop()` signals the condition variable the push waits on, so the push can resume first, which `--@seastar//:shuffle_task_queue=true` exposes about a quarter of the time. The queue itself is fine. Assert on the popped value instead, and drain the task queue before checking that the push has not completed. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31567",
        "createdAt": "2026-08-13T13:39:43Z",
        "updatedAt": "2026-08-13T16:08:57Z",
        "timestamp": "2026-08-13T16:08:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31568",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "storage: bump segment_concatenation_test timeout",
        "text": "`segment_concatenation_test` is declared `timeout = \"short\"` (60s), but the two 100-segment params take ~11s each locally, putting the whole target at 36s — 61% of its budget. The test is fsync-bound (`fsync::yes` appends plus sanitized file config), so on a CI host with slower syncs it runs past 60s and bazel kills it mid-test, which reads as a flake. Bumping to `moderate` (300s) gives ~3x headroom over the slowest observed run. `shuffle_task_queue` was incidental: measured locally it costs ~3%. The test carries the same `short` timeout and 100-segment params on all three release branches, so the flake applies there too. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes * none",
        "url": "https://github.com/redpanda-data/redpanda/pull/31568",
        "createdAt": "2026-08-13T13:40:03Z",
        "updatedAt": "2026-08-13T15:15:40Z",
        "timestamp": "2026-08-13T15:15:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [
          "area/build",
          "area/redpanda"
        ],
        "author": "nvartolomei",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31569",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v25.3.x] storage: bump segment_concatenation_test timeout",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
        "url": "https://github.com/redpanda-data/redpanda/pull/31569",
        "createdAt": "2026-08-13T15:17:14Z",
        "updatedAt": "2026-08-13T16:52:45Z",
        "timestamp": "2026-08-13T16:52:45Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31570",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] storage: bump segment_concatenation_test timeout",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
        "url": "https://github.com/redpanda-data/redpanda/pull/31570",
        "createdAt": "2026-08-13T15:17:15Z",
        "updatedAt": "2026-08-13T17:53:51Z",
        "timestamp": "2026-08-13T17:53:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31571",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.1.x] storage: bump segment_concatenation_test timeout",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
        "url": "https://github.com/redpanda-data/redpanda/pull/31571",
        "createdAt": "2026-08-13T15:17:19Z",
        "updatedAt": "2026-08-13T17:01:41Z",
        "timestamp": "2026-08-13T17:01:41Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/build",
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31572",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "[v26.2.x] kafka/client/test: deflake data_queue blocking push test",
        "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31567",
        "url": "https://github.com/redpanda-data/redpanda/pull/31572",
        "createdAt": "2026-08-13T16:10:51Z",
        "updatedAt": "2026-08-13T17:52:18Z",
        "timestamp": "2026-08-13T17:52:18Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [
          "area/redpanda",
          "kind/backport"
        ],
        "author": "vbotbuildovich",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:redpanda-data/redpanda:pull_request:31573",
        "source": "github",
        "group": "data-infrastructure",
        "project": "redpanda-data/redpanda",
        "kind": "pull_request",
        "title": "rpk: add stretch cluster grafana dashboard behind --cluster-type flag",
        "text": "Adds a `--cluster-type` flag to `rpk generate grafana-dashboard` (K8S-855, https://redpandadata.atlassian.net/browse/K8S-855). The default value (`default`) leaves the existing `--dashboard` flow untouched; `stretch` swaps in dashboard variants for stretch clusters. The stretch variant of the `operations` dashboard is a new dashboard focused on stretch cluster observability: cross-cluster raft health, StretchCluster member status, and operator reconcile health. ``` rpk generate grafana-dashboard --cluster-type stretch ``` The cluster type composes with `--dashboard` by overlaying the default dashboard set: any dashboard without a stretch-specific variant (e.g. `consumer-offsets`) is generated exactly as before. Stretch variants follow the same flow as the other dashboards — downloaded from the [observability](https://github.com/redpanda-data/observability) GitHub repository, with an embedded gzipped fallback pinned by sha256. Until the companion observability-repo change merges, the download 404s and rpk serves the byte-identical embedded copy. The dashboard is converted from the Redpanda operator repository's [`docs/operator-grafana-dashboard.json`](https://github.com/redpanda-data/redpanda-operator/blob/main/docs/operator-grafana-dashboard.json), with the Prometheus datasource parameterized as a `${DS_PROMETHEUS}` dashboard variable so it imports cleanly into any Grafana instance (matching the convention of the other embedded dashboards, e.g. serverless). Unit tests cover flag parsing, the GitHub download path (stubbed host), the embedded fallback, the overlay fallback to the default set, and the embedded gzip's integrity via sha256. ## Backports Required - [x] none - not a bug fix ## Release Notes ### Features * `rpk generate grafana-dashboard` accepts a new `--cluster-type` flag. Passing `--cluster-type stretch` generates a Grafana dashboard for stretch clusters managed by the Redpanda Operator, covering cross-cluster raft health, StretchCluster member status, and operator reconcile health. It composes with `--dashboard`: dashboards without a stretch-specific variant are generated as usual. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
        "url": "https://github.com/redpanda-data/redpanda/pull/31573",
        "createdAt": "2026-08-13T17:02:20Z",
        "updatedAt": "2026-08-13T17:51:51Z",
        "timestamp": "2026-08-13T17:51:51Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "area/rpk",
          "area/build"
        ],
        "author": "RafalKorepta",
        "state": "open",
        "assignees": [],
        "change": "updated"
      }
    ],
    "events": [
      {
        "id": "event:fbc54c873892d7beb341",
        "signalId": "github:redpanda-data/redpanda:pull_request:31568",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31568",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "storage: bump segment_concatenation_test timeout",
          "text": "`segment_concatenation_test` is declared `timeout = \"short\"` (60s), but the two 100-segment params take ~11s each locally, putting the whole target at 36s — 61% of its budget. The test is fsync-bound (`fsync::yes` appends plus sanitized file config), so on a CI host with slower syncs it runs past 60s and bazel kills it mid-test, which reads as a flake. Bumping to `moderate` (300s) gives ~3x headroom over the slowest observed run. `shuffle_task_queue` was incidental: measured locally it costs ~3%. ## Backports Required - [x] v26.2.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31568",
          "createdAt": "2026-08-13T13:40:03Z",
          "updatedAt": "2026-08-13T13:45:22Z",
          "timestamp": "2026-08-13T13:45:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:317da554addb62079e38",
        "signalId": "github:redpanda-data/redpanda:pull_request:31566",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31566",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "k/s/tests: deflake the offset_store producer lock test",
          "text": "`the_producer_lock_serializes_by_producer` asserted that two operations ran concurrently, which requires the first producer's operation to start before the second producer's runs. Both are scheduled as tasks in debug builds, and nothing ordered them: under `--@seastar//:shuffle_task_queue=true` the second can run to completion first (reproduced in 10/30 runs, 0/100 after the fix). Wait for the first operation to hold the lock instead. The mid-test checks also become non-fatal. `ASSERT_*_CORO` expands to `co_return`, which freed a coroutine frame the pending operations still referenced, so the assertion failure surfaced as a heap-use-after-free that masked it. Test-only change. ## Backports Required - [ ] none - papercut/not impactful enough to backport - [ ] none - not a bug fix - [ ] none - this is a backport - [x] none - issue does not exist in previous branches - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31566",
          "createdAt": "2026-08-13T13:39:22Z",
          "updatedAt": "2026-08-13T13:42:16Z",
          "timestamp": "2026-08-13T13:42:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:04f331d3a8756a705ad3",
        "signalId": "github:redpanda-data/redpanda:pull_request:31567",
        "event": "discovered",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31567",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "kafka/client/test: deflake data_queue blocking push test",
          "text": "`DataQueueTest.TestBlockingPushWhenFull` asserted the queue size between a `pop()` and the resumption of a push that was blocked on the queue being full. `pop()` signals the condition variable the push waits on, so the push can complete before the test thread regains control — which task ordering under `--@seastar//:shuffle_task_queue=true` makes happen roughly a quarter of the time. The queue itself behaves correctly. Assert on the popped value instead. The checks after `push_future.get()` still cover that the pop unblocked the push. 300 runs of the suite with shuffling enabled, no failures. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [x] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31567",
          "createdAt": "2026-08-13T13:39:43Z",
          "updatedAt": "2026-08-13T13:41:33Z",
          "timestamp": "2026-08-13T13:41:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:c126ef9429ff5ac1d3a7",
        "signalId": "github:redpanda-data/redpanda:pull_request:31411",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31411",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-15822] security/audit: make audit initialization controller-leader independent",
          "text": "## Problem On the RPC audit sink (the default since v25.3), `create_internal_topic()` is the **entire** configure step of `audit_client::initialize()`, and it unconditionally sends `CreateTopics` to the controller — so every audit enable/startup requires a reachable controller leader, even when the audit topic already exists. `topics_frontend::autocreate_topics` returns `no_leader_controller` immediately when no controller leader is elected (it does not wait for one), so initialization loops for as long as the leaderless state lasts. Transient churn (ordinary elections, rolling restarts with quorum) is absorbed by retries and the per-shard queues; but when the leaderless state outlasts that absorption (~90 s at default queue sizes — prolonged states of this kind are observed in production, e.g. CORE-15946), the drain timer is never armed, the queues fill and stay full, and with `audit_failure_policy=reject` authentication and the admin API are refused cluster-wide until a leader returns. This failure mode reproduces on HEAD. The Kafka-client sink (all pre-25.3 deployments, and the pre-25.3 → 25.3+ mixed-version upgrade window) has the same controller dependency twice over — its ACL write and its `CreateTopics` both need an elected leader — plus a credential-propagation design that this PR also fixes (below). Production history of this failure class: INC-2362 / CORE-15822 and CORE-12785 (full RCA with log evidence in [CORE-15822]). In both, the audit topic already existed and the produce path was viable the whole time. Those incidents ran the legacy Kafka-client sink, where initialization had an additional independent blocker (the client's all-or-nothing metadata bootstrap, fixed separately by the client rewrite in v25.2). ## Change ### RPC sink Check the local topic table (controller-log-backed, replicated to every node) before creating: if `_redpanda.audit_log` is already known, skip `CreateTopics` entirely. With the topic present — i.e. every enable/startup except the first in a cluster's life — RPC-sink initialization becomes dependency-free. ### Kafka-client sink The first version of this PR left this sink out, because its `CreateTopics` doubled as the ephemeral-credential bootstrap: the create's `sasl_authentication_failed` was the only trigger for `inform()`, which propagates the `__auditing` credential to brokers. Skipping the create without replacing that mechanism breaks authentication outright (an earlier skip-only revision failed 48 of 51 failing cases in a full audit-suite run with exactly that signature). Three changes make the sink leader-independent safely: 1. **Proactive `inform`-all** in `do_configure()`, right after minting the credential and before any broker contact. Credential propagation no longer rides on a create failure. This also fixes a standalone wedge reproduced while testing this PR: a broker that mints a fresh credential while its peers are unreachable (e.g. restarting into a leaderless window) keeps failing SASL against the returning peers (`Invalid credentials`); the kafka client aggregates those per-broker failures into `broker_error({node: -1}, broker_not_available)`, which is not `sasl_authentication_failed`, so `mitigate_error()` rethrows and `inform()` never runs — initialization loops forever (observed for minutes at the 75 s backoff cap) until a broker restart, which is exactly how the production incidents were remediated. Informs to unreachable peers are best-effort (logged, not thrown) and covered on demand by the existing mitigation path. 2. **ACL exists-skip**: `set_auditing_permissions()` returns early when both `__auditing` bindings (create + write on the audit topic) are readable in the local authorizer, eliding a raft-0 write that needs an elected leader. A stale-negative read is benign — `create_acls` is idempotent. The `create_acls` results are now also checked: its errors come back in-band (e.g. `no_leader_controller`) and were previously discarded, which the skip would have turned into a real hole — initialization could succeed with the topic present but no write ACL installed. Failing loudly keeps `initialize()` retrying until the ACL write lands. 3. **Topic exists-skip**, same as the RPC sink. Safe now that (1) owns credential propagation; and without a leader the request could not even be dispatched — `CreateTopics` must be routed to the controller broker. ### New error code handled in `update_status`: `topic_authorization_failed` Added to both the `_misconfigured` entry condition and `failed_codes`, alongside `illegal_sasl_state`. Why: if the `__auditing` ACL bindings are ever missing while auditing runs — the runtime deletion case, or the ACL exists-skip reading a stale positive — audit produces start failing authorization. Previously that state was **silent**: batches were dropped with only a WARN and the errors counter. Now it flips `_misconfigured`, so the enqueue gate degrades loudly (`Audit message rejected due to misconfigured authorization`, failure policy applied) and recovers automatically once the bindings are back. The exposure is narrow — `__auditing` is an `ephemeral_user` principal, a type the Kafka ACL API cannot express, so the bindings cannot be deleted by a principal-targeted filter — but a principal-less (wildcard) `DeleteAcls` on the topic does match them, and a silent-drop failure mode is the wrong default either way. Semantics notes: - `CreateTopics` never reconciled or altered an existing topic (`topic_already_exists` was a no-op), so skipping it cannot regress any re-creation path; a deleted topic is absent from the table and re-enable recreates it as before; - stale-negative table read (lagging, topic exists): falls through to the old path, controller answers `topic_already_exists` — degradation to today's behavior; - stale-positive read (topic deleted mid-enable): same exposure as today's create-then-delete race; recoverable by toggling `audit_enabled`; - stale-positive ACL read: degrades loudly via `topic_authorization_failed` above, recoverable by toggling `audit_enabled` (the skip then sees the bindings are gone and recreates them). ## The flow this fixes **Before** — a broker (re)starts, or auditing re-initializes, while the controller group has no elected leader: ``` init -> create_internal_topic -> needs controller-group leader -> no_leader_controller (immediately, no waiting) -> initialize() loops -> drain timer never armed -> per-shard queues fill (~90 s) -> reject refuses authn + admin API cluster-wide until a leader returns ``` **After** — same conditions, both sinks: ``` init -> exists-skips (local topic-table + authorizer reads, no network) -> configure() complete: no controller-group dependency -> drain armed -> produce goes to the audit partitions' own leaders (independent raft groups; commit = quorum of THEIR replicas) -> auditing fully functional with no controller leader at any step ``` The two paths use different leaders from different elections: `CreateTopics` needs the leader of the controller group (raft group 0), while writing audit records needs only the leaders of the `_redpanda.audit_log` partitions — which is why in both analyzed incidents the write path was healthy the entire time the init path was stuck. After this change the audit subsystem touches the controller group only for one-off events (the very first topic creation in a cluster's life, config changes), never for steady-state operation. Fine print: the very first enable still requires a controller leader (the topic and ACLs genuinely must be created), and flipping `audit_enabled` itself is a metadata mutation; the audit partitions must have elected leaders, but their elections are per-partition and do not involve the controller group. ## Testing All ducktape, all PASS on this branch: - `AuditLogLeaderlessControllerTest` (matrix over both transports): 5 brokers, the single audit partition pinned to the first three; stopping 3 of 5 leaves raft group 0 leaderless while the audit partition keeps 2/3 replicas. A surviving broker restarts into that window and must initialize through the exists-skips (no CreateTopics, no failure/retry loop, fibers start; the Kafka-client case additionally pins the ACL skip), and an admin request issued **during the window** must be consumed back from the audit topic **while the controller is still leaderless** — the end-to-end proof that a leaderless controller does not block auditing. After quorum returns, a controller-dependent event flows too, with no extra restart (the incidents required one). - `AuditLogTopicRecreateTest` (matrix over both transports): false-positive guard — with the audit topic genuinely deleted (auditing disabled, topic de-listed from `kafka_nodelete_topics`, then deleted), the skip must NOT fire; re-enable goes through `CreateTopics` and recreates it, and events flow. - `AuditLogTopicExistsTest` (matrix over both transports): re-enable with the topic present initializes without `CreateTopics` and still delivers; the Kafka-client case additionally pins the ACL skip and the proactive inform-all (`Informed:` logs). ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x v25.2.x / v25.1.x are not in the standard list, but if demand materializes they are now feasible as-is: this PR contains the Kafka-client variant (proactive inform + both skips), which those release lines run. ## Release Notes ### Improvements * Audit logging no longer requires controller availability on startup/enable when the audit topic already exists — on both the RPC transport (default since v25.3) and the Kafka-client transport. * A missing-ACL state for the audit principal now surfaces as rejected/permitted-per-policy audit events (misconfigured-authorization) instead of silently dropping records. 🤖 Generated with [Claude Code](https://claude.com/claude-code) [CORE-15822]: https://redpandadata.atlassian.net/browse/CORE-15822?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31411",
          "createdAt": "2026-08-04T15:30:22Z",
          "updatedAt": "2026-08-13T13:40:44Z",
          "timestamp": "2026-08-13T13:40:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "bartoszpiekny-redpanda",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:fae5e6898caf7750203b",
        "signalId": "github:redpanda-data/redpanda:pull_request:31301",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31301",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "metrics: remove one-shot IMDS connect probe",
          "text": "## What Removes the one-shot IMDS connect probe from the cloud instance-type detector (`instance_type_detector.cc`). ## Why The detector probed the EC2 IMDS endpoint with a single bounded `base_transport` connect before issuing any request. That was a workaround (added alongside the capacity-metrics feature in #30742) for `http::client::get_connected` retrying `connect()` in a tight **no-backoff loop** that spun hot whenever the link-local `169.254.169.254` address was unreachable, i.e. whenever we're not on EC2. The spin flooded the log and burned a shard's CPU for the whole timeout window (this was CORE-16451, and it caused the CI blowup in build #85550). #30796 (\"http: dial all resolved addresses in a single bounded attempt\") replaced that retry loop with a single bounded dial pass, so the request path no longer spins on an unreachable endpoint. The probe is now redundant: off EC2, issuing the request directly fails fast via `query_ec2_instance_type`'s normal failure path, which already returns `std::nullopt` on failure. JIRA: [CORE-16451](https://redpandadata.atlassian.net/browse/CORE-16451) ## Backports Required - [x] none - not a bug fix (removes a now-unnecessary workaround) ## Release Notes * none [CORE-16451]: https://redpandadata.atlassian.net/browse/CORE-16451?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31301",
          "createdAt": "2026-07-27T15:45:26Z",
          "updatedAt": "2026-08-13T13:20:46Z",
          "timestamp": "2026-08-13T13:20:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "travisdowns",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:792426cea408b7159595",
        "signalId": "github:redpanda-data/redpanda:pull_request:31536",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31536",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "bazel: add a remote_download_minimal config to cut CI download volume",
          "text": "Adds `--remote_download_minimal` in bazelrc that stops CI materialising outputs no CI step reads. This reduces downloaded and materialized-on-disk bytes by about 70% and shaves maybe ~1 minute off of a cached build. jira: [CORE-16997](https://redpandadata.atlassian.net/browse/CORE-16997) ## Activated later It is inert until something passes `--config=ci-remote-cache`; the vtools change to do that in CI is a separate PR. Keeping the flags here means the whitelist lives next to the rest of the repo's build config rather than in CI config (a pattern we should follow for other cache related flags when it makes sense). ## Why Bazel's default is `--remote_download_outputs=toplevel`, which materialises every top-level artifact of `//...`. On a warm-cache `bazel test //...` that is ~56 GB, and ~53 (ish) GB of it is never opened by any CI step: | what | files | size | |---|---|---| | rpbench `_binary` targets | 63 | 15.86 GB | | intermediate `.a` / `.so` | 2,411 | 12.96 GB | | duplicate copies under `bazel/packaging` | 40 | 5.27 GB | | parquet test-case JSONs | 64 | 4.42 GB | | other | 1,636 | 3.40 GB | The rpbench `_binary` targets are the largest single category at ~16 GB, some individually 2.3 GB. They are `cc_binary`, not test rules, so Bazel's TEST-mode exclusion for test rules does not cover them, and the unit-test job never runs them. ## Measured before / after Two same-commit warm-cache CI runs. Spawn count, cache-hit count and total output bytes were **identical across both** (13,864 spawns / 13,860 cache hits / 95,855,439,177 output bytes), so the deltas are attributable to the flags alone. | config | materialised | files | result | |---|---|---|---| | default (`toplevel`) | **56.05 GB** | 15,475 | green | | **`--config=ci-remote-cache`** | **17.15 GB** (**-69%**) | 11,274 | green | Bytes actually on the wire are ~75% below the materialised figures because `--remote_cache_compression` is on: **17.37 GB -> 4.66 GB, -73%**. ### Timing | config | bazel wall time | critical path | CI job | |---|---|---|---| | default | 65.70 s | 25.86 s | 4.43 min | | **`--config=ci-remote-cache`** | **38.96 s (-41%)** | 13.17 s | 3.52 min (-21%) | Single sample per arm and the arms ran on different agents of the same type, so treat the timing as indicative; the byte figures are the solid part. ## What is still materialised under the new config Of the remaining 17.15 GB: | what | files | size | share | |---|---|---|---| | test logs / xml | 2,007 | 13.97 GB | 81.5% | | the 3 whitelisted packaging outputs | 25 | 3.00 GB | 17.5% | | everything else | 9,242 | 0.18 GB | 0.9% | So after this change the residue is essentially just test logs plus the three artifacts the packaging step needs. **Test logs cannot be removed by any download flag.** `test.log` is the test runner's stdout, and Bazel downloads stdout/stderr unconditionally for every cache-hit action regardless of `--remote_download_outputs`. They are the floor. Some are enormous: ``` 3586.7 MB src/v/storage/tests/storage_e2e_single_thread/test.log 1474.6 MB src/v/cluster/tests/partition_balancer_simulator_test/test.log 855.0 MB src/v/cloud_topics/tests/end_to_end_test/test.log 687.2 MB src/v/raft/tests/raft_reconfiguration_test/test.log 588.6 MB src/v/kafka/client/direct_consumer/tests/direct_consumer_test/test.log ``` We should try to address this as well, it would result in another similar relative improvement in downloaded bytes. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none [CORE-16997]: https://redpandadata.atlassian.net/browse/CORE-16997?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31536",
          "createdAt": "2026-08-11T15:30:09Z",
          "updatedAt": "2026-08-13T13:17:16Z",
          "timestamp": "2026-08-13T13:17:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [],
          "author": "travisdowns",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1a8748853dde7d9a98d9",
        "signalId": "github:redpanda-data/redpanda:pull_request:31565",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31565",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "bazel: Apply Wno-backend-plugin to abseil crc_x86_arm_combined.cc",
          "text": "We were seeing: ``` error: external/abseil-cpp+/absl/crc/internal/crc_x86_arm_combined.cc: Function _ZNK4absl12lts_2025081412crc_internal12_GLOBAL__N_145CRC32AcceleratedX86ARMCombinedMultipleStreamsILm1ELm2ELNS2_14CutoffStrategyE1EE6ExtendEPjPKvm is annotated as a hot function but the profile is cold [-Werror,-Wbackend-plugin] ``` `CRC32AcceleratedX86ARMCombinedMultipleStreams::Extend` is marked `__hot__`. However, it's a template with non-type template parameters which are chosen based on the host CPU type. In training naturally only ever one of those instantiations gets trained on. The others stay cold. LLVM emits a warning when a function marked `__hot__` has cold profile data. With our new -Werror in external files this caused a build failure. Silence the warning in that file to fix the build. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31565",
          "createdAt": "2026-08-13T09:06:26Z",
          "updatedAt": "2026-08-13T13:11:52Z",
          "timestamp": "2026-08-13T13:11:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "StephanDollberg",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4bd5f3f7394d13acb865",
        "signalId": "github:redpanda-data/redpanda:pull_request:31366",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31366",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-16604] cluster_link: incremental Schema Registry sync by tailing _schemas",
          "text": "Schema Registry API sync only propagates changes on its periodic full sync, so a registration can wait a whole interval and every propagation costs a whole-registry scan of both sides. This adds a change feed: the link consumes the source's _schemas topic over the Kafka API and re-reads only the targets whose records it saw, reconciling them by the same per-subject path a full sync uses - so replay is harmless and semantics are identical. A CONTEXT record instead re-lists the source's contexts, since deletion is decided against the whole set. Full syncs still run on their interval as the backstop. A source that does not expose _schemas over Kafka - Confluent Cloud, Redpanda Serverless, or one too old for ListOffsets v4 - never arms. One that exposes it but refuses the fetch, e.g. a missing READ ACL, arms and then stops on its first poll. Either way the link keeps replicating on its full-sync interval alone. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Schema Registry API sync now propagates source changes by tailing the source's `_schemas` topic rather than waiting for the next full sync. Requires the link's Kafka credentials to have READ on `_schemas`; without it the link falls back to full syncs alone.",
          "url": "https://github.com/redpanda-data/redpanda/pull/31366",
          "createdAt": "2026-07-30T16:23:55Z",
          "updatedAt": "2026-08-13T12:39:57Z",
          "timestamp": "2026-08-13T12:39:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "mnajda-redpanda",
          "state": "open",
          "assignees": [
            "mnajda-redpanda"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:9e57712d41ce143b2732",
        "signalId": "github:redpanda-data/redpanda:pull_request:31139",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31139",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "bazel/seastar: enable task queue shuffling in debug builds",
          "text": "Seastar's CMake build defines `SEASTAR_SHUFFLE_TASK_QUEUE` in the Debug and Sanitize configurations, randomizing task queue execution order to surface latent task ordering assumptions. The Bazel debug build was missing this define; add it to restore the old CMake behavior. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.1.x - [ ] v25.3.x - [ ] v25.2.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31139",
          "createdAt": "2026-07-16T21:10:55Z",
          "updatedAt": "2026-08-13T12:06:00Z",
          "timestamp": "2026-08-13T12:06:00Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4b36cee392374df710b2",
        "signalId": "github:redpanda-data/redpanda:pull_request:31502",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31502",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "raft/tests: reproduce term span without configuration batch",
          "text": "The multi-term segments design assumes raft_configuration batches demarcate term boundaries. They do not. A candidate that wins an election but fails the local append of its initial configuration batch stays leader (`vote_stm.cc`: \"even if failed to replicate, don't step down\"). Nothing retries the append. Data of the new term then commits with no configuration batch for that term anywhere in the log. The test reproduces this and asserts the only durable record of the boundary is the file name of the segment rolled on term change — the safety net multi-term segments would remove. The first commit adds append failure injection to the fixture log. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31502",
          "createdAt": "2026-08-10T12:35:35Z",
          "updatedAt": "2026-08-13T10:45:49Z",
          "timestamp": "2026-08-13T10:45:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d72b120cdd3642614c6a",
        "signalId": "github:redpanda-data/redpanda:pull_request:31547",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31547",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "raft: fix election timer starvation and busy reply term loss",
          "text": "Two raft liveness gaps let a follower storm elections against a live leader, the family of symptoms previously reported in #30815: * A follower in active recovery rejects lightweight heartbeats without refreshing its election timer, and the full heartbeat fallback after a lightweight heartbeat failure can be suppressed indefinitely by stalled in-flight appends. The follower's timer expires and it keeps starting elections that are doomed while the leader is healthy. * Full heartbeat replies fabricated as `follower_busy` carried no term and the leader dropped busy replies before its reply-term check, so a follower wedged one term above a live leader (a leadership transfer whose vote round was lost) could never depose it. Both livelocks are reproduced by new `raft_fixture` tests that fail without the fixes. A companion test pins the pre-existing invariant that rejected append entries from the current term leader still refresh the election timer. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [x] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Bug Fixes * Fixed raft livelocks where a follower could repeatedly start elections against a live leader: a recovering follower starved of election timer refreshes, and a follower wedged at a higher term whose busy heartbeat replies carried no term.",
          "url": "https://github.com/redpanda-data/redpanda/pull/31547",
          "createdAt": "2026-08-12T14:41:38Z",
          "updatedAt": "2026-08-13T10:45:43Z",
          "timestamp": "2026-08-13T10:45:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0dccc97eaf86e9fd0b55",
        "signalId": "github:redpanda-data/redpanda:pull_request:31559",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31559",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "sr: fix broker abort when a request fails before its deferred authz check",
          "text": "Schema Registry brokers abort when a request fails before its deferred authorization check. With `schema_registry_enable_authorization` set, a few endpoints only learn which resource to check against partway through handling, so the check runs inside the handler. Three individually reasonable things interact badly: 1. `request_auth_result` protects against handlers that never run an authorization check: if it is destroyed without being checked, it throws and fails the request with a 500 so that a reply which skipped authorization cannot reach the client. The throw is suppressed when another exception is already in flight (the request is already failing for some other reason). 2. The deferred-authorization handlers take a `request_auth_result` as a by-value coroutine parameter. 3. A coroutine destroys its parameters during frame teardown, which runs after the handler's own exception has already been captured into its future (meaning the handler's exception is no longer \"in flight\"), and inside the reactor's noexcept task entry point. So when a deferred-check handler fails before it can check the `request_auth_result`, the handler's exception gets caught by the coroutine machinery and stored in the future it returned. By the time the `request_auth_result` parameter is destroyed, no exception is in flight: the suppression in `request_auth_result`'s destructor does not apply, and the destructor throws a second exception in a noexcept context. `std::terminate` aborts the broker instead of the request returning an error. The fix stops the destructor from throwing (it becomes noexcept and log-only) and moves enforcement into the route wrapper, which now shares the auth result with the handler and inspects it once the handler completes: a failed handler keeps its original error, and a reply produced without the check is discarded and replaced with a 500. The admin server is unaffected because it authorizes eagerly at dispatch. Covered by unit tests and two ducktape tests. One of the ducktape tests exercises the described failure (an unreachable `_schemas` partition makes the handler throw before the auth result can be checked) and asserts the request fails with an error response while the broker keeps serving; on an unfixed build it fails with the broker abort described above. Fixes [CORE-17034](https://redpandadata.atlassian.net/browse/CORE-17034) ## Backports Required - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes ### Bug Fixes * Fixed Schema Registry aborting the broker when a request failed before its deferred authorization check with `schema_registry_enable_authorization` enabled. Such requests now return an error response. [CORE-17034]: https://redpandadata.atlassian.net/browse/CORE-17034?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31559",
          "createdAt": "2026-08-12T23:01:13Z",
          "updatedAt": "2026-08-13T09:50:18Z",
          "timestamp": "2026-08-13T09:50:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "nguyen-andrew",
          "state": "open",
          "assignees": [
            "nguyen-andrew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:e83b3f64761d131b3530",
        "signalId": "github:redpanda-data/redpanda:pull_request:31564",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31564",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] [ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31461",
          "url": "https://github.com/redpanda-data/redpanda/pull/31564",
          "createdAt": "2026-08-13T07:20:04Z",
          "updatedAt": "2026-08-13T07:20:06Z",
          "timestamp": "2026-08-13T07:20:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7aed85b3e379ab8dcf54",
        "signalId": "github:redpanda-data/redpanda:pull_request:31563",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31563",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] [ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31461",
          "url": "https://github.com/redpanda-data/redpanda/pull/31563",
          "createdAt": "2026-08-13T07:20:01Z",
          "updatedAt": "2026-08-13T07:20:04Z",
          "timestamp": "2026-08-13T07:20:04Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:56734584fc7c0cfbc21b",
        "signalId": "github:redpanda-data/redpanda:pull_request:31562",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31562",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] [ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31461",
          "url": "https://github.com/redpanda-data/redpanda/pull/31562",
          "createdAt": "2026-08-13T07:19:59Z",
          "updatedAt": "2026-08-13T07:20:01Z",
          "timestamp": "2026-08-13T07:20:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cabaa506e668a97438c7",
        "signalId": "github:redpanda-data/redpanda:pull_request:31461",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31461",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[ CORE-14329] tests/node_ops: tolerate chaos testing in decommission progress check",
          "text": "`NodeDecommissionWaiter.wait_for_removal` asserts that a decommission has \"stopped making progress\" if `replicas_left`/bytes-left-to-move hasn't improved within `progress_timeout`. In tests that run `FailureInjectorBackgroundThread` alongside node operations (e.g. `ShadowLinkingRandomOpsTest.test_node_operations`), this assertion can trip even though the system is behaving correctly: the partition balancer legitimately stalls while nodes are being restarted, because each restart resets the ~30s disk-report freshness wait the balancer needs before it can compute new moves. `FailureInjectorBackgroundThread`'s default cadence (20-40s between injections) is faster than that recovery window, so a static `progress_timeout` can't distinguish \"still recovering from ongoing chaos\" from a genuine stall. This PR tracks failure-injection timestamps on `FailureInjectorBackgroundThread` and uses them in `NodeDecommissionWaiter` to hold off the \"stopped making progress\" check while chaos has been active in the last `chaos_settle_sec` (120s default), with an `absolute_max_wait_sec` (900s default) backstop so a genuine deadlock still fails in bounded time. <!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes * none <!-- If the changes in this PR do not need to be mentioned in the release notes, then don't add a sub-section and simply list `none`, e.g. * none Otherwise, adding a sub-section or `none` is REQUIRED if the PR is not a backport PR. If this is a backport PR, adding contents to this section will override the release notes section inherited from the original PR to dev. Add one or more of the sub-sections with a short description bullet point of the change, e.g. ### Bug Fixes * Short description of the bug fix if this is a PR to `dev` branch. ### Features * Short description of the feature. Explain how to configure. ### Improvements * Short description of how this PR improves existing behavior. -->",
          "url": "https://github.com/redpanda-data/redpanda/pull/31461",
          "createdAt": "2026-08-06T10:03:21Z",
          "updatedAt": "2026-08-13T07:18:28Z",
          "timestamp": "2026-08-13T07:18:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [],
          "author": "bartoszpiekny-redpanda",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c904cf81e6715a7e7545",
        "signalId": "github:redpanda-data/redpanda:pull_request:31544",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31544",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-12930] - Storage: Some observability improvements for unrecoverable segments",
          "text": "<!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> A .cannotrecover file is a segment that failed startup recovery. The presence of these is a consistent source of confusion for support engineers. This PR adds some new metrics, naming, and logging around unrecoverable segments. - gauge for unrecoverable segments on disk, labelled by whether the segment was the tail of its log at recovery time (tail, mid_log, or unknown for files an older version left behind) - counted during the open_segments directory walk during recovery - try to surface a reason we couldn't recover the segment (not enough bytes, zero header, crc error, etc.) - put the reason and the position in the filename: `<segment>.<stop_reason>.<tail|mid_log>.cannotrecover`, so both are still there after the log line that reported them rotates away - in addition to the filename, log the size on disk, `log_replayer` checkpoint (incl the stop reason), and where the segment sat in the log - for non-tail segments, log at warn, otherwise keep logging at info as before. The idea here is to give a bit more signal to support or oncallers trying to determine whether cannotrecover files might indicate data loss, where to look, etc. Also adds some code to clean up previously orphaned index files from unrecoverable segments that could otherwise prevent partition removal from fully removing partition directories. Covers .base_index and .compaction_index, plus the compaction staging variant of each. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Improves observability and logging around unrecoverable segments. <!-- If the changes in this PR do not need to be mentioned in the release notes, then don't add a sub-section and simply list `none`, e.g. * none Otherwise, adding a sub-section or `none` is REQUIRED if the PR is not a backport PR. If this is a backport PR, adding contents to this section will override the release notes section inherited from the original PR to dev. Add one or more of the sub-sections with a short description bullet point of the change, e.g. ### Bug Fixes * Short description of the bug fix if this is a PR to `dev` branch. ### Features * Short description of the feature. Explain how to configure. ### Improvements * Short description of how this PR improves existing behavior. -->",
          "url": "https://github.com/redpanda-data/redpanda/pull/31544",
          "createdAt": "2026-08-12T00:24:39Z",
          "updatedAt": "2026-08-13T06:06:07Z",
          "timestamp": "2026-08-13T06:06:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "oleiman",
          "state": "open",
          "assignees": [
            "oleiman"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:134ac03dd7b1060ba58d",
        "signalId": "github:redpanda-data/redpanda:pull_request:31483",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31483",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-16998, CORE-17000, CORE-17001] - Consumer Groups (Classic): Fix some bugs where the offset map could diverge from the log",
          "text": "CORE-16998, CORE-17000, CORE-17001 Three bugs in the classic group's offset accounting, all the same defect: live state and a replay of the log resolve the same key differently. The group's committed offset then changes on a leader change, a restart or compaction, and no client asked for it. None is a KIP-848 regression, and each commit carries its own unit test. - **CORE-16998 (High)** the retention sweep can tombstone an offset that an open transaction has staged, and when the tombstone lands last, recovery finds no offset for the partition. Fix: `filter_expired_offsets` consults the staged partitions. - **CORE-17000 (High)** a transactional commit ordered by its staging offset loses to a plain commit that raced it, so live state and the log disagree. Fix: stamp the staged offsets with the commit marker's offset. - **CORE-17001 (Medium)** a request that names one partition twice commits the first value in memory and the last in the log. Fix: commit only the last entry for a partition. **One divergence from Kafka, in 17000.** Kafka keeps the plain value for that race. It writes the transactional record at TxnOffsetCommit time, so the plain commit is the later record in its log. Both coordinators keep whatever compaction leaves, so the outcomes differ only in record placement. To match Kafka we would write our transactional records at staging time, which is a separate change. **Backports.** The release branches predate the `offset_store` extraction, so all three land in `group.cc` as `group::` methods there. We adapt each cherry-pick by hand. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes ### Bug Fixes * Consumer groups: Fix a bug where retention could expire a committed offset while an open transaction holds a new value for it. * Consumer groups: Fix a bug where a plain offset commit landing while a transaction is open would temporarily overrule the eventual transaction offset commit. * Consumer groups: Fix a bug where an OffsetCommit naming the same partition twice would lead to the in-memory store and the log to disagree on the final value.",
          "url": "https://github.com/redpanda-data/redpanda/pull/31483",
          "createdAt": "2026-08-07T07:48:46Z",
          "updatedAt": "2026-08-13T06:05:48Z",
          "timestamp": "2026-08-13T06:05:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/redpanda",
            "claude-review"
          ],
          "author": "oleiman",
          "state": "open",
          "assignees": [
            "oleiman"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:68b3616af65c192da453",
        "signalId": "github:redpanda-data/redpanda:pull_request:31365",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31365",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rptest: add OOM crash self-test; allow-list memory diagnostics",
          "text": "Revives #8570 (rebased onto dev, see #8565 for the original motivation): put the seastar memory diagnostics dump, which is emitted at ERROR level when a node runs out of memory, on the default log allow list so an OOM surfaces as the \"redpanda crashed\" diagnostic rather than a BadLogLines about the dump. Also clarifies in the raise_on_bad_logs docstrings that allow-list regexes are unanchored. The main value over the old PR is the meta-test: a new debug-only `oom` type for the `/v1/debug/trigger_crash` admin API that allocates until the seastar allocator gives out, plus a ducktape self-test (alongside the existing segfault/assert ones) that runs a node out of memory and checks the harness reports it as a NodeCrash. With the system allocator (sanitized builds) there is no seastar memory limit and unbounded allocation would exhaust host memory instead, so there the API refuses the request, and the self-test verifies the refusal and that the node survives. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.1.x - [ ] v25.3.x - [ ] v25.2.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31365",
          "createdAt": "2026-07-30T14:45:49Z",
          "updatedAt": "2026-08-13T04:53:33Z",
          "timestamp": "2026-08-13T04:53:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "travisdowns",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0e33ce6eeb475304185a",
        "signalId": "github:redpanda-data/redpanda:pull_request:31560",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31560",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "bazel: bump the toolchain sysroot to Ubuntu 24.04",
          "text": "The sysroot's `linux-libc-dev` decides which kernel uapi constants the build can see, no matter which kernel the broker actually runs on. Ours is built from Ubuntu 22.04, which pins that at 5.15 and costs us every addition since October 2021. Two of those we already reach for and do not find, and in both cases someone worked around it locally rather than the gap being caught by a build failure: - **`MADV_COLLAPSE` (6.1) is a live functional loss.** Seastar's memory prefaulter asks the kernel to synchronously collapse the freshly-populated seastar heap into transparent huge pages, guarded on `#ifdef MADV_COLLAPSE` in `src/core/smp.cc`. The constant is absent from the 22.04 sysroot, so that call is not compiled into anything we ship and the heap only gets whatever khugepaged does in the background. It is visible in the object code: `smp.pic.o` carries one `madvise` immediate (`$0x17`, `MADV_POPULATE_WRITE`) in every 22.04-built tree I have, and gains `$0x19` (25, `MADV_COLLAPSE`) with this change. Note `syschecks/hugepages.cc` already carries a local `#define MADV_COLLAPSE 25` for `promote_code_to_hugepages()`, so the gap was hit and patched in one translation unit while the seastar call site stayed dark. - **`STATX_DIOALIGN` (6.1) is no longer a functional loss, but its safety net is cut.** Seastar queries DIO alignment through a hand-rolled 256-byte `struct statx` stand-in with a local `0x2000`, so this works today (CORE-16295 / #30443). What does not work is the check on it: both `static_assert`s validating the stand-in's layout against the real struct sit inside the same `#ifdef`, so the hand-written padding has never been verified. Against real 6.8 headers they compile and pass, so there is no latent bug hiding behind them today, but nothing was keeping it that way. 24.04 carries 6.8 headers, which covers both. Two things worth stating explicitly because they look like risks and are not. **The glibc move is not one:** `//bazel/packaging` ships the sysroot's shared libraries and repoints the binaries' interpreter at the bundled loader, so the host's glibc is never consulted, and the loader's own minimum kernel is 3.2.0 on both 2.35 and 2.39. **io_uring is not gated by the sysroot** either, despite being the largest relevant slice of the header diff (74 new identifiers): seastar's backend goes through `@liburing`, which vendors its own `io_uring.h` already carrying `IORING_SETUP_SINGLE_ISSUER`, `DEFER_TASKRUN`, `REGISTER_RING_FDS` and `OP_SEND_ZC`. Bumping liburing is that lever, not this. 26.04 was measured too (glibc 2.43, 7.0 headers, same 3.2.0 floor) and deliberately skipped for now: it covers strictly more, but 24.04 clears both findings and is the smaller step for an input that invalidates every action in the build. The commits are (1) `Dockerfile.sysroot` onto `ubuntu:noble` with gcc's runtime dir moving 12 -> 13, plus README notes on the two floors a distro choice sets, and (2) the `repositories.bzl` / `MODULE.bazel.lock` pin. Verification, beyond the build passing: the tarball's file set outside `usr/include` is identical to the 22.04 sysroot on x86_64, and on aarch64 gains exactly gcc-13's `crtoffload{begin,end,table}.o` via the existing `*crt*.o` glob (additive, never linked unless offloading is used); no absolute symlinks are left for bazel's directory artifacts to reject; and `bazel build //src/v/base` builds all of seastar plus openssl and hwloc against it from scratch. **This is a draft because the tarballs are not published yet.** The third commit is marked DO NOT MERGE and carries them in-tree, handed to bazel through a distdir so the not-yet-existing `_SYSROOT_URL` is never fetched; it is here so CI can exercise the branch before anything is published. Reverting that one commit is the entire cleanup, because the pinned sha256s are already the final ones - a tarball's hash does not depend on where it is hosted. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31560",
          "createdAt": "2026-08-12T23:02:50Z",
          "updatedAt": "2026-08-13T04:32:01Z",
          "timestamp": "2026-08-13T04:32:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "area/build"
          ],
          "author": "travisdowns",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:844ab47390caab817b40",
        "signalId": "github:redpanda-data/redpanda:pull_request:31561",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31561",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] k/s/tests: deflake fetch_memory_units cross-shard test",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31512",
          "url": "https://github.com/redpanda-data/redpanda/pull/31561",
          "createdAt": "2026-08-13T00:34:49Z",
          "updatedAt": "2026-08-13T02:20:02Z",
          "timestamp": "2026-08-13T02:20:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:bdf1c25226b62ed3ab8e",
        "signalId": "github:redpanda-data/redpanda:pull_request:31512",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31512",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "k/s/tests: deflake fetch_memory_units cross-shard test",
          "text": "Units released back to their originating shard are returned via an unawaited cross-shard call, so they are not guaranteed to be visible to the immediately following read; under `--@seastar//:shuffle_task_queue=true` the read can run first (reproduced in 12/30 runs). Poll for the release instead. Test-only change. ## Backports Required - [ ] none - papercut/not impactful enough to backport - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [x] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31512",
          "createdAt": "2026-08-10T16:00:37Z",
          "updatedAt": "2026-08-13T00:33:19Z",
          "timestamp": "2026-08-13T00:33:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:987f18078764b0096710",
        "signalId": "github:redpanda-data/redpanda:pull_request:31557",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31557",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk: consolidate plugin version validation on pkg",
          "text": "Define the semver pattern once in pkg/redpanda: VersionFromString parses versions from command output, and a new ValidVersion strictly validates user-supplied flags. The connect, ai, check, and k8s install commands drop their private regexes and delegate to ValidVersion. The parser now accepts any prerelease or build suffix, so upgrade can parse any version that install accepts. The ai, check, and k8s validators previously capped segments at two digits and accepted trailing garbage; both bugs are gone with the shared pattern. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [X] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31557",
          "createdAt": "2026-08-12T22:12:08Z",
          "updatedAt": "2026-08-12T23:34:41Z",
          "timestamp": "2026-08-12T23:34:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/rpk",
            "area/build"
          ],
          "author": "r-vasquez",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5ab2833cb7d46a02ed0b",
        "signalId": "github:redpanda-data/redpanda:pull_request:31558",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31558",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "bazel/packaging: bake rpath at link time instead of patching",
          "text": "The packaging rules run two patchelf stages over every packaged binary, per consuming target: an rpath rewrite to `$ORIGIN/../lib` (OverrideBinaryRPath) and an interpreter rewrite (SetDynamicLoader). Because outputs are namespaced by the consuming target, a warm build of the packaging targets materializes nine patched copies of the 2.2 GB redpanda binary, each a distinct CAS blob that gets produced, uploaded to the remote cache, stored, and re-downloaded by consumers. This PR eliminates the rpath stage entirely by baking the rpath at link time instead: * `redpanda`, `rp_util` and `iotune` link with `-Wl,-rpath,$ORIGIN/../lib` via their `linkopts`. * The foreign_cc binaries get the same flag through their build systems: `hwloc` via an `LDFLAGS` env entry, `openssl` via a `Configure` command-line flag. The escaping in those two BUILD files is chosen so a literal `$ORIGIN` survives every evaluation layer (the foreign_cc script, configure, Makefile expansion, and the recipe shell); the resulting `DT_RUNPATH` is byte-identical to what patchelf used to write. The baked entry is inert in the build tree (`bin/../lib` does not exist next to the build output, so `bazel test` and dev workflows see no change), and the preexisting bazel `_solib` RUNPATH entries are inert in an install tree, so runtime behavior is identical in both contexts. With no binaries left to patch, the rpath patching machinery (`OverrideBinaryRPath`, the `rpath_override` attribute) is removed from the packaging rules; each packaged binary now produces exactly one patched output (the interpreter stage) instead of two chained ones. Savings, per warm build per configuration (sizes from the CI opt build; see the measurements in CORE-16997): * Patched redpanda copies drop from 9 (19.8 GB) to 4 (8.8 GB): the five rpath-only copies are no longer produced at all, about **11.0 GB less unique output** per build that is no longer written to disk, uploaded to and stored in the remote cache, or fetched by any cache consumer whose `--remote_download_regex` matches `bazel/packaging`. `iotune`, `hwloc-calc`, `hwloc-distrib` and `openssl` similarly lose their rpath copies (a further ~66 MB). * Five fewer patchelf passes over the 2.2 GB binary per build: roughly 11 GB fewer executor disk writes (about 22 GB less I/O counting the reads back), plus their wall clock off the packaging critical path. * On dev machines, the redpanda copies under `bazel-out` after a warm `bazel test //...` drop from 23.7 GB to about 12.7 GB. `oci_image_contents` only needed the rpath stage, so the container image now ships the binary byte-identical to the build output. The vtools release pipeline is unaffected: it takes the binaries from `bazel_out.tar.gz`, re-sets the interpreter itself, and relies on the `$ORIGIN/../lib` rpath, which is now baked in rather than patched in. Two behavioral notes: the shipped `libssl.so.3`/`libcrypto.so.3` gain a benign self-referential `$ORIGIN/../lib` runpath (their build flags changed, so expect a one-time openssl rebuild + relink), and redpanda/iotune keep their inert bazel `_solib` RUNPATH entries ahead of `$ORIGIN/../lib` where patchelf used to replace the whole list (verified those entries resolve nowhere in any install layout). The interpreter stage (SetDynamicLoader) is unchanged; deduplicating or eliminating the remaining four `redpanda_ld` copies is a follow-up. Verified locally: built all packaging targets and confirmed the rpath stage is gone from the action graph with only `_ld` outputs remaining; `readelf` shows the expected `DT_RUNPATH` and interpreter on every packaged binary (including the tuner deb's unpatched hwloc binaries); the in-tree binary runs normally; `redpanda --version` and `openssl version` run from the extracted vtools tarball tree via the shipped loader, with openssl resolving the bundled `libssl`/`libcrypto` purely via the baked rpath; an install-shaped tree without sysroot libs (the OCI configuration) resolves the bundled solibs via the baked rpath under the host loader. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31558",
          "createdAt": "2026-08-12T22:35:12Z",
          "updatedAt": "2026-08-12T23:23:40Z",
          "timestamp": "2026-08-12T23:23:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "travisdowns",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:80979c6d82094d461d9e",
        "signalId": "github:redpanda-data/redpanda:pull_request:31556",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31556",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[UX-1427] chore: bump franz-go to 1.21.6",
          "text": "Particularly interested in replacing golang.org/x/crypto/pbkdf2 with the stdlib one. <!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31556",
          "createdAt": "2026-08-12T18:44:39Z",
          "updatedAt": "2026-08-12T21:12:18Z",
          "timestamp": "2026-08-12T21:12:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/rpk"
          ],
          "author": "r-vasquez",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e00d1d1a864f5e92c4f9",
        "signalId": "github:redpanda-data/redpanda:pull_request:31517",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31517",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "hashing: back crc32c with abseil instead of google/crc32c",
          "text": "Backs `crc::crc32c` with abseil's CRC32C instead of google/crc32c. The wrapper becomes `basic_crc32c<Backend>` with backends for both libraries, and `crc::crc32c` is an alias for the abseil instantiation, so production switches over with no call site changes. CRC32C values are identical between the backends, which is what makes the switch safe: they are persisted in storage indices, snapshots and sst footers, and are part of the Kafka record batch wire format. The wrapper previously had no test coverage at all, so this adds a gtest suite typed over both backends that pins the absolute values to the RFC 3720 appendix B.4 vectors and the catalogued CRC-32/ISCSI check value — external references, not whatever the backing library computes — plus the properties callers rely on (incremental extends match one-shot at split points straddling both libraries' internal thresholds, the integral overload matches the equivalent bytes, `crc_extend_iobuf` is fragmentation independent). The final commit adds a benchmark that measures both backends through the wrapper, exactly as production calls them. It is the record of this decision and the way to revisit it; each fixture cross-checks that the backends agree before measuring. ## Benchmark `bazel run --config=release //src/v/hashing/tests:crc_bench_rpbench`. Times are the cost of one whole checksum, not one byte: | case | google/crc32c | abseil | abseil speedup | |---|--:|--:|--:| | batch header, 12 integer extends | 36.8ns | 2.19ns | **16.8x** | | 8B | 2.69ns | 1.34ns | 2.0x | | 57B | 4.29ns | 2.72ns | 1.6x | | 128B | 4.91ns | 5.13ns | 0.96x | | 512B | 27.0ns | 15.0ns | 1.8x | | 4032B | 113.6ns | 99.3ns | 1.14x | | 4096B | 99.4ns | 102.4ns | 0.97x | | 16256B | 447.0ns | 384.3ns | 1.16x | | 16384B | 382.9ns | 384.7ns | 1.0x | | 64KiB contiguous | 1.55µs | 1.52µs | 1.02x | | 64KiB iobuf, 512B fragments | 4.55µs | 2.46µs | **1.85x** | | 64KiB iobuf, 16KiB fragments | 1.56µs | 1.60µs | 0.98x | | 512B memcpy+crc, two-pass vs fused | 30.1ns | 20.7ns | 1.45x | | 64KiB memcpy+crc, two-pass vs fused | 2.36µs | 2.15µs | 1.10x | ## Why abseil is faster **Small extends are inlined behind a compile-time CPU check.** `absl::ExtendCrc32c` is defined in the header, and for buffers of 64 bytes or less expands in place to a straight-line sequence of hardware `crc32b/w/l/q` instructions (`crc_internal::ExtendCrc32cInline`), enabled at compile time by the SSE4.2 target baseline. `crc32c::Extend` is always an out-of-line library call that branches on a function-local `static bool can_use_sse42` and then calls the SSE4.2 kernel — for an 8-byte extend that is two calls and a dispatch check to reach a single `crc32q`, so per-call overhead dominates the actual CRC work. This is redpanda's hottest checksum shape: `model::internal_header_only_crc` feeds twelve 1–8 byte integers through the checksum for every replicated batch and pays the overhead twelve times. With abseil the entire sequence compiles to inlined crc32 instructions: 36.8ns → 2.2ns. The 128B row shows the flip side of the cutoff — just past 64B both libraries go out of line and google is marginally ahead. **Mid-size buffers get instruction-level parallelism sooner.** The `crc32` instruction has ~3-cycle latency but 1/cycle throughput, so a single dependency chain runs at a third of the hardware rate; full speed requires computing several independent streams and folding them together. google/crc32c engages its 3-way streams only in fixed block tiers — the smallest needs 3×336 = 1008 contiguous bytes, then 3×1360 and 3×5440 — and runs a serial 8-bytes-per-instruction loop below that and for whatever is left over. Abseil folds 3 streams starting at 256 bytes, and above 2KiB divides the whole buffer evenly across its streams. At 512 bytes google is fully serial while abseil is already 3-way: 27.0ns vs 15.0ns. Checksumming fragmented iobufs hits this shape repeatedly, hence 64KiB in 512B fragments going 4.55µs → 2.46µs. **Bulk throughput is a wash, but google sawtooths across its tiers.** Above ~4KiB both libraries run multi-stream at similar speed. Because google's tiers are fixed-size blocks, a length that just misses a tier drops its remainder to smaller tiers and finally the serial loop: 4032B (just under the 3×1360 tier) is ~14% slower per byte than 4096B (just over), and likewise 16256 vs 16384. Abseil's even split has no cliffs. This is also why the benchmark pins sizes on both sides of each tier boundary — a single size can rank the two libraries either way. Two footnotes recorded in the benchmark source: abseil's runtime CPU table does not recognize the benchmark host and falls back to its no-PCLMULQDQ streams for the out-of-line path, so abseil's mid/large numbers above are a floor rather than its best; and `absl::MemcpyCrc32c` (checksum fused with the copy in a single pass, non-temporal stores at large sizes) has no google/crc32c equivalent, so it is benchmarked raw against a memcpy-then-checksum two-pass. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Reduced the CPU cost of CRC32C checksumming (computed for every record batch header and for storage indices, snapshots and internal RPCs) by backing it with abseil's implementation. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/redpanda-data/redpanda/pull/31517",
          "createdAt": "2026-08-10T22:42:04Z",
          "updatedAt": "2026-08-12T20:45:32Z",
          "timestamp": "2026-08-12T20:45:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "dotnwat",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:af9d4800622cae7a3dfc",
        "signalId": "github:redpanda-data/redpanda:pull_request:31555",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31555",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] rpk: add load-factor dashboard to generate grafana-dashboard",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31554",
          "url": "https://github.com/redpanda-data/redpanda/pull/31555",
          "createdAt": "2026-08-12T18:15:04Z",
          "updatedAt": "2026-08-12T19:49:28Z",
          "timestamp": "2026-08-12T19:49:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/rpk",
            "area/build",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:accc39a58cb44d9e4d69",
        "signalId": "github:redpanda-data/redpanda:pull_request:31542",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31542",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "Tq more changes",
          "text": "WIP ## Backports Required - [X] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31542",
          "createdAt": "2026-08-11T21:16:40Z",
          "updatedAt": "2026-08-12T18:50:54Z",
          "timestamp": "2026-08-12T18:50:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "WillemKauf",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f24aa67c3640356a4f2d",
        "signalId": "github:redpanda-data/redpanda:issue:31552",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:issue:31552",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "issue",
          "title": "[v25.3.x] rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
          "text": "Backport https://github.com/redpanda-data/redpanda/issues/31545 to branch v25.3.x. Requested by PR https://github.com/redpanda-data/redpanda/pull/31546",
          "url": "https://github.com/redpanda-data/redpanda/issues/31552",
          "createdAt": "2026-08-12T15:55:59Z",
          "updatedAt": "2026-08-12T18:33:24Z",
          "timestamp": "2026-08-12T18:33:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7a72ea8306ce1af74fb8",
        "signalId": "github:redpanda-data/redpanda:pull_request:31553",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31553",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 - Command: git cherry-pick -x 8fc022e706 5f5ad1e384 - Commits backported: 2 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-31546-v25.3.x-1786550048 ## Conflict details - 8fc022e706 (src/go/rpk/pkg/redpanda/version.go): the `vMatch` regex line differed because `dev` also matches a `-nightly` suffix that v25.3.x does not have. Resolved by applying only this commit's intent — widening `\\d{1,2}` to `\\d+` for all three version segments and `-rc\\d{1,2}` to `-rc\\d+` — while keeping the target branch's suffix alternation (`\\s|-rc\\d+|-dev|$`), since `-nightly` predates this commit on `dev` and is not part of the change being backported. Fixes: https://github.com/redpanda-data/redpanda/issues/31552,",
          "url": "https://github.com/redpanda-data/redpanda/pull/31553",
          "createdAt": "2026-08-12T15:56:01Z",
          "updatedAt": "2026-08-12T18:30:33Z",
          "timestamp": "2026-08-12T18:30:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/rpk",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7185121036d0cf855fc4",
        "signalId": "github:redpanda-data/redpanda:pull_request:31554",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31554",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk: add load-factor dashboard to generate grafana-dashboard",
          "text": "The Redpanda Load Factor dashboard was recently added to the [observability repo](https://github.com/redpanda-data/observability) in redpanda-data/observability#51. It shows utilization relative to capacity (load factor) for the key broker resources: CPU (reactor), IO scheduler, disk IOPS, memory, network bandwidth, and client connections, so users can identify which resource a cluster will exhaust first. It is built entirely on public metrics (`/public_metrics`). This PR exposes it through `rpk generate grafana-dashboard --dashboard load-factor`, following the same pattern as the existing dashboards: the JSON is fetched from the observability repo's `main` branch at runtime, with an embedded gzipped copy (added here, hash-verified in the unit test) as the offline fallback. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Features * `rpk generate grafana-dashboard` now offers a `load-factor` dashboard showing utilization relative to capacity for key broker resources (CPU, IO scheduler, disk IOPS, memory, network bandwidth, client connections).",
          "url": "https://github.com/redpanda-data/redpanda/pull/31554",
          "createdAt": "2026-08-12T16:49:58Z",
          "updatedAt": "2026-08-12T18:13:24Z",
          "timestamp": "2026-08-12T18:13:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "area/rpk",
            "area/build"
          ],
          "author": "travisdowns",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1056189eac133080e5d2",
        "signalId": "github:redpanda-data/redpanda:issue:31550",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:issue:31550",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "issue",
          "title": "[v26.1.x] rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
          "text": "Backport https://github.com/redpanda-data/redpanda/issues/31545 to branch v26.1.x. Requested by PR https://github.com/redpanda-data/redpanda/pull/31546",
          "url": "https://github.com/redpanda-data/redpanda/issues/31550",
          "createdAt": "2026-08-12T15:52:53Z",
          "updatedAt": "2026-08-12T18:12:57Z",
          "timestamp": "2026-08-12T18:12:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:b4d29e2e8079db318628",
        "signalId": "github:redpanda-data/redpanda:issue:31548",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:issue:31548",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "issue",
          "title": "[v26.2.x] rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
          "text": "Backport https://github.com/redpanda-data/redpanda/issues/31545 to branch v26.2.x. Requested by PR https://github.com/redpanda-data/redpanda/pull/31546",
          "url": "https://github.com/redpanda-data/redpanda/issues/31548",
          "createdAt": "2026-08-12T15:51:46Z",
          "updatedAt": "2026-08-12T18:12:43Z",
          "timestamp": "2026-08-12T18:12:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7ce4f3cc7c9d75389edd",
        "signalId": "github:redpanda-data/redpanda:pull_request:29795",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:29795",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] Implement cross-segment prefetching for small segments in cloud storage reads",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/29496",
          "url": "https://github.com/redpanda-data/redpanda/pull/29795",
          "createdAt": "2026-03-11T10:34:42Z",
          "updatedAt": "2026-08-12T17:57:36Z",
          "timestamp": "2026-08-12T17:57:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a629219c8296452ee240",
        "signalId": "github:redpanda-data/redpanda:pull_request:30314",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30314",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] kafka/server: cap fetch memory allocation at max message size limit",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30023",
          "url": "https://github.com/redpanda-data/redpanda/pull/30314",
          "createdAt": "2026-04-28T02:54:07Z",
          "updatedAt": "2026-08-12T17:45:48Z",
          "timestamp": "2026-08-12T17:45:48Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:104433937b92c5ccc88a",
        "signalId": "github:redpanda-data/redpanda:pull_request:31549",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31549",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 Fixes: https://github.com/redpanda-data/redpanda/issues/31548,",
          "url": "https://github.com/redpanda-data/redpanda/pull/31549",
          "createdAt": "2026-08-12T15:51:49Z",
          "updatedAt": "2026-08-12T17:42:01Z",
          "timestamp": "2026-08-12T17:42:01Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/rpk",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5083a47506d8a0ab5a7a",
        "signalId": "github:redpanda-data/redpanda:pull_request:31551",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31551",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 Fixes: https://github.com/redpanda-data/redpanda/issues/31550,",
          "url": "https://github.com/redpanda-data/redpanda/pull/31551",
          "createdAt": "2026-08-12T15:52:55Z",
          "updatedAt": "2026-08-12T17:34:57Z",
          "timestamp": "2026-08-12T17:34:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/rpk",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:c11e43544126c02d7fc4",
        "signalId": "github:redpanda-data/redpanda:pull_request:31540",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31540",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] bazel: define an empty ci-remote-cache config",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
          "url": "https://github.com/redpanda-data/redpanda/pull/31540",
          "createdAt": "2026-08-11T17:40:18Z",
          "updatedAt": "2026-08-12T17:09:19Z",
          "timestamp": "2026-08-12T17:09:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:71e37e85f9bb42e7db7a",
        "signalId": "github:redpanda-data/redpanda:pull_request:31538",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31538",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] bazel: define an empty ci-remote-cache config",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
          "url": "https://github.com/redpanda-data/redpanda/pull/31538",
          "createdAt": "2026-08-11T17:37:31Z",
          "updatedAt": "2026-08-12T17:09:10Z",
          "timestamp": "2026-08-12T17:09:10Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2328768eb3982ab102ed",
        "signalId": "github:redpanda-data/redpanda:issue:4191",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:issue:4191",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "issue",
          "title": "Add rpk standalone installscript for OSX,Linux",
          "text": "### Who is this for and what problem do they have today? - Developer wants to interact with redpanda. The developer is not involved in Redpanda's hosting or installation. - Developer uses Linux distributions other than rhel/centos/fedora/ubuntu. ### What are the success criteria? - rpk can be installed with a shell one-liner on \"all\" Linux Distros and OSX - without installing redpanda server. ### Why is solving this problem impactful? - Developer Experience for the affected person is impacted. Shell one-liner is quicker than downloading the artifact; especially if the user does not use a currently supported Linux distro. ### Additional notes - This is only an improvement. The binaries are available for download from GH releases. - Binaries of rpk are already distributed in GitHub releases, so the heavy lifting is already done. Either some custom shell script, or maybe https://goreleaser.com/ can be used to provide an install script. JIRA Link: [CORE-879](https://redpandadata.atlassian.net/browse/CORE-879) [CORE-879]: https://redpandadata.atlassian.net/browse/CORE-879?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/issues/4191",
          "createdAt": "2022-04-05T07:09:39Z",
          "updatedAt": "2026-08-12T16:48:12Z",
          "timestamp": "2026-08-12T16:48:12Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [
            "kind/enhance",
            "area/rpk"
          ],
          "author": "birdayz",
          "state": "closed",
          "assignees": [
            "birdayz"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b4c774d2b9cba4d8b9e4",
        "signalId": "github:redpanda-data/redpanda:issue:2749",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:issue:2749",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "issue",
          "title": "Add healthcheck for Docker",
          "text": "Hi, I'm not sure where the Dockerfile lives and this might be the wrong repo but I'll throw the issue here in the hope that it's the best place. :) ## The goal (aka what we're trying to do): We are trying to set up a [docker-compose example with redpanda and tremor](https://github.com/tremor-rs/tremor-redpanda/) to make it easy for users to test out this integration. Where we're trying to get is that a user can clone the repo, run `docker compose up` (or `docker-compose up`) and have a working set of systems to try, then extend and play with. ## The issue (aka what we observe): At times when the startup of the different nodes is out of sync slightly, the topic is reported as an unknown topic or partition. ``` tremor_out_1 | 2021-10-22T09:05:02.208061694+00:00 ERROR tremor_runtime::source::kafka - [Source::tremor://localhost/onramp/redpanda-in/filter/out] Subscription failed: UnknownTopicOrPartition (Broker: Unknown topic or partition). Stopping. [11:08 AM] ``` ## The expectations (aka what we would have expected): The redpanda instance is started with `auto_create_topics_enabled=true` so we would have expected for unknown topics to be created and not throw an error. ## The guesswork (aka what we think happens): The `docker-compose.yaml` specifies the dependencies of the different containers so they should start in order. But looking at the [Dockerfile](https://hub.docker.com/layers/vectorized/redpanda/dev/images/sha256-ca167ad57167c52db1e16faa7e426f2510b434f8b63f86ca3cc85d8d10eab1ae?context=explore) for redpanda (let's hope it's the right one) there is no [`HEALTHCHECK`](https://docs.docker.com/engine/reference/builder/#healthcheck) specified in here. So our guess is that there is a race condition where the container is booted and docker thinks it's up but the redpanda system isn't fully bootstrapped yet. ## The dream (aka what we hope would fix this): Adding a [`HEALTHCHECK`](https://docs.docker.com/engine/reference/builder/#healthcheck) to redpandas `Dockerfile` - that said we're somewhat unsure what would be the best way to check this so would love someone with more knowledge of the internal workings to get some input. @mfelsche suggested `http://localhost:9644/v1/status/ready` so would this work: ```docker HEALTHCHECK --interval=30s --timeout=1s --start-period=5s --retries=3 CMD curl -f http://localhost:9644/v1/status/ready || exit 1 ``` ## Reproduction 1. Clone https://github.com/tremor-rs/tremor-redpanda/ 2. run drocker compose up JIRA Link: [CORE-770](https://redpandadata.atlassian.net/browse/CORE-770) [CORE-770]: https://redpandadata.atlassian.net/browse/CORE-770?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/issues/2749",
          "createdAt": "2021-10-22T09:35:30Z",
          "updatedAt": "2026-08-12T16:43:22Z",
          "timestamp": "2026-08-12T16:43:22Z",
          "metrics": {
            "reactions": 1,
            "comments": 7
          },
          "labels": [
            "kind/enhance",
            "area/build",
            "community"
          ],
          "author": "Licenser",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7b75ff0500a0150ecd25",
        "signalId": "github:redpanda-data/redpanda:pull_request:31546",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31546",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk/connect: don't cap version segments at two digits in VersionFromString",
          "text": "Fixes #31545 ## What this does `VersionFromString` (`pkg/redpanda/version.go`) parsed a version string through a regex that capped every segment at two digits: ```go regexp.MustCompile(`^v?(\\d{1,2})\\.(\\d{1,2})\\.(\\d{1,2})(?:\\s|-rc\\d{1,2}|-dev|-nightly|$)`) ``` It's called by `connect/upgrade.go`'s `connectVersion` helper to parse the *currently installed* Connect version before comparing it to latest. Connect's minor version crossed 100 in 2026, so `rpk connect upgrade` fails outright for any installed version >= 4.100.0 — i.e. every currently shipping version: ``` $ rpk connect --version Version: 4.103.1 Date: 2026-07-31T15:40:36Z $ rpk connect upgrade unable to determine current version of Redpanda Connect: unable to determine Redpanda version from \"4.103.1\": unable to get the redpanda version from \"4.103.1\" ``` This is the same bug class as #31432/CON-529 (fixed for `install.go`'s separate validator in #31438), in a sibling function that fix didn't touch. `VersionFromString` is exported from `pkg/redpanda`, so I widened `\\d{1,2}` to `\\d+` on all three segments plus the `-rc\\d{1,2}` suffix, matching #31438's fix. Filed as [CON-540](https://redpandadata.atlassian.net/browse/CON-540) / #31545. ## Test plan - Flipped the three existing test cases that asserted 3-digit segments *error* (`3 digits year/feature/patch`) to assert they now parse correctly — those cases were pinning the exact bug. - Added cases for the real-world trigger (`4.103.1`), a 3-digit segment with an `-rc` suffix, and a 3-digit segment with trailing build text. - `go test ./pkg/redpanda/...` and `go test ./pkg/cli/connect/...`: all pass. - Built `rpk` from this branch and reproduced the original failure end-to-end: `rpk connect upgrade` now correctly parses an installed `4.103.2` and upgrades to `4.104.0`, where it previously died before even fetching the manifest. ## Backports Required - [ ] none - not a bug fix - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x Same three branches #31438 was backported to — `VersionFromString` carries the defect on all of them, and this change touches only that one function, so it should cherry-pick cleanly. ## UX Changes ## Release Notes ### Bug Fixes * `rpk connect upgrade` no longer fails to determine the currently-installed Redpanda Connect version when that version has a segment of three or more digits, which had blocked upgrading any Connect install since 4.100.0. [CON-540]: https://redpandadata.atlassian.net/browse/CON-540?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31546",
          "createdAt": "2026-08-12T10:49:11Z",
          "updatedAt": "2026-08-12T15:51:36Z",
          "timestamp": "2026-08-12T15:51:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "area/rpk"
          ],
          "author": "JakeSCahill",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cd113426eecf6e6a767c",
        "signalId": "github:redpanda-data/redpanda:pull_request:29258",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:29258",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] Kgo Verifier Producer set linger to 0",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/29234",
          "url": "https://github.com/redpanda-data/redpanda/pull/29258",
          "createdAt": "2026-01-14T16:29:13Z",
          "updatedAt": "2026-08-12T15:49:42Z",
          "timestamp": "2026-08-12T15:49:42Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:59369871a07503a3b556",
        "signalId": "github:redpanda-data/redpanda:pull_request:30776",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30776",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] iceberg: Push Parquest column stats to Iceberg manifests",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30704 - Command: git cherry-pick -x 7f833530e9 b4e8232b4e - Commits backported: 2 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30704-v25.3.x-1781219157 ## Conflict details - 7f833530e9 (src/v/datalake/base_types.h): include block conflicted only on context; accepted the incoming includes (`base/format_to.h`, `bytes/bytes.h`, `container/chunked_vector.h`, `serde/envelope.h`, `serde/rw/bytes.h`) needed by the new `per_column_stats` struct. - 7f833530e9 (src/v/datalake/BUILD): accepted the incoming `base_types` library deps (`//src/v/base`, `//src/v/bytes`, `//src/v/container:chunked_vector`, `//src/v/serde`, `//src/v/serde:bytes`). - 7f833530e9 (src/v/datalake/coordinator/BUILD): kept the target's `//src/v/base` dep and added the incoming `//src/v/bytes` dep for `iceberg_file_committer`. - 7f833530e9 (src/v/serde/parquet/writer.cc): kept the two new file-level accumulators (`file_value_count`, `file_column_size_bytes`); the surrounding bloom-filter block in the incoming context does not exist on v25.3.x and was omitted (bloom filters are not present on the target branch). - 7f833530e9 (src/v/serde/parquet/column_writer.cc): the target branch lacks the stats-truncation feature, so the commit's `build_statistics` helper was adapted to the target's non-truncated form (`encode_for_stats(...)`, `is_exact=true`). Replaced per-page `record_value()` calls with the commit's `_flushed_stats.merge(_current_page_stats)`, added the new `_file_stats` collector and `file_column_stats()` accumulation, and omitted the `_bloom_filter` member (bloom filters absent on target). The second commit (b4e8232b4e, ducktape test) applied cleanly.",
          "url": "https://github.com/redpanda-data/redpanda/pull/30776",
          "createdAt": "2026-06-11T23:11:14Z",
          "updatedAt": "2026-08-12T15:49:40Z",
          "timestamp": "2026-08-12T15:49:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f8b39096b935a1bb7962",
        "signalId": "github:redpanda-data/redpanda:pull_request:30673",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30673",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] `compaction`: avoid recompression of unchanged batches",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30669 - Command: git cherry-pick -x c2abd24ceb 57e62ea1a7 - Commits backported: 2 - Conflicts resolved: 2 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30669-v25.3.x-1780368343 ## Conflict details - c2abd24ceb (src/v/compaction/types.cc): target branch still uses free `operator<<` instead of dev's `format_to` member, and lacks `expired_tombstones_discarded`; added only `compressed_batches_reused` to the existing format string and arg list. - c2abd24ceb (src/v/compaction/types.h): added `compressed_batches_reused` field; skipped the unrelated `expired_tombstones_discarded` field that exists only on dev. - c2abd24ceb (src/v/storage/compaction_reducers.cc): kept v25.3.x's narrower placeholder-install condition (only when last batch in segment — `_stm_mgr` / idempotent-window logic does not exist on this branch) while wrapping the return in the new `filtered_batch{ rebuilt, ... }`; in the second hunk switched the index entry to use `to_append` per the cherry-pick's intent. - c2abd24ceb (src/v/storage/tests/segment_deduplication_test.cc): merged the new include set with the v25.3.x-only `gmock`, `random/generators`, `storage/chunk_cache` headers; dropped the `disk_log.stm_hookset()` argument from the new test's `copy_data_segment_reducer` constructor call since v25.3.x's constructor doesn't take an stm hookset. - 57e62ea1a7 (src/v/compaction/filter.cc): merged the `operator()` decompress path with the cherry-pick's `decompress` + reuse logic; adapted `filter_and_rewrite_with_sink` to v25.3.x's two-arg sink — when the batch is unchanged and a compressed original is available, hand the original to the sink with `compression::none` so it isn't re-compressed; otherwise pass the (possibly rebuilt) batch with the original compression so the sink compresses it as before. - 57e62ea1a7 (src/v/compaction/tests/reducer_test.cc): merged the two include sets (`bytes/bytes.h` from HEAD plus the new `compaction/key.h` and `compaction/key_offset_map.h` from the cherry-pick).",
          "url": "https://github.com/redpanda-data/redpanda/pull/30673",
          "createdAt": "2026-06-02T02:52:55Z",
          "updatedAt": "2026-08-12T15:49:39Z",
          "timestamp": "2026-08-12T15:49:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8c758a54d3d426c6d017",
        "signalId": "github:redpanda-data/redpanda:pull_request:30716",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30716",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] build/deps: upgrade krb5 to 1.22.2",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30628 - Command: git cherry-pick -x 5fffb24bdc 5acac50623 ba231dd72d de492c91a7 - Commits backported: 4 - Conflicts resolved: 1 - Commits skipped (already on target): 1 - Backport branch: ai-backport-pr-30628-v25.3.x-1780530761 ## Conflict details - 5acac50623 (MODULE.bazel.lock): generated lockfile conflicted; accepted the incoming version from the PR (`git checkout --theirs`). Needs regeneration on the target branch (see below). - de492c91a7 (merge commit \"Merge branch 'dev' into fix/upgrade-kerberos-1.22\"): skipped. This is a sync merge of `dev` into the feature branch (first parent ba231dd72d was already applied, second parent is the `dev` tip). It carries no PR-specific delta, so cherry-picking it would pull unrelated `dev` changes onto the release branch. ## ⚠️ Generated files The following files were cherry-picked and may need regeneration: - MODULE.bazel.lock These files were accepted as-is from the source branch. Before merging, regenerate them on the target branch to ensure they're correct. For example: - MODULE.bazel.lock: run `bazel mod deps --lockfile_mode=update`",
          "url": "https://github.com/redpanda-data/redpanda/pull/30716",
          "createdAt": "2026-06-03T23:54:13Z",
          "updatedAt": "2026-08-12T15:49:37Z",
          "timestamp": "2026-08-12T15:49:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9f3a27c527ae18a63408",
        "signalId": "github:redpanda-data/redpanda:pull_request:30869",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30869",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] k/s/group_manager: fix inflated consumer group lag after log truncation",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30822 - Command: git cherry-pick -x a7da3be13f c3281ec0af b603df2d74 - Commits backported: 3 - Conflicts resolved: 2 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30822-v25.3.x-1782198716 ## Conflict details - a7da3be13f (src/v/cluster/BUILD): srcs/hdrs lists diverged around the insertion point; kept the target branch's `partition_balancer_types`/`partition_leaders_table` entries and added the new `partition_kafka_offsets.{cc,h}` in alphabetical order. - a7da3be13f (src/v/kafka/data/replicated_partition.cc): method ordering differs between branches (the commit deletes `partition_kafka_start_offset()` and `kafka_start_offset_with_override()`, which sit before the `stm_exact_offset_replicator`/`make_exact_offset_replicator` block on the target branch); removed the two now-extracted methods and kept the exact-offset-replicator block. The delegating edits to `start_offset()`/`high_watermark()`/`sync_effective_start()` applied cleanly. - c3281ec0af (src/v/cluster/health_monitor_backend.cc): include block diverged; the target branch had already dropped the unused `cluster/node_status_table.h` include, so only added `cluster/partition_kafka_offsets.h`. - c3281ec0af (src/v/cluster/health_monitor_types.cc): the target branch still prints `partition_status` via `operator<<` rather than the source branch's `format_to` member; applied the commit's intent by adding `log_start_offset` to the format string and argument list, keeping the `operator<<` form and `ps.` member prefixes.",
          "url": "https://github.com/redpanda-data/redpanda/pull/30869",
          "createdAt": "2026-06-23T07:16:58Z",
          "updatedAt": "2026-08-12T15:49:35Z",
          "timestamp": "2026-08-12T15:49:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:da2c9f4894b4d1fb9721",
        "signalId": "github:redpanda-data/redpanda:pull_request:31156",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31156",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] `storage`: make `_schemas` deletion exempt",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31130",
          "url": "https://github.com/redpanda-data/redpanda/pull/31156",
          "createdAt": "2026-07-17T13:30:48Z",
          "updatedAt": "2026-08-12T15:49:34Z",
          "timestamp": "2026-08-12T15:49:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ac2674adb7483347df8d",
        "signalId": "github:redpanda-data/redpanda:issue:31545",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:issue:31545",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "issue",
          "title": "rpk connect upgrade rejects all currently installed Connect versions (two-digit version regex in VersionFromString)",
          "text": "## What happened `rpk connect upgrade` fails for every currently installed Redpanda Connect version, because it parses the *current* version through a regex that caps every segment at two digits before comparing it to latest. Connect's minor version passed 100 in 2026, so upgrade is broken for every install since ~mid-2026. ``` $ rpk connect --version Version: 4.103.1 Date: 2026-07-31T15:40:36Z $ rpk connect upgrade unable to determine current version of Redpanda Connect: unable to determine Redpanda version from \"4.103.1\": unable to get the redpanda version from \"4.103.1\" ``` Reproduced on rpk **v26.1.15** (darwin arm64) — the release that already contains the #31438 fix — so this is not the same code path as #31432/CON-529, just the same bug class in a sibling function that fix didn't touch. Also present on **v26.2.1** and on **dev** as of 2026-08-12. ## Root cause `src/go/rpk/pkg/redpanda/version.go`, `VersionFromString` (called by `connect/upgrade.go`'s `connectVersion` helper): ```go vMatch := regexp.MustCompile(`^v?(\\d{1,2})\\.(\\d{1,2})\\.(\\d{1,2})(?:\\s|-rc\\d{1,2}|-dev|-nightly|$)`).FindStringSubmatch(s) ``` Same defect as #31432: `\\d{1,2}` caps every segment at two digits, so parsing the installed version string (e.g. `4.103.1`) fails outright. `upgrade.go` calls this on the *currently installed* version before it can even fetch/compare against latest, so the command dies before doing anything else. ## Suggested fix Apply the same fix #31438 used for `install.go`'s regex: ```go regexp.MustCompile(`^v?(\\d+)\\.(\\d+)\\.(\\d+)(?:\\s|-rc\\d+|-dev|-nightly|$)`) ``` `VersionFromString` is exported from `pkg/redpanda` — worth grepping for other call sites that may hit the same wall (it's used for more than just Connect version parsing). ## Impact - `rpk connect upgrade` cannot run against any Connect install >= 4.100.0 — the primary way most users are told to update (\"If you want to upgrade to the latest version, please run 'rpk connect upgrade'\" is `install`'s own suggested next step). - Anything that shells out to `rpk connect upgrade` to ensure a plugin is current before reading its version breaks the same way; docs automation worked around it by calling `rpk connect install --force` instead (`install`'s default `\"latest\"` path never reaches this regex). Found while reproducing and fixing CON-529/#31432 for docs automation (DOC-2417) — installing a pinned >= 4.100.0 version to unblock the first bug immediately exposed this second one on the very next command.",
          "url": "https://github.com/redpanda-data/redpanda/issues/31545",
          "createdAt": "2026-08-12T10:28:09Z",
          "updatedAt": "2026-08-12T15:49:26Z",
          "timestamp": "2026-08-12T15:49:26Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "JakeSCahill",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ef1acac078eb26d8d648",
        "signalId": "github:redpanda-data/redpanda:pull_request:31521",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31521",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk: fix grammar and wording defects in help text",
          "text": "## Summary Copyedits surfaced by auditing the docs team's rpk description overrides (`docs-data/rpk-overrides.json` in redpanda-data/docs) against rpk's source text — these are cases where the docs override exists only to paper over a source defect, so fixing the source lets the docs drop the override: - `profile clear`: \"unset **an prod** cluster profile\" → \"a production\" - `registry mode get`/`set`, `registry compatibility-level`: fix the \"global mode**, alternatively** you can…\" comma splice (same pattern in all three) - `security acl delete --cluster`: ACLs are removed *for*, not *to*, the cluster - `security acl list --registry-global`/`--registry-subject`: the list command **matches** ACLs, it doesn't **grant** them (text was copied from `acl create`, where \"grant\" is correct and unchanged) - `container start --retries`, `topic consume --fetch-max-bytes`/`--fetch-max-partition-bytes`: \"amount of\" → \"number of\" for countable nouns Help-text strings only; no behavior changes. Affected package tests pass. Related: #31520 (single-sources the `-X` flag docs from the same audit). 🤖 Generated with [Claude Code](https://claude.com/claude-code) ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31521",
          "createdAt": "2026-08-11T09:30:20Z",
          "updatedAt": "2026-08-12T15:48:46Z",
          "timestamp": "2026-08-12T15:48:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "area/rpk"
          ],
          "author": "JakeSCahill",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8f2b38a5901a8cf10ea4",
        "signalId": "github:redpanda-data/redpanda:pull_request:30777",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30777",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] iceberg: Push Parquest column stats to Iceberg manifests",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30704 - Command: git cherry-pick -x 7f833530e9 b4e8232b4e - Commits backported: 2 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-30704-v26.1.x-1781219179 ## Conflict details - 7f833530e9 (src/v/serde/parquet/writer.cc): only the two new file-level stat accumulation lines (`file_value_count`/`file_column_size_bytes`) were applied; the incoming bloom-filter block is a separate feature not present on v26.1.x and was omitted. - 7f833530e9 (src/v/serde/parquet/column_writer.cc): the commit was built atop a stats-truncation feature (`truncate_max`/`truncate_min`, `max_stats_truncate_length`, `is_utf8_string`) that does not exist on v26.1.x. Adapted `build_statistics()` and `flush_page()` to the target's non-truncating bound logic, dropped the redundant `_flushed_stats.record_value()` calls in favor of `merge()`, added the `_file_stats` collector, and omitted the unrelated `_bloom_filter` member. - 7f833530e9 (src/v/datalake/base_types.h): merged the new includes needed by `per_column_stats` (bytes, chunked_vector, serde envelope/rw); omitted `base/format_to.h` since v26.1.x's `local_file_metadata` still uses `operator<<` rather than `format_to`. - 7f833530e9 (src/v/datalake/BUILD): merged the new `base_types` deps (base, bytes, container:chunked_vector, serde, serde:bytes) into the target's deps list. - 7f833530e9 (src/v/datalake/coordinator/BUILD): kept both the target's `//src/v/base` dep and the incoming `//src/v/bytes` dep in `iceberg_file_committer`'s implementation_deps.",
          "url": "https://github.com/redpanda-data/redpanda/pull/30777",
          "createdAt": "2026-06-11T23:12:24Z",
          "updatedAt": "2026-08-12T15:43:24Z",
          "timestamp": "2026-08-12T15:43:24Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:cc5ec93541574a5f05e8",
        "signalId": "github:redpanda-data/redpanda:pull_request:29280",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:29280",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] Add comparison operators to iobuf fuzz test (and other enchancements)",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/29200",
          "url": "https://github.com/redpanda-data/redpanda/pull/29280",
          "createdAt": "2026-01-15T21:28:39Z",
          "updatedAt": "2026-08-12T15:42:49Z",
          "timestamp": "2026-08-12T15:42:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8a14960a553d2f8a86ea",
        "signalId": "github:redpanda-data/redpanda:pull_request:31539",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31539",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] bazel: define an empty ci-remote-cache config",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
          "url": "https://github.com/redpanda-data/redpanda/pull/31539",
          "createdAt": "2026-08-11T17:37:32Z",
          "updatedAt": "2026-08-12T15:42:46Z",
          "timestamp": "2026-08-12T15:42:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:68d1c7b2a9d6e828aa08",
        "signalId": "github:redpanda-data/redpanda:pull_request:30531",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:30531",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] config: refresh iceberg_enabled docstring",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/30529",
          "url": "https://github.com/redpanda-data/redpanda/pull/30531",
          "createdAt": "2026-05-19T15:12:12Z",
          "updatedAt": "2026-08-12T15:41:49Z",
          "timestamp": "2026-08-12T15:41:49Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3887c7341940d77c479e",
        "signalId": "github:redpanda-data/redpanda:pull_request:31186",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31186",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] storage: mark snapshot writer/reader closed regardless of close outcome",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31180",
          "url": "https://github.com/redpanda-data/redpanda/pull/31186",
          "createdAt": "2026-07-20T18:05:16Z",
          "updatedAt": "2026-08-12T15:41:46Z",
          "timestamp": "2026-08-12T15:41:46Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:5fcb4d714a95652301ca",
        "signalId": "github:redpanda-data/redpanda:pull_request:31192",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31192",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] kafka/client: authenticate under the reconnect mutex",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31152 - Command: git cherry-pick -x b08c929724 - Commits backported: 1 - Conflicts resolved: 1 - Commits skipped (already on target): 0 - Backport branch: ai-backport-pr-31152-v26.1.x-1784570815 ## Conflict details - b08c929724 (src/v/kafka/client/test/BUILD): the commit added the granular `//src/v/cluster:*` sub-targets (`client_quota_backend`, `cluster_utils`, `controller_log_limiter`, `data_migration_table`, `feature_backend`, `leader_balancer_strategy`, `node_status_backend`, `self_test`, `topic_metrics_watcher`) plus `//src/v/config` and `//src/v/kafka/client:api_types` to the `cluster_test` deps. On v26.1.x `//src/v/cluster` is still a single monolithic target and those fine-grained sub-targets do not exist, so the granular deps were collapsed to the single `//src/v/cluster` target (the pattern already used by the other test targets in this file), keeping `//src/v/config` and `//src/v/kafka/client:api_types`. This satisfies every include added to cluster_test.cc (`cluster/security_frontend.h` is provided by `//src/v/cluster`).",
          "url": "https://github.com/redpanda-data/redpanda/pull/31192",
          "createdAt": "2026-07-20T18:11:03Z",
          "updatedAt": "2026-08-12T15:41:44Z",
          "timestamp": "2026-08-12T15:41:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:960e083175b79352f733",
        "signalId": "github:redpanda-data/redpanda:pull_request:31572",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31572",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] kafka/client/test: deflake data_queue blocking push test",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31567",
          "url": "https://github.com/redpanda-data/redpanda/pull/31572",
          "createdAt": "2026-08-13T16:10:51Z",
          "updatedAt": "2026-08-13T16:14:06Z",
          "timestamp": "2026-08-13T16:14:06Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:c00935382a6598c12d82",
        "signalId": "github:redpanda-data/redpanda:pull_request:31567",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31567",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "kafka/client/test: deflake data_queue blocking push test",
          "text": "`TestBlockingPushWhenFull` asserted the queue size between a `pop()` and the resumption of a push blocked on a full queue. `pop()` signals the condition variable the push waits on, so the push can resume first, which `--@seastar//:shuffle_task_queue=true` exposes about a quarter of the time. The queue itself is fine. Assert on the popped value instead, and drain the task queue before checking that the push has not completed. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31567",
          "createdAt": "2026-08-13T13:39:43Z",
          "updatedAt": "2026-08-13T16:08:57Z",
          "timestamp": "2026-08-13T16:08:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0cf640c7b91748b5725e",
        "signalId": "github:redpanda-data/redpanda:pull_request:31566",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31566",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "k/s/tests: deflake the offset_store producer lock test",
          "text": "`the_producer_lock_serializes_by_producer` asserted that two operations ran concurrently, which requires the first producer's operation to start before the second producer's runs. Both are scheduled as tasks in debug builds, and nothing ordered them: under `--@seastar//:shuffle_task_queue=true` the second can run to completion first (reproduced in 10/30 runs, 0/100 after the fix). Wait for the first operation to hold the lock instead. The mid-test checks also become non-fatal. `ASSERT_*_CORO` expands to `co_return`, which freed a coroutine frame the pending operations still referenced, so the assertion failure surfaced as a heap-use-after-free that masked it. Test-only change. ## Backports Required - [ ] none - papercut/not impactful enough to backport - [ ] none - not a bug fix - [ ] none - this is a backport - [x] none - issue does not exist in previous branches - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31566",
          "createdAt": "2026-08-13T13:39:22Z",
          "updatedAt": "2026-08-13T15:32:39Z",
          "timestamp": "2026-08-13T15:32:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e12c31ff9e017de0be60",
        "signalId": "github:redpanda-data/redpanda:pull_request:31561",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31561",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] k/s/tests: deflake fetch_memory_units cross-shard test",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31512",
          "url": "https://github.com/redpanda-data/redpanda/pull/31561",
          "createdAt": "2026-08-13T00:34:49Z",
          "updatedAt": "2026-08-13T15:28:27Z",
          "timestamp": "2026-08-13T15:28:27Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7bc40b92ec820774c2b5",
        "signalId": "github:redpanda-data/redpanda:pull_request:31365",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31365",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rptest: add OOM crash self-test; allow-list memory diagnostics",
          "text": "Revives #8570 (rebased onto dev, see #8565 for the original motivation): put the seastar memory diagnostics dump, which is emitted at ERROR level when a node runs out of memory, on the default log allow list so an OOM surfaces as the \"redpanda crashed\" diagnostic rather than a BadLogLines about the dump. Also clarifies in the raise_on_bad_logs docstrings that allow-list regexes are unanchored. The main value over the old PR is the meta-test: a new debug-only `oom` type for the `/v1/debug/trigger_crash` admin API that allocates until the seastar allocator gives out, plus a ducktape self-test (alongside the existing segfault/assert ones) that runs a node out of memory and checks the harness reports it as a NodeCrash. With the system allocator (sanitized builds) there is no seastar memory limit and unbounded allocation would exhaust host memory instead, so there the API refuses the request, and the self-test verifies the refusal and that the node survives. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.1.x - [ ] v25.3.x - [ ] v25.2.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31365",
          "createdAt": "2026-07-30T14:45:49Z",
          "updatedAt": "2026-08-13T15:20:55Z",
          "timestamp": "2026-08-13T15:20:55Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "travisdowns",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:920e8cd1d140790de685",
        "signalId": "github:redpanda-data/redpanda:pull_request:31570",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31570",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] storage: bump segment_concatenation_test timeout",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
          "url": "https://github.com/redpanda-data/redpanda/pull/31570",
          "createdAt": "2026-08-13T15:17:15Z",
          "updatedAt": "2026-08-13T15:18:45Z",
          "timestamp": "2026-08-13T15:18:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:574e857005b4c4c2a6f4",
        "signalId": "github:redpanda-data/redpanda:pull_request:31569",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31569",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] storage: bump segment_concatenation_test timeout",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
          "url": "https://github.com/redpanda-data/redpanda/pull/31569",
          "createdAt": "2026-08-13T15:17:14Z",
          "updatedAt": "2026-08-13T15:18:38Z",
          "timestamp": "2026-08-13T15:18:38Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:3d2af8edf3bc19e902ce",
        "signalId": "github:redpanda-data/redpanda:pull_request:31571",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31571",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] storage: bump segment_concatenation_test timeout",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
          "url": "https://github.com/redpanda-data/redpanda/pull/31571",
          "createdAt": "2026-08-13T15:17:19Z",
          "updatedAt": "2026-08-13T15:18:29Z",
          "timestamp": "2026-08-13T15:18:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:b849a41311488ee6a139",
        "signalId": "github:redpanda-data/redpanda:pull_request:31568",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31568",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "storage: bump segment_concatenation_test timeout",
          "text": "`segment_concatenation_test` is declared `timeout = \"short\"` (60s), but the two 100-segment params take ~11s each locally, putting the whole target at 36s — 61% of its budget. The test is fsync-bound (`fsync::yes` appends plus sanitized file config), so on a CI host with slower syncs it runs past 60s and bazel kills it mid-test, which reads as a flake. Bumping to `moderate` (300s) gives ~3x headroom over the slowest observed run. `shuffle_task_queue` was incidental: measured locally it costs ~3%. The test carries the same `short` timeout and 100-segment params on all three release branches, so the flake applies there too. ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31568",
          "createdAt": "2026-08-13T13:40:03Z",
          "updatedAt": "2026-08-13T15:15:40Z",
          "timestamp": "2026-08-13T15:15:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f6613f0b0eed00f56805",
        "signalId": "github:redpanda-data/redpanda:pull_request:31411",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31411",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-15822] security/audit: make audit initialization controller-leader independent",
          "text": "## Problem On the RPC audit sink (the default since v25.3), `create_internal_topic()` is the **entire** configure step of `audit_client::initialize()`, and it unconditionally sends `CreateTopics` to the controller — so every audit enable/startup requires a reachable controller leader, even when the audit topic already exists. `topics_frontend::autocreate_topics` returns `no_leader_controller` immediately when no controller leader is elected (it does not wait for one), so initialization loops for as long as the leaderless state lasts. Transient churn (ordinary elections, rolling restarts with quorum) is absorbed by retries and the per-shard queues; but when the leaderless state outlasts that absorption (~90 s at default queue sizes — prolonged states of this kind are observed in production, e.g. CORE-15946), the drain timer is never armed, the queues fill and stay full, and with `audit_failure_policy=reject` authentication and the admin API are refused cluster-wide until a leader returns. This failure mode reproduces on HEAD. The Kafka-client sink (all pre-25.3 deployments, and the pre-25.3 → 25.3+ mixed-version upgrade window) has the same controller dependency twice over — its ACL write and its `CreateTopics` both need an elected leader — plus a credential-propagation design that this PR also fixes (below). Production history of this failure class: INC-2362 / CORE-15822 and CORE-12785 (full RCA with log evidence in [CORE-15822]). In both, the audit topic already existed and the produce path was viable the whole time. Those incidents ran the legacy Kafka-client sink, where initialization had an additional independent blocker (the client's all-or-nothing metadata bootstrap, fixed separately by the client rewrite in v25.2). ## Change ### RPC sink Check the local topic table (controller-log-backed, replicated to every node) before creating: if `_redpanda.audit_log` is already known, skip `CreateTopics` entirely. With the topic present — i.e. every enable/startup except the first in a cluster's life — RPC-sink initialization becomes dependency-free. ### Kafka-client sink The first version of this PR left this sink out, because its `CreateTopics` doubled as the ephemeral-credential bootstrap: the create's `sasl_authentication_failed` was the only trigger for `inform()`, which propagates the `__auditing` credential to brokers. Skipping the create without replacing that mechanism breaks authentication outright (an earlier skip-only revision failed 48 of 51 failing cases in a full audit-suite run with exactly that signature). Three changes make the sink leader-independent safely: 1. **Proactive `inform`-all** in `do_configure()`, right after minting the credential and before any broker contact. Credential propagation no longer rides on a create failure. This also fixes a standalone wedge reproduced while testing this PR: a broker that mints a fresh credential while its peers are unreachable (e.g. restarting into a leaderless window) keeps failing SASL against the returning peers (`Invalid credentials`); the kafka client aggregates those per-broker failures into `broker_error({node: -1}, broker_not_available)`, which is not `sasl_authentication_failed`, so `mitigate_error()` rethrows and `inform()` never runs — initialization loops forever (observed for minutes at the 75 s backoff cap) until a broker restart, which is exactly how the production incidents were remediated. Informs to unreachable peers are best-effort (logged, not thrown) and covered on demand by the existing mitigation path. 2. **ACL exists-skip**: `set_auditing_permissions()` returns early when both `__auditing` bindings (create + write on the audit topic) are readable in the local authorizer, eliding a raft-0 write that needs an elected leader. A stale-negative read is benign — `create_acls` is idempotent. The `create_acls` results are now also checked: its errors come back in-band (e.g. `no_leader_controller`) and were previously discarded, which the skip would have turned into a real hole — initialization could succeed with the topic present but no write ACL installed. Failing loudly keeps `initialize()` retrying until the ACL write lands. 3. **Topic exists-skip**, same as the RPC sink. Safe now that (1) owns credential propagation; and without a leader the request could not even be dispatched — `CreateTopics` must be routed to the controller broker. ### New error code handled in `update_status`: `topic_authorization_failed` Added to both the `_misconfigured` entry condition and `failed_codes`, alongside `illegal_sasl_state`. Why: if the `__auditing` ACL bindings are ever missing while auditing runs — the runtime deletion case, or the ACL exists-skip reading a stale positive — audit produces start failing authorization. Previously that state was **silent**: batches were dropped with only a WARN and the errors counter. Now it flips `_misconfigured`, so the enqueue gate degrades loudly (`Audit message rejected due to misconfigured authorization`, failure policy applied) and recovers automatically once the bindings are back. The exposure is narrow — `__auditing` is an `ephemeral_user` principal, a type the Kafka ACL API cannot express, so the bindings cannot be deleted by a principal-targeted filter — but a principal-less (wildcard) `DeleteAcls` on the topic does match them, and a silent-drop failure mode is the wrong default either way. Semantics notes: - `CreateTopics` never reconciled or altered an existing topic (`topic_already_exists` was a no-op), so skipping it cannot regress any re-creation path; a deleted topic is absent from the table and re-enable recreates it as before; - stale-negative table read (lagging, topic exists): falls through to the old path, controller answers `topic_already_exists` — degradation to today's behavior; - stale-positive read (topic deleted mid-enable): same exposure as today's create-then-delete race; recoverable by toggling `audit_enabled`; - stale-positive ACL read: degrades loudly via `topic_authorization_failed` above, recoverable by toggling `audit_enabled` (the skip then sees the bindings are gone and recreates them). ## The flow this fixes **Before** — a broker (re)starts, or auditing re-initializes, while the controller group has no elected leader: ``` init -> create_internal_topic -> needs controller-group leader -> no_leader_controller (immediately, no waiting) -> initialize() loops -> drain timer never armed -> per-shard queues fill (~90 s) -> reject refuses authn + admin API cluster-wide until a leader returns ``` **After** — same conditions, both sinks: ``` init -> exists-skips (local topic-table + authorizer reads, no network) -> configure() complete: no controller-group dependency -> drain armed -> produce goes to the audit partitions' own leaders (independent raft groups; commit = quorum of THEIR replicas) -> auditing fully functional with no controller leader at any step ``` The two paths use different leaders from different elections: `CreateTopics` needs the leader of the controller group (raft group 0), while writing audit records needs only the leaders of the `_redpanda.audit_log` partitions — which is why in both analyzed incidents the write path was healthy the entire time the init path was stuck. After this change the audit subsystem touches the controller group only for one-off events (the very first topic creation in a cluster's life, config changes), never for steady-state operation. Fine print: the very first enable still requires a controller leader (the topic and ACLs genuinely must be created), and flipping `audit_enabled` itself is a metadata mutation; the audit partitions must have elected leaders, but their elections are per-partition and do not involve the controller group. ## Testing All ducktape, all PASS on this branch: - `AuditLogLeaderlessControllerTest` (matrix over both transports): 5 brokers, the single audit partition pinned to the first three; stopping 3 of 5 leaves raft group 0 leaderless while the audit partition keeps 2/3 replicas. A surviving broker restarts into that window and must initialize through the exists-skips (no CreateTopics, no failure/retry loop, fibers start; the Kafka-client case additionally pins the ACL skip), and an admin request issued **during the window** must be consumed back from the audit topic **while the controller is still leaderless** — the end-to-end proof that a leaderless controller does not block auditing. After quorum returns, a controller-dependent event flows too, with no extra restart (the incidents required one). - `AuditLogTopicRecreateTest` (matrix over both transports): false-positive guard — with the audit topic genuinely deleted (auditing disabled, topic de-listed from `kafka_nodelete_topics`, then deleted), the skip must NOT fire; re-enable goes through `CreateTopics` and recreates it, and events flow. - `AuditLogTopicExistsTest` (matrix over both transports): re-enable with the topic present initializes without `CreateTopics` and still delivers; the Kafka-client case additionally pins the ACL skip and the proactive inform-all (`Informed:` logs). ## Backports Required - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x v25.2.x / v25.1.x are not in the standard list, but if demand materializes they are now feasible as-is: this PR contains the Kafka-client variant (proactive inform + both skips), which those release lines run. ## Release Notes ### Improvements * Audit logging no longer requires controller availability on startup/enable when the audit topic already exists — on both the RPC transport (default since v25.3) and the Kafka-client transport. * A missing-ACL state for the audit principal now surfaces as rejected/permitted-per-policy audit events (misconfigured-authorization) instead of silently dropping records. 🤖 Generated with [Claude Code](https://claude.com/claude-code) [CORE-15822]: https://redpandadata.atlassian.net/browse/CORE-15822?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31411",
          "createdAt": "2026-08-04T15:30:22Z",
          "updatedAt": "2026-08-13T14:54:39Z",
          "timestamp": "2026-08-13T14:54:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 9
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "bartoszpiekny-redpanda",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:366177146fd639f4fc68",
        "signalId": "github:redpanda-data/redpanda:pull_request:31366",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31366",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-16604] cluster_link: incremental Schema Registry sync by tailing _schemas",
          "text": "Schema Registry API sync only propagates changes on its periodic full sync, so a registration can wait a whole interval and every propagation costs a whole-registry scan of both sides. This adds a change feed: the link consumes the source's _schemas topic over the Kafka API and re-reads only the targets whose records it saw, reconciling them by the same per-subject path a full sync uses - so replay is harmless and semantics are identical. A CONTEXT record instead re-lists the source's contexts, since deletion is decided against the whole set. Full syncs still run on their interval as the backstop. A source that does not expose _schemas over Kafka - Confluent Cloud, Redpanda Serverless, or one too old for ListOffsets v4 - never arms. One that exposes it but refuses the fetch, e.g. a missing READ ACL, arms and then stops on its first poll. Either way the link keeps replicating on its full-sync interval alone. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Schema Registry API sync now propagates source changes by tailing the source's `_schemas` topic rather than waiting for the next full sync. Requires the link's Kafka credentials to have READ on `_schemas`; without it the link falls back to full syncs alone.",
          "url": "https://github.com/redpanda-data/redpanda/pull/31366",
          "createdAt": "2026-07-30T16:23:55Z",
          "updatedAt": "2026-08-13T14:13:02Z",
          "timestamp": "2026-08-13T14:13:02Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "mnajda-redpanda",
          "state": "open",
          "assignees": [
            "mnajda-redpanda"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:1211b2d767c71909ef01",
        "signalId": "github:redpanda-data/redpanda:pull_request:31559",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31559",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "sr: fix broker abort when a request fails before its deferred authz check",
          "text": "Schema Registry brokers abort when a request fails before its deferred authorization check. With `schema_registry_enable_authorization` set, a few endpoints only learn which resource to check against partway through handling, so the check runs inside the handler. Three individually reasonable things interact badly: 1. `request_auth_result` protects against handlers that never run an authorization check: if it is destroyed without being checked, it throws and fails the request with a 500 so that a reply which skipped authorization cannot reach the client. The throw is suppressed when another exception is already in flight (the request is already failing for some other reason). 2. The deferred-authorization handlers take a `request_auth_result` as a by-value coroutine parameter. 3. A coroutine destroys its parameters during frame teardown, which runs after the handler's own exception has already been captured into its future (meaning the handler's exception is no longer \"in flight\"), and inside the reactor's noexcept task entry point. So when a deferred-check handler fails before it can check the `request_auth_result`, the handler's exception gets caught by the coroutine machinery and stored in the future it returned. By the time the `request_auth_result` parameter is destroyed, no exception is in flight: the suppression in `request_auth_result`'s destructor does not apply, and the destructor throws a second exception in a noexcept context. `std::terminate` aborts the broker instead of the request returning an error. The fix stops the destructor from throwing (it becomes noexcept and log-only) and moves enforcement into the route wrapper, which now shares the auth result with the handler and inspects it once the handler completes: a failed handler keeps its original error, and a reply produced without the check is discarded and replaced with a 500. The admin server is unaffected because it authorizes eagerly at dispatch. Covered by unit tests and two ducktape tests. One of the ducktape tests exercises the described failure (an unreachable `_schemas` partition makes the handler throw before the auth result can be checked) and asserts the request fails with an error response while the broker keeps serving; on an unfixed build it fails with the broker abort described above. Fixes [CORE-17034](https://redpandadata.atlassian.net/browse/CORE-17034) ## Backports Required - [x] v26.2.x - [x] v26.1.x - [x] v25.3.x ## Release Notes ### Bug Fixes * Fixed Schema Registry aborting the broker when a request failed before its deferred authorization check with `schema_registry_enable_authorization` enabled. Such requests now return an error response. [CORE-17034]: https://redpandadata.atlassian.net/browse/CORE-17034?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31559",
          "createdAt": "2026-08-12T23:01:13Z",
          "updatedAt": "2026-08-13T14:09:57Z",
          "timestamp": "2026-08-13T14:09:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "area/build",
            "area/redpanda"
          ],
          "author": "nguyen-andrew",
          "state": "open",
          "assignees": [
            "nguyen-andrew"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2a560f580bc557bc037e",
        "signalId": "github:redpanda-data/redpanda:pull_request:31544",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "labels"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31544",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[CORE-12930] - Storage: Some observability improvements for unrecoverable segments",
          "text": "<!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> A .cannotrecover file is a segment that failed startup recovery. The presence of these is a consistent source of confusion for support engineers. This PR adds some new metrics, naming, and logging around unrecoverable segments. - gauge for unrecoverable segments on disk, labelled by whether the segment was the tail of its log at recovery time (tail, mid_log, or unknown for files an older version left behind) - counted during the open_segments directory walk during recovery - try to surface a reason we couldn't recover the segment (not enough bytes, zero header, crc error, etc.) - put the reason and the position in the filename: `<segment>.<stop_reason>.<tail|mid_log>.cannotrecover`, so both are still there after the log line that reported them rotates away - in addition to the filename, log the size on disk, `log_replayer` checkpoint (incl the stop reason), and where the segment sat in the log - for non-tail segments, log at warn, otherwise keep logging at info as before. The idea here is to give a bit more signal to support or oncallers trying to determine whether cannotrecover files might indicate data loss, where to look, etc. Also adds some code to clean up previously orphaned index files from unrecoverable segments that could otherwise prevent partition removal from fully removing partition directories. Covers .base_index and .compaction_index, plus the compaction staging variant of each. ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [x] v26.2.x - [x] v26.1.x - [ ] v25.3.x ## Release Notes ### Improvements * Improves observability and logging around unrecoverable segments. <!-- If the changes in this PR do not need to be mentioned in the release notes, then don't add a sub-section and simply list `none`, e.g. * none Otherwise, adding a sub-section or `none` is REQUIRED if the PR is not a backport PR. If this is a backport PR, adding contents to this section will override the release notes section inherited from the original PR to dev. Add one or more of the sub-sections with a short description bullet point of the change, e.g. ### Bug Fixes * Short description of the bug fix if this is a PR to `dev` branch. ### Features * Short description of the feature. Explain how to configure. ### Improvements * Short description of how this PR improves existing behavior. -->",
          "url": "https://github.com/redpanda-data/redpanda/pull/31544",
          "createdAt": "2026-08-12T00:24:39Z",
          "updatedAt": "2026-08-13T17:29:34Z",
          "timestamp": "2026-08-13T17:29:34Z",
          "metrics": {
            "reactions": 0,
            "comments": 2
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "claude-review"
          ],
          "author": "oleiman",
          "state": "open",
          "assignees": [
            "oleiman"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d50441a65def85ca13a9",
        "signalId": "github:redpanda-data/redpanda:pull_request:31560",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31560",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "bazel: bump the toolchain sysroot to Ubuntu 24.04",
          "text": "The sysroot's `linux-libc-dev` decides which kernel uapi constants the build can see, no matter which kernel the broker actually runs on. Ours is built from Ubuntu 22.04, which pins that at 5.15 and costs us every addition since October 2021. Two of those we already reach for and do not find, and in both cases someone worked around it locally rather than the gap being caught by a build failure: - **`MADV_COLLAPSE` (6.1) is a live functional loss.** Seastar's memory prefaulter asks the kernel to synchronously collapse the freshly-populated seastar heap into transparent huge pages, guarded on `#ifdef MADV_COLLAPSE` in `src/core/smp.cc`. The constant is absent from the 22.04 sysroot, so that call is not compiled into anything we ship and the heap only gets whatever khugepaged does in the background. It is visible in the object code: `smp.pic.o` carries one `madvise` immediate (`$0x17`, `MADV_POPULATE_WRITE`) in every 22.04-built tree I have, and gains `$0x19` (25, `MADV_COLLAPSE`) with this change. Note `syschecks/hugepages.cc` already carries a local `#define MADV_COLLAPSE 25` for `promote_code_to_hugepages()`, so the gap was hit and patched in one translation unit while the seastar call site stayed dark. - **`STATX_DIOALIGN` (6.1) is no longer a functional loss, but its safety net is cut.** Seastar queries DIO alignment through a hand-rolled 256-byte `struct statx` stand-in with a local `0x2000`, so this works today (CORE-16295 / #30443). What does not work is the check on it: both `static_assert`s validating the stand-in's layout against the real struct sit inside the same `#ifdef`, so the hand-written padding has never been verified. Against real 6.8 headers they compile and pass, so there is no latent bug hiding behind them today, but nothing was keeping it that way. 24.04 carries 6.8 headers, which covers both. Two things worth stating explicitly because they look like risks and are not. **The glibc move is not one:** `//bazel/packaging` ships the sysroot's shared libraries and repoints the binaries' interpreter at the bundled loader, so the host's glibc is never consulted, and the loader's own minimum kernel is 3.2.0 on both 2.35 and 2.39. **io_uring is not gated by the sysroot** either, despite being the largest relevant slice of the header diff (74 new identifiers): seastar's backend goes through `@liburing`, which vendors its own `io_uring.h` already carrying `IORING_SETUP_SINGLE_ISSUER`, `DEFER_TASKRUN`, `REGISTER_RING_FDS` and `OP_SEND_ZC`. Bumping liburing is that lever, not this. 26.04 was measured too (glibc 2.43, 7.0 headers, same 3.2.0 floor) and deliberately skipped for now: it covers strictly more, but 24.04 clears both findings and is the smaller step for an input that invalidates every action in the build. The commits are (1) `Dockerfile.sysroot` onto `ubuntu:noble` with gcc's runtime dir moving 12 -> 13, plus README notes on the two floors a distro choice sets, and (2) the `repositories.bzl` / `MODULE.bazel.lock` pin. Verification, beyond the build passing: the tarball's file set outside `usr/include` is identical to the 22.04 sysroot on x86_64, and on aarch64 gains exactly gcc-13's `crtoffload{begin,end,table}.o` via the existing `*crt*.o` glob (additive, never linked unless offloading is used); no absolute symlinks are left for bazel's directory artifacts to reject; and `bazel build //src/v/base` builds all of seastar plus openssl and hwloc against it from scratch. **This is a draft because the tarballs are not published yet.** The third commit is marked DO NOT MERGE and carries them in-tree, handed to bazel through a distdir so the not-yet-existing `_SYSROOT_URL` is never fetched; it is here so CI can exercise the branch before anything is published. Reverting that one commit is the entire cleanup, because the pinned sha256s are already the final ones - a tarball's hash does not depend on where it is hosted. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31560",
          "createdAt": "2026-08-12T23:02:50Z",
          "updatedAt": "2026-08-13T17:29:33Z",
          "timestamp": "2026-08-13T17:29:33Z",
          "metrics": {
            "reactions": 0,
            "comments": 11
          },
          "labels": [
            "area/build"
          ],
          "author": "travisdowns",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:595c669defaa22e5785c",
        "signalId": "github:redpanda-data/redpanda:pull_request:31556",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31556",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[UX-1427] chore: bump franz-go to 1.21.6",
          "text": "Particularly interested in replacing golang.org/x/crypto/pbkdf2 with the stdlib one. <!-- See https://github.com/redpanda-data/redpanda/blob/dev/CONTRIBUTING.md#pull-request-body for more details and examples of what is expected in a PR body. Content in this top section is REQUIRED. Describe, in plain language, the motivation behind the change (bug fix, feature, improvement) in this PR and how the included commits address it. Add the GitHub keyword `Fixes` to link to bug(s) this PR will fix, e.g. Fixes #ISSUE-NUMBER, Fixes #ISSUE-NUMBER, ... If this PR is a backport, link to the original with `Backport of PR`, e.g. Backport of PR #PR-NUMBER --> ## Backports Required <!-- Checking at least one of the checkboxes is REQUIRED if this PR is not a backport. --> - [ ] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31556",
          "createdAt": "2026-08-12T18:44:39Z",
          "updatedAt": "2026-08-13T17:14:35Z",
          "timestamp": "2026-08-13T17:14:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/rpk"
          ],
          "author": "r-vasquez",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:624d131a5b6decdfd1f0",
        "signalId": "github:redpanda-data/redpanda:pull_request:31365",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31365",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rptest: add OOM crash self-test; allow-list memory diagnostics",
          "text": "Revives #8570 (rebased onto dev, see #8565 for the original motivation): put the seastar memory diagnostics dump, which is emitted at ERROR level when a node runs out of memory, on the default log allow list so an OOM surfaces as the \"redpanda crashed\" diagnostic rather than a BadLogLines about the dump. Also clarifies in the raise_on_bad_logs docstrings that allow-list regexes are unanchored. The main value over the old PR is the meta-test: a new debug-only `oom` type for the `/v1/debug/trigger_crash` admin API that allocates until the seastar allocator gives out, plus a ducktape self-test (alongside the existing segfault/assert ones) that runs a node out of memory and checks the harness reports it as a NodeCrash. With the system allocator (sanitized builds) there is no seastar memory limit and unbounded allocation would exhaust host memory instead, so there the API refuses the request, and the self-test verifies the refusal and that the node survives. ## Backports Required - [x] none - not a bug fix - [ ] none - this is a backport - [ ] none - issue does not exist in previous branches - [ ] none - papercut/not impactful enough to backport - [ ] v26.1.x - [ ] v25.3.x - [ ] v25.2.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31365",
          "createdAt": "2026-07-30T14:45:49Z",
          "updatedAt": "2026-08-13T17:12:15Z",
          "timestamp": "2026-08-13T17:12:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "travisdowns",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:d171fb8b9d07ac31530a",
        "signalId": "github:redpanda-data/redpanda:pull_request:31573",
        "event": "discovered",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk: add stretch cluster grafana dashboard behind --cluster-type flag",
          "text": "Adds a `--cluster-type` flag to `rpk generate grafana-dashboard`. The default value (`default`) leaves the existing `--dashboard` flow untouched; `stretch` prints a new embedded dashboard focused on stretch cluster observability: cross-cluster raft health, StretchCluster member status, and operator reconcile health. ``` rpk generate grafana-dashboard --cluster-type stretch ``` The dashboard is converted from the Redpanda operator repository's [`docs/operator-grafana-dashboard.json`](https://github.com/redpanda-data/redpanda-operator/blob/main/docs/operator-grafana-dashboard.json), with the Prometheus datasource parameterized as a `${DS_PROMETHEUS}` dashboard variable so it imports cleanly into any Grafana instance (matching the convention of the other embedded dashboards, e.g. serverless). It is embedded-only — not downloaded from the [observability](https://github.com/redpanda-data/observability) GitHub repository — so combining `--dashboard` with `--cluster-type stretch` is rejected with a clear error, as is any unrecognized `--cluster-type` value. Unit tests cover flag parsing/routing (`TestClusterTypeStretch`) and the embedded gzip's integrity via sha256 (extended `Test_embeddedDecompressAndPrint`). ## Backports Required - [x] none - not a bug fix ## Release Notes ### Features * `rpk generate grafana-dashboard` accepts a new `--cluster-type` flag. Passing `--cluster-type stretch` generates a Grafana dashboard for stretch clusters managed by the Redpanda Operator, covering cross-cluster raft health, StretchCluster member status, and operator reconcile health. 🤖 Generated with [Claude Code](https://claude.com/claude-code) [K8S-855](https://redpandadata.atlassian.net/browse/K8S-855) [K8S-855]: https://redpandadata.atlassian.net/browse/K8S-855?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ",
          "url": "https://github.com/redpanda-data/redpanda/pull/31573",
          "createdAt": "2026-08-13T17:02:20Z",
          "updatedAt": "2026-08-13T17:10:58Z",
          "timestamp": "2026-08-13T17:10:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/rpk",
            "area/build"
          ],
          "author": "RafalKorepta",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:3997a8d83abd39851bb0",
        "signalId": "github:redpanda-data/redpanda:pull_request:31566",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31566",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "k/s/tests: deflake the offset_store producer lock test",
          "text": "`the_producer_lock_serializes_by_producer` asserted that two operations ran concurrently, which requires the first producer's operation to start before the second producer's runs. Both are scheduled as tasks in debug builds, and nothing ordered them: under `--@seastar//:shuffle_task_queue=true` the second can run to completion first (reproduced in 10/30 runs, 0/100 after the fix). Wait for the first operation to hold the lock instead. The mid-test checks also become non-fatal. `ASSERT_*_CORO` expands to `co_return`, which freed a coroutine frame the pending operations still referenced, so the assertion failure surfaced as a heap-use-after-free that masked it. Test-only change. ## Backports Required - [ ] none - papercut/not impactful enough to backport - [ ] none - not a bug fix - [ ] none - this is a backport - [x] none - issue does not exist in previous branches - [ ] v26.2.x - [ ] v26.1.x - [ ] v25.3.x ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31566",
          "createdAt": "2026-08-13T13:39:22Z",
          "updatedAt": "2026-08-13T17:02:37Z",
          "timestamp": "2026-08-13T17:02:37Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/redpanda"
          ],
          "author": "nvartolomei",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2cecc05030fcdb04b122",
        "signalId": "github:redpanda-data/redpanda:pull_request:31571",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31571",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] storage: bump segment_concatenation_test timeout",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
          "url": "https://github.com/redpanda-data/redpanda/pull/31571",
          "createdAt": "2026-08-13T15:17:19Z",
          "updatedAt": "2026-08-13T17:01:41Z",
          "timestamp": "2026-08-13T17:01:41Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:f985cae819c65124beed",
        "signalId": "github:redpanda-data/redpanda:pull_request:31569",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31569",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v25.3.x] storage: bump segment_concatenation_test timeout",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
          "url": "https://github.com/redpanda-data/redpanda/pull/31569",
          "createdAt": "2026-08-13T15:17:14Z",
          "updatedAt": "2026-08-13T16:52:45Z",
          "timestamp": "2026-08-13T16:52:45Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:99b1a3c62cf58d165cd0",
        "signalId": "github:redpanda-data/redpanda:pull_request:31540",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31540",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] bazel: define an empty ci-remote-cache config",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31537",
          "url": "https://github.com/redpanda-data/redpanda/pull/31540",
          "createdAt": "2026-08-11T17:40:18Z",
          "updatedAt": "2026-08-13T16:42:32Z",
          "timestamp": "2026-08-13T16:42:32Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:25de712cfcf86133808c",
        "signalId": "github:redpanda-data/redpanda:pull_request:31551",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31551",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.1.x] rpk/connect: don't cap version segments at two digits in VersionFromString",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31546 Fixes: https://github.com/redpanda-data/redpanda/issues/31550,",
          "url": "https://github.com/redpanda-data/redpanda/pull/31551",
          "createdAt": "2026-08-12T15:52:55Z",
          "updatedAt": "2026-08-13T16:42:31Z",
          "timestamp": "2026-08-13T16:42:31Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/rpk",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9773c951fc06b42c7935",
        "signalId": "github:redpanda-data/redpanda:pull_request:31521",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31521",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk: fix grammar and wording defects in help text",
          "text": "## Summary Copyedits surfaced by auditing the docs team's rpk description overrides (`docs-data/rpk-overrides.json` in redpanda-data/docs) against rpk's source text — these are cases where the docs override exists only to paper over a source defect, so fixing the source lets the docs drop the override: - `profile clear`: \"unset **an prod** cluster profile\" → \"a production\" - `registry mode get`/`set`, `registry compatibility-level`: fix the \"global mode**, alternatively** you can…\" comma splice (same pattern in all three) - `security acl delete --cluster`: ACLs are removed *for*, not *to*, the cluster - `security acl list --registry-global`/`--registry-subject`: the list command **matches** ACLs, it doesn't **grant** them (text was copied from `acl create`, where \"grant\" is correct and unchanged) - `container start --retries`, `topic consume --fetch-max-bytes`/`--fetch-max-partition-bytes`: \"amount of\" → \"number of\" for countable nouns Help-text strings only; no behavior changes. Affected package tests pass. Related: #31520 (single-sources the `-X` flag docs from the same audit). 🤖 Generated with [Claude Code](https://claude.com/claude-code) ## Release Notes * none",
          "url": "https://github.com/redpanda-data/redpanda/pull/31521",
          "createdAt": "2026-08-11T09:30:20Z",
          "updatedAt": "2026-08-13T16:41:50Z",
          "timestamp": "2026-08-13T16:41:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [
            "area/rpk"
          ],
          "author": "JakeSCahill",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0a49e9005ab8703359a4",
        "signalId": "github:redpanda-data/redpanda:pull_request:31570",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31570",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] storage: bump segment_concatenation_test timeout",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31568",
          "url": "https://github.com/redpanda-data/redpanda/pull/31570",
          "createdAt": "2026-08-13T15:17:15Z",
          "updatedAt": "2026-08-13T17:53:51Z",
          "timestamp": "2026-08-13T17:53:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/build",
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8c6f972773c0817f0479",
        "signalId": "github:redpanda-data/redpanda:pull_request:31572",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31572",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "[v26.2.x] kafka/client/test: deflake data_queue blocking push test",
          "text": "Backport of PR https://github.com/redpanda-data/redpanda/pull/31567",
          "url": "https://github.com/redpanda-data/redpanda/pull/31572",
          "createdAt": "2026-08-13T16:10:51Z",
          "updatedAt": "2026-08-13T17:52:18Z",
          "timestamp": "2026-08-13T17:52:18Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [
            "area/redpanda",
            "kind/backport"
          ],
          "author": "vbotbuildovich",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:8a854d322ff0dcd679c7",
        "signalId": "github:redpanda-data/redpanda:pull_request:31573",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "text",
          "updatedAt"
        ],
        "signal": {
          "id": "github:redpanda-data/redpanda:pull_request:31573",
          "source": "github",
          "group": "data-infrastructure",
          "project": "redpanda-data/redpanda",
          "kind": "pull_request",
          "title": "rpk: add stretch cluster grafana dashboard behind --cluster-type flag",
          "text": "Adds a `--cluster-type` flag to `rpk generate grafana-dashboard` (K8S-855, https://redpandadata.atlassian.net/browse/K8S-855). The default value (`default`) leaves the existing `--dashboard` flow untouched; `stretch` swaps in dashboard variants for stretch clusters. The stretch variant of the `operations` dashboard is a new dashboard focused on stretch cluster observability: cross-cluster raft health, StretchCluster member status, and operator reconcile health. ``` rpk generate grafana-dashboard --cluster-type stretch ``` The cluster type composes with `--dashboard` by overlaying the default dashboard set: any dashboard without a stretch-specific variant (e.g. `consumer-offsets`) is generated exactly as before. Stretch variants follow the same flow as the other dashboards — downloaded from the [observability](https://github.com/redpanda-data/observability) GitHub repository, with an embedded gzipped fallback pinned by sha256. Until the companion observability-repo change merges, the download 404s and rpk serves the byte-identical embedded copy. The dashboard is converted from the Redpanda operator repository's [`docs/operator-grafana-dashboard.json`](https://github.com/redpanda-data/redpanda-operator/blob/main/docs/operator-grafana-dashboard.json), with the Prometheus datasource parameterized as a `${DS_PROMETHEUS}` dashboard variable so it imports cleanly into any Grafana instance (matching the convention of the other embedded dashboards, e.g. serverless). Unit tests cover flag parsing, the GitHub download path (stubbed host), the embedded fallback, the overlay fallback to the default set, and the embedded gzip's integrity via sha256. ## Backports Required - [x] none - not a bug fix ## Release Notes ### Features * `rpk generate grafana-dashboard` accepts a new `--cluster-type` flag. Passing `--cluster-type stretch` generates a Grafana dashboard for stretch clusters managed by the Redpanda Operator, covering cross-cluster raft health, StretchCluster member status, and operator reconcile health. It composes with `--dashboard`: dashboards without a stretch-specific variant are generated as usual. 🤖 Generated with [Claude Code](https://claude.com/claude-code)",
          "url": "https://github.com/redpanda-data/redpanda/pull/31573",
          "createdAt": "2026-08-13T17:02:20Z",
          "updatedAt": "2026-08-13T17:51:51Z",
          "timestamp": "2026-08-13T17:51:51Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "area/rpk",
            "area/build"
          ],
          "author": "RafalKorepta",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      }
    ]
  }
}
