{
  "schemaVersion": 3,
  "dataset": {
    "version": 3,
    "date": "2026-08-13",
    "group": {
      "id": "data-infrastructure",
      "name": "Data / Messaging / Storage Infrastructure"
    },
    "repository": {
      "id": "dragonfly",
      "repo": "dragonflydb/dragonfly",
      "name": "Dragonfly",
      "keywords": [
        "DragonflyDB"
      ]
    },
    "context": {
      "repository": "dragonflydb/dragonfly",
      "url": "https://github.com/dragonflydb/dragonfly",
      "description": "A modern replacement for Redis and Memcached",
      "homepage": "https://www.dragonflydb.io/",
      "language": "C++",
      "topics": [
        "cache",
        "cpp",
        "database",
        "fibers",
        "in-memory",
        "in-memory-database",
        "key-value",
        "keydb",
        "memcached",
        "message-broker",
        "multi-threading",
        "nosql",
        "redis",
        "valkey",
        "vector-search"
      ],
      "license": "NOASSERTION",
      "defaultBranch": "main",
      "stars": 30992,
      "forks": 1226,
      "openIssues": 300,
      "archived": false,
      "collectedAt": "2026-08-13T18:02:04.758446+00:00"
    },
    "news": {
      "repository": "dragonflydb/dragonfly",
      "collectedAt": "2026-08-13T18:02:04.758446+00:00",
      "latestRelease": {
        "repository": "dragonflydb/dragonfly",
        "tag": "v1.40.1",
        "title": "v1.40.1",
        "url": "https://github.com/dragonflydb/dragonfly/releases/tag/v1.40.1",
        "publishedAt": "2026-08-06T06:54:05Z",
        "notes": "## This is a patch release.\r\n\r\n### What's Changed\r\n\r\n- Fixed connection-state handling in squashed pipelines. Commands following `AUTH`, `SELECT`, `HELLO`, `CLIENT`, or `RESET` now observe the updated connection state instead of stale authentication, database, or RESP protocol state ([#8016](https://github.com/dragonflydb/dragonfly/pull/8016)).\r\n- Fixed compressed QList node memory accounting and defragmentation ([#8011](https://github.com/dragonflydb/dragonfly/pull/8011), [#8014](https://github.com/dragonflydb/dragonfly/pull/8014)).\r\n\r\nFull Changelog: [v1.40.0...v1.40.1](https://github.com/dragonflydb/dragonfly/compare/v1.40.0...v1.40.1)",
        "highlights": [
          "This is a patch release.",
          "What's Changed",
          "Fixed connection-state handling in squashed pipelines. Commands following AUTH, SELECT, HELLO, CLIENT, or RESET now observe the updated connection state instead of stale authentication, database, or RESP protocol state (#8016).",
          "Fixed compressed QList node memory accounting and defragmentation (#8011, #8014)."
        ],
        "prerelease": false
      },
      "upcoming": [
        {
          "repository": "dragonflydb/dragonfly",
          "kind": "milestone",
          "title": "Cluster Search",
          "url": "https://github.com/dragonflydb/dragonfly/milestone/19",
          "description": "",
          "dueAt": "2025-12-31T00:00:00Z",
          "progress": 60,
          "openIssues": 2,
          "closedIssues": 3
        },
        {
          "repository": "dragonflydb/dragonfly",
          "kind": "milestone",
          "title": "v1.41",
          "url": "https://github.com/dragonflydb/dragonfly/milestone/24",
          "description": "",
          "dueAt": "2026-09-10T00:00:00Z",
          "progress": 12,
          "openIssues": 7,
          "closedIssues": 1
        }
      ],
      "communityDiscussions": []
    },
    "runs": [
      {
        "collectedAt": "2026-08-13T12:26:38.318Z",
        "since": "2026-08-12T12:26:38.318Z",
        "observedCount": 28,
        "changedCount": 28
      },
      {
        "collectedAt": "2026-08-13T13:48:00.446149Z",
        "since": "2026-08-12T13:48:00.446149Z",
        "observedCount": 27,
        "changedCount": 27
      },
      {
        "collectedAt": "2026-08-13T16:19:22.035158Z",
        "since": "2026-08-12T16:19:22.035158Z",
        "observedCount": 29,
        "changedCount": 8
      },
      {
        "collectedAt": "2026-08-13T17:43:20.785491Z",
        "since": "2026-08-12T17:43:20.785491Z",
        "observedCount": 29,
        "changedCount": 1
      },
      {
        "collectedAt": "2026-08-13T17:47:07.884300Z",
        "since": "2026-08-12T17:47:07.884300Z",
        "observedCount": 29,
        "changedCount": 0
      },
      {
        "collectedAt": "2026-08-13T18:01:55.420671Z",
        "since": "2026-08-12T18:01:55.420671Z",
        "observedCount": 29,
        "changedCount": 1
      }
    ],
    "signals": [
      {
        "id": "github:dragonflydb/dragonfly:issue:5662",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "test_replicaof_reject_on_load",
        "text": "\\https://github.com/dragonflydb/dragonfly/actions/runs/16884994243/job/47830310679#step:6:1077",
        "url": "https://github.com/dragonflydb/dragonfly/issues/5662",
        "createdAt": "2025-08-12T07:04:07Z",
        "updatedAt": "2026-08-13T07:05:50Z",
        "timestamp": "2026-08-13T07:05:50Z",
        "metrics": {
          "reactions": 0,
          "comments": 37
        },
        "labels": [
          "bug",
          "failing-test",
          "epoll",
          "iouring"
        ],
        "author": "kostasrim",
        "state": "open",
        "assignees": [
          "abhijat"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:7903",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "Blocked XREADGROUP is not woken when the watched stream is deleted or retyped",
        "text": "A client blocked with `XREADGROUP ... BLOCK 0 STREAMS <key> >` remains blocked when the watched stream becomes invalid. Operations such as `DEL`, expiration, `FLUSHDB`/`FLUSHALL`, or replacing the stream with another data type do not trigger reevaluation of the blocked command. With `BLOCK 0`, the connection can remain blocked indefinitely. Reproduction: Create a stream and consumer group: ```text XADD mystream 1-0 field value XGROUP CREATE mystream mygroup $ ``` In client A: ```text XREADGROUP GROUP mygroup consumer BLOCK 0 STREAMS mystream > ``` Wait until the client is blocked. In client B: ```text DEL mystream ``` Actual behavior: Client A remains blocked indefinitely. `INFO CLIENTS` continues to report it as a blocked client. Expected behavior: Client A should be woken and the command should be reevaluated. Since the stream and consumer group no longer exist, it should return a `NOGROUP` error. Equivalent cases: - `DEL`, `UNLINK`, expiration, `FLUSHDB`, and `FLUSHALL` should return `NOGROUP`. - Replacing the stream with a non-stream value should return `WRONGTYPE`. Plain `XREAD` should retain its existing semantics and remain blocked when a key disappears. The wake-on invalidation behavior is specifically required for `XREADGROUP`. Impact: - `BLOCK 0` connections can hang permanently. - Applications cannot recover when streams are deleted, expired, flushed, or recreated. - Blocked clients and their associated transaction resources remain retained. - The behavior differs from Valkey/Redis blocking-command semantics. Technical notes: The readiness checker can detect a missing key, missing consumer group, or wrong value type. However, deletion, expiration, flush, and retype paths do not notify the blocking controller, so the checker is never invoked. The underlying issue is that a blocked command is not reevaluated after invalidation of its watched key.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/7903",
        "createdAt": "2026-07-21T16:05:18Z",
        "updatedAt": "2026-08-13T10:39:52Z",
        "timestamp": "2026-08-13T10:39:52Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "vyavdoshenko",
        "state": "closed",
        "assignees": [
          "BorysTheDev"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:7953",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "FT.CREATE without SCHEMA returns OK but creates unusable listed index",
        "text": "## Summary `FT.CREATE` accepts an index definition that omits the required `SCHEMA` keyword. It returns `OK` and exposes the index in `FT._LIST`, but the resulting index is unusable: `FT.INFO` reports that the index does not exist. The command should be rejected during parsing/validation rather than creating a listed-but-non-operational index. ## Reproduction Started a local single-node Dragonfly instance: ```sh mkdir ~/local-data-dragonfly podman --connection dragon run -d --pull=always -p 6379:6379 \\ -v ~/local-data-dragonfly/:/data:Z --ulimit memlock=-1 \\ docker.dragonflydb.io/dragonflydb/dragonfly \\ --lua_allow_undeclared_auto_correct=true --nodf_snapshot_format & ``` Then: ```redis FT.CREATE idx_ts_test ON HASH PREFIX 1 ot:foo name TEXT timestamp NUMERIC SORTABLE # OK FT._LIST # includes \"idx_ts_test\" FT.INFO idx_ts_test # ERR Index with name 'idx_ts_test' not found ``` This is on a local single-node instance, so cluster request routing is not involved. ## Expected behavior The malformed `FT.CREATE` invocation should fail immediately because `SCHEMA` is missing, e.g. with a syntax error. A valid invocation is: ```redis FT.CREATE idx_ts_test ON HASH PREFIX 1 ot:foo SCHEMA name TEXT timestamp NUMERIC SORTABLE ```",
        "url": "https://github.com/dragonflydb/dragonfly/issues/7953",
        "createdAt": "2026-07-28T18:14:00Z",
        "updatedAt": "2026-08-13T15:59:09Z",
        "timestamp": "2026-08-13T15:59:09Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [],
        "author": "dragonclaw-dragonflydb",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8028",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "Disabled python tests in tests/dragonfly",
        "text": "## Unconditional skips (broken / flaky / WIP) - [ ] `tests/dragonfly/shutdown_test.py:15` — class `TestDflyAutoLoadSnapshot` (covers `test_gracefull_shutdown`) — \"Currently we can not guarantee that on shutdown if command is executed and value is written we response before breaking the connection\" @vyavdoshenko -> check if the invariants/what's written is true and if we can change somehow the test semantics to assert that we get a reply before we shutdown --- - [ ] `tests/dragonfly/memory_test.py:327` — `test_throttle_on_commands_squashing_replies_bytes` — \"Disabling test until improvements in squashing.\" @mkaruza --- - [ ] `tests/dragonfly/snapshot_test.py:222` — `test_cron_snapshot_failed_saving` — \"Fails and also causes all TLS tests to fail\" @BorysTheDev --- - [ ] `tests/dragonfly/snapshot_test.py:961` — `test_tiered_entries_throttle` — no reason given (also `@large`/`@opt_only`) @dranikpg --- - [ ] `tests/dragonfly/acl_family_test.py:271` — `test_acl_del_user_while_running_lua_script` — \"Flaky on CI, needs investigation\" @kostasrim --- - [ ] `tests/dragonfly/acl_family_test.py:297` — `test_acl_with_long_running_script` — \"Check TODO in the body below\" @kostasrim --- - [ ] `tests/dragonfly/connection_test.py:2483` — `test_pipeline_cache_size` — \"Flaky\" @abhijat --- - [ ] `tests/dragonfly/replication_specific_test.py:508` — `test_big_huge_streaming_restart` — \"fails, investigating\" @kostasrim --- - [x] `tests/dragonfly/replication_specific_test.py:730` — `test_replication_onmove_flow` — \"Fails constantly on CI\" Dead test https://github.com/dragonflydb/dragonfly/pull/7096 and https://github.com/dragonflydb/dragonfly/pull/7068/changes @kostasrim will remove it in a sec",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8028",
        "createdAt": "2026-08-07T09:05:39Z",
        "updatedAt": "2026-08-13T15:51:22Z",
        "timestamp": "2026-08-13T15:51:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug",
          "failing-test"
        ],
        "author": "kostasrim",
        "state": "open",
        "assignees": [
          "abhijat",
          "mkaruza",
          "kostasrim",
          "dranikpg",
          "vyavdoshenko"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8056",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "XREAD BLOCK replies in RESP2 array shape on a RESP3 connection",
        "text": "**Describe the bug** On a RESP3 connection, a blocking `XREAD` that is woken by a new entry replies in the RESP2 array shape instead of the RESP3 map shape. Real Redis replies with a map on the same path. This breaks clients that trust the negotiated protocol. `node-redis` selects its reply parser from the `HELLO` version, so it applies the RESP3 transform to the RESP2 array and throws `TypeError: Cannot read properties of undefined (reading 'map')`. The stream is then never consumed. The non-blocking path is correct. `XREAD` that returns immediately with data does reply with a map on RESP3. Only the blocking-wakeup path differs, which is what makes it easy to miss. `XREADGROUP` was fixed to reply as a map for RESP3 in v1.22.0 (#3639). Plain `XREAD` appears never to have been covered. **To Reproduce** ```js // npm i redis@6 import { createClient } from 'redis'; const url = 'redis://127.0.0.1:6379'; const KEY = 'repro:stream'; const writer = createClient({ url, RESP: 2 }); const reader = createClient({ url, RESP: 3 }); await writer.connect(); await reader.connect(); await writer.del(KEY); // Block from \"now\", then publish so the block wakes with an entry. const pending = reader.sendCommand(['XREAD', 'BLOCK', '3000', 'STREAMS', KEY, '$']); setTimeout(() => writer.xAdd(KEY, '*', { event: 'wake' }), 300); const reply = await pending; console.log(Array.isArray(reply) ? 'Array (RESP2 shape)' : 'Map/Object (RESP3 shape)'); ``` **Expected behavior** `Map/Object (RESP3 shape)`, matching Redis. **Actual behavior** `Array (RESP2 shape)`. Using the typed `reader.xRead({ key: KEY, id: '$' }, { BLOCK: 3000 })` instead of `sendCommand` throws `TypeError: Cannot read properties of undefined (reading 'map')`. **Results across servers** Same script, same client version, only the server changed: | Server | Blocking-wakeup reply on RESP3 | |---|---| | Dragonfly v1.38.1 | Array (RESP2 shape) | | Dragonfly v1.40.1 | Array (RESP2 shape) | | Redis 7.4 | Map (RESP3 shape) | | Redis 8 | Map (RESP3 shape) | RESP2 behaves identically on all four, so pinning the client to RESP2 is a viable workaround. **Environment** - Dragonfly: `df-v1.38.1` and `df-v1.40.1`, official images, default config, standalone - Client: `node-redis` 6.2.0 - Reproduced on both a local container and a deployed instance **Additional context** This surfaced in production as a service that publishes to a stream from several pods and reads it back with one blocking `XREAD` per pod. Reads failed continuously while writes succeeded, so the stream grew and nothing consumed it. The client's own error was the `TypeError` above, which points at the client rather than the server and made it slow to diagnose.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8056",
        "createdAt": "2026-08-12T02:40:13Z",
        "updatedAt": "2026-08-13T08:14:39Z",
        "timestamp": "2026-08-13T08:14:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [],
        "author": "BryanChow0112",
        "state": "closed",
        "assignees": [
          "mkaruza"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8063",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "test_replicate_old_master",
        "text": "Automated regression triage opened this issue. - Test: `dragonfly/replication_config_test.py::test_replicate_old_master[df_factory0--localhost-7000]` - Backend: iouring - Commit: f4019d7fec0ddcd1e6484dd6eeade7d52b146af6 - Run: https://github.com/dragonflydb/dragonfly/actions/runs/31641685740 - Marker: run/31641685740/attempt/1 ``` Reason: requests.exceptions.RetryError: HTTPSConnectionPool(host='github.com', port=443): Max retries exceeded with url: /dragonflydb/dragonfly/releases/download/v1.19.2/dragonfly-x86_64.tar.gz (Caused by ResponseError('too many 503 error responses')) ____________ test_replicate_old_master[df_factory0--localhost-7000] ____________ urllib3.exceptions.ResponseError: too many 503 error responses The above exception was the direct cause of the following exception: self = <requests.adapters.HTTPAdapter object at 0x7fdbe0bd0b30> request = <PreparedRequest [GET]>, stream = True, timeout = None, verify = True cert = None, proxies = OrderedDict() def send( self, request: PreparedRequest, stream: bool = False, timeout: _t.TimeoutType = None, verify: _t.VerifyType = True, cert: _t.CertType = None, proxies: dict[str, str] | None = None, ) -> Response: \"\"\"Sends PreparedRequest object. Returns Response object. :param request: The :class:`PreparedRequest <PreparedRequest>` being sent. :param stream: (optional) Whether to stream the request content. :param timeout: (optional) How long to wait for the server to send data before giving up, as a float, or a :ref:`(connect timeout, read timeout) <timeouts>` tuple. :type timeout: float or tuple or urllib3 Timeout object :param verify: (optional) Either a boolean, in which case it controls whether we verify the server's TLS certificate, or a string, in which case it must be a path to a CA bundle to use :param cert: (optional) Any user-provided SSL certificate to be trusted. :param proxies: (optional) The proxies dictionary to apply to the request. :rtype: requests.Response \"\"\" assert _is_prepared(request) try: conn = self.get_connection_with_tls_context( request, verify, proxies=proxies, cert=cert ) except LocationValueError as e: raise InvalidURL(e, request=request) self.cert_verify(conn, request.url, verify, cert) url = self.request_url(request, proxies) self.add_headers( request, stream=stream, timeout=timeout, verify=verify, cert=cert, proxies=proxies, ) chunked = not (request.body is None or \"Content-Length\" in request.headers) if isinstance(timeout, tuple): try: connect, read = timeout resolved_timeout = TimeoutSauce(connect=connect, read=read) except ValueError: raise ValueError( f\"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, \" f\"or a single float to set both timeouts to the same value.\" ) elif isinstance(timeout, TimeoutSauce): resolved_timeout = timeout else: resolved_timeout = TimeoutSauce(connect=timeout, read=timeout) try: > resp = conn.urlopen( method=request.method, url=url, body=request.body, # type: ignore[arg-type] # urllib3 stubs don't accept Iterable[bytes | str] headers=request.headers, # type: ignore[arg-type] # urllib3#3072 redirect=False, assert_same_host=False, preload_content=False, decode_content=False, retries=self.max_retries, timeout=resolved_timeout, chunked=chunked, ) cert = None chunked = False conn = <urllib3.connectionpool.HTTPSConnectionPool object at 0x7fdbe8249880> proxies = OrderedDict() request = <PreparedRequest [GET]> resolved_timeout = Timeout(connect=None, read=None, total=None) self = <requests.adapters.HTTPAdapter object at 0x7fdbe0bd0b30> stream = True timeout = None url = '/dragonflydb/dragonfly/releases/download/v1.19.2/dragonfly-x86_64.tar.gz' verify = True /usr/local/lib/python3.12/dist-packages/requests/adapters.py:696: _ _ _ _ _ _ _ _ _ _ ```",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8063",
        "createdAt": "2026-08-12T21:51:47Z",
        "updatedAt": "2026-08-12T21:51:59Z",
        "timestamp": "2026-08-12T21:51:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "failing-test",
          "iouring"
        ],
        "author": "dragonclaw-dragonflydb",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8067",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P1 — A heterogeneous blocking queue can hide XREADGROUP forever",
        "text": "#### 1. P1 — A heterogeneous blocking queue can hide `XREADGROUP` forever `BlockingController` keeps one FIFO queue per `(db, key)` for all blocking commands. `NotifyWatchQueue()` checks only the front waiter; when its checker returns `kKeyNotFound` (BLPOP / BZPOP\\* / BLMPOP on a missing key), the whole scan aborts (`blocking_controller.cc:259-265`) and the one-shot awakened event is cleared (`blocking_controller.cc:175`), so later waiters in the same queue are never checked. The TODO at `stream_family.cc:3037-3039` documents the limitation. The trigger is irrelevant: DEL, UNLINK, expiry, FLUSHDB/FLUSHALL, and FLUSHSLOTS all hit the same break. Scenario: ```text # client A: first waiter while k does not exist BLPOP k 0 # client B XGROUP CREATE k g 0 MKSTREAM # client C: later waiter in the same (db, key) queue XREADGROUP GROUP g c BLOCK 0 STREAMS k > # client B DEL k ``` Actual: client C receives nothing and stays blocked indefinitely; it is released only by a later unrelated mutation of `k` (e.g. `LPUSH k v`, which also serves client A). Expected: client A remains blocked, client C immediately receives `NOGROUP`. Same result with `BZPOPMIN` in place of `BLPOP`.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8067",
        "createdAt": "2026-08-13T10:00:27Z",
        "updatedAt": "2026-08-13T10:00:58Z",
        "timestamp": "2026-08-13T10:00:58Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8068",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P1 — `NotifyPending()` is reentrant through a suspending expiry checker",
        "text": "`NotifyPending()` iterates the live `awakened_indices_` / `awakened_keys` sets with no reentrancy guard and no snapshot (`blocking_controller.cc:150-182`). The stream readiness checker calls `FindReadOnly()` (`stream_family.cc:3032`), which can enter `ExpireIfNeeded()` (`db_slice.cc:1451-1493`); that path can suspend **before** deleting the entry — in `RecordExpiryBlocking()` when a replication journal is active (`db_slice.cc:1471-1473`), or in keyspace-notification delivery under publish-buffer backpressure (`db_slice.cc:1476-1484`, `channel_store.cc:120`). While the checker fiber is suspended, any concluding transaction on the shard calls `NotifyPending()` again (`transaction.cc:668`). The inner call clears and erases the containers the outer frames still reference — including the `DbWatchTable` and `WatchQueue` held by reference — so the outer call resumes on invalidated iterators and a stale `PrimeIterator`. Consequences: lost wake, CHECK/DCHECK failure, use-after-free, duplicate erase. Reachability: requires an active replication journal with a stalled streamer, or `notify-keyspace-events Ex` with a backpressured subscriber. A default standalone configuration has no suspension point on this path. The synchronous same-fiber lazy-expiry case is benign and already covered by `XReadGroupBlockLazyExpireDuringWakeDoesNotCrash`; the uncovered defect is the cross-fiber suspension window. Scenario (race-sensitive): ```text CONFIG SET notify-keyspace-events Ex PSUBSCRIBE __keyevent@0__:expired # subscriber stops consuming replies # fill the subscriber's output budget so an expired event must wait XGROUP CREATE s g 0 MKSTREAM XREADGROUP GROUP g c BLOCK 0 STREAMS s > # client A PEXPIRE s 10 # client B # after the TTL elapses, run another write on the same shard SET unrelated-key value ``` A deterministic test needs a latch inside `RecordExpiryBlocking()` or the notification send; a timing-only test would be flaky.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8068",
        "createdAt": "2026-08-13T10:01:22Z",
        "updatedAt": "2026-08-13T10:01:22Z",
        "timestamp": "2026-08-13T10:01:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8069",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P1 — Active expiry never runs for non-default namespaces",
        "text": "Heartbeat expiry and eviction scan only the default namespace (`engine_shard.cc:784-785`, `862-864`), and the post-expiry wake dispatch uses only the default namespace's `BlockingController` (`engine_shard.cc:959-964`). Keys in ACL namespaces are never actively expired at all — no deletion, no keyspace notification, no wake. Only a later tenant command that touches the key expires it lazily. Scenario: ```text ACL SETUSER u NAMESPACE:tenant ON >p +@all ~* # tenant connection 1 XGROUP CREATE s g 0 MKSTREAM XREADGROUP GROUP g c BLOCK 0 STREAMS s > # tenant connection 2 PEXPIRE s 100 ``` Actual: the blocked tenant client is never woken; with `BLOCK 0` it stays blocked indefinitely. A later tenant command touching `s` (even `EXISTS s`) lazily expires the key and only then the reader receives `NOGROUP`. The identical scenario in the default namespace correctly delivers `NOGROUP` shortly after the TTL elapses.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8069",
        "createdAt": "2026-08-13T10:01:39Z",
        "updatedAt": "2026-08-13T10:01:39Z",
        "timestamp": "2026-08-13T10:01:39Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8070",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P1 — Multi-stream `XREADGROUP` mutates group state and then returns an error",
        "text": "`HasEntries2()` creates or updates the consumer while checking readiness (`stream_family.cc:4042-4055`), and the single-shard path performs `OpRead()` and journals group side effects (`stream_family.cc:3262-3272`) before the accumulated error is checked (`stream_family.cc:3285-3287`). Pre-existing bug, not introduced by the wake fixes. Scenario: ```text XGROUP CREATE good g 0 MKSTREAM XADD good 1-0 f v XREADGROUP GROUP g c STREAMS good missing > > ``` Actual: the reply is `NOGROUP`, yet entry `1-0` is inserted into the PEL, consumer `c` is created, and the group's last-delivered ID advances to `1-0` — the client never receives the entry it now owns. This happens whenever both keys land on the same shard (any thread count, not only `--proactor_threads=1`). Expected: validation of all streams before any consumer-group mutation. Cross-shard variant: `OpRead()` is skipped on the error path, but the consumer is still created on the valid stream without any journal write — a replication-divergence risk.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8070",
        "createdAt": "2026-08-13T10:01:54Z",
        "updatedAt": "2026-08-13T10:01:54Z",
        "timestamp": "2026-08-13T10:01:54Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8071",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P1 — A blocked multi-stream read returns only one stream when several become ready together",
        "text": "The wake path resolves only `Transaction::GetWakeKey()` (`stream_family.cc:3139-3142`) and builds a reply with cardinality exactly one (`stream_family.cc:3221-3229`). The root cause is in the transaction layer: `NotifySuspended` claims a once-only blocking barrier and records a single wake key (`transaction.cc:1578`, `1598`), so the woken command structurally cannot see that several keys became ready — a fix needs to re-scan all requested streams on wake, as Valkey's unblock reprocessing does. Pre-existing bug, not introduced by #8047. Scenario (single shard for determinism): ```text XGROUP CREATE s1 g 0 MKSTREAM XGROUP CREATE s2 g 0 MKSTREAM # client A XREADGROUP GROUP g c BLOCK 0 STREAMS s1 s2 > > # client B MULTI XADD s1 1-0 f one XADD s2 1-0 f two EXEC ``` Actual: the reply contains only one stream (e.g. `s2`), and only that stream gets PEL treatment; `XPENDING s1 g` stays empty. The omitted entries are not lost (last-delivered on `s1` is unchanged, a later `>` read retrieves them), but the reply is incomplete vs Valkey and the PEL is asymmetric. Plain blocked `XREAD` has the same single-stream reply, without the PEL aspect.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8071",
        "createdAt": "2026-08-13T10:02:13Z",
        "updatedAt": "2026-08-13T10:02:13Z",
        "timestamp": "2026-08-13T10:02:13Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8072",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P2 — Multi-shard error precedence follows shard placement instead of argument order",
        "text": "For a woken blocked read, per-shard validation statuses are scanned in numeric shard-ID order (`stream_family.cc:3119-3135`), so the reported error is decided by key-to-shard hashing. Within a single shard, argument order is respected — the defect is strictly cross-shard. The initial non-blocked path is worse: the aggregate error (`AggregateValue`, `execution_state.h:21-28`) is written by whichever shard callback runs first, so identical invocations on identical state return different errors run to run (observed as a 10/10 split over 20 calls). Valkey deterministically reports the error of the first invalid argument in command order. Scenario (keys `a` and `b` on different shards, `a` listed first): ```text XGROUP CREATE a g 0 MKSTREAM XGROUP CREATE b g 0 MKSTREAM # client A XREADGROUP GROUP g c BLOCK 0 STREAMS a b > > # client B: invalidate both in one transaction MULTI XGROUP DESTROY a g DEL b SET b v EXEC ``` Actual: the returned error follows the lower-shard key regardless of its argument position (`WRONGTYPE` here although `NOGROUP` for `a` is expected; swapping which key is destroyed/retyped flips it accordingly).",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8072",
        "createdAt": "2026-08-13T10:02:29Z",
        "updatedAt": "2026-08-13T10:02:29Z",
        "timestamp": "2026-08-13T10:02:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8073",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P2 — `DFLYCLUSTER FLUSHSLOTS` operates on the default namespace regardless of the caller",
        "text": "`cluster_family.cc:519` takes the DbSlice from `namespaces->GetDefaultNamespace()` instead of the calling transaction's namespace. A namespace-bound ACL user with `+@admin` / `+dflycluster` can invoke the command; it then flushes default-namespace data while the tenant's keys in those slots survive indefinitely (including past slot migration), and blocked tenant readers on them are never woken. This is systemic in the cluster subsystem — slot migration, `GETSLOTINFO`, and slot key counts also hard-code the default namespace (`outgoing_slot_migration.cc`, `cluster_utility.cc:28`) — so either cluster commands should be namespace-aware, or namespace-bound users should be barred from them explicitly. Scenario: ```text # put keys of the same slot into both the default namespace and NAMESPACE:tenant # authenticate as the tenant user and run DFLYCLUSTER FLUSHSLOTS <slot> <slot> ``` Actual: default-namespace data in the slot is deleted; tenant data survives.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8073",
        "createdAt": "2026-08-13T10:02:46Z",
        "updatedAt": "2026-08-13T10:59:53Z",
        "timestamp": "2026-08-13T10:59:53Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": [
          "mkaruza"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8074",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "P3 — FLUSHSLOTS TOCTOU between the validation and action hops of a woken multi-stream read",
        "text": "The woken blocked read runs a cross-shard validation hop and then an action hop that revalidates only the wake-key shard (`stream_family.cc:3117-3137`, `3156-3164`); `FlushSlotsFb()` deletes slot keys in a detached fiber without transaction locks and yields between chunks (`db_slice.cc:955-985`). A flush landing between the two hops can delete a sibling stream after it passed validation, so the client is served data from the ready stream although the other requested group no longer exists. Impact is benign in practice: under cluster mode both streams share a slot (CROSSSLOT), the reply legally serializes before the flush acknowledgment, and there is no crash or replication divergence — listed for completeness. A deterministic test would need a latch between the two hops.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8074",
        "createdAt": "2026-08-13T10:03:03Z",
        "updatedAt": "2026-08-13T11:00:28Z",
        "timestamp": "2026-08-13T11:00:28Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": [
          "mkaruza"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:issue:8076",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "issue",
        "title": "Return RESP3 null for empty XREAD and XREADGROUP replies",
        "text": "Under RESP3, XREAD and XREADGROUP return *-1 when no entries are available or a BLOCK timeout occurs. Redis returns the RESP3 null value (_\\r\\n) instead. Update these paths to emit RESP3 null replies while preserving *-1 for RESP2.",
        "url": "https://github.com/dragonflydb/dragonfly/issues/8076",
        "createdAt": "2026-08-13T10:41:53Z",
        "updatedAt": "2026-08-13T10:45:19Z",
        "timestamp": "2026-08-13T10:45:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 0
        },
        "labels": [
          "bug"
        ],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": [
          "mkaruza"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:6991",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "Use pcre2 regex & enable auto async",
        "text": "Make auto async use pcre2 for regexes. PCRE2 uses limited stack space, whereas std::regex is not suited for the small fiber stacks that we have",
        "url": "https://github.com/dragonflydb/dragonfly/pull/6991",
        "createdAt": "2026-03-26T17:18:00Z",
        "updatedAt": "2026-08-12T13:48:36Z",
        "timestamp": "2026-08-12T13:48:36Z",
        "metrics": {
          "reactions": 0,
          "comments": 3
        },
        "labels": [],
        "author": "dranikpg",
        "state": "closed",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:7984",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "feat(geo): add GEOSEARCHSTORE command",
        "text": "<!-- **Commits Must Be Signed and Your PR title must conform to the conventional commit spec** * See: https://github.com/dragonflydb/dragonfly/blob/main/CONTRIBUTING.md * Please follow the section on `pre-commit hooks`, a linter will validate before you push Example PR Title: <type>(<scope>)!: <description> * `type` = bug, chore, feat, fix, docs, build, style, refactor, perf, test * `!` = OPTIONAL: signals a breaking change * `scope` = Optional when `type` is \"chore\" or \"docs\" * `description` = short description of the change Examples: * chore(examples): Clarify `docker` usage #120 * docs(readme): Fix Example Links #121 * feat(ingest)!: Add new ingest #122 * fix(ingest): Refactor for loop to list comprehension #123 --> Adds GEOSEARCHSTORE (dest + src keys, same search options as GEOSEARCH, optional STOREDIST). Most of the work was already in GeoSearchStoreGeneric() via GEORADIUS ... STORE — this wires up the dedicated command and parser. Closes #3883. Also fixed a couple of Redis mismatches in the shared store/search path: - STORE when the source key is missing now returns 0 and clears the dest key (was an empty array) - COUNT without ASC/DESC now defaults to ASC, same as Redis Tests in geo_family_test.cc. Ran geo_family_test locally + manual redis-cli checks.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/7984",
        "createdAt": "2026-08-03T03:51:25Z",
        "updatedAt": "2026-08-13T16:05:14Z",
        "timestamp": "2026-08-13T16:05:14Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [],
        "author": "rounaknandanwar",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:7997",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "chore: add docs/replication.md",
        "text": "* add docs on replication state machine and the internal functionally at each phase",
        "url": "https://github.com/dragonflydb/dragonfly/pull/7997",
        "createdAt": "2026-08-04T10:07:57Z",
        "updatedAt": "2026-08-13T11:57:25Z",
        "timestamp": "2026-08-13T11:57:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [],
        "author": "kostasrim",
        "state": "open",
        "assignees": [
          "kostasrim"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8002",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "docs: add design doc for replication fanout (NOT FOR MERGE)",
        "text": "This is a design doc for #7993.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8002",
        "createdAt": "2026-08-04T12:45:27Z",
        "updatedAt": "2026-08-13T15:48:22Z",
        "timestamp": "2026-08-13T15:48:22Z",
        "metrics": {
          "reactions": 0,
          "comments": 1
        },
        "labels": [],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8026",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(tiering): preserve offloaded hashes during serialization",
        "text": "Preserve the logical type of offloaded values when serializing snapshots, full-sync streams, and DUMP payloads. Previously, offloaded listpack hashes were read as strings. Debug builds aborted on a DCHECK, while release builds serialized the raw listpack bytes as a STRING. Root cause: `SerializerBase` routed every external value through `ReadTieredString` and stored the delayed result as a string. This assumed that all top-level tiered values were strings, which stopped being true after introducing experimental hash offloading. `DumpToString`, used by DUMP and cross-shard RENAME/COPY, made the same assumption. It could also request a `StringDecoder` while snapshot serialization requested a `ListpackMapDecoder` for the same coalesced disk read.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8026",
        "createdAt": "2026-08-06T19:44:34Z",
        "updatedAt": "2026-08-12T13:48:59Z",
        "timestamp": "2026-08-12T13:48:59Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [],
        "author": "vyavdoshenko",
        "state": "closed",
        "assignees": [
          "vyavdoshenko"
        ],
        "change": "updated"
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8033",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "chore(server): Prefetch + cid caching for speedup",
        "text": "1. Prefetching bucket data It requires really large pipelines to reach large differences <img width=\"1650\" height=\"660\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5b675ef9-35f4-4a82-b157-941f9a6a22ef\" /> 2. Caching last command id Helps avoiding absl::AsciiStrToUpper + hashtable lookup + ifs for simple cases like all GET pipeline. Adds additionally around +1.5%",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8033",
        "createdAt": "2026-08-07T15:58:04Z",
        "updatedAt": "2026-08-13T11:06:29Z",
        "timestamp": "2026-08-13T11:06:29Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [],
        "author": "dranikpg",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8039",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "feat: add time limitation for replication backlog",
        "text": "fixes: #7994 Summary: This PR replaces the fixed per-shard replication backlog entry count with time- and byte-based retention. Changes: - Adds a five-second default age target via --shard_repl_backlog_time_ms. - Adds --shard_repl_backlog_max_bytes, defaulting to 0.5% of maxmemory. - Deprecates --shard_repl_backlog_len and updates stale-partial-sync guidance. - Expands C++ coverage for byte limits, capacity growth, oversized records, and time expiry. - Updates replication-resilience tests to configure byte-based backlog behavior. Technical Notes: The byte allowance is divided among shards, while time eviction is evaluated on append and grouped into one-second buckets.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8039",
        "createdAt": "2026-08-10T11:55:56Z",
        "updatedAt": "2026-08-12T15:44:47Z",
        "timestamp": "2026-08-12T15:44:47Z",
        "metrics": {
          "reactions": 0,
          "comments": 10
        },
        "labels": [],
        "author": "BorysTheDev",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8043",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(tls): atomic context switch and correct tls_bytes accounting",
        "text": "Two coupled TLS defects: - The `tls` CONFIG callback reconfigured listeners one by one; a bad cert/key failed midway and left listeners split between old and new TLS state (a failed callback only rolls back the flag value). - `tls_bytes` summed per-thread counters, but OpenSSL objects are allocated on one thread and freed on another (startup contexts are built on the main thread, which the metric sum never includes), so after a reload the summed metric wrapped to ~UINT64_MAX. Deterministic repro: no TLS connections alive + `CONFIG SET tls false` -> `tls_bytes` = 18446744073709531520. Changes: - The callback builds the SSL_CTX once before touching any listener; the new infallible `Listener::ApplyTlsCtx()` then switches each listener (taking its own reference), so a bad cert fails with zero listeners touched. `ApplyTlsCtx` also refreshes the listener's `tls_cert_info_`, keeping INFO's certificate metadata correct after a reload. - TLS memory accounting is a single process-wide relaxed atomic counter (cross-thread alloc/free pairs make per-thread counters structurally wrong); the metric is read once instead of summed per thread. The realloc hook no longer adjusts the counter on a failed reallocation, which would otherwise subtract the still-live block twice. - Regression test with main+admin listeners covering disable/enable reload.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8043",
        "timestamp": "2026-08-12T13:01:24Z",
        "metrics": {
          "reactions": 0,
          "comments": 5
        },
        "labels": [],
        "author": "vyavdoshenko",
        "assignees": [],
        "change": "new"
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8053",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(generic): avoid data loss on RENAME when destination write fails",
        "text": "## Summary - `Renamer::FinalizeRename()` ran `DelSrc` and `DeserializeDest` in the same transaction hop with no ordering guarantee between shards. If `DeserializeDest` failed (e.g. OOM), the source key could already be deleted, silently losing its data. - Now deserializes into the destination first, and only deletes the source once that succeeds. `COPY` is unaffected since it never deletes the source. ## Test plan - [x] Added `GenericFamilyTest.RenameOOMPreservesSource`: drives real per-shard OOM via repeated `RESTORE`, then asserts a failed `RENAME` leaves the source key intact and does not create the destination. - [x] Full `generic_family_test` suite passes (84/84). - [x] `pre-commit` (clang-format) passes on changed files. Fixes #5296",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8053",
        "createdAt": "2026-08-11T17:02:15Z",
        "updatedAt": "2026-08-13T13:59:40Z",
        "timestamp": "2026-08-13T13:59:40Z",
        "metrics": {
          "reactions": 0,
          "comments": 8
        },
        "labels": [],
        "author": "Shikha-code36",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8054",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(search): reject FT.CREATE missing the SCHEMA keyword",
        "text": "## Summary - `FT.CREATE` accepted an index definition missing the required `SCHEMA` keyword, returning `OK` and listing the index in `FT._LIST`, but the index was unusable (`FT.INFO` reported it as not found). - Root cause: without `SCHEMA`, tokens that look like field definitions fall through the \"unsupported parameters are ignored\" branch in `CreateDocIndex` and get silently skipped one by one. - Fix: track whether any token was skipped this way; only reject the command when that happened *and* no fields were parsed. This keeps `FT.CREATE idx ON HASH` (no trailing tokens at all, populated later via `FT.ALTER`) legal, matching existing behavior covered by the `AlterIndex` test. ## Behavior notes - `FT.CREATE <existing-index> ON HASH` (with no `SCHEMA`) now returns the missing-SCHEMA error instead of \"Index already exists\" — the command is malformed regardless of whether the index name already exists, so the syntax error takes precedence. - During a mixed-version rolling upgrade window, a schema-less `FT.CREATE` journaled by an old master is rejected when replayed on an already-upgraded replica. That index was already unusable and non-durable under the old behavior (didn't survive serialization, diverged FT.INFO from FT._LIST), so this doesn't introduce new master/replica divergence — the replica simply converges at the next full sync. ## Test plan - [x] Added `SearchFamilyTest.CreateWithoutSchemaKeywordIsError`, covering both the bug repro (rejected with a syntax error, no phantom index left behind) and the still-legal empty-schema-for-later-ALTER case. - [x] Full `search_family_test` suite passes (330/330), confirming no regression in existing `FT.CREATE` tests (including `CreateDropListIndex` and `AlterIndex`, which rely on schema-less creation succeeding). - [x] `pre-commit` (clang-format) passes on changed files. Fixes #7953",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8054",
        "createdAt": "2026-08-11T18:32:41Z",
        "updatedAt": "2026-08-13T15:59:08Z",
        "timestamp": "2026-08-13T15:59:08Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [],
        "author": "Shikha-code36",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8058",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(cluster): scope slot migration finalize pause to migrated slots only",
        "text": "## Summary `OutgoingMigration::FinalizeMigration` currently pauses **every** client command on the node — via `dfly::Pause(..., ClientPause::ALL, ...)` — for the entire duration of each finalize attempt, even though the only reason for the pause is to stop *new transactions on the slots being migrated* from initializing with stale cluster-slot-ownership information. The code even carried a `// TODO implement blocking on migrated slots only` acknowledging this. In practice this means: every time a slot migration finalizes (which can happen repeatedly across retries during a live resharding operation), **all** read/write traffic on the source node freezes — not just traffic touching the slots actually moving. For clusters doing rolling resharding, this shows up as a latency/availability spike across the entire keyspace, not just the part being migrated. This PR scopes that pause down to just the slots being migrated. ## What changed **1. Slot-scoped pause primitive (`ServerState` / `server_state.{h,cc}`)** - `ServerState` already tracked global `CLIENT PAUSE ALL`/`WRITE` state as simple per-thread counters (`client_pauses_[2]`). Added a parallel, independent mechanism: `paused_slot_ranges_` — a small vector of `cluster::SlotRanges` currently under a slot-scoped pause. - New `SetSlotPauseState(SlotRanges, bool start)` / `IsSlotPaused(SlotId)` / `HasActiveSlotPause()`. - `AwaitPauseState(bool is_write, optional<SlotId> slot = nullopt)` now also blocks if the given slot falls inside any active slot-scoped pause, **regardless of `is_write`** — both reads and writes must wait out a migration finalize, since the concern is stale ownership metadata, not write safety. - Concurrent migrations covering disjoint slot ranges just add independent entries; a command only waits if its own resolved slot intersects one of them. **2. `dfly::Pause()` (`server_family.{h,cc}`)** - Gained an optional `cluster::SlotRanges slot_ranges = nullopt` parameter. When set, it calls `SetSlotPauseState` instead of the global `SetPauseState` on every shard thread, both when starting the pause and when the pause-completion fiber tears it down. - Existing callers (`CLIENT PAUSE` via `ClientPauseCmd`) are unaffected — they don't pass `slot_ranges`, so behavior there is unchanged (still pauses everything, as `CLIENT PAUSE` should). **3. `OutgoingMigration::FinalizeMigration` (`outgoing_slot_migration.cc`)** - Now passes `migration_info_.slot_ranges` (the migration's own slots, already tracked) instead of relying on the blanket `ClientPause::ALL` behavior. Closes the standing TODO. **4. Dispatch path (`main_service.cc`)** - `CheckPauseState` now accepts an optional resolved `SlotId` and additionally gates on `ServerState::IsSlotPaused`. - In `DispatchCommand`, the command's slot is only resolved (`ResolveCommandSlot`) when `ServerState::HasActiveSlotPause()` is true — i.e., the extra per-command key lookup only happens while a migration finalize is actually in progress, so the common case (no migration running) pays zero extra cost. - Resolution errors (cross-slot, bad key spec) are deliberately ignored at this point and left to surface from the existing `CheckKeysOwnership` check later in the same dispatch — this call is only about deciding whether to block, not about validating the command. **5. Blocking commands (`transaction.cc`)** - `Transaction::WaitOnWatch` (the post-wakeup pause check for commands like `BLPOP`) now passes its own already-resolved `GetUniqueSlotId()` into `AwaitPauseState`, so blocking commands are correctly slot-scoped too instead of always waiting out any pause unconditionally. **6. Refactor: `ResolveCommandSlot` (`main_service.{h,cc}`)** - `CheckKeysOwnership` and `TakenOverSlotError` had near-identical key→slot resolution logic (`FindKeys` → `UniqueSlotChecker` → cross-slot check), flagged by a standing `// TODO(kostas) refactor` comment. Extracted into one shared helper both now call, with no behavior change. This is also the helper reused by the new pause-scoping logic in `DispatchCommand`. ## Why this design - Slot-scoping only matters in cluster mode with an active migration — by gating the extra key resolution on `HasActiveSlotPause()`, the hot dispatch path is untouched in the overwhelmingly common case (no migration running). - Reusing the existing `ClientPause` counters for the global case (rather than folding everything into one mechanism) keeps `CLIENT PAUSE`'s existing, well-tested semantics completely unchanged — this PR only adds a new, additive blocking condition. - The one-time `DispatchTracker` synchronization step at the start of `Pause()` (waiting for already-in-flight commands to finish dispatching) is intentionally left blanket/unscoped — it's a bounded, short wait for correctness at the moment the pause begins, not the sustained blocking the original TODO was about. ## Test plan - Added `ClusterFamilyTest.SlotScopedPauseOnlyBlocksMatchingSlot` (`cluster_family_test.cc`): starts a slot-scoped pause via `dfly::Pause()` directly (the same code path `FinalizeMigration` uses), then verifies: - a `GET` on a key outside the paused slot returns immediately (not delayed by the active pause), and - a `GET` on a key inside the paused slot blocks until the pause's timeout elapses. - Full suite run, no regressions: - `cluster_family_test`: 40/40 passed - `cluster_config_test`: 41/41 passed - `dragonfly_test`: 49/49 passed, 1 pre-existing unrelated skip (`DefragDflyEngineTest.TestDefragOption`) ## Not in scope Left untouched, as separate follow-ups (surfaced during the initial codebase survey but out of scope here): - Global commands being hard-rejected during incoming migration replay (`incoming_slot_migration.cc`) - `FLUSHSLOTS` not canceling an in-flight outgoing migration for the same slots - No LRU/LFU eviction policy - `PSYNC CONTINUE` (partial resync) unimplemented - `Transaction::ScheduleInternal` always taking the pessimistic multi-shard path for `MULTI`/`EVAL`",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8058",
        "createdAt": "2026-08-12T07:50:27Z",
        "updatedAt": "2026-08-13T10:57:21Z",
        "timestamp": "2026-08-13T10:57:21Z",
        "metrics": {
          "reactions": 0,
          "comments": 6
        },
        "labels": [],
        "author": "Shikha-code36",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8062",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "ci(tests): add scheduled e2e workflow for ioredis client",
        "text": "## What Adds a scheduled GitHub Actions workflow, `.github/workflows/ioredis-tests.yml`, that runs the existing [ioredis](https://github.com/luin/ioredis) end-to-end test suite against a freshly built Dragonfly binary every day, and points to it from `tests/README.md`. ## Why (closes part of #383) `#383` asked for e2e coverage of the `ioredis` client. Looking at the history: - `tests/integration/ioredis.Dockerfile` and `tests/integration/run_ioredis_on_docker.sh` (added in #394 / #459) already give a local script + README instructions for running the curated ioredis test suite against Dragonfly (deliverables 1 & 2 from the issue). - What was still missing is what @romange asked for in the issue thread: *\"we just need to run them via the testing pipeline on github\"*. There was no CI workflow exercising these tests — `ioredis`/`node-redis`/`jedis` Dockerfiles are all local-only, and grepping every workflow in `.github/workflows` shows none of them touch `tests/integration`. - The repo already has exactly this pattern for another ioredis-based client, `bullmq-tests.yml` (nightly scheduled job, builds Dragonfly from source, installs Node, runs the client's own test suite with a skip-list). This PR mirrors that same structure for `ioredis` itself, reusing the pre-existing `tests/integration/.run_ioredis_valid_test.sh` skip/grep list rather than duplicating it. ## What this does NOT do - It does not touch the existing local Docker scripts — those keep working as-is for local development. - It does not attempt the issue's \"bonus\" deliverable (a bespoke Dragonfly-authored pub/sub regression test); that's a separate, larger piece of work and is left as a possible follow-up so this PR stays focused on the CI-wiring gap. ## Testing Since building the full C++ project wasn't practical in my environment, I validated the actual test-execution path the workflow runs: - Started the official `dragonfly` Docker image (`--cluster_mode=emulated --lock_on_hashtags --dbfilename=`, same flags as `bullmq-tests.yml` uses for Dragonfly). - Cloned `luin/ioredis`, ran `npm install`, then ran the exact command the new workflow runs: `npm run env -- TS_NODE_TRANSPILE_ONLY=true NODE_ENV=test ./run_tests.sh` using the repo's own `tests/integration/.run_ioredis_valid_test.sh` as `run_tests.sh`. - Result: **493 passing / 2 failing**, both failures being `Timeout of 8000ms exceeded` on tests unrelated to correctness (a large-pipeline timing test and a cluster startup-nodes teardown hook) — consistent with the constrained local sandbox rather than a real regression. I couldn't run the pre-commit C++/format hooks in this environment (no C++ toolchain available), but this PR only touches a workflow YAML file and a doc file, so `clang-format`/`black` are not applicable. Closes (partially) #383",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8062",
        "createdAt": "2026-08-12T15:06:03Z",
        "updatedAt": "2026-08-13T08:30:57Z",
        "timestamp": "2026-08-13T08:30:57Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [],
        "author": "vjymisal0",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8064",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(stream): Reply with RESP3 map for XREAD BLOCK",
        "text": "XReadBlock gated the RESP3 map shape on opts->read_group, so a plain blocking XREAD always fell back to the RESP2 array shape even on a RESP3 connection. The non-blocking path already applies the map shape to both XREAD and XREADGROUP, so drop the read_group condition. Fixes #8056 <!-- **Commits Must Be Signed and Your PR title must conform to the conventional commit spec** * See: https://github.com/dragonflydb/dragonfly/blob/main/CONTRIBUTING.md * Please follow the section on `pre-commit hooks`, a linter will validate before you push Example PR Title: <type>(<scope>)!: <description> * `type` = bug, chore, feat, fix, docs, build, style, refactor, perf, test * `!` = OPTIONAL: signals a breaking change * `scope` = Optional when `type` is \"chore\" or \"docs\" * `description` = short description of the change Examples: * chore(examples): Clarify `docker` usage #120 * docs(readme): Fix Example Links #121 * feat(ingest)!: Add new ingest #122 * fix(ingest): Refactor for loop to list comprehension #123 -->",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8064",
        "createdAt": "2026-08-13T06:36:45Z",
        "updatedAt": "2026-08-13T08:14:43Z",
        "timestamp": "2026-08-13T08:14:43Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [],
        "author": "mkaruza",
        "state": "closed",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8065",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(server): emit expired keyspace events for already-past expirations",
        "text": "Deleting a key because its newly set expiration is already in the past emitted no `__keyevent@<db>__:expired` notification, so a subscriber waiting for the event was blocked forever (found via the upstream Valkey pubsub TCL suite: \"expired event (Expiration time is already expired)\" hangs). This covers the `DbSlice::UpdateExpire` funnel: EXPIRE/PEXPIRE/(P)EXPIREAT, GETEX, and memcached GAT/GATS.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8065",
        "createdAt": "2026-08-13T06:51:10Z",
        "updatedAt": "2026-08-13T16:35:07Z",
        "timestamp": "2026-08-13T16:35:07Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [],
        "author": "vyavdoshenko",
        "state": "open",
        "assignees": [
          "vyavdoshenko"
        ]
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8066",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "server: Fix stream size memory accounting",
        "text": "It has been seen in certain cases that the stream size accounting breaks due to a negative delta. In that case, the metrics depending on this type go completely wrong and start showing object size in terrabytes instead of a few GiB. There are a couple of fixes and safeguards: * original check that size>=delta if delta<0 was not doing anythingi it was always true. it is removed * now the change requested in bytes is compared with the curr. size and the limit ie 56 bit int maximum * if the value does not fit, then we do slow recomputation of the real size instead of clamping to 0 etc. There is also removed a buggy cast, previously we did something like: `SetSize(Size() + size)` So first the sum was cast to uint64_t, which might already wrap a negative value (if size is negative of more magnitude than Size()). Then that value was stored from uint64_t into a 56 bit field. In addition a negative delta was first cast by C++ itself to uint64_t. In the current code the cast for negative delta is done correctly and the case where delta is at boundary INT64_MIN is also handled --- The added test was failing without the changes.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8066",
        "createdAt": "2026-08-13T08:48:45Z",
        "updatedAt": "2026-08-13T09:18:19Z",
        "timestamp": "2026-08-13T09:18:19Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [],
        "author": "abhijat",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8075",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(pubsub): preserve V2 message ordering",
        "text": "This PR fixes a known issue that was already fixed for V1 at 5ef35182. Defer V2 pipeline execution when queued control work predates the next parsed command. Let IoLoopV2 drain that work so UNSUBSCRIBE cannot discard earlier Pub/Sub messages. Tests: Extend the unsubscribe regression across V1 and V2 with standalone and deferred-GET commands. Bound the starvation publisher wave to assert fairness without treating ordered backlog as starvation.",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8075",
        "createdAt": "2026-08-13T10:10:27Z",
        "updatedAt": "2026-08-13T13:50:16Z",
        "timestamp": "2026-08-13T13:50:16Z",
        "metrics": {
          "reactions": 0,
          "comments": 7
        },
        "labels": [],
        "author": "glevkovich",
        "state": "open",
        "assignees": []
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8077",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "fix(tiering): Use TraverseBySegmentOrder and export more metrics/configs",
        "text": "The changes that I merged into 1.40 branch Only functional change: 1. Fixes small bins defrag by using bounded TraverseBySegmentOrder to avoid blowing up iteration costs when the table is largely empty Metrics & config: 1. Exports missing tiering fields from info to metrics 2. Adds new configurable values to allow limiting CPU limit consumption of offloading / background scanning",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8077",
        "createdAt": "2026-08-13T11:05:19Z",
        "updatedAt": "2026-08-13T17:54:15Z",
        "timestamp": "2026-08-13T17:54:15Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [],
        "author": "dranikpg",
        "state": "open",
        "assignees": [],
        "change": "updated"
      },
      {
        "id": "github:dragonflydb/dragonfly:pull_request:8078",
        "source": "github",
        "group": "data-infrastructure",
        "project": "dragonflydb/dragonfly",
        "kind": "pull_request",
        "title": "chore: remove dead test_replication_onmove_flow",
        "text": "no point in time was removed in https://github.com/dragonflydb/dragonfly/pull/7068/changes",
        "url": "https://github.com/dragonflydb/dragonfly/pull/8078",
        "createdAt": "2026-08-13T14:53:44Z",
        "updatedAt": "2026-08-13T16:04:25Z",
        "timestamp": "2026-08-13T16:04:25Z",
        "metrics": {
          "reactions": 0,
          "comments": 4
        },
        "labels": [],
        "author": "kostasrim",
        "state": "closed",
        "assignees": [
          "kostasrim"
        ]
      }
    ],
    "events": [
      {
        "id": "event:9c952001400244794edb",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8053",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8053",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(generic): avoid data loss on RENAME when destination write fails",
          "text": "## Summary - `Renamer::FinalizeRename()` ran `DelSrc` and `DeserializeDest` in the same transaction hop with no ordering guarantee between shards. If `DeserializeDest` failed (e.g. OOM), the source key could already be deleted, silently losing its data. - Now deserializes into the destination first, and only deletes the source once that succeeds. `COPY` is unaffected since it never deletes the source. ## Test plan - [x] Added `GenericFamilyTest.RenameOOMPreservesSource`: drives real per-shard OOM via repeated `RESTORE`, then asserts a failed `RENAME` leaves the source key intact and does not create the destination. - [x] Full `generic_family_test` suite passes (84/84). - [x] `pre-commit` (clang-format) passes on changed files. Fixes #5296",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8053",
          "createdAt": "2026-08-11T17:02:15Z",
          "updatedAt": "2026-08-13T13:46:44Z",
          "timestamp": "2026-08-13T13:46:44Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [],
          "author": "Shikha-code36",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e5120e911e074db8daac",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8075",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8075",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(pubsub): preserve V2 message ordering",
          "text": "This PR fixes a known issue that was already fixed for V1 at 5ef35182. Defer V2 pipeline execution when queued control work predates the next parsed command. Let IoLoopV2 drain that work so UNSUBSCRIBE cannot discard earlier Pub/Sub messages. Tests: Extend the unsubscribe regression across V1 and V2 with standalone and deferred-GET commands. Bound the starvation publisher wave to assert fairness without treating ordered backlog as starvation.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8075",
          "createdAt": "2026-08-13T10:10:27Z",
          "updatedAt": "2026-08-13T13:29:03Z",
          "timestamp": "2026-08-13T13:29:03Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [],
          "author": "glevkovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1a290b7e414c57d9bc41",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8065",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8065",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(server): emit expired keyspace events for already-past expirations",
          "text": "Deleting a key because its newly set expiration is already in the past emitted no `__keyevent@<db>__:expired` notification, so a subscriber waiting for the event was blocked forever (found via the upstream Valkey pubsub TCL suite: \"expired event (Expiration time is already expired)\" hangs). This covers the `DbSlice::UpdateExpire` funnel: EXPIRE/PEXPIRE/(P)EXPIREAT, GETEX, and memcached GAT/GATS.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8065",
          "createdAt": "2026-08-13T06:51:10Z",
          "updatedAt": "2026-08-13T13:04:35Z",
          "timestamp": "2026-08-13T13:04:35Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "vyavdoshenko",
          "state": "open",
          "assignees": [
            "vyavdoshenko"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:c74f46b05ae4203c0a55",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8054",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8054",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(search): reject FT.CREATE missing the SCHEMA keyword",
          "text": "## Summary - `FT.CREATE` accepted an index definition missing the required `SCHEMA` keyword, returning `OK` and listing the index in `FT._LIST`, but the index was unusable (`FT.INFO` reported it as not found). - Root cause: without `SCHEMA`, tokens that look like field definitions fall through the \"unsupported parameters are ignored\" branch in `CreateDocIndex` and get silently skipped one by one. - Fix: track whether any token was skipped this way; only reject the command when that happened *and* no fields were parsed. This keeps `FT.CREATE idx ON HASH` (no trailing tokens at all, populated later via `FT.ALTER`) legal, matching existing behavior covered by the `AlterIndex` test. ## Test plan - [x] Added `SearchFamilyTest.CreateWithoutSchemaKeywordIsError`, covering both the bug repro (rejected with a syntax error, no phantom index left behind) and the still-legal empty-schema-for-later-ALTER case. - [x] Full `search_family_test` suite passes (330/330), confirming no regression in existing `FT.CREATE` tests (including `CreateDropListIndex` and `AlterIndex`, which rely on schema-less creation succeeding). - [x] `pre-commit` (clang-format) passes on changed files. Fixes #7953",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8054",
          "createdAt": "2026-08-11T18:32:41Z",
          "updatedAt": "2026-08-13T12:42:13Z",
          "timestamp": "2026-08-13T12:42:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [],
          "author": "Shikha-code36",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:3b1494884e7621865163",
        "signalId": "github:dragonflydb/dragonfly:pull_request:7997",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:7997",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "chore: add docs/replication.md",
          "text": "* add docs on replication state machine and the internal functionally at each phase",
          "url": "https://github.com/dragonflydb/dragonfly/pull/7997",
          "createdAt": "2026-08-04T10:07:57Z",
          "updatedAt": "2026-08-13T11:57:25Z",
          "timestamp": "2026-08-13T11:57:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [],
          "author": "kostasrim",
          "state": "open",
          "assignees": [
            "kostasrim"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:73912d1d792bc51a0dfe",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8077",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8077",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(tiering): Use TraverseBySegmentOrder and export more metrics/configs",
          "text": "The changes that I merged into 1.40 branch Only functional change: 1. Fixes small bins defrag by using bounded TraverseBySegmentOrder to avoid blowing up iteration costs when the table is largely empty Metrics & config: 1. Exports missing tiering fields from info to metrics 2. Adds new configurable values to allow limiting CPU limit consumption of offloading / background scanning",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8077",
          "createdAt": "2026-08-13T11:05:19Z",
          "updatedAt": "2026-08-13T11:12:19Z",
          "timestamp": "2026-08-13T11:12:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "dranikpg",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:e2a182b5079f044bb1f8",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8033",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8033",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "chore(server): Prefetch + cid caching for speedup",
          "text": "1. Prefetching bucket data It requires really large pipelines to reach large differences <img width=\"1650\" height=\"660\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5b675ef9-35f4-4a82-b157-941f9a6a22ef\" /> 2. Caching last command id Helps avoiding absl::AsciiStrToUpper + hashtable lookup + ifs for simple cases like all GET pipeline. Adds additionally around +1.5%",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8033",
          "createdAt": "2026-08-07T15:58:04Z",
          "updatedAt": "2026-08-13T11:06:29Z",
          "timestamp": "2026-08-13T11:06:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [],
          "author": "dranikpg",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:ebfea7b9f0fcd8419e0a",
        "signalId": "github:dragonflydb/dragonfly:issue:8074",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8074",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P3 — FLUSHSLOTS TOCTOU between the validation and action hops of a woken multi-stream read",
          "text": "The woken blocked read runs a cross-shard validation hop and then an action hop that revalidates only the wake-key shard (`stream_family.cc:3117-3137`, `3156-3164`); `FlushSlotsFb()` deletes slot keys in a detached fiber without transaction locks and yields between chunks (`db_slice.cc:955-985`). A flush landing between the two hops can delete a sibling stream after it passed validation, so the client is served data from the ready stream although the other requested group no longer exists. Impact is benign in practice: under cluster mode both streams share a slot (CROSSSLOT), the reply legally serializes before the flush acknowledgment, and there is no crash or replication divergence — listed for completeness. A deterministic test would need a latch between the two hops.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8074",
          "createdAt": "2026-08-13T10:03:03Z",
          "updatedAt": "2026-08-13T11:00:28Z",
          "timestamp": "2026-08-13T11:00:28Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [
            "mkaruza"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:4e03c86b305a77824cf1",
        "signalId": "github:dragonflydb/dragonfly:issue:8073",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8073",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P2 — `DFLYCLUSTER FLUSHSLOTS` operates on the default namespace regardless of the caller",
          "text": "`cluster_family.cc:519` takes the DbSlice from `namespaces->GetDefaultNamespace()` instead of the calling transaction's namespace. A namespace-bound ACL user with `+@admin` / `+dflycluster` can invoke the command; it then flushes default-namespace data while the tenant's keys in those slots survive indefinitely (including past slot migration), and blocked tenant readers on them are never woken. This is systemic in the cluster subsystem — slot migration, `GETSLOTINFO`, and slot key counts also hard-code the default namespace (`outgoing_slot_migration.cc`, `cluster_utility.cc:28`) — so either cluster commands should be namespace-aware, or namespace-bound users should be barred from them explicitly. Scenario: ```text # put keys of the same slot into both the default namespace and NAMESPACE:tenant # authenticate as the tenant user and run DFLYCLUSTER FLUSHSLOTS <slot> <slot> ``` Actual: default-namespace data in the slot is deleted; tenant data survives.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8073",
          "createdAt": "2026-08-13T10:02:46Z",
          "updatedAt": "2026-08-13T10:59:53Z",
          "timestamp": "2026-08-13T10:59:53Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [
            "mkaruza"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:d2e584a76c80326371fd",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8058",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8058",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(cluster): scope slot migration finalize pause to migrated slots only",
          "text": "## Summary `OutgoingMigration::FinalizeMigration` currently pauses **every** client command on the node — via `dfly::Pause(..., ClientPause::ALL, ...)` — for the entire duration of each finalize attempt, even though the only reason for the pause is to stop *new transactions on the slots being migrated* from initializing with stale cluster-slot-ownership information. The code even carried a `// TODO implement blocking on migrated slots only` acknowledging this. In practice this means: every time a slot migration finalizes (which can happen repeatedly across retries during a live resharding operation), **all** read/write traffic on the source node freezes — not just traffic touching the slots actually moving. For clusters doing rolling resharding, this shows up as a latency/availability spike across the entire keyspace, not just the part being migrated. This PR scopes that pause down to just the slots being migrated. ## What changed **1. Slot-scoped pause primitive (`ServerState` / `server_state.{h,cc}`)** - `ServerState` already tracked global `CLIENT PAUSE ALL`/`WRITE` state as simple per-thread counters (`client_pauses_[2]`). Added a parallel, independent mechanism: `paused_slot_ranges_` — a small vector of `cluster::SlotRanges` currently under a slot-scoped pause. - New `SetSlotPauseState(SlotRanges, bool start)` / `IsSlotPaused(SlotId)` / `HasActiveSlotPause()`. - `AwaitPauseState(bool is_write, optional<SlotId> slot = nullopt)` now also blocks if the given slot falls inside any active slot-scoped pause, **regardless of `is_write`** — both reads and writes must wait out a migration finalize, since the concern is stale ownership metadata, not write safety. - Concurrent migrations covering disjoint slot ranges just add independent entries; a command only waits if its own resolved slot intersects one of them. **2. `dfly::Pause()` (`server_family.{h,cc}`)** - Gained an optional `cluster::SlotRanges slot_ranges = nullopt` parameter. When set, it calls `SetSlotPauseState` instead of the global `SetPauseState` on every shard thread, both when starting the pause and when the pause-completion fiber tears it down. - Existing callers (`CLIENT PAUSE` via `ClientPauseCmd`) are unaffected — they don't pass `slot_ranges`, so behavior there is unchanged (still pauses everything, as `CLIENT PAUSE` should). **3. `OutgoingMigration::FinalizeMigration` (`outgoing_slot_migration.cc`)** - Now passes `migration_info_.slot_ranges` (the migration's own slots, already tracked) instead of relying on the blanket `ClientPause::ALL` behavior. Closes the standing TODO. **4. Dispatch path (`main_service.cc`)** - `CheckPauseState` now accepts an optional resolved `SlotId` and additionally gates on `ServerState::IsSlotPaused`. - In `DispatchCommand`, the command's slot is only resolved (`ResolveCommandSlot`) when `ServerState::HasActiveSlotPause()` is true — i.e., the extra per-command key lookup only happens while a migration finalize is actually in progress, so the common case (no migration running) pays zero extra cost. - Resolution errors (cross-slot, bad key spec) are deliberately ignored at this point and left to surface from the existing `CheckKeysOwnership` check later in the same dispatch — this call is only about deciding whether to block, not about validating the command. **5. Blocking commands (`transaction.cc`)** - `Transaction::WaitOnWatch` (the post-wakeup pause check for commands like `BLPOP`) now passes its own already-resolved `GetUniqueSlotId()` into `AwaitPauseState`, so blocking commands are correctly slot-scoped too instead of always waiting out any pause unconditionally. **6. Refactor: `ResolveCommandSlot` (`main_service.{h,cc}`)** - `CheckKeysOwnership` and `TakenOverSlotError` had near-identical key→slot resolution logic (`FindKeys` → `UniqueSlotChecker` → cross-slot check), flagged by a standing `// TODO(kostas) refactor` comment. Extracted into one shared helper both now call, with no behavior change. This is also the helper reused by the new pause-scoping logic in `DispatchCommand`. ## Why this design - Slot-scoping only matters in cluster mode with an active migration — by gating the extra key resolution on `HasActiveSlotPause()`, the hot dispatch path is untouched in the overwhelmingly common case (no migration running). - Reusing the existing `ClientPause` counters for the global case (rather than folding everything into one mechanism) keeps `CLIENT PAUSE`'s existing, well-tested semantics completely unchanged — this PR only adds a new, additive blocking condition. - The one-time `DispatchTracker` synchronization step at the start of `Pause()` (waiting for already-in-flight commands to finish dispatching) is intentionally left blanket/unscoped — it's a bounded, short wait for correctness at the moment the pause begins, not the sustained blocking the original TODO was about. ## Test plan - Added `ClusterFamilyTest.SlotScopedPauseOnlyBlocksMatchingSlot` (`cluster_family_test.cc`): starts a slot-scoped pause via `dfly::Pause()` directly (the same code path `FinalizeMigration` uses), then verifies: - a `GET` on a key outside the paused slot returns immediately (not delayed by the active pause), and - a `GET` on a key inside the paused slot blocks until the pause's timeout elapses. - Full suite run, no regressions: - `cluster_family_test`: 40/40 passed - `cluster_config_test`: 41/41 passed - `dragonfly_test`: 49/49 passed, 1 pre-existing unrelated skip (`DefragDflyEngineTest.TestDefragOption`) ## Not in scope Left untouched, as separate follow-ups (surfaced during the initial codebase survey but out of scope here): - Global commands being hard-rejected during incoming migration replay (`incoming_slot_migration.cc`) - `FLUSHSLOTS` not canceling an in-flight outgoing migration for the same slots - No LRU/LFU eviction policy - `PSYNC CONTINUE` (partial resync) unimplemented - `Transaction::ScheduleInternal` always taking the pessimistic multi-shard path for `MULTI`/`EVAL`",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8058",
          "createdAt": "2026-08-12T07:50:27Z",
          "updatedAt": "2026-08-13T10:57:21Z",
          "timestamp": "2026-08-13T10:57:21Z",
          "metrics": {
            "reactions": 0,
            "comments": 6
          },
          "labels": [],
          "author": "Shikha-code36",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:28e9d0e1e73717455f09",
        "signalId": "github:dragonflydb/dragonfly:issue:8076",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8076",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "Return RESP3 null for empty XREAD and XREADGROUP replies",
          "text": "Under RESP3, XREAD and XREADGROUP return *-1 when no entries are available or a BLOCK timeout occurs. Redis returns the RESP3 null value (_\\r\\n) instead. Update these paths to emit RESP3 null replies while preserving *-1 for RESP2.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8076",
          "createdAt": "2026-08-13T10:41:53Z",
          "updatedAt": "2026-08-13T10:45:19Z",
          "timestamp": "2026-08-13T10:45:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [
            "mkaruza"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:2def1bbf244c4c0db9fe",
        "signalId": "github:dragonflydb/dragonfly:issue:7903",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:7903",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "Blocked XREADGROUP is not woken when the watched stream is deleted or retyped",
          "text": "A client blocked with `XREADGROUP ... BLOCK 0 STREAMS <key> >` remains blocked when the watched stream becomes invalid. Operations such as `DEL`, expiration, `FLUSHDB`/`FLUSHALL`, or replacing the stream with another data type do not trigger reevaluation of the blocked command. With `BLOCK 0`, the connection can remain blocked indefinitely. Reproduction: Create a stream and consumer group: ```text XADD mystream 1-0 field value XGROUP CREATE mystream mygroup $ ``` In client A: ```text XREADGROUP GROUP mygroup consumer BLOCK 0 STREAMS mystream > ``` Wait until the client is blocked. In client B: ```text DEL mystream ``` Actual behavior: Client A remains blocked indefinitely. `INFO CLIENTS` continues to report it as a blocked client. Expected behavior: Client A should be woken and the command should be reevaluated. Since the stream and consumer group no longer exist, it should return a `NOGROUP` error. Equivalent cases: - `DEL`, `UNLINK`, expiration, `FLUSHDB`, and `FLUSHALL` should return `NOGROUP`. - Replacing the stream with a non-stream value should return `WRONGTYPE`. Plain `XREAD` should retain its existing semantics and remain blocked when a key disappears. The wake-on invalidation behavior is specifically required for `XREADGROUP`. Impact: - `BLOCK 0` connections can hang permanently. - Applications cannot recover when streams are deleted, expired, flushed, or recreated. - Blocked clients and their associated transaction resources remain retained. - The behavior differs from Valkey/Redis blocking-command semantics. Technical notes: The readiness checker can detect a missing key, missing consumer group, or wrong value type. However, deletion, expiration, flush, and retype paths do not notify the blocking controller, so the checker is never invoked. The underlying issue is that a blocked command is not reevaluated after invalidation of its watched key.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/7903",
          "createdAt": "2026-07-21T16:05:18Z",
          "updatedAt": "2026-08-13T10:39:52Z",
          "timestamp": "2026-08-13T10:39:52Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "vyavdoshenko",
          "state": "closed",
          "assignees": [
            "BorysTheDev"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:78324869ec8cc301ff23",
        "signalId": "github:dragonflydb/dragonfly:issue:8072",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8072",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P2 — Multi-shard error precedence follows shard placement instead of argument order",
          "text": "For a woken blocked read, per-shard validation statuses are scanned in numeric shard-ID order (`stream_family.cc:3119-3135`), so the reported error is decided by key-to-shard hashing. Within a single shard, argument order is respected — the defect is strictly cross-shard. The initial non-blocked path is worse: the aggregate error (`AggregateValue`, `execution_state.h:21-28`) is written by whichever shard callback runs first, so identical invocations on identical state return different errors run to run (observed as a 10/10 split over 20 calls). Valkey deterministically reports the error of the first invalid argument in command order. Scenario (keys `a` and `b` on different shards, `a` listed first): ```text XGROUP CREATE a g 0 MKSTREAM XGROUP CREATE b g 0 MKSTREAM # client A XREADGROUP GROUP g c BLOCK 0 STREAMS a b > > # client B: invalidate both in one transaction MULTI XGROUP DESTROY a g DEL b SET b v EXEC ``` Actual: the returned error follows the lower-shard key regardless of its argument position (`WRONGTYPE` here although `NOGROUP` for `a` is expected; swapping which key is destroyed/retyped flips it accordingly).",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8072",
          "createdAt": "2026-08-13T10:02:29Z",
          "updatedAt": "2026-08-13T10:02:29Z",
          "timestamp": "2026-08-13T10:02:29Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:194dd7c0a5c50c2aee43",
        "signalId": "github:dragonflydb/dragonfly:issue:8071",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8071",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P1 — A blocked multi-stream read returns only one stream when several become ready together",
          "text": "The wake path resolves only `Transaction::GetWakeKey()` (`stream_family.cc:3139-3142`) and builds a reply with cardinality exactly one (`stream_family.cc:3221-3229`). The root cause is in the transaction layer: `NotifySuspended` claims a once-only blocking barrier and records a single wake key (`transaction.cc:1578`, `1598`), so the woken command structurally cannot see that several keys became ready — a fix needs to re-scan all requested streams on wake, as Valkey's unblock reprocessing does. Pre-existing bug, not introduced by #8047. Scenario (single shard for determinism): ```text XGROUP CREATE s1 g 0 MKSTREAM XGROUP CREATE s2 g 0 MKSTREAM # client A XREADGROUP GROUP g c BLOCK 0 STREAMS s1 s2 > > # client B MULTI XADD s1 1-0 f one XADD s2 1-0 f two EXEC ``` Actual: the reply contains only one stream (e.g. `s2`), and only that stream gets PEL treatment; `XPENDING s1 g` stays empty. The omitted entries are not lost (last-delivered on `s1` is unchanged, a later `>` read retrieves them), but the reply is incomplete vs Valkey and the PEL is asymmetric. Plain blocked `XREAD` has the same single-stream reply, without the PEL aspect.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8071",
          "createdAt": "2026-08-13T10:02:13Z",
          "updatedAt": "2026-08-13T10:02:13Z",
          "timestamp": "2026-08-13T10:02:13Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1fd815ae0f08d0fbd26f",
        "signalId": "github:dragonflydb/dragonfly:issue:8070",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8070",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P1 — Multi-stream `XREADGROUP` mutates group state and then returns an error",
          "text": "`HasEntries2()` creates or updates the consumer while checking readiness (`stream_family.cc:4042-4055`), and the single-shard path performs `OpRead()` and journals group side effects (`stream_family.cc:3262-3272`) before the accumulated error is checked (`stream_family.cc:3285-3287`). Pre-existing bug, not introduced by the wake fixes. Scenario: ```text XGROUP CREATE good g 0 MKSTREAM XADD good 1-0 f v XREADGROUP GROUP g c STREAMS good missing > > ``` Actual: the reply is `NOGROUP`, yet entry `1-0` is inserted into the PEL, consumer `c` is created, and the group's last-delivered ID advances to `1-0` — the client never receives the entry it now owns. This happens whenever both keys land on the same shard (any thread count, not only `--proactor_threads=1`). Expected: validation of all streams before any consumer-group mutation. Cross-shard variant: `OpRead()` is skipped on the error path, but the consumer is still created on the valid stream without any journal write — a replication-divergence risk.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8070",
          "createdAt": "2026-08-13T10:01:54Z",
          "updatedAt": "2026-08-13T10:01:54Z",
          "timestamp": "2026-08-13T10:01:54Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4f2e1b315a1894013211",
        "signalId": "github:dragonflydb/dragonfly:issue:8069",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8069",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P1 — Active expiry never runs for non-default namespaces",
          "text": "Heartbeat expiry and eviction scan only the default namespace (`engine_shard.cc:784-785`, `862-864`), and the post-expiry wake dispatch uses only the default namespace's `BlockingController` (`engine_shard.cc:959-964`). Keys in ACL namespaces are never actively expired at all — no deletion, no keyspace notification, no wake. Only a later tenant command that touches the key expires it lazily. Scenario: ```text ACL SETUSER u NAMESPACE:tenant ON >p +@all ~* # tenant connection 1 XGROUP CREATE s g 0 MKSTREAM XREADGROUP GROUP g c BLOCK 0 STREAMS s > # tenant connection 2 PEXPIRE s 100 ``` Actual: the blocked tenant client is never woken; with `BLOCK 0` it stays blocked indefinitely. A later tenant command touching `s` (even `EXISTS s`) lazily expires the key and only then the reader receives `NOGROUP`. The identical scenario in the default namespace correctly delivers `NOGROUP` shortly after the TTL elapses.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8069",
          "createdAt": "2026-08-13T10:01:39Z",
          "updatedAt": "2026-08-13T10:01:39Z",
          "timestamp": "2026-08-13T10:01:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:524da54d71df9e42daef",
        "signalId": "github:dragonflydb/dragonfly:issue:8068",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8068",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P1 — `NotifyPending()` is reentrant through a suspending expiry checker",
          "text": "`NotifyPending()` iterates the live `awakened_indices_` / `awakened_keys` sets with no reentrancy guard and no snapshot (`blocking_controller.cc:150-182`). The stream readiness checker calls `FindReadOnly()` (`stream_family.cc:3032`), which can enter `ExpireIfNeeded()` (`db_slice.cc:1451-1493`); that path can suspend **before** deleting the entry — in `RecordExpiryBlocking()` when a replication journal is active (`db_slice.cc:1471-1473`), or in keyspace-notification delivery under publish-buffer backpressure (`db_slice.cc:1476-1484`, `channel_store.cc:120`). While the checker fiber is suspended, any concluding transaction on the shard calls `NotifyPending()` again (`transaction.cc:668`). The inner call clears and erases the containers the outer frames still reference — including the `DbWatchTable` and `WatchQueue` held by reference — so the outer call resumes on invalidated iterators and a stale `PrimeIterator`. Consequences: lost wake, CHECK/DCHECK failure, use-after-free, duplicate erase. Reachability: requires an active replication journal with a stalled streamer, or `notify-keyspace-events Ex` with a backpressured subscriber. A default standalone configuration has no suspension point on this path. The synchronous same-fiber lazy-expiry case is benign and already covered by `XReadGroupBlockLazyExpireDuringWakeDoesNotCrash`; the uncovered defect is the cross-fiber suspension window. Scenario (race-sensitive): ```text CONFIG SET notify-keyspace-events Ex PSUBSCRIBE __keyevent@0__:expired # subscriber stops consuming replies # fill the subscriber's output budget so an expired event must wait XGROUP CREATE s g 0 MKSTREAM XREADGROUP GROUP g c BLOCK 0 STREAMS s > # client A PEXPIRE s 10 # client B # after the TTL elapses, run another write on the same shard SET unrelated-key value ``` A deterministic test needs a latch inside `RecordExpiryBlocking()` or the notification send; a timing-only test would be flaky.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8068",
          "createdAt": "2026-08-13T10:01:22Z",
          "updatedAt": "2026-08-13T10:01:22Z",
          "timestamp": "2026-08-13T10:01:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1476aa6257a4e7cdf370",
        "signalId": "github:dragonflydb/dragonfly:issue:8067",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8067",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "P1 — A heterogeneous blocking queue can hide XREADGROUP forever",
          "text": "#### 1. P1 — A heterogeneous blocking queue can hide `XREADGROUP` forever `BlockingController` keeps one FIFO queue per `(db, key)` for all blocking commands. `NotifyWatchQueue()` checks only the front waiter; when its checker returns `kKeyNotFound` (BLPOP / BZPOP\\* / BLMPOP on a missing key), the whole scan aborts (`blocking_controller.cc:259-265`) and the one-shot awakened event is cleared (`blocking_controller.cc:175`), so later waiters in the same queue are never checked. The TODO at `stream_family.cc:3037-3039` documents the limitation. The trigger is irrelevant: DEL, UNLINK, expiry, FLUSHDB/FLUSHALL, and FLUSHSLOTS all hit the same break. Scenario: ```text # client A: first waiter while k does not exist BLPOP k 0 # client B XGROUP CREATE k g 0 MKSTREAM # client C: later waiter in the same (db, key) queue XREADGROUP GROUP g c BLOCK 0 STREAMS k > # client B DEL k ``` Actual: client C receives nothing and stays blocked indefinitely; it is released only by a later unrelated mutation of `k` (e.g. `LPUSH k v`, which also serves client A). Expected: client A remains blocked, client C immediately receives `NOGROUP`. Same result with `BZPOPMIN` in place of `BLPOP`.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8067",
          "createdAt": "2026-08-13T10:00:27Z",
          "updatedAt": "2026-08-13T10:00:58Z",
          "timestamp": "2026-08-13T10:00:58Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug"
          ],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:2f478666381e30160c78",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8066",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8066",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "server: Fix stream size memory accounting",
          "text": "It has been seen in certain cases that the stream size accounting breaks due to a negative delta. In that case, the metrics depending on this type go completely wrong and start showing object size in terrabytes instead of a few GiB. There are a couple of fixes and safeguards: * original check that size>=delta if delta<0 was not doing anythingi it was always true. it is removed * now the change requested in bytes is compared with the curr. size and the limit ie 56 bit int maximum * if the value does not fit, then we do slow recomputation of the real size instead of clamping to 0 etc. There is also removed a buggy cast, previously we did something like: `SetSize(Size() + size)` So first the sum was cast to uint64_t, which might already wrap a negative value (if size is negative of more magnitude than Size()). Then that value was stored from uint64_t into a 56 bit field. In addition a negative delta was first cast by C++ itself to uint64_t. In the current code the cast for negative delta is done correctly and the case where delta is at boundary INT64_MIN is also handled --- The added test was failing without the changes.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8066",
          "createdAt": "2026-08-13T08:48:45Z",
          "updatedAt": "2026-08-13T09:18:19Z",
          "timestamp": "2026-08-13T09:18:19Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "abhijat",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:46e06921b2f12cee81a4",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8062",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8062",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "ci(tests): add scheduled e2e workflow for ioredis client",
          "text": "## What Adds a scheduled GitHub Actions workflow, `.github/workflows/ioredis-tests.yml`, that runs the existing [ioredis](https://github.com/luin/ioredis) end-to-end test suite against a freshly built Dragonfly binary every day, and points to it from `tests/README.md`. ## Why (closes part of #383) `#383` asked for e2e coverage of the `ioredis` client. Looking at the history: - `tests/integration/ioredis.Dockerfile` and `tests/integration/run_ioredis_on_docker.sh` (added in #394 / #459) already give a local script + README instructions for running the curated ioredis test suite against Dragonfly (deliverables 1 & 2 from the issue). - What was still missing is what @romange asked for in the issue thread: *\"we just need to run them via the testing pipeline on github\"*. There was no CI workflow exercising these tests — `ioredis`/`node-redis`/`jedis` Dockerfiles are all local-only, and grepping every workflow in `.github/workflows` shows none of them touch `tests/integration`. - The repo already has exactly this pattern for another ioredis-based client, `bullmq-tests.yml` (nightly scheduled job, builds Dragonfly from source, installs Node, runs the client's own test suite with a skip-list). This PR mirrors that same structure for `ioredis` itself, reusing the pre-existing `tests/integration/.run_ioredis_valid_test.sh` skip/grep list rather than duplicating it. ## What this does NOT do - It does not touch the existing local Docker scripts — those keep working as-is for local development. - It does not attempt the issue's \"bonus\" deliverable (a bespoke Dragonfly-authored pub/sub regression test); that's a separate, larger piece of work and is left as a possible follow-up so this PR stays focused on the CI-wiring gap. ## Testing Since building the full C++ project wasn't practical in my environment, I validated the actual test-execution path the workflow runs: - Started the official `dragonfly` Docker image (`--cluster_mode=emulated --lock_on_hashtags --dbfilename=`, same flags as `bullmq-tests.yml` uses for Dragonfly). - Cloned `luin/ioredis`, ran `npm install`, then ran the exact command the new workflow runs: `npm run env -- TS_NODE_TRANSPILE_ONLY=true NODE_ENV=test ./run_tests.sh` using the repo's own `tests/integration/.run_ioredis_valid_test.sh` as `run_tests.sh`. - Result: **493 passing / 2 failing**, both failures being `Timeout of 8000ms exceeded` on tests unrelated to correctness (a large-pipeline timing test and a cluster startup-nodes teardown hook) — consistent with the constrained local sandbox rather than a real regression. I couldn't run the pre-commit C++/format hooks in this environment (no C++ toolchain available), but this PR only touches a workflow YAML file and a doc file, so `clang-format`/`black` are not applicable. Closes (partially) #383",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8062",
          "createdAt": "2026-08-12T15:06:03Z",
          "updatedAt": "2026-08-13T08:30:57Z",
          "timestamp": "2026-08-13T08:30:57Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "vjymisal0",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4062fe75b80b86a7bfc2",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8064",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8064",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(stream): Reply with RESP3 map for XREAD BLOCK",
          "text": "XReadBlock gated the RESP3 map shape on opts->read_group, so a plain blocking XREAD always fell back to the RESP2 array shape even on a RESP3 connection. The non-blocking path already applies the map shape to both XREAD and XREADGROUP, so drop the read_group condition. Fixes #8056 <!-- **Commits Must Be Signed and Your PR title must conform to the conventional commit spec** * See: https://github.com/dragonflydb/dragonfly/blob/main/CONTRIBUTING.md * Please follow the section on `pre-commit hooks`, a linter will validate before you push Example PR Title: <type>(<scope>)!: <description> * `type` = bug, chore, feat, fix, docs, build, style, refactor, perf, test * `!` = OPTIONAL: signals a breaking change * `scope` = Optional when `type` is \"chore\" or \"docs\" * `description` = short description of the change Examples: * chore(examples): Clarify `docker` usage #120 * docs(readme): Fix Example Links #121 * feat(ingest)!: Add new ingest #122 * fix(ingest): Refactor for loop to list comprehension #123 -->",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8064",
          "createdAt": "2026-08-13T06:36:45Z",
          "updatedAt": "2026-08-13T08:14:43Z",
          "timestamp": "2026-08-13T08:14:43Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "mkaruza",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:7c089932b3257caf4f99",
        "signalId": "github:dragonflydb/dragonfly:issue:8056",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8056",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "XREAD BLOCK replies in RESP2 array shape on a RESP3 connection",
          "text": "**Describe the bug** On a RESP3 connection, a blocking `XREAD` that is woken by a new entry replies in the RESP2 array shape instead of the RESP3 map shape. Real Redis replies with a map on the same path. This breaks clients that trust the negotiated protocol. `node-redis` selects its reply parser from the `HELLO` version, so it applies the RESP3 transform to the RESP2 array and throws `TypeError: Cannot read properties of undefined (reading 'map')`. The stream is then never consumed. The non-blocking path is correct. `XREAD` that returns immediately with data does reply with a map on RESP3. Only the blocking-wakeup path differs, which is what makes it easy to miss. `XREADGROUP` was fixed to reply as a map for RESP3 in v1.22.0 (#3639). Plain `XREAD` appears never to have been covered. **To Reproduce** ```js // npm i redis@6 import { createClient } from 'redis'; const url = 'redis://127.0.0.1:6379'; const KEY = 'repro:stream'; const writer = createClient({ url, RESP: 2 }); const reader = createClient({ url, RESP: 3 }); await writer.connect(); await reader.connect(); await writer.del(KEY); // Block from \"now\", then publish so the block wakes with an entry. const pending = reader.sendCommand(['XREAD', 'BLOCK', '3000', 'STREAMS', KEY, '$']); setTimeout(() => writer.xAdd(KEY, '*', { event: 'wake' }), 300); const reply = await pending; console.log(Array.isArray(reply) ? 'Array (RESP2 shape)' : 'Map/Object (RESP3 shape)'); ``` **Expected behavior** `Map/Object (RESP3 shape)`, matching Redis. **Actual behavior** `Array (RESP2 shape)`. Using the typed `reader.xRead({ key: KEY, id: '$' }, { BLOCK: 3000 })` instead of `sendCommand` throws `TypeError: Cannot read properties of undefined (reading 'map')`. **Results across servers** Same script, same client version, only the server changed: | Server | Blocking-wakeup reply on RESP3 | |---|---| | Dragonfly v1.38.1 | Array (RESP2 shape) | | Dragonfly v1.40.1 | Array (RESP2 shape) | | Redis 7.4 | Map (RESP3 shape) | | Redis 8 | Map (RESP3 shape) | RESP2 behaves identically on all four, so pinning the client to RESP2 is a viable workaround. **Environment** - Dragonfly: `df-v1.38.1` and `df-v1.40.1`, official images, default config, standalone - Client: `node-redis` 6.2.0 - Reproduced on both a local container and a deployed instance **Additional context** This surfaced in production as a service that publishes to a stream from several pods and reads it back with one blocking `XREAD` per pod. Reads failed continuously while writes succeeded, so the stream grew and nothing consumed it. The client's own error was the `TypeError` above, which points at the client rather than the server and made it slow to diagnose.",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8056",
          "createdAt": "2026-08-12T02:40:13Z",
          "updatedAt": "2026-08-13T08:14:39Z",
          "timestamp": "2026-08-13T08:14:39Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [],
          "author": "BryanChow0112",
          "state": "closed",
          "assignees": [
            "mkaruza"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:b427fe3af625f89e223f",
        "signalId": "github:dragonflydb/dragonfly:issue:5662",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:5662",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "test_replicaof_reject_on_load",
          "text": "\\https://github.com/dragonflydb/dragonfly/actions/runs/16884994243/job/47830310679#step:6:1077",
          "url": "https://github.com/dragonflydb/dragonfly/issues/5662",
          "createdAt": "2025-08-12T07:04:07Z",
          "updatedAt": "2026-08-13T07:05:50Z",
          "timestamp": "2026-08-13T07:05:50Z",
          "metrics": {
            "reactions": 0,
            "comments": 37
          },
          "labels": [
            "bug",
            "failing-test",
            "epoll",
            "iouring"
          ],
          "author": "kostasrim",
          "state": "open",
          "assignees": [
            "abhijat"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:bab1bc95a2b9ba4040bc",
        "signalId": "github:dragonflydb/dragonfly:issue:8063",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8063",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "test_replicate_old_master",
          "text": "Automated regression triage opened this issue. - Test: `dragonfly/replication_config_test.py::test_replicate_old_master[df_factory0--localhost-7000]` - Backend: iouring - Commit: f4019d7fec0ddcd1e6484dd6eeade7d52b146af6 - Run: https://github.com/dragonflydb/dragonfly/actions/runs/31641685740 - Marker: run/31641685740/attempt/1 ``` Reason: requests.exceptions.RetryError: HTTPSConnectionPool(host='github.com', port=443): Max retries exceeded with url: /dragonflydb/dragonfly/releases/download/v1.19.2/dragonfly-x86_64.tar.gz (Caused by ResponseError('too many 503 error responses')) ____________ test_replicate_old_master[df_factory0--localhost-7000] ____________ urllib3.exceptions.ResponseError: too many 503 error responses The above exception was the direct cause of the following exception: self = <requests.adapters.HTTPAdapter object at 0x7fdbe0bd0b30> request = <PreparedRequest [GET]>, stream = True, timeout = None, verify = True cert = None, proxies = OrderedDict() def send( self, request: PreparedRequest, stream: bool = False, timeout: _t.TimeoutType = None, verify: _t.VerifyType = True, cert: _t.CertType = None, proxies: dict[str, str] | None = None, ) -> Response: \"\"\"Sends PreparedRequest object. Returns Response object. :param request: The :class:`PreparedRequest <PreparedRequest>` being sent. :param stream: (optional) Whether to stream the request content. :param timeout: (optional) How long to wait for the server to send data before giving up, as a float, or a :ref:`(connect timeout, read timeout) <timeouts>` tuple. :type timeout: float or tuple or urllib3 Timeout object :param verify: (optional) Either a boolean, in which case it controls whether we verify the server's TLS certificate, or a string, in which case it must be a path to a CA bundle to use :param cert: (optional) Any user-provided SSL certificate to be trusted. :param proxies: (optional) The proxies dictionary to apply to the request. :rtype: requests.Response \"\"\" assert _is_prepared(request) try: conn = self.get_connection_with_tls_context( request, verify, proxies=proxies, cert=cert ) except LocationValueError as e: raise InvalidURL(e, request=request) self.cert_verify(conn, request.url, verify, cert) url = self.request_url(request, proxies) self.add_headers( request, stream=stream, timeout=timeout, verify=verify, cert=cert, proxies=proxies, ) chunked = not (request.body is None or \"Content-Length\" in request.headers) if isinstance(timeout, tuple): try: connect, read = timeout resolved_timeout = TimeoutSauce(connect=connect, read=read) except ValueError: raise ValueError( f\"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, \" f\"or a single float to set both timeouts to the same value.\" ) elif isinstance(timeout, TimeoutSauce): resolved_timeout = timeout else: resolved_timeout = TimeoutSauce(connect=timeout, read=timeout) try: > resp = conn.urlopen( method=request.method, url=url, body=request.body, # type: ignore[arg-type] # urllib3 stubs don't accept Iterable[bytes | str] headers=request.headers, # type: ignore[arg-type] # urllib3#3072 redirect=False, assert_same_host=False, preload_content=False, decode_content=False, retries=self.max_retries, timeout=resolved_timeout, chunked=chunked, ) cert = None chunked = False conn = <urllib3.connectionpool.HTTPSConnectionPool object at 0x7fdbe8249880> proxies = OrderedDict() request = <PreparedRequest [GET]> resolved_timeout = Timeout(connect=None, read=None, total=None) self = <requests.adapters.HTTPAdapter object at 0x7fdbe0bd0b30> stream = True timeout = None url = '/dragonflydb/dragonfly/releases/download/v1.19.2/dragonfly-x86_64.tar.gz' verify = True /usr/local/lib/python3.12/dist-packages/requests/adapters.py:696: _ _ _ _ _ _ _ _ _ _ ```",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8063",
          "createdAt": "2026-08-12T21:51:47Z",
          "updatedAt": "2026-08-12T21:51:59Z",
          "timestamp": "2026-08-12T21:51:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "failing-test",
            "iouring"
          ],
          "author": "dragonclaw-dragonflydb",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:a8dc5206623262254b56",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8039",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8039",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "feat: add time limitation for replication backlog",
          "text": "fixes: #7994 Summary: This PR replaces the fixed per-shard replication backlog entry count with time- and byte-based retention. Changes: - Adds a five-second default age target via --shard_repl_backlog_time_ms. - Adds --shard_repl_backlog_max_bytes, defaulting to 0.5% of maxmemory. - Deprecates --shard_repl_backlog_len and updates stale-partial-sync guidance. - Expands C++ coverage for byte limits, capacity growth, oversized records, and time expiry. - Updates replication-resilience tests to configure byte-based backlog behavior. Technical Notes: The byte allowance is divided among shards, while time eviction is evaluated on append and grouped into one-second buckets.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8039",
          "createdAt": "2026-08-10T11:55:56Z",
          "updatedAt": "2026-08-12T15:44:47Z",
          "timestamp": "2026-08-12T15:44:47Z",
          "metrics": {
            "reactions": 0,
            "comments": 10
          },
          "labels": [],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:9a4e866189cb2bc38d83",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8026",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8026",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(tiering): preserve offloaded hashes during serialization",
          "text": "Preserve the logical type of offloaded values when serializing snapshots, full-sync streams, and DUMP payloads. Previously, offloaded listpack hashes were read as strings. Debug builds aborted on a DCHECK, while release builds serialized the raw listpack bytes as a STRING. Root cause: `SerializerBase` routed every external value through `ReadTieredString` and stored the delayed result as a string. This assumed that all top-level tiered values were strings, which stopped being true after introducing experimental hash offloading. `DumpToString`, used by DUMP and cross-shard RENAME/COPY, made the same assumption. It could also request a `StringDecoder` while snapshot serialization requested a `ListpackMapDecoder` for the same coalesced disk read.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8026",
          "createdAt": "2026-08-06T19:44:34Z",
          "updatedAt": "2026-08-12T13:48:59Z",
          "timestamp": "2026-08-12T13:48:59Z",
          "metrics": {
            "reactions": 0,
            "comments": 5
          },
          "labels": [],
          "author": "vyavdoshenko",
          "state": "closed",
          "assignees": [
            "vyavdoshenko"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:05eea9a6d2b372b5aa2c",
        "signalId": "github:dragonflydb/dragonfly:pull_request:6991",
        "event": "changed",
        "observedAt": "2026-08-13T13:48:00.446149Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:6991",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "Use pcre2 regex & enable auto async",
          "text": "Make auto async use pcre2 for regexes. PCRE2 uses limited stack space, whereas std::regex is not suited for the small fiber stacks that we have",
          "url": "https://github.com/dragonflydb/dragonfly/pull/6991",
          "createdAt": "2026-03-26T17:18:00Z",
          "updatedAt": "2026-08-12T13:48:36Z",
          "timestamp": "2026-08-12T13:48:36Z",
          "metrics": {
            "reactions": 0,
            "comments": 3
          },
          "labels": [],
          "author": "dranikpg",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:76b455386e56daab9e5f",
        "signalId": "github:dragonflydb/dragonfly:pull_request:7984",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:7984",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "feat(geo): add GEOSEARCHSTORE command",
          "text": "<!-- **Commits Must Be Signed and Your PR title must conform to the conventional commit spec** * See: https://github.com/dragonflydb/dragonfly/blob/main/CONTRIBUTING.md * Please follow the section on `pre-commit hooks`, a linter will validate before you push Example PR Title: <type>(<scope>)!: <description> * `type` = bug, chore, feat, fix, docs, build, style, refactor, perf, test * `!` = OPTIONAL: signals a breaking change * `scope` = Optional when `type` is \"chore\" or \"docs\" * `description` = short description of the change Examples: * chore(examples): Clarify `docker` usage #120 * docs(readme): Fix Example Links #121 * feat(ingest)!: Add new ingest #122 * fix(ingest): Refactor for loop to list comprehension #123 --> Adds GEOSEARCHSTORE (dest + src keys, same search options as GEOSEARCH, optional STOREDIST). Most of the work was already in GeoSearchStoreGeneric() via GEORADIUS ... STORE — this wires up the dedicated command and parser. Closes #3883. Also fixed a couple of Redis mismatches in the shared store/search path: - STORE when the source key is missing now returns 0 and clears the dest key (was an empty array) - COUNT without ASC/DESC now defaults to ASC, same as Redis Tests in geo_family_test.cc. Ran geo_family_test locally + manual redis-cli checks.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/7984",
          "createdAt": "2026-08-03T03:51:25Z",
          "updatedAt": "2026-08-13T16:05:14Z",
          "timestamp": "2026-08-13T16:05:14Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [],
          "author": "rounaknandanwar",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:caa9f5d7ff48632e3ccb",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8078",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8078",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "chore: remove dead test_replication_onmove_flow",
          "text": "no point in time was removed in https://github.com/dragonflydb/dragonfly/pull/7068/changes",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8078",
          "createdAt": "2026-08-13T14:53:44Z",
          "updatedAt": "2026-08-13T16:04:25Z",
          "timestamp": "2026-08-13T16:04:25Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "kostasrim",
          "state": "closed",
          "assignees": [
            "kostasrim"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:4b5a0cf5e9829c8aea63",
        "signalId": "github:dragonflydb/dragonfly:issue:7953",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:7953",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "FT.CREATE without SCHEMA returns OK but creates unusable listed index",
          "text": "## Summary `FT.CREATE` accepts an index definition that omits the required `SCHEMA` keyword. It returns `OK` and exposes the index in `FT._LIST`, but the resulting index is unusable: `FT.INFO` reports that the index does not exist. The command should be rejected during parsing/validation rather than creating a listed-but-non-operational index. ## Reproduction Started a local single-node Dragonfly instance: ```sh mkdir ~/local-data-dragonfly podman --connection dragon run -d --pull=always -p 6379:6379 \\ -v ~/local-data-dragonfly/:/data:Z --ulimit memlock=-1 \\ docker.dragonflydb.io/dragonflydb/dragonfly \\ --lua_allow_undeclared_auto_correct=true --nodf_snapshot_format & ``` Then: ```redis FT.CREATE idx_ts_test ON HASH PREFIX 1 ot:foo name TEXT timestamp NUMERIC SORTABLE # OK FT._LIST # includes \"idx_ts_test\" FT.INFO idx_ts_test # ERR Index with name 'idx_ts_test' not found ``` This is on a local single-node instance, so cluster request routing is not involved. ## Expected behavior The malformed `FT.CREATE` invocation should fail immediately because `SCHEMA` is missing, e.g. with a syntax error. A valid invocation is: ```redis FT.CREATE idx_ts_test ON HASH PREFIX 1 ot:foo SCHEMA name TEXT timestamp NUMERIC SORTABLE ```",
          "url": "https://github.com/dragonflydb/dragonfly/issues/7953",
          "createdAt": "2026-07-28T18:14:00Z",
          "updatedAt": "2026-08-13T15:59:09Z",
          "timestamp": "2026-08-13T15:59:09Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [],
          "author": "dragonclaw-dragonflydb",
          "state": "closed",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:2fc0acbdfab52094d069",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8054",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "text",
          "updatedAt",
          "metrics",
          "state"
        ],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8054",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(search): reject FT.CREATE missing the SCHEMA keyword",
          "text": "## Summary - `FT.CREATE` accepted an index definition missing the required `SCHEMA` keyword, returning `OK` and listing the index in `FT._LIST`, but the index was unusable (`FT.INFO` reported it as not found). - Root cause: without `SCHEMA`, tokens that look like field definitions fall through the \"unsupported parameters are ignored\" branch in `CreateDocIndex` and get silently skipped one by one. - Fix: track whether any token was skipped this way; only reject the command when that happened *and* no fields were parsed. This keeps `FT.CREATE idx ON HASH` (no trailing tokens at all, populated later via `FT.ALTER`) legal, matching existing behavior covered by the `AlterIndex` test. ## Behavior notes - `FT.CREATE <existing-index> ON HASH` (with no `SCHEMA`) now returns the missing-SCHEMA error instead of \"Index already exists\" — the command is malformed regardless of whether the index name already exists, so the syntax error takes precedence. - During a mixed-version rolling upgrade window, a schema-less `FT.CREATE` journaled by an old master is rejected when replayed on an already-upgraded replica. That index was already unusable and non-durable under the old behavior (didn't survive serialization, diverged FT.INFO from FT._LIST), so this doesn't introduce new master/replica divergence — the replica simply converges at the next full sync. ## Test plan - [x] Added `SearchFamilyTest.CreateWithoutSchemaKeywordIsError`, covering both the bug repro (rejected with a syntax error, no phantom index left behind) and the still-legal empty-schema-for-later-ALTER case. - [x] Full `search_family_test` suite passes (330/330), confirming no regression in existing `FT.CREATE` tests (including `CreateDropListIndex` and `AlterIndex`, which rely on schema-less creation succeeding). - [x] `pre-commit` (clang-format) passes on changed files. Fixes #7953",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8054",
          "createdAt": "2026-08-11T18:32:41Z",
          "updatedAt": "2026-08-13T15:59:08Z",
          "timestamp": "2026-08-13T15:59:08Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [],
          "author": "Shikha-code36",
          "state": "closed",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:4fcf7b3b3e488a41170e",
        "signalId": "github:dragonflydb/dragonfly:issue:8028",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:issue:8028",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "issue",
          "title": "Disabled python tests in tests/dragonfly",
          "text": "## Unconditional skips (broken / flaky / WIP) - [ ] `tests/dragonfly/shutdown_test.py:15` — class `TestDflyAutoLoadSnapshot` (covers `test_gracefull_shutdown`) — \"Currently we can not guarantee that on shutdown if command is executed and value is written we response before breaking the connection\" @vyavdoshenko -> check if the invariants/what's written is true and if we can change somehow the test semantics to assert that we get a reply before we shutdown --- - [ ] `tests/dragonfly/memory_test.py:327` — `test_throttle_on_commands_squashing_replies_bytes` — \"Disabling test until improvements in squashing.\" @mkaruza --- - [ ] `tests/dragonfly/snapshot_test.py:222` — `test_cron_snapshot_failed_saving` — \"Fails and also causes all TLS tests to fail\" @BorysTheDev --- - [ ] `tests/dragonfly/snapshot_test.py:961` — `test_tiered_entries_throttle` — no reason given (also `@large`/`@opt_only`) @dranikpg --- - [ ] `tests/dragonfly/acl_family_test.py:271` — `test_acl_del_user_while_running_lua_script` — \"Flaky on CI, needs investigation\" @kostasrim --- - [ ] `tests/dragonfly/acl_family_test.py:297` — `test_acl_with_long_running_script` — \"Check TODO in the body below\" @kostasrim --- - [ ] `tests/dragonfly/connection_test.py:2483` — `test_pipeline_cache_size` — \"Flaky\" @abhijat --- - [ ] `tests/dragonfly/replication_specific_test.py:508` — `test_big_huge_streaming_restart` — \"fails, investigating\" @kostasrim --- - [x] `tests/dragonfly/replication_specific_test.py:730` — `test_replication_onmove_flow` — \"Fails constantly on CI\" Dead test https://github.com/dragonflydb/dragonfly/pull/7096 and https://github.com/dragonflydb/dragonfly/pull/7068/changes @kostasrim will remove it in a sec",
          "url": "https://github.com/dragonflydb/dragonfly/issues/8028",
          "createdAt": "2026-08-07T09:05:39Z",
          "updatedAt": "2026-08-13T15:51:22Z",
          "timestamp": "2026-08-13T15:51:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 0
          },
          "labels": [
            "bug",
            "failing-test"
          ],
          "author": "kostasrim",
          "state": "open",
          "assignees": [
            "abhijat",
            "mkaruza",
            "kostasrim",
            "dranikpg",
            "vyavdoshenko"
          ],
          "change": "new"
        }
      },
      {
        "id": "event:689a671572ee5dd75d4c",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8002",
        "event": "discovered",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8002",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "docs: add design doc for replication fanout (NOT FOR MERGE)",
          "text": "This is a design doc for #7993.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8002",
          "createdAt": "2026-08-04T12:45:27Z",
          "updatedAt": "2026-08-13T15:48:22Z",
          "timestamp": "2026-08-13T15:48:22Z",
          "metrics": {
            "reactions": 0,
            "comments": 1
          },
          "labels": [],
          "author": "BorysTheDev",
          "state": "open",
          "assignees": [],
          "change": "new"
        }
      },
      {
        "id": "event:d69921d1fd1adb21ed7c",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8053",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt",
          "metrics"
        ],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8053",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(generic): avoid data loss on RENAME when destination write fails",
          "text": "## Summary - `Renamer::FinalizeRename()` ran `DelSrc` and `DeserializeDest` in the same transaction hop with no ordering guarantee between shards. If `DeserializeDest` failed (e.g. OOM), the source key could already be deleted, silently losing its data. - Now deserializes into the destination first, and only deletes the source once that succeeds. `COPY` is unaffected since it never deletes the source. ## Test plan - [x] Added `GenericFamilyTest.RenameOOMPreservesSource`: drives real per-shard OOM via repeated `RESTORE`, then asserts a failed `RENAME` leaves the source key intact and does not create the destination. - [x] Full `generic_family_test` suite passes (84/84). - [x] `pre-commit` (clang-format) passes on changed files. Fixes #5296",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8053",
          "createdAt": "2026-08-11T17:02:15Z",
          "updatedAt": "2026-08-13T13:59:40Z",
          "timestamp": "2026-08-13T13:59:40Z",
          "metrics": {
            "reactions": 0,
            "comments": 8
          },
          "labels": [],
          "author": "Shikha-code36",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:0c40eae3420f1415cb13",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8075",
        "event": "changed",
        "observedAt": "2026-08-13T16:19:22.035158Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8075",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(pubsub): preserve V2 message ordering",
          "text": "This PR fixes a known issue that was already fixed for V1 at 5ef35182. Defer V2 pipeline execution when queued control work predates the next parsed command. Let IoLoopV2 drain that work so UNSUBSCRIBE cannot discard earlier Pub/Sub messages. Tests: Extend the unsubscribe regression across V1 and V2 with standalone and deferred-GET commands. Bound the starvation publisher wave to assert fairness without treating ordered backlog as starvation.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8075",
          "createdAt": "2026-08-13T10:10:27Z",
          "updatedAt": "2026-08-13T13:50:16Z",
          "timestamp": "2026-08-13T13:50:16Z",
          "metrics": {
            "reactions": 0,
            "comments": 7
          },
          "labels": [],
          "author": "glevkovich",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      },
      {
        "id": "event:1780f97d8c5d88a2a40c",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8065",
        "event": "changed",
        "observedAt": "2026-08-13T17:43:20.785491Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8065",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(server): emit expired keyspace events for already-past expirations",
          "text": "Deleting a key because its newly set expiration is already in the past emitted no `__keyevent@<db>__:expired` notification, so a subscriber waiting for the event was blocked forever (found via the upstream Valkey pubsub TCL suite: \"expired event (Expiration time is already expired)\" hangs). This covers the `DbSlice::UpdateExpire` funnel: EXPIRE/PEXPIRE/(P)EXPIREAT, GETEX, and memcached GAT/GATS.",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8065",
          "createdAt": "2026-08-13T06:51:10Z",
          "updatedAt": "2026-08-13T16:35:07Z",
          "timestamp": "2026-08-13T16:35:07Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "vyavdoshenko",
          "state": "open",
          "assignees": [
            "vyavdoshenko"
          ],
          "change": "updated"
        }
      },
      {
        "id": "event:ebe7ae1502bd820bff25",
        "signalId": "github:dragonflydb/dragonfly:pull_request:8077",
        "event": "changed",
        "observedAt": "2026-08-13T18:01:55.420671Z",
        "changedFields": [
          "updatedAt"
        ],
        "signal": {
          "id": "github:dragonflydb/dragonfly:pull_request:8077",
          "source": "github",
          "group": "data-infrastructure",
          "project": "dragonflydb/dragonfly",
          "kind": "pull_request",
          "title": "fix(tiering): Use TraverseBySegmentOrder and export more metrics/configs",
          "text": "The changes that I merged into 1.40 branch Only functional change: 1. Fixes small bins defrag by using bounded TraverseBySegmentOrder to avoid blowing up iteration costs when the table is largely empty Metrics & config: 1. Exports missing tiering fields from info to metrics 2. Adds new configurable values to allow limiting CPU limit consumption of offloading / background scanning",
          "url": "https://github.com/dragonflydb/dragonfly/pull/8077",
          "createdAt": "2026-08-13T11:05:19Z",
          "updatedAt": "2026-08-13T17:54:15Z",
          "timestamp": "2026-08-13T17:54:15Z",
          "metrics": {
            "reactions": 0,
            "comments": 4
          },
          "labels": [],
          "author": "dranikpg",
          "state": "open",
          "assignees": [],
          "change": "updated"
        }
      }
    ]
  }
}
