Skip to content

perf(api): batch indexing status segment counts - #42191

Merged
fatelei merged 1 commit into
langgenius:mainfrom
CoralGarden52:fix/indexing-status-batch-counts
Sep 11, 2026
Merged

fatelei merged 1 commit into
langgenius:mainfrom
CoralGarden52:fix/indexing-status-batch-counts

Conversation

@CoralGarden52

Copy link
Copy Markdown
Contributor

Summary

  • Batch-load completed_segments and total_segments for indexing-status responses with one grouped query keyed by document ownership.
  • Apply the bounded query path to the console dataset status endpoint, console batch status endpoint, and Service API batch status endpoint.
  • Preserve the response shape, RE_SEGMENT exclusion, completed-at semantics, and tenant/dataset ownership filtering.

Fixes #42188

Motivation and scope

The web UI polls batch indexing status every 2.5 seconds. The previous implementation executed two COUNT(DocumentSegment.id) queries inside the response loop for every document. This made the segment-count query count grow as 2N for a batch of N documents.

This change is limited to the indexing-status endpoints. It does not change the document-list endpoint addressed by #40123 and #40125, and it does not depend on that still-open PR.

Verification evidence

Environment: Linux, CPython 3.12.14, SQLAlchemy SQLite test engine, using real controller methods and database rows. No Docker or external service was required for this query-count verification. The measurement verifies query shape and returned values; it is not a production latency benchmark.

The same harness populated 1, 5, and 20 documents with three segments per document and counted emitted SELECT statements:

Endpoint / measured phase Before, 20 documents After, 20 documents
Console dataset indexing-status 41 total (1 + 2N) 2 total (document load + grouped count)
Console batch indexing-status segment-count phase 40 (2N) 1 grouped count
Service API batch indexing-status segment-count phase 40 (2N) 1 grouped count

After-fix harness output was:

documents=1 dataset_indexing_status_selects=2 console_batch_indexing_status_selects=1 service_batch_indexing_status_selects=2
documents=5 dataset_indexing_status_selects=2 console_batch_indexing_status_selects=1 service_batch_indexing_status_selects=2
documents=20 dataset_indexing_status_selects=2 console_batch_indexing_status_selects=1 service_batch_indexing_status_selects=2

The Service API total includes its unchanged dataset lookup; the batch document lookup was isolated in the harness so the status phase could be measured separately. The console batch document lookup is likewise unchanged.

The regression tests also verify that each document keeps its completed and total counts, and that RE_SEGMENT rows and rows belonging to another tenant are excluded.

Screenshots

Not applicable (backend-only change).

Tests

  • DEBUG=false uv run --project api python -m pytest -o addopts='' api/tests/unit_tests/controllers/console/datasets/test_datasets.py api/tests/unit_tests/controllers/console/datasets/test_datasets_document.py api/tests/unit_tests/controllers/service_api/dataset/test_document.py -q — 255 passed.
  • make check — passed.
  • uv run --project api --dev ruff format --check ./api — 3523 files already formatted.
  • make api-contract-lint — 528 valid, 0 mismatch, 0 refactorable.
  • uv run --directory api --dev lint-imports — 31 contracts kept, 0 broken.
  • make type-check — no issues found in 1796 source files.

Checklist

  • This change requires a documentation update, included: Dify Document
  • I understand that this PR may be closed in case there was no previous discussion or issues. (This doesn't apply to typos!)
  • I've added a test for each change that was introduced, and I tried as much as possible to make a single atomic change.
  • I've updated the documentation accordingly.
  • I ran the backend Ruff, response-contract, import, unit-test, and type-check validations listed above.

@github-actions

Copy link
Copy Markdown
Contributor

Pyrefly Diff

base → PR
--- /tmp/pyrefly_base.txt	2026-09-11 07:17:05.996001224 +0000
+++ /tmp/pyrefly_pr.txt	2026-09-11 07:16:57.851813687 +0000
@@ -3421,15 +3421,15 @@
 ERROR Argument `scoped_session[flask_sqlalchemy.session.Session]` is not assignable to parameter `session` with type `sqlalchemy.orm.session.Session` in function `_add_binding` [bad-argument-type]
   --> tests/unit_tests/controllers/console/datasets/test_data_source.py:85:28
 ERROR Missing argument `name` in function `controllers.console.datasets.datasets.DatasetCreatePayload.__init__` [missing-argument]
-   --> tests/unit_tests/controllers/console/datasets/test_datasets.py:581:49
+   --> tests/unit_tests/controllers/console/datasets/test_datasets.py:591:49
 ERROR Object of class `NoneType` has no attribute `token` [missing-attribute]
-    --> tests/unit_tests/controllers/console/datasets/test_datasets.py:1448:16
+    --> tests/unit_tests/controllers/console/datasets/test_datasets.py:1517:16
 ERROR Argument `Literal['custom']` is not assignable to parameter `mode` with type `ProcessRuleMode | SQLCoreOperations[ProcessRuleMode]` in function `models.dataset.DatasetProcessRule.__init__` [bad-argument-type]
-   --> tests/unit_tests/controllers/console/datasets/test_datasets_document.py:245:18
+   --> tests/unit_tests/controllers/console/datasets/test_datasets_document.py:246:18
 ERROR Argument `Literal['custom']` is not assignable to parameter `mode` with type `ProcessRuleMode | SQLCoreOperations[ProcessRuleMode]` in function `models.dataset.DatasetProcessRule.__init__` [bad-argument-type]
-   --> tests/unit_tests/controllers/console/datasets/test_datasets_document.py:280:67
+   --> tests/unit_tests/controllers/console/datasets/test_datasets_document.py:281:67
 ERROR Argument `Literal['custom']` is not assignable to parameter `mode` with type `ProcessRuleMode | SQLCoreOperations[ProcessRuleMode]` in function `models.dataset.DatasetProcessRule.__init__` [bad-argument-type]
-    --> tests/unit_tests/controllers/console/datasets/test_datasets_document.py:1505:67
+    --> tests/unit_tests/controllers/console/datasets/test_datasets_document.py:1543:67
 ERROR Argument `Literal['opendal']` is not assignable to parameter `storage_type` with type `StorageType` in function `models.model.UploadFile.__init__` [bad-argument-type]
    --> tests/unit_tests/controllers/console/datasets/test_datasets_segments.py:137:22
 ERROR Argument `Literal['account']` is not assignable to parameter `created_by_role` with type `CreatorUserRole` in function `models.model.UploadFile.__init__` [bad-argument-type]

@github-actions

Copy link
Copy Markdown
Contributor

Pyrefly Type Coverage

Metric Base PR Delta
Type coverage 64.06% 64.06% -0.00%
Strict coverage 63.67% 63.67% -0.00%
Typed symbols 45,933 45,943 +10
Untyped symbols 25,924 25,934 +10
Modules 3398 3398 0

@fatelei
fatelei added this pull request to the merge queue Sep 11, 2026
Merged via the queue into langgenius:main with commit 5b58f51 Sep 11, 2026
45 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: indexing-status executes two segment-count queries per document

2 participants