With local models you're limited by context/memory. If your system can't support parallel context, tasks end up taking far longer since it has to reprocess context each time it switches between tasks.
EDIT: For comparison, a task I would expect to take 5 minutes when worked on by a single agent ends up taking 45 minutes when I have 5 in parallel.
slot update_slots: id 0 | task 6389 | Checking checkpoint with [46119, 46119] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [45607, 45607] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [44756, 44756] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [44244, 44244] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [42590, 42590] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [42233, 42233] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [42126, 42126] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [41614, 41614] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [39843, 39843] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [39331, 39331] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [39059, 39059] against 11855...
slot update_slots: id 0 | task 6389 | Checking checkpoint with [38916, 38916] against 11855...
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 12287, pos_max = 12287, n_tokens = 12288, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 16383, pos_max = 16383, n_tokens = 16384, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 18099, pos_max = 18099, n_tokens = 18100, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 18611, pos_max = 18611, n_tokens = 18612, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 22897, pos_max = 22897, n_tokens = 22898, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 25805, pos_max = 25805, n_tokens = 25806, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 26317, pos_max = 26317, n_tokens = 26318, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 30008, pos_max = 30008, n_tokens = 30009, n_swa = 0, pos_next = 8192, size = 149.626 MiB)
slot update_slots: id 0 | task 6389 | erased invalidated context checkpoint (pos_min = 30520, pos_max = 30520, n_tokens
Feature hasn't been suggested before.
Describe the enhancement you want to request
With local models you're limited by context/memory. If your system can't support parallel context, tasks end up taking far longer since it has to reprocess context each time it switches between tasks.
If an agent starts 5 parallel subagents:
EDIT: For comparison, a task I would expect to take 5 minutes when worked on by a single agent ends up taking 45 minutes when I have 5 in parallel.
llama-serverlogs are filled with stuff like this: