Skip to content

Striping seems unnecessary at this point #48

Description

@BurningWitness

As noted in the 0.5 benchmark comment, swapping all MVars for TVars in #38 did make the single-stripe solution significantly faster. The benchmark itself is however quite weird, as it assumes incredibly tiny jobs (fib 10), roughly in the range of dozens to hundreds of nanoseconds.

I wrote a more comprehensive benchmark over a fork (link).

For tiny jobs the difference indeed looks substantive:

Benchmark
$ cabal bench resource-pool-bench:bench --benchmark-options='--items=10000000 --item-size=1ns:5ns --stripes=4 +RTS -N4'
[..]
Arguments: 
  4 workers over 4 capabilities
  10000000 items, each between 1ns and 5ns
  4 stripes
[..]
Results: 
  Expected time per item: 3.00ns ± 0.00ns
  Extra time spent per item: 0.64us ± 46.26ns

$ cabal bench resource-pool-bench:bench --benchmark-options='--items=10000000 --item-size=1ns:5ns --stripes=1 +RTS -N4'
[..]
Arguments: 
  4 workers over 4 capabilities
  10000000 items, each between 1ns and 5ns
  1 stripes
[..]
Results: 
  Expected time per item: 3.00ns ± 0.00ns
  Extra time spent per item: 1.33us ± 17.87ns

Bump runtimes a thousandfold and now the difference is wholly inconsequential:

Benchmark
$ cabal bench resource-pool-bench:bench --benchmark-options='--items=10000 -
-item-size=10us:100us --stripes=4 +RTS -N4'
[..]
Arguments: 
  4 workers over 4 capabilities
  10000 items, each between 10us and 100us
  4 stripes
[..]
Results: 
  Expected time per item: 54.98us ± 94.05ns
  Extra time spent per item: 0.65us ± 0.47us

Benchmark bench: FINISH

$ cabal bench resource-pool-bench:bench --benchmark-options='--items=10000 --item-size=10us:100us --stripes=1 +RTS -N4'
[..]
Arguments: 
  4 workers over 4 capabilities
  10000 items, each between 10us and 100us
  1 stripes
[..]
Results: 
  Expected time per item: 54.98us ± 94.05ns
  Extra time spent per item: 0.73us ± 87.03ns

And this is the busy-waiting case; if threads get to sleep instead the margins are way narrower (in my case both above are at 0.95ms ± 5.91us and 0.96ms ± 11.94us respectively).


On top of that, the very existence of striping complicates library design substantially:

  • Running the collector thread every second is egregiously inefficient;

  • "Maximum number of resources" is misleading. There's no guarantee incoming requests will be spread uniformly among all stripes.


I think the correct approach would be to remove striping and then optimize the rest of the library. Anyone who still wants striping should be able to create one pool per capability themselves, avoiding the need for whatever magic currently allows this to work across all possible configurations.

Activity

  1. arybczak commented on Aug 3, 2026

    @arybczak
    Contributor

    Running the collector thread every second is egregiously inefficient;

    Yeah, but this has nothing to do with stripes. I made it better in #49.

    I'm not particularly opposed to removal of striping (I think it's more or less a historical accident), but I don't see how its removal would allow you to "optimize the rest of the library", considering that most functions already operate on a single stripe.

    In the default case (a single stripe) you'd save a single array lookup with it, not particularly exciting if you ask me ;)

  2. BurningWitness commented on Aug 4, 2026

    @BurningWitness
    Author

    Naively (by which I mean I haven't made a prototype), I expect a clean separation between the three actors in this library, which are:

    1. Incoming request(s);
    2. Collector thread.
    3. Managing requests that exceed max-connections to reduce TVar congestion. I feel like this should be a thread to be able to time entries out, but I'm not sure.

    None of these require access to the entire pool state (Stripe) at the same time, so it can be split into:

    1. Current connection count (TVar Int). Must be incremented to proceed, decremented as a last action.
    2. Queue of pending requests (TQueue ?). Definitely doesn't need to be parametrized by resource type.
    3. Cached entries (TVar [Entry a]). Cons/uncons to add/retrieve; cleanup by cutting the list lazily, so memory-wise it's swapping to a thunk.

    There's also a separate discussion on whether forever + exceptions is a good way to notify threads, I'd argue for messaging via TQueues instead.

  3. arybczak commented on Aug 4, 2026

    @arybczak
    Contributor

    And what problems does this solve when compared to the current design?

    whether forever + exceptions is a good way to notify threads, I'd argue for messaging via TQueues instead.

    Why?

  4. BurningWitness commented on Aug 4, 2026

    @BurningWitness
    Author

    And what problems does this solve when compared to the current design?

    None in particular, thank you for your attention.

  5. arybczak commented on Aug 5, 2026

    @arybczak
    Contributor

    Well, at the very least your inquiry about the collector prompted me to have a closer look at it which resulted in its significant improvement, so thanks for doing the work 🙂

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions