With the current timer setup, Quinn destroys a tokio timer each time
the timer is disabled, and recreates it when a timeout is required again.
This has a certain attached cost, since timer creation requires a memory
allocation, and both creating the timer as well as resetting and polling
it incur some synchronization costs.
With this change, we keep one timer around and reuse it - which saves
some of these costs.
In addition to this, the change checks the connection first on whether
a timer is still required before polling the timeout. This can remove
some unnecessary wakeups.
This change adds some structs which track connection-level statistics.
These are rather helpful to determine how the implementation performs,
and which impact certain changes have.
The implementation is not complete. It's just a first step on getting
some more visibility into the transmitting side of the library. More
statistics - e.g. around congestion control, retransmission, timing, etc
could be added.
The stats are currently private, and I just debug printed them so far
where found useful. They could however be public if the API is fine.
Another possibility is to have the stats behind a feature flag if people
are concerned about extra memory usage.
The current version of Quinn tries to enqueue a MAX_STREAM_DATA
update after every read from the stream. Those could potentially be
tiny. Besides wasting network capacity with sending tiny packets, that
behavior also causes a wakeup on the connection task.
This change triggers MAX_STREAM_DATA frames only to be enqueued
if they are deemed significant enough. In the version here, this means
is bound to 12.5% of the overall window - but this could be changed.
Note that the general connection task wakeup after reads is still there.
A similar change would need to be performed for updating the
connection window, in order to determine whether the connection task
wakeup is still necessary.
I gathered some example metrics for the amount and increment of window
updates in the benchmark:
Before:
Num max stream announcements: 1663, Total window announced: 58500130, Avg window diff: 35177
After:
Num max stream announcements: 301, Total window announced: 58195055, Avg window diff: 193339
==> This sends 5x less window updates
It wasn't super clear what `more` means here. The purpose of this
parameter is to signal whether a `MAX_STREAM_DATA` frame needs to be
sent. This change tries to improve that.
Use vectors instead of boxed slices
The conversion from Vec<u8> into Box<[u8]> is not always for free.
It will call `Vec::into_boxed_slice`, which will call `shrink_to_fit`.
https://github.com/rust-lang/rust/blob/28f03ac4c08fc7ec62428d0b914e1510ce7ee2cb/library/alloc/src/vec.rs#L689
If the `Vec` isn't fully utilized before, this will cause a reallocation
and a copy of all data. This might currently happen with every
outgoing packet.
This change simply keeps things as `Vec<u8>`, which works just fine since
the IO layer can deal with it.
This leaves data flow control to be governed purely at the stream
level by default, removing a surprising footgun for users that need to
tune flow control and who aren't aware of all the knobs.