fix(deps): pin hyper to flush-before-shutdown fix for large-GET unexpected EOF (#4797)

* fix(deps): pin hyper to flush-before-shutdown fix for large-GET unexpected EOF

hyper <= 1.10.1 can call poll_shutdown() on an HTTP/1 socket while response
bytes are still buffered (a prior poll_flush() returned Poll::Pending and the
result was discarded), so a backpressured/slow-reading peer receives a graceful
FIN before the full Content-Length body is flushed. Standard S3 clients
(minio-go/warp) report this as sporadic `unexpected EOF` on large-object GET
under load (rustfs/backlog#1232). This is the transport-layer bug from
Cloudflare's "hyper-bug" writeup — distinct from the EC-reconstruct-desync EOF
already fixed via lockstep decode; here the app layer delivers the full body and
truncation is purely hyper->socket.

Fixed upstream in hyperium/hyper#4018 (commit 72046cc7, "fix(http1): flush
buffered data before shutdown"), which is not in any crates.io release yet as of
hyper 1.10.1. Pin hyper via [patch.crates-io] to git rev ccc1e850 (a descendant
of the fix that still declares version 1.10.1).

Use [patch.crates-io], not a git+rev on the workspace `hyper` dependency: the
server connection is driven by the transitive hyper-util (conn::auto /
GracefulShutdown), and a git source would not unify with the crates.io hyper
hyper-util resolves, leaving two hyper copies with the server path still on the
buggy one. The patch rewrites the crates.io source globally, so the lockfile
holds a single git-sourced hyper.

Add rustfs/tests/hyper_h1_shutdown_flush_regression.rs, a deterministic guard
(mirrors hyper's own h1_shutdown_while_buffered) that fails if the pin is dropped
or hyper is downgraded below the fix. Its load-bearing assertion is a
synchronously-set flag, so a slow runner can only under-detect, never false-red.

Closes rustfs/backlog#1232.

* build(deny): allow the hyperium/hyper git source for the flush-before-shutdown pin

cargo-deny's [sources] check denies unknown git sources. The hyper
[patch.crates-io] pin added for rustfs/backlog#1232 uses a github.com/hyperium
git source, so add it to allow-git with an owner/review note and a removal
condition (drop once a released hyper > 1.10.1 carries commit 72046cc7).
This commit is contained in:
Zhengchao An
2026-07-13 23:36:38 +08:00
committed by GitHub
parent 0ac7f0d0cf
commit 0f83a27f6a
4 changed files with 181 additions and 2 deletions
Generated
+1 -2
View File
@@ -5050,8 +5050,7 @@ dependencies = [
[[package]]
name = "hyper"
version = "1.10.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "55281c53a1894c864990125767da440a4e630446785086f52523b20033b74498"
source = "git+https://github.com/hyperium/hyper.git?rev=ccc1e850dc0cda3e71b0acd11f60ca3d48d09034#ccc1e850dc0cda3e71b0acd11f60ca3d48d09034"
dependencies = [
"atomic-waker",
"bytes",
+18
View File
@@ -362,3 +362,21 @@ codegen-units = 1
[profile.profiling]
inherits = "release"
debug = true
# Pin hyper to a revision that carries the HTTP/1 "flush buffered data before
# shutdown" fix (hyperium/hyper#4018, commit 72046cc7). This lands as a
# `[patch.crates-io]` entry — not on the `hyper` workspace dependency — so that
# every consumer in the tree, including the transitive `hyper-util` server path
# (`conn::auto` / `GracefulShutdown`) that actually drives our connections,
# resolves to the fixed hyper rather than the buggy crates.io copy.
#
# hyper <= 1.10.1 can call `poll_shutdown()` on the socket while response bytes
# are still buffered (a prior `poll_flush()` returned `Poll::Pending` and the
# result was discarded). A backpressured / slow-reading peer then receives a
# graceful FIN before the full Content-Length body is flushed, which standard S3
# clients (minio-go / warp) report as `unexpected EOF` on large-object GET under
# load. The fix is not in any crates.io release yet as of hyper 1.10.1; drop
# this patch once a released version (> 1.10.1) contains commit 72046cc7.
# See rustfs/backlog#1232.
[patch.crates-io]
hyper = { git = "https://github.com/hyperium/hyper.git", rev = "ccc1e850dc0cda3e71b0acd11f60ca3d48d09034" }
+6
View File
@@ -40,6 +40,12 @@ allow-git = [
# Pinned to a specific commit in workspace Cargo.toml.
# "https://github.com/rustfs/s3s",
"https://github.com/apache/datafusion.git",
# hyper pinned (via [patch.crates-io]) to a rev carrying the HTTP/1
# flush-before-shutdown fix (hyperium/hyper#4018, commit 72046cc7) that is
# not yet in a crates.io release. Fixes sporadic large-GET `unexpected EOF`
# (rustfs/backlog#1232). Remove once a released hyper > 1.10.1 contains it.
# owner: rustfs-maintainers review: 2026-07
"https://github.com/hyperium/hyper.git",
]
[bans]
@@ -0,0 +1,156 @@
// Regression guard for rustfs/backlog#1232 — sporadic large-object GET
// `unexpected EOF` under load.
//
// Root cause was in hyper, not RustFS: hyper <= 1.10.1 could call
// `poll_shutdown()` on the HTTP/1 socket while response bytes were still
// buffered (a prior `poll_flush()` returned `Poll::Pending` and the result was
// discarded). A backpressured / slow-reading peer then received a graceful FIN
// before the full Content-Length body was flushed, which S3 clients
// (minio-go / warp) report as `unexpected EOF`. Fixed upstream by
// hyperium/hyper#4018 (commit 72046cc7, "fix(http1): flush buffered data
// before shutdown"), which RustFS pulls in via the `[patch.crates-io]` hyper
// pin in the workspace Cargo.toml.
//
// This test reproduces the exact trigger deterministically and fails if the
// pin is ever dropped / hyper is downgraded below the fix — the one realistic
// regression vector (an accidental `cargo update` or a lost patch). It mirrors
// hyper's own upstream regression test `tests/h1_shutdown_while_buffered.rs`.
//
// The load-bearing assertion is on `shutdown_called_with_buffered`, a flag set
// synchronously inside `poll_shutdown`. It is timing-independent: a slow CI
// runner can only fail to reach the shutdown at all (a false pass), never flip
// the flag spuriously (a false red). Drop this test once the hyper pin is
// replaced by a released hyper > 1.10.1 that contains commit 72046cc7.
use std::pin::Pin;
use std::sync::{Arc, Mutex};
use std::task::{Context, Poll};
use std::time::Duration;
use bytes::Bytes;
use http::{Request, Response};
use http_body_util::Full;
use hyper::body::Incoming;
use hyper::service::service_fn;
use hyper_util::rt::TokioIo;
use tokio::io::{AsyncRead, AsyncWrite, ReadBuf};
use tokio::net::{TcpListener, TcpStream};
use tokio::time::{sleep, timeout};
const RESP_BODY_LEN: usize = 500_000;
// Accept one partial write, then pend forever — the boundary that fills hyper's
// write buffer and used to trip the premature shutdown. Same value hyper uses.
const WRITE_CHUNK: usize = 212_992;
#[derive(Debug, Default)]
struct Stats {
bytes_written: usize,
total_attempted: usize,
shutdown_called_with_buffered: bool,
buffered_at_shutdown: usize,
}
/// A socket wrapper that accepts a single partial write and then pends on every
/// further write, simulating a backpressured peer whose receive window is full
/// exactly at the write-buffer boundary.
struct PendingStream {
inner: TcpStream,
write_count: usize,
stats: Arc<Mutex<Stats>>,
}
impl AsyncRead for PendingStream {
fn poll_read(mut self: Pin<&mut Self>, cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<std::io::Result<()>> {
Pin::new(&mut self.inner).poll_read(cx, buf)
}
}
impl AsyncWrite for PendingStream {
fn poll_write(mut self: Pin<&mut Self>, cx: &mut Context<'_>, buf: &[u8]) -> Poll<std::io::Result<usize>> {
self.write_count += 1;
self.stats.lock().unwrap().total_attempted += buf.len();
if self.write_count == 1 {
let partial = std::cmp::min(buf.len(), WRITE_CHUNK);
let result = Pin::new(&mut self.inner).poll_write(cx, &buf[..partial]);
if let Poll::Ready(Ok(n)) = result {
self.stats.lock().unwrap().bytes_written += n;
}
return result;
}
Poll::Pending
}
fn poll_flush(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
let buffered = {
let s = self.stats.lock().unwrap();
s.total_attempted - s.bytes_written
};
// Cannot flush while the peer refuses more data.
if buffered > 0 {
return Poll::Pending;
}
Pin::new(&mut self.inner).poll_flush(cx)
}
fn poll_shutdown(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
{
let mut s = self.stats.lock().unwrap();
let buffered = s.total_attempted - s.bytes_written;
if buffered > 0 {
s.shutdown_called_with_buffered = true;
s.buffered_at_shutdown = buffered;
}
}
Pin::new(&mut self.inner).poll_shutdown(cx)
}
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn hyper_h1_does_not_shutdown_with_buffered_response() {
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let addr = listener.local_addr().unwrap();
let stats = Arc::new(Mutex::new(Stats::default()));
let stats_srv = stats.clone();
let server = tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
let ps = PendingStream {
inner: stream,
write_count: 0,
stats: stats_srv,
};
let service = service_fn(|_req: Request<Incoming>| async move {
Ok::<_, hyper::Error>(Response::new(Full::new(Bytes::from(vec![b'X'; RESP_BODY_LEN]))))
});
hyper::server::conn::http1::Builder::new()
.serve_connection(TokioIo::new(ps), service)
.await
});
sleep(Duration::from_millis(50)).await;
// Client: send one complete request, then hold the connection open. It does
// not need to read — the backpressure is modelled by PendingStream.
let client = tokio::spawn(async move {
use tokio::io::AsyncWriteExt;
let mut s = TcpStream::connect(addr).await.unwrap();
s.write_all(b"GET / HTTP/1.1\r\nHost: localhost\r\n\r\n").await.unwrap();
s.flush().await.unwrap();
sleep(Duration::from_secs(2)).await;
});
// On buggy hyper the connection future resolves after the wrongful shutdown;
// on fixed hyper it parks (flush pends), so a timeout here is expected and
// harmless — the verdict is the flag below, not the future's completion.
let _ = timeout(Duration::from_millis(1500), server).await;
client.abort();
let s = stats.lock().unwrap();
assert!(
!s.shutdown_called_with_buffered,
"hyper shut the HTTP/1 socket down with {} response bytes still buffered — the backlog#1232 \
premature-FIN bug is back; the hyper flush-before-shutdown pin (hyperium/hyper#4018) was likely dropped",
s.buffered_at_shutdown
);
}