Files
Eliah Rusin 7643f4f541 rust: a seaweed-common crate for the address and TLS helpers both crates carry (#11358)
* rust: a seaweed-common crate for the address and TLS helpers both crates carry

seaweed-volume and seaweed-worker are separate cargo trees with separate
lockfiles and no root manifest, so anything both of them need has had to be
written twice. Two of those copies are a correctness risk rather than a typing
cost, and this crate is where they stop being copies.

address.rs is the HTTP<->gRPC port rule: `host:port` means gRPC on port+10000,
`host:port.grpcPort` names it outright. The two copies had already drifted —
the worker's bracketed IPv6 literals, the volume server's did not — so the rule
lives here once, returning a typed AddressError whose Display text is the volume
server's original wording, with join_host_port public beside it. A test asserts
two of those messages in full rather than by substring, because the wording is
the contract its callers hand to a Status or an io::Error; the other three end
in a std ParseIntError message, which is std's to reword. The enum is
#[non_exhaustive] so a future variant is not a breaking change for either
consumer. The tests are both crates' cases together, plus the IPv6,
already-bracketed and normalisation cases neither copy covered on its own.

tls.rs is install_default_crypto_provider. Both binaries link aws-lc-rs and ring
transitively, so rustls cannot auto-select and tonic's client TLS panics on
first use; each binary has to pin one and it has to be the same one, which is
exactly the kind of choice that should not exist twice. It is safe to share
because `cargo tree -i rustls` resolves a single rustls in each tree (0.23.37 in
seaweed-volume, 0.23.43 in seaweed-worker) and cargo unifies all
semver-compatible `rustls = "0.23"` requirements into one crate per binary, so
this crate writes the same process-wide static its consumer reads. rustls is
already in both graphs — directly in the volume server, through tonic's
tls-aws-lc in seaweed-worker-core — so the dependency adds no crate to either.

rust-version is 1.91.1, the lower of the two consumers' floors, so depending on
this crate cannot raise either tree's MSRV; verified with
`cargo +1.91.1 check --all-targets`. The lockfile is committed even though this
is a library: CI builds it directly, so a committed lock is what makes those
runs reproducible and their caches stable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust: take the address and TLS helpers from seaweed-common

Both public signatures are kept, so no caller outside the two wrapper files
changes. parse_grpc_address stays `Result<String, String>` and maps the typed
error through Display; server_to_grpc_address stays `Option<String>` and drops
it with .ok(). Their doc comments and the volume server's 13 call sites are
otherwise untouched.

Three behaviours change, each in the direction of the copy that was already
right:

- The volume server now brackets IPv6 literals. `::1:19333` used to come back as
  `::1:29333`, which build_grpc_endpoint rejects with "invalid gRPC endpoint
  http://::1:19333: invalid authority" — an IPv6 master or EC peer could not be
  dialled at all. Two tests in grpc_client.rs pin it, one on the string and one
  on the endpoint the string builds.
- The volume server now emits the *parsed* gRPC port of the dotted form instead
  of the original text it had just validated, so `host:8080.018080` and
  `host:8080.+18080` come back as `host:18080` rather than as authorities the
  URI parser rejects. Same port either way; only malformed spellings change.
- The worker's dotted form now validates the HTTP port it discards.
  `server_to_grpc_address("host:abc.18080")` used to answer Some("host:18080");
  it now answers None, which is what the volume server's copy has always done.

install_default_crypto_provider becomes a re-export in both trees, so
`crate::security::tls::install_default_crypto_provider` and
`weed_lance_worker::tls::install_default_crypto_provider` still resolve. The
lance crate's `rustls = "0.23"` was its only direct use of rustls and goes away
with the body; seaweed-common states the same requirement, so neither the
resolved version nor the enabled features move in either lockfile.

The PEM test fixtures stay where they are. The two tests that use them are not
duplicates: the volume server's exercises build_grpc_endpoint, and the lance one
exists precisely because aws-lc-rs and ring are both linked in that crate's
graph. Only the literals are shared, and exporting test fixtures from a library
to dedupe two constants costs more than it saves.

A path dependency outside both trees means every build context that copies one
crate directory has to copy the other. The repo has one: the Rust source-build
stage of docker/Dockerfile.go_build, which now copies seaweed-common beside
seaweed-volume. Every workflow whose `paths:` filter keys on a crate directory
gains `seaweed-common/**` — the two Rust test workflows, rust_binaries_dev,
container_dev and performance. The tag- and dispatch-triggered ones
(rust_binaries_release, container_release_unified, container_latest) have no
`paths:` filter and need nothing.

The two Rust test workflows also run `cargo test` in seaweed-common, from their
unit-test job, because a path dependency is not a workspace member and neither
tree's own `cargo test` reaches it. Each step builds into its job's cached
target directory, and both cache keys now hash seaweed-common/Cargo.lock as well
so a change there invalidates the cache it would otherwise silently reuse.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docker: keep go_build working for BRANCH revisions without seaweed-common

The rust_builder stage copies seaweed-common unconditionally now that seaweed-volume path-depends on it, but BRANCH can name any revision — including ones that predate the crate. Create the directory in the builder stage so the COPY always has a source; an empty dir beside an old seaweed-volume is harmless.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-20 13:30:58 -07:00

471 lines
15 KiB
YAML

name: "Performance"
on:
push:
branches: [ master ]
paths:
- '**/*.go'
- 'go.mod'
- 'go.sum'
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/perf/**'
- '.github/workflows/performance.yml'
workflow_dispatch:
inputs:
profile_duration:
description: "CPU profiling duration in seconds"
required: false
default: "30"
type: string
benchmark_files:
description: "Number of files for the throughput benchmark"
required: false
default: "100000"
type: string
benchmark_concurrency:
description: "Concurrent read/write workers"
required: false
default: "16"
type: string
benchmark_size:
description: "Simulated file size in bytes"
required: false
default: "1024"
type: string
s3_objects:
description: "Number of objects for the S3 benchmark"
required: false
default: "20000"
type: string
s3_size:
description: "S3 object size in bytes"
required: false
default: "4096"
type: string
concurrency:
group: ${{ github.head_ref || github.ref }}/performance
cancel-in-progress: true
permissions:
contents: read
env:
VOL_SIZE_LIMIT: "1024"
jobs:
performance-profile:
name: CPU and Heap Profile
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Build weed
run: go build -o weed_bin ./weed
- name: Start server with profiling enabled
run: |
mkdir -p ./perfdata
./weed_bin -v=1 server -debug -debug.port=6060 -dir=./perfdata \
-s3 -filer -volume.max=0 -master.volumeSizeLimitMB=100 \
-s3.port=8000 -s3.config=./docker/compose/s3.json \
> weed.log 2>&1 &
echo "WEED_PID=$!" >> "$GITHUB_ENV"
for i in $(seq 1 60); do
if curl -sf http://localhost:9333/dir/status >/dev/null 2>&1; then
echo "master is ready"
break
fi
sleep 1
done
# give the volume server a moment to register with the master
sleep 3
- name: Start memory sampler
run: |
bash test/perf/mem_sample.sh mem-profile.csv "server=${WEED_PID}" &
echo "SAMPLER_PID=$!" >> "$GITHUB_ENV"
- name: Capture profiles under load
run: |
DURATION="${{ github.event.inputs.profile_duration || '30' }}"
# drive write load so the sampled profile reflects real work
./weed_bin benchmark -master=localhost:9333 -writeOnly \
-c=16 -n=5000000 -size=1024 > benchmark-load.log 2>&1 &
echo "Sampling CPU profile for ${DURATION}s..."
curl -s "http://localhost:6060/debug/pprof/profile?seconds=${DURATION}" -o cpu.pprof
curl -s "http://localhost:6060/debug/pprof/heap" -o heap.pprof
curl -s "http://localhost:6060/debug/pprof/goroutine?debug=1" -o goroutine.txt
go tool pprof -top -nodecount=50 cpu.pprof > cpu-top.txt 2>/dev/null || true
go tool pprof -top -nodecount=50 -sample_index=inuse_space heap.pprof > heap-top.txt 2>/dev/null || true
- name: Record memory usage
if: always()
run: |
kill -TERM "${SAMPLER_PID}" 2>/dev/null || true
sleep 2
{
echo "## Memory usage (peak RSS)"
echo '```'
if [ -f mem-profile.csv.peak ]; then
awk -F'\t' '{printf "%-10s %8d KB (%.1f MB)\n", $1, $2, $2/1024}' mem-profile.csv.peak
else
echo "no memory samples captured"
fi
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
- name: Profile summary
if: always()
run: |
{
echo "## CPU profile (top functions)"
echo '```'
head -45 cpu-top.txt 2>/dev/null || echo "no cpu profile captured"
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
- name: Stop server
if: always()
run: kill "${WEED_PID}" 2>/dev/null || true
- name: Show server log on failure
if: failure()
run: tail -200 weed.log || true
- name: Upload profiles
if: always()
uses: actions/upload-artifact@v7
with:
name: performance-profile-${{ github.run_number }}
path: |
cpu.pprof
heap.pprof
cpu-top.txt
heap-top.txt
goroutine.txt
mem-profile.csv
mem-profile.csv.peak
weed.log
retention-days: 30
benchmark:
name: Throughput Benchmark (${{ matrix.impl }} volume)
runs-on: ubuntu-latest
timeout-minutes: 60
strategy:
fail-fast: false
matrix:
impl: [go, rust]
env:
IMPL: ${{ matrix.impl }}
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Install Rust toolchain
if: matrix.impl == 'rust'
uses: dtolnay/rust-toolchain@stable
- name: Cache cargo registry and target
if: matrix.impl == 'rust'
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-
- name: Build weed
run: go build -o weed_bin ./weed
- name: Build Rust volume server
if: matrix.impl == 'rust'
run: cd seaweed-volume && cargo build --release
- name: Start master and volume server
run: |
mkdir -p ./perfdata/master ./perfdata/vol
./weed_bin -v=1 master -ip=127.0.0.1 -port=9333 \
-mdir=./perfdata/master -peers=none \
-volumeSizeLimitMB="${VOL_SIZE_LIMIT}" -defaultReplication=000 \
> master.log 2>&1 &
echo "MASTER_PID=$!" >> "$GITHUB_ENV"
for i in $(seq 1 60); do
if curl -sf http://localhost:9333/dir/status >/dev/null 2>&1; then
echo "master is ready"
break
fi
sleep 1
done
if [ "${IMPL}" = "rust" ]; then
./seaweed-volume/target/release/weed-volume \
--master 127.0.0.1:9333 --ip 127.0.0.1 --ip.bind 127.0.0.1 \
--port 8080 --dir ./perfdata/vol --max 100 --preStopSeconds 0 \
> volume.log 2>&1 &
else
./weed_bin -v=1 volume -master=127.0.0.1:9333 -ip=127.0.0.1 \
-port=8080 -dir=./perfdata/vol -max=100 \
> volume.log 2>&1 &
fi
echo "VOLUME_PID=$!" >> "$GITHUB_ENV"
for i in $(seq 1 60); do
if curl -sf http://localhost:8080/status >/dev/null 2>&1; then
echo "volume server is ready"
break
fi
sleep 1
done
# let the volume server register with the master via heartbeat
sleep 3
- name: Start memory sampler
run: |
bash test/perf/mem_sample.sh mem-benchmark.csv \
"master=${MASTER_PID}" "volume=${VOLUME_PID}" &
echo "SAMPLER_PID=$!" >> "$GITHUB_ENV"
- name: Run throughput benchmark
run: |
N="${{ github.event.inputs.benchmark_files || '100000' }}"
C="${{ github.event.inputs.benchmark_concurrency || '16' }}"
SIZE="${{ github.event.inputs.benchmark_size || '1024' }}"
./weed_bin benchmark -master=localhost:9333 \
-c="${C}" -n="${N}" -size="${SIZE}" 2>&1 | tee benchmark-results.txt
- name: Run Go micro-benchmarks
if: matrix.impl == 'go'
continue-on-error: true
run: |
go test -run='^$' -bench=. -benchmem -benchtime=10x \
./weed/topology/... ./weed/util/log_buffer/... ./weed/util/buffered_queue/... \
2>&1 | tee go-benchmarks.txt
- name: Record memory usage
if: always()
run: |
kill -TERM "${SAMPLER_PID}" 2>/dev/null || true
sleep 2
{
echo "## Memory usage (peak RSS, ${IMPL} volume)"
echo '```'
if [ -f mem-benchmark.csv.peak ]; then
awk -F'\t' '{printf "%-10s %8d KB (%.1f MB)\n", $1, $2, $2/1024}' mem-benchmark.csv.peak
else
echo "no memory samples captured"
fi
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
- name: Benchmark summary
if: always()
run: |
{
echo "## Throughput benchmark (${IMPL} volume)"
echo '```'
grep -E "Concurrency Level|Time taken|Completed requests|Failed requests|Requests per second|Transfer rate" \
benchmark-results.txt 2>/dev/null || echo "no benchmark results captured"
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
- name: Stop processes
if: always()
run: |
kill "${VOLUME_PID}" "${MASTER_PID}" 2>/dev/null || true
- name: Show logs on failure
if: failure()
run: |
echo "=== master.log ==="; tail -100 master.log 2>/dev/null || true
echo "=== volume.log ==="; tail -200 volume.log 2>/dev/null || true
- name: Upload benchmark results
if: always()
uses: actions/upload-artifact@v7
with:
name: benchmark-results-${{ matrix.impl }}-${{ github.run_number }}
path: |
benchmark-results.txt
go-benchmarks.txt
mem-benchmark.csv
mem-benchmark.csv.peak
master.log
volume.log
retention-days: 7
s3-benchmark:
name: S3 Read/Write Benchmark (${{ matrix.impl }} volume)
runs-on: ubuntu-latest
timeout-minutes: 60
strategy:
fail-fast: false
matrix:
impl: [go, rust]
env:
IMPL: ${{ matrix.impl }}
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Install Rust toolchain
if: matrix.impl == 'rust'
uses: dtolnay/rust-toolchain@stable
- name: Cache cargo registry and target
if: matrix.impl == 'rust'
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-
- name: Build weed and S3 load tool
run: |
go build -o weed_bin ./weed
go build -o s3bench ./test/s3/benchmark
- name: Build Rust volume server
if: matrix.impl == 'rust'
run: cd seaweed-volume && cargo build --release
- name: Start cluster with S3 gateway
run: |
mkdir -p ./perfdata/master ./perfdata/vol ./perfdata/filer
./weed_bin -v=1 master -ip=127.0.0.1 -port=9333 \
-mdir=./perfdata/master -peers=none \
-volumeSizeLimitMB="${VOL_SIZE_LIMIT}" -defaultReplication=000 \
> master.log 2>&1 &
echo "MASTER_PID=$!" >> "$GITHUB_ENV"
for i in $(seq 1 60); do
if curl -sf http://localhost:9333/dir/status >/dev/null 2>&1; then
echo "master is ready"
break
fi
sleep 1
done
if [ "${IMPL}" = "rust" ]; then
./seaweed-volume/target/release/weed-volume \
--master 127.0.0.1:9333 --ip 127.0.0.1 --ip.bind 127.0.0.1 \
--port 8080 --dir ./perfdata/vol --max 100 --preStopSeconds 0 \
> volume.log 2>&1 &
else
./weed_bin -v=1 volume -master=127.0.0.1:9333 -ip=127.0.0.1 \
-port=8080 -dir=./perfdata/vol -max=100 \
> volume.log 2>&1 &
fi
echo "VOLUME_PID=$!" >> "$GITHUB_ENV"
for i in $(seq 1 60); do
if curl -sf http://localhost:8080/status >/dev/null 2>&1; then
echo "volume server is ready"
break
fi
sleep 1
done
sleep 3
./weed_bin -v=1 filer -master=127.0.0.1:9333 -ip=127.0.0.1 -port=8888 \
-s3 -s3.port=8000 -s3.config=./docker/compose/s3.json \
> filer.log 2>&1 &
echo "FILER_PID=$!" >> "$GITHUB_ENV"
for i in $(seq 1 30); do
if nc -z localhost 8000 2>/dev/null; then
echo "s3 gateway is ready"
break
fi
sleep 1
done
sleep 2
- name: Start memory sampler
run: |
bash test/perf/mem_sample.sh mem-s3.csv \
"master=${MASTER_PID}" "volume=${VOLUME_PID}" "filer=${FILER_PID}" &
echo "SAMPLER_PID=$!" >> "$GITHUB_ENV"
- name: Run S3 read/write benchmark
run: |
OBJECTS="${{ github.event.inputs.s3_objects || '20000' }}"
C="${{ github.event.inputs.benchmark_concurrency || '16' }}"
SIZE="${{ github.event.inputs.s3_size || '4096' }}"
./s3bench -endpoint=http://localhost:8000 \
-access-key=some_access_key1 -secret-key=some_secret_key1 \
-objects="${OBJECTS}" -size="${SIZE}" -concurrency="${C}" -mode=both \
2>&1 | tee s3-benchmark-results.txt
- name: Record memory usage
if: always()
run: |
kill -TERM "${SAMPLER_PID}" 2>/dev/null || true
sleep 2
{
echo "## Memory usage (peak RSS, ${IMPL} volume)"
echo '```'
if [ -f mem-s3.csv.peak ]; then
awk -F'\t' '{printf "%-10s %8d KB (%.1f MB)\n", $1, $2, $2/1024}' mem-s3.csv.peak
else
echo "no memory samples captured"
fi
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
- name: S3 benchmark summary
if: always()
run: |
{
echo "## S3 read/write benchmark (${IMPL} volume)"
echo '```'
grep -E "results:|Concurrency Level|Time taken|Completed requests|Failed requests|Requests per second|Transfer rate|Latency" \
s3-benchmark-results.txt 2>/dev/null || echo "no S3 benchmark results captured"
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
- name: Stop processes
if: always()
run: |
kill "${FILER_PID}" "${VOLUME_PID}" "${MASTER_PID}" 2>/dev/null || true
- name: Show logs on failure
if: failure()
run: |
echo "=== master.log ==="; tail -100 master.log 2>/dev/null || true
echo "=== volume.log ==="; tail -200 volume.log 2>/dev/null || true
echo "=== filer.log ==="; tail -200 filer.log 2>/dev/null || true
- name: Upload S3 benchmark results
if: always()
uses: actions/upload-artifact@v7
with:
name: s3-benchmark-results-${{ matrix.impl }}-${{ github.run_number }}
path: |
s3-benchmark-results.txt
mem-s3.csv
mem-s3.csv.peak
master.log
volume.log
filer.log
retention-days: 7