* test: stop swapping http.DefaultTransport in the Happ network test
The Happ local-generation test replaced the http.DefaultTransport global to
refuse sockets. resolvePublicIPsInBackground fires a fire-and-forget resolver
(ServerService.GetStatus -> resolvePublicIPs -> getPublicIP), and that
goroutine can outlive the test that scheduled it, so a later test writes the
global while the resolver is reading it: a data race that `make race`
(-shuffle=on) only reports when the status test happens to run first.
Route the lookup transport through an atomic swap point and count in-flight
lookups, so the test stubs its own path and waits for earlier resolvers to
settle before installing it. That also stops another test's dials from
landing in this test's zero-network tally.
Reproduced on the parent commit with -shuffle=6 (read in getPublicIP,
previous write in TestHappGenerateLocallyWithoutNetwork); the same seed is
clean with this change, across repeats and for the whole package under -race.
* test: prove the Happ network guard intercepts panel egress
The PR review pointed out that HappService.Generate never reaches a stub
installed only on the public-IP lookup, so the "zero network attempts"
assertion could not fail. Move the override onto the panel's shared egress
seam — getPublicIP and SettingService.NewProxiedHTTPClient — and gate the
test on a canary request that must be refused and must move the tally, so a
future outbound call from local generation is reported instead of passing
unnoticed.
* test(service): stop the cold-status test leaking a public-IP resolver
The race CI hit (run 37345781756) came from TestCurrentStatusSamplesBeforeFirstTick:
its CurrentStatus call starts resolvePublicIPsInBackground, a goroutine that
makes real internet requests and outlives the test. On a runner without
IPv6 every lookup waits out its 3s timeout, so the goroutine is still
building clients from http.DefaultTransport when the Happ test swaps it.
Pre-settling the IP cache keeps that test from starting the resolver at
all, so no unit test reaches the internet or leaks the goroutine. This
replaces the earlier panelEgressTransport / panelEgressLookups seam, which
added test-only hooks to production code without removing the leak.
Reproduced in golang:1.27-bookworm with egress routed to a blackhole
(HTTPS_PROXY=http://10.255.255.1:9, -race -count=3 on the two tests): the
same server.go:406 race as CI before, clean over -count=5 after.
---------
Co-authored-by: MHSanaei <ho3ein.sanaei@gmail.com>
Most tests opened a throwaway panel DB with database.InitDB, which runs the
full AutoMigrate + seed on an empty file every time: ~230ms, and ~850ms under
-race because GORM's reflection-heavy migration is what the detector slows
most. internal/web/service does this in ~550 of its 830 tests, so the CI race
job spent ~10 of its ~14.6 minutes re-migrating empty databases.
internal/database/dbtest.InitDB migrates once per test process, then hands
each test its own copy of that file (~130ms under -race) and registers the
CloseDB cleanup. The copy then goes through InitDB like a panel restart, so
every test still starts from the state a fresh install has. Tests that reopen
an existing file, migrate a hand-built legacy DB or target Postgres keep
calling database.InitDB.
Locally under -race: internal/web/service 626s (last CI run) -> 114s,
internal/sub 246s -> 35s.
Adding a node fails right after that node's panel restarts. nodes/add
probes the node's /panel/api/server/status first, and that endpoint
returns whatever the @2s ticker last sampled - nil until the first tick
lands, so the master reads a healthy panel as unreachable and rejects it
with "Add node (remote returned success=false: )", an error whose
message is empty because the node answered success with a null obj.
The window is far wider than one tick: GetStatus resolved the public
IPv4/IPv6 addresses inline and held s.mu across every lookup, so a box
with no IPv6 route spent 3s per service - about 15s of nil status after
each restart, and the same stall on a fresh panel's first sample.
- status now answers from CurrentStatus, which samples on demand when
the ticker has not run yet instead of returning a null obj
- the public-IP lookups run in the background and outside s.mu, so a
status sample never waits on them
- probe tells "no status yet" apart from a genuine success=false, so the
master's error says something when it meets an older node