Commit Graph
592 Commits
Author SHA1 Message Date
Security Fix b79a6fd7f2 ssh: don't open port 22 for the loopback-only sshd
sshd-localhost.nix turns sshd on for every role so that "ssh root@localhost"
works, and binds it to 127.0.0.1 only ("zero network exposure", as its
own comment says, and as sshd.nix expects: the sshd feature "extends this
to 0.0.0.0 and opens port 22 on the firewall when the user enables remote
SSH").

It never turned off services.openssh.openFirewall, which NixOS defaults
to true and applies whether or not sshd listens on the ports. So port 22
was open in the firewall on every role, Desktop Only included, with
nothing behind it. No listener means no live exposure today; it does mean
the firewall was not saying what the documentation says, and the day
something does bind 22 on a wider address (a listenAddresses change, a
second daemon) it would be reachable without anyone having opened it.

openFirewall is now mkDefault false there. Nothing that wants SSH
published loses it: sshd.nix (feature sshd) and remote-deploy.nix both
add 22 explicitly, and only when enabled. Found by evaluating the real
module set; grepping for allowedTCPPorts does not see a NixOS default.

Evaluated with nix eval, TCP firewall ports:

                            before     after
  Server + Desktop          22 80 443 3051 8937   80 443 3051 8937
  Bitcoin Node Only         22 3051 8937 60847    3051 8937 60847
  Desktop Only              22                    (none)
  Desktop + features.sshd   22                    22
  Desktop + deploy.enable   22 3389               22 3389

and sshd's listen addresses are unchanged: 127.0.0.1 by default,
127.0.0.1 and 0.0.0.0 with the sshd feature. Desktop Only now opens no
TCP port at all; UDP 5353 (mDNS) is the only port open there, and
SECURITY.md says so.

Add tests/test_ssh_exposure.py.
2026-10-02 02:24:29 -05:00
Security Fix 3694ea6489 bitcoin: drop the stray UDP 3051 firewall rule
allowedUDPPorts was set to [ 3051 ] alongside allowedTCPPorts, which looks
like it was copied from the line above. Caddy serves Ride The Lightning
over TCP on 3051; nothing listens for UDP there, so the rule only opened a
port for no reason.

The comment above the pair also said "Hub management port"; 3051 is RTL.

Evaluated with nix eval, UDP firewall ports per role: Server + Desktop
80 443 3051 5353 -> 80 443 5353, Bitcoin Node Only 3051 5353 -> 5353,
Desktop Only 5353. (80 and 443 are Caddy's; 5353 is mDNS.)
2026-10-02 02:24:29 -05:00
Security Fix 6987d9bf2c installer: raise generated password entropy from ~23 to ~33 bits
generate_diceware_password() built the password from 3 words out of a 96
word list plus a single digit: 96^3 x 10 = 8,847,360 combinations, about
23 bits.

That one password is the desktop login, the 'free' account password, and
the only thing standing in front of the Hub, which runs as root and
displays the root password, the SSH passphrase, the RTL password and the
Vaultwarden admin token. 23 bits is thin for something that valuable, and
the rate limiting in front of it was weaker than intended (see "hub: make
the login lockout that LOGIN_FAIL_MAX described").

Now 4 words plus 2 digits: 96^4 x 100 = 8,493,465,600, about 33 bits,
for the cost of one more word to write down.

- iso/installer.py: generate_diceware_password().
- modules/credentials.nix: the three fallback generators in
  root-password-setup, free-password-setup and free-password-migration,
  so a machine provisioned without the installer gets the same strength.

Affects new installs only; existing passwords are untouched.

Checked by running the real thing. The installer function was exercised
2000 times: 96 words in the list, always word-word-word-word-NN, 33.0
bits. For the three services, the generator lines were taken from the
script the evaluated module really produces (nix eval on the nixpkgs
revision flake.lock pins) and run 1500 times each: 96 words in each
list, always word-word-word-word-NN, every two-digit suffix from 00 to 99
seen, and no repeated password.
2026-10-02 02:24:29 -05:00
Security Fix ebcc17ae3c caddy: stop filtering the RTL and Mempool sites by client address
c33457f put an address check (sovran_lan_only) on the Hub, RTL and
Mempool sites. It was written for ports 80/443 being forwarded and a
Host header selecting a site, and that only ever applied to the Hub. RTL
and Mempool are sites on ports of their own (:3051, :60847): a request
on 80/443 cannot select them, whatever Host it carries.

Checked with Caddy 2.9.1 and the Caddyfile this module generates for
Server + Desktop with every domain configured, serving the public sites
on a stand-in port: Host: sovransystemsos.local, localhost:8937,
127.0.0.1 and x:3051 all get an empty 200, and Host: matrix.example.org
gets the Synapse stand-in. With the Hub off Caddy the guard has nothing
left to guard, and it could not be made right for the two sites that
remain:

- IPv6. A laptop's global address on the LAN looks exactly like a
  stranger's. The choice was between letting all of 2000::/3 through,
  which is the whole IPv6 internet and is what c33457f does, and
  refusing every LAN device that connects over a global address unless
  the operator copies the ISP's prefix into a Nix option.
- They do not need it. RTL has a random 20-character password
  (pwgen -s 20, about 119 bits) and, since 0.15.12, which Sovran_Bitcoin
  pins, a 30-minute lockout keyed on the client address. Mempool shows
  public chain data. If someone forwards 3051 or 60847 that is the same
  exposure as any other port on the machine, and SECURITY.md says not to.

Remove the snippet and its two imports, and tests/test_caddy_lan_only.py
with them. What is still worth pinning moves to test_hub_direct.py: no
address filter anywhere in caddy.nix, the two sites are plain proxies to
their loopback ports, and no option for a declared prefix is left
half-wired.

Behaviour change: RTL and Mempool answer any client that can reach :3051
or :60847, as they did before c33457f. In practice that is the local
network, because nothing asks you to forward those ports.

Checked with the real generator and Caddy 2.9.1: Node Only generates
`auto_https off` and the two plain sites and validates. Run live next to
the real Hub, a LAN client gets the Hub, RTL and Mempool; a stranger's
address gets RTL and Mempool (by design) and a 403 from the Hub; port 80
is not listening on Node Only.
2026-10-02 02:24:29 -05:00
Security Fix 78bfc5b408 hub: serve the Hub on its own port instead of through Caddy
Caddy fronted the Hub at http://sovransystemsos.local, but the Hub
already listens on 0.0.0.0:8937 itself, and nothing Caddy added is
something it needs:

- Not the name. That is avahi's: mDNS advertises a hostname, not a port,
  so the name resolves wherever the Hub listens.
- Not TLS (the site was plain http), not authentication, not cache
  headers. The header block duplicated NoCacheMiddleware, and its
  Clear-Site-Data ("cache") overrode the app's stronger ("cache",
  "storage").
- Not access control, and this is the point. With ports 80/443 forwarded
  for public services, a Host header on those ports reached the Hub. That
  second door is how the reported bug happened, and c33457f guards it
  with an address check instead of closing it.

The Hub is now served on port 8937 only, at
http://sovransystemsos.local:8937, and Caddy has no site for it. The
only thing Caddy answers on 80/443 is the public sites. Caddy keeps
Ride The Lightning (:3051) and Mempool (:60847), because those do need
it: Sovran_Bitcoin binds both to 127.0.0.1 and RTL's unit is sandboxed to
loopback besides, so Caddy is how the local network reaches them.

- caddy.nix: no Hub site. Caddy runs wherever RTL and Mempool do, which
  includes Bitcoin Node Only. There it did not run at all (enable was
  needsHttpsPorts || extraVhosts != ""), so :3051 and :60847 were open
  in the firewall with nothing listening. Ports 80/443 still follow
  needsHttpsPorts alone, so Node Only does not open them. The two sites
  are written only where their service exists; they were unconditional.
- sovran-hub.nix: 8937 follows the new hub.directPort, 60847 follows
  Mempool. It used to be `[ 8937 60847 ]` on every role, Desktop Only
  included.
- roles.nix: hub.directPort defaults to !roles.desktop: open on Server +
  Desktop and Bitcoin Node Only, closed on Desktop Only, where the Hub is
  reached from the machine itself through the desktop window on
  localhost.
- The bind stays 0.0.0.0, which is IPv4 only: with that bind [::1]:8937
  is refused and "localhost" falls back to 127.0.0.1. That is on purpose
  and is now said in the comment. An IPv6 listener would let in clients
  whose global addresses the Hub cannot tell from a stranger's, which is
  the question the previous commit declines to answer by guessing.
- README, SECURITY.md and two strings in index.html give the new URL.

Behaviour changes: the Hub's address gains :8937, and http://sovransystemsos.local
on port 80 no longer reaches it. Bitcoin Node Only now runs Caddy.

Evaluated with nix eval (nixpkgs as flake.lock pins it, Sovran_Bitcoin at
the locked revision), firewall TCP ports per role:

                      c33457f                      this commit
  Server + Desktop    22 80 443 3051 8937 60847    22 80 443 3051 8937
  Bitcoin Node Only   22 3051 8937 60847           22 3051 8937 60847   (Caddy now runs)
  Desktop Only        22 8937 60847                22

Port 22 is open on every role although sshd listens on loopback only;
the last commit of this series deals with that.

The Caddyfile the module really generates (the evaluated generator
script, run, then `caddy validate` with Caddy 2.9.1): Node Only gets the
two sites and nothing else; Server + Desktop with every domain
configured gets the seven domain sites plus :3051 and :60847 and no
mention of the Hub; with Bitcoin off there are no local-network sites;
Node Only with Bitcoin off and no domains leaves Caddy off.

Add tests/test_hub_direct.py and keep tests/test_caddy_lan_only.py for
the two sites it still covers.
2026-10-02 02:24:29 -05:00
Security Fix 361b25a8bd hub: answer the local network only, whichever way a client arrives
The reported bug was the Hub being reachable from outside the local
network when Server + Desktop is active. c33457f guards the Caddy site
for sovransystemsos.local, but the Hub is an application that also
listens on a port of its own (8937), and a check in Caddy does nothing
for a client that never goes through Caddy. Whether a client could reach
the Hub depended on which door it used.

The Hub now asks the question itself. LanOnlyMiddleware is registered
outermost, so a client that is not on this computer or the local network
gets a bare 403 before authentication is considered; it never sees the
login page.

- Local means loopback, 10/8, 172.16/12, 192.168/16, 100.64/10
  (Tailscale and other CGNAT/VPN ranges) and 169.254/16 over IPv4, and
  ::1, fc00::/7 and fe80::/10 over IPv6.
- IPv6 global addresses are not on the list. A global address belonging
  to a laptop on the LAN cannot be told apart from a stranger's by the
  address alone, and 2000::/3 is every public IPv6 address there is.
- ::ffff:a.b.c.d is read as the IPv4 address inside it.
- sovran_systemsOS.hub.extraLanNetworks adds networks (IPv4 or IPv6 CIDR)
  for setups whose own devices use addresses outside those ranges. It is
  checked at build time. The app ignores an entry it cannot parse and
  refuses 0.0.0.0/0 and ::/0: it must never widen its policy by
  guessing, and "everyone" is hub.lanOnly = false, asked for by name.
- sovran_systemsOS.hub.lanOnly (default true) turns the check off.
- The first refusal from each address is logged, naming the option to
  change, so an operator whose own device is refused can find out why.
  The list is capped so a scanner cannot fill the journal or memory.

Behind Caddy the policy applies to the real client, not to Caddy: uvicorn
takes the address from X-Forwarded-For only when the peer is 127.0.0.1.

Checked, not only reviewed:
- The real app with the config.json the module really generates (nix
  eval on the nixpkgs revision flake.lock pins, read back from the
  derivation), on a real socket with the source address chosen per
  request: 127.0.0.1, 192.168/16, 100.64/10 and a declared extra
  network are served; 203.0.113.9, 8.8.4.4 and an address just outside
  the declared /28 get 403 on /login, / and /api/ping. /auto-login still
  answers 303 to loopback and 403 to everyone else.
- With c33457f's Caddy guard in front, an IPv6 client at 2001:db8::9
  passes Caddy (it is inside 2000::/3) and is refused by the Hub; an
  IPv4 stranger has the connection closed by Caddy; a LAN client is
  served.
- hub.extraLanNetworks accepts 203.0.113.0/28, 2001:db8:abcd::/48 and
  bare hosts, and fails the build for /33, 300.1.1.1/8, 0.0.0.0/0, ::/0,
  2001:db8::/129 and junk, with a message that says what to write.

Add tests/test_lan_policy.py and tests/test_hub_lan_only.py, and a note
in SECURITY.md.
2026-10-02 02:24:29 -05:00
Arena.ai Agent c33457fff2 caddy: serve the Hub, RTL and Mempool sites to local clients only
The Hub (sovransystemsos.local), Ride The Lightning (:3051) and Mempool
(:60847) sites are meant for the home network. With ports 80/443
forwarded for public services, Caddy also receives requests from other
clients, so these sites now check the client address as well as the Host
header.

A new snippet, sovran_lan_only, closes the connection unless the client
is on this computer or the local network: private_ranges, 100.64.0.0/10
(Tailscale), 169.254.0.0/16, fe80::/10 and fc00::/7. IPv6 global
addresses (2000::/3) are not filtered: computers on the network often
connect over their own global address, which cannot be told apart from
one on the internet by the address alone. Only the three local sites
import the snippet; the domain sites for public services are unchanged.

Clients with a public IPv4 address on the local network are no longer
served on these sites. The Hub is still available on port 8937.

Checked with Caddy 2.11.4 and the Caddyfile the generator writes: public
IPv4 clients get the connection closed on all three sites, local clients
are served, and the public domain sites answer as before.

Add tests/test_caddy_lan_only.py and a note in SECURITY.md.
2026-10-01 21:43:47 -05:00
Arena.ai Agent 2d777450e1 ddns: take the public IP from Njal.la only and give it to LiveKit
The Hub no longer asks a STUN server, a public DNS resolver or a "what
is my IP" service for the home IP address. The DDNS update asks Njal.la
to use the address the request comes from ("&auto"), reads back the
address Njal.la says it recorded and saves it to
/var/lib/secrets/external-ip. Njal.la is the only third party that
learns the address; it has to, to publish it.

- Add app/sovran_systemsos_web/ddns_update.py, installed as
  /etc/sovran/ddns-update.py and run by sovran-ddns-update.service. It
  runs curl --ipv4 without redirects, accepts only a public IPv4
  address, and rewrites the file atomically and only when the address
  changes. Stored "&a=${IP}" URLs are converted when they are used and
  "&quiet" is dropped.
- Rewrite modules/core/njalla.nix around that runner and delete
  modules/core/public-ip.nix. Setting a sovran_systemsOS.publicIP.*
  option now fails with a message that says where the address comes
  from. Activation removes the old scripts in /var/lib/sovran. The
  existing external-ip file keeps working.
- server.py reads the saved address and starts
  sovran-ddns-update.service after a domain is saved, instead of looking
  the address up itself.
- Element calling uses sovran_systemsOS.elementCalling.externalIP if
  set, otherwise the saved address, and fails with a clear message when
  neither exists or the address is not public. It no longer falls back
  to STUN. livekit-external-ip.path re-runs livekit-turn-setup and
  starts LiveKit when the address changes.
- Add tests/test_ddns_update.py.
2026-10-01 21:43:47 -05:00
Sovran Contributor 8bc325b148 postgresql: drop per-database autovacuum ALTERs (rejected by Postgres)
Follow-up to the previous two commits: the ALTER DATABASE ... SET
autovacuum_* commands fail at boot with

  ERROR: parameter "autovacuum_vacuum_scale_factor" cannot be changed now

and take matrix-synapse-db-tune.service (and nextcloud-db-init.service)
down with them.

Root cause: ALTER DATABASE/ROLE ... SET validates through
set_config_option() with an interactive context, and guc.c rejects any
PGC_SIGHUP parameter set that way. All autovacuum_* GUCs are
SIGHUP-context, so per-database scoping is impossible for them — only
USERSET-level parameters (e.g. work_mem, statement_timeout) can be set
per-database.

Nothing is lost: the same values are already set cluster-wide in
configuration.nix, which covers both nextclouddb and matrix-synapse.
Remove the ALTERs from nextcloud-db-init and delete the now-purposeless
matrix-synapse-db-tune service.
2026-09-21 14:21:01 -05:00
Sovran Contributor e31094c194 synapse: performance tuning for 32 GB Server+Desktop hosts
Monolith Synapse spends most of its RAM on caches to avoid Postgres
round-trips, but Sovran ships stock cache settings (global_factor 0.5,
10K event cache, no autotuning) and a default 5-connection DB pool.

- caches.global_factor 4.0 + 100K event cache + autotuning capped at
  2G (target 1G), with boosts for the /sync and room-join hot paths.
- DB pool cp_min 5 / cp_max 15, txn_limit 10000 (fewer reconnects).
- gc_thresholds raised to cut GC pauses on a 32 GB box.
- cache-memory extra for cache-size statistics.
- Per-database autovacuum (ALTER DATABASE, scoped to matrix-synapse)
  matching the nextclouddb tuning.

Deliberately unchanged: presence and URL previews stay enabled
(disabling them is faster but user-visible), and no workers — monolith
is the right call under ~100 users. Workers would need Redis
replication, the redis extra, and Caddy reverse-proxy rework; revisit
if federation load ever justifies it.
2026-09-21 14:00:19 -05:00
Sovran Contributor 09d4cc9b83 nextcloud, postgresql: fix Nextcloud 35 DB warnings on 32 GB hosts
Nextcloud 35's Database checks flag three Performance issues out of the
box: buffer cache hit ratio ~96% (wants 99%+), 100k+ dead tuples, and
million-plus sequential scans on oc_mail_tags / oc_guests_users.

Root causes in Sovran: stock 128MB shared_buffers, stock 60s autovacuum
naptime, APCu file locking, and db:add-missing-indices running exactly
once at install time (never on upgrades or app installs).

Size Postgres for the README's Server + Desktop recommendation (32 GB
RAM, NVMe): 2GB shared_buffers, 12GB effective_cache_size, 512MB
maintenance_work_mem, 32MB work_mem, 4GB max_wal_size, 30s autovacuum
naptime with 4 workers. shared_buffers stays below the 25% rule because
Postgres shares the box with bitcoind, Electrs, LND, MariaDB and PHP-FPM.

Scope the aggressive autovacuum to nextclouddb via ALTER DATABASE so the
shared matrix-synapse DB keeps the milder cluster defaults.

Add a local Redis (127.0.0.1:6379, Nextcloud only) and move
memcache.distributed/locking to Redis; migrate existing installs with a
one-shot since nextcloud-init never re-runs.

Add a weekly nextcloud-db-maintenance timer (VACUUM ANALYZE +
db:add-missing-*) so upgrades and later app installs can't regress the
checks again.

Note: shared_buffers needs one 'systemctl restart postgresql', which
briefly takes down both Nextcloud and Matrix. Everything else is
reload-only or scoped to nextclouddb.
2026-09-21 13:59:52 -05:00
naturallaw777 3341659a0c updated to php85 and fixes 2026-09-19 16:24:57 -05:00
naturallaw777 859f25f0c1 hub: point RTL credentials at /rtl/ and bump dev version to 0.15.12
Sovran_Bitcoin 0.15.12 serves the RTL UI under /rtl/ (upstream Angular
<base href="/rtl/"> + PathLocationStrategy); the package redirects / to
/rtl/ only for the exact root path.

Update the Hub's RTL tile so the Tor and Local Network credentials show
the canonical /rtl/ URLs (with trailing slash, which the redirect does
not cover) instead of relying on the root-path redirect. Bump the dev
versions.json fallback for rtl.service from 0.15.10 to 0.15.12 to match
the flake (deployed systems already read pkgs.sovran-bitcoin.rtl.version).
2026-09-08 21:39:27 -05:00
Sovran Systems d1e226a687 fix(hub): self-heal truncated/corrupt Nix downloads and keep failed updates retryable
Updater/rebuild self-heal:
- Add a shared run_step wrapper used by both the update and rebuild
  scripts. On the first failure matching a transient fetch/cache signature
  (truncated tarball, corrupt NAR, hash mismatch, network timeout,
  interrupted download), clear Nix's fetch caches and repair the store,
  then retry once. Real config errors do not match and still fail loudly.
- The kernel-change boot fallback in the rebuild path is also wrapped.
- Fixes the reported 'cannot read file from tarball: Truncated tar archive
  detected' failure, which a plain re-run cannot clear because Nix reuses
  the corrupt cached archive.

Failed-update recovery / reporting:
- check_for_updates() now compares the running Hub version against the
  branch VERSION, so a failed 'nix flake update' (lock advanced but no
  generation staged) can no longer masquerade as 'up to date' and block
  retries.
- /api/updates/check surfaces a persistent 'failed' state; /api/updates/run
  never blocks a retry after a failure.
- Dashboard shows a red 'Update failed - click to retry' tile; the modal
  offers a Retry Update button and stops offering a reboot on failure.
2026-09-03 11:54:29 -05:00
Sovran_SystemsOS f34d1533c3 element-calling: fix Nix string interpolation of LAN_CIDR echo
The debug echo used bash ${LAN_CIDR:-<none>} syntax, but inside a Nix
indented string ${...} is Nix interpolation, not bash. Nix parsed
'LAN_CIDR:-<none>' as a lambda and failed the build with 'cannot coerce a
function to a string'. Rewrite the echo without brace expansion.
2026-09-01 12:29:47 -05:00
Sovran_SystemsOS 81ab3b2280 element-calling: fix Wi-Fi calls and tighten media/TURN ports
Root cause of 'calls fail on Wi-Fi but work on mobile data': LiveKit only
advertised the public/WAN IP (rtc.node_ip), so LAN clients had to hairpin
through the router for media. Fixes and cleanup:

- rtc.advertise_internal_ip: true — also advertise the primary interface's
  LAN host candidate, so Wi-Fi callers connect directly (no hairpin).
- Drop rtc.port_range_start/end (30000-40000) and keep the single UDP mux
  (udp_port: 7882). In LiveKit 1.13.x the range takes precedence over
  udp_port, so media was actually spread over 10000 ports.
- Drop turn.tls_port: 5349 — LiveKit advertises turns:<domain>:443 to clients
  regardless of tls_port, so a 5349 TURN/TLS listener was unreachable dead
  config (and needless attack surface).
- Pin TURN relay allocation to 40000-40099 (disjoint from the media mux) and
  open/forward that range; the old default overlapped RTC media.
- turn.allow_restricted_peer_cidrs with the LAN subnet derived from the
  primary interface: without it the relay refuses to deliver to the private
  LAN host candidate and its final hop would fall back to WAN hairpin.
- Update Hub port guidance (server.py) to the new list.
2026-09-01 12:15:16 -05:00
naturallaw777 ddf87a1c1c refactor(nwc): dedupe NWC tooling — use Sovran_Bitcoin's sovran-nwc
The Hub vendored a second copy of the NWC stack
(app/sovran_systemsos_web/nwc_hub_manager.py, nwc_audit.py,
nwc_lnurl_service.py, nwc_wallet_cli.py) and built its own nwc-wallet /
nwc-lnurl binaries from it. That copy drifted from the pinned Alby Hub
API contract (appId vs toAppId) and duplicated code that Sovran_Bitcoin
already ships and fixes.

Changes:
- Delete the four vendored modules; server.py now imports the canonical
  implementation directly (from sovran_nwc import nwc_hub_manager) from
  the sovran-nwc package (pkgs.sovran-bitcoin.nwc). API fixes in
  Sovran_Bitcoin now propagate to the Hub web app automatically.
- sovran-hub-web launcher: add <sovran-nwc>/lib/sovran-nwc to
  sys.path so the import resolves.
- Stop shipping nwc-wallet / nwc-lnurl binaries from sovran-hub-web:
  the flake already provides them (env-wrapped nwc-wallet with
  NWC_* vars via albyhub.nix, and nwc-lnurl.service via lnurl.nix).

Requires a Sovran_Bitcoin rev containing the toAppId fix (and the
LNURL module audit-log fix); bump the flake input afterwards:
  nix flake update sovran-bitcoin

Test:
  - nixos-rebuild switch
  - Hub Wallet Connections tab still lists/creates wallets
  - nwc-wallet list works from the operator shell
  - journalctl -u nwc-lnurl shows no import/contract errors
2026-08-31 20:28:10 -05:00
naturallaw777 f0e4c33a5f fix: set lnurl domainFile for Hub-managed Lightning Address domain 2026-08-31 10:58:04 -05:00
naturallaw777 ef1c045e0e fix: correct disablewallet casing 2026-08-31 10:46:19 -05:00
naturallaw777 e27bf0ef1f fixed typo 2026-08-31 10:27:11 -05:00
Sovran Patch 362fa0b36c refactor: extract bitcoin stack into Sovran_Bitcoin flake input
Decouple the Bitcoin/Lightning modules and packages into the standalone
Sovran_Bitcoin flake, consumed as a NixOS module input.

Deleted (now in Sovran_Bitcoin):
  - modules/bitcoin/          (19 files — vendored nix-bitcoin modules)
  - modules/bitcoinecosystem.nix
  - modules/nwc-wallets.nix
  - modules/mempool.nix
  - packages/{albyhub,mempool,rtl,build-support}/
  - tests/bitcoin-btcpay-hardening.nix

Created:
  - modules/sovran-bitcoin-integration.nix — the OS-specific bridge that
    maps sovran_systemsOS.* options to sovran-bitcoin.* and applies
    Second_Drive paths, operator 'free', forced wallet, firewall 3051,
    and Sovran Hub NWC environment wiring.

Modified:
  - flake.nix — added sovran-bitcoin flake input, updated module imports
  - modules/modules.nix — removed deleted imports
  - modules/core/sovran-hub.nix — version metadata now reads from
    pkgs.sovran-bitcoin.* overlay instead of local packages/
  - tests/test_bitcoin_tor_gossip.py — updated to check integration layer

The sovran_systemsOS.* option namespace is preserved. The Hub, roles,
and custom.nix continue to work unchanged.
2026-08-31 10:17:14 -05:00
Arena Agent 60b55f715e fix(hub): evaluate local vendored package versions directly
Fix the version metadata generation in sovran-hub.nix for Alby Hub, RTL, and Mempool. Previously, the build would incorrectly fall back to older upstream nixpkgs versions because the package names exist upstream, despite the OS deploying custom vendored forks locally. This replaces the fragile checks with direct evaluations of the local packages.

Also updates the development fallback versions.json to reflect the current vendored Alby Hub version (1.24.0).
2026-08-27 09:05:20 -05:00
naturallaw777 d600198049 feat(albyhub): vendor v1.24.0 LND-only, no-frontend build
Replace the nixpkgs albyhub overrideAttrs patch-chain with a fully
vendored package at packages/albyhub for v1.24.0.

- modules/core/sovran-hub.nix: bump fallback version 1.8.0 -> 1.24.0
- modules/nwc-wallets.nix: build via pkgs.callPackage ../packages/albyhub
- packages/albyhub:
  - drop 0002-isolated-invoice-app-id.patch (fixed upstream)
  - add 0004-lnd-only.patch: strip LDK/Bark/Cashu/CLN/Phoenix backends
    from service/start.go, leaving only the LND case
  - add 0005-no-frontend.patch: remove //go:embed dist and the
    frontend handler registration
  - add default.nix: buildGoModule for v1.24.0 with no nodejs/yarn/
    bark-ffi-go/ldk-node deps (only stdenv.cc.cc), subPackages cmd/http

Keeps 0001-private-route-hints and 0003-loopback-bind-host. Does not
touch flake.nix, VERSION, or CHANGELOG.
2026-08-25 19:11:25 -05:00
naturallaw777 47e9bb99b0 fix(public-ip): move system.activationScripts under the config attribute
The module mixed the `options` keyword attribute with a bare top-level
`system.*` setting. Once a module declares `options` (or `config`), every
other top-level attribute must be a reserved module keyword — the nixpkgs
unifyModuleSyntax check rejects anything else, so every nixos-rebuild
aborted at evaluation time with:

  error: Module '.../modules/core/public-ip.nix' has an unsupported
  attribute `system'. ... move all of them (namely: system) into the
  `config' attribute.

Prefix the activation script with `config.` (equivalent to wrapping it in
`config = { ... };`) so the module evaluates again. The detector script
itself is unchanged.

Fixes: ac6c498615 (-feat(public-ip): unify public-IP detection into one privacy-first script-)
2026-08-20 17:00:23 -05:00
naturallaw777 ac6c498615 feat(public-ip): unify public-IP detection into one privacy-first script
The public IP was previously detected independently in three places,
each contacting a different third party: the Hub (HTTPS echo via
api.ipify.org / ifconfig.me / icanhazip.com on every API call and
background tick), DDNS (myip.opendns.com via OpenDNS), and LiveKit
(embedded STUN). Consolidate into a single detector with one shared
cache so every consumer reads the same value with minimal exposure.

- add modules/core/public-ip.nix: installs /var/lib/sovran/public-ip.py
  (pure Python stdlib, no new deps) writing /var/lib/secrets/external-ip
- detection chain (first success wins): explicit pin, fresh cache
  (default TTL 300s), STUN binding request over UDP (one packet, no
  metadata), DNS myip.opendns.com query, then OPT-IN HTTPS echo
  (publicIP.httpsEcho, empty by default — never contacted unless listed)
- privacy: while the cache is fresh zero third parties are contacted;
  at most one party learns the IP per refresh interval, via the least
  exposing mechanism available
- hub (server.py): _get_external_ip() now reads the shared detector /
  cache instead of calling ipify/ifconfig/icanhazip directly
- ddns (njalla.nix): use the shared detector instead of a separate
  OpenDNS dig; allow the hardened service to write /var/lib/secrets
- element-calling: livekit-turn-setup falls back to the shared
  detector on cold boot; add LiveKit webhooks to lk-jwt-service
  (sfu_webhook) so abrupt disconnects are cleaned up immediately;
  set LIVEKIT_SANITY_CHECK_INTERVAL_SECONDS=60 as a missed-webhook
  guard; drop the dead services.livekit.settings block and set
  openFirewall=false (Caddy fronts the SFU; no public 7880/tcp)
- new options: sovran_systemsOS.publicIP.{stunServer,stunPort,
  dnsResolver,httpsEcho,cacheTTL}
2026-08-20 16:40:00 -05:00
naturallaw777 224ea99ce4 fix(element-calling): reuse Hub external IP for LiveKit, drop egress service calls 2026-08-20 16:12:00 -05:00
naturallaw777 c54dbfe2a5 feat(element-calling): fix Element X discovery and harden federated calling
The element-calling feature only advertised the LiveKit focus via the
well-known org.matrix.msc4143.rtc_foci file, and relied on STUN
auto-detection for the public IP. Element X queries the MatrixRTC
transports registry endpoint and fails with MISSING_MATRIX_RTC_TRANSPORT
when it is absent, and blocked STUN egress silently left LiveKit
advertising a private IP (call connects but no video across servers).

- synapse: enable msc4143_enabled and advertise matrix_rtc.transports
  (MSC4519) with the site's element-calling URL, so Element X can
  discover the LiveKit focus instead of erroring out
- livekit: determine the public IP to advertise at runtime —
  explicit pin, then HTTPS egress detection (api.ipify.org /
  checkip.amazonaws.com / ifconfig.me), then STUN fallback with a
  warning; reject non-routable results (private/loopback/CGNAT)
- lk-jwt-service: append optional extra homeservers to
  LIVEKIT_FULL_ACCESS_HOMESERVERS via the new
  sovran_systemsOS.elementCalling.fullAccessHomeservers option
- add sovran_systemsOS.elementCalling.externalIP option to pin the
  advertised public IP for multi-WAN/VPN setups
- add element-calling-public-check.service: boot-time diagnostics for
  public DNS (via 1.1.1.1, bypassing local loopback overrides), JWT
  healthz through Caddy and via the public IP, and the transports
  endpoint — turns the silent -no media- failure into a visible error
- add restartTriggers so livekit/lk-jwt-service pick up regenerated
  runtime configs on rebuild
2026-08-20 14:07:06 -05:00
naturallaw777 a1fa40cacf fix(hub): derive Restart required from boot default vs running system
NixOS already knows whether a reboot is pending: /nix/var/nix/profiles/
system vs /run/current-system. Marker files only the Hub's own updater
wrote desynced for terminal-updated machines (and markers from older
updaters could never clear), pinning the badge on forever. Reconcile
REBOOT_REQUIRED against live state on every read; the stale marker
self-heals to IDLE. The .generation marker write is now informational.
2026-08-19 11:31:29 -05:00
naturallaw77 48dacbeef3 refactor(lnd): use pkgs.lndinit, drop local packages/lndinit
The local packages/lndinit/default.nix is a verbatim copy of the
upstream Nixpkgs expression, frozen at v0.1.3-beta (the version that
was vendored in from nix-bitcoin before the Aug 10 2026 refactor in
commit 1fbeafd). It carries no Sovran-specific patches, no local
overrides, and no behavioral modifications - it is byte-for-byte
identical to what Nixpkgs ships, except ~19 minor versions older
(Nixpkgs currently ships 0.1.22-beta; the developer upstream
lightninglabs/lndinit is at v0.1.36-beta as of June 10 2026).

Why this matters
----------------

The Aug 10 2026 refactor (1fbeafd, "refactor: move vendor/nix-bitcoin
to modules/bitcoin, remove overlays") stated the new convention:

    No more random vendor/ or pkgs/ dirs - follows Sovran convention:
    modules/ for NixOS modules, packages/ for packages

That refactor successfully removed:
  * pkgs/sovran-overlay.nix
  * pkgs/nbxplorer.nix
  * pkgs/README.md
  * modules/vendor/ (entire directory)
  * overlay-sovran from flake.nix

It moved the lndinit expression into packages/lndinit/default.nix as
an intermediate step, but the file is still a verbatim upstream copy
and therefore still incurs the maintenance burden the refactor was
meant to eliminate: manual version bumps, manual hash refreshes, and
no upstream security or bug-fix flow. Removing it completes the
intent of 1fbeafd.

The change
----------

modules/bitcoin/lnd.nix (line 153):

    -  lndinit = "${(pkgs.callPackage ../../packages/lndinit {})}/bin/lndinit";
    +  lndinit = "${pkgs.lndinit}/bin/lndinit";

The two later uses of `lndinit` in the same file (lines 243 and 247,
inside the systemd.services.lnd.preStart block that calls
`lndinit gen-seed` and `lndinit init-wallet`) are unchanged because
they reference the let-bound `lndinit` value, not the callPackage
expression. They continue to work with the new pkgs.lndinit binary
path transparently.

Removed:
  * packages/lndinit/default.nix
  * packages/lndinit/ (now empty directory)

No other files in the repository reference packages/lndinit. Verified
by:
  * Git tree search for "packages/lndinit" -> only the file and its
    parent directory match
  * Content grep of flake.nix, configuration.nix,
    modules/bitcoin/default.nix, modules/bitcoin/common.nix, and
    iso/common.nix -> zero matches

Why this is safe
----------------

1. CLI compatibility. The preStart script only invokes two
   lndinit subcommands:
     * `lndinit gen-seed`
     * `lndinit -v init-wallet --file.seed=... --file.wallet-password=... --init-file.output-wallet-dir=...`
   Both subcommands and all four flags have been stable since the
   0.1.x line. The Nixpkgs 0.1.22-beta binary produces a wallet.db
   and admin.macaroon in the same on-disk format that 0.1.3-beta did
   for the same LND version (LND is pinned separately by pkgs.lnd
   from Nixpkgs and is unaffected by this change).

2. No coupled Go modules or shared vendor tree. The local
   packages/lndinit/default.nix is a self-contained buildGoModule
   derivation; it has no shared state with any other Sovran package.

3. Nixpkgs pin is current. flake.nix pins
   github:NixOS/nixpkgs/nixos-unstable, which has shipped pkgs.lndinit
   since 2022 and is currently at 0.1.22-beta. There is no
   "missing attribute" risk.

4. Wallet data is forward-compatible. The wallet.db format is owned
   by LND, not lndinit. lndinit is only used at first boot to create
   the seed and initialize the wallet; subsequent LND restarts do
   not invoke lndinit. So even if a user already initialized a
   wallet with 0.1.3-beta, the binary being upgraded to 0.1.22-beta
   is irrelevant - LND owns the wallet from that point on.

5. Single call site. Only modules/bitcoin/lnd.nix references
   lndinit. No other modules, scripts, or tests need to change.

Operational notes
-----------------

* After this commit, lndinit updates flow through the normal
  `nix flake update` workflow (or whatever automated dependency
  tooling is already in use, e.g. for the recent "chore(deps):
  update RTL to 0.15.10" commits). No Sovran-side action is needed
  to pick up future lndinit versions.

* If a future LND version requires a specific lndinit version, the
  pin can be done in flake.nix via a one-line overlay:

      nixpkgs.overlays = [ (final: prev: {
        lndinit = prev.lndinit.overrideAttrs (o: {
          version = "X.Y.Z-beta";
          src = prev.fetchFromGitHub { ... };
          vendorHash = "...";
        });
      }) ];

  This keeps the upgrade path explicit without bringing the entire
  expression back into the Sovran tree.

* This drops ~20 lines of frozen derivation code, eliminates one
  source of upstream drift, and reduces the surface area of what
  Sovran needs to keep current.
2026-08-18 12:30:59 -05:00
naturallaw777 b6e6adbe31 chore(deps): update RTL to 0.15.10
Refresh the RTL source and Node dependency hashes, and align Hub version metadata with the vendored package.
2026-08-18 11:19:23 -05:00
naturallaw777 64624002bb fix(hub): reconcile completed updates after polling stalls
The full-system updater runs as a detached systemd service and can finish
successfully even when the browser loses its status connection. In that
case the update log and status file correctly report REBOOT_REQUIRED, but
the Hub modal can remain on "Updating..." with its controls disabled.

There were four independent ways for the frontend to get stuck:

* update status fetches had no deadline, so a request that stayed pending
  never rejected and never advanced the existing failure counter;
* setInterval started async polls without waiting for the previous poll,
  allowing slow requests to overlap and responses to arrive out of order;
* each log chunk used textContent +=, replacing the complete and growing
  Nix build log every two seconds, which could stall browser rendering and
  was especially visible over RDP; and
* page reload, tab resume, and RDP reconnect did not reattach the modal to
  the update status persisted by the backend.

This produced a dangerous UX mismatch: the machine had a fully staged
NixOS generation and was ready to reboot, while the Hub continued telling
the user that the update was still running.

Bound status requests with AbortController, prevent overlapping polls, and
replace the endless spinner after sustained failures with an explicit
"Update status unavailable" state and Retry Status action. Reconcile state
immediately on focus, visibility, online, page startup, and before starting
a new update. Use no-store requests and render verbose logs incrementally
with a bounded visible tail while retaining the complete report in memory.
Apply the same timeout and single-flight protection to rebuild polling.

Record the exact generation produced by `nixos-rebuild boot`. The Hub now
keeps REBOOT_REQUIRED visible until that generation matches
/run/current-system, then clears the marker after reboot. For an update
started by an older updater that did not write the marker, recover the
staged generation from the final nixos-rebuild log line. The dashboard
sidebar also distinguishes update-in-progress and restart-required states.

Regression coverage verifies generation marker/log recovery, pre- versus
post-reboot detection, request timeout wiring, single-flight polling,
connection-loss UX, RDP/tab resume reconciliation, bounded log rendering,
page-reload recovery, and JavaScript syntax.

Validation:
* python3 -m unittest discover -s tests -p 'test_*.py' -v (170 passed)
* node --check app/sovran_systemsos_web/static/js/*.js
* python3 -m py_compile for changed Python modules
* git diff --check

A Nix evaluation was not available in the development sandbox; the NixOS
module should still be evaluated and built in CI or on a test machine before
release.
2026-08-18 10:31:28 -05:00
naturallaw777 1ccce429a5 fix(bitcoin): migrate i2pd SAM settings for nixpkgs 26.11
nixpkgs commit c8f9654 refactored the services.i2pd module to use
an RFC42-style settings attribute set and removed services.i2pd.proto.
After updating the root nixpkgs input from f13ff45 to ec2d622, the
vendored bitcoind module failed evaluation on the obsolete
services.i2pd.proto.sam.enable definition.

The error occurred even with services.bitcoind.i2p at its false default:
bitcoind was enabled, so NixOS still validated the obsolete option path
inside the conditional i2pd integration.

Read the SAM endpoint from services.i2pd.settings.sam and configure its
new upstream-style fields explicitly. Keep 127.0.0.1:7656, matching the
old typed option defaults that bitcoind uses to generate its i2psam
setting.

This preserves optional I2P support without activating it by default.
i2pd remains disabled until services.bitcoind.i2p is set to true or
"only-outgoing".

Nixpkgs migration: https://github.com/NixOS/nixpkgs/commit/c8f965411e812060a9377fa4c2d7d0f84e8b10e0
2026-08-18 09:19:15 -05:00
naturallaw777 db5b9f6b60 fix(livekit): order turn-setup after network-online to fix boot-time red dot
livekit-turn-setup.service detects the primary interface from the IPv4
default route, but had no ordering against network-online.target. With
NetworkManager+DHCP the default route is applied late at boot, so the
oneshot could run before it existed, exit 1, and — being a hard
dependency of livekit.service — take livekit down with it. The Hub then
showed a 'failed' red dot until livekit was restarted manually.

Order both livekit.service and livekit-turn-setup.service after
network-online.target. Also add a bounded retry when copying Caddy's ACME
cert so we never write an empty turn.crt/turn.key on a fresh boot.
2026-08-17 19:07:54 -05:00
Sovran Systems 2ea1427766 fix: restore a Zeus-scannable LND REST connect QR
The LND-only rewrite of lndconnect.nix shipped a wrapper Zeus cannot
use: unknown flags (--cert/--macaroon), a non-existent onion path
(free/lnd.onion), and a REST hidden service that collided with LND's
P2P onion. Restore the nix-bitcoin contract — dedicated lnd-rest
onion on port 8080, --nocert over Tor, admin macaroon in the URI —
and only persist a valid lndconnect:// URI for the Hub QR.
2026-08-17 18:22:36 -05:00
naturallaw777 a9ff168fd6 security: prevent LND admin macaroon exposure in curl argv 2026-08-15 23:00:59 -05:00
naturallaw777 2a1d73af25 Migrate dock/folder/mime entries from brave to brave-origin on upgrade 2026-08-15 17:31:25 -05:00
naturallaw777 587c19c2a5 fix(hub): persistent browser profile so logout survives window reopen
The Hub launcher used an ephemeral /tmp profile deleted on exit, which
wiped the hub_manual_logout marker cookie. On reopen, /auto-login minted a
new session and logged the user straight back in without a password.

Use a persistent per-user profile under XDG_STATE_HOME and drop the
deletion trap so the logout marker survives close/reopen. Keep
--skip-origin-startup-dialog. Adds regression tests.
2026-08-15 17:22:58 -05:00
naturallaw777 c862ed5806 Skip Brave Origin startup dialog in Hub launcher 2026-08-15 17:15:12 -05:00
naturallaw777 5edf594eac Switch default browser to Brave Origin (brave-origin) 2026-08-15 16:58:33 -05:00
naturallaw777 30b753ec40 feat: add Bitcoin Core Tor IBD gossip control 2026-08-15 14:44:13 -05:00
naturallaw777 3541f6baa1 Replace Bitcoin Knots with Bitcoin Core 2026-08-13 13:43:17 -05:00
naturallaw777 8f89a4350a fix: prevent Bitcoin Core switch from hanging the Hub UI 2026-08-11 18:47:40 -05:00
naturallaw777 4193c56397 fix: correct RTL and Mempool Hub versions 2026-08-11 13:07:28 -05:00
copilot-swe-agent[bot]andnaturallaw777 3f233beea0 njalla.nix: fail on ImportError; fix redundant except clause
Co-authored-by: naturallaw777 <99053422+naturallaw777@users.noreply.github.com>
2026-08-11 15:40:04 +00:00
copilot-swe-agent[bot]andnaturallaw777 947c04834d Fix all 8 security hardening blockers for PR #423
Co-authored-by: naturallaw777 <99053422+naturallaw777@users.noreply.github.com>
2026-08-11 15:38:08 +00:00
naturallaw777 894707a87c Correctly escape DDNS placeholder in Nix string 2026-08-11 10:11:44 -05:00
naturallaw777 43fc01d350 Fix Nix interpolation in DDNS runner 2026-08-11 10:06:32 -05:00
copilot-swe-agent[bot]andnaturallaw777 a111de1ece Security hardening: fix all 8 blocking findings for PR #419
Fix 1: Update support.js to collect SSH public key and POST JSON
Fix 2: Legacy njalla.sh migration - parse safely, archive non-executable, replace cron with systemd timer
Fix 3: DDNS SSRF prevention - allowlist only njal.la, reject other hosts, disable curl redirects
Fix 4: Legacy root support-key removal migration (_remove_legacy_root_support_key)
Fix 5: Automatic support-key expiration (expires_at + _expire_support_if_stale)
Fix 6: Move security helpers to security_helpers.py, tests import production code
Fix 7: Real NIP-19/Bech32 npub validation (_bech32_decode + _validate_npub)
Fix 8: Replace journalctl sudo wildcard with restricted sovran-journal-helper.py
Also: Make _write_hub_overrides() atomic with tempfile+os.replace
94 tests passing

Co-authored-by: naturallaw777 <99053422+naturallaw777@users.noreply.github.com>
2026-08-11 12:07:18 +00:00
copilot-swe-agent[bot]andnaturallaw777 9b77b04741 Fix IP validation in DDNS and document journalctl sudo rule
Co-authored-by: naturallaw777 <99053422+naturallaw777@users.noreply.github.com>
2026-08-11 10:45:50 +00:00
copilot-swe-agent[bot]andnaturallaw777 f2ad9c1f17 Security hardening: fix DDNS injection, Nix injection, reboot auth, support key, sudo rules
Co-authored-by: naturallaw777 <99053422+naturallaw777@users.noreply.github.com>
2026-08-11 10:44:26 +00:00