Files
tf_provider/HISTORY/2026-09-28_vpn-transit-213-vultr.md
T

167 lines
9.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# VPN transit via VM 213 and Vultr
Date: 2026-09-27 to 2026-09-28
## Goal
Provide access from Russian residential/mobile networks to services restricted by Russian network filtering, while retaining the existing foreign egress on Vultr.
## Verified network facts
- Test host `3060`: `46.39.251.163`, connection from Khimki / Iskratelecom.
- Transit VM `213`: `5.172.178.213`, public egress observed as `5.172.178.65`; hosted in NUBES data centre.
- Vultr addresses: primary `95.179.252.111`; secondary `104.238.177.67`.
- `3060 -> 213`: ICMP approximately 3 ms, 0% loss.
- `213 -> Vultr`: ICMP approximately 34 ms, 0% loss; HTTPS response returned in about 0.07-0.11 s.
- Direct `213 -> Vultr` test file transfer: 10 MiB in 1.59 s, about 6.27 MiB/s / 50.2 Mbit/s.
- Direct `3060 -> Vultr` test file transfer timed out / was throttled.
- Direct `213 -> OVH proof endpoint`: 10 MiB in 1.18 s, about 8.5 MiB/s.
- Direct access from `213` to YouTube and Telegram failed with `HTTP=000` and timeout/SSL errors, while OVH and Google returned HTTP 200. Therefore a foreign egress remains required for those services.
## Persistent changes on VM 213
- Created backup:
- `/etc/nginx/sites-available/check.kube5s.ru.bak_vpn`
- Modified:
- `/etc/nginx/sites-available/check.kube5s.ru`
- Added an Nginx `/ws` reverse-proxy location with:
- upstream `https://95.179.252.111:443`
- SNI `vipien.kube5s.ru`
- upstream Host header `vipien.kube5s.ru`
- WebSocket upgrade headers
- 3600-second proxy timeouts
- Ran `nginx -t` successfully and reloaded Nginx.
- Existing unrelated Nginx warnings about duplicate `contracts.kube5s.ru` server names remained.
## Persistent/previously existing changes on Vultr
The following configuration was read or used during validation:
- `/etc/nginx/conf.d/vipien.conf`: TLS/WebSocket endpoint for `vipien.kube5s.ru`.
- `/etc/v2ray-agent/xray/conf/08_VLESS_ws_inbound.json`: VLESS WebSocket inbound on `127.0.0.1:10086`, path `/ws`.
- `/etc/systemd/system/hysteria-server.service`: Hysteria service was stopped and disabled; it was not changed in this work.
- Xray service was confirmed active.
- Nginx service was confirmed active.
- Cloudflared tunnel configuration was inspected earlier, but it is not used by the final working route.
- A temporary 10 MiB test file was created on Vultr and removed after testing.
## Temporary files on test VM 3060
The following temporary client files were created under `/tmp/xray-test/` for validation and are not repository files:
- `client-cf.json`
- `client-213.json`
- `client-directip.json`
- temporary log/test artifacts where applicable
The files contained test Xray client configurations. They were used only to verify the route from `3060`; no permanent system service was installed there.
## Final tested route
`client in Russia -> 5.172.178.213:443 -> Nginx WebSocket proxy -> 95.179.252.111:443 -> Xray -> Internet`
Final test from `3060` through the route:
- observed outbound IP: `95.179.252.111`
- 10 MiB OVH download: 1.76-1.91 s
- measured speed: approximately 5.5-6.0 MiB/s
## Final client parameters
- Address: `5.172.178.213`
- Port: `443`
- UUID: existing UUID used by the Vultr Xray inbound
- TLS SNI: `check.kube5s.ru`
- WebSocket path: `/ws`
- WebSocket Host: `vipien.kube5s.ru`
The final direct-IP test used Xray 26.3.27. The client-side `allowInsecure` option was not used because this Xray version reports that the option was removed.
## Secondary Vultr IP
Before removal, the Nginx upstream on VM 213 was switched from `104.238.177.67` to `95.179.252.111`. A post-switch end-to-end test succeeded, with outbound IP `95.179.252.111` and approximately 6.0 MiB/s.
No Vultr IP deletion was performed in this work. The secondary address was only confirmed as no longer referenced by the transit configuration.
## Scope audit
- No repository source/configuration files were edited before this record.
- `git status` was clean before this documentation file was created.
- This documentation file is the only workspace file created by the current documentation action.
- Server-side files were changed on VM 213 and earlier on Vultr; temporary test files were also created on VM 3060.
- No commit was created for this record.
## Important limitations
The measurements prove the route worked at test time. They do not guarantee permanent availability: NUBES, Vultr, upstream providers, or network filtering policy can change independently.
## Later the same day: optimisation attempt and its outcome
### Automation created
A reusable, idempotent tool was created outside this repository:
```text
/home/naeel/nubes/HowTo/vpn-transit/vpn-setup.sh check | apply | verify | passthrough | verify-passthrough | client-config | rollback
/home/naeel/nubes/HowTo/vpn-transit/client-config.json generated client config (chmod 600, contains UUID)
/home/naeel/nubes/HowTo/vpn-transit/README.md description, measurements, rollback
/home/naeel/nubes/HowTo/howto-vpn-transit-213-vultr-2026-09-28.md full report
```
Every change is preceded by a timestamped backup and followed by a config test (`nginx -t`, `xray run -test`) with automatic rollback on failure.
### Changes applied
| Host | File | Change | Backup |
|---|---|---|---|
| 213 | `/etc/nginx/sites-available/check.kube5s.ru` | `proxy_buffering off;` added inside `location /ws`, marked `# vpn-transit: proxy_buffering off` | `check.kube5s.ru.bak.1790601681` |
| Vultr | `/etc/v2ray-agent/xray/conf/00_log.json` | `loglevel`: `debug` → `warning` (log had grown to 76 MB), service restarted | `00_log.json.bak.1790601723` |
| 213 | `/usr/local/sbin/vpn-transit-dnat.sh`, `/etc/systemd/system/vpn-transit-dnat.service` | DNAT `213:8443 → 95.179.252.111:443` plus FORWARD rules, enabled at boot | none (rules tagged `vpn-transit`) |
### Measurements after the changes
- Outbound IP: `95.179.252.111`
- Throughput: `5.6–7.3 MiB/s` (10 MiB in 1.4–1.9 s)
- Per-connection latency: `0.23–0.37 s`
- WebSocket upgrade success rate on 213: `14569 / 14573` (99.97%), one `upstream timed out` error
### Hypothesis that was disproved: mux
`verify` compared the tunnel with and without `"mux": {"enabled": true, "concurrency": 8}`:
| Mode | 10 MiB download | Connection behaviour |
|---|---|---|
| without mux | 7.32 MiB/s in 1.43 s | stable |
| with mux | **0 B/s, failed** | after 4 requests connections hang for 15 s |
Conclusion: mux is harmful in the `VLESS + WebSocket behind nginx` combination. It is excluded from the client config. The test remains in the script for re-checking on future Xray versions.
### Optimisation that could not be delivered: removing the second TLS layer
The intended speed fix was to drop one TLS handshake (`client → 213`, then `213 → Vultr`) by forwarding TCP straight through to Vultr.
- `ngx_stream_module.so` is absent on 213, so nginx cannot do SNI-based passthrough without installing `libnginx-mod-stream`.
- Kernel-level DNAT on port 8443 was installed instead, but **does not work**: from outside, port 8443 returns `Connection refused` and the DNAT counter on 213 stays at 0 packets — traffic never reaches the machine.
- Cause: the provider firewall in front of 213 exposes only ports 80 and 443. Measured from `3060`: `3001, 8080, 8443, 8766, 8767, 8888, 18080, 40229` are closed.
- Therefore the second TLS layer can only be removed after the provider opens an additional port. The rules are already installed and would start working immediately once that happens.
### Errors made during this work
1. **Recommended `mux` before measuring it.** The recommendation was given as the main fix and was later disproved by measurement. Correct order: measure first, recommend after.
2. **Changed server configuration before measuring the benefit.** `proxy_buffering off` has no effect on a WebSocket connection after the `101 Switching Protocols` upgrade, and `loglevel` affects only log size. Neither change improves speed, so from the user's point of view nothing changed.
3. **Changed the client config to port 8443 before verifying the port was reachable from outside.** The config was regenerated back to port 443 immediately.
### Net result for the user
Nothing changed for the client: address `5.172.178.213`, port `443`, SNI `check.kube5s.ru`, path `/ws` and the UUID are unchanged, and the previously used link still works. No client-side reconfiguration is required.
The only actionable finding is client-side: the Xray log on Vultr contained **331** `connect: connection refused` to `127.0.0.1:45987`, i.e. the client requested a loopback address, plus Telegram advertises AAAA records while the tunnel is IPv4-only. The generated `client-config.json` addresses both (remote DNS, `queryStrategy: UseIPv4`), but the device itself was not modified.
Separately: **10170** `reset by peer` entries to `157.240.0.13` (Meta infrastructure) are blocking by those sites, unrelated to the transit.
### Scope audit (this action)
- Repository files changed: this document only. `git status` also showed unrelated pre-existing changes (`DEV_STAND/FullPipe/shturval.tf` deletion, `TMP/*` files) that were **not** touched or committed.
- Server-side files changed: as listed in the table above.
- Temporary test files on 3060: `/tmp/xray-test/*` (no permanent service installed).