Corrects a claim from the previous update: Satellite Phase 3 shipped, so discovery scans now run through satellite-covered sites too (only SNMP health-metric polling and WebSSH remain Central-only). New capabilities added to the feature list: per-device availability windows (suppress false OFFLINE warnings during expected downtime), per-SSID MAC access-control lists with a dedicated Wireless ACL tab, a new RADIUS Management section (global FreeRADIUS server/NAS/user management), and three OPNsense monitoring additions (BGP neighbors, TLS certificate/Trust-store monitoring, DDNS-down warning). Also notes that netOrk's own config pushes are now auto-recognized so they're never mistaken for an unauthorized change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
306 lines
15 KiB
Markdown
306 lines
15 KiB
Markdown
# netOrk — Product Description
|
||
|
||
## One-liner
|
||
|
||
**netOrk is a self-hosted network orchestration platform that discovers,
|
||
monitors, and manages heterogeneous network infrastructure from a single UI.**
|
||
|
||
## Elevator Pitch (3 sentences)
|
||
|
||
netOrk connects to your routers, switches, access points, firewalls, and servers
|
||
via NAPALM and custom vendor drivers — regardless of manufacturer. It continuously
|
||
polls device state, detects configuration drift, and lets you push corrections in
|
||
one click. All findings are synced to NetBox as the source of truth, and security
|
||
integrations with Wazuh and Graylog give you visibility across the full stack.
|
||
|
||
---
|
||
|
||
## Target Audience
|
||
|
||
**Primary:** Network engineers and IT administrators managing small to medium
|
||
heterogeneous environments (10–500 devices) — mixed vendor, mixed OS.
|
||
|
||
**Secondary:** Serious homelab operators who run "prosumer" or enterprise-grade
|
||
hardware and want operational visibility beyond what consumer dashboards offer.
|
||
|
||
**Pain points this solves:**
|
||
- Multiple vendor-specific management UIs open at once
|
||
- No single view of "what's running where"
|
||
- Config changes made directly on devices — nobody knows what changed
|
||
- Manual SSH into every device to check interface status or VLAN membership
|
||
- Security tooling (Wazuh agents, syslog) not consistently deployed
|
||
|
||
---
|
||
|
||
## Value Propositions
|
||
|
||
1. **One UI for everything** — OpenWRT APs, OPNsense firewalls, HP ProCurve
|
||
switches, Proxmox hosts, Linux servers, Fritz!Boxes, NAS devices, and more,
|
||
all managed in one place.
|
||
|
||
2. **Intent-based configuration** — Define desired state in netOrk (VLAN names,
|
||
SSID settings, AP radio profiles). netOrk pushes the config to devices and
|
||
corrects drift automatically or on demand.
|
||
|
||
3. **Automatic drift detection** — Every poll compares device config against the
|
||
DB. Drifted devices get a warning; a one-click fix stream applies the
|
||
correction via SSH/UCI/REST and shows live output.
|
||
|
||
4. **Scheduled automation** — Automatic reboots for OpenWRT APs (GTK key rotation
|
||
workaround), scheduled config fixes, package updates — all with time windows
|
||
and per-site concurrency limits.
|
||
|
||
5. **Security visibility** — Wazuh agent tracking with CVE counts and alert history
|
||
per device. Graylog syslog forwarding status and auto-fix. CrowdSec org-level
|
||
threat summary per device.
|
||
|
||
6. **Deep NetBox integration** — Devices, interfaces, IP addresses, prefixes,
|
||
VLANs synced to NetBox automatically. netOrk uses NetBox as the canonical
|
||
documentation target.
|
||
|
||
7. **Plugin system** — Integrations (Wazuh, Graylog, CrowdSec, apt-cacher-ng)
|
||
are plugins that can be enabled/disabled per deployment. Adding a new
|
||
integration follows a documented pattern.
|
||
|
||
8. **Self-hosted, no SaaS** — Runs in Docker Compose. Your data stays on your
|
||
infrastructure. No telemetry, no cloud dependency.
|
||
|
||
9. **NIS2 evidence foundation** — NIS2 Art. 21 mandates asset inventory, patch
|
||
management, access control, and audit trails. netOrk produces all of these as
|
||
day-to-day operational outputs: full device inventory, per-device update status,
|
||
Wazuh CVE tracking, EOL firmware/OS flagging, RBAC with MFA, Git-backed config
|
||
snapshots with diff/restore, config drift detection, and a complete audit log.
|
||
|
||
10. **Build your own view** — Configurable, shareable dashboards: pick from 13
|
||
widgets, arrange them on a WYSIWYG grid, and share the result with colleagues
|
||
who can subscribe to the live version or clone their own copy.
|
||
|
||
11. **From zero to managed in one flow** — Provision a Cloud-Init VM on a
|
||
Proxmox hypervisor, assign Ansible roles to configure it, and netOrk
|
||
auto-links it as a Device — no separate tools, no manual SSH-and-copy.
|
||
|
||
12. **Reach sites netOrk can't touch directly** — Deploy a lightweight
|
||
Satellite agent to poll devices locally at a disconnected or firewalled
|
||
site and sync results back over HTTPS; scheduled fixes route through it
|
||
the same way they do for directly reachable devices.
|
||
|
||
---
|
||
|
||
## Feature List
|
||
|
||
### Device Management
|
||
- CRUD for devices with credential profiles and SSH key management
|
||
- Per-device poll intervals (minutes) or manual-only
|
||
- Status tracking: planned / staged / active / decommissioning / offline / disabled
|
||
- Vendor/model/OS auto-populated from NAPALM `get_facts()`
|
||
- Site assignment with FK to structured Site records
|
||
- AP Profile assignment for grouped OpenWRT config
|
||
|
||
### Discovery
|
||
- ICMP ping sweep, SNMP scan, HTTP/HTTPS probing
|
||
- Device fingerprinting: vendor + platform confidence scoring
|
||
- FQDN resolution (reverse DNS)
|
||
- Manual adoption from scan results (no auto-create to avoid inventory noise)
|
||
|
||
### VM Provisioning
|
||
- Cloud-Init based VM creation directly from a hypervisor's VMs tab — no
|
||
manual template or VMID setup
|
||
- Multi-distro image catalog: Debian 12, Ubuntu 22.04/24.04/26.04,
|
||
Fedora 42/43/44, with Ubuntu and Fedora releases synced automatically as
|
||
new versions ship
|
||
- Pick a target VLAN and an IP from its subnet — netOrk creates the DHCP
|
||
reservation automatically
|
||
- Cloud-init provisions a real Linux user with an SSH key, plus configurable
|
||
bootstrap toggles (SNMP, QEMU guest agent)
|
||
- Reusable provisioning templates for repeatable bootstrap settings
|
||
- The new VM is auto-linked as a netOrk Device and its hostname assigned to
|
||
a DNS zone once bootstrap finishes
|
||
- Deploy progress shown as a live step checklist in the UI
|
||
- Delete a VM and its linked netOrk Device together, gated behind a
|
||
name-confirmation prompt
|
||
|
||
### Supported Device Drivers
|
||
Custom NAPALM drivers for all of the following:
|
||
|
||
| Driver | Device type |
|
||
|---|---|
|
||
| `openwrt` | OpenWRT access points |
|
||
| `opnsense` | OPNsense firewalls |
|
||
| `proxmox` | Proxmox VE hypervisors |
|
||
| `linux` | Generic Linux servers |
|
||
| `procurve` | HP ProCurve / Aruba switches |
|
||
| `tplink_jetstream` | TP-Link Jetstream managed switches |
|
||
| `netgear` | Netgear switches |
|
||
| `fritzbox` | AVM Fritz!Box routers |
|
||
| `zyxel` | Zyxel switches |
|
||
| `openmediavault` | OpenMediaVault NAS |
|
||
| `sonos` | Sonos speakers |
|
||
|
||
Plus all built-in NAPALM drivers: Cisco IOS/IOS-XE/NX-OS, Arista EOS, Juniper JunOS.
|
||
|
||
### Networking & Inventory
|
||
- Interface browser with IPv4/IPv6 addresses, MAC, speed, MTU
|
||
- LLDP neighbor discovery and topology graph
|
||
- ARP table and DHCP lease browser per device
|
||
- Subnet browser with interface-to-subnet assignments
|
||
- VLAN list grouped by site; per-VLAN device membership view
|
||
- SSID management with push to OpenWRT APs via UCI
|
||
- Per-SSID MAC access control lists (whitelist / blacklist) pushed to every
|
||
AP broadcasting the SSID, quick-add straight from the Connected Clients list
|
||
- MAC ACL state is a first-class drift item — covered by the same drift
|
||
detection, scheduled auto-fix, and warning aggregation as any other config
|
||
drift
|
||
- Dedicated Access Control Lists tab on the Wireless page listing every SSID
|
||
with its ACL editor inline
|
||
|
||
### Configuration Management
|
||
- Config drift detection: desired state (DB) vs device state (poll snapshot)
|
||
- One-click drift fix stream with live SSH output in the browser
|
||
- UCI-based config push for OpenWRT (VLAN names, SSID settings, radio config)
|
||
- AP profile system: country code, HT/VHT mode, 802.11r, NTP, syslog, SSH port
|
||
- Configuration backup & versioning: every poll captures a config snapshot into
|
||
a local Git repository, with full history and a side-by-side diff viewer
|
||
between any two points in time
|
||
- One-click config restore for OPNsense from any prior snapshot
|
||
- Unauthorised configuration changes are surfaced as a device warning
|
||
- netOrk's own config pushes (drift fixes, ACL provisioning) are recognized
|
||
and auto-accepted as the new baseline — never mistaken for an unauthorized
|
||
change
|
||
|
||
### Configuration Automation (Ansible)
|
||
- Reusable Ansible roles and playbooks stored and edited directly in
|
||
netOrk — no separate git checkout
|
||
- 11 built-in roles ready to assign: base, ubuntu, docker, adguard, zoraxy,
|
||
portainer, watchtower, uptime-kuma, vaultwarden, wireguard, fail2ban
|
||
- Automatic dependency resolution — assigning `docker` pulls in `base`
|
||
automatically, no manual role ordering
|
||
- Built-in roles can't be deleted but are fully editable; customizations
|
||
survive upgrades, and only untouched files auto-heal on bugfixes
|
||
- `ansible-doc`-backed autocomplete while writing roles and playbooks
|
||
- Upload your own role as an archive
|
||
- Device-level role assignment with a dedicated Ansible tab on the device
|
||
detail page
|
||
- Run history per device, snapshotting the exact role/playbook content
|
||
that was executed
|
||
- Wired into VM provisioning: assign roles at VM-creation time and they
|
||
run automatically after boot
|
||
|
||
### Scheduled Operations
|
||
- Scheduled reboots for OpenWRT APs with per-site concurrency lock
|
||
- Failback cron script written to device for netOrk-unreachable scenarios
|
||
- Scheduled config drift fixes with time-window enforcement
|
||
- Package update scheduling and one-click apply
|
||
- Wake-on-LAN via a firewall's driver (OPNsense today) — saved WOL targets
|
||
with on-demand "Wake now" and recurring schedules; save a seen host as a
|
||
target directly from the DHCP/ARP tabs
|
||
|
||
### Satellite Deployments
|
||
- Lightweight Docker agent deployed at a site netOrk can't reach directly —
|
||
polls devices locally and syncs results back to Central over HTTPS
|
||
- Deployed in one flow via VM provisioning: pick a hypervisor and site,
|
||
netOrk provisions the VM and installs the satellite container automatically
|
||
- Central automatically skips direct polling for any device at a site with
|
||
an online, heartbeating satellite — no manual per-site toggling
|
||
- Scheduled/on-demand reboots and the SNMP auto-fix flow run through the
|
||
same command channel whether a device is directly reachable or behind a
|
||
satellite
|
||
- Discovery jobs at a satellite-covered site scan locally through the same
|
||
command channel, instead of failing to reach the subnet from Central
|
||
- Not yet satellite-covered: SNMP health-metric polling still runs from
|
||
Central, and WebSSH console access isn't available through a satellite
|
||
|
||
### Monitoring & Health
|
||
- SNMP health metrics (CPU, memory, interface counters) via `get_health_metrics()`
|
||
- Per-device warning system with severity levels (error / warning / info)
|
||
- One-click Ack on any warning — clears it immediately and writes an audit log
|
||
entry; for config-change warnings the current state is accepted as the new
|
||
baseline
|
||
- Docker container and image status (Proxmox/Linux)
|
||
- Service status and start/stop/restart (systemd)
|
||
- VM/container list with OS device cross-linking (Proxmox)
|
||
- Per-device availability windows — suppress OFFLINE status and poll-failure
|
||
warnings during expected downtime (e.g. a nightly power-off); opt-in,
|
||
unconfigured devices are unaffected
|
||
- OPNsense: BGP neighbor status polling and display, with a peer-down warning
|
||
- OPNsense: TLS certificate monitoring for the Trust store, with
|
||
expiring-soon / expired warnings
|
||
- OPNsense: Dynamic DNS service-down warning (os-ddclient)
|
||
|
||
### Dashboards
|
||
- Configurable, shareable dashboards — build your own from a widget picker
|
||
instead of a fixed layout
|
||
- WYSIWYG grid-layout editor: drag, resize, and arrange widgets on a canvas
|
||
- 13 widget types: stats, device warnings, recently updated devices, network
|
||
topology, EOL status, config drift summary, Wazuh security alerts, audit log
|
||
activity, discovery jobs status, upcoming scheduled actions, DNS zones
|
||
overview, site overview, config snapshot history
|
||
- Multi-instance widgets with independent per-widget settings
|
||
- Share a dashboard with specific users; recipients can subscribe to the
|
||
owner's live version or clone it into their own editable copy
|
||
- Favorite dashboards for quick access from the main menu; set any dashboard
|
||
as your home view
|
||
|
||
### Security Integrations (plugins)
|
||
- **Wazuh** — agent enrollment tracking, vulnerability counts (by severity),
|
||
recent alert history, CIS benchmark scores, one-click agent install fix stream
|
||
- **Graylog** — rsyslog forwarding status per device, one-click fix to write rule
|
||
- **CrowdSec** — org-level decisions, remediation metrics, top attack scenarios
|
||
- **EOL Tracking** — flags devices running end-of-life or soon-to-be-end-of-life
|
||
firmware/OS via the endoflife.date API, checked daily
|
||
|
||
### DNS
|
||
- DNS zone management with authoritative device assignment
|
||
- Forward record provisioning (A records from device interfaces)
|
||
- PTR record provisioning to reverse zones
|
||
- Pending job queue for zone changes when device is unreachable
|
||
|
||
### RADIUS Management
|
||
- Global FreeRADIUS server, NAS client, and user management — no site
|
||
scoping, usable from any site's SSIDs/APs
|
||
- Changes push to the device via the driver and are only stored after the
|
||
device confirms — avoids drift between netOrk's view and the actual
|
||
FreeRADIUS config
|
||
- Dedicated server list and detail page (NAS clients / users tabs) under
|
||
the Wireless section
|
||
- SSID 802.1X integration (auto-provisioning a NAS client from an SSID's
|
||
RADIUS server) is a deliberate follow-up, not included yet
|
||
|
||
### Access Control
|
||
- JWT authentication with remember-me (localStorage) or session-only (sessionStorage)
|
||
- Two-factor authentication (MFA/TOTP) — authenticator app at login, backup
|
||
codes for emergencies, session invalidation on TOTP changes, enforceable
|
||
per role
|
||
- RBAC with four built-in roles: viewer / operator / engineer / administrator
|
||
- Custom roles with any permission combination
|
||
- Full audit log of all orchestration actions, filterable by date range,
|
||
user, action, or resource — export to CSV or PDF
|
||
|
||
### NetBox Sync
|
||
- Pushes vendor, model, OS version, status to NetBox dcim.devices
|
||
- Syncs interfaces, IP addresses, prefixes, VLANs
|
||
- VM interfaces and disks for Proxmox hosts
|
||
|
||
### Developer / Operator Experience
|
||
- OpenAPI / Swagger at `/docs`
|
||
- Plugin system: new integrations follow a documented pattern
|
||
- Celery task queue with dedicated queues per workload type
|
||
- Docker Compose deployment (single command)
|
||
- Alembic migrations run automatically on deploy
|
||
|
||
---
|
||
|
||
## Architecture in One Paragraph
|
||
|
||
netOrk runs as a set of Docker containers: a FastAPI API server, three Celery
|
||
worker pools (general, poll, and Ansible), a Celery Beat scheduler, and an
|
||
nginx UI server. Redis is the broker. PostgreSQL stores all state. Device
|
||
communication is always blocking I/O executed in Celery workers — FastAPI
|
||
request handlers are async-only for DB and quick operations. Custom NAPALM
|
||
drivers live in `vendor/` as editable packages and self-register via
|
||
`@register_driver`. The plugin system (`netork/plugins/`) provides a hook bus,
|
||
a plugin registry with enable/disable state in the DB, and a documented
|
||
pattern for adding integrations. For sites Central can't reach directly, a
|
||
separate Satellite container polls devices locally and syncs results back
|
||
over HTTPS; Central dispatches actions (reboots, SNMP fixes) to it through a
|
||
generic command channel, transparently to the UI.
|