Commit Graph
25 Commits
Author SHA1 Message Date
Christian Manivong a13af149d9 feat(vm-provision): let Proxmox download cloud images into an import storage
Provisioning downloaded every cloud image over SSH into
/var/lib/vz/template/netork-images, on the node's root filesystem. On a
small root that fills up and takes Proxmox down with it (netOrk #480).

When the node has an active storage with content type "import" (Proxmox
8.2+), Proxmox now does it itself: download-url with checksum
verification into that storage, then import-from as the root disk. The
file is named after a hash of the full URL and reused when present.
Proxmox takes the format from the extension and has no ".img", so
Ubuntu's qcow2 .img is stored as .qcow2 -- a wrong guess fails at import
instead of attaching a qcow2 container as a raw disk.

Without an import storage, or for an image type Proxmox cannot import,
the SSH download is used as before.
2026-10-02 08:59:17 +02:00
christianmanivong a9f4cd249f Merge pull request 'feat: implement the HypervisorDriver VM contract' (#1) from feature/hypervisor-contract into master 2026-10-01 18:59:36 +00:00
Christian Manivong abce85d6ef fix: resolve the node the connection landed on, not the first cluster member
_resolve_node() took the first entry of GET /nodes. In a cluster that
lists every member, so a node polled without an explicit `node` driver
argument talked to whichever member came first: pve-dual reported
pve-02's name and VMs, and netOrk's VM sync moved pve-02's VM devices
over to it.

Resolve through GET /cluster/status instead: the entry marked local,
then a match by IP or (short) name, then the sole node of a standalone
host, and otherwise raise rather than guess.

The lookup no longer swallows API errors either. A TLS verification
failure used to leave the IP as the node name, so open() succeeded and
every getter failed quietly while the poll reported success with empty
data. It now surfaces as a ConnectionException from open().

Refs NetOrk/netork#417, NetOrk/netork#418
2026-09-29 10:37:01 +02:00
Christian Manivong 08bfb5c1c0 feat: VM snapshots and reboot_host through the API
get_vm_snapshots, create_vm_snapshot, delete_vm_snapshot and
rollback_vm_snapshot for VMs and containers, so netOrk's snapshot view
works on Proxmox as it does on VMware. Proxmox lists the live state as a
pseudo-snapshot named "current"; it is never reported or addressable.
Containers have no RAM state, so include_memory is ignored for them.

reboot_host() restarts the node with POST /nodes/{node}/status
command=reboot instead of /sbin/reboot over SSH.
2026-09-24 10:00:12 +02:00
Christian Manivong dd48d3c1e5 feat: implement the HypervisorDriver VM contract
start_vm, stop_vm, reboot_vm, suspend_vm and get_vm_config existed only
as declarations. netOrk called Proxmox's own power_vm and read a VM's
raw config through _node_api(), so no other hypervisor could serve the
same endpoints. These let netOrk talk to every hypervisor alike.

The power methods accept a VM's name or vmid, wait for the Proxmox task,
and raise ValueError/RuntimeError as the contract says instead of
returning a result dict. A forced reboot of a container is stop + start,
since LXC has no reset; suspending a container is refused. power_vm is
unchanged for existing callers.

get_vm_config moves the config parsing netOrk did in
_parse_proxmox_hw_config into the driver and returns a VMConfigDict:
disks with storage and size, NICs with model, MAC, bridge and VLAN, CPU
topology, firmware, machine type and PCI/USB passthrough.

get_vms reports vmid as a string ("100"), following
napalm-device-types 2.0, still ordered numerically.
2026-09-24 09:06:59 +02:00
Christian Manivong fc9f1426be fix: report the source package's version, not only its name
`get_packages` named the Debian source package and never its version, so a
consumer was handed two numbers on different axes and no way to tell.

OSV states Debian ranges in *source* versions. libldb2 is
2:2.11.0+samba4.22.11+dfsg-… while its source, samba, is 2:4.22.11+dfsg-…;
comparing the first against a samba range is meaningless, and dpkg reads
ldb's 2.11.0 as older than the 2:4.17.4+dfsg-1 that fixed CVE-2022-44640.
Reporting the source without its version is worse than reporting neither,
because it looks usable.

Measured on three live Proxmox nodes: every one of their 2 349 packages
was in that state — 802 of 802, 774 of 774, 773 of 773 — while twenty
non-Proxmox hosts had both fields. It was not a parsing bug. The
dpkg-query format string never asked for ${source:Version}, so nothing
downstream could have recovered it.

Now asked for and reported, with the same fallback napalm-linux uses:
dpkg leaves the field empty when it equals Version, and an older dpkg
leaves it empty because it does not know the field at all. Neither may
produce a package without a coordinate.

tests/test_packages.py covers all of it, including that the *query* names
the field — the assertion that would have caught this.
2026-09-20 17:23:03 +02:00
Christian Manivong 20fcf2ebc3 fix: four real defects the fourteen failing tests were pointing at
Closes netork#115.

The suite had been red long enough that it stopped being read. Four of the
fourteen failures were the tests being right.

`interfaces_mixin.py` used `re.match` without importing `re`, so
`get_mac_address_table` raised NameError against any node with a Linux bridge.
The tests never reached that line: they mocked the API call underneath
`_exec_ssh_command`, which takes two positional arguments where the doubles
accepted one, and which base64-wraps the command — so a fixture keyed on
"bridge fdb" appearing in the text matched nothing and the helper returned "".
They mock `_exec_ssh_command` itself now, which is the driver's own seam.

`is_alive` called `_resolve_node()`, which returns early without touching the
API whenever a node was configured through optional_args. A dead connection
reported itself alive. It probes `GET /version` now.

The documented `realm` optional_arg was read into `self._realm` in `__init__`
and then never used. Proxmox authenticates against "<user>@<realm>" and rejects
a bare username, so the option had no effect and callers had to know to type the
realm themselves.

`get_vlans` filtered out entries with no member ports on one return path while
the OVS path returned them, so a configured SDN VNet was visible or invisible
depending on which branch ran. A VNet exists on the node whether or not anything
is attached to it, and netOrk's VLAN discovery reads this.

`get_ipv6_neighbors_table` was simply missing and fell through to NAPALM's stub;
it is implemented against `ip -6 neigh show`, dropping FAILED entries.

The rest were stale tests. The DNS fixture put an FQDN where a search domain
belongs, which made `get_facts` build "pve1.pve1.example.com" and look like a
driver bug. The LLDP fixture was a simplified shape that real `lldpcli show
neighbors summary` does not produce — the parser matches on the ", via: LLDP"
that follows the interface name. And `test_bridge_vlan_show_parsing` covered a
fallback that was replaced by VM-config scanning, asserting an "interfaces" key
this method has never returned; it is now a test of the fallback that exists.
2026-08-21 13:22:03 +07:00
Christian Manivong 38f0c0a656 fix(vm_provision_mixin): stale same-named cloud image cache causes checksum mismatch
_download_cloud_image() cached downloaded images under just the URL's
basename (e.g. ubuntu-26.04-server-cloudimg-amd64.img). Ubuntu's per-build
download URLs change daily under that same stable basename
(.../release-20260713/... vs .../release-20260714/...), so a previous
day's cached file satisfied the "already cached" check and got checksum-
verified against the *new* day's expected hash from NetOrk's daily catalog
sync — failing outright and aborting the whole provisioning job, even
though a plain retry would have re-downloaded and succeeded (the bad file
was already being deleted on mismatch, just never re-fetched).

Found live during a NetOrk deploy: "Checksum mismatch for
https://cloud-images.ubuntu.com/.../release-20260713/
ubuntu-26.04-server-cloudimg-amd64.img: expected 0826c500..., got
3ee4f67f...".

Fix: key the cache path on a hash of the full URL (not just the
basename), and retry the download once after a checksum-mismatch cleanup
before raising.
2026-07-14 16:23:13 +02:00
Christian ManivongandClaude Sonnet 5 3181ade728 fix(vm_provision_mixin): destroy_vm's delete call rejected by Proxmox (400)
Passed destroy_unreferenced_disks (underscore) as a kwarg to proxmoxer's
delete(), but Proxmox's actual DELETE /nodes/{node}/qemu/{vmid} parameter
is hyphenated (destroy-unreferenced-disks). proxmoxer forwards kwargs to
the request verbatim with no underscore-to-hyphen translation, so Proxmox
rejected every call with "property is not defined in schema" before ever
touching the VM — the VM stayed fully intact (config, disks) despite the
caller believing destroy had at least been attempted. Fixed by building
the params as a dict (bypassing the Python-identifier restriction) with
the correct hyphenated key.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 19:54:05 +02:00
Christian Manivong 7038494b49 fix(vm_provision_mixin): get_vm_status hit a non-existent guest-agent path
agent.network_get_interfaces.get() built the URL path segment literally
("network_get_interfaces"), but the real Proxmox REST endpoint uses
hyphens ("network-get-interfaces") and must be reached via agent(...) as
a callable resource — the underscored attribute path 404ed silently on
every poll, so wait_for_ip always ran out the full timeout even though
the guest agent was reporting the IP to Proxmox correctly the whole time.

Also stopped assuming interfaces[0] is the real NIC — the guest agent
commonly reports "lo" first, matching the working pattern already used
in vm_mixin.py (skip "lo", require ip-address-type == "ipv4").
2026-07-08 14:09:40 +02:00
Christian Manivong 40d36b4b99 feat(vm_provision_mixin): set agent=1 on VM create for guest-agent channel
Proxmox only opens the virtio-serial channel qemu-guest-agent needs when
agent=1 is set at VM creation — without it, the agent package can be
installed but never actually reachable.
2026-07-08 12:56:09 +02:00
Christian Manivong bcadd77420 fix(vm_provision_mixin): cloud-init drive (ide2) needs images storage, not snippets
Proxmox's cloud-init drive is a disk image and requires a storage with
content='images' — the same requirement as the root disk — not the
snippets storage. These are commonly different storages (e.g. 'local'
with content=snippets-only, 'local-zfs' with content=images), and real
Proxmox now creates the VM fine but fails at *start* time with "storage
'X' does not support content-type 'images'" once it tries to generate
the cloud-init ISO.

Found live: a real deployment created the VM successfully, and only
failed when the user started it manually on the Proxmox side.
2026-07-08 11:07:06 +02:00
Christian Manivong 1d6aabb5a1 feat(vm_provision_mixin): pin explicit NIC MAC when given
nics[i]['mac'] is set via virtio=<mac>,bridge=... instead of the bare
virtio,bridge=... form, so a caller-supplied MAC actually takes effect
(needed for DHCP reservations created before the VM exists).
2026-07-08 09:09:28 +02:00
Christian Manivong 685d9b67ae fix(vm_provision_mixin): write Cloud-Init snippet via SSH, not the upload API
Real Proxmox's POST /nodes/{node}/storage/{storage}/upload only accepts
content in {iso, vztmpl, import} — content='snippets' is rejected
outright with a 400 ("does not have a value in the enumeration").
Snippets can only be written directly to the storage's filesystem path.

Found live, right after the previous multipart-upload fix: the VM
shell, disk import, and node-scoped storage selection all succeeded,
then create_vm_from_cloud_init failed with a 400 at the snippet write
step. Resolves the storage's path via the cluster storage config and
writes the file over SSH (base64-piped, to survive arbitrary YAML
content safely).
2026-07-07 23:24:05 +02:00
Christian Manivong 80d9c7335f fix(vm_provision_mixin): upload Cloud-Init snippet as a real multipart file
proxmoxer only builds a multipart request for io.IOBase values passed
as kwargs; a plain filename string (plus a nonexistent "data" field,
as the old code sent) goes out as an ordinary form-urlencoded POST
instead. Real Proxmox's /storage/{s}/upload endpoint expects an actual
file upload for "filename" and responds to anything else by closing
the connection with no HTTP response at all.

Found live: the VM shell, disk import, and node-scoped storage
selection all succeeded, then create_vm_from_cloud_init failed with
requests.exceptions.ConnectionError / RemoteDisconnected right at the
snippet upload step.
2026-07-07 23:01:27 +02:00
Christian Manivong 4d568bc6dc fix(vm_provision_mixin): query node-scoped storage, not cluster-wide
The cluster-wide /storage endpoint lists every storage regardless of
its "nodes" restriction, so _find_default_image_storage (and the
snippet-storage lookup) could pick a storage not actually available on
the node the VM is being created on. On a real server this stranded a
freshly-created VM shell with no disk attached: "qm importdisk" failed
with "storage 'local-lvm' is not available on node 'pve-02'" after the
VM (VMID 103) already existed. Querying /nodes/{node}/storage instead
fixes this, since Proxmox itself only lists what's available there.

Also adds get_image_storages() and an optional storage= override on
create_vm_from_cloud_init, so callers aren't stuck with auto-detection.
2026-07-07 22:36:14 +02:00
Christian Manivong 12135735cb fix(vm_provision_mixin): storage 'enabled' absent means enabled, not disabled
Proxmox's /storage API omits the "enabled" key entirely for storages that
were never explicitly toggled, rather than defaulting it to 1 — it isn't
present-and-falsy, it's just absent. Both _find_default_image_storage and
the snippet-storage discovery treated storage.get("enabled") as truthy-check,
so every storage without an explicit "enabled": 1 was silently excluded.

Confirmed live against a real Proxmox test server: local-lvm, local-zfs, and
fast-zfs all had content=images with no "enabled" key at all, causing
create_vm_from_cloud_init to always fail with "No storage with
content='images' found" despite multiple valid storages existing. All prior
tests used "enabled": 1 explicitly in their fixtures, masking the bug.

Fix: storage.get("enabled", 1) != 0 — absent or truthy means enabled, only
an explicit 0 excludes it. 4 new regression tests, 28 total pass.
2026-07-07 12:12:44 +02:00
Christian Manivong 9264cdcba9 feat(vm_provision_mixin): create_vm_from_cloud_init downloads cloud images directly
Replaces the template-clone flow with: create empty VM shell, download the
cloud image on the node (cached by filename, optional checksum verification),
qm importdisk, attach as scsi0. NIC config, snippet upload, ssh keys, disk
resize, and start remain unchanged (already generic).

New helpers: _run_node_command (strict SSH exec with custom timeout and
non-zero-exit detection, unlike the best-effort _exec_ssh_command),
_download_cloud_image (idempotent download + checksum check),
_find_default_image_storage (content=images discovery, mirrors the existing
snippet-storage discovery).

24 tests pass (10 new: _run_node_command x2, _download_cloud_image x4, plus
rewrites of the 4 existing create_vm_from_cloud_init tests for the new flow).
2026-07-07 10:37:36 +02:00
Christian Manivong ddd4e6fc03 feat(vm_provision_mixin): expose fixed_vlan_tag for SDN vnets
vnet's SDN tag (VLAN ID) is now surfaced in get_network_targets() output
instead of being silently discarded. Bridges never set this field.
2026-07-07 10:22:56 +02:00
Christian Manivong c7289fa674 feat(vm_provision_mixin): implement get_network_targets()
Filters _get_node_network() to bridge/OVSBridge types only (excludes physical
NICs, bonds), plus SDN vnets from _get_sdn_vnets(). vlan_aware: Linux bridge
reflects its bridge_vlan_aware config flag; OVS bridge always true; SDN vnet
always false (VLAN already fixed by the vnet's zone/tag).

4 new tests: bridge/vnet filtering, Linux bridge vlan_aware flag, OVS bridge
always vlan_aware, SDN vnet never vlan_aware. All 15 tests in the file pass.
2026-07-07 09:08:21 +02:00
Christian Manivong a5a5b634b0 Reapply "Merge feature/generic-vm-provisioning: generalize vm_provision_mixin for arbitrary NIC configs"
This reverts commit 6ede48d244.
2026-07-07 08:21:00 +02:00
Christian Manivong 6ede48d244 Revert "Merge feature/generic-vm-provisioning: generalize vm_provision_mixin for arbitrary NIC configs"
This reverts commit 1d2006f9fb, reversing
changes made to 7bdac4c496.
2026-07-07 00:53:13 +02:00
Christian ManivongandClaude Haiku 4.5 ae23208eac test(vm_provision): cover generic NIC config, dual-NIC trunk, disk resize
Test cases: single/dual NIC with VLAN tags or trunk config, per-NIC DHCP control,
disk resize parameter, snippet storage validation, IP wait timeout, destroy paths.
All TDD cases green.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-07-07 00:00:09 +02:00
Christian ManivongandClaude Haiku 4.5 7bdac4c496 feat(provisioning): implement VM provisioning mixin for Proxmox
Add ProxmoxVMProvisionMixin with three methods:
- create_vm_from_cloud_init(): clone template → dual-NIC config → Cloud-Init → start
- destroy_vm(): stop → delete VM → cleanup snippets
- get_vm_status(): poll guest-agent for IP with optional wait-for-IP polling

Tests (9 cases):
- _wait_for_task success/error/timeout handling
- create_vm happy path + missing snippet storage error
- get_vm_status with/without wait-for-IP, timeout handling
- destroy_vm on running or already-stopped VM

All tests pass (100% coverage on mixin code paths).

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-07-06 22:03:41 +02:00
Christian Manivong f3ecf14c8d initial commit 2026-05-29 09:24:39 +02:00