A VM created without a cpu argument gets Proxmox's API default, kvm64: no
AES-NI and no AVX. MongoDB 5.0 and later exit with "Illegal instruction" on
it, which is how a Graylog provisioned through netOrk failed (netork#494).
Every VM built by hand in the same cluster uses host or x86-64-v2-AES; only
the ones this driver created were left on kvm64.
create_vm_from_cloud_init now always passes cpu=, defaulting to x86-64-v2-AES
as the Proxmox GUI does since PVE 8, and takes cpu_type for a different one.
get_vm_cpu_types() lists x86-64-v2-AES (default), x86-64-v3 and host. The
named models carry the cpuinfo flags they add over qemu64, following
Proxmox's own definitions (CPUConfig.pm), and are available when the node's
CPU has all of them, read from /nodes/{node}/status. host lists the node's
own flags, so on a CPU without AVX it does not pretend to offer any.
A model asked for by name is checked against the node before a VMID is
allocated: one the CPU cannot run would only fail at VM start, after the disk
import, leaving a half-built VM behind. The default is not checked, so a
plain create costs no extra API call. VMCpuTypeDict is imported for type
checking only, so this works with an older napalm_device_types.
Provisioning downloaded every cloud image over SSH into
/var/lib/vz/template/netork-images, on the node's root filesystem. On a
small root that fills up and takes Proxmox down with it (netOrk #480).
When the node has an active storage with content type "import" (Proxmox
8.2+), Proxmox now does it itself: download-url with checksum
verification into that storage, then import-from as the root disk. The
file is named after a hash of the full URL and reused when present.
Proxmox takes the format from the extension and has no ".img", so
Ubuntu's qcow2 .img is stored as .qcow2 -- a wrong guess fails at import
instead of attaching a qcow2 container as a raw disk.
Without an import storage, or for an image type Proxmox cannot import,
the SSH download is used as before.
_download_cloud_image() cached downloaded images under just the URL's
basename (e.g. ubuntu-26.04-server-cloudimg-amd64.img). Ubuntu's per-build
download URLs change daily under that same stable basename
(.../release-20260713/... vs .../release-20260714/...), so a previous
day's cached file satisfied the "already cached" check and got checksum-
verified against the *new* day's expected hash from NetOrk's daily catalog
sync — failing outright and aborting the whole provisioning job, even
though a plain retry would have re-downloaded and succeeded (the bad file
was already being deleted on mismatch, just never re-fetched).
Found live during a NetOrk deploy: "Checksum mismatch for
https://cloud-images.ubuntu.com/.../release-20260713/
ubuntu-26.04-server-cloudimg-amd64.img: expected 0826c500..., got
3ee4f67f...".
Fix: key the cache path on a hash of the full URL (not just the
basename), and retry the download once after a checksum-mismatch cleanup
before raising.
Passed destroy_unreferenced_disks (underscore) as a kwarg to proxmoxer's
delete(), but Proxmox's actual DELETE /nodes/{node}/qemu/{vmid} parameter
is hyphenated (destroy-unreferenced-disks). proxmoxer forwards kwargs to
the request verbatim with no underscore-to-hyphen translation, so Proxmox
rejected every call with "property is not defined in schema" before ever
touching the VM — the VM stayed fully intact (config, disks) despite the
caller believing destroy had at least been attempted. Fixed by building
the params as a dict (bypassing the Python-identifier restriction) with
the correct hyphenated key.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
agent.network_get_interfaces.get() built the URL path segment literally
("network_get_interfaces"), but the real Proxmox REST endpoint uses
hyphens ("network-get-interfaces") and must be reached via agent(...) as
a callable resource — the underscored attribute path 404ed silently on
every poll, so wait_for_ip always ran out the full timeout even though
the guest agent was reporting the IP to Proxmox correctly the whole time.
Also stopped assuming interfaces[0] is the real NIC — the guest agent
commonly reports "lo" first, matching the working pattern already used
in vm_mixin.py (skip "lo", require ip-address-type == "ipv4").
Proxmox only opens the virtio-serial channel qemu-guest-agent needs when
agent=1 is set at VM creation — without it, the agent package can be
installed but never actually reachable.
Proxmox's cloud-init drive is a disk image and requires a storage with
content='images' — the same requirement as the root disk — not the
snippets storage. These are commonly different storages (e.g. 'local'
with content=snippets-only, 'local-zfs' with content=images), and real
Proxmox now creates the VM fine but fails at *start* time with "storage
'X' does not support content-type 'images'" once it tries to generate
the cloud-init ISO.
Found live: a real deployment created the VM successfully, and only
failed when the user started it manually on the Proxmox side.
nics[i]['mac'] is set via virtio=<mac>,bridge=... instead of the bare
virtio,bridge=... form, so a caller-supplied MAC actually takes effect
(needed for DHCP reservations created before the VM exists).
Real Proxmox's POST /nodes/{node}/storage/{storage}/upload only accepts
content in {iso, vztmpl, import} — content='snippets' is rejected
outright with a 400 ("does not have a value in the enumeration").
Snippets can only be written directly to the storage's filesystem path.
Found live, right after the previous multipart-upload fix: the VM
shell, disk import, and node-scoped storage selection all succeeded,
then create_vm_from_cloud_init failed with a 400 at the snippet write
step. Resolves the storage's path via the cluster storage config and
writes the file over SSH (base64-piped, to survive arbitrary YAML
content safely).
proxmoxer only builds a multipart request for io.IOBase values passed
as kwargs; a plain filename string (plus a nonexistent "data" field,
as the old code sent) goes out as an ordinary form-urlencoded POST
instead. Real Proxmox's /storage/{s}/upload endpoint expects an actual
file upload for "filename" and responds to anything else by closing
the connection with no HTTP response at all.
Found live: the VM shell, disk import, and node-scoped storage
selection all succeeded, then create_vm_from_cloud_init failed with
requests.exceptions.ConnectionError / RemoteDisconnected right at the
snippet upload step.
The cluster-wide /storage endpoint lists every storage regardless of
its "nodes" restriction, so _find_default_image_storage (and the
snippet-storage lookup) could pick a storage not actually available on
the node the VM is being created on. On a real server this stranded a
freshly-created VM shell with no disk attached: "qm importdisk" failed
with "storage 'local-lvm' is not available on node 'pve-02'" after the
VM (VMID 103) already existed. Querying /nodes/{node}/storage instead
fixes this, since Proxmox itself only lists what's available there.
Also adds get_image_storages() and an optional storage= override on
create_vm_from_cloud_init, so callers aren't stuck with auto-detection.
Proxmox's /storage API omits the "enabled" key entirely for storages that
were never explicitly toggled, rather than defaulting it to 1 — it isn't
present-and-falsy, it's just absent. Both _find_default_image_storage and
the snippet-storage discovery treated storage.get("enabled") as truthy-check,
so every storage without an explicit "enabled": 1 was silently excluded.
Confirmed live against a real Proxmox test server: local-lvm, local-zfs, and
fast-zfs all had content=images with no "enabled" key at all, causing
create_vm_from_cloud_init to always fail with "No storage with
content='images' found" despite multiple valid storages existing. All prior
tests used "enabled": 1 explicitly in their fixtures, masking the bug.
Fix: storage.get("enabled", 1) != 0 — absent or truthy means enabled, only
an explicit 0 excludes it. 4 new regression tests, 28 total pass.
Replaces the template-clone flow with: create empty VM shell, download the
cloud image on the node (cached by filename, optional checksum verification),
qm importdisk, attach as scsi0. NIC config, snippet upload, ssh keys, disk
resize, and start remain unchanged (already generic).
New helpers: _run_node_command (strict SSH exec with custom timeout and
non-zero-exit detection, unlike the best-effort _exec_ssh_command),
_download_cloud_image (idempotent download + checksum check),
_find_default_image_storage (content=images discovery, mirrors the existing
snippet-storage discovery).
24 tests pass (10 new: _run_node_command x2, _download_cloud_image x4, plus
rewrites of the 4 existing create_vm_from_cloud_init tests for the new flow).
Filters _get_node_network() to bridge/OVSBridge types only (excludes physical
NICs, bonds), plus SDN vnets from _get_sdn_vnets(). vlan_aware: Linux bridge
reflects its bridge_vlan_aware config flag; OVS bridge always true; SDN vnet
always false (VLAN already fixed by the vnet's zone/tag).
4 new tests: bridge/vnet filtering, Linux bridge vlan_aware flag, OVS bridge
always vlan_aware, SDN vnet never vlan_aware. All 15 tests in the file pass.
- Replace fixed mgmt/capture dual-NIC parameters with generic nics list
- Loop over NICs to build net0, net1, ... config strings (access VLAN or trunk)
- Remove hardcoded 'ip link set eth1 up' Cloud-Init hack (caller responsibility)
- Add disk_resize_gb parameter for post-clone disk expansion (scsi0/virtio0/ide0/sata0)
- Make DHCP configuration per-NIC with sensible defaults (primary NIC only)
- Update docstrings and logging to reflect generic NIC architecture