_resolve_node() took the first entry of GET /nodes. In a cluster that
lists every member, so a node polled without an explicit `node` driver
argument talked to whichever member came first: pve-dual reported
pve-02's name and VMs, and netOrk's VM sync moved pve-02's VM devices
over to it.
Resolve through GET /cluster/status instead: the entry marked local,
then a match by IP or (short) name, then the sole node of a standalone
host, and otherwise raise rather than guess.
The lookup no longer swallows API errors either. A TLS verification
failure used to leave the IP as the node name, so open() succeeded and
every getter failed quietly while the poll reported success with empty
data. It now surfaces as a ConnectionException from open().
Refs NetOrk/netork#417, NetOrk/netork#418
get_vm_snapshots, create_vm_snapshot, delete_vm_snapshot and
rollback_vm_snapshot for VMs and containers, so netOrk's snapshot view
works on Proxmox as it does on VMware. Proxmox lists the live state as a
pseudo-snapshot named "current"; it is never reported or addressable.
Containers have no RAM state, so include_memory is ignored for them.
reboot_host() restarts the node with POST /nodes/{node}/status
command=reboot instead of /sbin/reboot over SSH.
start_vm, stop_vm, reboot_vm, suspend_vm and get_vm_config existed only
as declarations. netOrk called Proxmox's own power_vm and read a VM's
raw config through _node_api(), so no other hypervisor could serve the
same endpoints. These let netOrk talk to every hypervisor alike.
The power methods accept a VM's name or vmid, wait for the Proxmox task,
and raise ValueError/RuntimeError as the contract says instead of
returning a result dict. A forced reboot of a container is stop + start,
since LXC has no reset; suspending a container is refused. power_vm is
unchanged for existing callers.
get_vm_config moves the config parsing netOrk did in
_parse_proxmox_hw_config into the driver and returns a VMConfigDict:
disks with storage and size, NICs with model, MAC, bridge and VLAN, CPU
topology, firmware, machine type and PCI/USB passthrough.
get_vms reports vmid as a string ("100"), following
napalm-device-types 2.0, still ordered numerically.
`get_packages` named the Debian source package and never its version, so a
consumer was handed two numbers on different axes and no way to tell.
OSV states Debian ranges in *source* versions. libldb2 is
2:2.11.0+samba4.22.11+dfsg-… while its source, samba, is 2:4.22.11+dfsg-…;
comparing the first against a samba range is meaningless, and dpkg reads
ldb's 2.11.0 as older than the 2:4.17.4+dfsg-1 that fixed CVE-2022-44640.
Reporting the source without its version is worse than reporting neither,
because it looks usable.
Measured on three live Proxmox nodes: every one of their 2 349 packages
was in that state — 802 of 802, 774 of 774, 773 of 773 — while twenty
non-Proxmox hosts had both fields. It was not a parsing bug. The
dpkg-query format string never asked for ${source:Version}, so nothing
downstream could have recovered it.
Now asked for and reported, with the same fallback napalm-linux uses:
dpkg leaves the field empty when it equals Version, and an older dpkg
leaves it empty because it does not know the field at all. Neither may
produce a package without a coordinate.
tests/test_packages.py covers all of it, including that the *query* names
the field — the assertion that would have caught this.
Closes netork#115.
The suite had been red long enough that it stopped being read. Four of the
fourteen failures were the tests being right.
`interfaces_mixin.py` used `re.match` without importing `re`, so
`get_mac_address_table` raised NameError against any node with a Linux bridge.
The tests never reached that line: they mocked the API call underneath
`_exec_ssh_command`, which takes two positional arguments where the doubles
accepted one, and which base64-wraps the command — so a fixture keyed on
"bridge fdb" appearing in the text matched nothing and the helper returned "".
They mock `_exec_ssh_command` itself now, which is the driver's own seam.
`is_alive` called `_resolve_node()`, which returns early without touching the
API whenever a node was configured through optional_args. A dead connection
reported itself alive. It probes `GET /version` now.
The documented `realm` optional_arg was read into `self._realm` in `__init__`
and then never used. Proxmox authenticates against "<user>@<realm>" and rejects
a bare username, so the option had no effect and callers had to know to type the
realm themselves.
`get_vlans` filtered out entries with no member ports on one return path while
the OVS path returned them, so a configured SDN VNet was visible or invisible
depending on which branch ran. A VNet exists on the node whether or not anything
is attached to it, and netOrk's VLAN discovery reads this.
`get_ipv6_neighbors_table` was simply missing and fell through to NAPALM's stub;
it is implemented against `ip -6 neigh show`, dropping FAILED entries.
The rest were stale tests. The DNS fixture put an FQDN where a search domain
belongs, which made `get_facts` build "pve1.pve1.example.com" and look like a
driver bug. The LLDP fixture was a simplified shape that real `lldpcli show
neighbors summary` does not produce — the parser matches on the ", via: LLDP"
that follows the interface name. And `test_bridge_vlan_show_parsing` covered a
fallback that was replaced by VM-config scanning, asserting an "interfaces" key
this method has never returned; it is now a test of the fallback that exists.
_download_cloud_image() cached downloaded images under just the URL's
basename (e.g. ubuntu-26.04-server-cloudimg-amd64.img). Ubuntu's per-build
download URLs change daily under that same stable basename
(.../release-20260713/... vs .../release-20260714/...), so a previous
day's cached file satisfied the "already cached" check and got checksum-
verified against the *new* day's expected hash from NetOrk's daily catalog
sync — failing outright and aborting the whole provisioning job, even
though a plain retry would have re-downloaded and succeeded (the bad file
was already being deleted on mismatch, just never re-fetched).
Found live during a NetOrk deploy: "Checksum mismatch for
https://cloud-images.ubuntu.com/.../release-20260713/
ubuntu-26.04-server-cloudimg-amd64.img: expected 0826c500..., got
3ee4f67f...".
Fix: key the cache path on a hash of the full URL (not just the
basename), and retry the download once after a checksum-mismatch cleanup
before raising.
Passed destroy_unreferenced_disks (underscore) as a kwarg to proxmoxer's
delete(), but Proxmox's actual DELETE /nodes/{node}/qemu/{vmid} parameter
is hyphenated (destroy-unreferenced-disks). proxmoxer forwards kwargs to
the request verbatim with no underscore-to-hyphen translation, so Proxmox
rejected every call with "property is not defined in schema" before ever
touching the VM — the VM stayed fully intact (config, disks) despite the
caller believing destroy had at least been attempted. Fixed by building
the params as a dict (bypassing the Python-identifier restriction) with
the correct hyphenated key.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
agent.network_get_interfaces.get() built the URL path segment literally
("network_get_interfaces"), but the real Proxmox REST endpoint uses
hyphens ("network-get-interfaces") and must be reached via agent(...) as
a callable resource — the underscored attribute path 404ed silently on
every poll, so wait_for_ip always ran out the full timeout even though
the guest agent was reporting the IP to Proxmox correctly the whole time.
Also stopped assuming interfaces[0] is the real NIC — the guest agent
commonly reports "lo" first, matching the working pattern already used
in vm_mixin.py (skip "lo", require ip-address-type == "ipv4").
Proxmox only opens the virtio-serial channel qemu-guest-agent needs when
agent=1 is set at VM creation — without it, the agent package can be
installed but never actually reachable.
Proxmox's cloud-init drive is a disk image and requires a storage with
content='images' — the same requirement as the root disk — not the
snippets storage. These are commonly different storages (e.g. 'local'
with content=snippets-only, 'local-zfs' with content=images), and real
Proxmox now creates the VM fine but fails at *start* time with "storage
'X' does not support content-type 'images'" once it tries to generate
the cloud-init ISO.
Found live: a real deployment created the VM successfully, and only
failed when the user started it manually on the Proxmox side.
nics[i]['mac'] is set via virtio=<mac>,bridge=... instead of the bare
virtio,bridge=... form, so a caller-supplied MAC actually takes effect
(needed for DHCP reservations created before the VM exists).
Real Proxmox's POST /nodes/{node}/storage/{storage}/upload only accepts
content in {iso, vztmpl, import} — content='snippets' is rejected
outright with a 400 ("does not have a value in the enumeration").
Snippets can only be written directly to the storage's filesystem path.
Found live, right after the previous multipart-upload fix: the VM
shell, disk import, and node-scoped storage selection all succeeded,
then create_vm_from_cloud_init failed with a 400 at the snippet write
step. Resolves the storage's path via the cluster storage config and
writes the file over SSH (base64-piped, to survive arbitrary YAML
content safely).
proxmoxer only builds a multipart request for io.IOBase values passed
as kwargs; a plain filename string (plus a nonexistent "data" field,
as the old code sent) goes out as an ordinary form-urlencoded POST
instead. Real Proxmox's /storage/{s}/upload endpoint expects an actual
file upload for "filename" and responds to anything else by closing
the connection with no HTTP response at all.
Found live: the VM shell, disk import, and node-scoped storage
selection all succeeded, then create_vm_from_cloud_init failed with
requests.exceptions.ConnectionError / RemoteDisconnected right at the
snippet upload step.
The cluster-wide /storage endpoint lists every storage regardless of
its "nodes" restriction, so _find_default_image_storage (and the
snippet-storage lookup) could pick a storage not actually available on
the node the VM is being created on. On a real server this stranded a
freshly-created VM shell with no disk attached: "qm importdisk" failed
with "storage 'local-lvm' is not available on node 'pve-02'" after the
VM (VMID 103) already existed. Querying /nodes/{node}/storage instead
fixes this, since Proxmox itself only lists what's available there.
Also adds get_image_storages() and an optional storage= override on
create_vm_from_cloud_init, so callers aren't stuck with auto-detection.
Proxmox's /storage API omits the "enabled" key entirely for storages that
were never explicitly toggled, rather than defaulting it to 1 — it isn't
present-and-falsy, it's just absent. Both _find_default_image_storage and
the snippet-storage discovery treated storage.get("enabled") as truthy-check,
so every storage without an explicit "enabled": 1 was silently excluded.
Confirmed live against a real Proxmox test server: local-lvm, local-zfs, and
fast-zfs all had content=images with no "enabled" key at all, causing
create_vm_from_cloud_init to always fail with "No storage with
content='images' found" despite multiple valid storages existing. All prior
tests used "enabled": 1 explicitly in their fixtures, masking the bug.
Fix: storage.get("enabled", 1) != 0 — absent or truthy means enabled, only
an explicit 0 excludes it. 4 new regression tests, 28 total pass.
Replaces the template-clone flow with: create empty VM shell, download the
cloud image on the node (cached by filename, optional checksum verification),
qm importdisk, attach as scsi0. NIC config, snippet upload, ssh keys, disk
resize, and start remain unchanged (already generic).
New helpers: _run_node_command (strict SSH exec with custom timeout and
non-zero-exit detection, unlike the best-effort _exec_ssh_command),
_download_cloud_image (idempotent download + checksum check),
_find_default_image_storage (content=images discovery, mirrors the existing
snippet-storage discovery).
24 tests pass (10 new: _run_node_command x2, _download_cloud_image x4, plus
rewrites of the 4 existing create_vm_from_cloud_init tests for the new flow).
Filters _get_node_network() to bridge/OVSBridge types only (excludes physical
NICs, bonds), plus SDN vnets from _get_sdn_vnets(). vlan_aware: Linux bridge
reflects its bridge_vlan_aware config flag; OVS bridge always true; SDN vnet
always false (VLAN already fixed by the vnet's zone/tag).
4 new tests: bridge/vnet filtering, Linux bridge vlan_aware flag, OVS bridge
always vlan_aware, SDN vnet never vlan_aware. All 15 tests in the file pass.
Test cases: single/dual NIC with VLAN tags or trunk config, per-NIC DHCP control,
disk resize parameter, snippet storage validation, IP wait timeout, destroy paths.
All TDD cases green.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>