Provisioning downloaded every cloud image over SSH into
/var/lib/vz/template/netork-images, on the node's root filesystem. On a
small root that fills up and takes Proxmox down with it (netOrk #480).
When the node has an active storage with content type "import" (Proxmox
8.2+), Proxmox now does it itself: download-url with checksum
verification into that storage, then import-from as the root disk. The
file is named after a hash of the full URL and reused when present.
Proxmox takes the format from the extension and has no ".img", so
Ubuntu's qcow2 .img is stored as .qcow2 -- a wrong guess fails at import
instead of attaching a qcow2 container as a raw disk.
Without an import storage, or for an image type Proxmox cannot import,
the SSH download is used as before.
_resolve_node() took the first entry of GET /nodes. In a cluster that
lists every member, so a node polled without an explicit `node` driver
argument talked to whichever member came first: pve-dual reported
pve-02's name and VMs, and netOrk's VM sync moved pve-02's VM devices
over to it.
Resolve through GET /cluster/status instead: the entry marked local,
then a match by IP or (short) name, then the sole node of a standalone
host, and otherwise raise rather than guess.
The lookup no longer swallows API errors either. A TLS verification
failure used to leave the IP as the node name, so open() succeeded and
every getter failed quietly while the poll reported success with empty
data. It now surfaces as a ConnectionException from open().
Refs NetOrk/netork#417, NetOrk/netork#418
get_vm_snapshots, create_vm_snapshot, delete_vm_snapshot and
rollback_vm_snapshot for VMs and containers, so netOrk's snapshot view
works on Proxmox as it does on VMware. Proxmox lists the live state as a
pseudo-snapshot named "current"; it is never reported or addressable.
Containers have no RAM state, so include_memory is ignored for them.
reboot_host() restarts the node with POST /nodes/{node}/status
command=reboot instead of /sbin/reboot over SSH.
start_vm, stop_vm, reboot_vm, suspend_vm and get_vm_config existed only
as declarations. netOrk called Proxmox's own power_vm and read a VM's
raw config through _node_api(), so no other hypervisor could serve the
same endpoints. These let netOrk talk to every hypervisor alike.
The power methods accept a VM's name or vmid, wait for the Proxmox task,
and raise ValueError/RuntimeError as the contract says instead of
returning a result dict. A forced reboot of a container is stop + start,
since LXC has no reset; suspending a container is refused. power_vm is
unchanged for existing callers.
get_vm_config moves the config parsing netOrk did in
_parse_proxmox_hw_config into the driver and returns a VMConfigDict:
disks with storage and size, NICs with model, MAC, bridge and VLAN, CPU
topology, firmware, machine type and PCI/USB passthrough.
get_vms reports vmid as a string ("100"), following
napalm-device-types 2.0, still ordered numerically.
`get_packages` named the Debian source package and never its version, so a
consumer was handed two numbers on different axes and no way to tell.
OSV states Debian ranges in *source* versions. libldb2 is
2:2.11.0+samba4.22.11+dfsg-… while its source, samba, is 2:4.22.11+dfsg-…;
comparing the first against a samba range is meaningless, and dpkg reads
ldb's 2.11.0 as older than the 2:4.17.4+dfsg-1 that fixed CVE-2022-44640.
Reporting the source without its version is worse than reporting neither,
because it looks usable.
Measured on three live Proxmox nodes: every one of their 2 349 packages
was in that state — 802 of 802, 774 of 774, 773 of 773 — while twenty
non-Proxmox hosts had both fields. It was not a parsing bug. The
dpkg-query format string never asked for ${source:Version}, so nothing
downstream could have recovered it.
Now asked for and reported, with the same fallback napalm-linux uses:
dpkg leaves the field empty when it equals Version, and an older dpkg
leaves it empty because it does not know the field at all. Neither may
produce a package without a coordinate.
tests/test_packages.py covers all of it, including that the *query* names
the field — the assertion that would have caught this.
dpkg-query now also reports ${source:Package} as source_package, so consumers
can match installed binaries to the correct Debian source (e.g. openssh-server
-> openssh) for accurate OSV vulnerability lookups.
Closes netork#115.
The suite had been red long enough that it stopped being read. Four of the
fourteen failures were the tests being right.
`interfaces_mixin.py` used `re.match` without importing `re`, so
`get_mac_address_table` raised NameError against any node with a Linux bridge.
The tests never reached that line: they mocked the API call underneath
`_exec_ssh_command`, which takes two positional arguments where the doubles
accepted one, and which base64-wraps the command — so a fixture keyed on
"bridge fdb" appearing in the text matched nothing and the helper returned "".
They mock `_exec_ssh_command` itself now, which is the driver's own seam.
`is_alive` called `_resolve_node()`, which returns early without touching the
API whenever a node was configured through optional_args. A dead connection
reported itself alive. It probes `GET /version` now.
The documented `realm` optional_arg was read into `self._realm` in `__init__`
and then never used. Proxmox authenticates against "<user>@<realm>" and rejects
a bare username, so the option had no effect and callers had to know to type the
realm themselves.
`get_vlans` filtered out entries with no member ports on one return path while
the OVS path returned them, so a configured SDN VNet was visible or invisible
depending on which branch ran. A VNet exists on the node whether or not anything
is attached to it, and netOrk's VLAN discovery reads this.
`get_ipv6_neighbors_table` was simply missing and fell through to NAPALM's stub;
it is implemented against `ip -6 neigh show`, dropping FAILED entries.
The rest were stale tests. The DNS fixture put an FQDN where a search domain
belongs, which made `get_facts` build "pve1.pve1.example.com" and look like a
driver bug. The LLDP fixture was a simplified shape that real `lldpcli show
neighbors summary` does not produce — the parser matches on the ", via: LLDP"
that follows the interface name. And `test_bridge_vlan_show_parsing` covered a
fallback that was replaced by VM-config scanning, asserting an "interfaces" key
this method has never returned; it is now a test of the fallback that exists.
Both were facts about this driver that netOrk kept in hardcoded driver-name
sets, each duplicated across a file pair (netork#113). The driver is the right
place for them: everything runs over the PVE REST API, and a node reboots
through a full init sequence plus storage checks before it is worth polling.
get_device_warnings() now returns only {code, meta} — severity, title,
message, and action are resolved centrally by netork's
WARNING_CATALOG (netork/core/device_warnings.py), not by the driver.
Keeps this driver independent of netork and avoids per-vendor drift in
how the same warning code is presented.
_download_cloud_image() cached downloaded images under just the URL's
basename (e.g. ubuntu-26.04-server-cloudimg-amd64.img). Ubuntu's per-build
download URLs change daily under that same stable basename
(.../release-20260713/... vs .../release-20260714/...), so a previous
day's cached file satisfied the "already cached" check and got checksum-
verified against the *new* day's expected hash from NetOrk's daily catalog
sync — failing outright and aborting the whole provisioning job, even
though a plain retry would have re-downloaded and succeeded (the bad file
was already being deleted on mismatch, just never re-fetched).
Found live during a NetOrk deploy: "Checksum mismatch for
https://cloud-images.ubuntu.com/.../release-20260713/
ubuntu-26.04-server-cloudimg-amd64.img: expected 0826c500..., got
3ee4f67f...".
Fix: key the cache path on a hash of the full URL (not just the
basename), and retry the download once after a checksum-mismatch cleanup
before raising.
Passed destroy_unreferenced_disks (underscore) as a kwarg to proxmoxer's
delete(), but Proxmox's actual DELETE /nodes/{node}/qemu/{vmid} parameter
is hyphenated (destroy-unreferenced-disks). proxmoxer forwards kwargs to
the request verbatim with no underscore-to-hyphen translation, so Proxmox
rejected every call with "property is not defined in schema" before ever
touching the VM — the VM stayed fully intact (config, disks) despite the
caller believing destroy had at least been attempted. Fixed by building
the params as a dict (bypassing the Python-identifier restriction) with
the correct hyphenated key.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
agent.network_get_interfaces.get() built the URL path segment literally
("network_get_interfaces"), but the real Proxmox REST endpoint uses
hyphens ("network-get-interfaces") and must be reached via agent(...) as
a callable resource — the underscored attribute path 404ed silently on
every poll, so wait_for_ip always ran out the full timeout even though
the guest agent was reporting the IP to Proxmox correctly the whole time.
Also stopped assuming interfaces[0] is the real NIC — the guest agent
commonly reports "lo" first, matching the working pattern already used
in vm_mixin.py (skip "lo", require ip-address-type == "ipv4").
Proxmox only opens the virtio-serial channel qemu-guest-agent needs when
agent=1 is set at VM creation — without it, the agent package can be
installed but never actually reachable.
Proxmox's cloud-init drive is a disk image and requires a storage with
content='images' — the same requirement as the root disk — not the
snippets storage. These are commonly different storages (e.g. 'local'
with content=snippets-only, 'local-zfs' with content=images), and real
Proxmox now creates the VM fine but fails at *start* time with "storage
'X' does not support content-type 'images'" once it tries to generate
the cloud-init ISO.
Found live: a real deployment created the VM successfully, and only
failed when the user started it manually on the Proxmox side.
nics[i]['mac'] is set via virtio=<mac>,bridge=... instead of the bare
virtio,bridge=... form, so a caller-supplied MAC actually takes effect
(needed for DHCP reservations created before the VM exists).
Real Proxmox's POST /nodes/{node}/storage/{storage}/upload only accepts
content in {iso, vztmpl, import} — content='snippets' is rejected
outright with a 400 ("does not have a value in the enumeration").
Snippets can only be written directly to the storage's filesystem path.
Found live, right after the previous multipart-upload fix: the VM
shell, disk import, and node-scoped storage selection all succeeded,
then create_vm_from_cloud_init failed with a 400 at the snippet write
step. Resolves the storage's path via the cluster storage config and
writes the file over SSH (base64-piped, to survive arbitrary YAML
content safely).
proxmoxer only builds a multipart request for io.IOBase values passed
as kwargs; a plain filename string (plus a nonexistent "data" field,
as the old code sent) goes out as an ordinary form-urlencoded POST
instead. Real Proxmox's /storage/{s}/upload endpoint expects an actual
file upload for "filename" and responds to anything else by closing
the connection with no HTTP response at all.
Found live: the VM shell, disk import, and node-scoped storage
selection all succeeded, then create_vm_from_cloud_init failed with
requests.exceptions.ConnectionError / RemoteDisconnected right at the
snippet upload step.
The cluster-wide /storage endpoint lists every storage regardless of
its "nodes" restriction, so _find_default_image_storage (and the
snippet-storage lookup) could pick a storage not actually available on
the node the VM is being created on. On a real server this stranded a
freshly-created VM shell with no disk attached: "qm importdisk" failed
with "storage 'local-lvm' is not available on node 'pve-02'" after the
VM (VMID 103) already existed. Querying /nodes/{node}/storage instead
fixes this, since Proxmox itself only lists what's available there.
Also adds get_image_storages() and an optional storage= override on
create_vm_from_cloud_init, so callers aren't stuck with auto-detection.
Proxmox's /storage API omits the "enabled" key entirely for storages that
were never explicitly toggled, rather than defaulting it to 1 — it isn't
present-and-falsy, it's just absent. Both _find_default_image_storage and
the snippet-storage discovery treated storage.get("enabled") as truthy-check,
so every storage without an explicit "enabled": 1 was silently excluded.
Confirmed live against a real Proxmox test server: local-lvm, local-zfs, and
fast-zfs all had content=images with no "enabled" key at all, causing
create_vm_from_cloud_init to always fail with "No storage with
content='images' found" despite multiple valid storages existing. All prior
tests used "enabled": 1 explicitly in their fixtures, masking the bug.
Fix: storage.get("enabled", 1) != 0 — absent or truthy means enabled, only
an explicit 0 excludes it. 4 new regression tests, 28 total pass.
Replaces the template-clone flow with: create empty VM shell, download the
cloud image on the node (cached by filename, optional checksum verification),
qm importdisk, attach as scsi0. NIC config, snippet upload, ssh keys, disk
resize, and start remain unchanged (already generic).
New helpers: _run_node_command (strict SSH exec with custom timeout and
non-zero-exit detection, unlike the best-effort _exec_ssh_command),
_download_cloud_image (idempotent download + checksum check),
_find_default_image_storage (content=images discovery, mirrors the existing
snippet-storage discovery).
24 tests pass (10 new: _run_node_command x2, _download_cloud_image x4, plus
rewrites of the 4 existing create_vm_from_cloud_init tests for the new flow).
Filters _get_node_network() to bridge/OVSBridge types only (excludes physical
NICs, bonds), plus SDN vnets from _get_sdn_vnets(). vlan_aware: Linux bridge
reflects its bridge_vlan_aware config flag; OVS bridge always true; SDN vnet
always false (VLAN already fixed by the vnet's zone/tag).
4 new tests: bridge/vnet filtering, Linux bridge vlan_aware flag, OVS bridge
always vlan_aware, SDN vnet never vlan_aware. All 15 tests in the file pass.
Test cases: single/dual NIC with VLAN tags or trunk config, per-NIC DHCP control,
disk resize parameter, snippet storage validation, IP wait timeout, destroy paths.
All TDD cases green.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
- Replace fixed mgmt/capture dual-NIC parameters with generic nics list
- Loop over NICs to build net0, net1, ... config strings (access VLAN or trunk)
- Remove hardcoded 'ip link set eth1 up' Cloud-Init hack (caller responsibility)
- Add disk_resize_gb parameter for post-clone disk expansion (scsi0/virtio0/ide0/sata0)
- Make DHCP configuration per-NIC with sensible defaults (primary NIC only)
- Update docstrings and logging to reflect generic NIC architecture
_get_vm_disk_and_boot() now returns a list of disk dicts (name, size_mb)
instead of a single total. Each disk entry is read from the VM config
via /qemu/{vmid}/config or /lxc/{vmid}/config. The onboot flag is also
read from the same config endpoint.
Both QEMU VMs and LXC containers are covered. The 'disks' and 'onboot'
keys are added to every entry in the vms_snapshot.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
product_version is used as model only when it contains a space, indicating
a human-readable marketing name (e.g. "ThinkCentre M910x"). Part numbers
like "J26843-409" have no space and are skipped — product_name is used
instead (e.g. "NUC6CAYH" for Intel NUC).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The combined printf approach silently produced empty values when
product_name/version contained special chars or the shell split tokens
incorrectly. Read each /sys/class/dmi/id/ file via a separate cat,
collect lines, then apply vendor-specific model name selection:
- Intel NUC: product_name='NUC6CAYH' (marketing) preferred over
product_version='J26843-409' (part number)
- Lenovo: product_name='10MYS03U00' (type code, all-caps+digits) →
prefer product_version='ThinkCentre M910x' (marketing name)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
On Lenovo (and some other vendors) product_name contains the machine-type
code (e.g. "10MYS03U00") while product_version holds the marketing name
(e.g. "ThinkCentre M910x"). Read both and prefer product_version when it
is set and different from product_name.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
vendor, model and serial_number now come from /sys/class/dmi/id/
(sys_vendor, product_name, product_serial) via SSH, reflecting the
actual physical server rather than the Proxmox software layer.
Falls back to "Proxmox Server Solutions GmbH" / status.model if SSH
or DMI files are unavailable (e.g. bare-metal without SSH creds).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Implements get_system_config() in ProxmoxSystemMixin:
- cluster_name from /cluster/status (type=cluster entry)
- cluster_nodes list of online node hostnames
- timezone from /nodes/{node}/time
- hostname, ssh_port, ssh_password_auth with safe defaults
Used by NetOrk's sync task to name the NetBox Cluster after the
actual Proxmox cluster rather than falling back to the site name.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
After renaming /etc/hostname, Proxmox boots under the new node name and
looks for VM/CT configs in /etc/pve/nodes/<new>/qemu-server/. Without
migrating the directory first, all VMs appear missing after the reboot.
Now renames /etc/pve/nodes/<old>/ to /etc/pve/nodes/<new>/ while
pve-cluster is running (pmxcfs supports live rename). Also replaces
pvecm updatecerts -f with pvenode cert create --overwrite — pvecm
updatecerts restarts pve-cluster, unmounts /etc/pve and can crash VMs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>