LYNXLAB.DEV
GLOBAL REPOSITORY // QUESTIONS & ANSWERS

Questions & Answers

Interactive engineering question repository focusing on internal mechanics, crisis reflexes, and diagnostic methodology.

Platform:
Category:
Level:
100 questions found
01
PROXMOX•PVE / Cluster & HA•[Senior]

In a 3-node cluster, how does quorum behave if 1 node drops? What happens if 2 nodes drop?

When 1 node drops, quorum is maintained (2/3 majority), so the cluster continues operating normally. When 2 nodes drop (1/3), quorum is lost; pmxcfs switches to read-only mode, and all VM starts, configuration changes, and cluster operations freeze.

[+] Details
02
PROXMOX•PVE / Cluster & HA•[Senior]

When is a QDevice used? What does it solve, and what does it NOT solve?

It provides an external 1-vote tie-breaker for 2-node or even-numbered clusters to prevent split-brain ties. It provides only a quorum vote—it does NOT provide storage or compute resources.

[+] Details
03
PROXMOX•PVE / Cluster & HA•[Senior]

Why is a dedicated network always recommended for Corosync?

Corosync heartbeat packets are extremely sensitive to millisecond latency spikes and jitter. If shared with VM, Storage, or Migration traffic, buffer bloat will cause corosync timeouts and freeze the cluster.

[+] Details
04
PROXMOX•PVE / Cluster & HA•[Senior]

How do you diagnose packet loss and latency issues in Corosync?

Standard ICMP ping is insufficient; you must inspect UDP ring latency, jitter, and Corosync retransmission logs.

[+] Details
05
PROXMOX•PVE / Cluster & HA•[Mid]

A node appears offline in GUI but is reachable via SSH. Where do you start troubleshooting?

Even if SSH and ICMP work, the `pve-cluster` or `corosync` service may be crashed, the isolated cluster ring network severed, or pmxcfs locked.

[+] Details
06
PROXMOX•PVE / Cluster & HA•[Senior]

What is pmxcfs and why is it critical?

It is a FUSE-based cluster filesystem mounted at `/etc/pve`, backed by SQLite and replicated in real-time across all nodes via Corosync.

[+] Details
07
PROXMOX•PVE / Cluster & HA•[Senior]

The cluster filesystem `/etc/pve` became read-only. What causes this?

Quorum majority was lost due to network partition or node failures.

[+] Details
08
PROXMOX•PVE / Cluster & HA•[Senior]

How do you safely remove a node from a Proxmox cluster?

Power off and physically disconnect the node to be removed first; then execute `pvecm delnode <node>` from one of the remaining active nodes.

[+] Details
09
PROXMOX•PVE / Cluster & HA•[Expert]

What is split-brain and how is it prevented in PVE?

It is a catastrophic condition where network isolation leads both partitioned sides to believe they are the authoritative cluster, writing to the same shared VM disk simultaneously. PVE prevents this via Quorum Majority voting and Watchdog Fencing.

[+] Details
10
PROXMOX•PVE / Cluster & HA•[Senior]

What is the architectural difference between HA and Live Migration?

Live Migration is zero-downtime memory transfer during planned maintenance. HA is automated crash recovery (reboot on another node) after unexpected hardware failure.

[+] Details
11
PROXMOX•PVE / Cluster & HA•[Senior]

Why does an HA VM restart on another node instead of seamlessly migrating when a node dies?

Because when physical hardware crashes or loses power, volatile RAM is instantly erased; there is no live memory state left to transfer.

[+] Details
12
PROXMOX•PVE / Cluster & HA•[Expert]

Why is fencing mandatory in high-availability clusters?

To ensure a non-responsive or failed host is guaranteed dead and can no longer write to shared storage before another node starts its VMs.

[+] Details
13
PROXMOX•PVE / Cluster & HA•[Mid]

What is an HA group and when is it used?

Rules that constrain VMs to run only on designated nodes with specific failover priority orders.

[+] Details
14
PROXMOX•PVE / Cluster & HA•[Senior]

You added a VM to HA, but when its node died the VM did not start on another node. How do you debug?

1) Did the surviving cluster have quorum? 2) Did fencing complete? 3) Does the VM reference local storage (local ISO/disk) or PCI hardware missing on the target? 4) Inspect CRM logs.

[+] Details
15
PROXMOX•Live Migration•[Mid]

Why is live migration fast when shared storage (Ceph/NFS/iSCSI) is in use?

Because virtual disks already reside on shared storage, only volatile RAM and CPU register state are streamed over the network (downtime < 500ms).

[+] Details
16
PROXMOX•Live Migration•[Senior]

What actually happens under the hood during live migration of a VM with local disk storage?

In addition to streaming RAM state, all virtual disk blocks are mirrored in real-time to the target host using QEMU NBD (Network Block Device) mirror protocol.

[+] Details
17
PROXMOX•Live Migration•[Senior]

What happens during live migration if the VM memory dirty rate is higher than the network transfer rate?

If the VM modifies RAM faster than network throughput can transfer dirty pages, live migration cannot converge, leading to an infinite migration loop.

[+] Details
18
PROXMOX•Live Migration•[Senior]

Why can live migration get stuck at 95% for an extended duration?

The target downtime threshold required to pause the VM and transfer the final batch of dirty RAM pages cannot be satisfied due to heavy memory write load.

[+] Details
19
PROXMOX•Live Migration•[Senior]

What is the architectural disadvantage of using CPU type `host` for live migration?

If the target host processor is older or lacks specific instruction flags present on the source, live migration will instantly trigger a guest kernel panic.

[+] Details
20
PROXMOX•Live Migration•[Mid]

How do you choose a CPU model across cluster nodes with different CPU generations?

Select the highest common instruction set supported by the oldest CPU in the cluster (e.g., `x86-64-v2-AES` or `x86-64-v3`).

[+] Details
21
PROXMOX•Live Migration•[Mid]

Why is live migration problematic if a VM uses PCI passthrough (e.g., GPU, NIC)?

Physical hardware registers and DMA memory mappings are tied to a specific physical card and cannot be serialized and streamed across the network.

[+] Details
22
PROXMOX•Live Migration•[Senior]

Why might a VM fail to initialize on the target node during live migration?

Due to missing VLAN/bridge interfaces, insufficient free host RAM, missing storage pools, or incompatible CPU instruction sets on the target host.

[+] Details
23
PROXMOX•CPU / RAM / Performance•[Mid]

What is CPU overcommit? When does it become hazardous?

Allocating more total vCPUs than available physical CPU threads. It becomes dangerous when multiple VMs simultaneously spike to 100% load, causing extreme scheduling latency and unresponsive guests.

[+] Details
24
PROXMOX•CPU / RAM / Performance•[Senior]

What is CPU steal time (`%st`) and how do you interpret it in Proxmox?

The percentage of time a virtual machine wants to execute instructions on CPU but is forced to wait by the hypervisor scheduler because physical cores are busy serving other VMs.

[+] Details
25
PROXMOX•CPU / RAM / Performance•[Expert]

When is NUMA architecture truly critical for performance?

On multi-socket servers hosting large-memory database VMs, NUMA is critical to avoid cross-socket UPI memory latency penalties.

[+] Details
26
PROXMOX•CPU / RAM / Performance•[Mid]

Is it good practice to assign all physical CPU cores as vCPUs to a single VM?

Generally no. Oversizing vCPUs forces the Linux hypervisor scheduler to pay severe co-scheduling penalties, reducing overall VM performance.

[+] Details
27
PROXMOX•CPU / RAM / Performance•[Senior]

When does CPU pinning boost performance, and when does it degrade system flexibility?

It maximizes L3 CPU cache hits and eliminates context switching for real-time/database workloads. However, on overcommitted hosts, it eliminates scheduler flexibility and creates severe queuing delays.

[+] Details
28
PROXMOX•CPU / RAM / Performance•[Senior]

How does RAM ballooning work? What happens if the guest lacks a balloon driver?

The hypervisor inflates a balloon driver inside the guest to reclaim unused RAM for the host. If the guest lacks the VirtIO balloon driver, the setting is completely inoperative.

[+] Details
29
PROXMOX•CPU / RAM / Performance•[Mid]

When should Hugepages be utilized?

On large-memory VMs (32GB+) to reduce CPU Translation Lookaside Buffer (TLB) cache misses and accelerate memory access speeds.

[+] Details
30
PROXMOX•CPU / RAM / Performance•[Senior]

A VM displays "100% CPU utilization" but application throughput is crawling. What is the real cause?

The application is not CPU-bound; it is blocked waiting on disk storage I/O or memory locks (I/O Wait / `%wa` or CPU steal `%st`).

[+] Details
31
PROXMOX•CPU / RAM / Performance•[Senior]

A VM's RAM usage keeps climbing steadily. How do you prove whether it is a true memory leak or normal Linux buffer/cache allocation?

By distinguishing between Linux filesystem buffer/cache (`buff/cache`) and actual process anonymous memory (RSS / `Active(anon)`).

[+] Details
32
PROXMOX•Storage (Ceph & ZFS)•[Expert]

How do you choose between Ceph, local ZFS, and LVM-Thin storage architectures?

Ceph: 5+ node clusters requiring shared resilient storage for seamless HA. Local ZFS: Maximum single-node NVMe IOPS with PBS backup replication. LVM-Thin: Lightweight local block storage with thin-provisioning.

[+] Details
33
PROXMOX•Storage (Ceph & ZFS)•[Senior]

Why can Ceph be problematic in clusters smaller than 4-5 nodes?

With 3x replication, if 1 node drops, the remaining 2 nodes cannot find a 3rd distinct host to satisfy CRUSH failure domain rules, leaving the cluster permanently degraded.

[+] Details
34
PROXMOX•Storage (Ceph & ZFS)•[Senior]

Why is Monitor (MON) quorum critical in Ceph?

If MON quorum is lost, clients cannot retrieve the active cluster map, causing all storage I/O operations across the entire datacenter to freeze immediately.

[+] Details
35
PROXMOX•Storage (Ceph & ZFS)•[Senior]

What is the difference between an OSD being "down" versus "out" in Ceph?

Down: The OSD daemon is stopped/unreachable, but its data is still expected to return. Out: Ceph declares the disk dead and triggers active data rebalancing to other OSDs.

[+] Details
36
PROXMOX•Storage (Ceph & ZFS)•[Senior]

When Ceph reports `HEALTH_WARN`, what is your immediate first diagnostic command?

Execute `ceph health detail` to determine the exact root cause (e.g. degraded PGs, down OSDs, near-full capacity thresholds, or clock skews).

[+] Details
37
PROXMOX•Storage (Ceph & ZFS)•[Expert]

An OSD is continuously flapping (down/up/down). How do you diagnose and triage?

Investigate physical drive SMART read timeouts, OSD network interface packet drops, controller queue hangs, or OSD memory exhaustion.

[+] Details
38
PROXMOX•Storage (Ceph & ZFS)•[Mid]

What is a Placement Group (PG) in Ceph?

A logical subdivision that groups millions of objects into manageable collections for mapping to OSDs, preventing individual object tracking overhead.

[+] Details
39
PROXMOX•Storage (Ceph & ZFS)•[Expert]

What do degraded, undersized, and inactive PG states mean in Ceph?

Degraded: Object replica count is below target but readable. Undersized: Fewer OSDs than pool size policy. Inactive: Minimum replica count cannot be satisfied, causing complete I/O freeze.

[+] Details
40
PROXMOX•Storage (Ceph & ZFS)•[Senior]

Why do client VMs slow down during Ceph recovery and backfill operations?

Because background data reconstruction consumes massive disk I/O and network bandwidth, forcing client VM requests into deep queues.

[+] Details
41
PROXMOX•Storage (Ceph & ZFS)•[Senior]

Why are both high bandwidth and ultra-low latency essential in Ceph network architecture?

Because a write transaction is not acknowledged to the VM until it commits across 3 OSDs over the network; network latency directly dictates disk write latency.

[+] Details
42
PROXMOX•Storage (Ceph & ZFS)•[Mid]

What is ZFS ARC and how does it impact Proxmox host RAM?

The Adaptive Replacement Cache (read cache) which defaults to consuming up to 50% of total host RAM. If unconstrained, spawning new VMs can trigger kernel OOM kills.

[+] Details
43
PROXMOX•Storage (Ceph & ZFS)•[Senior]

What is the architectural trade-off between ZFS Mirrors (RAID10) and RAIDZ for VM workloads?

Mirrors provide aggregate random IOPS multiplied by disk pairs. RAIDZ provides only the random IOPS of a single disk regardless of drive count due to parity calculation lock-step.

[+] Details
44
PROXMOX•Storage (Ceph & ZFS)•[Mid]

What is the fundamental difference between a ZFS snapshot and a PBS backup?

A ZFS snapshot is an internal block pointer on the same physical pool (not an external backup; if the disk dies, snapshots die). A PBS backup is an encrypted, deduplicated copy residing on an independent server.

[+] Details
45
PROXMOX•Storage (Ceph & ZFS)•[Senior]

What catastrophic failures occur when an LVM-Thin pool reaches 100% capacity?

All virtual machine write I/O immediately suspends; guest filesystems switch to read-only mode and data corruption can occur.

[+] Details
46
PROXMOX•Storage (Ceph & ZFS)•[Mid]

What happens to active VMs when underlying NFS shared storage drops offline?

VM I/O processes hang in uninterruptible sleep (`D state`), causing applications to freeze until NFS connectivity is restored.

[+] Details
47
PROXMOX•Storage (Ceph & ZFS)•[Senior]

How does Proxmox react when an iSCSI storage path fails?

If Multipath I/O is not configured, SCSI commands timeout, the block device is dropped, and guest filesystems switch to read-only.

[+] Details
48
PROXMOX•Enterprise Networking•[Junior]

How does a Linux Bridge operate under the hood in Proxmox?

It functions as a software Layer-2 Ethernet switch in kernel space, connecting physical NICs with virtual machine `tap` interfaces.

[+] Details
49
PROXMOX•Enterprise Networking•[Mid]

What is the operational benefit of a VLAN-aware Linux Bridge?

It eliminates creating separate bridges for every VLAN, dynamically filtering and trunking up to 4096 802.1Q VLAN tags over a single `vmbr0`.

[+] Details
50
PROXMOX•Enterprise Networking•[Senior]

You assigned a VLAN tag to a VM but it has no network connectivity. How do you debug layer-by-layer?

1) Guest IP/Gateway, 2) VM tap port VLAN tag, 3) Host `bridge-vlan-aware` configuration, 4) Physical switch port trunk VLAN allowed list.

[+] Details
51
PROXMOX•Enterprise Networking•[Senior]

The physical switch trunk is properly configured, but the VM still cannot communicate on its VLAN. What do you inspect?

Verify whether `bridge-vlan-aware` is enabled on the host bridge and check if `pve-firewall` or ebtables rules are dropping DHCP/ARP broadcast packets.

[+] Details
52
PROXMOX•Enterprise Networking•[Junior]

What is the architectural difference between an Access port and a Trunk port?

An Access port belongs to a single VLAN carrying untagged frames. A Trunk port multiplexes multiple VLANs carrying 802.1Q tagged frames.

[+] Details
53
PROXMOX•Enterprise Networking•[Mid]

What is the difference between Active-Backup (Mode 1) bonding and LACP (802.3ad / Mode 4)?

Active-Backup requires no switch configuration and keeps one link idle for failover. LACP negotiates with the switch to provide both link redundancy and aggregated multi-flow bandwidth.

[+] Details
54
PROXMOX•Enterprise Networking•[Senior]

Why does an LACP bond NOT split a single TCP stream across two physical interfaces?

To prevent out-of-order packet delivery which destroys TCP throughput, a single TCP flow is pinned to a single physical link via packet header hashing.

[+] Details
55
PROXMOX•Enterprise Networking•[Senior]

You configured MTU 9000 (Jumbo Frames) but Ceph and storage throughput dropped severely. What is the cause?

An intermediary switch port, bridge, or interface in the path remains at MTU 1500, causing silent packet drops and massive TCP retransmission storms.

[+] Details
56
PROXMOX•Enterprise Networking•[Senior]

How do you definitively verify that Jumbo Frames are functioning end-to-end without fragmentation?

Send an 8972-byte ICMP payload with the Don't Fragment (DF) bit enabled (8972 payload + 28 bytes IP/ICMP header = 9000 bytes).

[+] Details
57
PROXMOX•Enterprise Networking•[Mid]

The Proxmox host has Internet connectivity, but virtual machines cannot reach the outside network. Where do you start?

Inspect VM default gateway configuration, Linux bridge binding, and verify whether NAT masquerading or firewall forwarding rules are missing on the host.

[+] Details
58
PROXMOX•Enterprise Networking•[Junior]

A VM can successfully ping its default gateway but cannot reach Internet hosts. What is the issue?

1) Missing or invalid DNS configuration in `/etc/resolv.conf`, or 2) The upstream firewall/gateway lacks NAT forwarding rules for the VM's IP subnet.

[+] Details
59
PROXMOX•Enterprise Networking•[Senior]

Only VMs hosted on a specific node experience network connectivity failures. What is your diagnosis?

A switch port configuration mismatch (missing VLAN trunk allowed list), physical cabling/transceiver failure on that node, or a local `interfaces` misconfiguration.

[+] Details
60
PROXMOX•PBS (Backup Server)•[Senior]

Why is Proxmox Backup Server (PBS) fundamentally superior to classical `vzdump` backups?

Classical vzdump writes full compressed disk images every time. PBS splits data into 4MB SHA-256 chunks, achieving global deduplication, client-side encryption, and instant incremental backups.

[+] Details
61
PROXMOX•PBS (Backup Server)•[Senior]

How does deduplication work internally in PBS?

Data streams are sliced into 4MB chunks, hashed with SHA-256, and indexed. If a chunk hash already exists on disk, it is not rewritten; only its reference is appended to the backup manifest.

[+] Details
62
PROXMOX•PBS (Backup Server)•[Junior]

What is a "Chunk" in PBS terminology?

The atomic storage block in PBS; an encrypted and compressed slice of data averaging 4MB in size, addressed uniquely by its SHA-256 cryptographic digest.

[+] Details
63
PROXMOX•PBS (Backup Server)•[Senior]

Explain the difference between Prune and Garbage Collection (GC) with a real-world analogy.

Prune = Deleting the catalog index card of a book from the library. Garbage Collection (GC) = Walking the bookshelves and physically shredding orphaned pages that no catalog card references.

[+] Details
64
PROXMOX•PBS (Backup Server)•[Senior]

Your datastore is 95% full; you ran Prune but utilization remains at 95%. Why?

Because Prune only deletes metadata manifests; physical chunks are not freed until Garbage Collection (GC) runs and passes the 24-hour safety grace window.

[+] Details
65
PROXMOX•PBS (Backup Server)•[Mid]

Why can Garbage Collection take an extensive time to complete?

Because datastores contain millions of small 4MB chunk files, forcing random disk I/O bottlenecks during metadata status checks (especially on mechanical HDDs).

[+] Details
66
PROXMOX•PBS (Backup Server)•[Senior]

What is the critical distinction between a Verification Job and a Restore Test?

Verification tests cryptographic chunk integrity on disk (SHA-256 match). A Restore Test proves that the operating system and database actually boot, mount, and function correctly in production.

[+] Details
67
PROXMOX•PBS (Backup Server)•[Senior]

A backup passes verification with 100% success, but restoring the VM fails to boot. What happened?

During backup, the QEMU Guest Agent filesystem freeze (`fs-freeze`) failed, capturing the database in an inconsistent state, or the guest OS bootloader was corrupted.

[+] Details
68
PROXMOX•PBS (Backup Server)•[Senior]

Why is filesystem selection critical for a PBS datastore?

For high-speed metadata lookups across millions of chunks and built-in protection against silent data corruption (bit-rot) via ZFS checksumming.

[+] Details
69
PROXMOX•PBS (Backup Server)•[Senior]

How do you design a storage RAID pool for a high-performance PBS datastore?

Deploy ZFS RAID10 (Striped Mirrors) for maximum random IOPS and fast rebuilds; for large capacity RAIDZ2 pools, attach an NVMe Special Metadata vdev.

[+] Details
70
PROXMOX•PBS (Backup Server)•[Mid]

SSD versus HDD for PBS: Which hardware fits which workload?

For enterprise production and high VM densities, SSDs are mandatory. For multi-petabyte cold archives, enterprise HDDs backed by NVMe metadata caching can be utilized.

[+] Details
71
PROXMOX•PBS (Backup Server)•[Mid]

What is the architectural benefit of isolating the PBS backup network from the PVE cluster network?

To prevent heavy backup traffic (especially initial baseline transfers) from saturating corosync heartbeats, storage networks, and live VM interfaces.

[+] Details
72
PROXMOX•PBS (Backup Server)•[Mid]

What is the operational difference between a Backup Job and a Sync Job?

Backup: PVE clients push VM data to PBS. Sync Job: A secondary PBS server pulls missing deduplicated chunks directly from a primary PBS server over WAN.

[+] Details
73
PROXMOX•PBS (Backup Server)•[Senior]

Why is an offsite Remote PBS vastly more valuable than a second datastore in the same physical datacenter?

Because site-wide physical disasters (fire, flood, electrical catastrophe, physical theft) or ransomware outbreaks destroy local secondary stores simultaneously.

[+] Details
74
PROXMOX•PBS (Backup Server)•[Expert]

What happens if the client-side encryption key for a PBS backup is lost?

The backups are mathematically irrecoverable; all data is permanently and irrevocably lost.

[+] Details
75
PROXMOX•PBS (Backup Server)•[Senior]

Where should PBS encryption keys be securely stored and archived?

1) Active PVE host `/etc/pve/priv/storage/<id>.enc`, 2) Corporate Secrets Manager (HashiCorp Vault), 3) Physical printed offline Paperkey.

[+] Details
76
PROXMOX•PBS (Backup Server)•[Senior]

A PBS datastore suffered disk corruption. How do you identify which backup snapshots are intact?

Execute a comprehensive `Verify Job` across the entire datastore to scan and report corrupted chunk hashes and affected snapshot manifests.

[+] Details
77
PROXMOX•PBS (Backup Server)•[Senior]

Only one specific VM fails verification while all other VMs pass. What is the root cause?

A specific chunk unique to that VM suffered bit-rot corruption, or a previous backup write was interrupted mid-transfer.

[+] Details
78
PROXMOX•PBS (Backup Server)•[Expert]

Primary PBS is completely destroyed, but Remote PBS offsite is healthy. How do you execute disaster recovery?

Connect PVE directly to the Remote PBS server over WAN using the master encryption key, and restore VMs with `--live-restore 1` for immediate boot.

[+] Details
79
PROXMOX•Disaster Recovery•[Expert]

A PVE cluster is completely lost, but PBS backups are intact. What are your step-by-step recovery actions?

1) Install a clean Proxmox VE host, 2) Configure base network bridges/VLANs, 3) Add PBS storage target and provide encryption keys, 4) Restore all VMs sequentially via `qmrestore --live-restore 1`.

[+] Details
80
PROXMOX•Disaster Recovery•[Mid]

What is the architectural difference between a PVE Configuration backup and a VM backup?

VM backup captures virtual disks and VM `.conf` files. Host config backup captures `/etc/pve`, `/etc/network/interfaces`, and Corosync cluster definitions.

[+] Details
81
PROXMOX•Disaster Recovery•[Mid]

You installed a fresh PVE host. How do you restore VMs from an existing PBS server?

Add the PBS storage via web GUI or CLI (`pvesm add pbs ...`), navigate to Datacenter -> Storage -> Backups, select the VM snapshot, and click Restore.

[+] Details
82
PROXMOX•Disaster Recovery•[Expert]

Why does an entire Disaster Recovery plan fail if encryption keys are missing?

Because encrypted PBS chunks cannot be decrypted without the secret key; even with 100% of the storage chunks present, data recovery is mathematically impossible.

[+] Details
83
PROXMOX•Disaster Recovery•[Mid]

Define Recovery Point Objective (RPO) and Recovery Time Objective (RTO) in infrastructure architecture.

RPO: Maximum acceptable data loss duration (e.g., last 15 minutes of transactions). RTO: Maximum acceptable downtime duration required to bring systems back online (e.g., 30 minutes).

[+] Details
84
PROXMOX•Disaster Recovery•[Senior]

What is the profound difference between "Having Backups" and "Being DR-Ready"?

Having backups means data copies exist. Being DR-ready means having an audited, rehearsed, documented, and timed playbook proven to restore operational business services.

[+] Details
85
PROXMOX•Disaster Recovery•[Senior]

How do you architect an automated monthly restore testing pipeline?

Restore randomly sampled production VMs into an isolated sandbox VLAN/bridge on a dedicated staging host, and run automated health-check scripts without IP collision risks.

[+] Details
86
PROXMOX•Disaster Recovery•[Mid]

If Production PBS and DR PBS are in the same building, why is it NOT true Disaster Recovery?

Because site-wide physical disasters (fire, structural collapse, flood, power surge, theft) destroy both backup systems simultaneously.

[+] Details
87
PROXMOX•PDM (Datacenter Manager)•[Senior]

What is the architectural difference between PDM and a PVE Cluster?

PVE Cluster is a tightly-coupled local cluster requiring shared Corosync quorum. PDM is a loosely-coupled multi-datacenter management layer that centrally monitors independent clusters without shared quorum dependencies.

[+] Details
88
PROXMOX•PDM (Datacenter Manager)•[Senior]

If the PDM server goes down, what is the operational impact on constituent PVE clusters?

Zero operational impact! All PVE clusters continue executing local VM workloads, HA failovers, and storage replication without interruption.

[+] Details
89
PROXMOX•PDM (Datacenter Manager)•[Mid]

How does PDM communicate securely with disparate PVE clusters across WAN?

Via HTTPS/TLS using PVE REST APIs and scoped API Tokens with cryptographic authentication.

[+] Details
90
PROXMOX•PDM (Datacenter Manager)•[Junior]

What is the security advantage of API Tokens over user passwords in automation?

Passwords grant broad interactive access and are difficult to revoke cleanly. API Tokens have restricted privileges (least-privilege), expiration dates, and can be instantly revoked without changing user credentials.

[+] Details
91
PROXMOX•PDM (Datacenter Manager)•[Junior]

What is the fundamental difference between Authentication (AuthN) and Authorization (AuthZ)?

Authentication: "Who are you?" (Identity verification via PAM, LDAP, OIDC, 2FA). Authorization: "What are you permitted to do?" (RBAC permissions to start, modify, or delete VMs).

[+] Details
92
PROXMOX•PDM (Datacenter Manager)•[Senior]

How do you design RBAC so a user can only access designated clusters within PDM?

Assign granular PDM RBAC roles (e.g. `PVE.Auditor` or `PVE.PowerUser`) scoped strictly to specific Cluster Path ACLs.

[+] Details
93
PROXMOX•PDM (Datacenter Manager)•[Mid]

Why is centralized visibility in PDM operationally essential for enterprise infrastructure?

It aggregates global VM inventories, cross-datacenter storage utilization, and active alerts on a single unified pane of glass instead of requiring separate logins to dozens of disparate clusters.

[+] Details
94
PROXMOX•PDM (Datacenter Manager)•[Senior]

Why is preserving individual cluster autonomy essential when designing multi-datacenter infrastructure with PDM?

To ensure that WAN outages, cloud partition events, or centralized PDM failures do not disrupt local cluster HA failover, storage I/O, or VM execution.

[+] Details
95
PROXMOX•Crisis Scenarios (Triage)•[Expert]

Scenario 1: 5-node Ceph cluster. Node 3 suddenly crashed. Ceph reports HEALTH_WARN. 6 of 20 VMs failed to restart under HA. Simultaneously, some VMs lost network. What is your exact triage sequence in the first 10 minutes?

1) Verify Quorum via `pvecm status` (4/5 active). 2) Confirm Node 3 Watchdog Fencing completed via `ha-manager status` and check if failed VMs have local ISO/disk dependencies. 3) Inspect `ceph health detail` for inactive PGs. 4) Verify switch VLAN trunking on target nodes. 5) Safely boot remaining failed VMs.

[+] Details
96
PROXMOX•Crisis Scenarios (Triage)•[Senior]

Scenario 2: VM live migration stalls indefinitely at 97%. VM is operating normally, but migration will not finish. Storage is Ceph. Network is 10 GbE. What do you inspect and tune?

1) Check VM RAM dirty rate against network bandwidth in QEMU task logs, 2) Verify MTU and interface packet drops on the 10 GbE link, 3) Enable QEMU `auto-converge` and temporarily increase downtime tolerance.

[+] Details
97
PROXMOX•Crisis Scenarios (Triage)•[Expert]

Scenario 3: PBS datastore is 100% full. Prune completed. You want to run Garbage Collection, but active backup jobs are currently executing. What do you do and why?

Temporarily pause or abort running backup jobs and immediately initiate Garbage Collection (GC). When datastore is 100% full, new chunk writes fail immediately, breaking all active backups until orphaned chunks are cleared.

[+] Details
98
PROXMOX•Crisis Scenarios (Triage)•[Senior]

Scenario 4: Last night's backup reported success. Today's scheduled verification failed. Can you trust the backup archive? What actions do you take?

You CANNOT trust it! Verification failure proves on-disk chunk corruption (bit-rot or storage controller error). Attempt test restore to an isolated sandbox and immediately trigger a Force Full Backup for the affected VM.

[+] Details
99
PROXMOX•Crisis Scenarios (Triage)•[Expert]

Scenario 5: All 3 PVE nodes suffered simultaneous drive failure. Cluster is gone. Remote PBS offsite is healthy. Backups are encrypted and you have the master encryption key. How do you rebuild production from scratch?

1) Install clean Proxmox VE on fresh hardware, 2) Configure identical switch VLAN bridges, 3) Add Remote PBS storage with encryption key, 4) Execute `qmrestore` with `--live-restore 1` to boot all mission-critical VMs within seconds.

[+] Details
100
PROXMOX•Crisis Scenarios (Triage)•[Expert]

Scenario 6 (Master Triage): In a 3-node PVE + Ceph + PBS + LACP/VLAN + HA + PDM architecture: "Node 2 crashed, Ceph is degraded, 2 VMs lost network, HA failed to start 1 VM, and the latest PBS backup failed verification." In what exact order do you triage and why?

Triage Sequence: 1) Quorum (2/3 majority valid?), 2) Node 2 Fencing (isolated by watchdog?), 3) Ceph I/O (any inactive PGs blocking disk writes?), 4) HA Failed VM (local ISO/storage blocking start?), 5) Network (VLAN trunk permissions on target switch port?), 6) PBS (trigger replacement full backup for corrupted snapshot).

[+] Details