Skip to content

Log WireGuard command errors to syslog - #2401

Merged
openipc-ai merged 12 commits into
OpenIPC:masterfrom
usa-:log-wireguard-errors
Sep 13, 2026
Merged

openipc-ai merged 12 commits into
OpenIPC:masterfrom
usa-:log-wireguard-errors

Conversation

@usa-

@usa- usa- commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Problem

WireGuard initialization errors are currently printed to the console only, which makes them easy to miss when the camera is accessed remotely.

Solution

Add a run_cmd() helper that captures command output and exit codes. When a command fails, the error, exit code, and command output are written to syslog before the script terminates with the original exit code.

This helper could be moved to a common shell library in the future and reused by multiple OpenIPC scripts. The global variable names would need to be adjusted first to avoid conflicts with variables in the calling scripts.

Hardware tested on

  • gk7205v300, G6S
  • t31x, don't know the board name
  • ssc337de, don't know the board name
  • ssc378de, don't know the board name, multiple cameras

The fix has been tested on these platforms.

Evidence

See OpenIPC/firmware#2319 for the problem reproduction, testing, and discussion.

Scope

  • No kernel patches under general/package/all-patches/linux/ (those go to https://github.com/OpenIPC/linux)
  • No files specific to a single retail camera model (those go to https://github.com/OpenIPC/builder)
  • No probing or bring-up tooling (that goes to https://github.com/OpenIPC/ipctool)
  • Nothing under general/overlay/ or in a shared load_<vendor> script hardcodes a value specific to my board
  • Package sources come from an OpenIPC repository, and any version bump keeps at least the specificity of the pin it replaces (a new package should pin a full 40-character SHA)
  • No LD_PRELOAD, and no binaries that cannot be rebuilt from source
  • New code is selected by a defconfig, so CI actually builds it

@qodo-free-for-open-source-projects

Copy link
Copy Markdown

PR Summary by Qodo

Log WireGuard command failures to syslog

🐞 Bug fix ✨ Enhancement 🕐 10-20 Minutes

Grey Divider

AI Description

• Capture WireGuard setup command output and exit status through a shared helper.
• Log failures and command output to syslog before terminating with the original status.
• Apply consistent failure handling across environment, interface, configuration, and route
 commands.
Diagram

graph TD
  A["WireGuard init"] --> B["run_cmd helper"] --> C["System command"] --> D{"Command succeeds?"}
  D -->|Yes| E["Return output"]
  D -->|No| F["Syslog error"] --> G["Terminate script"]
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Shared shell command library
  • ➕ Provides consistent error logging across multiple OpenIPC scripts.
  • ➕ Centralizes future fixes to command execution and exit propagation.
  • ➖ Introduces shared global-variable collision risks in existing scripts.
  • ➖ Expands this narrowly scoped fix into a cross-cutting migration.
  • ➖ Requires defining and validating a stable shell helper interface.

Recommendation: Keep the helper local to the WireGuard script for this PR. It resolves the immediate observability problem without introducing shared-library compatibility risks; extraction should follow only after variable scoping and caller conventions are standardized.

Files changed (1) +56 / -28

Bug fix (1) +56 / -28
wireguardCapture and syslog WireGuard command failures +56/-28

Capture and syslog WireGuard command failures

• Adds a 'run_cmd()' helper that captures combined output and exit status, logs failures with command context to syslog, and terminates initialization. Routes module loading, environment reads, WireGuard configuration, interface setup, and route creation through the helper.

general/overlay/usr/sbin/wireguard

@qodo-free-for-open-source-projects

qodo-free-for-open-source-projects Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Cameras without shared keys do not start ✓ Resolved 🐞 Bug ≡ Correctness
Description
run_cmd sends an untrapped TERM to $$ on any nonzero command result, terminating the invoking
shell before exit "$rc" can preserve the captured status. An absent optional wg_sharkey value
therefore stops startup despite the trailing || true, while direct failures in module loading,
interface creation, configuration application, address setup, and route setup report a
signal-derived status instead of the original command code.
Code

general/overlay/usr/sbin/wireguard[31]

+WG_PRESHARED_KEY=$(run_cmd "" fw_printenv -n wg_sharkey 2>/dev/null || true)
Evidence
The wrapper saves the failed command's status at line 10, but signals $$ at line 20 before line 21
can exit with that status; command substitution retains the invoking shell's process ID, so this
signal still targets the main script. Repository code also shows that fw_printenv -n returns
failure for an undefined variable, the preshared key is emitted only when nonempty, and startup
requires only the private key, proving that a missing optional wg_sharkey value should be
tolerated. Because multiple newly wrapped setup commands invoke the same failure path directly,
their failures are likewise terminated by the signal before the captured status can be returned.

general/overlay/usr/sbin/wireguard[9-21]
general/overlay/usr/sbin/wireguard[30-32]
general/overlay/usr/sbin/wireguard[52-52]
general/package/all-patches/uboot-tools/0011-env-partition-autosearch.patch[17-25]
general/overlay/etc/init.d/S98wireguard[3-8]
general/overlay/usr/sbin/wireguard[27-28]
general/overlay/usr/sbin/wireguard[57-69]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`run_cmd` signals the current script on failure before it can return or exit with the command status stored in `rc`. This prevents an unset optional `wg_sharkey` value from being tolerated by `|| true` and causes mandatory setup failures to report a signal-derived status rather than the original command status.
## Fix Focus Areas
- general/overlay/usr/sbin/wireguard[5-21]
- general/overlay/usr/sbin/wireguard[9-21]
- general/overlay/usr/sbin/wireguard[27-28]
- general/overlay/usr/sbin/wireguard[30-31]
- general/overlay/usr/sbin/wireguard[50-53]
- general/overlay/usr/sbin/wireguard[57-69]
## Recommended Fix
Remove process signaling from `run_cmd`; after logging the failure, use `return "$rc"` so callers receive the captured command status. Refactor mandatory configuration lookups currently embedded in command substitutions into variable assignments whose failures are checked by the main shell with explicit termination using the returned status, and make each mandatory top-level setup caller explicitly terminate with that status. Keep the optional preshared-key lookup nonfatal so a missing `wg_sharkey` produces an empty value and configuration generation continues.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Tip of the day
💡 Did you know, you can start a comment with 'qodo' or '@qodo' to chat about any finding

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread general/overlay/usr/sbin/wireguard Outdated
@usa-

usa- commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Tested on a real camera with BusyBox 1.36.1. A missing required U-Boot environment variable causes fw_printenv to return exit code 1; run_cmd() logs the error to syslog and terminates the script before it can reach the DNS wait. The optional wg_sharkey variable is intentionally handled separately.

@usa-

usa- commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor Author

For human reviewers:

One more thought: some of the explicit error descriptions may not be necessary. The syslog entry already includes the command that failed, which often makes it clear what operation could not be completed. The additional message is useful where it adds context, but perhaps not every command needs one.

@openipc-ai openipc-ai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this — the syslog capture is the right idea, and the local output / local rc=$? split is correct (checked under busybox ash: rc really does get the command's status, which is the trap most versions of this helper fall into). Reading the environment before modprobe/ip link add is a real improvement too: master creates and configures wg0 and only then exits on a missing wg_address, leaving a half-built interface behind.

Qodo's kill -s TERM $$ finding is genuinely fixed in c6e4104 — I checked the commit history rather than trusting the resolved marker.

Three things to fix before this lands, left inline below. Two of them are reachable on a stock configuration, and this file is in general/overlay/, so it ships to all 99 boards (ci-matrix.py --stdin widens to the full matrix).

Evidence

CLAUDE.md asks a behaviour-changing PR for pasted before/after output rather than a description of it, and here the output is the feature. Could you add the actual logread lines from one of the four boards — a failing run showing the three lines run_cmd writes? The link to #2319 covers the reproduction, but not what the new logging looks like on a camera.

Separately, for a maintainer: the five workflows on this PR are all sitting at action_required, so nothing has built yet.

Comment thread general/overlay/usr/sbin/wireguard Outdated
Comment on lines +19 to +22
printf '%s\n' "$output" |
while IFS= read -r line; do
logger -p user.err -t "$prog[$$]" "$line"
done

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

$output is empty on the most common failure here, and printf '%s\n' "" still feeds one empty line to logger.

This tree patches fw_printenv so that -n suppresses ## Error: "x" not defined (general/package/all-patches/uboot-tools/0011-env-partition-autosearch.patch), so an unset variable exits 1 with nothing on stderr. Each of the six environment reads below therefore logs a trailing blank line:

Error: Failed to read wg_alive U-Boot environment variable.
Command `fw_printenv -n wg_alive` finished with code 1:
<blank>

Wrapping the loop in [ -n "$output" ] drops it.

Comment on lines +27 to +31
if [ -n "$var" ]; then
eval "$var=\"\$output\""
else
printf '%s\n' "$output"
fi

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This printf runs on success too, so every run_cmd call with an empty var writes a newline to stdout. Byte-exact stdout of a fully successful run (five void commands plus one per wg_allowed entry):

this PR:  0000000  \n  \n  \n  \n  \n  \n  \n
master:   0000000

That breaks the WebUI. majestic-webui's www/cgi-bin/wireguard.cgi runs this script synchronously and unredirected inside a haserl <% %> block, and the next thing it calls is redirect_to, which writes the raw status line itself (p/common.cgi:517):

  /usr/sbin/wireguard
  sleep 1
  redirect_to "$SCRIPT_NAME" "success" "WireGuard is up"

The blank lines land ahead of HTTP/1.1 303 See Other, so the "WireGuard is up" redirect stops working after a save.

Nothing reads run_cmd's stdout any more now that every value-returning call passes var, so the simplest fix is to drop the else branch and keep only the assignment.

Comment thread general/overlay/usr/sbin/wireguard Outdated
Comment on lines +70 to +75
run_cmd "" "" ip address add dev wg0 "$WG_ADDRESS"
run_cmd "" "" ip link set up dev wg0

for i in $(echo "$WG_ALLOWED" | tr ',' ' '); do
run_cmd "" "" ip -4 route add "$i" dev wg0
done

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Making these fatal is a regression against #1897, which added the route loop with per-entry failures deliberately tolerated.

Both failure modes are reachable on a stock camera — checked against real iproute2:

ip -4 route add ::/0 dev wg0        ->  rc=1  Error: inet prefix is expected rather than "::/0".
ip -4 route add 0.0.0.0/0 dev wg0   ->  rc=2  RTNETLINK answers: File exists   (eth0 default route already present)

So with wg_allowed=0.0.0.0/0,::/0,10.9.0.0/24:

  • master installs 0.0.0.0/0, ignores the ::/0 failure, installs 10.9.0.0/24
  • this version installs 0.0.0.0/0, exits on ::/0, and never installs 10.9.0.0/24

S98wireguard backgrounds the script, so the non-zero exit is invisible. What shows up on the camera is a missing route and a tunnel that half works — exactly the kind of failure this PR is trying to make visible. ip address add has the same shape on a re-run, where it fails with File exists and now takes ip link set up down with it.

Suggest keeping these three non-fatal — log and carry on — and reserving exit for the steps that genuinely cannot be survived.

@usa-

usa- commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

@openipc-ai,

Why do you recommend making ip address add, ip link set up, and route creation non-fatal? If any of these commands fails, the WireGuard tunnel is not fully operational. Wouldn't treating them as fatal be more appropriate?

@openipc-ai openipc-ai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do you recommend making ip address add, ip link set up, and route creation non-fatal? If any of these commands fails, the WireGuard tunnel is not fully operational. Wouldn't treating them as fatal be more appropriate?

Because exiting there doesn't prevent a broken tunnel — it builds a more broken one. By that point wg0 exists, carries the peer config, and (in the route loop) already has its address and is up. exit rolls none of that back; it only skips the work that is left. Failing on entry 2 of 3 leaves entry 1 installed and entry 3 missing, which is strictly worse than installing 1 and 3 and logging 2.

And the failures that dominate here are not breakage. RTNETLINK answers: File exists means the route or address is already present — that is the desired end state, and it is what a re-run produces; the WebUI runs /usr/sbin/wireguard on every save. An IPv6 entry in wg_allowed is valid WireGuard configuration too — it belongs in AllowedIPs, it just is not an IPv4 route — so ip -4 route add ::/0 failing is the script misreading a correct config as an error. wg_allowed is a list; "operational" is not one bit, and one rejected entry should not cost the others.

There is also an argument specific to this PR: the usual reason to abort early is that nobody would otherwise notice the failure. Your change removes that. Once it is in syslog, doing as much as possible and logging the rest beats stopping.

Where I agree with you: fatal is right when a failure makes everything after it meaningless — modprobe, ip link add, wg setconf — and those should stay fatal. ip address add is the genuinely arguable one, since a tunnel with no address is dead. If you would rather keep it fatal, tolerate EEXIST first, otherwise the re-run path kills it.


On 8d912d1: all three earlier findings are fixed — checked under busybox ash, all three routes now install where the previous revision stopped at ::/0, and both test_shell_parse.sh and test_strip_shell_comments.sh still pass. Three new things came in with the rework, inline below.

For a maintainer: the five workflow runs on c6e4104 all finished green, including 100/100 board build jobs. The runs on 8d912d1 are back at action_required; probably worth approving them after this round rather than spending the matrix twice.

Comment thread general/overlay/usr/sbin/wireguard Outdated
Comment on lines +23 to +34
if [ -n "$msg" ]; then
logger -p user.err -t "$prog[$$]" "$msg"
elif [ -z "$output" ]; then
logger -p user.err -t "$prog[$$]" "Command \`$*\` finished with code $rc."
else
logger -p user.err -t "$prog[$$]" "Command \`$*\` finished with code $rc:"

printf '%s\n' "$output" |
while IFS= read -r line; do
logger -p user.err -t "$prog[$$]" "$line"
done
fi

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch is now either/or, so setting msg suppresses the command's own output in syslog — which undoes what the PR set out to do. Same failing wg setconf, both revisions:

8d912d15:  Failed to apply wireguard configuration

c6e4104a:  Error: Failed to apply wireguard configuration.
           Command `wg setconf wg0 /tmp/wireguard.conf` finished with code 1:
           Line unrecognized: `PersistentKeepalive='
           Configuration parsing error

Line unrecognized / Configuration parsing error is the diagnostic #2319 was chasing — it is the difference between "WireGuard did not start" and "PersistentKeepalive is empty". It still reaches stderr via the block below, but S98wireguard backgrounds this script at boot, so stderr goes to the console: the thing the PR description opens by calling easy to miss when the camera is remote.

Suggest logging the message and the output — msg as the headline, then the command line and its output as before.

Comment thread general/overlay/usr/sbin/wireguard Outdated

if [ "$fatal" = "true" ]; then
exit "$rc"
fi

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Trailing tab after fi.

Comment on lines +45 to +49
if [ -n "$var" ]; then
eval "$var=\"\$output\""
elif [ -n "$output" ]; then
echo "$output"
fi

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

--non-fatal clears var at the top, so after a non-fatal failure execution falls through to here and prints the error a second time — once to stderr at line 37, once to stdout here.

The stdout copy is the WebUI problem from my earlier comment, still live. Byte-exact stdout of a run where one route fails:

0000000   E   r   r   o   r   :       i   n   e   t       p   r   e   f   i   x ...

A first save through the WebUI on a camera that already has a default route hits RTNETLINK answers: File exists on 0.0.0.0/0, and that text lands in the CGI body ahead of redirect_to's HTTP/1.1 303 See Other.

Since the failure path already writes to stderr and syslog, this branch only needs to run when rc is 0.

Two smaller notes on the same block: --non-fatal occupying the var slot means a command can never be both non-fatal and value-returning, which will bite when this helper moves to a shared library; and echo "$output" is worth keeping as printf '%s\n' "$output", since busybox echo can eat backslashes in some builds.

@openipc-ai openipc-ai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All findings addressed, and the rework is cleaner than what I suggested — folding the default message into msg keeps the message and the command output travelling together, and the Command ... finished with code N line carries the exit status either way.

Re-verified on the rebased head under busybox ash: stdout stays empty on a non-fatal route failure so the WebUI's redirect_to is unaffected, all three entries of a mixed v4/v6 wg_allowed are attempted, and a missing variable logs two clean lines with no stray blank. test_shell_parse.sh, test_strip_shell_comments.sh and ci-matrix.py --self-test all pass.

CI is green on ac915ea: 114 checks, 100/100 board builds, including hi3519v101_lite now that #2404 has landed.

Rebased onto master for you — #2404 touched only hi3519v101.generic.config, so nothing conflicted and the diff is unchanged.

@openipc-ai
openipc-ai merged commit 3769227 into OpenIPC:master Sep 13, 2026
116 checks passed
@usa-
usa- deleted the log-wireguard-errors branch September 18, 2026 08:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants