Conversation
himci_request() has ten ways out and nine of them reach request_end, which calls himci_finish_request() and through it mmc_request_done(). The tenth, in the FIFO-reset loop on the data path, is a bare return. Nothing else completes the request. The MMC core is sitting in mmc_wait_for_req() on an uninterruptible wait_for_completion(), so the caller becomes an unkillable D state, and because the host is never released every later request queues behind it. On a camera that is recording stopping and staying stopped until someone cuts the power -- the block layer, any MMC_IOC_CMD passthrough, and the card's own rescan all block on the same claim. Set the data error the way every other timeout in this driver does and jump to request_end. There is no DMA to unwind: himci_idma_start() is below this point. Found by reading, while working out why an MMC_IOC_CMD passthrough had once been reported as hanging a camera unkillably. That report matches this signature exactly -- a data command on a card that did not answer -- but I could not re-trigger it to confirm, so this is offered as a correctness fix rather than a diagnosis. The path leaves "fifo reset is timeout!" in dmesg, which is the thing to grep for if it is ever seen in the field. Compile-tested for hi3516av300 (CONFIG_HIMCI=y, arm-openipc-linux-gnueabi, kernel 4.9.37); no new warnings. The same bare return is in the hisilicon-hi3516ev200, hisilicon-hi3516cv200, hisilicon-hi3516av100 and hisilicon-hi3536dv100 branches and wants the same one-line change there.
PR Summary by QodoComplete HiMCI requests after FIFO reset timeouts
AI Description
Diagram
High-Level Assessment
Files changed (1)
|
Code Review by Qodo
1.
|
From review of the previous commit, and a fault in it. himci_setup_data() does two things that have to be undone: it maps the scatterlist for the device, and it takes host->data. himci_data_done() is the only thing that gives either back. Jumping to request_end from the FIFO-reset timeout skipped it, so the buffer went back to the MMC core still mapped for the device, and host->data was left pointing at a request that had just been completed. Call it, with no status bits, so it unmaps, clears host->data, zeroes bytes_xfered and keeps the -ETIMEDOUT set above rather than deciding an error of its own. Still no himci_idma_stop(): the engine is genuinely not running there, because himci_idma_start() is below this point. Worth noting while this is open: the command-error paths further down have the same shape. When himci_exec_cmd() fails or the command completes with an error, the data branch calls himci_idma_stop() and falls through to request_end without himci_data_done(), so those leak the mapping too. That is older than this change and is left alone here rather than folded into a one-line fix. Compile-tested as before, no new warnings.
|
Right, and it is a fault in the fix rather than in the old code. Addressed in
It now calls Worth recording while this is open, since it is the same class and I looked: Compile-tested as before, no new warnings. |
himci_request()has ten ways out. Nine reachrequest_end, which callshimci_finish_request()and through itmmc_request_done(). The tenth — inthe FIFO-reset loop on the data path — is a bare
return.Why it matters more than one lost request
The MMC core is parked in
mmc_wait_for_req()on an uninterruptiblewait_for_completion(). Nothing will ever complete it, so the caller is anunkillable D state — and because the host is never released, every later
request queues behind the same claim. The block layer, any
MMC_IOC_CMDpassthrough, and the card's own rescan all stop together.
On a camera that is recording stopping and staying stopped until someone cuts
the power.
The fix
Set the data error the way every other timeout in this driver does
(
-ETIMEDOUT, ashimci_wait_data_complete()andhimci_wait_card_complete()both do) and
goto request_end. There is no DMA to unwind —himci_idma_start()is below this point.
What this is and is not
Found by reading, while establishing whether the
MMC_IOC_CMDpassthrough isusable for SD vendor-health commands on HiSilicon. There is an older report of
exactly this shape — a data command on a card that did not answer leaving an
unkillable process and a camera that needed a power cycle — and the signature
matches, but I could not re-trigger it, so this is offered as a correctness
fix rather than a diagnosis.
I did look for a way to force it honestly.
retry_countis amodule_paramwith 0600 permissions, but the driver is built in and exposes no
/sys/module/himci/parameters/, so the only lever is ahimci.retry_count=1bootarg — and on a board whose rootfs is on NOR while the SD card is mounted
during init, that wedges the mount on every boot. The test board has neither
altbootcmdnor a wired serial console, so that experiment ends in a boardnobody can reach. Not worth it to prove what the control flow already states.
The path leaves
fifo reset is timeout!in dmesg, which is what to grep for ifit is ever seen in the field.
Testing
Compile-tested for hi3516av300 —
CONFIG_HIMCI=y,arm-openipc-linux-gnueabi,kernel 4.9.37,
make drivers/mmc/host/himci/himci.oclean with no newwarnings. Behaviour on that board is unchanged in normal use: SD recording, a
CMD13passthrough and a 512-byteCMD56data read all behave exactly asbefore this patch, because none of them reaches the FIFO-reset timeout.
Other branches
The identical bare return is in
hisilicon-hi3516ev200,hisilicon-hi3516cv200,hisilicon-hi3516av100andhisilicon-hi3536dv100.hisilicon-hi3516cv300andhisilicon-hi3519v101do not have it. Happy tosend the same one-line change to the other four if you would like it in one go
or as separate PRs — say which.