English | 简体中文
A dynamic-programming-based stochastic optimization framework to derive cost-optimal locking strategies for the weapon-refining mechanics in Identity V, minimizing the expected number of feathers spent.
- Slots: 5 slots total (1 color slot + 4 ordinary slots).
- Independent Re-rolls: Each wash re-rolls all unlocked slots independently.
-
Dynamic Probabilities: The per-slot gold probability rises linearly with cumulative wash count
$n$ from 1% ($n=0$ ) to a 10% cap ($n \ge 200$ ). -
Exponential Resource Scaling: Each wash costs
$2^t$ feathers, where$t$ is the number of locked slots. - Permanent Rollback: The game permits players to restore the slot state to immediately before any prior wash at zero cost, preserving the cumulative wash count.
Comparing the DP-optimal policies against the empirical baseline (the naive "lock everything you already hold" heuristic, evaluated via 5,000 Monte Carlo trials):
| Target | Description | Expected Cost | Strategy | Feathers Saved |
|---|---|---|---|---|
| A | All 5 slots gold (any series) | 388.06 | Defer locking until |
33.4% |
| B | All 5 slots gold in 1 specific series | 864.05 | Lock color immediately ( |
51.5% |
| C | 4 ordinary same series + 1 color any gold | 627.63 | Smart-route locks to current max ordinary series; never lock color slot alone ( |
43.2% |
-
Never Lock Early (
$n < 150$ ): Because rollback keeps past progress safe for free, the exponential cost$2^t$ severely penalizes locking during early low-probability stages ($p < 7.7%$ ). -
Target B Bottleneck: The color slot (
$p_c = p/10$ ) is the extreme bottleneck; once rolled, it must be locked immediately from$n=0$ , while ordinary locks should still be deferred. -
Target C Color Degeneracy: Holding a color slot before completing the 4 ordinary same-series slots provides zero cost reduction (
$V(1, m) = V(0, m)$ for$m < 4$ ). Hence, locking the color slot alone is strictly suboptimal.
-
Target A (1D State): Admits an upper-triangular state transition structure. Solved via closed-form back-substitution at the boundary and reverse DP across
$n$ . -
Target B (2D State): Bi-directional transitions break the triangular structure. Solved via boundary Value Iteration and reverse DP across
$n$ . -
Target C (4D State): High-dimensional state space (
$2 \times 35 = 70$ states per wash count). Solved via exact Policy Iteration (linear system matrix inversion$(I - P)V = c$ ) at the boundary to eliminate extreme convergence iterations, followed by reverse DP.
Run any target script to compute the exact expected costs and policy tables:
python target_a.py
python target_b.py
python target_c.py
Run 5,000 simulation trials per target comparing naive play against optimal DP values:
python baseline.py
Query the optimal locking action
# Example: c=1 (color held), na=2, nb=1, nc=0 at wash count n=150
python query.py -c 1 --na 2 --nb 1 --nc 0 -n 150
Programmatic usage in Python:
from query import query
# query(c, na, nb, nc, n) -> (t_c, t_n)
tc, tn = query(1, 0, 0, 0, 100)
print(f"Lock color: {tc}, Lock ordinary: {tn}")
# Output: Lock color: 0, Lock ordinary: 0 (defer locks at low n)├── baseline.py # Monte Carlo simulator for naive baseline policies
├── common.py # Shared game constants and dynamic probability function p(n)
├── target_a.py # Target A DP (closed-form back-substitution)
├── target_b.py # Target B DP (value iteration)
├── target_c.py # Target C DP (policy iteration + matrix inversion)
├── query.py # O(1) policy lookup CLI and module interface
├── report.md # Comprehensive technical report (mathematical derivations)
└── README_zh.md # Chinese documentation
For full mathematical derivations, Bellman formulations, boundary condition proofs, and detailed value tables, refer to:
This project is licensed under the MIT License - see the LICENSE file for details.