Deep Hedging Under Realistic Market Frictions: A Regime-Conditional Empirical Study of Dynamic Option Hedging on Bitcoin Options¶
Bibliographic record. Follow the original-source link for the publication.
| Field | Value |
|---|---|
| Primary domain | Hedging Exposure Risk |
| Other domains | Execution Costs |
| Methods | Financial Ml |
| Facets | — |
| Authors | Sheryan Kumar |
| Published | 2026-08-29 |
| Source | arXiv Quantitative Finance History |
| Identifiers | arxiv:2608.29025 |
| URL | Open original source |
Editorial synthesis¶
Why it matters¶
The study puts cost-aware classical benchmarks and neural policies under the same hourly short-option hedging and cost rules, providing a concrete check on benchmark fairness and sparse trading. It is a boundary test for particular BTC data and implementations, not evidence that deep hedging generally fails. (full_text:S19, full_text:S77, full_text:S80, full_text:S122, full_text:S188, full_text:S205, full_text:S207)
Main author claims¶
- The author reports a Whalley-Wilmott minus BS cost difference of -1.79 USD per episode, with 95% CI [-2.21, -1.39] and p < 0.0001; the mean P&L difference is +8.66 USD, with 95% CI [-3.36, 20.74] and p = 0.164. Reported evidence is stronger for cost savings than for a P&L advantage. (
full_text:S129,full_text:S136,full_text:S138) - The author reports about 20.44 trades for each neural configuration versus 2.47 for Whalley-Wilmott. Adding the LSTM turnover penalty lowers mean turnover from 0.685 to 0.647 without reducing trading frequency. Limited data and missing structural sparsity are proposed explanations for the failure to learn no-trade behavior, not identified causes. (
full_text:S127,full_text:S156,full_text:S174,full_text:S175,full_text:S178,full_text:S184) - The author reports lower Whalley-Wilmott mean cost in the calmer validation period but no mean P&L advantage over the two classical baselines, interpreting the P&L result as regime-linked. The author explicitly describes this comparison as in-sample evidence. (
full_text:S163,full_text:S164,full_text:S165,full_text:S189)
Data, method, or discussion scope¶
The author reports 61 first-of-month day windows, with each contract-day forming a short BTC option episode of 4-24 hourly steps. Reported chronological splits are train 2019-12 through 2022-12, validation 2023-01 through 2023-08, and test 2023-09 through 2024-12; the 11,546 test episodes cover only 16 sampled days. Standardization uses training statistics and checkpoint selection uses validation loss. Terminal P&L includes fixed BTC hedge costs of 5 bp round-trip, or 2.5 bp per trade; training uses CVaR(95%) of loss = -P&L, with turnover-penalty variants. Inference uses 5,000 day-block bootstrap resamples. The target is short-horizon mark-to-market hedging error, not continuous portfolio returns through expiry. (full_text:S47, full_text:S51, full_text:S52, full_text:S73, full_text:S74, full_text:S77, full_text:S78, full_text:S79, full_text:S80, full_text:S99, full_text:S106, full_text:S107, full_text:S109, full_text:S111, full_text:S118, full_text:S119, full_text:S121)
Main limitations¶
The author acknowledges one rally-dominated test window, sparse day coverage, fixed costs, by-inspection lambda = 60, limited data and two neural architectures; no all-asset or all-deep-hedging conclusion is warranted. Reader-inferred limits: day-block resampling handles shared intraday paths but does not verify cross-day dependence; validation already used for checkpoint selection is not untouched regime replication. The feedforward comparison changes both architecture and penalty strength, so a nonsignificant difference is neither equivalence nor exclusion of those explanations. Equal cost rules do not establish equally controlled tuning; the data boundary for lambda inspection is unspecified. Drawdown from cumulative daily mean episode P&L is not a continuous executable portfolio drawdown. The reported CVaR sign convention needs clarification: Eq. 9 defines CVaR using loss = -P&L, whereas Tables 1 and 4 show negative values. That ambiguity alone does not establish invalid inference. (full_text:S93, full_text:S106, full_text:S110, full_text:S111, full_text:S114, full_text:S115, full_text:S116, full_text:S118, full_text:S127, full_text:S149, full_text:S150, full_text:S164, full_text:S165, full_text:S174, full_text:S186, full_text:S187, full_text:S192, full_text:S194, full_text:S195, full_text:S197, full_text:S200, full_text:S201, full_text:S203, full_text:S204, full_text:S205, full_text:S207, full_text:S208)
Relationships¶
- None recorded.