Skip to content

Deep Hedging Under Realistic Market Frictions: A Regime-Conditional Empirical Study of Dynamic Option Hedging on Bitcoin Options

Bibliographic record. Follow the original-source link for the publication.

Field Value
Primary domain Hedging Exposure Risk
Other domains Execution Costs
Methods Financial Ml
Facets —
Authors Sheryan Kumar
Published 2026-08-29
Source arXiv Quantitative Finance History
Identifiers arxiv:2608.29025
URL Open original source

Editorial synthesis

Why it matters

The study puts cost-aware classical benchmarks and neural policies under the same hourly short-option hedging and cost rules, providing a concrete check on benchmark fairness and sparse trading. It is a boundary test for particular BTC data and implementations, not evidence that deep hedging generally fails. (full_text:S19, full_text:S77, full_text:S80, full_text:S122, full_text:S188, full_text:S205, full_text:S207)

Main author claims

  • The author reports a Whalley-Wilmott minus BS cost difference of -1.79 USD per episode, with 95% CI [-2.21, -1.39] and p < 0.0001; the mean P&L difference is +8.66 USD, with 95% CI [-3.36, 20.74] and p = 0.164. Reported evidence is stronger for cost savings than for a P&L advantage. (full_text:S129, full_text:S136, full_text:S138)
  • The author reports about 20.44 trades for each neural configuration versus 2.47 for Whalley-Wilmott. Adding the LSTM turnover penalty lowers mean turnover from 0.685 to 0.647 without reducing trading frequency. Limited data and missing structural sparsity are proposed explanations for the failure to learn no-trade behavior, not identified causes. (full_text:S127, full_text:S156, full_text:S174, full_text:S175, full_text:S178, full_text:S184)
  • The author reports lower Whalley-Wilmott mean cost in the calmer validation period but no mean P&L advantage over the two classical baselines, interpreting the P&L result as regime-linked. The author explicitly describes this comparison as in-sample evidence. (full_text:S163, full_text:S164, full_text:S165, full_text:S189)

Data, method, or discussion scope

The author reports 61 first-of-month day windows, with each contract-day forming a short BTC option episode of 4-24 hourly steps. Reported chronological splits are train 2019-12 through 2022-12, validation 2023-01 through 2023-08, and test 2023-09 through 2024-12; the 11,546 test episodes cover only 16 sampled days. Standardization uses training statistics and checkpoint selection uses validation loss. Terminal P&L includes fixed BTC hedge costs of 5 bp round-trip, or 2.5 bp per trade; training uses CVaR(95%) of loss = -P&L, with turnover-penalty variants. Inference uses 5,000 day-block bootstrap resamples. The target is short-horizon mark-to-market hedging error, not continuous portfolio returns through expiry. (full_text:S47, full_text:S51, full_text:S52, full_text:S73, full_text:S74, full_text:S77, full_text:S78, full_text:S79, full_text:S80, full_text:S99, full_text:S106, full_text:S107, full_text:S109, full_text:S111, full_text:S118, full_text:S119, full_text:S121)

Main limitations

The author acknowledges one rally-dominated test window, sparse day coverage, fixed costs, by-inspection lambda = 60, limited data and two neural architectures; no all-asset or all-deep-hedging conclusion is warranted. Reader-inferred limits: day-block resampling handles shared intraday paths but does not verify cross-day dependence; validation already used for checkpoint selection is not untouched regime replication. The feedforward comparison changes both architecture and penalty strength, so a nonsignificant difference is neither equivalence nor exclusion of those explanations. Equal cost rules do not establish equally controlled tuning; the data boundary for lambda inspection is unspecified. Drawdown from cumulative daily mean episode P&L is not a continuous executable portfolio drawdown. The reported CVaR sign convention needs clarification: Eq. 9 defines CVaR using loss = -P&L, whereas Tables 1 and 4 show negative values. That ambiguity alone does not establish invalid inference. (full_text:S93, full_text:S106, full_text:S110, full_text:S111, full_text:S114, full_text:S115, full_text:S116, full_text:S118, full_text:S127, full_text:S149, full_text:S150, full_text:S164, full_text:S165, full_text:S174, full_text:S186, full_text:S187, full_text:S192, full_text:S194, full_text:S195, full_text:S197, full_text:S200, full_text:S201, full_text:S203, full_text:S204, full_text:S205, full_text:S207, full_text:S208)

Relationships

  • None recorded.