Doubt-driven review · 26 August 2026
A complete, sourced, adversarial review of the Bial Foundation application on infraslow brain rhythms and the conscious registration of nocturnal hot flashes. Twenty agents, 197 raw findings, every one put to a refuter whose job was to prove it wrong.
The verdict
Not fundable as written, and the reasons are structural rather than editorial. The idea is genuinely original and squarely psychophysiological: substitute an endogenous, frequent, peripherally measurable interoceptive event for the experimenter's tone, and ask whether an infraslow brain rhythm decides which such events are noticed. But the confirmatory instrument cannot answer it in either direction. A positive result is manufactured by the design — a zero-phase filter estimates the predictor from EEG recorded after the outcome, and the outcome is defined by a change in the same signal — and a null is uninterpretable, because there is no positive control, no bounded onset error, and the predictor's measurability at frontopolar derivations is contradicted by both papers cited as feasibility support. Underneath that, the application contains no expected event count, no effect size and no power figure, and its single quantitative power claim is arithmetically the reverse of how jitter acts; on the review's own yield model the design is short by roughly an order of magnitude for any effect size comparable to published phase-gating magnitudes. Finally, the 222-woman "cohort" is an undisclosed four-arm randomised trial of hormone therapy and CBT-I in women meeting a clinical insomnia threshold, the analysed timepoint is never named, and the primary outcome measure does not appear in that trial's approved protocol — so the reviewer discovers from the applicants' own reference list both the eligibility question and the possibility that the outcome variable does not yet exist in the data. All four panellists recommend revise and resubmit; so do I, and what is required is a redesign plus pilot numbers, not a rewrite.
Mock Bial panel
| Reviewer | Significance | Innovation | Approach | Feasibility | Team | Fit to Bial | Overall | Recommendation |
|---|---|---|---|---|---|---|---|---|
| Sleep neurophysiologist | 3 | 3 | 8 | 8 | 4 | 6 | 7 | revise and resubmit |
| Consciousness scientist | 4 | 3 | 8 | 8 | 5 | 6 | 7 | revise and resubmit |
| Methodologist / statistician | 3 | 3 | 8 | 8 | 4 | 5 | 7 | revise and resubmit |
| Menopause clinician | 3 | 2 | 7 | 8 | 4 | 6 | 7 | revise and resubmit |
NIH convention: 1 is exceptional, 9 is poor. Lower is better.
Where they agreed and where they split
Unanimous on the shape of the problem and on the recommendation. All four score sheets read identically in direction: significance and innovation at the strong end (2-4), approach and feasibility at the weak end (7-8), team 4-5, fit 5-6, overall 7 in every case, revise and resubmit from all four. The spread between innovation and approach is the whole review — a strong idea attached to an instrument that cannot make the measurement. Four independent convergences carry the most weight, because the reviewers reached them by different routes. Every reviewer concluded that a POSITIVE result would be uninterpretable: the neurophysiologist and the statistician via the non-causal filter, the consciousness reviewer via the press outcome, the clinician via unbounded onset error. All four flagged the absent power arithmetic. All four flagged the undisclosed nesting inside the MHT/CBT-I trial. Three of four independently reached the same conclusion about the inverted confirmatory hierarchy. Where they genuinely split, and who is right: (1) What should be primary. The statistician and the neurophysiologist want the hierarchy inverted so cortical arousal is primary, on the ground that it carries roughly thirteen times the information at the same event count. The consciousness reviewer resists this in substance: arousal-primary is a replication of Lecci et al. with a spontaneous rather than a delivered event, teaches nothing about conscious access, and so weakens rather than strengthens the fit to Bial's remit. Both are right about different things, and the fix must satisfy both: invert the hierarchy for what is estimable at this yield AND add a no-report cortical outcome (a flash-locked evoked response and a flash-locked change in the heartbeat evoked potential, following the draft's own cited Cataldi et al.) as a third rung. Inverting without the third rung leaves a well-powered replication; adding the third rung without inverting leaves an underpowered claim. The consciousness reviewer is right that this is the change that decides whether the project belongs in this competition, and it is cheap because it uses recordings already made — though frontal-only EEG with no EOG limits it, so it must be presented as an exploratory rung and a feasibility outcome, not as a solved measure. (2) Whether the frontal montage is fatal or revisable. The neurophysiologist calls it fatal; the other three call it major. The neurophysiologist is right and is the domain authority: the evidence is one-directional (Lazar reports frontopolar lowest of any region, Lecci recorded C3/C4, Dimitriades locates the peak centro-parietally), and the draft's stated fallback is computed from the same failing signal. This is not fixable by writing. Either a central derivation is added, or the primary aim is demoted. (3) Whether the Article 16(6) therapy exclusion is decisive. The clinician says it is arguable and that the real problem is candour rather than scope; the verified findings support that, noting that Bial's own funded portfolio includes clinically adjacent work. The clinician is right. The exclusion is a high-probability screening risk that the draft simply fails to argue, and the damage is done by non-disclosure rather than by the nesting itself. (4) Form compliance. All four treat it as minor or moderate. On the regulation they are all wrong on one item: the absence of a designated Host Entity is an eligibility bar under Art. 16(7), not a scoring deduction, so it outranks every scientific objection in the fix order despite being the cheapest thing in the document to fix. That collective under-weighting is the most dangerous blind spot in the four reviews. One further point no reviewer made forcefully enough: the clinician's proposal to screen on measured nocturnal flash rate rather than on self-reported symptom interference is the cleanest single improvement in the entire panel. It removes a selection criterion that is the very measure whose unreliability is the proposal's premise, and it aligns the analysis sample with what the analysis actually needs. Adopt it.
Each reviewer in full
Summary
This is a good question attached to a measurement that, as specified, cannot be made. The conceptual move is genuinely original and I want it to succeed: use a spontaneous, frequent, peripherally measurable interoceptive event to test infraslow gating, which escapes the stimulus-randomisation design that the whole arousability literature is built on, and delivers the one cell no delivered stimulus can — the quiet event with neither arousal nor report. I would fund that idea. I cannot fund this instrument. Three defects are, in my judgement, disqualifying as written. First, the predictor is estimated from the one scalp location where the phenomenon is known to be weakest: the infraslow sigma fluctuation is centro-parietal, and Lazar et al. (2019) — cited in the draft *as feasibility support* — reports that frontopolar sites have significantly lower integrated infraslow power than every other region, with parietal highest; Lecci's human recordings were C3/C4 and Carro-Domínguez et al. (2025) measured sigma at Cz. The ZMax gives F7-Fpz and F8-Fpz. The fallback for weak frontal expression is the amplitude of the same frontal signal, which inherits the failure it is meant to escape. Second, the phase estimate is non-causal. A zero-phase band-pass at 0.01-0.04 Hz has support of tens of seconds on both sides of the onset sample, so post-onset EEG determines "phase at onset" — and post-onset EEG is exactly where a registered or aroused flash produces spindle loss, movement and muscle artefact, on two frontal derivations with no EOG or EMG, with the arousal-scoring band (>16 Hz) abutting the predictor band (11-16 Hz). That single choice manufactures an association between "fragility at onset" and registration with zero true gating, in precisely the predicted direction. The draft's long circularity discussion is about the flash detector and never touches this, which is the larger circularity by far. Third, onset timing. The draft correctly names a ~12 s tolerance as its own dealbreaker and then supplies no evidence that any link in the chain achieves it: the cited detector emits a decision every 15 s, computes features over ±250 s and ±500 s windows, and was scored against expert annotations within a ±90 s matching window — seven times the budget — against a gold standard rule defined on a 30 s rise. The proposed precision check (dispersion of intervals between peripheral markers) measures agreement between two effectors, not error against the central event, and is entirely blind to systematic lag, which rotates rather than shrinks the estimate at 7.2° per second. Beyond these: the directional prediction is probably inverted, and the draft's own named collaborator is second author on the human paper that contradicts it (Dimitriades et al. 2026, n=154: microarousals cluster at the ISFS *peak*, p<0.0001 in every age group); "fragility = low sigma power" misreads Lecci and Osorio-Forero, both of whom define the arousable phase by the sigma *derivative*. The inclusion criterion that governs the entire analysis sample ("stable NREM") is never defined, and its relaxed fallback (two cycles, ~100 s) is shorter than the one-cycle period of the passband's own low edge and less than half the ≥256-280 s minimum both published human ISF studies require. There is no event count anywhere in the document, and the single quantitative power claim is mathematically backwards. The confirmatory test is the least informative of the three analyses proposed and cannot demonstrate the advertised claim. And the cohort is an undisclosed four-arm randomised trial of hormone therapy and CBT-I in women with ISI≥10, described as "physically healthy", studied "without intervention", with a primary outcome measure (the button press) that does not appear in the parent protocol at all.
Objections raised
Feasibility is supported: sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023), and the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019). [DRAFT.md L271-274]
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase. [DRAFT.md L267-269]
Because the ISF cycle is only about 50 s, onset uncertainty translates directly into phase uncertainty... Intervals between this inflection and the accompanying temperature and pulse-rate changes index achievable precision. [DRAFT.md L258-263]
Extending the exogenous literature, flashes arriving in fragility (low sigma power) are predicted to be more likely registered, and followed by an arousal, than those in continuity. [DRAFT.md L355-357]
Analyses are restricted to stable (and non-artifactual) NREM. [DRAFT.md L270] ... Low event yield is met by relaxing the stable-bout criterion from three ISF cycles to two, pre-specified as secondary. [DRAFT.md L343-345]
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle. [DRAFT.md L315-318]
The primary registration model is confirmatory; all else is exploratory. [DRAFT.md L310-311]
They are physically healthy and take no hormone-altering medication [DRAFT.md L199-200] ... in healthy humans and without intervention [DRAFT.md L369-370]
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration. [DRAFT.md L215-217]
Flashes are detected from multi-sensor features of a research-grade wristband (Tsiartas et al., 2021). [DRAFT.md L65-66] ... reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance [DRAFT.md L244-245]
the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019) [DRAFT.md L273-274]
Recordings are scored in 30-s epochs (wake, REM, N1-N3) by an experienced scorer... Arousals follow standard criteria: transient > 16 Hz activity lasting 3-15 s (Wassing et al., 2019). [DRAFT.md L234-237]
the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis; weak frontal ISF expression triggers the same rule. [DRAFT.md L341-343]
EEG and wristband are synchronised by cross-correlating accelerometer and PPG signals (Sikder et al., 2026), with a fixed offset and linear drift estimated per recording; tap markers validate this and residual error is reported. [DRAFT.md L227-230]
Because the cycle is short, the binding constraint is not event count but the precision with which onset can be timed against it. [DRAFT.md L79-80]
Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing. [DRAFT.md L259-261]
the flash-arousal ordering reverses across the night: flashes precede arousals in the first half, arousals precede flashes in the second. [DRAFT.md L145-147]
This literature rests entirely on stimuli delivered by an experimenter, a design that deliberately randomises stimulus timing against brain state. [DRAFT.md L115-117]
Lecci et al. (2017) described an oscillation of sigma-band (11-16 Hz) power at approximately 0.02 Hz [DRAFT.md L95-96] ... the Hilbert transform to calculate the IAF phase [DRAFT.md L269]
why women report far fewer nocturnal hot flashes than they have, and why treating the flashes does not reliably resolve the sleep complaint that accompanies them. [DRAFT.md L371-373]
What would change my mind
Pilot data, not promises, on the three things the study rests on — and I mean numbers from these women on these recordings, in the application. (1) Frontal ISF, demonstrated: per-participant infraslow spectral peak in the log sigma envelope from F7/F8-Fpz in at least ten of the cohort's existing nights, tested against interval-shuffled surrogates that preserve spindle count and individual refractoriness, and against a concurrent central derivation in a small simultaneous-PSG subsample, with the phase-error distribution between frontal and central reported. If the frontal envelope does not carry a peak that beats surrogates, no amount of statistical care rescues this design and the honest move is to add a central channel or a device that offers one. (2) Onset error, measured against a reference: one laboratory night per a handful of participants with simultaneous sternal skin conductance, reporting the bias and SD of the change-point back-dated onset against expert-scored gold standard, expressed as a fraction of each participant's own period rather than a nominal 50 s. Bias matters more than variance here and only a reference standard can see it: a systematic lag rotates the estimated preferred phase 7.2° per second, so without a bias bound under a few seconds the direction of your result is not interpretable, only its presence. (3) A causal phase estimator plus a negative control: phase estimated from a pre-onset window only, and the primary model recomputed with phase taken from a matched epoch one full cycle earlier, where a true gating effect must vanish. If the earlier-cycle control shows the same association, the finding is filter leakage from the outcome and we will both know it before any data are unblinded. Alongside those: invert the analysis hierarchy so cortical arousal among all detected flashes and registration among aroused flashes are the two pre-registered confirmatory tests, with a directionally agnostic prediction citing Dimitriades et al. (2026) and Carro-Domínguez et al. (2025) rather than a prediction of low-sigma fragility that the human literature contradicts; define stable NREM as ≥280 s pre-onset only, following the published human ISF studies, and meet low yield with more nights rather than a shorter window; supply the yield table and a real power curve with a named minimum effect of interest; put an operational definition on the button-press outcome; state that scoring is blinded and locked before the merge; and disclose the parent trial, name the timepoint, name the Host Entity, and give me a budget in Euro that makes clear Bial is funding the analysis and the added measurement rather than data collection NWO has already paid for. Do that and this becomes a proposal I would argue for in the room — the question is worth it, and nobody else is positioned to ask it.
Summary
This proposal has a genuinely good idea inside it: replace the experimenter's tone with an event the body generates on its own schedule, and ask whether an endogenous brain rhythm decides which of those events is noticed. That is a real advance in paradigm, and it is the kind of mind-body question Bial exists to fund — a mental event (noticing) set against concurrently recorded physical responses (sudomotor, vascular, cardiac) and brain state. I want to be persuaded by it. I am not, as written, for three reasons that sit squarely in my remit rather than in the statistics. First, the proposal claims to join "consciousness science" and never engages any theory of conscious access. It names no framework, derives no prediction from one, and specifies no result that would adjudicate between any two. Run the frameworks through it and each absorbs a positive result without learning anything: on a global-workspace account report requires ignition and NREM suppresses it, so a phase effect on a press is a phase effect on the transition to wake; on higher-order accounts a null on the press says nothing about first-order interoceptive representation; on recurrent-processing accounts most flashes are felt and never pressed, so press-versus-no-press indexes access and attention, not experience. The draft itself concedes the collapse — "Because a button press requires sufficient arousal, registered flashes are expected to be largely a subset of aroused ones" — and then designates that near-nested outcome as the confirmatory test while relegating the one analysis that could separate access from arousal to the exploratory tier. Second, the measure cannot bear the interpretation. A single binary press "made as soon as the participant thinks she is having a flash" is a conjunction of interoceptive signal, arousal sufficient for volitional action, correct causal attribution of a bodily state to a flash, motor execution and compliance, with no confidence rating, no catch trials, and no rule for presses with no matched flash — therefore no false-alarm estimate and no way to separate perceptual sensitivity from response criterion. In the one adjacent literature that has done this properly — cardiac and respiratory phase modulation of conscious tactile detection — the authors needed d-prime, criterion and meta-d-prime/d-prime to make the claim at all, and found effects of roughly four percentage points in 41 fully awake participants. This design cannot compute any of those quantities. A phase effect on presses is exactly as well explained by a phase-dependent willingness to wake up and act as by a phase-dependent gate on awareness, and the proposal has no instrument that tells the two apart. Third, a construct problem the proposal never sees. A hot flash is not an afferent signal arriving at a gate; it is a heat-dissipation effector response commanded by the hypothalamus, accompanied by a fourfold rise in skin sympathetic nerve activity. What the woman notices is the re-afferent consequence of her own brain's autonomic output. The proposal then hypothesises that the same infraslow noradrenergic system issues that command, sets the arousal threshold, and gates the percept. On its own mechanistic story, "the body reaching the mind" is the brain detecting its own echo — and a positive result is not distinguishable from the brain having shouted louder at itself. The draft half-registers this as an occurrence confound and states the right estimand in the literature review ("comparable events have different probabilities"), then abandons it in the analysis section by making the magnitude-unadjusted total effect confirmatory. Around these, the technical work has already found the failures that would break this on contact with data, and I endorse the important ones: a zero-phase filter estimates the predictor from EEG recorded after the outcome, which manufactures the predicted association; the ISF is centro-parietal and this device is frontopolar; the human evidence puts arousal markers at the sigma peak, so the pre-registered direction is probably inverted; the detector comes from three postmenopausal women on different hardware evaluated at plus-or-minus 90 seconds against a 50-second cycle; and there is no power analysis, no budget, no schedule, no Host Entity, and no section called Specific Aims. The cohort is an undisclosed four-arm hormone-therapy and CBT-I trial in women with clinical insomnia, which engages Bial's therapy exclusion and makes "physically healthy and take no hormone-altering medication" and "without intervention" untrue as written. The mechanistic section is the weakest kind of padding — mechanism-shopping. The KNDy-to-median-preoptic noradrenergic step is not in the source cited for it, all five major neuromodulators oscillate synchronously at 0.02 Hz so nothing selects noradrenaline, and the sign of the noradrenaline-sigma relation is misread. None of it is needed. The question stands or falls on the psychophysics, and I would rather read a page of psychophysics than a page of mouse locus coeruleus. Revise and resubmit. The paradigm is worth funding; this instrumentation of it is not.
Objections raised
It brings together sleep neuroscience, neuroendocrinology and consciousness science around a question about the sleeping human: when does the body reach the mind during sleep?
Conscious registration is indexed by a time-stamped button press made as soon as the participant thinks she is having a flash; morning reports give a complementary measure of recall and compliance.
Because a button press requires sufficient arousal, registered flashes are expected to be largely a subset of aroused ones; the joint distribution is reported and registration additionally modelled among aroused flashes.
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase.
That nucleus is also where autonomic thermoregulation and sleep-wake control converge, and it receives noradrenergic input, raising the possibility that a single infraslow rhythm times arousability and thermoeffector drive together (Rothhaas & Chung, 2021).
Cataldi et al. (2026) compared auditory and heartbeat evoked potentials across wakefulness and REM microstates and found a dissociation, with auditory responses declining progressively into phasic REM while cardiac responses were preserved and enhanced.
Extending the exogenous literature, flashes arriving in fragility (low sigma power) are predicted to be more likely registered, and followed by an arousal, than those in continuity.
the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis; weak frontal ISF expression triggers the same rule.
Interestingly enough, just as with external signals, they are sometimes consciously registered, and other times leave no trace (Gombert-Labedens et al., 2025).
In rodents this fluctuation tracks locus coeruleus (LC) activity.
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration.
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle.
Feasibility is supported: sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023), and the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019).
They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms.
It would also give a mechanistic account of a clinical puzzle - why women report far fewer nocturnal hot flashes than they have, and why treating the flashes does not reliably resolve the sleep complaint that accompanies them.
This literature rests entirely on stimuli delivered by an experimenter, a design that deliberately randomises stimulus timing against brain state.
Because the cycle is short, the binding constraint is not event count but the precision with which onset can be timed against it.
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference [number]).
What would change my mind
Four changes, in priority order, and I would then advocate for this. 1. Add a no-report outcome to the same events. Time-lock an evoked response to the electrodermal change point, and measure a flash-locked change in the heartbeat-evoked potential, following the draft's own cited Cataldi et al. (2026) — whose authors propose exactly this class of index for contexts where behavioural responsiveness cannot be assessed. That converts a one-rung design into three rungs (cortical response to the interoceptive event, arousal, report), gives a measure of the signal reaching cortex that does not require waking, and is the only version of this project that can address interoceptive access rather than arousal threshold. The relevant expertise is in Lausanne, where the claimed rodent collaborator already sits. 2. Make the report measure criterion-controlled. Add a confidence rating to each press, define the matching window and the treatment of unmatched presses so a false-alarm rate exists, and pre-specify a sensitivity-versus-criterion analysis. Without a false-alarm estimate there is no d-prime, no criterion, and therefore no defensible claim that phase moved awareness rather than the threshold for acting. The wakefulness literature on cardiac-phase gating of conscious detection did this and found effects of a few percentage points; anchor the power simulation to that magnitude. 3. Invert the confirmatory hierarchy and estimate the predictor causally. Pre-register two confirmatory tests: phase to cortical arousal among all detected flashes (the replication, and the informative one), and phase to registration among aroused flashes or as a natural direct effect in a mediation model with arousal as mediator (the quantity that isolates access). Estimate ISF phase from pre-onset data only with a causal estimator, and demonstrate on an outcome-shuffled surrogate and on a phase estimate taken one full cycle earlier that no association survives. As it stands the confirmatory test is a replication of phase-dependent arousability, and the predictor is contaminated by the outcome. 4. Say what theory this adjudicates, in three sentences, and name a result that would embarrass one account and favour another. Then delete the KNDy and locus coeruleus apparatus, which is decorative, partly unsupported by the sources cited for it, and does not select noradrenaline over the four other neuromodulators oscillating at the same frequency. Also fix the direction: state the phases by the derivative of sigma as the source work does, cite the human evidence placing arousal markers at the peak, and let the sin/cos test remain two-sided. Alongside these: disclose the host randomised trial, restrict to the pre-randomisation block, add a Host Entity, a euro budget, a schedule, CVs, a Specific Aims section, the real ethics reference with the amendment status of the button press, and the missing Lecci reference. Add the two zero-cost positive controls so that a null is interpretable. If the onset-precision gate cannot be closed with existing evidence, submit this as a feasibility and effect-estimation study with a stated kill criterion — I would score that far more generously than a confirmatory framing the applicants themselves cannot yet support.
Summary
A genuinely good question wrapped around a confirmatory design that cannot answer it in either direction. The proposal asks whether the phase of the ~0.02 Hz infraslow sigma fluctuation at nocturnal hot flash onset predicts whether the flash is consciously registered. That is novel, mechanistically motivated, and squarely psychophysiological. But the statistical core is absent and the identification is broken. Three things sink it as written. First, there is no expected number of analysable events anywhere in the document. The exclusion cascade is named at L204-206 and never computed; no event count, no effect size, no power figure, no minimum detectable effect appears. I built the yield chain myself from the draft's own numbers (137 women x 4 nights x ~0.6 usable EEG nights x 3.5 flashes/night x 0.75 not-already-awake x 0.85 NREM x ~0.5 stable-bout x 0.8 artefact/onset) and get roughly 300 analysable flashes, of which 45-125 registered. My own calculation of the 2-df sin/cos likelihood-ratio test gives, for an odds ratio of 2 between best and worst phase, a requirement of 1,258 events at zero jitter and 1,867 at 5 s onset jitter (press rate 0.15), or 654 and 971 at the draft's implied 43% report rate. The minimum detectable effect at N=300 with 5 s jitter is an odds ratio of 3.5 to 5.6. That is an implausibly large phase-gating effect: the closest published benchmark, cardiac and respiratory phase modulation of conscious tactile detection in wakefulness, is circular R around 0.34. The design is short by a factor of three to nine for any effect size a reviewer would find credible. Second, the one quantitative power claim in the document is mathematically false, and it is precisely the claim that licenses stating no N. Independent onset jitter of SD s against period T attenuates the first circular harmonic by exp(-(2*pi*s/T)^2/2); it shrinks the effect but leaves the estimator consistent, so power is strictly increasing in N at every finite jitter and there is no threshold past which events stop helping. I computed the inflation factors: 1.48x at 5 s, 4.04x at 9.4 s, 11.79x at the draft's own quarter-cycle (12.5 s), 34.9x at 15 s. Jitter does not bound power; it multiplies the required event count. Removing the false ceiling removes the argument for not stating N — and both constraints then bind simultaneously, contradicting the draft's own L191-192. Third, and worst, the predictor is estimated with a zero-phase filter. A non-causal band-pass at 0.01-0.04 Hz has an impulse response spanning tens of seconds in both directions, so the phase assigned to a flash onset depends on EEG recorded after that onset. Both outcomes — a cortical arousal and a button press requiring wakefulness — produce post-onset spindle suppression and movement artefact, i.e. a sigma drop, which the filter propagates backwards into the estimated onset phase and pulls it toward the "fragility" pole. The design therefore manufactures the predicted association with no gating whatever, and it manufactures it in exactly the direction the Expected Outcomes section pre-commits to. Combined with the sigma upper edge (16 Hz) abutting the arousal-defining band (>16 Hz) on the same two frontal derivations, this is a validity failure that makes a positive result as uninterpretable as a null. Layered on top: the designated confirmatory test is the least informative analysis in the proposal. Registration is largely a subset of arousal, as the draft itself states at L299-302. I verified the consequence: logit-scale amplitude transmits from the arousal model to the registration model with a factor of 0.353, and after the base-rate difference the registration test carries 7.6% of the non-centrality of the arousal test — it needs 13.2x the events. So the proposal nominates its weakest test as confirmatory, while the only analysis that isolates the advertised claim (registration among aroused flashes) is declared exploratory at L310-311 and conditions on a mediator the draft forbids adjusting for two paragraphs earlier at L306-309. There is no stated estimand, test statistic, or decision rule for the comparison the Summary promises at L54-56. Then the escape hatches. I count eleven named-but-empty decision thresholds, plus a twelfth that is worse: the primary outcome variable has no operational definition at all — no window within which a button press counts as registering a given flash, no rule for unmatched presses, multiply-matched presses, or presses during scored wake, against a flash lasting one to five minutes and a cycle of fifty seconds. Defensible choices of that one window produce materially different confirmatory results from the same data. And the pre-registered fallback (L341-343) substitutes pre-flash ISF amplitude while asserting it tests "the same hypothesis". It does not: amplitude is symmetric across the rising and falling limbs, which is the distinction the mechanism is about, and it is read off the same frontal envelope whose weakness triggers the switch. That sentence converts a null into something reportable as a positive. Finally, a null here would be uninterpretable, which is the design property I am least willing to fund. There is no positive control. A null is equally consistent with no gating, insufficient onset precision, unmeasurable frontal ISF, cross-device synchronisation error, and detector false positives — and the proposal acknowledges at L275-276 that two of those are unresolved. Two cheap controls are available in these exact recordings (recovering the human sigma-to-heart-rate infraslow coupling from ZMax sigma against EmbracePlus PPG; recovering the phase distribution of spontaneous microarousals) and neither is proposed. The feasibility gap at the binding constraint is unbridged: the draft sets its own tolerance at a quarter cycle (~12.5 s, and only ~6.25 s for a participant at the fast end of its own 0.01-0.04 Hz band) while the cited detector decides every 15 s and was scored against a ±90 s matching tolerance in three postmenopausal women on a bench-wired sensor array that is not the study device. The rhythm itself is reported centro-parietally maximal and frontopolar-minimal by both papers cited as feasibility support. And the cohort is not a cohort: 222 is the parent trial's enrolment target, 65 have been screened, and that parent trial is an undisclosed four-arm randomisation to estradiol/progesterone and CBT-I in women with ISI>=10 — so the timepoint supplying the nights determines both the yield (by a factor of three) and whether half the sample is on a drug that suppresses the outcome event by about 75%. The draft never states which timepoint it uses. The science is worth doing. The application is not fundable in this form, and the required changes are not editorial.
Objections raised
Simulations use the observed counts and precision, and the minimum detectable effect is reported.
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle.
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase.
The primary registration model is confirmatory; all else is exploratory.
if it exceeds a pre-registered threshold, the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis; weak frontal ISF expression triggers the same rule.
the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration.
Because the ISF cycle is only about 50 s, onset uncertainty translates directly into phase uncertainty.
Feasibility is supported: sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023), and the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019).
Frontal expression and achievable phase precision are themselves feasibility outcomes.
Analyses are restricted to stable (and non-artifactual) NREM.
Labelling rests on objective physiology, not self-report: the detector is never trained on button presses, which are themselves an outcome, so it cannot be circular with respect to registration.
Flash magnitude is a mediator, not a confounder: if ISF phase modulates the size of the autonomic response, adjusting for it would bias the total effect of phase on registration.
Participants are perimenopausal women (irregular cycles and climacteric symptoms) reporting sleep disturbances, from an ongoing cohort of 222 women already being recorded.
They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms.
Extending the exogenous literature, flashes arriving in fragility (low sigma power) are predicted to be more likely registered, and followed by an arousal, than those in continuity.
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference [number]).
Random intercepts are included for participant and for night within participant, with covariates for sleep stage, time since the preceding flash, and half of the night
Lecci et al. (2017) described an oscillation of sigma-band (11-16 Hz) power at approximately 0.02 Hz, present in both mice and humans during NREM sleep.
Women have on average 3.5 objectively recorded flashes per night but only report 1.5 (De Zambotti et al., 2014).
What would change my mind
Six specific changes, and I would move from "revise and resubmit" to "fund" on the strength of them. They are all specifiable now, without new data. (1) A yield table and a power curve, in the application. One row per exclusion — usable EEG night, onset during sleep, NREM, stable pre-onset bout, artefact-free, reliable onset — each with a retention fraction and a stated source (pilot, published estimate, or declared assumption), ending in a single expected count of analysable flashes and expected registered events, with pessimistic/central/optimistic bounds. Then name a smallest effect worth detecting (an odds ratio between best and worst phase; anchor it to the wakefulness benchmark of circular R around 0.34 rather than picking a convenient number) and report power for it across a jitter grid of 2, 5, 9 and 12.5 s, and a minimum detectable effect at the central yield. If the central yield is a few hundred flashes, as my own arithmetic suggests, then say so and reframe: make cortical arousal the primary outcome, which by my calculation carries 13.2 times the information of registration, and present registration as effect estimation plus feasibility rather than as a confirmatory test. I would fund that honest version. I will not fund a confirmatory test whose detectability the applicants have not computed. (2) A causal phase estimator, plus a negative control that can fail. Estimate ISF phase from the pre-onset log sigma envelope only — a forward-only filter or a windowed sinusoid fit on the interval ending at onset minus a stated guard — so that no post-onset EEG, no flash-related arousal, no spindle suppression and no movement artefact can enter the predictor. Quantify the bias and variance of that estimator against the non-causal one on flash-free stable NREM. Then add the control I would actually look for: recompute the primary model with phase taken from a matched epoch one full ISF cycle earlier. A true gating effect must vanish there; the artefact will not. Separate the arousal-scoring band from the predictor band, and have arousals scored by a rater blind to the sigma envelope, to the wristband channels and to the button presses — none of which the draft currently promises. (3) The estimand stated, and the hierarchy inverted. Pre-register two confirmatory tests: phase to cortical arousal among all analysable flashes (the powered replication), and phase to registration among aroused flashes, or the natural direct effect in a mediation model with arousal as mediator (the test that actually isolates conscious access). Say in one sentence that a phase effect on registration among all detected flashes is produced by arousal gating alone and is therefore not evidence about access. Resolve the inconsistency by which magnitude adjustment is refused on mediator grounds while arousal is conditioned on, and align the estimand with the aim: if the aim is gating of comparable events, the magnitude-matched contrast is the confirmatory quantity, not a sensitivity analysis. (4) Numbers on all eleven blanks, plus the twelfth. Every threshold in the Risks section and the exclusion cascade gets a value: the onset-dispersion criterion in seconds, "weak frontal ISF expression" as a spectral criterion against phase-shuffled surrogates, the low-yield trigger as an event count, the compliance floor as a rate, the maximum acceptable synchronisation residual, the artefact rule, the estimation window, the stable-bout definition (pre-onset only, in seconds and in cycles of the participant's own period), and the minimum detectable effect. Most importantly, define the primary outcome: the window within which a press counts as registering a flash, with pre-registered sensitivity analyses at two other windows, and rules for unmatched, multiply-matched and wake-period presses. And replace the amplitude fallback with a phase-preserving one — the binary sign of the sigma derivative at onset, the variable the team's own rodent work manipulated — or state plainly that amplitude tests a different hypothesis and is secondary. (5) A positive control that gates the interpretation of a null. Reproduce, within participant and before outcomes are inspected, the human infraslow sigma-to-heart-rate coupling from ZMax sigma against EmbracePlus PPG, and the phase distribution of spontaneous microarousals. Both are free in these recordings and between them they validate the frontal phase estimate, the cross-device synchronisation and the achievable resolution — the three things a null would otherwise be blamed on. Pre-commit: if the controls pass, a null is a null and is interpretable; if they fail, the study reports as a feasibility study. State an equivalence bound. A design whose null is interpretable is a fundable design; this one currently is not. (6) The timepoint, and the missing form. One sentence stating that the confirmatory analysis uses the pre-randomisation baseline block only, with the yield recomputed on that basis, and disclosing that the recordings sit inside a randomised trial of hormone therapy and CBT-I in women meeting an insomnia threshold — then arguing eligibility explicitly rather than by adjective, and dropping both "physically healthy" and "without intervention". Plus the elements the regulation requires and the document lacks entirely: a section headed Specific Aims with three numbered, falsifiable objectives; a Host Entity and named PI; CVs; a schedule with a start date inside the 2027 window; and a Euro budget that says plainly that another funder pays for collection and this grant buys the analysis, the amendment and the scoring — which is the strongest argument this application has and the one it never makes. What would NOT move me: better prose, more citations, or a promise to compute the power analysis after funding. The whole problem is that the two quantities the design turns on — expected analysable events and achievable onset precision — are deferred past the funding decision, and the one place a number does appear it is arithmetically wrong in the direction that excuses the deferral.
Summary
This proposal asks a question I have wanted someone to ask for fifteen years. The gap between the nocturnal hot flashes women have and the ones they report is one of the oldest unresolved oddities in my field, and it has been treated almost exclusively as a measurement nuisance or a recall failure. Reframing it as a question about state-dependent access to awareness — and testing it against an endogenous, stereotyped, objectively measurable bodily event rather than an experimenter's tone — is genuinely original and squarely psychophysiological. The choice of the nocturnal flash as the model event is, in my judgement, the best available and nobody has used it this way. What is proposed to test it does not yet exist. The proposal sets its own dealbreaker at roughly 12 s of onset-timing jitter against a 50 s cycle, and there is no method in the vasomotor-symptom literature — including the one cited — that times a flash onset to anything like that resolution. The field's reference definition of a flash is a 2 µS sternal skin-conductance rise within a 30-second window; the cited detector decides in 15 s frames, computes features over windows extending ±500 s, and was scored against expert annotations within a ±90 s matching tolerance in three postmenopausal women contributing 27 laboratory flashes on a bespoke sensor array wired into a PSG amplifier. Not the study device, not the study population, not the study setting, and two orders of magnitude away from the required precision. The peripheral sudomotor inflection is also not the central event, and the latency between them has never been measured. Two further things trouble me as a clinician. The cohort is not disclosed: it is the team's own four-arm randomised trial of transdermal estradiol plus progesterone and online CBT-I in perimenopausal women meeting an insomnia threshold, and the draft calls those women "physically healthy", "with self-reported sleep disturbances", studied "without intervention". Menopausal hormone therapy reduces flash frequency by around 75% and CBT-I directly targets the symptom monitoring on which the primary outcome depends. And the analysis sample is selected on self-reported flash interference with sleep, which is the very report whose unreliability is the proposal's premise, while the analysis needs objectively detected flashes with onsets inside consolidated, arousal-free NREM — a fraction the draft's own cited review puts at well under a third of nocturnal events. No yield figure, no event count, no power figure appears anywhere. I would like to see this study run. I cannot recommend funding it in this form.
Objections raised
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle.
They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms.
The analysis sample comprises those whose nocturnal flashes interfere with sleep: of the first 65 screened, 40 (62%, 95% CI 49-72) meet this criterion, projecting to roughly 137 women.
Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier, reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance, with sleep-onset performance reported separately.
Expected yield is estimated by applying the planned exclusions sequentially - EEG night, NREM, stable NREM, estimation window, artefact-free recording, reliable onset timing.
Both are fitted to every detected flash, so events passing without arousal or report contribute.
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration.
It would also give a mechanistic account of a clinical puzzle - why women report far fewer nocturnal hot flashes than they have, and why treating the flashes does not reliably resolve the sleep complaint that accompanies them.
The design is within-subject, observational and multi-night, conducted at home.
Oestrogen withdrawal hyperactivates arcuate KNDy neurons, which signal through the neurokinin 3 receptor to the median preoptic nucleus and narrow the thermoneutral zone (Gombert-Labedens et al., 2025).
Freedman and Roehrs (2006) showed that hot flashes are suppressed during REM sleep, when thermoregulatory effector responses are largely inactive, and that the flash-arousal ordering reverses across the night: flashes precede arousals in the first half, arousals precede flashes in the second.
Flash magnitude is a mediator, not a confounder: if ISF phase modulates the size of the autonomic response, adjusting for it would bias the total effect of phase on registration.
Extending the exogenous literature, flashes arriving in fragility (low sigma power) are predicted to be more likely registered, and followed by an arousal, than those in continuity.
if it exceeds a pre-registered threshold, the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference \[number\]).
has access to an ongoing cohort of 222 perimenopausal women with self-reported sleep disturbances, many of whom experience regular nocturnal hot flashes
Recent research has given insight into how external events (e.g., sounds) are gated during sleep (Lecci et al., 2017).
What would change my mind
Five specific changes, in order of weight, would move me from revise-and-resubmit to fund. First, and non-negotiable: an onset-precision validation study, run first, gated, and costed. A small in-laboratory or first-night subsample of perimenopausal women recorded on the actual EmbracePlus with simultaneous sternal skin conductance as reference, reporting the full distribution of change-point-to-reference onset error — bias and dispersion separately, because a systematic lag rotates the estimated preferred phase rather than shrinking it, at roughly 7 degrees per second against a 50 s cycle. Then a pre-registered kill criterion: the phase analysis proceeds only if bias is under about 6 s and dispersion under a quarter of each participant's own empirical period. If the applicants told me they could not afford that subsample, I would want them to say so and propose the study as a precision-and-yield feasibility project with effect estimation, which I would look on kindly. A proposal that names its own kill criterion earns far more trust than one that defers it. Second, full disclosure of the parent trial in the body text, in the Participants section, with the timepoint stated. I want one sentence saying the analysis uses the pre-randomisation baseline block only, four EEG nights per woman, no participant on study medication or having completed CBT-I; the phrase 'physically healthy' replaced by an accurate description of the inclusion criteria; 'without intervention' deleted; and the yield recomputed on a single four-night block. I would also want the button-press and tap-marker amendment status stated separately from the parent approval, with a real reference number and a count of women already recorded under the amended protocol. Third, a yield table with a number and a source at every step, ending in an expected count of analysable flashes and of registered flashes, with pessimistic, central and optimistic bounds, and a power curve over that count crossed with a jitter grid for a named minimum effect of interest. Selection changed from self-reported symptom interference to measured nocturnal flash rate from the wristband week. If the central yield is a few hundred flashes and a few dozen registered events, say so and reframe the primary aim accordingly — I would rather fund an honest estimation study than an underpowered confirmatory one. Fourth, the primary outcome fully operationalised: the matching window between a press and an estimated onset with sensitivity analyses at other windows, rules for multiple, unmatched and wake-period presses, the press-latency distribution reported descriptively, and the stable-NREM criterion stated as pre-onset only with arousal and awakening explicitly never used as exclusions. Plus the reactivity check the design already contains for free — flash rate and magnitude on button-press nights versus the three wristband-only nights. Fifth, the smaller but telling corrections: bedside temperature and humidity logging with ambient temperature and a temperature-by-half-of-night term in both models; 'inter-threshold zone' in place of 'thermoneutral zone' with the mixed evidence on sweating thresholds acknowledged; the flash–arousal ordering reversal hedged rather than asserted; the phases restated as rising versus declining sigma with the direction treated as open; the amplitude fallback replaced by a rising/falling binary and relabelled; the closing therapeutic sentence deleted; and the Lecci reference added, the aims section renamed and numbered as Specific Aims, a host entity designated, and a euro budget and schedule supplied that make explicit that the parent grant funds collection and Bial funds the analysis. If those five arrived, I would argue for this in panel. The question is good enough to be worth the wait.
Ranked by what would sink it
A zero-phase band-pass at 0.01-0.04 Hz has support of tens of seconds on both sides of the onset sample, so "phase at onset" depends on post-onset EEG. Both outcomes (a cortical arousal; a press requiring wakefulness) produce post-onset spindle suppression and movement artefact, i.e. a sigma drop, which the filter propagates backwards and pulls the estimated onset phase toward the low-sigma pole the draft labels fragility and predicts should be associated with registration. Compounded by the predictor band (11-16 Hz) abutting the arousal-scoring band (>16 Hz) on the same two frontal derivations, with no EOG or EMG, and by the fact that the word 'blind' appears nowhere in the document.
Why it is decisive It invalidates a positive result as thoroughly as it invalidates a null. The draft devotes eight lines to circularity in the flash detector and never touches this, which is the larger circularity by an order of magnitude. Against a 50 s cycle a few seconds of leakage is tens of degrees of phase.
F63 F180 F130 F129 F49
The infraslow sigma fluctuation is maximal centro-parieto-occipitally; Lazar et al. (2019) specifically report frontopolar sites as having the lowest infraslow power of any region, Lecci's human recordings were C3/C4, and Dimitriades et al. locate the peak centro-parietally. The ZMax provides only F7-Fpz and F8-Fpz and the study has no other EEG. The Esfahani bias statistic is close to the wrong statistic: a static offset in mean relative sigma power is removed entirely by the band-pass, that validation reports no time-resolved or infraslow analysis at all, and gamma's bias is smaller than sigma's, so 'smallest bias of any band' is also false. The stated fallback for weak frontal expression is the amplitude of the same weak frontal envelope, so it inherits the failure it is meant to escape. Separately, whether the fluctuation is a continuous oscillation at all is contested by a PNAS paper (n=1,025) the draft does not cite.
Why it is decisive If the frontal envelope carries no infraslow peak that beats interval-shuffled surrogates, there is no phase to have, the filter will nonetheless return one, and no amount of statistical care rescues the design. This cannot be fixed by writing: it requires a central derivation, a concurrent-PSG subsample, or demotion of the primary aim.
F43 F103 F133 F42 F120 F37 F134 F125
A quarter of a 50 s cycle is 12.5 s, and only 6.3 s for a participant at the fast end of the draft's own 0.01-0.04 Hz band. The cited detector decides in 15 s frames, computes features over windows extending to plus or minus 500 s, and was scored against expert annotations within a plus-or-minus 90 s matching tolerance; the field's reference criterion is itself a 2 microsiemens rise within 30 s. The proposed precision check measures the dispersion of intervals between peripheral markers, which quantifies agreement among effectors, cannot see error common to all of them, and is entirely blind to systematic lag — and systematic lag is the mode that destroys the study, because it rotates the estimated preferred phase at 7.2 degrees per second rather than shrinking it. Wrist tonic electrodermal level during sleep has moreover been reported (in E4 data) to track skin hydration rather than sympathetic tone.
Why it is decisive Onset error enters the power calculation as a multiplier and enters the directional interpretation as a rotation. Without a bias bound below roughly 6 s the Expected Outcomes direction is not recoverable; without a dispersion estimate no power figure can be computed at all. Both require a validation subsample against a reference standard that the design does not currently contain or cost.
F47 F137 F161 F89 F188 F90 F27 F74 F168 F117
Independent onset jitter attenuates the first circular harmonic by exp(-(2*pi*sigma/T)^2/2): it shrinks the effect but leaves the estimator consistent, so power is strictly increasing in N at every jitter level and no threshold exists past which events stop helping. Jitter multiplies the required count (1.5x at 5 s, 4x at 9.4 s, 11.8x at a quarter cycle). Removing the false ceiling removes the licence to state no event count — and the draft states none anywhere in 464 lines, names six sequential exclusions with no retention rate for any of them, and slides between 'simulation shows' (past) and 'simulations use the observed counts' (future). It also contradicts itself: Study design justifies four nights precisely because power depends on event count. On the review's yield model the central scenario is roughly 350 analysable flashes and about 50 registered events, against a requirement near 1,900-3,000 for an odds ratio of 2 at 5 s jitter.
Why it is decisive A confirmatory test is being proposed whose detectability the applicants have not computed, and the closest published magnitude for phase gating of conscious detection is modest (circular R about 0.34 in wakefulness). At the plausible yield the minimum detectable effect is an odds ratio in the range 4-8, which is not a credible gating effect.
F86 F114 F28 F65 F87 F115 F165 F88 F150 F96 F95
P(registered | phase) factorises as P(aroused | phase) x P(registered | aroused, phase). The draft concedes registration is largely a subset of arousal, so a significant primary result is produced by phase-dependent arousability alone — which Lecci et al. published in 2017. Only the second factor isolates access to report over and above arousal, and it is declared exploratory. The information cost is measurable: logit-scale amplitude transmits with a factor of about 0.35, so the registration test carries roughly 8% of the non-centrality of the arousal test and needs about 13x the events. The draft also refuses to adjust for magnitude on mediator grounds and then freely conditions on arousal, an equally post-onset mediator, and its Literature Review states the correct estimand ('comparable events') which the Statistical Analyses section abandons by making the magnitude-unadjusted total effect confirmatory. The Summary promises that 'comparing the two' distinguishes cortical gating from gating of access, and names no estimand, statistic or decision rule.
Why it is decisive As specified, the confirmatory result cannot support the interpretation the proposal is built on, and the proposal is internally inconsistent about its own estimand. This is a pre-registration choice, so it must be settled before submission, not after data.
F64 F93 F67 F185 F155
The 222 women are the applicants' own four-arm 2x2 randomised trial (NCT06306404) of transdermal estradiol plus progesterone and online CBT/circadian therapy for insomnia, enrolling on Insomnia Severity Index >=10 and Greene Climacteric Scale >=13, with the identical device package collected at three timepoints. So 'physically healthy', 'take no hormone-altering medication' and 'in healthy humans and without intervention' are all untrue as written for two of four arms at two of three timepoints; hormone therapy reduces flash frequency by roughly 75%, removing the events the study is powered on, and CBT-I perturbs both the predictor and the symptom monitoring on which the outcome depends. Further, 222 is a recruitment target inclusive of a dropout allowance while 65 have been screened, and the button press, the tap markers and the morning flash count appear nowhere in the parent protocol, so they require an amendment and can exist only for participants recorded after it — while the ethics reference is the literal string '[number]'.
Why it is decisive The timepoint answer changes the available data by a factor of three, determines whether the eligibility argument is available at all, and determines whether the confirmatory outcome variable exists in the recordings the draft describes as 'already being recorded'. Non-disclosure converts a scope question a reviewer might have decided generously into a candour question, because they find the trial in the applicants' own reference list.
F2 F3 F8 F36 F38 F56 F109 F136 F158 F159 F167 F181 F82 F4 F39 F174 F170
Nothing states the window within which a press counts as registering a given flash, the rule for multiple, unmatched or wake-period presses, or the treatment of press latency — against a flash lasting one to five minutes, a cycle of 50 s, and an orienting time that will itself vary with sleep depth, which is the predictor. Unmatched presses are exactly the observations that would yield a false-alarm rate, and they are unaddressed, so neither perceptual sensitivity nor response criterion is identifiable. Meanwhile 'as soon as she thinks she is having a flash' requires waking, noticing warmth, and causally attributing it to a flash rather than to the room or the insomnia. No theory of conscious access is named or tested, and the one no-report interoceptive measure the draft itself cites is not adopted.
Why it is decisive Defensible alternative windows produce materially different confirmatory results from identical data, which is unconstrained analytic freedom on the primary endpoint of a study calling itself pre-registered. And without a no-report rung, every framework of conscious access re-reads a positive result as arousal-threshold modulation, which is the finding the arousal model already delivers.
F70 F196 F126 F186 F76 F99
The only quantity appears seventy lines later in the contingencies as 'three ISF cycles', relaxable to two, with no statement of whether the bout precedes onset, follows it, or brackets it — and because the estimator is non-causal, the bracketing reading is the natural one. Under that reading any flash producing an arousal, stage shift or awakening terminates its own bout and is excluded, deleting most of the events that carry the effect, since roughly 70% of nocturnal flashes are associated with an arousal. The relaxed fallback (about 100 s) equals one period of the passband's own low edge and is under half the >=280 s minimum both published human ISF studies require and justify on frequency-resolution grounds — in a population selected for insomnia, where long uninterrupted N2 bouts are scarce.
Why it is decisive This single undefined clause determines the analysis sample, the yield, and whether the outcome has any variance. It also collides with the Summary's claim that every detected flash contributes.
F18 F71 F94 F148 F80 F54 F149
The word 'host' does not appear in the document; no PI is named; the implied team spans Amsterdam and Lausanne with the rodent-expertise claim asserted rather than evidenced by CVs; there is no Euro figure anywhere against a EUR 60,000 cap under which overheads and wearable-device use costs are non-reimbursable; there is no start date inside the 1 January to 31 October 2027 window; and the objectives sit in flowing prose under 'Research aims' rather than in the section the regulation names as the basis of assessment and of the final report. The Procedure section is written in the present tense as though the applicants run a measurement protocol that NWO has already funded, so a reviewer cannot tell what the money buys — while the single largest plausible cost, manual staging and arousal scoring of roughly 550 nights, appears in no budget.
Why it is decisive One item here can retire the application before any science is read. The complementarity argument that would make a EUR 60k ask unusually strong — another funder pays for collection, Bial buys the analysis, the amendment and the scoring — is the best argument this application has and it is never made.
F5 F6 F7 F9 F57 F127 F162 F164 F172 F173 F177
Fragility in the source work is defined by the sign of the sigma derivative, not its level: arousals followed stimuli delivered during declining sigma, and thalamic noradrenaline peaks before sigma declines. Every subsequent human dataset places arousal markers at or after the sigma peak, including Dimitriades et al. (2026, n=154), whose second author is the collaborator the draft invokes as its expertise claim and whose paper it does not cite. Meanwhile pre-flash ISF amplitude is symmetric across the rising and falling limbs that carry opposite arousability, is confounded with N2/N3, spindle density and time of night, and is read off the same envelope whose weakness triggers the switch — so it neither tests the phase claim nor survives the failure it is invoked for.
Why it is decisive The sin/cos test is direction-agnostic and survives, so this is not fatal to the analysis; it is fatal to the narrative and to the Expected Outcomes commitment against which the final report will be checked. And a contingency plan with eleven blanks plus a hypothesis substitution is not a pre-registration.
F132 F183 F62 F169 F48 F116 F143 F189 F69 F154
Causal, not thematic
Some fixes are pointless until a prior one is settled. Rewriting the power section before the event count exists means rewriting it twice. Item 12 is last in logic and first in wall-clock time, because it is a screening bar that takes an afternoon.
Do not lose these in revision
Only the applicants can answer
Every source checked
For each reference: does it exist as cited, and does the draft's claim survive contact with it. Where a figure could not be verified it is marked so rather than guessed.
| Reference | Metadata | Claim as used | Verdict |
|---|---|---|---|
| Lecci et al. 2017 | Absent from the reference list entirely — cited five times | Sigma ISF at 0.02 Hz; 11–16 Hz band; continuity/fragility; human arousability; memory prediction | Major |
| The load-bearing citation of the whole proposal has no entry. Three separate problems in how it is used: the paper explicitly disclaims the human arousability finding attributed to it (no stimuli were delivered to the human sample); the 11–16 Hz band matches none of the cited sources; and “establishing that it is functionally consequential” rests on a single r = 0.45, p = .027, n = 24 correlation in a nine-electrode subsample of young men. | |||
| Osorio-Forero et al. 2021 | Listed twice as near-identical entries, plus a third copy under own publications | LC noradrenergic activity “rises and dips with this phase”; 0.02 Hz / 50 s | Minor |
| Substance verified against the full text: the 50 s / 0.02 Hz timescale and spindle clustering are correctly reported. Two small things: noradrenaline is anticorrelated with sigma, which “rises and dips with this phase” obscures; and the mechanism rests on thalamic biosensor photometry, which the draft states without qualification. | |||
| Osorio-Forero et al. 2025 | Cited in text but present only under Previous own publications, not in References; volume and pages missing | ~50 s NREM partition; high LC facilitates microarousals; low LC required for NREM→REM | Holds |
| All three claims verified verbatim against the abstract. Nature Neuroscience 28(1):84–96. The “-024-” DOI stem is the online-first year (25 Nov 2024), not an error — do not “fix” it. | |||
| Tsiartas et al. 2021 | Correct | Detector reaches “over 90% sensitivity at 95.6% specificity”; “research-grade wristband” | Major |
| The quotation is faithful and the classifier was indeed a decision tree. What the words omit is the problem: three postmenopausal participants, 27 events, a temperature-controlled laboratory, cross-validation that is not subject-wise, a ±90 s matching tolerance against expert annotation, and a custom array of consumer-grade sensors — which is the inverse of “research-grade wristband”. The source also reports 95.6% specificity in its abstract and 96.5% in its results. No positive predictive value is given at any prevalence. | |||
| De Zambotti et al. 2014 | Correct | 3.5 flashes recorded per night, only 1.5 reported; 19.8% occur without disturbing sleep | Major |
| The 3.5/night figure is corroborated (n = 34), and 19.8% is consistent with the ~20% reported second-hand. The 1.5 figure could not be verified in any accessible source, including the draft's own review of that paper — every route returned 403, 429, CAPTCHA or a robots block. Its provenance decides a lot: if it is a morning retrospective count rather than a per-event report, the motivating discrepancy is about recall, not about online access, and the primary outcome measures the other thing. Also: the draft's own Baker 2019 citation reports 28.6% undisturbed from n = 86, and the discrepancy goes unmentioned. | |||
| Freedman & Roehrs 2006 | Correct | Flashes suppressed in REM; flash–arousal ordering reverses across the night | Minor |
| Both quoted correctly — the reversal sentence is verbatim. But n = 18 for the hot-flash group, the ordering claim carries no inferential statistic in the abstract, and the reversal is contested in later literature, including in the review the draft cites elsewhere. Using half-of-night as a covariate is the right response; presenting the reversal as established is a stretch. | |||
| Baker et al. 2019 | Correct in every field | Flashes that arouse show a larger cardiac response | Holds |
| Metadata and substance both check out: 51.1% arousal-associated, ~20% HR increase versus marginal, χ² = 158.7, p < .001. The draft's use of this to pre-specify a skin-conductance-and-temperature-only detector is a genuinely good methodological move, and it does not overclaim causality. | |||
| Esfahani et al. 2023 | Still a preprint, cited without a DOI | Sigma-band power “shows the smallest bias of any band” against PSG (0.005 ± 0.012) | Major |
| Two problems. Factually, gamma's bias is smaller than sigma's, so “smallest bias of any band” is false. More importantly it is close to the wrong statistic: a static offset in mean relative band power is removed entirely by the 0.01–0.04 Hz band-pass, and that validation reports no time-resolved or infraslow analysis at all. It is not evidence that the infraslow modulation of sigma is recoverable from this device. | |||
| Lazar et al. 2019 | Correct in every field | “The infraslow modulation of sigma is established in humans” | Major |
| The claim is supported, but the paper is cited as frontal feasibility support while reporting that frontopolar sites carry the lowest infraslow power of any region. It also reports human timescales of 20–40 s, not the 50 s that is the denominator of the draft's entire precision argument. | |||
| Cataldi et al. 2026 | Exists, not fabricated; page range 4151–4159 unconfirmed | Auditory responses decline into phasic REM while cardiac responses are preserved and enhanced | Holds |
| Verified on the web, direction accurate, n = 25, high-density EEG, two nights. Crossref carries no volume, issue or pages for the DOI, so that part of the entry could not be checked. Worth noting the draft cites the paper that supplies a no-report interoceptive measure and then does not adopt it. | |||
| Kjaerby et al. 2022 | Correct | Memory benefit depends on the oscillatory amplitude of norepinephrine “rather than on its mean level” | Minor |
| The paper exists and the amplitude finding is real. The contrast against mean level is not a comparison the paper reports testing. | |||
| Fernandez & Lüthi 2020 | Correct | “In humans, fragility is marked by absent sleep spindles and spontaneous arousals” | Minor |
| UNVERIFIED — Physiological Reviews returned 403 at every route and no repository copy was reachable. Flagged structurally rather than asserted: a review is cited as the source of a specific human empirical claim that the primary human study disclaims. | |||
| Rothhaas & Chung 2021 | Entry ends with a stray asterisk | The preoptic nucleus “receives noradrenergic input” | Minor |
| The paper does not report noradrenergic input to the median preoptic nucleus, which is the nucleus the draft names. The shared-rhythm speculation this supports needs a different source — and a better one sits unused in a review the draft already cites. | |||
| Freeman & Sherif 2007 | Correct in every field | A majority of women experience flashes across the transition | Holds |
| Fine as used. A 2007 systematic review is dated for a 2026 prevalence claim; newer data exist. | |||
| Gombert-Labedens et al. 2025 | Correct | KNDy → NK3R → median preoptic nucleus; narrowed thermoneutral zone; flashes sometimes registered, sometimes not | Minor |
| “Median preoptic nucleus” is the correct anatomy — verified in the review's own abbreviation list — not a slip for medial. Add “(MnPO)” so no reviewer makes the mistake. Two real problems: “thermoneutral zone” contradicts the review's explicit terminological ruling in favour of inter-threshold zone, and the narrowing claim needs hedging; and the conscious-registration sentence is the draft's own inference attributed to the review. This same review also classifies the Tsiartas detector as development-stage and never validated for the use proposed. | |||
| Sikder et al. 2026 | Correct in every field | EEG and wristband synchronised by cross-correlating accelerometer and PPG | Major |
| The citation exists and is perfectly formatted, but it does not support the method: no PPG, a different device pair, and the authors abandoned automated cross-correlation for manual alignment and report no residual error. The synchronisation approach needs either a different source or its own validation. | |||
| Wassing et al. 2019 | Formatted as a section heading, not a reference entry | Arousals follow “standard criteria”: transient > 16 Hz activity lasting 3–15 s | Minor |
| Quoted verbatim and correctly. The problem is calling it standard: the AASM rule includes alpha and theta, a spindle exclusion and a preceding-sleep requirement, and the device has no EOG or EMG to support AASM-conformant scoring at all. | |||
Cited nowhere, and should be
| Work | Why it matters | ||
|---|---|---|---|
| Dimitriades et al. 2026 | n = 154. Locates the ISF peak centro-parietally and places arousal markers at or after the sigma peak. Co-authored by the collaborator the draft invokes as its expertise claim, and uncited. | ||
| The PNAS autocorrelation critique | n = 1,025. Questions whether the infraslow fluctuation is a continuous oscillation at all rather than an autocorrelation consequence of spindle clustering and refractoriness. Uncited, and it goes to the existence of the predictor. | ||
| Carro-Dominguez et al. 2025 | Human infraslow–LC work. The proposal motivates a human study from rodent LC photometry while the human literature now exists. | ||
| Grund et al. 2022 | The only available anchor for how large a bodily-rhythm phase effect on conscious detection actually is — circular R ≈ 0.34 in wakefulness. Needed to name a smallest effect of interest. | ||
| The CAP and PLMS literature | An entire prior literature on infraslow NREM microstructure gating endogenous events. It is the closest existing precedent for the study's own premise, and it is absent. | ||
| Chen et al. 2025 | Flagged by the literature sweep as directly relevant and uncited. | ||
151 findings that survived
Grouped by kind, major first. Each carries the draft quote, the evidence, the fix, and what the refuter said when trying to knock it down. “Confirmed” means verified against a primary source or an airtight contradiction; “Plausible” means likely right but not fully verifiable.
Whether Bial can fund this at all, and whether the application clears its formal bars.
The problem
The draft describes the sample as physically healthy women "reporting sleep disturbances". The host trial's inclusion criteria are Insomnia Severity Index ≥10 and Greene Climacteric Scale ≥13 — that is a clinical insomnia population enrolled into a randomised trial of two treatments, with change in ISI as the primary outcome. Bial excludes "Projects involving clinical or experimental models of human disease and therapy". An unsympathetic reviewer can retire the application on scope alone: it is an ancillary analysis inside a treatment trial for a disorder. The underlying science is genuinely psychophysiological and the study is genuinely observational, so this is winnable — but only if argued, and the draft does not argue it because it does not disclose the trial. Compounding this, the draft's softening of "insomnia disorder" to "self-reported sleep disturbances" (lines 61-62, 197-198) is the kind of understatement that, once a reviewer finds the host protocol, colours everything else in the application.
In the draft — line 196-199
Participants are perimenopausal women (irregular cycles and climacteric symptoms) reporting sleep disturbances, from an ongoing cohort of 222 women already being recorded. They are physically healthy
Evidence
/root/grantreview/BIAL_CONTEXT.md line 10: "Projects involving clinical or experimental models of human disease and therapy" listed under EXCLUSIONS. Host trial inclusion, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: "Insomnia Severity Index (ISI) ≥10", "Greene Climacteric Scale (GCS) ≥13"; primary outcome "Change in insomnia severity via ISI at T1 (8 weeks) and T2 (15 weeks)"; four arms including MHT and CBCTi.
Fix
Disclose and then argue scope, in the Participants section: "Measurements are nested in an ongoing randomised trial (NCT06306404) but this project is observational and non-therapeutic: it uses the pre-treatment measurement block only, its outcome is the conscious registration of a normal physiological event of the menopausal transition, and neither its question nor its analysis concerns disease mechanism or treatment efficacy." Then state the psychophysiological framing Bial's scope requires — the relation between a measurable bodily response and reportable conscious experience — and drop "physically healthy" in favour of an accurate description: perimenopausal women with clinically significant insomnia symptoms and climacteric symptoms, free of the host trial's exclusions.
What the refuter said
Verified independently and it is stronger than filed. I fetched the host protocol: inclusion is 'Insomnia Severity Index (ISI) >=10' and 'Greene Climacteric Scale (GCS) score >=13'; the primary outcome is 'Change in insomnia severity (Insomnia Severity Index) at 8 weeks (T1) and 15 weeks (T2)'; there are four arms — CBCTi, MHT, CBCTi+MHT, control; n = 222; and the device schedule is 'Z-Max ... four nights per timepoint' plus 'EmbracePlus ... seven days per timepoint', i.e. exactly the draft's design. Bial Art. 16(6) verified verbatim from the Regulation PDF. STRONGEST DEFENCES: the proposed study is genuinely observational psychophysiology, not a therapy study, and 'physically healthy' is defensible against the protocol's medical exclusions (cancer, liver disease, thromboembolism) — insomnia is not a physical illness. So the eligibility argument is arguable, and that is precisely the finding's point: it is arguable, and the draft does not argue it because it never discloses the trial. One thing sharper than the finding states: MHT is one of the four arms, and L199-200 asserts participants 'take no hormone-altering medication' with 'attention to agents suppressing vasomotor symptoms' — a statement that is true only of baseline recordings, and the draft never says which timepoint its four nights come from. That is a substantive, checkable gap in the sample description, not just a framing issue. Major.
The problem
The draft never discloses in its body text that the 222-woman cohort IS the sample of a 2x2 randomised controlled trial of transdermal estradiol plus oral progesterone and online CBT for insomnia. The only trace is a citation buried in PREVIOUS OWN PUBLICATIONS. I fetched that protocol: inclusion is Insomnia Severity Index >=10 and Greene Climacteric Scale >=13, with recruitment partly from menopause clinic waiting lists — i.e. care-seeking women meeting a clinical insomnia threshold, whose own protocol title calls them 'perimenopausal women suffering from insomnia'. The draft's counter-assertions are three scattered phrases — 'a normal stage of reproductive life' (L37-38), 'physically healthy' (L199), and 'in healthy humans and without intervention' (L369) — and the last is flatly contradicted: half the trial arms receive estradiol/progesterone and half receive CBCTi. Bial's own priority language is 'research focused on the healthy human being'. An assessor screening 400+ applications against Art. 16(6) has a clean, defensible route to rejection here, and the draft mounts no argument against it. Worse, non-disclosure means the assessor discovers the RCT themselves from the reference list, which converts a fit question into a candour question.
In the draft — line 197-200
from an ongoing cohort of 222 women already being recorded. They are physically healthy and take no hormone-altering medication
Evidence
Parent protocol, fetched 26 Aug 2026 from https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: target 222 women, n=50 per arm in a 2x2 design; inclusion 'Ages 40–55; Insomnia Severity Index (ISI) >=10; Greene Climacteric Scale (GCS) >=13'; recruitment 'via general population self-enrollment and menopause clinic waiting lists'; four arms — CBCTi, MHT (Systen 50 mcg transdermal estradiol + Utrogestan 200 mg progesterone, 15 weeks), CBCTi+MHT, control; NCT06306404; funded by NWO project NWA.1518.22.104. Regulation Art. 16(6), verbatim: "Projects involving clinical or experimental models of human disease and therapy" are ineligible. Bial call language: "encourage research into the healthy human being, both from the physical and spiritual point of view" (https://www.fundacaobial.com/en-GB/news/bial-foundation-opens-new-call-for-grants-for-scientific-research-2026-2027). NOTE the counter-precedent: Bial's own funded portfolio includes clinically adjacent work, e.g. the grantee presentation 'COping with PAin through Hypnosis, mindfulness and Spirituality (COPAHS)' and 'Mind-shaped body: A new conceptual framework beyond the placebo effect' (14th symposium oral poster programme, https://assets.ctfassets.net/osaht2ckekgb/61NFK3eSb04C7ep5pELKa3/0aecdba734992fd582ae3a700aa7ed68/14th-symposium-oral-poster-presentations.pdf). So the exclusion is applied with some latitude — this is a high-probability screening risk, not a certainty.
Fix
Disclose the parent trial in the body, in the Participants section, and make the eligibility argument explicitly rather than by adjective. Add a dedicated short paragraph, e.g.: 'Participants are drawn from the baseline (pre-randomisation) assessment of an ongoing trial (NCT06306404; van Baarzel et al., 2026). This project administers no intervention, tests no treatment, and models no disease: menopause is a normal reproductive-life stage, and the outcome is the conscious registration of a spontaneous physiological event in unmedicated participants. Data are analysed only from the pre-randomisation timepoint, before any participant receives hormone therapy or behavioural therapy.' That last clause is only usable if it is true — see the timepoint finding. Replace 'physically healthy' with an accurate description of the inclusion criteria; asserting 'healthy' against ISI>=10 and clinic-waiting-list recruitment is the kind of gloss that destroys reviewer trust when they open the cited protocol.
What the refuter said
AGAINST: Bial's Art. 16(6) bars "Projects involving clinical or experimental models of human disease and therapy" (BIAL_CONTEXT.md L10). The *project applied for* is observational, delivers no therapy, and menopause is not a disease — the draft says so at L37-38 ("a normal stage of reproductive life"). Nor is the host trial concealed: the PREVIOUS OWN PUBLICATIONS entry (DRAFT.md L449-454) carries the full title including "randomized controlled trial of menopausal hormone therapy and online guided cognitive behavioural and circadian therapy for insomnia in perimenopausal women suffering from insomnia" — an assessor cannot miss it, so the charge of non-disclosure-as-concealment is weaker than stated. The finding also concedes latitude in Bial's portfolio. SURVIVES, partly. I fetched the protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full) and it verifies every factual premise: ISI ≥10 and GCS ≥13 inclusion, recruitment via "OLVG hospital menopause clinic waiting lists", four arms (MHT / CBCTi / both / control), Systen 50mcg + Utrogestan 200mg, NCT06306404, NWA.1518.22.104. Against that, DRAFT.md L199 ("physically healthy and take no hormone-altering medication") and L369-370 ("in healthy humans and without intervention") are affirmatively false of the host sample as recorded, and the body text never names insomnia or the RCT. So the verified defect is not "ineligible" but "an eligibility question is raised by the applicants' own reference list and the draft mounts no answer to it, while making two statements a reader can check and find wrong." That is real and decision-relevant. Downgraded from fatal: the exclusion applies to the proposed project, which is non-interventional, and the fix is a disclosure paragraph, not a redesign.
The problem
Art. 16(7) makes an application without a designated Host Entity ineligible — this is a hard screening bar, not a scoring deduction. The word 'host' does not appear in the draft; 'Amsterdam UMC' appears only inside the ethics sentence. Meanwhile the team is implicitly split: the data, the cohort and the ethics approval sit at Amsterdam UMC, while the claimed rodent-mechanism expertise (Osorio-Forero, Lüthi) sits in Lausanne. The PREVIOUS OWN PUBLICATIONS list compounds this by claiming Osorio-Forero et al. 2021 and 2025 as the applicant's own publications alongside van Baarzel et al. 2026 — an assertion that is only legitimate if those authors are formally on the team with CVs and consent. As drafted, the assessor cannot tell who the applicant is, which institution hosts the grant, or whether the headline expertise claim is real.
In the draft — line 58-62
Our team is uniquely placed to run this test: it includes the researchers who characterised the infraslow noradrenergic mechanism in rodents (Osorio-Forero et al., 2021, 2025), and has access to an ongoing cohort of 222 perimenopausal women
Evidence
Regulation Art. 16(7): applications without a Host Entity are ineligible (fetched from the official 2026 Regulation PDF; corroborated in /root/grantreview/BIAL_CONTEXT.md line 19: "A Host Entity MUST be designated; applications without one are ineligible"). Art. 3(3): "The Principal Investigator, considered the Grant Holder ... shall be responsible for coordinating the Research Project and shall act as the interlocutor with Bial Foundation, representing the project and the other team members". Grep of DRAFT.md: 'host'/'Host' = 0 occurrences. Amsterdam UMC affiliations for van Baarzel, Koning, van Someren, Broekman confirmed at https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full; Osorio-Forero and Lüthi are the Lausanne authors of Current Biology 31(22):5009-5023.
Fix
Name the PI, designate Amsterdam UMC (with the specific department/Research Centre) as Host Entity in the form and in the narrative, and list every team member with role and institution. Then either (a) confirm Osorio-Forero and/or Lüthi as CV-bearing team members and keep the expertise claim, or (b) drop 'it includes the researchers who characterised...' and replace with an accurate collaboration statement. Remove Osorio-Forero et al. 2021/2025 from PREVIOUS OWN PUBLICATIONS unless the relevant author is genuinely on the team.
What the refuter said
STRONGEST DEFENCE: the document under review carries form-field headings ('LITERATURE REVIEW (MAX. 6000 CHARACTERS INCLUDING SPACES)'), so it is plausibly only the narrative prose sections of the BF-GMS online form. Host Entity, Research Centre, schedule and budget are structured fields elsewhere in that form, so their absence from this file may be an artefact of what was handed for review rather than an ineligibility. That defence disposes of the eligibility prong. It does not dispose of the substantive prong. VERIFIED: grep over DRAFT.md returns zero occurrences of 'host'/'Host'; 'Amsterdam' appears only at L332 inside the ethics sentence; no budget or schedule section exists. Art. 16(7) verified verbatim from the official Regulation PDF: 'Applications without a Host Entity are not eligible'; Art. 16(3) requires 'detailed project description, schedule, budget, CV(s), and Host Entity identification'; Art. 3(3) PI-as-interlocutor language verified. The host-trial authorship is verified independently: van Baarzel, Broekman, van Dijken, van Someren and Koning are all Amsterdam UMC (medRxiv protocol), while Osorio-Forero and Lüthi are the Lausanne authors of Current Biology 31(22). The load-bearing claim at L58-60 — 'it includes the researchers who characterised the infraslow noradrenergic mechanism in rodents' — is attributable to no named team member anywhere in the document, and PREVIOUS OWN PUBLICATIONS (L447-464) claims two Lausanne-authored papers as the applicant's own alongside the Amsterdam protocol. That is not a form field: an assessor cannot verify the single differentiating expertise claim in the proposal. Decision-relevant. Severity held at major rather than fatal because the eligibility half is probably recoverable in the form.
The problem
Bial excludes "Projects involving clinical or experimental models of human disease and therapy". The host cohort is a four-arm randomised trial of menopausal hormone therapy and online cognitive behavioural and circadian therapy, enrolling women with an Insomnia Severity Index of 10 or higher — a clinical insomnia population in a therapy trial. The draft never discloses this, describing participants only as 'physically healthy' women 'reporting sleep disturbances'. A reviewer who reads the PREVIOUS OWN PUBLICATIONS entry will see it immediately, and the mismatch between how the cohort is described and what it is will read as concealment rather than framing. The psychophysiological question itself is squarely in scope; the vehicle is the problem.
In the draft — line 197-198
from an ongoing cohort of 222 women already being recorded
Evidence
BIAL_CONTEXT.md line 10: "Projects involving clinical or experimental models of human disease and therapy" are excluded. Parent protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): inclusion requires "insomnia severity index (ISI) of 10 or higher" and "green climacteric scale (GCS) score of 13 or higher"; the four arms are MHT alone, CBCTi alone, MHT + CBCTi, and control. DRAFT.md 198-199 describes the same participants as "physically healthy".
Fix
Disclose the host study and frame the boundary: 'Participants are recruited from an ongoing randomised trial of menopausal hormone therapy and cognitive behavioural therapy for insomnia in perimenopausal women (van Baarzel et al., 2026). This project uses only the pre-randomisation baseline recordings and tests no intervention, no clinical outcome and no therapeutic hypothesis; it is a psychophysiological study of the relation between an endogenous bodily event and its conscious registration.' Then keep to that restriction throughout, including in the yield calculation.
What the refuter said
AGAINST: Bial's exclusion (BIAL_CONTEXT.md L10) attaches to the project applied for, and this project administers nothing and treats no one; "reporting sleep disturbances" is not a false description of women with ISI ≥10, merely a euphemistic one. Nor is the vehicle hidden: the own-publications entry at L449-454 states in its title that it is a randomised controlled trial of MHT and CBT-I in "perimenopausal women suffering from insomnia", so an assessor learns it from the same document. The finding is also largely a restatement of F2. SURVIVES on the same verified basis. The protocol confirms ISI ≥10 and GCS ≥13 inclusion and the four therapy arms, while DRAFT.md L198-199 calls the same women "physically healthy" and L369-370 claims the work is done "without intervention". The mismatch between how the sample is described and what it is stands unargued, and the finding's own framing is the right one: the question is in scope, the vehicle raises the risk. Major, and it should be merged with F2 rather than reported twice — the remedy is a single disclosure passage that names the host trial, the timepoint used, and why Art. 16(6) does not bite.
The problem
Four compliance problems. (1) The time-stamped button press — the source of the primary outcome — does not appear in the parent RCT protocol, which specifies only objective multi-sensor flash estimation; adding a participant-facing procedure requires an amendment, so the blanket claim of approval is not established. (2) The ethics reference is left as a placeholder; Bial Art. 19(4) requires documentary evidence of submission and states no Agreement is signed without it. (3) The draft has no schedule and no budget, both of which the regulation requires, and no Host Entity is designated — applications without one are ineligible. (4) The regulation names the "Specific Aims" section as essential for assessment; the draft's heading is "Research aims". Separately, the scope exclusion for "projects involving clinical or experimental models of human disease and therapy" is a live risk: participants are drawn from a randomised therapy trial in women with clinical insomnia, and the Expected Outcomes section leans on clinical framing.
In the draft — line 331-332
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference \[number\]).
Evidence
BIAL_CONTEXT.md: "Host Entity MUST be designated; applications without one are ineligible"; "Detailed description of the research project + schedule + budget"; "No Agreement will be issued and signed unless such document has been duly provided (Art. 19(4))"; "The objectives ... namely in the 'Specific Aims' section. These objectives are essential for the assessment"; exclusion of "Projects involving clinical or experimental models of human disease and therapy". Parent protocol lists no button press and gives reference NL87156.018.24, NCT06306404 (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full).
Fix
State the actual METC reference and the amendment status of the button-press procedure separately from the parent approval, or drop the claim of existing approval for it. Add a Host Entity, a schedule mapped to the 1 Jan - 31 Oct 2027 start window and the cohort's recruitment state (65 of 222 screened), and a budget in Euro. Rename the aims section "Specific Aims". Reframe Expected Outcomes away from clinical benefit toward the psychophysiological question, and delete or heavily qualify "why treating the flashes does not reliably resolve the sleep complaint".
What the refuter said
Four prongs, of unequal quality. (1) is the substantive one and I verified the absence: the protocol at https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full describes only that "objective occurrence of nocturnal hot flashes will be estimated from a multisensory approach... assessed with the EmbracePlus smartwatch", with no participant-initiated button press or event marker anywhere. Against that: the EmbracePlus carries a physical event-marker button as standard, a published protocol is not an exhaustive inventory of participant instructions, and amendments are routine and unpublished - so "the procedure is not approved" is unproven, and the honest statement is that the blanket approval claim at L331-332 is not *established* by the public record. (2) stands: the reference is literally "[number]", and BIAL_CONTEXT.md L34-36 requires documentary evidence of ethics submission with "No Agreement will be issued and signed unless such document has been duly provided (Art. 19(4))". (3) belongs to F1 and is largely a category error about narrative versus form fields. (4) is refuted outright - "Specific Aims" versus "Research aims" is a heading supplied by the online form, and grading a section title as a compliance problem is exactly the kind of nitpick that discredits a review. PLAUSIBLE at moderate, on prongs (1) and (2) only.
The problem
The parent protocol collects exactly this device package — ZMax for four consecutive weeknights, EmbracePlus for seven days — at THREE timepoints: T0 baseline (week 1), T1 post-intervention (week 8), T2 follow-up (week 15). The draft describes the package as if it happened once and never says which timepoint(s) it analyses. This is not a cosmetic omission. If T1/T2 nights are pooled in, participants are on transdermal estradiol plus progesterone or have completed CBCTi — both of which directly alter vasomotor symptom frequency, sleep continuity and sigma-band activity, and both of which make the analysis interventional clinical research. If only T0 is used, the yield calculation halves or thirds relative to what a reader might assume, and the draft should say so plainly because it is also the eligibility defence.
In the draft — line 63-65
Each completes four nights of ambulatory frontal EEG alongside seven days and nights of wrist-worn autonomic monitoring at home.
Evidence
Parent protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): "EEG headband: Z-max triple-electrode (Hypnodyne); 4 consecutive weeknights per timepoint"; "Smartwatch: Empatica EmbracePlus; 7 days per timepoint"; "Assessment timepoints: Baseline (T0, week 1), post-intervention (T1, week 8), follow-up (T2, week 15)". The draft's own claim of no intervention, L369: "in healthy humans and without intervention", is compatible only with T0.
Fix
State it in one sentence in Study design: 'Analyses use the pre-randomisation baseline assessment only (T0: four EEG nights, seven days of wrist monitoring per participant); on-treatment timepoints are excluded.' Then propagate that constraint into the yield calculation and the power section, and cite it again in the eligibility paragraph. If the team actually intends to use T1/T2 for extra events, say so and accept that the eligibility argument becomes much harder.
What the refuter said
AGAINST: the draft is not silent by accident — L200-201 ("They are physically healthy and take no hormone-altering medication") and L369 ("in healthy humans and without intervention") both implicitly exclude on-treatment nights, and the draft consistently describes ONE timepoint's package (L63-65 and L188-193: seven days wearable, four of those nights EEG), which is exactly the T0 quantity. So the finding's claim that "the yield calculation halves or thirds relative to what a reader might assume" is wrong: a reader of this draft assumes 4 nights, which IS the per-timepoint number. That limb is refuted. SURVIVES: I verified the parent protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): three timepoints (T0 week 1, T1 week 8, T2 week 15), Z-max "at home for four nights during each timepoint", EmbracePlus "for seven days at each timepoint", 2x2 arms MHT / CBCTi / both / control, ISI>=10. A grep of DRAFT.md for baseline|timepoint|T0|randomis|intervention|hormone|trial returns nothing in the body except L200 and L370 — the draft never once discloses that its cohort is an MHT+CBCTi randomised trial, never names a timepoint, and never states whether control-arm or post-treatment nights enter. Against a Bial regulation that excludes "projects involving clinical or experimental models of human disease and therapy", non-disclosure of the RCT context is a live eligibility question the draft leaves a reviewer to discover from its own reference at L451. Downgraded from major to moderate: this is a disclosure gap fixable in two sentences, and the implicit off-treatment signals at L200/L369 show the intent is T0.
The problem
The draft never identifies the 222-woman cohort, but the sole listed own publication is a protocol for a randomised controlled trial of menopausal hormone therapy and CBT for insomnia in perimenopausal women with insomnia. If that trial is the cohort, then a substantial fraction of participants receive hormone therapy — which suppresses vasomotor symptoms and directly contradicts the stated exclusion — and the participants have a diagnosed sleep disorder receiving therapy rather than the "self-reported sleep disturbances" of the summary. That also engages Bial's exclusion of projects involving clinical or experimental models of human disease and therapy. Either the cohort is a different, untreated sample, in which case the draft must say so, or the eligibility statement cannot stand as written. As drafted, a reviewer cannot tell which, and the ambiguity sits on the project's feasibility claim.
In the draft — line 197-201
Participants are perimenopausal women (irregular cycles and climacteric symptoms) reporting sleep disturbances, from an ongoing cohort of 222 women already being recorded. They are physically healthy and take no hormone-altering medication
Evidence
Draft L449-454 lists as own work: "Sleeping Through Menopause: study protocol of a randomized controlled trial of menopausal hormone therapy and online guided cognitive behavioural and circadian therapy for insomnia in perimenopausal women suffering from insomnia." Against draft L199-200: "They are physically healthy and take no hormone-altering medication". BIAL_CONTEXT.md L10: "Projects involving clinical or experimental models of human disease and therapy" is listed among the exclusions. Logical contradiction: a participant randomised to menopausal hormone therapy takes hormone-altering medication by design.
Fix
Name the cohort and its relationship to the trial explicitly. If participants are drawn from the observational arm or from screening before randomisation, state that: "Participants are recruited from the pre-randomisation screening cohort of [cohort name] (n = 222); women who have started menopausal hormone therapy or any other agent suppressing vasomotor symptoms are excluded, and treatment status is recorded at each recording night." If treated women are included, drop the exclusion sentence, add treatment as a covariate, and address the Bial therapy exclusion head-on by stating that no intervention is delivered or evaluated within this project.
What the refuter said
Strongest case against: two of the three limbs are over-reads. I verified the parent protocol (medRxiv 2026.07.22.26358657): n=222 exactly matching the draft's cohort, EmbracePlus for seven days per timepoint, four arms (CBCTi, MHT, CBCTi+MHT, control), METC AUMC NL87156.018.24, NWO NWA.1518.22.104. But its exclusion criteria include "current MHT use" — so nobody is on hormone therapy at entry, and if the draft's four EEG nights are the baseline timepoint (which "already being recorded" and "of the first 65 screened" both suggest), then "take no hormone-altering medication" is literally true and there is no contradiction. Second, inclusion is "Insomnia severity index (ISI) of 10 or higher" — a self-report questionnaire threshold in the subthreshold range, not a diagnosed disorder under treatment, so the draft's "reporting sleep disturbances" is accurate and the finding's "diagnosed sleep disorder receiving therapy" is wrong. What survives, and is the finding's own fallback: the draft never identifies the cohort or which timepoint the recordings come from, so a reviewer cannot check the eligibility claim at all, and the only clue in the document is a listed own-publication that is a hormone-therapy RCT. That ambiguity sits directly on the feasibility claim. Moderate, not the contradiction alleged.
The problem
Two distinct problems. (1) A bare '[number]' placeholder in a sentence claiming completed approval is unverifiable by the assessor and, if left in the submitted form, reads as a claim the applicant could not substantiate. (2) More seriously: the primary outcome measure of this proposal — a time-stamped wristband button press for prospective conscious registration of a flash — does not appear in the parent protocol's measures, which list EmbracePlus-based hot flash detection and a daily consensus sleep diary. Nor does the lights-off/waking five-tap synchronisation procedure. Adding a prospective nocturnal self-report action to a randomised trial protocol is a substantive protocol change requiring an METC amendment. Whether such an amendment exists is UNVERIFIED, but the draft asserts blanket approval for 'the study' without distinguishing the parent trial's approval from this sub-study's additions. Under Art. 19(4) no Agreement is signed until the approval document is provided, and the PI signs a declaration — so an over-broad claim here is an integrity exposure, not just a tidiness issue.
In the draft — line 331-332
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference \[number\]).
Evidence
Parent protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full) gives the approval as "Medical Research Ethics Committee, Amsterdam University Medical Center (METC AUMC); approval ref. NL87156.018.24", and lists measurement instruments as ZMax EEG, EmbracePlus smartwatch ("heart rate, electrodermal activity, movement, hot flash detection") and a daily consensus sleep diary — no nocturnal button press, no tap-marker synchronisation. Regulation Art. 19(4), verbatim: "The applicant shall also provide Bial Foundation, prior to signature of the Agreement, with a copy of the approval of the Research Project by the competent Ethics committee(s)/authority(ies). No Agreement will be issued and signed unless such document has been duly provided to Bial Foundation." Art. 16(4) requires at application stage "documentary evidence of its submission for review and approval" plus a signed PI declaration.
Fix
Replace the sentence with an accurate, scope-explicit version, e.g.: 'The host trial is approved by the Medical Research Ethics Committee of Amsterdam UMC (NL87156.018.24; NCT06306404). The nocturnal event-marking procedure and device-synchronisation markers introduced for this sub-study are covered by amendment [ref, date] / have been submitted as amendment [ref] on [date].' Attach the METC letter (or the submission receipt for the amendment) to the application. Do not submit the form with a '[number]' placeholder — Art. 16(4) wants a document, and an unsubstantiated approval claim is worse than an honest 'submitted, decision pending'.
What the refuter said
STRONGEST DEFENCE: (a) a bracketed placeholder is a drafting artefact, not a substantive claim — the real reference is public (NL87156.018.24) and costs nothing to insert; (b) Bial's application-stage bar is only "documentary evidence of its submission" (BIAL_CONTEXT.md L34-36), so a pending amendment would satisfy the form; (c) the draft's sentence "The study is approved" could be read as referring to the parent protocol, which genuinely is approved. WHY IT SURVIVES: I verified the placeholder is literally in the file (DRAFT.md L332: "reference \[number\]"), and I fetched the parent protocol (medrxiv 10.64898/2026.07.22.26358657): approval is "Medical Research Ethics Committee of the Amsterdam University Medical Center (METC AUMC)", ref NL87156.018.24, and its instrument list is ZMax (4 nights/timepoint), EmbracePlus (7 days/timepoint, "multisensory nocturnal hot flash estimation"), consensus sleep diary and questionnaires. The fetch explicitly returned "No button-press marker mentioned" and no tap/synchronisation marker. So the draft's two novel participant-facing procedures — the nocturnal button press (the PRIMARY outcome) and the five-tap synchronisation — are not in the approved parent protocol, and the draft nonetheless asserts blanket approval for "the study". Whether an amendment exists is correctly flagged UNVERIFIED by the finding, which is the honest position. DOWNGRADE: from major to moderate. The placeholder is trivially fixable; the scope issue is a disclosure/precision defect rather than a demonstrated eligibility failure, and Art. 19(4) bites at Agreement signature, not assessment.
The problem
Bial funds psychophysiology and excludes "projects involving clinical or experimental models of human disease and therapy". The draft handles one part of this well - it says explicitly that the menopausal transition is "a normal stage of reproductive life" (L37-38), which is the right move. Then it undercuts itself in four places. The analysis sample is defined by a symptom-interference criterion, i.e. selected on clinical burden. The parent cohort is a randomised trial of hormone therapy and CBT-I in women screened for insomnia severity (see cohort-is-a-treatment-rct), which the draft does not disclose. The ethics route is a Medical Ethics Review Committee. And the closing paragraph promises to explain a treatment failure. A reviewer reading only the opening and the closing sees a consciousness question at the top and a therapy question at the bottom, which invites the exclusion. The psychophysiological framing is also weaker than it needs to be: the study's real strength for this funder is that it measures a mental event (reported awareness) against a simultaneously recorded physical response (autonomic flash) and a brain state, which is psychophysiology in the strict sense, and the draft never says so.
In the draft — line 202-204
The analysis sample comprises those whose nocturnal flashes interfere with sleep: of the first 65 screened, 40 (62%, 95% CI 49-72) meet this criterion
Evidence
/root/grantreview/BIAL_CONTEXT.md: scope is "Psychophysiology - relationships between mental/psychological processes and physical responses"; exclusion: "Projects involving clinical or experimental models of human disease and therapy"; also "Projects in fundamental neuroscience ... that are not directly and unequivocally associated with a psychophysiological measure". Draft L370-373 makes the therapy claim; L202-203 selects on symptom interference; L331-334 routes ethics through Amsterdam UMC's Medical Ethics Review Committee; the parent trial screens on "Insomnia severity index (ISI) of 10 or higher" (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full).
Fix
Frame consistently as psychophysiology and let the clinical relevance be implicit. Concretely: (a) add one explicit sentence early - "The study relates a mental event (whether a bodily change is noticed and reported) to concurrently recorded physical responses (sudomotor, vascular and cardiac activation) and to brain state, in healthy women undergoing a normal reproductive transition; no disease model and no therapy is involved." (b) Reframe the sample criterion physiologically rather than clinically: select on nocturnal flash frequency (a minimum detected event rate per night), not on whether flashes "interfere with sleep". (c) Disclose the parent cohort and state that the analysis uses recordings taken while participants are not receiving hormone therapy. (d) Cut the therapy sentence at the end (see closing-clinical-overclaim). (e) Justify the rodent noradrenergic material by its link to the human psychophysiological measure in one clause, so it cannot be read as fundamental neuroscience for its own sake.
What the refuter said
Strongest case against: the exclusion is not engaged, and the finding half-admits it. Bial excludes "projects involving clinical or experimental models of human disease and therapy"; this project is observational, administers nothing, models no disease, and studies women in what the draft twice calls a normal state — "a normal stage of reproductive life" (L37-38) and "in healthy humans and without intervention" (L369). Selecting on symptom interference is a sampling criterion for event yield, not a disease model, and routing ethics through a Medical Ethics Review Committee is what Dutch law requires of any human-subjects sleep study; neither is evidence of clinical scope. The claim that the psychophysiological framing is "weaker than it needs to be" and that "the draft never says so" is wrong on its face: a mental event measured against a simultaneously recorded physical response is the entire architecture of the proposal (title, L12-15, L48-56, L85-87, L175-177). What survives is presentational and modest: the undisclosed nesting in an MHT/CBT-I randomised trial (F36's verified point) and the closing sentence at L370-373 promising to explain why treating flashes fails to resolve the sleep complaint — an avoidable invitation to the exclusion. Positioning advice, not a defect; minor.
The problem
The regulation makes a designated Host Entity a condition of eligibility, requires a detailed schedule and a budget in Euro, requires CVs for the applicant and every team member, and states that the objectives in the 'Specific Aims' section are what the application is assessed against. The draft contains none of these: no Host Entity or Research Centre, no schedule, no budget, no CVs, and a section headed 'Research aims' rather than 'Specific Aims'. It also does not say when the project would start, though Bial requires a start between 1 January and 31 October 2027 — which interacts with the cohort question, since the draft never states whether the recordings it needs already exist or will be collected during the grant.
Evidence
BIAL_CONTEXT.md line 19: "A Host Entity MUST be designated; applications without one are ineligible." Lines 31-33: "CV of applicant and every team member. Detailed description of the research project + schedule + budget. Amounts in Euro. Host Entity and Research Centre identified." Lines 42-44: "The objectives to be achieved within the scope of the Research Project are those defined in the submitted application, namely in the 'Specific Aims' section. These objectives are essential for the assessment of the application." Line 22: "Maximum duration 3 years. Project must start between 1 Jan and 31 Oct 2027." DRAFT.md contains no matching sections.
Fix
Add the missing sections before submission: Host Entity and Research Centre (Amsterdam UMC, named department), a Gantt-style schedule against a start date inside the 1 Jan - 31 Oct 2027 window, a budget in Euro noting that overheads and equipment-use costs are not reimbursable, and CVs for every named team member. Rename 'Research aims' to 'Specific Aims' and state the aims as numbered, assessable objectives, since these are the criteria the project will be judged against on completion.
What the refuter said
AGAINST — and this is most of the finding. Its sole evidence is "DRAFT.md contains no matching sections", which proves nothing for four of its five items: Host Entity, Research Centre, CVs and the budget table are BF-GMS form fields and uploads (I confirmed Art. 16(3) in the regulation PDF: "the Curriculum Vitae of the applicant ... a detailed description of the Research Project and corresponding schedule and budget; the identification of the Host Entity and the Research Centre"). None of them would appear in a narrative-sections draft, so their absence here is expected. The 'Specific Aims' heading complaint duplicates F9 and is cosmetic. WHAT SURVIVES: the schedule and the project-timing question are genuinely narrative and genuinely absent. I grepped DRAFT.md for schedule|timeline|work package|month|year and found nothing but the idiom "the body's own schedule". The draft says the cohort is "already being recorded" (L198) but the parent trial runs three timepoints and is ongoing, and the Bial start window is 1 Jan-31 Oct 2027 — so a reviewer cannot tell whether the analysed nights exist, are mid-collection, or are yet to be recorded. That is decision-relevant, and it is the same gap F3 and F159 approach from other directions. PLAUSIBLE at minor because the finding's own framing is largely a checklist against the wrong document.
The problem
The Bial regulation excludes "Projects involving clinical or experimental models of human disease and therapy". The proposal is careful in places — it calls the menopausal transition "a normal stage of reproductive life" (l.37-38), recruits healthy women, and involves no intervention (l.369-370: "in healthy humans and without intervention"). But the closing paragraph pivots to a therapeutic rationale ("why treating the flashes does not reliably resolve the sleep complaint"), participants are selected for sleep disturbance and for flashes that "interfere with sleep" (l.201-203), and the cohort appears to derive from a therapy trial in women "suffering from insomnia" (l.449-454). A reviewer applying the exclusion literally could read the sample as a clinical population and the motivation as therapeutic. The psychophysiological core — an endogenous bodily signal and its access to conscious report — is squarely in scope and is the frame that should dominate.
In the draft — line 370-373
It would also give a mechanistic account of a clinical puzzle - why women report far fewer nocturnal hot flashes than they have, and why treating the flashes does not reliably resolve the sleep complaint that accompanies them.
Evidence
Draft l.370-373 (therapeutic framing) and l.201-203 ("those whose nocturnal flashes interfere with sleep") and l.449-454 (parent protocol is a therapy RCT in women with insomnia), against /root/grantreview/BIAL_CONTEXT.md line 10: "Projects involving clinical or experimental models of human disease and therapy" listed under EXCLUSIONS. Also BIAL_CONTEXT.md line 6: the in-scope field is "Psychophysiology — relationships between mental/psychological processes and physical responses".
Fix
Keep the psychophysiological frame dominant and demote the clinical implication to a single subordinate clause. Suggested replacement for l.370-373: "Because the flash is a normally occurring bodily event in a normal stage of reproductive life, the study asks a psychophysiological question — how a measurable peripheral event becomes a reportable experience — in healthy women and without any intervention. A secondary implication is that the well-documented gap between the flashes women have and the flashes they report may reflect gating of access rather than faulty recall." Remove the reference to treatment efficacy, and state explicitly in Participants that no participant is being studied as a patient and that no therapy is administered or evaluated within this project.
What the refuter said
Strongest case against: identical in substance to F13, and the exclusion clause does not bite. BIAL_CONTEXT.md L10 excludes "Projects involving clinical or experimental models of human disease and therapy" - a clause about design, and this design administers nothing to anyone. Every draft-side quote checks out (L37-38 "a normal stage of reproductive life"; L201-203 "those whose nocturnal flashes interfere with sleep"; L369-370 "in healthy humans and without intervention"; L449-454 the parent RCT in women "suffering from insomnia"), and I confirmed the parent trial's inclusion criteria at https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: "Insomnia Severity Index >=10" and "Greene Climacteric Scale >=13". So the sample genuinely is drawn from a clinical population enrolled in a therapy trial, and an assessor applying the exclusion literally has material to work with. What makes this finding acceptable where F13 is not is calibration: it grades itself minor, credits the draft's careful framing, and recommends a presentational fix (let the psychophysiological frame dominate the close) rather than alleging ineligibility. PLAUSIBLE rather than CONFIRMED because whether any assessor would actually invoke the clause is unknowable and no comparable rejection is on record.
What the design can and cannot detect. All numbers here were produced by running the simulations, not by reading the draft.
The problem
Sensitivity and specificity are prevalence-free; PPV is not, and PPV is what determines whether the analysed events are flashes. On the draft's own event rate the base rate is a few per cent per 15-s frame, so 4.4% false-positive rate over a whole night swamps the true events. Worse, Tsiartas oversampled the flash class to a 1:2 ratio for training, so the reported operating point was measured at roughly 33% prevalence, not the ~3% of a real night. Every false positive enters the primary model with a fabricated onset time, and its apparent phase is set by whatever produced the spurious EDA excursion — which is not random with respect to sigma.
In the draft — line 244-245
reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance
Evidence
My computation on the draft's and its sources' numbers: 8 h in bed = 28,800 s / 15 s = 1,920 decision frames; 3.5 flashes/night (Gombert-Labedens 2025, /root/grantreview/GombertLabedens2025.txt ~line 967) at ~4 min each = 56 flash frames, prevalence 2.9%. TP = 0.90 x 56 = 50; FP = 0.044 x 1,864 = 82; PPV = 50/(50+82) = 38%, i.e. 62% of positive frames false. Even at 99% specificity, PPV = 73%. Tsiartas 2021 (/root/grantreview/Tsiartas2021.txt): "we oversampled the HF class to 1:2 ratio between the HF and non-HF regions"; "we randomly split our data in two sets, 80% of data for training the Decision Tree and 20% for evaluating" — no subject-wise held-out validation across only three participants.
Fix
Report the expected PPV at the study's own event base rate, and pre-specify the minimum PPV required before the phase analysis proceeds. Re-validate with subject-wise cross-validation, not random 80/20 splits over three people. Add a manual adjudication step: have a blinded scorer confirm each candidate event against the raw multi-signal trace, and report the adjudicated event count as the analysis N.
What the refuter said
AGAINST: frame-level specificity does not map one-to-one onto event counts — Tsiartas "extracted the HF regions" from the frame decisions and evaluated against annotations with a ±90 s matching window, so scattered false frames need not become analysed events, and the PPV figure is an upper-bound illustration rather than a measurement. SURVIVES. The arithmetic checks (1,920 frames per 8 h; 3.5 flashes at ~4 min ≈ 56 frames, prevalence 2.9%; TP 50 vs FP 82; PPV 38%), the 3.5 flashes/night is verified in GombertLabedens2025.txt ("an average of 3.5 hot flashes per night"), and the two structural facts are verbatim in Tsiartas2021.txt: "we oversampled the HF class to 1:2 ratio between the HF and non-HF regions" and "we randomly split our data in two sets, 80% of data for training the Decision Tree and 20% for evaluating" — a frame-level random split over three participants, so no subject-wise held-out validation exists. The draft quotes "over 90% sensitivity at 95.6% specificity" (L244-245) as a settled operating characteristic with no mention of prevalence, PPV, n=3, 27 flashes, or that the wrist SC channel was a wired lab sensor on a PSG amplifier rather than an EmbracePlus. Every false positive enters the primary model with a fabricated onset time. Major.
The problem
Six sequential exclusions are named and none is quantified. Published figures make the survival rate small: about half of nocturnal flashes occur in wake epochs and a fifth in N1, leaving under a third for consolidated NREM; ZMax and Empatica each lose roughly a third of recordings to quality criteria; and the stable-bout, estimation-window and onset-reliability filters cut further. A rough chain on the draft's own optimistic 137 women x 4 nights x 3.5 flashes gives order 100-250 analysable events and perhaps 30-80 registered ones — thin for a two-degree-of-freedom circular test that is additionally attenuated by timing jitter. The Bial regulation makes the aims section decisive for assessment; an unquantified N in the section that is supposed to establish feasibility is the first thing a reviewer will ask for.
In the draft — line 204-206
Expected yield is estimated by applying the planned exclusions sequentially - EEG night, NREM, stable NREM, estimation window, artefact-free recording, reliable onset timing.
Evidence
Gombert-Labedens 2025 (/root/grantreview/GombertLabedens2025.txt ~line 995) reporting the GnRH-agonist model study: "hot flashes occurred primarily in association with light (stage N1) sleep (20%) and wake epochs (51%)". Quality attrition: Esfahani 2023 report only "55-71%" of a clinical dataset met ZMax quality thresholds (https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full); Parry & Briganti 2026 excluded 31% (69 of 100) on quality (https://www.medrxiv.org/content/10.64898/2026.06.10.26355348v2). BIAL_CONTEXT.md: the Specific Aims "are essential for the assessment of the application".
Fix
Insert a numbered attrition table with an assumed retention fraction and source for each of the six steps, ending in expected total events, expected registered events, and expected events per participant. State the power at that N for a stated effect size under the jitter attenuation above. If the yield is under a few hundred events, say so and reframe the aims.
What the refuter said
Strongest case against the arithmetic, which does not hold. The 51%-in-wake figure is from the wrong population: I read GombertLabedens2025.txt in context and it describes "an experimental model of new-onset hot flashes in young pre-menopausal individuals, treated with a gonadotropin-releasing hormone agonist", where "hot flashes occurred primarily in association with light (stage N1) sleep (20%) and wake epochs (51%)". The draft's own headline source, Baker 2019 [179] in the same review, gives a very different split for the relevant population: "most of the detected hot flashes (51%) were associated with sleep disruption, and 29% of hot flashes occurred in undisturbed sleep; the remainder happened during wakefulness or were ambiguous" - i.e. roughly 80% with onset in sleep, not under a third. Arousal-association is not the same as wake-onset, and the finding conflates them. So the 100-250 estimate is not defensible. The core criticism nonetheless survives untouched and matters more than the arithmetic: L204-206 names six sequential exclusions and quantifies none; the power section defers everything to data not yet in hand ("Simulations use the observed counts and precision", L318-320); and no expected analysable-event count appears anywhere - in a design whose own Study Design section says "power depends on event count" (L192-193). Under BIAL_CONTEXT.md L42-44, where the stated objectives are "essential for the assessment", that is the first thing an assessor will ask for. Major on the gap, not the numbers.
The problem
The draft names six sequential exclusions and reports the result of none of them. "Expected yield is estimated by" is a promise, not an estimate: there is no expected number of analysable flashes, no expected number of participants, no expected number of nights, and no stated assumption for any step. Four things make the promise worse than empty. (1) 222 is the host trial's enrolment *target*, inclusive of a 10% dropout allowance (50 completers per cell × 4 = 200); 137 is 62% of a number not yet achieved. (2) Device participation is optional and revocable in the host protocol — the EEG/wearable denominator is not 222 and is not knowable from the trial's design. (3) The screening criterion is self-reported flash interference with sleep, not objectively detected flashes; 62% of women saying flashes disturb their sleep does not imply 62% yield detectable nocturnal flashes on the recorded nights. (4) The one hard number available for the largest single loss — usable ZMax nights — comes from the draft's *own* cited validation paper and is far worse than the draft implies: usable-recording rates of 68%, 63%, 55% and 9/23 across its datasets. Doing the cascade even roughly: ~137 women × 4 baseline nights = 548 nights; × ~0.6 usable EEG = ~330; × ~0.9 concurrent wristband = ~300; × 3.5 objectively recorded flashes per night (the draft's own De Zambotti figure) = ~1000; minus REM-suppressed flashes = ~800; × an unstated stable-NREM-bout fraction that will be low in an ISI≥10 insomnia population = a few hundred; × an unstated reliable-onset fraction. The result may still be adequate for a 2-df test — but the draft must show the arithmetic, because the reviewer's default assumption in its absence is that it was not done.
In the draft — line 200-206
The analysis sample comprises those whose nocturnal flashes interfere with sleep: of the first 65 screened, 40 (62%, 95% CI 49-72) meet this criterion, projecting to roughly 137 women. Expected yield is estimated by applying the planned exclusions sequentially - EEG night, NREM, stable NREM, estimation window, artefact-free recording, reliable onset timing.
Evidence
Host trial, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: "Target N: 222 participants (accounting for 10% dropout)", "n=50 per arm post-10% dropout allowance"; "Participants are free to opt out of EEG and smartwatch measurements at any time without consequences for study participation, treatment or otherwise"; inclusion "Insomnia Severity Index (ISI) ≥10". ZMax usable-recording rates, Esfahani et al. 2023 (cited by the draft at line 273 for a *favourable* statistic), https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full.pdf: "Dataset 2: 66/97 recordings included", "Dataset 4: 43/96 included", "Dataset 5 (IBD cohort): 280 (55%) were deemed entirely useful of 512 recordings". Draft's own event rate, line 40-41: "Women have on average 3.5 objectively recorded flashes per night".
Fix
Replace the promise with a table: one row per exclusion, each with an assumed retention fraction, its source, and the running count of participants / nights / flashes. Anchor the EEG-night row on Esfahani's own 55-68% usable rate rather than on its bias statistic. Propagate the screening CI rather than the point estimate (48-71% of an achieved enrolment of N, not 62% of a target of 222). State the resulting expected flash count as a range and carry the lower bound into the power simulation. Add the one number the cascade most needs and cannot get from the literature — the stable-NREM-bout fraction in fragmented insomnia sleep — as a quantity to be established in the first work package.
What the refuter said
AGAINST: only that the finding's own rough cascade applies de Zambotti's 3.5 flashes/night (from an unselected n=34 sample) to a cohort deliberately selected for flashes that interfere with sleep, so its arithmetic is illustrative rather than authoritative — which the finding itself says. That is the only weakness I can find. Everything else verified, on all four limbs. (1) Parent protocol: "Expecting a drop-out of 10%, the total number of participants to be recruited is n=222" — so 137 is 62% of a recruitment TARGET, not of achieved completers. (2) "Participants are free to opt out of EEG and smartwatch measurements at any time without consequences" — the device denominator is not 222 and is unknowable from the design. (3) The screening criterion at L201-203 is self-reported interference ("those whose nocturnal flashes interfere with sleep"), not objectively detected nocturnal flashes. (4) Esfahani et al. 2023, cited by the draft at L272 for a favourable bias statistic, reports usable-recording rates of 9/23, 66/97, 17/23, 43/96 and "280/512 recordings (55%) deemed entirely useful" — the single largest loss in the cascade, omitted. And the core is airtight from the draft: L204-206 names six sequential exclusions and reports the result of none, so "Expected yield is estimated by" is a promise, not an estimate. Major, and among the most actionable findings in this batch.
The problem
The quoted figures are accurate to the source abstract but are not generalisable performance estimates and the draft presents them without any of the qualifiers that make them interpretable. (i) N=3 postmenopausal women, 27 events, one ~12 h lab session in a sound-attenuated temperature-controlled bedroom. (ii) Cross-validation was a random 80/20 split of 15 s frames repeated 5-fold — with only 3 participants and 27 events, training and test folds necessarily contain frames from the same subjects and the same flashes, so the estimate is inflated by within-event and within-subject leakage; there is no subject-wise validation and none is possible at N=3. (iii) The operating point was produced by oversampling the positive class to 1:2, a class balance nothing like an overnight recording, so the specificity figure does not describe behaviour at natural prevalence. (iv) The body of the paper reports the constrained specificity as 96.5±1% and "At 96.5% specificity", not 95.6% — the number the draft quotes to one decimal place is internally inconsistent in its own source, and no numeric sensitivity value appears anywhere in the text (only in unlabelled figures).
In the draft — line 244-245
reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance, with sleep-onset performance reported separately
Evidence
Tsiartas2021.txt L29-31: "Three women (Age, mean ± SD: 55.6 ± 0.6 y) who reported having daily HFs participated in a ~12h lab-based study, that encompassed an overnight. A total of 27 physiological HFs were recorded from the women." L125-128: "In all the analyses, we randomly split our data in two sets, 80% of data for training the Decision Tree and 20% for evaluating the system. We repeated this process using a 5-fold, cross-validation setup until all data were used for testing." L135-136: "To ensure similar specificity (96.5±1%) across analyses, we oversampled the HF class to 1:2 ratio between the HF and non-HF regions." L167: "At 96.5% specificity, the SC+ set showed better HF classification performance than the SC set". Abstract L24-25: "the multi-sensor approach achieved above 90% sensitivity at 95.6% specificity".
Fix
Replace the performance clause with the provenance: "Detection follows the multi-sensor approach of Tsiartas et al. (2021), a feasibility study in three postmenopausal women (27 laboratory-recorded flashes) that reported above 90% sensitivity at ~96% specificity under event-wise cross-validation and an oversampled 1:2 class balance. Because that estimate is not subject-wise, not at natural overnight prevalence, and not obtained with our hardware, we treat detector performance as an outcome of this study rather than an assumption, and report subject-wise sensitivity, false alarms per hour of NREM, and onset-timing error in a validation subsample."
What the refuter said
Strongest case against: the draft quotes the abstract accurately, it does flag state-dependence ("with sleep-onset performance reported separately", L245-246), and citing a published detector by its headline operating point is ordinary practice in a character-limited methods section. Sub-claim (iv) also fails on inspection: Tsiartas2021.txt L135-136 gives the constrained specificity as "96.5+-1%" and the abstract (L24-25) reports 95.6%, which falls inside 96.5+-1 - these are a target band and an achieved value, not an internal inconsistency, so that part of the finding is wrong. Everything else I verified verbatim in /root/grantreview/Tsiartas2021.txt: "Three women (Age, mean +- SD: 55.6 +- 0.6 y)" and "A total of 27 physiological HFs" (L29-31); one "~12h lab-based study" at the SRI sleep laboratory; "we randomly split our data in two sets, 80% of data for training... 5-fold, cross-validation" (L125-128) with no subject-wise holdout and none possible at N=3, so training and test folds necessarily share subjects and events; "we oversampled the HF class to 1:2 ratio" (L135-136). The entire measurement chain of the proposal rests on this classifier, quoted with none of these qualifiers, on different hardware from the study's EmbracePlus. A hostile expert reviewer will read the four-page EMBC paper and find N=3 in the first column. Major stands.
The problem
The design's advertised advantage over stimulus studies is that it captures quiet flashes — events with neither arousal nor report. That cell is also exactly where detector false positives accumulate, because a false positive has no flash to arouse or be felt. Tsiartas reports sensitivity and specificity only, on an oversampled class balance; no PPV, no F1, no false alarms per hour. Illustrative arithmetic on the draft's own numbers (my calculation, stated as such): an 8 h night is ~1920 fifteen-second frames; at ~4% false-positive rate that is ~77 false-positive frames per night, against 3.5 true flashes of 1-5 min duration, i.e. ~14-70 true frames. Under those assumptions false positives equal or exceed true positives and PPV falls below 50%. Since the primary contrast is registration probability across ISF phase, and false positives are unregistered by construction, any phase-dependence of the false-positive rate — which is expected, because the detector's cardiac and electrodermal inputs are themselves coupled to the 0.02 Hz rhythm — will appear as the hypothesised effect.
In the draft — line 283 (with 52-54, 246-249)
Both are fitted to every detected flash, so events passing without arousal or report contribute.
Evidence
Tsiartas2021.txt L137-140: "HF performance was evaluated in terms of system sensitivity (percent of true positives) and specificity (percent of true negatives) in HF detection compared to gold standard sternum SC expert evaluation" — no PPV, no per-hour false-alarm rate, and L135-136 confirms the evaluation set was rebalanced: "we oversampled the HF class to 1:2 ratio between the HF and non-HF regions." That the detector's own inputs carry the rhythm is established, not speculative: OsorioForero2021.txt L399-401: "1-Hz-LC stimulation during NREMS disrupted the infraslow HR variations (Figure 7G) and decreased their anticorrelation with sigma power (Figure 7H)."
Fix
Add a pre-registered detector-characterisation step with a hard gate: estimate false alarms per hour of NREM against a sternal SC reference in a lab or first-night subsample, and pre-specify a maximum tolerable false-alarm rate below which the primary analysis proceeds. Additionally pre-specify a negative-control analysis: test whether *detector-positive-but-physiologically-implausible* events (e.g. no accompanying temperature rise) show the same phase distribution as the full set. Add: "Because a false positive is unregistered by construction, phase-dependent false alarms would mimic the predicted effect; we therefore estimate the false-alarm rate and its phase distribution before testing the primary hypothesis."
What the refuter said
AGAINST: the draft already anticipates detector-phase coupling: "It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent" (L249-251), and pre-specifies a fallback detector. The finding's arithmetic is explicitly "illustrative", frame specificity does not translate cleanly to event counts (Tsiartas "extracted the HF regions" from contiguous frame decisions and matched at ±90 s), so isolated false frames may never become events. Survives anyway, because the pre-emption addresses the wrong quantity. The draft's safeguard concerns phase-dependent *sensitivity* (misses); the finding concerns *false positives*, which the draft never mentions. I verified in Tsiartas2021.txt that only sensitivity and specificity are reported ("HF performance was evaluated in terms of system sensitivity (percent of true positives) and specificity (percent of true negatives)") and that the operating point was measured on a rebalanced set ("we oversampled the HF class to 1:2 ratio between the HF and non-HF regions") — no PPV, no false alarms per hour, at a prevalence far above a real night's. The logical core is airtight and unaddressed: a false positive has no flash to arouse or be felt, so it lands by construction in the quiet-flash cell that carries the hypothesis, and any phase structure in the false-positive rate appears as the predicted effect. Kept at major.
The problem
"Simulation shows" presents as an obtained result something the next sentence says depends on counts and precision that do not yet exist — a conflation of a planned analysis with a finding. Consequently the proposal contains no expected event count for the primary analysis cell, no minimum detectable effect, and no power figure of any kind. It also contradicts itself on what limits power: the Study design section says power depends on event count, and the Power section says it does not. The arithmetic matters and is not hard: 137 women × 4 EEG nights × 3.5 flashes ≈ 1900 detected events before exclusions, but 20-30% of nocturnal flashes occur when the participant is already awake, and the draft then applies six sequential exclusions (EEG night, NREM, stable NREM, estimation window, artefact-free, reliable onset) for which no retention rates are given. At a plausible 50% retention per stage the primary cell falls to the low hundreds, and the registered subset lower still — a two-degree-of-freedom test on sin/cos with participant and night random intercepts and a sparse binary outcome would then detect only large effects. Bial's regulation makes the Specific Aims the basis of assessment; an aim with no feasibility arithmetic is the weakest possible version of this proposal.
In the draft — line 317-320 (with 191-192)
simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle. Simulations use the observed counts and precision, and the minimum detectable effect is reported.
Evidence
Internal contradiction in the draft: L191-192 "The unit of analysis is the flash, not the participant: power depends on event count, hence multiple nights" versus L315 "The binding constraint is measurement precision rather than event count". Input numbers are available from the cited literature — GombertLabedens2025.txt L917-923: "an average of 3.5 hot flashes per night and that 70% of hot flashes were associated with an arousal from sleep, with only a minority (20%) occurring without disturbance to sleep; the remaining hot flashes occurred when the participant had already been awake for at least one minute" and L925-928: "most of the detected hot flashes (51%) were associated with sleep disruption, and 29% of hot flashes occurred in undisturbed sleep". (Event-count cascade above is my arithmetic on the draft's own figures, not a source claim.)
Fix
Run the simulation now and report numbers. Give the expected-yield table the draft already promises at L204-207 with an explicit retention rate at each of the six exclusion steps, the resulting expected count of detected flashes in stable NREM, the expected registered and aroused subsets, and the minimum detectable odds ratio for the 2-df sin/cos test at 80% power under jitter of 5, 10 and 15 s. Change "simulation shows" to "simulations reported here show" only once they have been run; otherwise write "we will simulate". Resolve the contradiction by stating that both count and precision bind, and that the design is powered for the worse of the two.
What the refuter said
STRONGEST DEFENCE: the two tenses can be read as two different simulations — a generic/analytic one establishing the attenuation relationship (available now) and a final one using observed counts (later). That is not strictly a contradiction. And the draft does promise the MDE will be reported. WHY IT SURVIVES, HARDER THAN THE FINDING ARGUES: the claim itself is mathematically FALSE, which I verified independently of the reviewers' code. Onset jitter attenuates the first-harmonic logit amplitude by λ = exp(−σ_φ²/2) but never nullifies it, so the 2-df noncentrality n·p̄(1−p̄)(λA)²/2 is strictly increasing in n and power → 1 for every finite jitter. The required N is inflated by exactly 1/λ² = exp((2πσ_t/T)²) — at σ_t = T/4 that is exp(π²/4) = 11.8x. Power never "ceases to improve". So DRAFT.md L317-318 ("simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle") states a false result, and it is the sole justification offered for the Power section containing no number at all: no expected event count, no MDE, no power figure. Bial's regulation makes the Specific Aims "essential for the assessment of the application" (BIAL_CONTEXT.md L42-44) and the call requires expected-output indicators. The finding's input numbers check out (GombertLabedens2025.txt: 3.5/night, 70% aroused, 20% undisturbed; Baker n=86: 51%/29%, remainder wakeful/ambiguous — so "20-30% already awake" is right for Baker, ~10% for de Zambotti, a small overstatement). The retention cascade is explicitly labelled the reviewer's arithmetic, not a source claim, which is the correct discipline. Severity stands at major: a confirmatory pre-registered primary test with no feasibility arithmetic is the single most likely reason an assessor rejects this.
The problem
This sentence describes a calculation in the passive future and never performs it. The proposal states nowhere how many analysable flash events it expects. It also states nowhere what onset-timing precision is achievable — precision is itself deferred to a feasibility outcome (l.275-276: "Frontal expression and achievable phase precision are themselves feasibility outcomes"). Since the draft asserts that power is jointly determined by these two unknowns (l.315-320), power is entirely unspecified. The one quantitative power claim is presented as an already-obtained result but no simulation is reported: "simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle" gives no parameters, no assumed effect size, no jitter grid, no target power, and no figure. "Simulations use the observed counts and precision, and the minimum detectable effect is reported" (l.319-320) again defers everything to after funding. A reviewer cannot distinguish this design from one with 80 analysable events and one with 800. Rough arithmetic from the draft's own numbers gives the order of magnitude that should have been supplied and then discounted: 3.5 detected flashes per night (l.40) x 4 EEG nights (l.65) x 137 women (l.204) = about 1918 detected flashes before ANY of the six named exclusions, before false positives, and before the further restriction that the flash fall inside a stable NREM bout of three ISF cycles (l.345-347). The cascade could plausibly remove 50-90% of these; the proposal must say which.
In the draft — line 204-206
Expected yield is estimated by applying the planned exclusions sequentially - EEG night, NREM, stable NREM, estimation window, artefact-free recording, reliable onset timing.
Evidence
Draft l.204-206 (cascade named, never applied), l.275-276 ("achievable phase precision are themselves feasibility outcomes"), l.317-319 ("simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle" — asserted, not reported), l.319-320 ("Simulations use the observed counts and precision" — future tense). No number of analysable events appears anywhere in DRAFT.md.
Fix
Add a yield table with an explicit number at each step and a stated source for each retention rate (pilot data, published estimate, or assumption), ending in a single expected count of analysable flashes with a pessimistic and optimistic bound; then report a power curve over that count crossed with a jitter grid (e.g. 2, 5, 8, 12 s SD) for a named minimum effect of interest (e.g. an odds ratio of 1.8 between the preferred and anti-preferred phase, or a 15 percentage-point difference in registration probability). State the target power. If pilot retention rates do not yet exist, say so explicitly and present the power curve over a range of retention rates rather than deferring the whole calculation.
What the refuter said
I tried hard to break this and could not. Verified by direct inspection: L204-206 names the exclusion cascade in the passive future ('Expected yield is estimated by applying...') and never applies it; I grepped every numeral in the entire methods block L171-350 and the only sample-related numbers are 222, 137, 65, 40, 62, and the CI 49-72 — there is no analysable event count anywhere in DRAFT.md. Precision is deferred to a feasibility outcome at L275-276. The one quantitative power statement, L317-319 'simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle', is presented as an obtained result with no parameters, no assumed effect size, no jitter grid, no target power and no attribution; L319-320 then defers the calculation itself ('Simulations use the observed counts and precision'). The finding's sanity arithmetic checks: 3.5 x 4 x 137 = 1918. STRONGEST DEFENCE: for a secondary analysis of a cohort still recording, event yield is genuinely unknown in advance, and pre-committing to report the MDE is honest. That defence fails on the point that matters — the draft itself asserts that power is jointly determined by count and precision and then supplies neither, so a reviewer deciding today cannot distinguish this design from one with 80 events or one with 800. Downgraded fatal to major: under-specified rather than invalid, and fixable in a paragraph.
The problem
Onset jitter acts as a multiplicative attenuation on the first circular harmonic, not as a variance floor. It shrinks the effect size but leaves the estimator consistent, so power is strictly increasing in N at every jitter level and approaches 1. There is no jitter threshold at which additional events stop helping. The claim is not a conservative simplification; it is the opposite of what the model does, and it is the only quantitative statement in the Power and precision section.
In the draft — line 316-318
onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle.
Evidence
Analytic: for wrapped-normal onset error of SD sigma_t against a cycle T, the attenuation is lambda = |E[e^{i*eps}]| = exp(-(2*pi*sigma_t/T)^2/2). Simulation (400,000 events, /root/grantreview/sim/01_attenuation.py) reproduces it to 2 decimals: sigma_t = 5 s -> 0.805 sim vs 0.821 analytic; 12.5 s -> 0.291 vs 0.291; 25 s -> 0.007 vs 0.007. Power of the 2-df LRT at sigma_t = 12.5 s (the draft's quarter-cycle), OR(best:worst phase)=2, base rate 0.20 (/root/grantreview/sim/02_power_vs_N.py): N=500 -> 0.082; N=2,000 -> 0.190; N=11,823 -> 0.800; N=50,000 -> 1.000. Even at sigma_t = 25 s (half the cycle), power reaches 0.089 at N=1,000,000. Required N is inflated by exactly 1/lambda^2 = exp((2*pi*sigma_t/T)^2): 1.5x at 5 s, 4.0x at 9.4 s, 11.8x at 12.5 s, 554x at 20 s, 19,334x at 25 s.
Fix
Replace with: "Onset jitter attenuates the first circular harmonic by a factor exp(-(2*pi*sigma_t/T)^2/2), so it does not bound power but inflates the required event count by exactly the reciprocal of that factor squared. At an onset-timing SD of 5 s against the 50-s cycle the required N rises 1.5-fold; at 9.4 s it rises 4-fold and half the true effect is lost; at a quarter of the cycle (12.5 s) it rises 11.8-fold. The design is therefore feasible only if onset timing is held below about 5 s SD, and the achievable precision is itself a feasibility outcome."
What the refuter said
AGAINST: read charitably, "power ceasing to improve" could be practical shorthand — at 12.5 s jitter the finding's own numbers put required N at ~11,800 events against a design that will yield hundreds, so within reachable N power really does barely move, and the sentence's *conclusion* (precision, not count, is binding) is correct. SURVIVES, because it is stated as a simulation result and it is analytically false. I re-derived the attenuation independently: for wrapped-normal onset error, σ_ε = 2πσ_t/T radians and λ = exp(−σ_ε²/2); at σ_t = 5, 12.5, 25 s against T = 50 s this gives 0.821, 0.291, 0.0072, matching the finding's simulation output in /root/grantreview/sim/out_01_attenuation.json (sim 0.805 vs analytic 0.821 at 5 s). Jitter is a multiplicative shrinkage on the first circular harmonic, so the estimator stays consistent and power is strictly increasing in N at every jitter level — there is no threshold beyond which events stop helping. Required N inflates by exactly 1/λ² (11.8× at 12.5 s). This is the only quantitative claim in Power and precision, it is the opposite of what the model does, and a statistician on the Scientific Board will catch it. Downgraded from fatal because the corrected statement supports rather than undermines the proposal's own thesis; the damage is to credibility, not to the design.
The problem
The entire power argument is deferred to after data collection. No expected event count, no target effect size and no minimum detectable effect appear anywhere in the document, so the reviewer is asked to fund a confirmatory test whose detectability is unstated. When the numbers are supplied the design does not clear a plausible effect. Note this is THIS REVIEW's yield model, built because the draft gives none; every assumption is stated and varied.
In the draft — line 318-320
Simulations use the observed counts and precision, and the minimum detectable effect is reported.
Evidence
Yield model (/root/grantreview/sim/08_yield_v2.py), central scenario, all factors stated: 137 women x 4 nights x 0.80 usable EEG nights = 438 nights; x 3.5 objectively detected flashes/night (de Zambotti et al. 2014 abstract, verified: "Women had an average of 3.5 (95%CI:2.8-4.2, range=1-9) objective hot flashes per night") = 1,534 detected; x 0.80 onset in sleep; x 0.88 NREM; x 0.75 N2/N3 = 810; x bout geometry exp(-150/240 s) = 0.535; x 0.80 artefact-free and reliably timed = 347 analysable flashes, of which ~52 registered at an assumed 15% press rate. Power for OR(best:worst phase)=2 at that yield: 0.189 with zero onset jitter, 0.141 at 5 s, 0.081 at 9.4 s, 0.060 at 12.5 s. Minimum detectable effect at 80% power, 5 s jitter (including the measured lambda_edge=0.78 for a pre-onset window): OR = 7.86. Pessimistic scenario (109 women, 2.5 flashes/night, 150-s bouts): N = 71, ~4 registered events, MDE unreachable. Optimistic (160 women, 5 flashes/night, 420-s bouts, 30% press rate): N = 1,249, MDE OR = 2.33 at 5 s jitter. To reach 80% power for OR=2 at 5 s jitter and a 15% press rate requires N = 3,069 analysable flashes: a 9-fold shortfall, equivalent to ~35 EEG nights per woman at the present cohort size, or ~1,200 women at 4 nights.
Fix
State the yield model explicitly in the proposal with its assumptions, give the expected number of analysable flashes and registered events with a pessimistic/central/optimistic range, name a target effect size (an OR between best and worst phase), and report the resulting power before submission rather than after data collection. If the central yield is a few hundred flashes, either reframe the primary outcome as cortical arousal (13x more informative, see finding nested-hierarchy-inverted) or present the study as a feasibility and effect-estimation study rather than a confirmatory test.
What the refuter said
AGAINST the quantitative half, which does not stand as fact. The 9-fold shortfall is driven by an invented 15% press rate; on the finding's own alternative (43%, the draft's implied figure) the stated requirement falls to 971 against a central yield of 347 — a ~2.8x gap, not an order of magnitude — and it shrinks further at OR=3 or if T1/T2 nights are admissible. Several factors are assumptions the draft cannot be blamed for (exp(-150/240) bout geometry; 0.80 artefact-free; 0.80 usable EEG, which is generous against Esfahani's verified 55-68%). The finding is honest that this is its own model, but the "order of magnitude" figure must not be reported to the applicant as established. The lead claim is CONFIRMED by reading and is what matters. L318-320 is entirely prospective — "Simulations use the observed counts and precision, and the minimum detectable effect is reported" — and I confirmed no expected event count, no target effect size and no MDE appears anywhere in DRAFT.md. A Bial call with a verified ~19% success rate (434 received, 82 approved) is being asked to fund a confirmatory test whose detectability is stated nowhere, in a design the draft itself says is precision-limited. Major: the fix is a scenario table with explicit assumptions, not a redesign.
The problem
Registration is largely a subset of arousal, as the draft itself states. Two consequences the draft does not draw. (a) Information: if registration is arousal thinned by an independent press probability, the arousal-model amplitude transmits to the registration model with a factor of 0.35 on the logit scale, and the lower base rate compounds it, so the registration test carries 7.6% of the non-centrality of the arousal test and needs 13x the events. The proposal thus nominates its least informative analysis as the confirmatory one. (b) Specificity: a significant registration effect is produced by a pure arousal-gating truth, so rejecting the primary null does not support the interpretive claim at line 53-56 that comparing the two analyses distinguishes cortical gating from gating of access. The only analysis that does distinguish them - registration among aroused flashes - is designated exploratory, even though it is as powerful as the primary under the access-gating truth.
In the draft — line 310-311
The primary registration model is confirmatory; all else is exploratory.
Evidence
Analytic (/root/grantreview/sim/11_nested.py), P(arousal)=0.70 (de Zambotti et al. 2014: "69.4% of hot flashes were associated with an awakening"), P(registered)=0.15: d logit(q*p)/d logit(p) = 0.353, non-centrality ratio registration:arousal = 0.0756, i.e. 13.2x the events required. Simulation at N=347, no onset jitter, 1,200 replicates. Truth = arousal gating only: primary registration test rejects 0.077 (OR=2), 0.135 (OR=3), 0.194 (OR=5); arousal test rejects 0.434 / 0.845 / 0.996; registration-among-aroused correctly stays at 0.061 / 0.054 / 0.046. Truth = access gating only: primary 0.247 / 0.534 / 0.895; arousal test correctly 0.065 / 0.045 / 0.058; registration-among-aroused 0.267 / 0.562 / 0.918. Correlation between the primary and supporting LRT statistics is only +0.00 to +0.23, so reporting both is not a multiplicity problem - the problem is that neither identifies the target construct.
Fix
Invert the hierarchy and pre-register two confirmatory tests: (1) phase -> cortical arousal among all detected flashes, as the primary (13x the information); (2) phase -> registration among aroused flashes, as the co-primary that isolates access to awareness, with the joint distribution reported. Demote registration-among-all-detected to a descriptive summary, and state explicitly: "Because registered flashes are largely a subset of aroused ones, a phase effect on registration among all detected flashes is produced by arousal gating alone (simulated rejection rate 0.14 at OR=3 under an arousal-only truth) and is therefore not interpretable as evidence about conscious access."
What the refuter said
AGAINST: nominating the scientifically novel outcome as confirmatory is a legitimate and common choice — you pre-register the question you care about, not the best-powered proxy. And the draft is not blind to the nesting: L299-302 states registration is "expected to be largely a subset of aroused ones" and commits to "registration additionally modelled among aroused flashes", while L54-56 says "Comparing the two indicates...", implying the pair is read jointly. So the discriminating analysis exists in the plan. Both limbs survive anyway. I checked the analytic core by hand: d logit(qp)/d logit(p) = (1-p)/(1-qp), which at p=0.694 (de Zambotti's verified "69.4% of hot flashes were associated with an awakening") and q=0.15 gives ~0.34, matching the claimed 0.353; combining the squared attenuation with the lower outcome variance puts the registration test's non-centrality at roughly 5-8% of the arousal test's, so "the confirmatory test carries a small fraction of the information of the one called supporting" is sound. Limb (b) is the stronger point and is pure logic: if registration is nested in arousal, a pure arousal-gating truth produces a significant primary result, so rejecting the primary null cannot distinguish the two accounts L54-56 claims to distinguish — and the only test that can, registration-among-aroused, is demoted to exploratory by L310-311 ("The primary registration model is confirmatory; all else is exploratory"). Major: promote it, or drop the interpretive claim.
The problem
Independent timing error attenuates a circular effect by a multiplicative factor, it does not impose a ceiling. For Gaussian onset error with SD sigma and angular frequency w = 2*pi/50 = 0.1257 rad/s, the estimated sin/cos coefficients are attenuated by exp(-sigma^2 w^2 / 2): 0.82 at sigma = 5 s, 0.45 at 10 s, 0.29 at 12.5 s (the draft's quarter-cycle), 0.04 at 20 s. Required N scales as the inverse square of that: about 1.5x at 5 s, 4.9x at 10 s, 12x at 12.5 s, 550x at 20 s. Power keeps improving at every finite sigma. The false ceiling claim is what allows the draft to assert that event count is not the binding constraint and thereby to state no N at all. Two tenses also conflict: "simulation shows" (done) versus "Simulations use the observed counts and precision" (not yet possible).
In the draft — line 317-320
simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle. Simulations use the observed counts and precision, and the minimum detectable effect is reported.
Evidence
Standard attenuation result for circular regressors under independent phase jitter; computed above from the draft's own 50-s cycle and quarter-cycle threshold. Draft states no event count, no effect size, no power anywhere; line 204-206 promises only that "Expected yield is estimated by applying the planned exclusions sequentially".
Fix
Delete the ceiling claim. Replace with: "Independent onset jitter of SD sigma attenuates the circular effect by exp(-sigma^2 w^2/2), inflating the required event count by the inverse square of that factor: 1.5x at sigma = 5 s and 12x at sigma = 12.5 s against a 50-s cycle." Then give the numbers a reviewer needs: expected analysable events after each exclusion, expected registered events, the assumed effect size on the log-odds scale, and the resulting power. Present the simulation now, with plausible ranges, not as a future deliverable.
What the refuter said
Strongest case against: read charitably, "power ceasing to improve" is a practical statement — at a quarter-cycle of jitter the attenuation is so severe that reachable N buys almost nothing — and the draft does promise a data-based power analysis rather than denying one ("Simulations use the observed counts and precision, and the minimum detectable effect is reported", L318-320). That is defensible for a study whose event yield is genuinely unknown before data. But the claim as written is false and I re-derived it independently: with independent Gaussian onset error the sin/cos coefficients are attenuated by exp(-sigma^2 omega^2 / 2), omega = 2*pi/50 = 0.1257 rad/s; at sigma = 12.5 s that is exp(-(pi/2)^2/2) = 0.291, so required N inflates by 1/lambda^2 = 11.8 — a constant multiplier, not a ceiling. The reviewers' simulation confirms it directly: "power is strictly increasing in N at every jitter level; it never plateaus" (N=500 -> 0.082; N=11,825 -> 0.800; N=50,000 -> 1.000). And the consequence is verified: grep finds no event count, no effect size and no power figure anywhere in DRAFT.md, with L204-206 promising only a sequential-exclusion yield estimate. A false mathematical claim that licenses omitting every quantitative design figure; moderate rather than major because the analysis is deferred, not refused.
The problem
The performance figures are quoted as if they were device-level validation. They come from five-fold cross-validation over 27 hot flashes contributed by three women in a single ~12 h laboratory session, with the hot-flash class oversampled to a 1:2 ratio. The participants were postmenopausal after natural menopause, not perimenopausal. The sensors were a customised array of consumer-grade components wired into a Compumedics recording system — not the Empatica EmbracePlus the draft will use, which has different sensors, placements and sampling rates. Transferring a three-subject decision tree to 137 perimenopausal women recorded at home on different hardware, without any recalibration plan, is the largest undisclosed risk in the proposal, and a reviewer who opens the paper will find the sample size in the first paragraph of the Methods.
In the draft — line 241-245
Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier, reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance
Evidence
/root/grantreview/Tsiartas2021.txt lines 28-35: "Three women (Age, mean ± SD: 55.6 ± 0.6 y) who reported having daily HFs participated in a ~12h lab-based study, that encompassed an overnight. A total of 27 physiological HFs were recorded from the women." Lines 40-42: "Women were free from major mental and medical conditions, had undergone natural menopause, and none of them was currently taking hormone therapy." Lines 59-63: "Signals from a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700 - 1024 Hz; T sensor: TI-TMP36GT9Z – 16 Hz) were collected from each participant's wrist". Line 135-136: "To ensure similar specificity (96.5±1%) across analyses, we oversampled the HF class to 1:2 ratio".
Fix
State the provenance and add a transfer step: 'Detection follows the multi-sensor feature approach of Tsiartas et al. (2021), which reached over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance — in a three-participant, 27-event laboratory pilot on bespoke sensors. We therefore re-fit and validate the classifier on our own device in a calibration subsample with concurrent sternal skin conductance in N participants, and report its sensitivity, specificity and onset latency before any phase analysis.' Budget the concurrent sternal-SC validation nights explicitly.
What the refuter said
STRONGEST DEFENCE: the four sensing modalities transfer exactly (the parent protocol's EmbracePlus provides "multisensory nocturnal hot flash estimation (skin conductance, temperature, heart rate, movement)"), so this is not an exotic method transplant; and the draft does pre-specify a re-detection from skin conductance and temperature alone (DRAFT.md L251-254) to guard against a cardiac-weighted detector preferentially finding aroused flashes — evidence the applicants have thought about detector bias. WHY IT SURVIVES: every factual claim checks out verbatim in Tsiartas2021.txt — "Three women (Age, mean ± SD: 55.6 ± 0.6 y) ... a ~12h lab-based study"; "A total of 27 physiological HFs"; "had undergone natural menopause" (postmenopausal, not the draft's perimenopausal population); "a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700 - 1024 Hz; T sensor: TI-TMP36GT9Z – 16 Hz) ... integrated with the Compumedics recording system"; "we oversampled the HF class to 1:2 ratio"; and "we randomly split our data in two sets, 80% ... 20%" with 5-fold CV — i.e. frame-level random splitting over three participants, so the 90%/95.6% figures are not subject-independent. The draft quotes those figures at L242-246 with no qualification and calls the source a "research-grade wristband" at L65, which it was not. DOWNGRADE: major to moderate — this is a strict subset of F160, which states the same problem more completely and adds the missing-ground-truth point. Merge into F160.
The problem
Specificity is only interpretable against a candidate rate, and "candidate event" is never defined — how candidates are generated, at what rate per hour, over what window. With about 3.5 true flashes per night (l.40), even a modest candidate rate makes false positives a substantial fraction of the analysis sample. False positives have, by construction, near-zero probability of a button press and no flash-related arousal, so they enter the models as guaranteed zeros on both outcomes, attenuating any true phase effect. Worse, a false positive is an autonomic transient misclassified as a flash, and autonomic transients during NREM are plausibly phase-structured for the same brain-autonomic reason the draft invokes throughout (l.103-105, l.107) — so the contaminant is not noise independent of the predictor but is correlated with it, and can generate or mask a phase effect on registration. The draft's own hedge undermines its feasibility claim: "with sleep-onset performance reported separately" concedes that the quoted operating point is not the sleep operating point, while the entire analysis is sleep-only. The source confirms that performance differs by state: Tsiartas2021.txt reports "System performance for hot flashes (HFs) onsets occurring during sleep vs wake" as a separate figure, and its body text gives the matched operating point as 96.5% specificity rather than the 95.6% abstract figure the draft quotes.
In the draft — line 241-245
Each candidate event is classified as a hot flash or not, a binary indicator. Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier, reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance, with sleep-onset performance reported separately.
Evidence
Draft l.241-245. Source: /root/grantreview/Tsiartas2021.txt line 24 ("sensor approach achieved above 90% sensitivity at 95.6%" / line 25 "specificity in HF classification") versus line 135 ("To ensure similar specificity (96.5±1%) across") and line 167 ("At 96.5% specificity, the SC+ set showed better HF"), and line 196 ("Fig. 4 System performance for hot flashes (HFs) onsets occurring during sleep vs wake"). Logical point: a specificity figure without a candidate rate does not yield an expected false-positive count.
Fix
State the candidate-generation rule and the expected false-positive count: "Candidate events are generated by [stated rule] at an observed rate of [X] per hour of NREM; at the sleep-specific operating point of the classifier ([sensitivity]/[specificity] for onsets during sleep, Tsiartas et al. 2021 Fig. 4), this implies approximately [N] false positives per participant-night against [M] true flashes. Because false positives cannot be registered and carry no flash-related arousal, and because autonomic transients may themselves be phase-structured, the primary analysis is repeated at a higher decision threshold trading sensitivity for specificity, and the positive predictive value is estimated in the sternal-reference validation subsample." Quote the sleep-specific operating point rather than the pooled abstract figure.
What the refuter said
Strongest case against: false positives are, by the finding's own construction, guaranteed zeros on both outcomes, which attenuates toward the null - conservative for a confirmatory test, not a validity threat. The draft also does not claim a PPV, and the quoted figures are accurate to the source. And "with sleep-onset performance reported separately" is a disclosure, not a concealed concession - Tsiartas2021.txt L196 confirms "Fig. 4 System performance for hot flashes (HFs) onsets occurring during sleep vs wake" exists as a separate result, and the body text (L232-235) adds "We observed a greater contribution from the non-SC features in the HF classification performance for HFs with onsets occurring during sleep vs wake". What survives is the core, verified by reading L241-245 in full: "candidate event" is never defined - no generation rule, no rate per hour, no window - so no expected false-positive count per night can be derived from a specificity figure, and the expected composition of the analysis sample is unknown. The genuinely worrying part is the second-order argument: an autonomic transient misclassified as a flash is plausibly phase-structured for exactly the brain-autonomic reason the draft invokes at L103-105, so the contaminant is correlated with the predictor rather than independent of it, and can generate or mask an effect. That is unaddressed. Moderate rather than major: the first-order effect is conservative, and the fix is one sentence defining the candidate stage.
The problem
The mediation logic is formally correct for a total effect and it is the wrong estimand. "Gating" means the same signal has different access depending on brain state. A total effect of phase on registration is satisfied by phase making the flash bigger, which is not gating at all - it is a stronger stimulus. The draft's Literature Review states the correct estimand and the Statistical Analyses section then abandons it: the review says the informative comparison is whether "comparable events have different probabilities", and comparability is exactly what magnitude adjustment or matching delivers. So the confirmatory analysis tests a quantity the proposal's own stated logic says is uninformative, and the analysis that would be informative is demoted to "a sensitivity analysis only". Worse, magnitude is known to differ between aroused and unaroused flashes, so the unadjusted contrast is guaranteed to be partly a magnitude contrast.
In the draft — line 307-309
Flash magnitude is a mediator, not a confounder: if ISF phase modulates the size of the autonomic response, adjusting for it would bias the total effect of phase on registration. The primary model therefore does not adjust for magnitude
Evidence
Draft L166-169: "The informative comparison for conscious gating is whether, among flashes that do occur, comparable events have different probabilities of cortical arousal or conscious registration depending on the phase at which they arise." Draft L309-311: "a magnitude-adjusted model is a sensitivity analysis only. The primary registration model is confirmatory; all else is exploratory." Direct internal contradiction. Magnitude differs by arousal status in the draft's own cited data - Gombert-Labedens et al. 2025 (/root/grantreview/GombertLabedens2025.txt, Figure 5 legend): "Hot flashes that occurred in undisturbed sleep were accompanied by a drop in systolic blood pressure and an increase in heart rate, which was of smaller magnitude than that for a hot flash associated with an arousal."
Fix
Make the magnitude-matched or magnitude-adjusted model the confirmatory one and report the unadjusted total effect alongside it as the descriptive quantity. Rewrite L307-311 as: "Two estimands are distinguished. The total effect of phase on registration (magnitude unadjusted) is reported descriptively. The confirmatory test of gating is the direct effect of phase on registration holding peripheral magnitude fixed, since only that quantity distinguishes differential access to a comparable signal from a differently sized signal. Both are pre-specified; the direct effect carries the confirmatory claim." Pre-specify the magnitude index (electrodermal rise amplitude and rate, temperature excursion, heart-rate increment) and note the caveat that skin-conductance amplitude does not track self-reported severity (Gombert-Labedens et al., 2025), so magnitude must be treated as a multivariate index, not one number.
What the refuter said
The claimed 'direct internal contradiction' does not survive reading the source sentence in context. L162-169 is contrasting the OCCURRENCE question with the gating question: 'The informative comparison for conscious gating is whether, among flashes that do occur, comparable events have different probabilities ...' — 'comparable events' there means flashes rather than non-events, which is what the whole paragraph is about. Reading it as 'magnitude-matched' is the finding's gloss, not the draft's word. The estimand choice is also defensible on its own terms, and the finding's preferred analysis is the more dangerous one: a 'direct effect' obtained by covariate adjustment for a mediator is biased by anything affecting both mediator and outcome, and sleep depth affects both flash magnitude and registration — so the draft's decision to keep the total effect confirmatory and the adjusted model as sensitivity (L309-311) is the conservative and defensible choice. The Figure 5 evidence the finding adduces (aroused flashes carried larger heart-rate and blood-pressure changes) is verified verbatim, but it is consistent with magnitude being a mediator, i.e. it supports the draft's model rather than undermining it. One real residual: the interpretive claim at L366-367 that the study would show 'an infraslow rhythm sets when an internal event becomes an experience' outruns what a total effect establishes — but the draft's arousal-versus-registration comparison (L54-56) is its stated route to that separation, not magnitude adjustment. Downgraded major to minor.
The problem
The 50 s figure is the mouse value. The paper the draft cites for the noradrenergic mechanism notes in its own discussion that the corresponding human intervals are ~20 s (fMRI, N3) and 20-40 s (cyclic alternating pattern arousability). The draft's own analysis band, 0.01-0.04 Hz, spans periods of 25-100 s and it explicitly expects a per-participant empirical peak — i.e. it concedes the period varies fourfold — while its power argument uses a single number. If the human period is 25-40 s, the quarter-cycle jitter budget falls from ~12.5 s to 6-10 s, and the cited detector's 15 s decision grid alone exceeds half a cycle at the fast end. The proposal's central feasibility claim is calibrated to the most favourable possible period. Separately, a 0.01-0.04 Hz passband is two octaves wide; a Hilbert instantaneous phase from a two-octave band is not well defined, and a zero-phase filter with a 0.01 Hz lower edge cannot be applied meaningfully to the 100-150 s bouts the draft proposes without edge effects dominating.
In the draft — line 315-318 (with 46, 266-270)
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle.
Evidence
OsorioForero2021.txt L398-406 (right column): "In human functional imaging studies of deep NREMS (stage N3), the LC and surrounding areas increase activity in intervals of ~20 s.24 Moreover, fast spindle activity appears clustered over tens of seconds, in particular during light NREMS (stage N2).27,32,33 These are timescales compatible with infraslow LC activity in human NREMS." And: "These events occur preferentially at the beginning and end of a NREMS period and lead to higher arousability every 20-40 s.54" The ~50 s figure is stated as the mouse thalamic measurement — L32-33: "Thalamic NA fluctuates over ~50 s". Internal inconsistency in the draft: L266-268 specifies "zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant)" — a band whose fastest admitted period is 25 s — against the fixed "50-s cycle" of L316.
Fix
Recompute the precision budget across the plausible human range and report the worst case. Rewrite as: "Onset jitter translates directly into phase uncertainty. Human infraslow periods are reported between roughly 20 and 50 s (Lecci et al., 2017; Osorio-Forero et al., 2021), so the tolerable jitter is 5-12 s depending on the participant's empirical peak. We therefore estimate each participant's period first and express jitter as a fraction of that period, excluding participants for whom achieved timing precision exceeds one quarter of their own period." Also narrow the passband (e.g. ±0.5 octave around each participant's empirical peak) so the Hilbert phase is well defined, and state the minimum segment length the filter requires.
What the refuter said
AGAINST — and this largely lands. The headline ("the 50 s figure is the mouse value") is wrong: Lecci et al. 2017 established the ~0.02 Hz sigma oscillation in HUMANS as well as mice (I verified the human core dataset, 27 men, all-night NREM sigma ISF), Lazar et al. 2019 report human infraslow sigma oscillations in the same range, and Dimitriades et al. 2026 describe the human ISFS as spindle clustering "over 10-100 s" in N=154. The two Osorio-Forero passages the finding quotes are verbatim correct (L398-406: "increase activity in intervals of ~20 s"; "higher arousability every 20-40 s") but they concern LC fMRI activity intervals and cyclic alternating pattern — not a competing estimate of the human sigma ISF period. The draft's 50 s for humans is properly sourced. WHAT SURVIVES, as a minor drafting point: the internal inconsistency is real and readable off the draft. L266-269 specifies "0.01-0.04 Hz, empirical peak reported per participant" (fastest admitted period 25 s) while L315-318 fixes the jitter budget against "a 50-s cycle" — so for a fast-ISF participant the quarter-cycle budget is 6 s, not 12.5 s. Partially pre-empted by L318-319 ("Simulations use the observed counts and precision"). The filter-edge sub-point is also legitimate — a 0.01 Hz lower edge against the 100-150 s stable bouts of L344-347 will be edge-dominated — but it is a clause inside a finding whose main claim is a mis-attribution, so minor here.
The problem
The draft presents 19.8% as a settled quantity, sourced from the smaller of two studies by overlapping authors, and cites the larger study elsewhere (line 253) without noting that it gives a substantially different figure. Baker et al. 2019 analysed 542 flashes from 86 women and found 28.6% occurred in undisturbed sleep against 51.1% with arousal or awakening; De Zambotti's n=34 gave roughly 20% and 70%. The draft's expected event yield for the quiet, unaroused flashes — the events it says an experimenter-delivered stimulus can never provide (line 53-54) — differs by nearly half depending on which figure is used, and the draft never says which it planned on.
In the draft — line 149-151
Separately, 19.8% of objectively detected nocturnal flashes occur without disturbing sleep at all (De Zambotti et al., 2014)
Evidence
Baker et al. (2019), Sleep 42(11):zsz175, abstract (https://academic.oup.com/sleep/article/42/11/zsz175/5549533): "HFs associated with arousals/awakenings (51.1%), were accompanied by an increase in systolic (SBP; ~6 mmHg) and diastolic (DBP; ~5 mmHg) BP and HR (~20% increase)... In contrast, HFs occurring in undisturbed sleep (28.6%)..." (n = 86 for HR, 542 HFs). Gombert-Labedens et al. 2025 (/root/grantreview/GombertLabedens2025.txt lines 924-929) states both side by side: the n=34 study with "only a minority (20%) occurring without disturbance to sleep" and "A subsequent study in a larger sample of 86 individuals also found that most of the detected hot flashes (51%) were associated with sleep disruption, and 29% of hot flashes occurred in undisturbed sleep".
Fix
Give the range and use the larger study for planning: 'Between roughly 20% and 29% of objectively detected nocturnal flashes occur without disturbing sleep (De Zambotti et al., 2014, n = 34; Baker et al., 2019, n = 86, 542 flashes).' Use the lower bound for the expected-yield calculation.
What the refuter said
STRONGEST DEFENCE, LARGELY SUCCESSFUL: the draft's 19.8% is faithful to the source it cites. GombertLabedens2025.txt independently corroborates the de Zambotti n=34 figure ("only a minority (20%) occurring without disturbance to sleep") and the Baker n=86 figure ("most of the detected hot flashes (51%) were associated with sleep disruption, and 29% of hot flashes occurred in undisturbed sleep"), and the review attributes the divergence to definitional and sampling differences, not to one study being wrong. Crucially, the draft quotes 19.8% alongside the 3.5/1.5 counts from the SAME study (L149-152) — using one source consistently for three linked numbers is good practice, not cherry-picking. And the finding's core consequence is invented: the draft contains NO event-yield calculation anywhere, so nothing in it "differs by nearly half depending on which figure is used". Both figures support the identical qualitative point the draft is making (a substantial minority of flashes pass without disturbance). WHY A KERNEL SURVIVES: the draft does cite Baker et al. 2019 at L253 for a different purpose and never notes that the larger study gives a materially different proportion. A reviewer who knows both will want one clause acknowledging it. DOWNGRADE: moderate to minor — a one-clause addition, no consequence for design or inference.
The problem
With only random intercepts, the sin and cos coefficients are fixed effects: the model assumes a single preferred phase and a single modulation depth common to all participants. Two things in the draft argue against that assumption. First, the ISF peak frequency is acknowledged to vary between women (l.267-268: "empirical peak reported per participant"), and phase is defined relative to that individually estimated rhythm, so both the estimate and any true preference are participant-specific. Second, onset-timing precision will vary between participants (skin conductance quality, device fit, artefact load), so the attenuation of any phase effect varies too. If preferred phases differ across women, averaging sin/cos over participants attenuates the group effect toward zero — a false negative — while the intercept-only random structure leaves standard errors on the fixed phase terms understated relative to a model that allows participant-level heterogeneity. Neither the heterogeneity nor the choice not to model it is discussed, and there is no stated plan for how many participants or events are needed to fit a random-slope model if one is warranted.
In the draft — line 288-291
Circular phase enters each model as sin(phase) and cos(phase) ... Random intercepts are included for participant and for night within participant
Evidence
Draft l.288-291 (sin/cos as fixed effects; random intercepts only) against l.267-268 ("empirical peak reported per participant"), which establishes that the phase reference is individually estimated and therefore individually variable.
Fix
Pre-specify the random-effects structure and a fallback: "The primary model includes participant-level random slopes for sin(phase) and cos(phase) in addition to random intercepts for participant and night, allowing the preferred phase and modulation depth to vary between women; the fixed-effect test of the joint sin/cos term is then a test of a consistent group-level preference. If the random-slope model fails to converge or is singular, the intercept-only model is used and the resulting assumption of a common preferred phase is stated as a limitation, with between-participant heterogeneity in preferred phase reported descriptively (per-participant circular mean and a second-level test)."
What the refuter said
AGAINST: the finding's inferential bridge is broken. Individual variation in ISF peak FREQUENCY (L267-268) does not imply variation in preferred PHASE — phase is defined relative to each woman's own rhythm, which normalises frequency differences away rather than propagating them. And a single population-level preferred phase is the natural first hypothesis for a conserved brain mechanism, so fixed sin/cos is the standard and most powerful test of it. Random slopes on sin/cos need many events per participant; with a handful of analysable flashes each, such a model would routinely fail to converge, making random intercepts the defensible choice rather than an oversight. The dominant bias the finding names — attenuation toward zero if preferred phases differ — is conservative, i.e. a false negative, which is not what sinks a proposal. WHAT SURVIVES, as a minor point: the description of the model is accurate (L288-291 gives random intercepts only), the omission is not discussed, and if there IS true slope heterogeneity the standard errors on the fixed phase terms are anticonservative for a test the draft designates confirmatory (L310-311). A sentence acknowledging the assumption and saying when a random-slope model would be attempted would close it. Minor.
The problem
At the achievable yield a night contributes well under one analysable flash, so the night-level variance is estimated from a small minority of nights and collapses to the boundary in about half of fits. A boundary variance component makes the reported model not the fitted model, and the 2-df LRT on the fixed sin/cos pair is measurably liberal at these counts. This is not fatal - it is a specification and reference-distribution issue with a clean fix - but the draft names the LRT as the confirmatory inference without qualification.
In the draft — line 287-291
Random intercepts are included for participant and for night within participant, with covariates for sleep stage, time since the preceding flash, and half of the night
Evidence
Arithmetic: 347 analysable flashes over 438 usable nights = 0.79 per night; under Poisson, 198 nights have 0 events, 157 have 1, and only 83 of 438 (19%) have the >=2 events needed to inform a night-level variance. Simulation of the draft's exact model, fitted by nested Gauss-Hermite quadrature with an analytic gradient verified against numerical differentiation to 1e-5 (/root/grantreview/sim/common.py, 04_clustering.py; 400 replicates, 137 participants, 4 nights, ~347 events, ~48 registered, sigma_participant=1.0, sigma_night=0.5, no phase effect): convergence 100%, SINGULAR FITS 52.3%. Type-I error at alpha=0.05: proposed GLMM with both random intercepts 0.0775 [95% CI 0.055, 0.108]; participant intercept only 0.070 [0.049, 0.099]; pooled logistic ignoring clustering 0.085 [0.061, 0.116]; GEE exchangeable, Wald 0.075 [0.053, 0.105]; permutation of phase within participant 0.0625 [0.043, 0.091]. With the same counts but a higher base rate (500 events, ~123 registered) all five methods were nominal (0.060, 0.060, 0.064, 0.068, 0.060), so the inflation is driven by the small number of registered events, not by the clustering per se.
Fix
Add to Statistical analyses: "Because a night contributes fewer than one analysable flash on average, the night-within-participant variance is not identifiable; only the participant random intercept is retained, and the night-level term is reported as a sensitivity analysis. Confirmatory inference uses the likelihood-ratio statistic of the joint sin and cos term referred to a permutation distribution obtained by shuffling ISF phase within participant, because in simulation at the expected counts the asymptotic chi-square reference is liberal (type-I 0.078, 95% CI 0.055-0.108) while the within-participant permutation reference is nominal (0.063, 0.043-0.091)."
What the refuter said
The simulation outputs check out: /root/grantreview/sim/out_04_A.json gives singular_rate 0.5225 and typeI_glmm2 0.0775 with conv_rate 1.0 over 400 replicates, matching the finding. The Poisson arithmetic (347/438 = 0.79 per night, 19% of nights with >=2 events) is correct. But the finding's own evidence undercuts its own headline, and a second file it does not mention undercuts it further. In out_04_A the pooled logistic that ignores clustering entirely is the MOST inflated method at 0.085, above the two-random-intercept GLMM at 0.0775 — so the inflation cannot be attributed to the night-level random intercept. The finding concedes this itself ('the inflation is driven by the small number of registered events, not by the clustering per se'), which contradicts its own title. And out_04_B.json, the phase-clustered scenario, gives typeI_glmm2 0.045 — nominal. With 400 replicates the 0.0775 estimate carries a 95% interval of roughly [0.055, 0.108], so 'measurably liberal' rests on a margin of half a percentage point. Finally the whole exercise conditions on an assumed yield of 347 analysable flashes and ~48 registered events, a number the draft never states (that is F65's point, not evidence for this one). What survives is one legitimate, verified technical note: at plausible densities the night-level variance will sit on the boundary in about half of fits, so the reported model will not be the fitted model. Downgraded moderate to minor.
The problem
The unit-of-analysis reasoning is correct for a within-cluster predictor, but it hides the fact that a participant whose analysable flashes are all unregistered contributes essentially nothing to a within-participant phase contrast. At the expected press rate the great majority of women fall in that category, so 'multiple nights' does not convert into information for the women who never press the button - a point the draft's own low-compliance contingency (line 347-349) half-recognises without quantifying.
In the draft — line 191-194
The unit of analysis is the flash, not the participant: power depends on event count, hence multiple nights.
Evidence
Simulation, 4,000 draws (/root/grantreview/sim/05 replaced by inline computation, reproduced in results_table.txt): central scenario N=347 flashes across 137 women at a 15% press rate - 38.3 women are discordant (some registered, some not) and therefore informative; 82.8 women have analysable flashes but never press; only 139 of 347 flashes (40%) sit inside an informative participant. Pessimistic scenario (N=71, 5% press rate): 1.6 informative women, 5% of flashes. Optimistic (N=1,249, 30%): 144 informative women, 93% of flashes. The design is therefore hypersensitive to the press-rate assumption, which the draft never states.
Fix
Add to Power and precision: "Because a participant with no registered flash carries no within-participant phase contrast, the informative sample is smaller than the flash count implies. At an assumed per-event registration rate of p, roughly 137*(1 - exp(-n_i*p) - exp(-n_i*(1-p))) participants are discordant; at n_i = 2.5 analysable flashes and p = 0.15 this is about 38 of 137 women and 40% of flashes. Power is reported against this informative subset as well as against the raw event count, and the assumed value of p is pre-registered with the pilot data supporting it."
What the refuter said
Strongest case against: the statistical premise is a conditional-likelihood intuition applied to a model that is not conditional. The draft specifies "a generalised linear mixed model with a logistic link" and "Random intercepts ... for participant and for night within participant" (L280, L287-289). In a random-intercept GLMM, clusters with all-zero outcomes are not dropped as they would be in conditional/fixed-effects logistic regression; they contribute through partial pooling. Their leverage on the phase slope is small, so the direction of the point is right, but "contributes essentially nothing to a within-participant phase contrast" overstates it — and the finding itself concedes "The unit-of-analysis reasoning is correct for a within-cluster predictor." Second, every number is assumption-driven: the 15%/5%/30% press rates are the reviewer's own, and the finding says so. The simulation output in /root/grantreview/sim/results_table.txt is internally consistent (central scenario 347 flashes, 52 registered, ~43/137 women contributing at least one registered flash), but it validates the reviewer's yield model, not a fact about the study. The genuine defect — no stated press rate, no stated event count — is F114's. Minor.
The problem
The concern is correctly identified and the causal structure is real: phase -> detection <- flash magnitude / arousal propensity -> registration, so conditioning on detection opens a path from phase to registration under the null. The draft's mitigation (re-detect without cardiac features) addresses one route but never states how large a phase-dependence in detection the design can tolerate, and there is no phase-independent reference against which to bound it, because the gold standard is itself an autonomic signal. Simulation shows the bias is tolerable at the draft's implicit N and intolerable at the N the design would need.
In the draft — line 249-252
It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent; a pre-specified analysis repeats detection from skin conductance and temperature alone
Evidence
Simulation (/root/grantreview/sim/10_collider.py; magnitude -> detection OR 3.0 per SD, magnitude -> registration OR 3.0 per SD, registration rate 0.15, no true phase effect). At N=347 detected: with 70% detection sensitivity, OR_detect(best:worst phase) of 1.5/3/5/10/20 gives rejection 0.054/0.059/0.069/0.086/0.117 and induced spurious OR 1.04/1.17/1.20/1.33/1.41; at 90% detection, rejection stays 0.047-0.068 throughout. At N=3,000 detected (the count needed for adequate power): 70% detection gives rejection 0.075/0.128/0.240/0.445 at OR_detect 1.5/3/5/10; 50% detection gives 0.077/0.193/0.443/0.652. The tolerance is therefore OR_detect <= about 1.5. Note the induced spurious OR of 1.2-1.4 is already the same order as the effect the study hopes to find.
Fix
Add: "Conditioning on detection opens a path from phase to registration whenever detection sensitivity is phase-dependent and flash magnitude predicts registration. Simulation of that structure shows type-I error stays below 0.075 only while the detection odds ratio between the best and worst ISF phase is under about 1.5; at an odds ratio of 5 and 70% sensitivity, type-I error at an adequately powered event count reaches 0.24. Phase-dependent detection sensitivity is therefore estimated directly - by re-running detection on phase-shuffled surrogates and, in the laboratory sub-sample, against simultaneous sternal skin conductance - and the primary analysis is conditioned on that odds ratio falling below a pre-registered bound of 1.5."
What the refuter said
Strongest case against: this finding largely exonerates the draft and then reports the exoneration as a defect. I ran down /root/grantreview/sim/10_collider.py and out_10.log and the simulation is competently built (correct collider DAG, calibrated intercepts, LRT on the joint sin/cos term) and its N=347 numbers reproduce - rejection 0.051 to 0.117 across OR_detect 1.0 to 20 at 70% detection, i.e. type-I inflation from mild to modest. So at the design's own scale the mitigation question is not urgent, which is what the finding concedes. The claim that the induced spurious OR of 1.2-1.4 is "the same order as the effect the study hopes to find" is speculative, since the draft nowhere states a target effect size - the finding is comparing a computed bias against an unstated quantity. The N=3,000 figures I could not reproduce from the log I have. What is true and airtight is narrow: the draft never states how much phase-dependence in detection the design can tolerate, and there is no phase-independent reference to bound it against, because the gold standard is itself autonomic. That point is made better, with primary-source evidence rather than simulation, by F193. PLAUSIBLE at minor.
Every link from a thermoregulatory event in the body to a number in a model.
The problem
Four separable problems, all in the source file. (1) Temporal resolution: the classifier "makes a decision every 15 s"; features are computed over windows of ±30 s (SC), ±250-500 s (temperature) and ±120 s (PPG); labels were matched to expert annotations within a "±90 s matching window"; and the SC feature itself is defined as "the ±2 minutes around the HF predicted onset." Against a 50 s ISF cycle, a ±90 s matching tolerance is ±648° of phase. (2) The >90%/95.6% figures were obtained under that tolerance, so they do not certify onset timing at all, and they certainly do not transfer to the draft's re-derived change-point onset, which is a different quantity from a different pipeline. (3) N=3 women, 27 flashes, one lab night each. (4) The sensors are a custom research array (Grove SC, SEN-11574 PPG, NXP accelerometer, TI temperature) wired into a Compumedics PSG — not the EmbracePlus the study will use. Calling this a "research-grade wristband" elides that.
In the draft — line 242-245
Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier, reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance, with sleep-onset performance reported separately.
Evidence
Tsiartas et al. 2021 (source file /root/grantreview/Tsiartas2021.txt): "Three women (Age, mean ± SD: 55.6 ± 0.6 y)... A total of 27 physiological HFs were recorded"; "a Decision Tree classifier, which makes a decision every 15 s"; "We time-aligned the features with the HF expert annotations for prediction and evaluation (±90 s matching window)"; "We designate the HF onset as the ±2 minutes around the HF predicted onset"; sensor list "PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700...T sensor: TI-TMP36GT9Z" "integrated with the Compumedics recording system." Also note the gold-standard event rule is itself 30-s grained: Gombert-Labedens et al. 2025 (source file line ~896): "a rapid rise of at least 2 microSiemens in sternal skin conductance within a 30-second period."
Fix
Report Tsiartas's actual temporal specification and stop using its accuracy figures as evidence about timing. Replace the sentence with: "Detection follows the multi-sensor approach of Tsiartas et al. (2021), which reached >90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance — but under a ±90 s event-matching tolerance, in three women, on a custom sensor array rather than the EmbracePlus. Neither the classifier's 15 s decision grid nor that tolerance is adequate against a 50 s cycle; onset timing is therefore re-derived independently by change-point detection on the electrodermal rise, and its precision is validated as a primary feasibility outcome before any phase analysis." Add a pre-registered validation of the re-derived onset against simultaneous sternal skin conductance in a subsample; without that, the precision claim rests on nothing. Also consider Naghavi et al. 2025 (Psychophysiology, https://doi.org/10.1111/psyp.70056), a more recent flash detection/prediction pipeline.
What the refuter said
All four prongs verified verbatim in Tsiartas2021.txt, which I read in full: 'Three women (Age, mean ± SD: 55.6 ± 0.6 y)' and 'A total of 27 physiological HFs were recorded' (L29-31); 'a Decision Tree classifier, which makes a decision every 15 s' (L86-89); '(±90 s matching window)' (L82-84); 'We designate the HF onset as the ±2 minutes around the HF predicted onset' (L120-121); the custom sensor list wired into Compumedics (L59-74). This finding is more careful than F47 on the key point and I credit it: it correctly says the >90%/95.6% figures 'do not certify onset timing at all, and they certainly do not transfer to the draft's re-derived change-point onset'. STRONGEST DEFENCES, which force the downgrade from fatal. The draft quotes the performance figures accurately and volunteers the sleep/wake split ('with sleep-onset performance reported separately'), so there is no misreporting of results. The '±648 degrees of phase' framing is rhetoric, since the draft does not use Tsiartas's onset. And the pipeline is not speculative: the host protocol I fetched already prescribes EmbracePlus 'objective hot flash occurrence via multisensory approach (skin conductance, temperature, heart rate, movement)' across the cohort. What survives, and is decision-relevant: the measurement that defines the entire exposure set is validated on three postmenopausal women, 27 events, one laboratory night, on different hardware in a temperature-controlled bedroom, and none of that is disclosed anywhere in the draft.
The problem
The causal reasoning is correct and the decision not to adjust is right. But the argument presumes that the measured magnitude sits on the causal path from phase to registration — i.e. that a bigger autonomic response is a stronger interoceptive signal and hence more likely reported. The draft's own cited review states that skin-conductance amplitude does not represent self-reported flash severity. If magnitude as measured does not track the perceptual quantity, the sensitivity analysis is weaker evidence than the draft implies, and the mediation framing should not be leaned on.
In the draft — line 306-308
Flash magnitude is a mediator, not a confounder: if ISF phase modulates the size of the autonomic response, adjusting for it would bias the total effect of phase on registration.
Evidence
Gombert-Labedens et al. 2025 (source file, lines ~888-890): "Also, the amplitude of the skin conductance does not represent the self-reported severity of a hot flash [193]." Supporting the general direction, however: Baker et al. 2019 (as described in the same review, lines 985-995) found that flashes associated with arousal/awakening were accompanied by larger heart-rate and blood-pressure increases than flashes in undisturbed sleep (86 individuals, 542 flashes).
Fix
Keep the decision, weaken the claim. Rewrite: "Flash magnitude is treated as a possible mediator rather than a confounder, so the primary model does not adjust for it; a magnitude-adjusted model is a sensitivity analysis only. Because electrodermal amplitude does not track self-reported flash severity (Gombert-Labedens et al., 2025), magnitude is an imperfect proxy for interoceptive signal strength, and the sensitivity analysis is interpreted accordingly."
What the refuter said
The quote is verified verbatim in GombertLabedens2025.txt L888-890: 'Also, the amplitude of the skin conductance does not represent the self-reported severity of a hot flash [193].' The finding is also honest in conceding that the draft's causal reasoning is correct and the decision not to adjust is right. But its premise about the draft is an assumption, not a reading: L306-308 says 'the size of the autonomic response', never skin-conductance amplitude, and the draft's own cited Baker et al. 2019 evidence points the other way — I verified the review's Figure 5 legend, which reports that flashes associated with arousal/awakening carried larger heart-rate and blood-pressure increases than flashes in undisturbed sleep (86 individuals, 542 flashes). That is direct support for magnitude sitting on the path from event to arousal. And the review's statement is about a subjective severity RATING, which is a different quantity from the probability of noticing an event at all. What survives is a real underspecification the finding stumbles onto without naming it: 'flash magnitude' is never operationally defined anywhere in the draft, so a reviewer cannot tell whether the mediator is measurable at all. Minor, as filed.
The problem
Four separate breaks between what Tsiartas et al. establish and what the draft needs. (1) Sample: N=3 women, mean age 55.6 ± 0.6, who "had undergone natural menopause" — postmenopausal, not the draft's perimenopausal 40-55 population — yielding 27 flashes in a single ~12 h lab session. (2) Device: the signals came from a custom array of loose consumer components (Grove 101020052 skin-conductance sensor, TI TMP36GT9Z temperature sensor, SEN-11574 PPG, NXP FXOS8700 accelerometer) taped to the wrist and wired into a Compumedics PSG amplifier — not an EmbracePlus, not any integrated wristband. The draft's "multi-sensor features of a research-grade wristband (Tsiartas et al., 2021)" describes a device that paper did not use. Sampling rates, electrode geometry, contact pressure and EDA units all differ. (3) Timescale: the classifier makes a binary decision every 15 s from features computed over ±30 s, ±120 s (PPG), ±250 s (SC AUC) and ±500 s (temperature) windows, and was scored against expert annotations using a "±90 s matching window". A method whose own evaluation tolerance is ±90 s cannot time onsets against a 50 s cycle. The paper reports no onset-accuracy metric of any kind. (4) No ground truth here: performance was measured against sternal skin conductance scored by experts. The host trial collects no sternal SC, so the 90%/95.6% figures cannot be reproduced, checked, or recalibrated in this sample. Detection is therefore an unvalidated derived quantity presented as a measurement, and false positives dilute any phase effect toward null.
In the draft — line 242-246 (also 64-66)
Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier, reaching over 90% sensitivity at 95.6% specificity against expert-scored sternal skin conductance, with sleep-onset performance reported separately.
Evidence
/root/grantreview/Tsiartas2021.txt: "Three women (Age, mean ± SD: 55.6 ± 0.6 y) who reported having daily HFs participated in a ~12h lab-based study"; "A total of 27 physiological HFs were recorded"; "had undergone natural menopause, and none of them was currently taking hormone therapy"; "Signals from a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700 - 1024 Hz; T sensor: TI-TMP36GT9Z – 16 Hz) were collected from each participant's wrist… and integrated with the Compumedics recording system"; "We time-aligned the features with the HF expert annotations for prediction and evaluation (±90 s matching window)"; "we randomly split our data in two sets, 80% of data for training the Decision Tree and 20% for evaluating the system" (frame-level random splitting across only three participants — not subject-independent). Independent corroboration that this class of method is not ready: Gombert-Labedens et al. 2025, /root/grantreview/GombertLabedens2025.txt lines 899-901: "however, these techniques are mostly in the development phase and have not been applied to track hot flashes in clinical trials."
Fix
Do not present detection as settled. Add a detector-validation work package as a stated deliverable with its own resource: (a) report that the published performance derives from N=3 with a non-equivalent sensor array and a ±90 s tolerance; (b) either bring in a sternal-SC reference sub-study in a small validation sample on EmbracePlus, or state that detection is validated only against manually reviewed multi-sensor traces with double-scoring and reported agreement; (c) report the detector's onset-timing error distribution against that reference — a number the whole proposal depends on — before claiming feasibility. Replace "a research-grade wristband (Tsiartas et al., 2021)" with "features analogous to those of Tsiartas et al. (2021), re-derived for the EmbracePlus and validated in this population".
What the refuter said
STRONGEST DEFENCE: sub-point (3) is refuted — the draft does not use the classifier to time onsets; L259-261 back-dates onset by "change-point detection to the inflection at which the electrodermal rise begins", a separate procedure, and Tsiartas's own reference [4] (Forouzanfar et al., "Automatic detection of hot flash occurrence and TIMING from skin conductance activity") shows onset estimation is an established, distinct capability in this lineage. The four modalities also transfer exactly: the parent protocol commits EmbracePlus to "multisensory nocturnal hot flash estimation (skin conductance, temperature, heart rate, movement)", and the classic 2 µS/30 s skin-conductance rule remains available as a fallback — so degradation, not impossibility. And the draft does pre-specify a skin-conductance-plus-temperature re-detection (L251-254) against phase-dependent detector bias, which is more than most proposals do. WHY THE REST SURVIVES: sub-points (1), (2) and (4) are all verbatim-verified in Tsiartas2021.txt — "Three women (Age, mean ± SD: 55.6 ± 0.6 y) ... a ~12h lab-based study"; "A total of 27 physiological HFs"; "had undergone natural menopause"; the custom Grove/TMP36/SEN-11574/FXOS8700 array "integrated with the Compumedics recording system"; "we randomly split our data in two sets, 80% ... 20%" (frame-level, not subject-independent, across three participants); and oversampling "to 1:2 ratio". The paper's own conclusion says "initial feasibility". Sub-point (4) is the most consequential and I confirmed it against the protocol: no sternal skin conductance is collected, so the 90%/95.6% figures can never be reproduced, checked or recalibrated in this sample, and the draft calls Tsiartas a "research-grade wristband" (L65) when it was a bench array wired to a PSG amplifier. Detection is a derived quantity presented as a measurement, and false positives dilute the phase effect toward null. DOWNGRADE: fatal to major.
The problem
The draft correctly identifies the binding constraint and then declines to address it. A quarter of a 50 s cycle is 12.5 s: the draft is stating that the study fails unless total onset jitter is under roughly 12 s. Total jitter is the sum of at least three error sources, none quantified anywhere in the document: (i) detector onset error — the source paper's own evaluation tolerance is ±90 s, 7× the budget; (ii) device synchronisation residual — the draft promises "a fixed offset and linear drift estimated per recording; tap markers validate this and residual error is reported" but pre-specifies no acceptable maximum, and a 5 s residual is already a 36° phase error; (iii) the physiological lag between central onset and the peripheral electrodermal inflection, which the draft proposes to *measure* rather than bound ("Intervals between this inflection and the accompanying temperature and pulse-rate changes index achievable precision"). Worse, the cited simulation is not shown, not parameterised, and not attributed — "simulation shows" with no figure, no assumed jitter distribution, no event count and no result. And "Simulations use the observed counts and precision" defers the entire power analysis until after the data exist, which is the one thing a power section cannot do. A reviewer is being asked to fund a study whose feasibility the applicants agree is unknown.
In the draft — line 315-320
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle. Simulations use the observed counts and precision, and the minimum detectable effect is reported.
Evidence
Draft line 317-318 sets the tolerance ("a quarter of the cycle" of a "50-s cycle" = 12.5 s). /root/grantreview/Tsiartas2021.txt: "(±90 s matching window)" and features computed over windows of ±250 s (SC AUC) and ±500 s (temperature) — no onset-accuracy figure is reported in the paper. Draft line 229-231 promises sync residual reporting with no threshold. Draft line 260-263 makes the physiological lag an outcome rather than a bounded input. Logical contradiction: the power section conditions on "the observed counts and precision", i.e. the study's feasibility can only be established after the money is spent.
Fix
Present the error budget as a table before the power section: stated tolerance (12.5 s), and for each of detector onset error, EEG–wristband synchronisation residual, and central-to-peripheral lag, a numeric estimate with its source and the resulting phase error in degrees. Then run and report the simulation now, with explicit assumptions: e.g. "With 300 analysable flashes, a von Mises concentration of κ=0.4 and Gaussian onset jitter of SD 8 s against a 50 s cycle, the joint sin/cos likelihood-ratio test has 80% power; at SD 15 s power falls to 41% and no achievable event count recovers it." State the minimum detectable effect as a number. If the error budget cannot be closed to under a quarter cycle from existing evidence, make a pilot precision study the first work package and gate the confirmatory analysis on its result — and say that plainly, because a reviewer will trust a proposal that names its own kill criterion far more than one that defers it.
What the refuter said
NOTE: revised_severity below should read major, not the schema label — see reasoning. Verified textually at every point. L317-318 sets the draft's own tolerance ('a quarter of the cycle' of a '50-s cycle' = 12.5 s). Tsiartas2021.txt reports no onset-accuracy figure at all, only the ±90 s evaluation window and windowed features. L228-230 promises 'residual error is reported' with no acceptable maximum pre-specified. L260-263 makes the physiological lag an outcome rather than a bounded input. The two sharpest prongs are pure textual observations and are airtight: L317-319's 'simulation shows ...' is asserted with no parameters, no effect size, no jitter distribution, no event count, no figure and no attribution; and L319-320's 'Simulations use the observed counts and precision' defers the entire power analysis until after the money is spent. STRONGEST DEFENCES, which force the downgrade from fatal. The draft names the risk itself rather than hiding it, pre-registers a threshold and a fallback (L338-343), and its confirmatory 2-df sin/cos test is invariant to systematic lag, so the failure mode is lost power, not a false result. This is an unquantified feasibility risk, not a demonstrated impossibility. Heavy overlap with F65 and F47 — the parent should merge these three into one objection rather than three.
The problem
A zero-phase (non-causal) band-pass at 0.01-0.04 Hz has an impulse response spanning tens of seconds in BOTH directions. The phase assigned to a flash onset is therefore partly determined by EEG recorded AFTER that onset. The two outcomes - cortical arousal and a button press - are both events that raise broadband EEG power (the arousal criterion is literally a transient in the band immediately adjacent to sigma), and a button press requires the participant to be awake for some period after onset. So registered/aroused flashes systematically have post-onset EEG that differs from unregistered ones, and that difference leaks backwards into the estimated pre-onset phase. The design therefore manufactures a phase-registration association with no gating whatever. This is not a small bias: against a 50 s cycle, a few seconds of leakage is tens of degrees of phase. Combined with the fact that the sigma upper bound (16 Hz) abuts the arousal-defining band (>16 Hz), spectral leakage from the arousal itself into the sigma envelope is a second route to the same artefact.
In the draft — line 269-270
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase.
Evidence
Draft L236-237 defines the outcome as "transient > 16 Hz activity lasting 3-15 s (Wassing et al., 2019)" while L23 defines the predictor band as "sigma-band (11-16 Hz)" - adjacent bands. The predictor is described as zero-phase filtered (L269), i.e. non-causal by construction. Logical contradiction: a predictor whose value depends on data recorded after the outcome cannot test whether the predictor causes the outcome.
Fix
Estimate ISF phase causally: use a forward-only (minimum-phase or one-sided) filter, or fit the ISF on a pre-onset window only (e.g. the 100-150 s ending at onset) and extrapolate phase to onset, with the extrapolation error quantified. Pre-specify that all data from onset onward are excluded from the phase estimate. Then demonstrate on shuffled-outcome surrogates that the causal estimator produces no phase-registration association. Add a hard negative control: recompute the primary model with the phase estimate taken from a matched epoch one full ISF cycle earlier; a true gating effect must vanish there, an artefact will not. Also move the sigma band away from the arousal band (e.g. 11-15 Hz) or notch the arousal transient before enveloping.
What the refuter said
Strongest case against: zero-phase filtering plus Hilbert on the sigma envelope is standard practice in this exact literature (Lecci, Osorio-Forero, Carro-Dominguez, Dimitriades all do it), and the draft does restrict to "stable (and non-artifactual) NREM" (L270), which would exclude the post-flash wake segment itself; a 3-15 s arousal transient also has little spectral energy inside a 0.01-0.04 Hz passband, so leakage from the arousal alone would be small. But the objection stands, because this is not the descriptive case those papers ran: here the analysis is event-locked and the outcome perturbs the very signal used to build the predictor. The specification is verified — L269 is explicitly "zero-phase", hence non-causal, and the predictor band (11-16 Hz, L23) abuts the arousal-defining band (>16 Hz, L236-237). The magnitude is worse than the finding claims: the reviewers' filter simulation shows a filtfilt Butterworth at this band has 99% of its impulse-response energy spanning 72-136 s, i.e. +/-36 to +/-68 s about each sample — more than a full ISF cycle of *future* data entering the phase estimate at onset. A registered flash implies being awake for a period afterwards, and wake sigma-band power differs greatly from NREM, so this is not a small perturbation. Nowhere does the draft say the estimate uses pre-onset data only, and for a confirmatory pre-registered predictor that is a real gap. Not fatal, because the honest fix is available but costly: the same simulation shows a pre-onset-only window leaves ~40 degrees circular SD (lambda ~ 0.78) at the terminal sample.
The problem
Both human sources for the phenomenon measured it away from the frontal pole. Lazar et al. found the infraslow modulation most prominent over centro-parieto-occipital sites; Lecci et al.'s human recordings were C3 and C4. The ZMax records F7-Fpz and F8-Fpz — the site where the cited evidence is weakest. The draft acknowledges this only as a 'feasibility outcome', and its contingency is circular: 'weak frontal ISF expression triggers the same rule' (line 343), where the rule is to switch the predictor from infraslow phase to pre-flash infraslow *amplitude*. If the frontal infraslow oscillation is weakly expressed, its amplitude is equally unreliable — the fallback is derived from the same failed signal. There is no fallback that recovers a posteriorly expressed rhythm from a frontal-only device.
In the draft — line 273-276
the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019). Frontal expression and achievable phase precision are themselves feasibility outcomes.
Evidence
Lazar, Dijk & Lazar (2019), J Neurosci Methods 316:22-34, abstract (https://research-portal.uea.ac.uk/en/publications/infraslow-oscillations-in-human-sleep-spindle-activity): "These ISO are most prominent in the high sigma band and over the centro-parieto-occipital regions." "ISO in sleep spindles are most prominent in the centro-parieto-occipital regions, left hemisphere and second half of the night". Lecci et al. 2017 (PMC5298853): "Polysomnographic recordings included EEG from C3 and C4 electrode sites". Esfahani et al. (bioRxiv 2023.08.18.553744) on the device: "two frontal EEG channels (F7-Fpz, F8-Fpz), a tri-axial accelerometer, and a PPG sensor".
Fix
State the topography problem openly and give a real contingency: 'The human infraslow sigma rhythm has been characterised centrally (C3/C4; Lecci et al., 2017) and is most prominent centro-parieto-occipitally (Lazar et al., 2019); this device records frontally. We therefore treat frontal expression as a gating feasibility criterion, quantified in the first N participants against phase-shuffled surrogates, and we will add a central derivation (or a device offering one) if frontal amplitude falls below a pre-registered threshold.' Delete the claim that switching to amplitude rescues weak frontal expression.
What the refuter said
Strongest case against: the draft does not claim frontal ISF is established - it says the modulation "is established in humans" (true) and immediately makes frontal expression a feasibility outcome (L275-276). Attenuation is not absence. Both defences fail to dispose of the finding. I verified the topography from primary sources. Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853/): human core montage was "C3 and C4 electrode sites", and "the 0.02-Hz oscillations showed a maximum over parietal derivations for power in both the sigma and the FSP band and declined toward anterior central and frontal areas". Lazar, Dijk & Lazar 2019 (https://pmc.ncbi.nlm.nih.gov/articles/PMC6390176/): "Frontopolar region had significantly (adjusted P < 0.05) lower integrated EIP compared to all other brain regions except the temporal brain region", with the explicit ranking parietal > occipital > central > frontal > temporal > frontopolar. The ZMax montage is F7-Fpz / F8-Fpz (Esfahani, verified) - frontopolar-referenced, spanning the two weakest regions in Lazar's ranking. Citing these two papers as feasibility support without disclosing the gradient is a real reviewer-facing problem. The circularity of the fallback is also sound: if frontal ISF expression is weak, its amplitude is derived from the same weak signal. Major, not fatal: the ISO is measurable frontopolarly (Lazar computed integrated EIP there), and four nights of NREM give ~thousands of ISF cycles to average over.
The problem
A zero-phase (non-causal, e.g. filtfilt) band-pass at 0.01-0.04 Hz has an effective support of tens of seconds on BOTH sides of any sample. The phase estimated "at onset" therefore depends on sigma-band power AFTER the flash. The draft's own account of the outcome is that arousals occur in, and are marked by, sigma/spindle loss: "low sigma power marks a fragility phase" (l.99) and "in humans, fragility is marked by absent sleep spindles and spontaneous arousals" (l.29-30). A flash followed by an arousal thus produces a post-onset sigma drop, which the non-causal filter propagates backwards into the estimated onset phase, pulling it toward fragility. The result is a mechanical, artefactual association between "fragility phase at onset" and "arousal/registration" that would appear with zero true phase gating. The draft's entire circularity discussion (l.246-254) is about the flash detector and never touches this, which is the far larger circularity. It is compounded by band adjacency: the predictor is sigma 11-16 Hz power (l.96) while the arousal outcome is scored as "transient > 16 Hz activity lasting 3-15 s" (l.237) from the same two frontal derivations, and by movement/EMG artefact at arousal, which also perturbs the sigma envelope.
In the draft — line 267-269
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase.
Evidence
Draft l.267-269 (zero-phase filtering) against draft l.29-30 ("in humans, fragility is marked by absent sleep spindles and spontaneous arousals") and l.99 ("low sigma power marks a fragility phase"). Logical contradiction: the predictor is estimated from a time window that contains the outcome, and the outcome is defined by a change in the same signal that defines the predictor. Draft l.237 places the arousal scoring band (>16 Hz) immediately adjacent to the sigma band (11-16 Hz, l.96) on the same frontal derivations (l.216-217).
Fix
State that ISF phase at onset is estimated causally, from pre-onset data only: e.g. fit the ISF using a window ending at onset minus a stated guard interval, using a forward-only (causal) filter or a windowed sinusoid/complex-demodulation fit to the pre-onset sigma envelope, and report the phase-estimation error this introduces. Replace the sentence with: "Instantaneous ISF phase at flash onset is estimated from the pre-onset log sigma envelope only, using a causal estimator applied to the window [onset - 3 cycles, onset], so that no post-onset EEG — including any flash-related arousal, spindle suppression or movement artefact — can influence the phase estimate; the bias and variance of this causal estimator relative to the non-causal estimate are quantified on flash-free stable NREM segments and reported." Additionally, pre-specify that epochs containing the flash-related arousal are excluded from sigma-envelope estimation, and separate the arousal-scoring band from the predictor band (or score arousals with a rater blind to the sigma envelope).
What the refuter said
AGAINST, three ways. (1) Zero-phase filtering is near-universal in this literature, so this could be standard practice mistaken for an error. (2) The draft restricts to "stable (and non-artifactual) NREM" (L270) over bouts of three ISF cycles (L344), which may exclude arousal-terminated bouts and thus the contaminating cases. (3) Magnitude is entirely unquantified — the finding offers no simulation of the induced bias. None of these rescues it. On (1): the field's own causal-manipulation work explicitly avoids non-causal estimation — Osorio-Forero et al. detected phase ONLINE ("We detected these phases online through a machine-learning algorithm and triggered optogenetic activation based on whether sigma power started to rise or decline"), i.e. causally, precisely because the outcome follows the phase estimate. So a causal alternative is the field norm where it matters. On (2): the escape is a trap — if stable-bout selection removes arousal-terminated bouts, the supporting arousal outcome loses its variance, and if it truncates them, the 0.01 Hz lower edge makes the retained segment edge-dominated. Either branch is a defect. The logical contradiction is airtight and un-pre-empted: L267-269 estimates the predictor with a filter whose support extends tens of seconds past onset, while L29-30 and L99 define the outcome as loss of the very sigma power that defines the predictor, and L236-237 scores arousal as ">16 Hz activity" immediately adjacent to the 11-16 Hz predictor band on the same two frontal derivations with no EMG or EOG. The draft's circularity paragraph (L246-254) addresses only the detector. Major rather than fatal: one sentence in the pre-registration (causal filter, or phase from a strictly pre-onset window) fixes it — but as written the confirmatory analysis would manufacture the predicted effect.
The problem
Three distinct problems. (1) It is not the same hypothesis. The hypothesis under test is that WHEN within an endogenous cycle an interoceptive event arrives determines whether it reaches awareness — a claim about temporal alignment, and the reason the design is interesting is that phase is orthogonal to mean sleep depth within a bout. "Pre-flash ISF amplitude" is a state-level quantity and is also ambiguous between two different constructs the draft never distinguishes: the depth of infraslow modulation over the preceding window (analogous to the noradrenaline oscillation amplitude of Kjaerby et al., 2022, cited at l.111-113) versus the instantaneous sigma level before onset. The first tests whether the rhythm's strength matters; the second tests whether sigma power level matters. Neither tests timing. Both are heavily confounded with N2 versus N3, spindle density, homeostatic pressure and time of night, and neither can distinguish "the rhythm times access" from "deeper sleep is less permeable" — the very confound the phase design exists to escape. (2) The stated reason it "tolerates far more timing error" is precisely that it discards the temporal information the hypothesis concerns; tolerance to jitter is a symptom of testing a different, coarser question, not a virtue. (3) The switching rule is not pre-specified in any operational sense: the threshold value for marker dispersion is absent, the criterion for "weak frontal ISF expression" is absent, the "pre-specified floor" for reporting compliance (l.348-349) is absent, and the relaxation from three ISF cycles to two (l.345-347) has no trigger value. A pre-registration whose decision thresholds are not stated in the proposal cannot be evaluated and gives full post hoc latitude. Computing the decision statistic before inspecting outcomes (l.338-340) is good practice and does not by itself break pre-registration — the problem is that the numbers are missing.
In the draft — line 340-347
if it exceeds a pre-registered threshold, the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis; weak frontal ISF expression triggers the same rule.
Evidence
Draft l.340-343 ("pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis") against the aims at l.177-179 ("whether the phase of the infraslow sigma fluctuation (ISF) at nocturnal hot flash onset predicts conscious registration") and l.79-80 ("Because the cycle is short, the binding constraint is not event count but the precision with which onset can be timed against it"). Missing values: l.339-340 "a pre-registered threshold" (no number), l.347 "weak frontal ISF expression" (no criterion), l.348-349 "a pre-specified floor" (no number), l.345-347 relaxation to two cycles (no trigger).
Fix
Either state the numbers or drop the claim of equivalence. Suggested rewrite: "Insufficient onset-timing precision is the principal analytic risk. Before outcomes are inspected we compute the dispersion of onset estimates across peripheral markers; if the standard deviation exceeds [X] s (one eighth of the mean ISF period), the phase analysis is reported as underpowered and the pre-registered fallback predictor is pre-flash ISF amplitude, defined as [state which: the peak-to-trough modulation depth of the log sigma envelope over the [N] s preceding onset]. This tests a related but distinct hypothesis — that the strength of the infraslow rhythm, rather than the phase within it, governs registration — and is labelled as such; it does not substitute for the phase test. Weak frontal ISF expression is defined as an infraslow modulation index below [Y] in more than [Z]% of stable NREM bouts. The stable-bout criterion is relaxed from three to two ISF cycles if fewer than [N] analysable events survive. Participants contributing fewer than [K] button presses across [M] nights, or whose morning diary counts exceed their press counts by more than [R]-fold, are excluded from the registration model."
What the refuter said
AGAINST: on limb (1) a defender can say the switch is honest risk management, not a bait-and-switch — the draft flags the direction question openly at L358-361 and the fallback is pre-specified rather than post hoc, which is better discipline than most proposals show. And the finding fairly credits the draft for computing the decision statistic before inspecting outcomes (L338-340). But both substantive limbs hold. On (1): "pre-flash ISF amplitude ... testing the same hypothesis" (L341-343) is false against the draft's own aim at L177-179 ("whether the phase of the infraslow sigma fluctuation (ISF) at nocturnal hot flash onset predicts conscious registration") and its own rationale at L79-80 ("the binding constraint is not event count but the precision with which onset can be timed against it"). Amplitude is a state quantity; the hypothesis is about timing within a cycle. The ambiguity the finding identifies is real too — "ISF amplitude" could mean Hilbert modulation depth (Kjaerby-style rhythm strength, cited at L111-113) or pre-onset sigma level, and the draft never says which; under either reading it is confounded with N2/N3, spindle density and time of night, which is the confound the phase design exists to escape. On (3), verified by reading: L339-340 "a pre-registered threshold" (no value), L343 "weak frontal ISF expression" (no criterion), L344-345 "Low event yield" (no trigger), L347-348 "a pre-specified floor" (no number). Four decision rules with zero operational content in a document that calls itself confirmatory. Major.
The problem
This is the inclusion criterion that determines the entire analysis sample, and it is defined nowhere in Data processing or Statistical analyses. The only quantitative hint appears in the contingency section — "relaxing the stable-bout criterion from three ISF cycles to two" (l.345-347) — which tells a reviewer that the criterion involves a bout of some length without saying whether the bout is measured before onset, after onset, or symmetrically around it. That choice is decisive. If stability must hold after onset, then any flash that produces an arousal, a stage shift or an awakening terminates its own bout and is excluded — the sample is selected against the outcome, and the study systematically deletes the events that carry the effect, biasing the registration and arousal models toward null while leaving a residual sample enriched for quiet flashes. If stability is assessed before onset only, this is avoided, but the phase estimate then has no post-onset support (which is the correct choice, see the non-causal filtering finding) and the draft's non-causal estimator is inconsistent with it. Two further undefined items sit in the same cascade: "estimation window" (l.205-206) is named as an exclusion and never defined, and "artefact-free recording" (l.206) has no stated rejection criterion or rater.
In the draft — line 270
Analyses are restricted to stable (and non-artifactual) NREM.
Evidence
Draft l.270 (criterion asserted, undefined) and l.205-206 ("stable NREM, estimation window, artefact-free recording, reliable onset timing" listed as exclusions with no definitions), against l.345-347 ("relaxing the stable-bout criterion from three ISF cycles to two"), the only place a quantity appears. Logical consequence: if the bout requirement extends past onset, inclusion depends on the outcome.
Fix
Define the criterion in Data processing and make it pre-onset: "A flash enters the analysis if the [N] s preceding its estimated onset fall entirely within contiguous N2/N3 epochs free of scored arousals, wake or artefact, where N corresponds to three cycles of the participant's empirical ISF peak. Stability is assessed on pre-onset data only, so that a flash-related arousal, stage shift or awakening after onset never affects inclusion. Artefact rejection uses [stated automatic criterion plus visual confirmation]. The estimation window for the sigma envelope and ISF phase is the pre-onset interval defined above." State the number of events lost at this step (see the yield-table finding).
What the refuter said
Verified by grep: 'stable' appears at L77, L205, L270 and L344 and is defined at none of them; the only quantity is L344-345 'relaxing the stable-bout criterion from three ISF cycles to two', in the contingency section. 'estimation window' appears exactly once (L206) and is never defined; 'artefact' appears at L206 and L270 with no rejection criterion and no rater. So the textual claim is airtight. STRONGEST DEFENCE: 'stable NREM' is a term of art in this literature — the team's own Current Biology paper says 'we detected all infraslow sigma power cycles taking place during NREMS (excluding transitional periods)' (OsorioForero2021.txt L272-274) — so a competent analyst knows roughly what is meant. That defence does not survive, because the finding's fork is decisive and the draft's own methods resolve it in the damaging direction: L267-268 specifies 'zero-phase band-pass filtering', which is non-causal and therefore requires data on both sides of onset. Post-onset support is thus required, which means an arousal, stage shift or awakening terminates its own bout and the event is excluded — inclusion depends on the outcome, and the residual sample is enriched for quiet flashes in both the registration and the arousal model. That is a genuine selection problem that would surface only when the data arrive. Major, as filed.
The problem
A hot flash is defined by a tonic conductance rise (2 uS/30 s), and the draft's onset definition is the inflection of that tonic rise. But wrist tonic EDA during sleep shows an inverted stage pattern attributed to sweat accumulation under the device in deep sleep. That confound generates slow conductance rises indistinguishable in shape from a flash, and it is stage-dependent — therefore correlated with sigma power, therefore potentially correlated with the predictor itself. The draft's only validity anchor for wrist EDA is a paper on a different, hard-wired sensor in a lab.
In the draft — line 260-262
Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing.
Evidence
Parry & Briganti 2026, medRxiv, on the Wearanize+ dataset the draft cites: tonic EDA showed an "inverted stage pattern" due to "wrist sweat accumulation during deep sleep, representing a known confound for wrist-worn EDA during sleep" and is not recommended; only phasic EDA "plausible patterns and may be used with caution" (https://www.medrxiv.org/content/10.64898/2026.06.10.26355348v2). Wrist-vs-palm validity: van der Mee et al. 2021, Int J Psychophysiol 168:52-64 — within-subject wrist-palm r = 0.31 for SCL and 0.42 for ns.SCR frequency, recommended "at least for epidemiology-sized ambulatory studies" (https://research.vu.nl/en/publications/validity-of-electrodermal-activity-based-measures-of-sympathetic-/).
Fix
Add a sternal skin-conductance channel on the EEG nights and define both detection and onset from it, using the wrist only for adjunct features. If sternal recording is refused, the draft must (a) cite and confront the tonic-EDA sleep confound, (b) report the shape-discrimination performance of the change-point detector against non-flash sweat-accumulation ramps, and (c) restrict onset definition to phasic components with the resulting loss of timing precision stated.
What the refuter said
Strongest case against: "fatal" does not survive, because the cited paper measures a different quantity. Parry & Briganti concerns *tonic* level by stage — a slowly varying baseline whose inverted N3 pattern they attribute to hydration; I verified the wording, including "Tonic EDA during sleep therefore reflects skin hydration rather than autonomic sympathetic activity" and "should not be used as a sympathetic arousal proxy in sleep studies". A hot flash is a rapid rise of at least 2 uS within 30 s (verified in GombertLabedens2025 and Tsiartas), i.e. an event whose detection depends on the derivative, not the baseline — and Tsiartas detected exactly such events from *wrist* skin conductance during sleep at above 90% sensitivity, with sleep-versus-wake performance reported separately. So the draft does have a wrist-EDA-in-sleep anchor; the finding's "only validity anchor for wrist EDA is a paper on a different, hard-wired sensor in a lab" describes Tsiartas accurately but wrongly implies it is not wrist EDA in sleep. The stage-dependence-implies-correlation-with-sigma path is also speculative and would act on detection and onset timing, not on the predictor itself. What survives: a hydration-driven slow rise can mimic a flash's shape, the confound is stage-dependent, and the draft cites nothing for its change-point method. Moderate.
The problem
Ambient temperature is the best-established environmental modulator of nocturnal flashes and it is absent from the draft entirely — no logging, no covariate, no mention in Risks. Freedman and Roehrs, whom the draft cites twice for the REM result, showed in the same paper that 18 degrees C cut flash count from 2.2 to 1.5 and that the effect was confined to the first half of the night. Ambient temperature also raises tonic skin conductance directly, so it moves the detector's own input and its false-positive rate, and it interacts with the wrist sweat-accumulation confound. Recording spread across seasons in Amsterdam homes, with bedding, sleepwear and a possible bed partner unrecorded, leaves a large uncontrolled between- and within-participant source of variance in both exposure supply and detection.
In the draft — line 188-189
The design is within-subject, observational and multi-night, conducted at home.
Evidence
Freedman & Roehrs 2006, Menopause 13(4):576-583 (https://pubmed.ncbi.nlm.nih.gov/16837879/): three conditions at 30, 23 and 18 degrees C; "the 18 degrees C ambient temperature significantly reduced the number of hot flashes, from 2.2 +/- 0.4 to 1.5 +/- 0.4", an effect present only in the first half of the night. The parent protocol also logs no room temperature (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full).
Fix
Add a bedside temperature and humidity logger for every recorded night (cheap, no participant burden) and enter mean and within-night bedroom temperature as covariates in both models, plus a temperature-by-half-of-night term. Record bedding, sleepwear, window/heating state and bed-sharing in the morning diary. Add ambient temperature to the Risks section as a modifier of both flash rate and detector specificity.
What the refuter said
Strongest case against: ambient temperature cannot confound the primary predictor. The ISF cycle is ~50 s; room temperature is effectively constant over that window, so it cannot covary with phase within a flash and cannot bias the phase-registration estimate. The primary analysis is also within-subject and within-night with random intercepts for participant and night (L285-289), which absorbs between-night thermal variation. That reduces this from a threat to validity to a threat to yield and generalisability, and takes major off the table. The factual claims are all verified. Freedman & Roehrs 2006 (https://pubmed.ncbi.nlm.nih.gov/16837879/): "Ambient conditions varied across nights 2-4: 30C, 23C, and 18C in randomized order" and 18C "significantly reduced the number of hot flashes, from 2.2 +/- 0.4 to 1.5 +/- 0.4", with temperature effects confined to the first half. And I confirmed the parent protocol logs no ambient temperature. So the draft cites this paper twice for its REM result while never mentioning the other half of its title, and records at home across seasons with no thermal logging, no covariate, and no mention in Risks - while ambient temperature also shifts tonic skin conductance, the detector's own primary input. A real, cheap, missed control. Moderate.
The problem
Instantaneous phase from a narrowband analytic signal is only interpretable when the signal is near-sinusoidal and the band is narrow relative to its centre. Neither holds here. The team's own Current Biology paper reports that the noradrenaline-sigma relationship produces "a non-symmetrical U-shaped time course", and their operational definition of infraslow phase was a binary rising-versus-declining classification from an online detector plus explicit cycle detection, not a continuous Hilbert phase. The draft's band, 0.01-0.04 Hz, has a bandwidth of 0.03 Hz against a centre of 0.025 Hz — a fractional bandwidth of 1.2, which is not narrowband in any sense, so the Bedrosian condition underlying the Hilbert construction is grossly violated. The band is also mis-centred: Lecci found the human peak at 0.019 Hz and Lazar found it below 0.02 Hz, both in the lower third of the draft's band, and filter centre misalignment is a documented driver of rapidly increasing phase error. "IAF phase" is also the wrong acronym for the study's central quantity.
In the draft — line 269
the Hilbert transform to calculate the IAF phase
Evidence
Osorio-Forero et al. 2021 (/root/grantreview/OsorioForero2021.txt): "Across animals, NA had already declined when sigma levels started rising, producing a non-symmetrical U-shaped time course"; "We detected these phases online through a machine-learning algorithm and triggered optogenetic activation based on whether sigma power started to rise or decline"; "we detected all infraslow sigma power cycles taking place during NREMS (excluding transitional periods)". Wodeyar et al. 2023, eNeuro 10(11):ENEURO.0507-22.2023 (https://www.eneuro.org/content/10/11/ENEURO.0507-22.2023): "the Hilbert transform should only be applied when the spectral support of the amplitude time series does not overlap with the spectral support of the phase time series"; "Brain rhythms can consist of nonsinusoidal waveforms, exhibit multiple nearby peak frequencies, and persist for short periods"; misaligned filter centre frequencies cause phase error to "rapidly increase"; "when the sinusoid is absent there is a greater cross-method difference in phase estimates". Lecci human peak "0.019 ± 0.001 Hz"; Lazar: ISO "with a frequency below the previously reported 0.02 Hz".
Fix
Use the team's own validated operationalisation as primary: detect infraslow sigma cycles on the unfiltered log envelope and classify each onset as falling on the rising or the declining limb (or into quartiles of the detected cycle). Report a per-event phase uncertainty measure and exclude events whose phase is not identifiable, following Wodeyar et al.; narrow the band to something centred on the empirical peak (e.g. 0.010-0.030 Hz) and justify it. Fix "IAF" to "ISF". If continuous phase is retained, show cross-method agreement between Hilbert and cycle detection as a validity check.
What the refuter said
This is one of the two best findings in the batch and it strengthened under checking. Team's own data verified in OsorioForero2021.txt: 'Across animals, NA had already declined when sigma levels started rising, producing a non-symmetrical U-shaped time course' (L277-279); 'We detected these phases online through a machine-learning algorithm and triggered optogenetic activation based on whether sigma power started to rise or decline' (L269-272); 'we detected all infraslow sigma power cycles taking place during NREMS (excluding transitional periods)' (L272-274). So the team's operational definition of infraslow phase was binary rising/declining plus cycle detection, not continuous Hilbert phase. Wodeyar et al. 2023 verified by fetching eNeuro: 'the Hilbert transform ... should only be applied when the spectral support ... of the amplitude time series does not overlap with the spectral support of the phase time series', 'Brain rhythms can consist of nonsinusoidal waveforms', and 'As the central frequency of the filter deviates from the correct central frequency, phase differences ... rapidly increase'. And I verified the mis-centring independently, which the finding only asserted: Lecci et al. 2017 gives the human peak as '0.019 ± 0.001 Hz ... corresponding to a cycle length of 52.6 ± 2.6 s', and Lazar et al. 2019 puts the fast-sigma ISO peak 'around 10 mHz' — both at or below the lower third of the draft's 0.01-0.04 Hz band, whose fractional bandwidth is indeed 1.2. STRONGEST DEFENCE: L268 says 'empirical peak reported per participant', so the applicants know the peak varies. But reporting the peak is not centring the filter on it, and the band is stated as fixed. Moderate, as filed.
The problem
The press is treated as if it timestamps awareness, but it aggregates awareness, waking enough to act, locating the button, motor capacity, sleep inertia and expectancy. The draft reports no latency distribution between subjective awareness and press, no false-press or missed-press rate, and no precedent for the measure. The closest published precedent — self-initiated nocturnal button pressing to mark wakefulness — achieved only poor epoch-level agreement with actigraphy. Reactivity is unaddressed: instructing women to monitor for and report flashes plausibly increases interoceptive attention and lowers the arousal threshold, altering the outcome. The design contains an unexploited control for exactly this — seven wearable nights against four EEG nights, so three nights have flash detection without the button-press instruction.
In the draft — line 216-219
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration.
Evidence
Keller et al. 2020, PLOS ONE 15(6):e0234060 (https://journals.plos.org/plosone/article?id=10.1371%2Fjournal.pone.0234060): self-initiated pressing showed poor epoch-by-epoch agreement with actigraphy, kappa = 0.23, versus moderate agreement for vibration-prompted pressing, kappa = 0.46; compliance self-reported as "usually" or "always" on 88% of nights, with 12% reporting "seldom"; no press-latency data reported. The parent protocol does not include a button press at all, using only "objective occurrence of nocturnal hot flashes ... estimated from a multisensory approach" (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full).
Fix
Add three things. (1) A within-participant reactivity test: compare detected flash rate, flash magnitude and (where available) arousal rate on button-press versus non-button-press nights using the three wearable-only nights. (2) A press-latency and compliance characterisation: an evening and morning calibration in which participants press on cue, plus a per-night count of presses with no nearby detected flash (false presses) and detected flashes with no press. (3) A complementary awareness measure not gated on volitional motor action — prompted awakenings with immediate report on a subset of events — and pre-register how the two are combined. Report the press-to-detection interval distribution as a feasibility outcome.
What the refuter said
Strongest case against: the Keller precedent is weak evidence and the arousal conjunction is already conceded. I verified Keller (kappa = 0.23 self-initiated versus 0.46 vibration-prompted, 88% compliance, no latency data), but that kappa is epoch-by-epoch agreement with *actigraphy*, itself a poor reference for wake, so a low value indicts both measures and says little about whether a press marks a felt flash. The draft also states the conjunction the finding raises, in nearly the same words — L299-300 "Because a button press requires sufficient arousal, registered flashes are expected to be largely a subset of aroused ones" — and then models registration among aroused flashes (L301-302), which is the decomposition that isolates access beyond arousal; compliance has a pre-specified floor with participants dropped from the registration model but retained for the arousal model (L346-349). What survives and is genuinely unaddressed: reactivity. Instructing women to monitor for and report flashes plausibly raises interoceptive attention and lowers the arousal threshold, altering the outcome, and nothing in the draft addresses it — while the design contains the obvious control (three wristband-only nights). And I confirmed the parent protocol has no button press and no event marker, so the press is this project's addition with no latency, false-press or missed-press analysis proposed. Moderate.
The problem
Two defects. (1) It does not test the same hypothesis. Lecci's fragility is a phase (declining sigma), not a level; and the ISF's functional signature in Lecci et al. is its spectral STRENGTH across the whole night (correlated with memory), a between-subject quantity, not a within-event predictor. Swapping instantaneous phase for pre-flash amplitude converts a test of "when in the cycle" into a test of "how much sigma just before" — which is a sleep-depth/spindle-density proxy already partly captured by the stage covariate. (2) The second trigger is self-defeating: if frontal ISF expression is too weak to give reliable phase, it is also too weak to give a meaningful amplitude, because both are read off the same frontopolar sigma envelope. The fallback inherits the failure it is meant to escape.
In the draft — line 340-343
if it exceeds a pre-registered threshold, the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis; weak frontal ISF expression triggers the same rule.
Evidence
Lecci et al. 2017: continuity and fragility are defined by "declining and rising sigma power levels", and the memory result is "Recall correlated with the individual peak of 0.02-Hz oscillations in the fast spindle band during all-night non-REM sleep (r = 0.45, P = 0.027)" — a per-subject spectral strength, not a per-event amplitude. Logical contradiction: DRAFT.md line 343 makes weak frontal ISF expression trigger a fallback that is computed from the same weak frontal ISF.
Fix
Replace the fallback with one that does not depend on the failing signal. Use the peripheral route: "If onset-timing precision or frontal ISF expression falls below the pre-registered threshold, the primary predictor switches to infraslow phase estimated from very-low-frequency heart-rate variability, which is phase-locked to the sigma ISF in both mice and humans (Jacobsen et al., 2026) and is recorded independently by PPG on both devices." If a level-based fallback is kept anyway, state plainly that it tests a different hypothesis (sleep-depth dependence, not phase gating) and demote it to secondary.
What the refuter said
Verified independently by fetching Lecci et al. 2017, and the finding is right on the point that matters: fragility corresponds to the 'declining portions of the 0.02-Hz oscillation in spindle activity' and offline/continuity to the rising portions — a phase, not a level; and the memory result is 'Recall correlated with the individual peak of 0.02-Hz oscillations in the fast spindle band during all-night non-REM sleep (r = 0.45, P = 0.027; n = 24)', a per-subject spectral strength, not a per-event amplitude. So L340-343's 'while testing the same hypothesis' is not established. The deeper problem the finding exposes is that 'pre-flash ISF amplitude' is never defined: the amplitude of the ISF oscillation (how strongly the rhythm is expressed) and the pre-flash sigma level (where in the cycle you are) are different constructs with different meanings, and the draft does not say which it means. The second prong is a fair logical observation — L343 makes weak frontal ISF expression trigger a fallback computed from the same weak frontal ISF envelope. STRONGEST DEFENCE, and it is why I downgrade: the draft elsewhere defines fragility as a level, 'flashes arriving in fragility (low sigma power)' at L355-356, under which a pre-flash-level fallback is internally coherent even if it misreads Lecci. And this is a contingency, not the confirmatory analysis. Downgraded major to moderate.
The problem
Sleep stage is entered as a covariate on the outcome, but the problem is upstream: in N3 the ISF is largely absent, so instantaneous phase in N3 is not a measurement of a weaker version of the same thing — it is filter output on a signal that isn't there. Lazar et al. found ISO power in SWS reduced relative to N2 independent of spindle density (F=108.7, p<0.0001), and Lecci et al. reported the human alternation as prominent in stage 2 rather than deep sleep. Both human ISF characterisation papers restrict to N2 bouts. Including N3 flashes with a stage covariate adds noise phase estimates that will attenuate any true effect and cannot be corrected by a covariate.
In the draft — line 287-291
Random intercepts are included for participant and for night within participant, with covariates for sleep stage, time since the preceding flash, and half of the night
Evidence
Lazar, Dijk & Lazar 2019 (https://pmc.ncbi.nlm.nih.gov/articles/PMC6390176/): "The ISO power in SWS was smaller compared to NREM stage 2 sleep independent of the incidence rate of sleep spindles"; "SWS has significantly decreased integrated EIP compared to stage NREM2 independent of sleep spindle density"; main effect of sleep stage F=108.7, p<0.0001. — Lecci et al. 2017: human continuity/fragility alternation was prominent during light (stage 2) sleep rather than deep sleep. — Dimitriades et al. 2024 and 2025 both restrict bouts to N2.
Fix
Make N2 the primary analysis stage and pre-register N3 as a separate, secondary stratum with its own ISF-presence gate. Rewrite the analysis restriction from "restricted to stable NREM" to "restricted to stable N2 bouts of at least 280 s in which ISF presence exceeds the surrogate threshold; N3 flashes are analysed separately and only where ISF presence is established, since infraslow sigma power is markedly reduced in slow-wave sleep (Lazar et al., 2019)." This interacts with the yield problem — say so in the yield table.
What the refuter said
Both sources verified independently. Lazar, Dijk & Lazar 2019 (PMC6390176): main effect of sleep stage on integrated excess infra power F(1,95.1) = 108.7, p < 0.0001, with the reduction in SWS confirmed 'independent of sleep spindle density' after controlling for detected spindle counts. Lecci et al. 2017: 'The 0.02-Hz oscillation of sigma power appeared to be more prominent in S2 than in SWS.' Both human ISF characterisation papers analyse N2 as the primary stage. The draft stages N1-N3 (L234), restricts only to 'stable ... NREM' (L270), and enters sleep stage as a covariate on the outcome (L288-289). The methodological point is sound and not a nitpick: this is measurement error in the predictor — in N3 the filter output is phase on a signal that is largely absent — and a covariate on the outcome cannot repair it. STRONGEST DEFENCES, which cap the severity. The bias is toward the null, so it costs power rather than manufacturing a false positive, and the finding says so. And the undefined 'stable NREM' (F71) might already be intended to mean N2 — in which case this is a documentation gap rather than a design error, which is itself an argument for fixing F71. Moderate, as filed.
The problem
Everything downstream — the phase assignment, and hence the primary test — depends on the wrist EDA inflection being a faithful, low-latency marker of flash onset. A 2026 analysis of the Wearanize+ dataset, with concurrent PSG, reports that tonic Empatica wrist EDA during sleep shows an inverted stage pattern (highest in N3) and concludes it "reflects skin hydration rather than autonomic sympathetic activity" and "should not be used as a sympathetic arousal proxy in sleep studies"; phasic EDA "may be used with caution" but did not correlate with micro-arousal frequency (r=−0.015, p=0.904). Wrist PPG-derived HRV in the same analysis was unusable (only "45.1% of epochs survived" plausibility filtering), which matters because pulse-rate is one of the four detector features and because the draft proposes cross-correlating heart-rate series for device synchronisation. Skin temperature was the one robust channel — and temperature is precisely the slowest-responding feature, with a ±500 s window in the detector. The caveat is real and I state it: that analysis used the Empatica E4, not the EmbracePlus, and concerned tonic level rather than change-point detection of a rising phasic event, so it does not prove the draft's method fails. But it shifts the burden: the draft asserts the electrodermal inflection as an established anchor and cites nothing for it.
In the draft — line 260-262
Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing.
Evidence
Parry & Briganti, "Validity and Limitations of the Empatica E4 Wristband for Autonomic and Thermoregulatory Sleep Monitoring Against Concurrent Polysomnography: A Wearanize+ Dataset Study", medRxiv 2026, https://www.medrxiv.org/content/10.64898/2026.06.10.26355348v1.full: tonic EDA "highest during N3 (2.93 μS) rather than lowest, opposite of expected physiology"; "sweat accumulates under the wristband during periods of reduced movement… increasing apparent skin conductance through a hydration mechanism rather than sympathetic arousal"; "Tonic EDA during sleep therefore reflects skin hydration rather than autonomic sympathetic activity"; "should not be used as a sympathetic arousal proxy in sleep studies"; phasic EDA vs micro-arousals "Spearman r=−0.015, p=0.904"; BVP-derived HRV "45.1% of epochs survived" and Bland-Altman limits "−2,557 to +3,679 ms"; skin temperature "the most robust Empatica feature". Draft lines 260-262 cite no source for the change-point method or its precision. Draft line 229-230 relies on "cross-correlating accelerometer and PPG signals" for synchronisation.
Fix
Cite this evidence and answer it. Add: "Wrist tonic EDA during sleep has been reported to track skin hydration rather than sympathetic tone (Parry & Briganti, 2026, E4); our anchor is therefore not tonic level but the change point of a phasic rise, whose validity and latency against sternal skin conductance we establish in a validation subsample before any outcome analysis." State the fallback if the EDA inflection proves unusable — most plausibly a temperature- or pulse-based anchor, whose slower dynamics must then be entered into the error budget explicitly rather than assumed away. For synchronisation, prefer the accelerometer cross-correlation and the tap markers over the heart-rate series given the PPG epoch-survival figure, and pre-specify a maximum acceptable residual (in seconds) beyond which a night is excluded.
What the refuter said
Strongest case against: the caveats the finding itself states do real work — E4 rather than EmbracePlus, tonic level rather than change-point detection of a rising event — and the HRV limb is weaker than presented. Cross-correlating a heart-rate *series* to estimate a fixed offset plus linear drift is far less demanding than beat-to-beat HRV validity, so 45.1% epoch survival under RMSSD/MeanNN plausibility filters does not show the synchronisation will fail, and the draft additionally has accelerometry and five-tap markers with residual error reported (L219-221, L227-230). What makes this finding stand where F108 does not is calibration: it concedes it "does not prove the draft's method fails" and correctly identifies the burden as the issue. All quotes are verified at source: "Tonic EDA during sleep therefore reflects skin hydration rather than autonomic sympathetic activity"; "should not be used as a sympathetic arousal proxy in sleep studies"; phasic EDA versus micro-arousals r = -0.015, p = 0.904; skin temperature "the most robust Empatica feature". And the burden point is right: L260-262 asserts the electrodermal inflection as the onset anchor and cites nothing at all for the change-point method or its precision, while the same Wearanize+ dataset the draft cites for synchronisation (Sikder et al. 2026) is where the tonic-EDA warning comes from. Moderate.
The problem
In the source literature the fragility phase is *defined* by the occurrence of spontaneous arousals, absent spindles and transitions to lighter sleep. A flash that is consciously registered involves waking enough to execute a long-press on the wrist, generating both a wake/arousal epoch and a large motion artefact within seconds to a minute of onset. If the stable-bout requirement (three ISF cycles, i.e. ~150 s; relaxed to two, ~100 s) must be satisfied around the onset, then registered flashes and aroused flashes are systematically excluded, and what remains is disproportionately the quiet flashes for which the outcome is constant. The draft never states whether the stable-bout window is centred on onset, precedes it, or follows it — and that unstated choice determines whether the primary outcome has any variance at all. The exposure-defining rhythm cannot be estimated on segments that exclude the very events being predicted.
In the draft — line 270 (with 76-77, 191-192, 344-347)
Analyses are restricted to stable (and non-artifactual) NREM.
Evidence
The source itself excludes exactly these periods when characterising the rhythm: OsorioForero2021.txt L272-273: "we detected all infraslow sigma power cycles taking place during NREMS (excluding transitional periods)". That exclusion is appropriate for describing the oscillation and destructive for predicting arousal from it. In humans the fragility phase is characterised by the very events being excluded — cf. GombertLabedens2025.txt L917-923: "perimenopausal individuals (n = 34) had an average of 3.5 hot flashes per night and that 70% of hot flashes were associated with an arousal from sleep" — i.e. 70% of the candidate events carry the arousal that the stability criterion penalises.
Fix
Define the stable-bout window explicitly as *pre-onset only* (e.g. ≥2 ISF cycles of artefact-free NREM immediately preceding onset, with no requirement on the post-onset epoch), state that arousals and awakenings following onset are outcomes and never exclusion criteria, and pre-specify that artefact rejection is applied only to the pre-onset estimation window. Add one sentence: "Because arousal and awakening are outcomes, no post-onset epoch is used for inclusion; stability and artefact criteria apply exclusively to the pre-onset phase-estimation window." Then report, per exclusion step, how many registered flashes survive.
What the refuter said
Strongest case against: the central premise is wrong. The finding asserts "In the source literature the fragility phase is *defined* by the occurrence of spontaneous arousals, absent spindles and transitions to lighter sleep." Lecci et al. 2017 (PMC5298853) defines it by the sigma oscillation itself: "offline periods correspond to raising, whereas fragility periods correspond to declining portions of the 0.02-Hz oscillation" and "wake-ups and sleep-throughs occur during declining and rising sigma power levels, respectively." Fragility is a phase of the ISF cycle inside consolidated NREM; restricting to stable NREM therefore does not delete it. Nor does the Osorio-Forero quote support the finding: "excluding transitional periods" means NREM-to-REM/wake transitions, not the fragility half-cycle. And a microarousal, per the draft's own criterion (L236-237, "transient > 16 Hz activity lasting 3-15 s"), does not exit NREM staging. So "deletes the fragility phase" is refuted. What survives is the finding's own best sentence: "The draft never states whether the stable-bout window is centred on onset, precedes it, or follows it." That is true and material — the reviewers' own yield model shows a centred, arousal-free +/-75 s requirement leaving ~5 registered events in the central scenario versus ~52 for a pre-onset window. Same underspecification that F180 exploits from the opposite side. Not fatal; moderate.
The problem
The draft names the failure mode and then leaves it unmitigated. If phase does not predict registration, the team cannot distinguish (a) no gating, (b) onset timing too imprecise, (c) frontal ISF phase from a two-channel headband too noisy, (d) cross-device synchronisation error, (e) the fallback amplitude predictor testing the wrong thing. Every one of these produces the same null. "Frontal expression and achievable phase precision are themselves feasibility outcomes" (L275-276) reports the inputs but does not validate the pipeline end to end, and the interval-dispersion check in Risks (L338-341) measures peripheral marker agreement, not whether the EEG-derived phase means anything. So a funded three-year project has one publishable outcome and one uninterpretable one. This is fixable cheaply, and the fixes are already reproducible in these exact recordings.
In the draft — line 79-80
Because the cycle is short, the binding constraint is not event count but the precision with which onset can be timed against it.
Evidence
Draft Risks section L338-343 mitigates only onset-timing dispersion; the draft contains no analysis that validates the phase estimate against a known phase-dependent phenomenon, and no equivalence bound (L318-320 reports "the minimum detectable effect" but no pre-specified inferiority/equivalence margin). Two positive controls are available in the same data. (i) Human sigma-heart-rate infraslow coupling: Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853) - "In humans, heart rate alterations also correlated with sigma power, but with a clear time lag", with heart rate rising at sigma minima and leading the next sigma peak by about 5 s. Both devices record PPG (draft L214-217), so this coupling can be recovered from ZMax EEG against EmbracePlus PPG. (ii) Human microarousal phase distribution: Dimitriades et al. 2026 (doi 10.1038/s41598-026-58423-z) - "electrophysiological markers of arousal and memory reactivation are organized within the spindle-rich ISFS peak"; preprint: "they clustered around the peak of the ISFS in all age groups".
Fix
Add a Positive controls subsection with three pre-specified checks, all run before outcome data are inspected. (1) Reproduce the sigma-to-heart-rate infraslow coupling within-participant from ZMax sigma and EmbracePlus PPG, and require the phase lag to fall within the range reported by Lecci et al. 2017. This simultaneously validates the frontal ISF phase estimate, the cross-device synchronisation and the achievable phase resolution - the three things a null would otherwise be blamed on. (2) Reproduce the phase distribution of spontaneous microarousals against the ISF, and require concentration near the peak as in Dimitriades et al. 2026. This validates the phase estimate against an arousal outcome with no dependence on hot flashes at all. (3) Pre-register an equivalence bound: state the smallest phase modulation the team would regard as mechanistically meaningful (e.g. an odds ratio of registration between opposing half-cycles) and pre-commit to reporting equivalence when the confidence interval falls inside it, conditional on controls (1) and (2) passing. State explicitly that a null is only interpretable if the controls pass, and that if they fail the study reports as a feasibility study.
What the refuter said
Strongest case against: the draft has more validation than the finding allows. The supporting arousal model is itself a partial positive control - it asks whether ISF phase predicts cortical arousal, which is a replication of the established human arousability finding in this montage; a positive result there validates the phase pipeline and rescues the interpretation of a null on registration. And L275-276 commits to reporting frontal expression and achievable phase precision, so two of the five null-producing mechanisms the finding lists are directly measured. That is a real mitigation the finding does not credit. It nonetheless survives, because both-null remains uninterpretable and a genuine in-sample validation against a known phase-dependent phenomenon is missing and nearly free. Both proposed controls are verified real. Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853/): human "heart rate declined rapidly once sigma power had reached a peak and increased gradually during sigma power minima... before subsequent sigma peaks by ~5 s" - recoverable here from ZMax EEG against EmbracePlus PPG, both of which the draft records (L213-217). Dimitriades et al. 2026 (https://www.nature.com/articles/s41598-026-58423-z): "electrophysiological markers of arousal and memory reactivation are organized within the spindle-rich ISFS peak" - the phase distribution of *all* spontaneous arousals is available in-sample and is not proposed. Worth noting for the applicants: that same result places arousal markers at the sigma peak, in tension with the fragility prediction, though the draft's two-sided 2-df test and L357-361 pre-empt the direction. Moderate.
The problem
The draft states this as an open possibility. It is a documented human fact for the very signal the detector weights: heart rate is coupled to the infraslow sigma rhythm in humans, rising at sigma minima and leading the next sigma peak. Since the detector's PPG feature is a heart-rate differential across the event, its sensitivity varies systematically with ISF phase, and the detected event set is a phase-biased sample of the true event set. The proposed remedy - re-run detection from skin conductance and temperature alone - does not fix it, because the infraslow rhythm is explicitly a brain-autonomic, sympathetically coupled rhythm, and electrodermal activity is a sudomotor sympathetic output. The right move is to measure the bias rather than to argue it away by dropping one channel.
In the draft — line 249-252
It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent; a pre-specified analysis repeats detection from skin conductance and temperature alone
Evidence
Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853): "In humans, heart rate alterations also correlated with sigma power, but with a clear time lag", heart rate rising during sigma minima and leading the next sigma peak by about 5 s; the paper's own framing is of a coordinated neural-and-cardiac rhythm. The draft acknowledges the coupling in its own words at L103-105: "The oscillation is coordinated with an infraslow cardiac rhythm, indexing a brain-autonomic rather than a purely cortical state" - which makes the phase-dependence of an autonomically driven detector a prediction, not a possibility. Tsiartas et al. 2021 confirms PPG is a substantial contributor and more so during sleep (/root/grantreview/Tsiartas2021.txt: "We observed a greater contribution from the non-SC features in the HF classification performance for HFs with onsets occurring during sleep vs wake").
Fix
Rewrite as an expected bias with a measurement plan: "Because heart rate is itself coupled to the infraslow sigma rhythm in humans (Lecci et al., 2017), detector sensitivity is expected to vary with ISF phase. We quantify this directly rather than assuming it away." Then add the quantification: in the in-lab validation sub-study (see onset-precision-exceeds-detector-resolution), estimate detection sensitivity as a function of ISF phase against sternal skin conductance, and carry that sensitivity function into the primary model as an offset or inverse-probability weight. Keep the skin-conductance-only re-run as a secondary check, but state that it does not eliminate the problem because electrodermal activity is itself a sympathetic output of the same rhythm.
What the refuter said
Strongest case against: the draft identifies the problem itself ("phase-dependent sensitivity cannot be assumed absent", L249-252) and pre-specifies a mitigation, and re-running detection without cardiac features does test the largest single route - Tsiartas gives PPG a substantial Shapley contribution, so dropping it is not a token gesture. Nor is the draft's hedge dishonest; it is merely weaker than the evidence warrants. The finding survives that defence on primary sources. Lecci et al. 2017, verified: in humans "heart rate declined rapidly once sigma power had reached a peak and increased gradually during sigma power minima... before subsequent sigma peaks by ~5 s" - so cardiac phase-dependence is documented, not hypothetical, and the draft asserts the coupling in its own voice at L103-105 ("coordinated with an infraslow cardiac rhythm, indexing a brain-autonomic rather than a purely cortical state"). And Tsiartas2021.txt L232-235, verified verbatim: "We observed a greater contribution from the non-SC features in the HF classification performance for HFs with onsets occurring during sleep vs wake" - the reliance is greatest in exactly the state analysed. The fallback also does not escape the mechanism, since electrodermal activity is sudomotor sympathetic output under the same brain-autonomic rhythm. One consequence the finding understates and I would add: phase-biased detection is not separable from genuinely phase-structured occurrence, so the prerequisite aim's result cannot be interpreted as "a separate finding about flash expression" (L296-297) as the draft claims. Moderate.
The problem
The draft defines continuity and fragility by sigma power *level*. Lecci's arousal data distinguish the *derivative*: wake-ups occurred when the noise fell in a phase of declining power, sleep-throughs when it fell in a phase of rising power. Level and slope are 90 degrees apart in phase, so the two framings make different predictions about where in the cycle registration should peak. Kjaerby's noradrenaline data point the same way — micro-arousals ride the peaks, spindles the descending phase. The draft's sin/cos parameterisation is agnostic and will find whatever preferred phase exists, so this does not break the analysis; but the Expected outcomes section commits to a direction ('flashes arriving in fragility (low sigma power) are predicted to be more likely registered', lines 356-358) that the primary source does not support in that form.
In the draft — line 23-25
alternating roughly every 25 s between a continuity phase of high sigma power and a fragility phase of low sigma power
Evidence
Lecci et al. 2017 (PMC5298853): "wake-ups and sleep-throughs occur during declining and rising sigma power levels, respectively"; "In a wake-up trial from a single mouse, sigma power was at its maximum before noise onset, such that noise exposure fell within a phase of declining power. In contrast, for a sleep-through trial of the same mouse, sigma power had just exited the trough, and noise was played within the phase of incrementing power". Kjaerby et al. 2022 (Nat Neurosci 25:1059-1070): "micro-arousals are generated in a periodic pattern during NREM sleep, riding on the peak of locus-coeruleus-generated infraslow oscillations of extracellular NE, whereas descending phases of NE oscillations drive spindles".
Fix
State the prediction in the terms the source uses: 'In mice, arousals followed stimuli delivered during declining rather than rising sigma power (Lecci et al., 2017), and micro-arousals coincide with peaks of the noradrenaline oscillation (Kjaerby et al., 2022). We therefore predict registration to be most likely for flashes arriving on the descending limb of the sigma envelope. The sin/cos parameterisation tests for any preferred phase without committing to this, and the estimated preferred phase is reported in radians relative to the sigma maximum.'
What the refuter said
AGAINST: the finding concedes its own defusal — the sin/cos parameterisation is agnostic and "will find whatever preferred phase exists", so the analysis is not broken. And high-sigma-continuity / low-sigma-fragility is the framing in general circulation, which the draft can reasonably adopt. SURVIVES. I fetched Lecci (PMC5298853) and the source's operational definition is the derivative, not the level: "offline periods correspond to raising, whereas fragility periods correspond to declining portions of the 0.02-Hz oscillation in spindle activity", with the arousal result "wake-ups and sleep-throughs occur during declining and rising sigma power levels, respectively". The draft instead defines the phases by level ("a continuity phase of high sigma power and a fragility phase of low sigma power", L24-25; repeated L98-100), which is 90° out from the source, and then commits to a level-based direction in Expected outcomes ("flashes arriving in fragility (low sigma power) are predicted to be more likely registered", L355-357). So the load-bearing citation is mischaracterised and the stated prediction points at the wrong quarter of the cycle. Moderate, as claimed: it misstates the source and the prediction, but does not invalidate the test.
The problem
The confirmatory analysis predicts "whether a detected flash is consciously registered" (l.280-281), and the proposal never states what makes a detected flash registered. Missing: (a) the time window after estimated onset within which a press counts — a flash lasts 1-5 minutes and a woman may wake, orient, and press minutes later, or press only at a later awakening; (b) the rule when one press falls near two detected flashes, or several presses fall near one; (c) the treatment of presses with no detected flash, which are direct evidence about the false-negative rate of the detector and about reporting behaviour and are simply unaddressed; (d) whether presses during scored wake count, given the analysis is "restricted to stable NREM" (l.270) — a press necessarily follows some arousal, so the press timestamp itself will usually fall outside the NREM bout that supplied the phase estimate; (e) the treatment of accidental presses in sleep. Different defensible choices of window (30 s, 2 min, 5 min, next awakening) will produce materially different outcome vectors and therefore different confirmatory results from the same data. The same gap applies to the supporting outcome: "a supporting model predicts cortical arousal" (l.281-282) with no stated interval after onset within which an arousal counts as flash-related, and an incomplete arousal criterion (l.237 gives "transient > 16 Hz activity lasting 3-15 s" but omits the preceding-stable-sleep requirement and the alpha/theta frequency shifts of the standard rule).
In the draft — line 216-218
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration.
Evidence
Draft l.216-218 (measure described) and l.280-281 ("predicting whether a detected flash is consciously registered") — no window, matching rule, or exclusion appears anywhere in DRAFT.md. Draft l.270 ("Analyses are restricted to stable (and non-artifactual) NREM") versus the fact that a button press requires arousal (l.299-300), so the outcome event is systematically outside the analysis window that defines the predictor.
Fix
Pre-specify the outcome fully, in the Procedure or Statistical analyses section: "A detected flash is coded as registered if a button press occurs within [W] s of the estimated onset (primary W = [value], with sensitivity analyses at [values]). Where one press falls within W of more than one detected flash it is assigned to the nearest onset; additional presses within W of the same flash are collapsed. Presses with no detected flash within W are counted, reported per participant as a false-alarm/miss index, and used in the detector-sensitivity analysis; they do not enter the registration model. Presses occurring during scored wake are retained provided the estimated onset falls in stable NREM. Press latency relative to onset is reported descriptively and is not used as an outcome." Add the matching interval and the full arousal criterion for the supporting outcome, and cite the AASM scoring manual rather than a substantive research paper for a scoring rule.
What the refuter said
STRONGEST DEFENCE: DRAFT.md L322-327 explicitly commits outcome definitions to a pre-registration before outcome data are inspected — "Hypotheses, outcome definitions, model specifications, inclusion criteria and the sensitivity analyses above are registered on the Open Science Framework". Deferring the matching window to the pre-registration is legitimate and standard; grant proposals are not analysis plans. Sub-point (d) is also wrong: "restricted to stable NREM" is most naturally an inclusion criterion on the PHASE ESTIMATE at onset, not a requirement that the press occur in NREM, so a press during subsequent wake is not excluded. And the arousal-criterion complaint is a nitpick — "Arousals follow standard criteria" plus a citation is adequate compression at this length. WHY THE CORE SURVIVES: I confirmed by reading L216-218, L280-302 and the whole file that no matching window, no multiple-press rule, no unmatched-press rule and no latency bound appears anywhere. The consequence is real and not hypothetical: a flash lasts 1-5 minutes, waking and locating a button takes time that plausibly correlates with sleep depth (the predictor), and the cycle is 50 s — so 30 s, 2 min, 5 min and next-awakening windows would produce materially different outcome vectors and different confirmatory results from identical data. Unmatched presses are also the cheapest available evidence about the detector's false-negative rate and are simply unused. DOWNGRADE: major to moderate, because the pre-registration clause covers the process even though the reviewer cannot evaluate feasibility without the window.
The problem
The dispersion of intervals between two peripheral markers measures their relative timing, not the error of either against the central event whose gating by ISF phase is the hypothesis. Any error component common to all peripheral effectors (sudomotor conduction delay, sweat-gland response latency, wrist-versus-sternal propagation) cancels in the difference and is invisible to this check. Worse, the check is entirely insensitive to systematic lag, and systematic lag is the error mode that destroys the study's directional interpretation: it rotates the estimated preferred phase rather than shrinking it.
In the draft — line 261-263
Intervals between this inflection and the accompanying temperature and pulse-rate changes index achievable precision.
Evidence
Simulation with pure systematic lag and zero jitter (/root/grantreview/sim/01_attenuation.py): the estimated preferred phase rotates by exactly 360*L/50 degrees per L seconds of lag, i.e. 7.2 deg/s. Measured rotations: 1 s -> 7.0 deg; 5 s -> 35.5 deg; 6.25 s -> 44.9 deg; 12.5 s -> 89.6 deg; 25 s -> 179.8 deg. A 12.5 s systematic back-dating error moves a true fragility peak onto the fragility/continuity boundary; 25 s inverts the label entirely. A zero-mean but SKEWED error also rotates the estimate: gamma(shape 2) onset error with SD 12.5 s and mean zero rotates the preferred phase by -31.7 deg (SD 15 s -> -50.2 deg). The draft's directional prediction at line 355-357 ('flashes arriving in fragility (low sigma power) are predicted to be more likely registered') therefore requires the onset estimate to be unbiased to within about +/-6 s, which the proposed check cannot establish.
Fix
Replace with: "Onset-timing error is decomposed into a variance component and a bias component. The variance component is estimated from the dispersion of intervals between peripheral markers; the bias component cannot be estimated this way and is bounded instead against a within-participant reference of known latency (for example, simultaneous sternal skin conductance in a laboratory sub-sample of the same women). Because a systematic lag of L seconds rotates the estimated preferred phase by 7.2*L degrees, the fragility-versus-continuity interpretation is pre-specified as conditional on a demonstrated bias below 6 s; absent that demonstration only the presence of phase dependence, not its direction, is interpreted."
What the refuter said
The core is an airtight logical point and it survives every defence I could build. The difference between two peripheral markers cancels any error component common to all peripheral effectors (sudomotor conduction delay, sweat-gland latency, wrist-versus-sternal propagation) and is identically zero under a pure systematic lag, so L261-263 cannot bound the error against the central event whose phase-gating is the hypothesis. Two corrections to the finding. Its 'simulation' of systematic lag is tautological — 360*L/50 = 7.2 deg/s is arithmetic, and reading /root/grantreview/sim/01_attenuation.py confirms the 'bias' branch simply adds a constant offset, so the measured rotations restate the input. The skew result is real: out_01_attenuation.json gives skew_phase_shift_deg -31.69 at SD 12.5 s and -50.17 at SD 15 s, matching the finding. STRONGEST DEFENCE, which forces the downgrade: the draft's confirmatory test is rotation-invariant by construction — L285-287 enters phase as sin and cos 'giving a two-degrees-of-freedom test for any preferred phase without fixing the fragility-continuity boundary in advance' — and L358-361 explicitly makes the fragility direction 'predicted rather than defining'. Systematic lag therefore cannot produce a false positive; it can only mislabel which phase. Also the draft's verb is 'index', a hedge, not a claim to measure absolute error. What remains: the interpretive claim at L366-367 does need a known phase, and the proposed check cannot supply it. Downgraded major to moderate.
The problem
Neither this sentence, nor the exclusion cascade at line 205-206, nor the contingency at line 346-347 says whether the required stable NREM bout precedes onset or brackets it. It matters decisively. If the window is centred on onset, inclusion requires roughly 75 s of arousal-free NREM AFTER onset, which removes the ~69% of flashes that end in an awakening - that is, conditioning on the negation of both outcomes. The arousal model then has almost no outcome variance and the button-press outcome loses exactly the events that could be pressed. Because the draft prescribes zero-phase filtering plus the Hilbert transform, which are non-causal, the centred reading is the natural one, so the ambiguity is not benign.
In the draft — line 270
Analyses are restricted to stable (and non-artifactual) NREM.
Evidence
de Zambotti et al. 2014 abstract (verified via Semantic Scholar, DOI 10.1016/j.fertnstert.2014.08.016): "69.4% of hot flashes were associated with an awakening." Yield under the two readings (/root/grantreview/sim/08_yield_v2.py, central scenario): pre-onset-only window, N = 347 analysable flashes and ~52 registered; centred +/-75 s arousal-free window, N = 104 and ~5 registered (only ~5 of 137 women contribute a registered flash). The pre-onset reading is affordable: measured phase error at the final sample of a pre-onset-only window (/root/grantreview/sim/07_edge_terminal.py, 2nd-order Butterworth 0.01-0.04 Hz, filtfilt+Hilbert, 600 replicates) is a circular SD of 38.8-43.2 deg (5.4-6.0 s equivalent) with lambda = 0.75-0.80 and a bias of +4 to +6 deg, essentially independent of window length from 100 s to 600 s.
Fix
Specify: "The estimation window is the 150 s of stable, artefact-free NREM immediately PRECEDING onset; no requirement is placed on sleep continuity after onset, since arousal and registration are the outcomes and conditioning on post-onset stability would select on them. Phase at the onset sample is taken from the terminal sample of the pre-onset window, which carries a measured circular SD of about 40 deg (attenuation factor 0.78) from the filter edge; this factor is included in the power calculation."
What the refuter said
STRONGEST DEFENCE: the centred reading is not compelled. Zero-phase filtering needs only filter settling after the sample — order-2 filtfilt in this band spans about ±36 s of 99% impulse energy — and that trailing data need only EXIST, not be arousal-free NREM; you can filter across an arousal and still take the phase at onset. The stable-bout requirement is stated for phase estimability, which is a pre-onset property, and the draft's own contingency ("relaxing the stable-bout criterion from three ISF cycles to two", L344-346) reads naturally as being about the bout containing the onset. So the finding's claim that the outcome-destroying reading is "the natural one" is argument, not demonstration. WHY IT SURVIVES: the ambiguity itself is real and verifiable — L270 says only "Analyses are restricted to stable (and non-artifactual) NREM", and neither the exclusion cascade (L204-206) nor the contingency (L344-346) says whether the bout must precede onset or bracket it. The de Zambotti figure is confirmed from the abstract via Europe PMC: "69.4% of hot flashes were associated with an awakening", so under the centred reading the criterion would condition on the negation of both outcomes and remove most events — the stakes are as high as claimed. The affordability check is also fair: pre-onset-only terminal-sample phase error of roughly 39-43° circular SD (5-6 s equivalent) is tolerable. This is the honest version of the point that F184 asserts as fact. DOWNGRADE: major to moderate — it is a one-sentence specification gap, not a demonstrated defect.
The problem
Two independent gradients run in opposite directions across the night and the draft accounts for neither. The infraslow spindle modulation is reported as more prominent in the second half of the night, while flashes are more frequent and more ambient-temperature-sensitive in the first half and are suppressed by the REM that dominates the second. So measurement quality of the predictor and supply of the events are anti-correlated, which means the effective information per night is concentrated in a narrow window and the night-half covariate is absorbing at least three distinct mechanisms at once. A single linear covariate cannot separate them, and if the phase effect itself differs by half of night the model is misspecified.
In the draft — line 288-291
with covariates for sleep stage, time since the preceding flash, and half of the night - the last because the flash-arousal relationship may reverse across the night
Evidence
Lazar 2019 abstract: "ISO in sleep spindles are most prominent in the centro-parieto-occipital regions, left hemisphere and second half of the night independent of the number of spindles" (https://research-portal.uea.ac.uk/en/publications/infraslow-oscillations-in-human-sleep-spindle-activity). Freedman & Roehrs 2006: the 18 degrees C reduction in flash count occurred only in the first half, and "In the second half of the night, rapid eye movement sleep suppresses hot flashes and associated arousals and awakenings" (https://pubmed.ncbi.nlm.nih.gov/16837879/). Sano et al. 2014: longer EDA storms cluster in the first half of sleep.
Fix
Report the joint distribution of analysable events and measured ISF strength by half of the night, or better by NREM cycle number, before the confirmatory test. Add a pre-specified phase-by-half-of-night interaction rather than a main effect only, and state that the covariate stands in for at least three mechanisms (flash-arousal ordering, REM suppression, ISF amplitude gradient). Consider using ISF amplitude at the event as a per-event weight or precision variable rather than treating all events as equally informative.
What the refuter said
AGAINST: the draft does the right thing — it includes half-of-night as a covariate (L288-291) — and its stated rationale is sourced, not wrong: Freedman & Roehrs's own conclusion is about the second half of the night. Criticising a covariate for being justified by one valid reason rather than three is weak, and "a single linear covariate cannot separate three mechanisms" is true of every covariate in every model; the draft explicitly labels everything beyond the primary registration model exploratory (L310-311). Survives only as an unaddressed feasibility observation, and its two anchor facts do verify: Lazar 2019's abstract states ISO in spindles are "most prominent in the centro-parieto-occipital regions, left hemisphere and second half of the night", and Freedman & Roehrs's cooling effect was confined to the first half. Predictor quality and event supply therefore run in opposite directions across the night, which the draft nowhere reckons with. Worth a sentence; changes nothing about validity. Note the same Lazar quote raises a sharper point the finding does not make — the draft records frontal derivations only, while Lazar locates the effect centro-parieto-occipitally. Minor.
The problem
The draft defines sigma as 11-16 Hz (line 23 and line 96). Lecci used 10-15 Hz in both species; Osorio-Forero used 10-15 Hz in mouse S1; Lazar's ISO was most prominent in the high sigma band, 13-15 Hz, i.e. the fast centro-parietal spindle population. At frontal derivations the dominant spindle population is the slower frontal type, so an 11-16 Hz frontal band maximally samples the spindle class least associated with the ISF while extending into the frequency range where frontalis muscle and eye-movement artefact live — with no EMG or EOG channel to detect them. The band choice is unsourced and works against the design in both directions.
In the draft — line 22-23
an infraslow fluctuation of sigma-band power at approximately 0.02 Hz
Evidence
Lecci 2017: sigma defined as "10 to 15 Hz" for primary analyses (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853). Osorio-Forero 2021 (/root/grantreview/OsorioForero2021.txt): "the sigma (10-15 Hz) and the delta (1.5-4 Hz) frequency bands". Lazar 2019 abstract: "The analyses focused on fast sleep spindle and sigma activity (13-15 Hz)"; ISO "most prominent in the high sigma band".
Fix
Justify the band or align it with the source literature, and report the phase result across at least two band definitions (10-15 Hz and a participant-specific fast-spindle band) as a robustness check. Add an explicit high-frequency artefact-rejection step before the envelope is computed, and report what fraction of stable-NREM time is lost to it.
What the refuter said
AGAINST: the artefact half is weak. 16 Hz is a conventional sigma upper edge — I confirmed Esfahani et al. use 13-16 Hz and Dimitriades et al. 2026 use 10-16 Hz — so the draft's upper limit is unremarkable, and the ZMax has no EMG or EOG channel regardless of band, making that a device limitation rather than a band-choice error. The analysis operates on the log-sigma ENVELOPE's infraslow phase, which is robust to a 1 Hz shift in a band edge. The frontal-expression concern is also pre-empted twice: L274-276 ("Frontal expression and achievable phase precision are themselves feasibility outcomes") and L343 ("weak frontal ISF expression triggers the same rule"). SURVIVES as a citation-consistency fact, and my checks sharpen it. Lecci 2017 uses "the sigma (10 to 15 Hz) power band"; OsorioForero2021.txt gives "the sigma (10-15 Hz) and the delta (1.5-4 Hz) frequency bands"; Lazar et al. 2019's abstract gives fast spindle/sigma as 13-15 Hz and adds that the ISO are "most prominent in the high sigma band and over the centro-parieto-occipital regions". So 11-16 Hz matches none of the three cited sources, and Lazar's topography is a real argument against frontal-only recording — which is why the draft's two pre-emptions matter. Minor, as the finding itself rates it: source the band or align it.
The problem
The supporting outcome is whether an arousal follows a flash, scored by a human from two frontal channels with no EOG or EMG, in the presence of movement generated by the event. If the scorer can see the flash detection, the wristband markers or the button presses, an ambiguous frontal transient near a known flash will be scored as an arousal more often than the same transient elsewhere. That single unblinded judgement can manufacture the supporting result and, because arousal and registration are near-nested, contaminate the primary one. Inter-rater agreement between two equally unblinded scorers does not address this.
In the draft — line 233-235
Recordings are scored in 30-s epochs (wake, REM, N1-N3) by an experienced scorer, with a double-scored subsample giving inter-rater agreement.
Evidence
Absence: no occurrence of blind, blinded or masked anywhere in DRAFT.md. Structural: staging and arousal scoring are performed on the same recordings that carry the accelerometer transients from the taps and from flash-related movement, and the button press is on a synchronised device.
Fix
State explicitly that all EEG scoring is performed on recordings stripped of wristband channels, flash detections and button-press markers, that scoring is completed and locked before the flash and registration data are merged, and that the merge is performed by a different person. Report this as a pre-registered procedural safeguard, and add a manipulation check: verify that the scorer cannot predict flash timing above chance in a blinded subsample.
What the refuter said
STRONGEST DEFENCE, LARGELY SUCCESSFUL: the finding's own claim of consequence is wrong. The PRIMARY outcome is a device button press with a hardware timestamp — no scorer touches it — so an unblinded scorer cannot "contaminate the primary one", and the near-nesting of arousal and registration does not transmit scorer bias into a press event. Structurally, flash detection lives on the EmbracePlus and staging on the ZMax record, so a scorer working from the EEG file has no flash markers unless someone deliberately merges them; the tap transients are at lights-off and waking, not at flashes. So the bias channel is latent and avoidable rather than built in. WHY A KERNEL SURVIVES: the absence is real (grep returns zero occurrences of blind, blinded or masked anywhere in DRAFT.md), the supporting outcome IS a human judgement about a transient near a known event scored from two frontal channels with no EOG or EMG, and blinding sleep scoring to hot-flash markers is standard in exactly this literature. Inter-rater agreement between two equally unblinded scorers genuinely does not address it. One sentence fixes it. DOWNGRADE: moderate to minor — the stated consequence overreaches and the primary outcome is immune.
The problem
Three ISF cycles at 0.02 Hz is 150 s; two is 100 s. Both are below the minimum NREM bout length used in the two published human ISF characterisations that define the method the draft adopts: 280 s (chosen explicitly to contain two full cycles of the LOWEST frequency of interest, 0.0075 Hz) and 300 s. This matters for a concrete reason the draft does not address: a band-pass at 0.01-0.04 Hz cannot be estimated, and its "empirical peak reported per participant" cannot be located, from a 100 s segment — the frequency resolution is not there, and edge effects from zero-phase filtering consume a large fraction of a short bout. The relaxed fallback is therefore not a weaker version of the same analysis; it is an analysis whose predictor is undefined.
In the draft — line 343-345
Low event yield is met by relaxing the stable-bout criterion from three ISF cycles to two, pre-specified as secondary.
Evidence
Dimitriades et al., bioRxiv 2024.11.06.620875 (https://www.biorxiv.org/content/10.1101/2024.11.06.620875v1.full): bouts required "≥280 seconds continuous N2 sleep" because "The minimum bout length of 280 seconds was chosen to ensure two full cycles of the lowest frequency of interest (0.0075 Hz)." — Dimitriades et al., bioRxiv 2025.04.23.650209: "N2 sleep bouts lasting at least 300 seconds" to "ensure adequate frequency resolution." Contrast DRAFT.md lines 267-270, which specify a 0.01-0.04 Hz band-pass with "empirical peak reported per participant."
Fix
Raise the criterion and drop the fallback. Set the minimum stable-NREM bout to ≥280 s, matching Dimitriades et al., and cite that rationale. Delete the two-cycle relaxation, or replace it with a different lever that does not corrupt the predictor — e.g. "low event yield is met by extending the number of EEG nights, not by shortening the estimation window, since the ISF band-pass cannot be estimated from bouts shorter than ~280 s." Report edge-effect handling (discarded filter transient at each bout boundary) explicitly, since it further shortens the usable window.
What the refuter said
STRONGEST DEFENCE, PARTLY SUCCESSFUL: the sharpest sub-claim is refuted by the draft's own wording. "Empirical peak reported per participant" (L268-269) is PER PARTICIPANT, not per bout — so the peak is estimated over all of a participant's data and only instantaneous phase is taken within a bout. The frequency-resolution objection to "locating the peak from a 100 s segment" therefore attacks something the draft does not do, and "an analysis whose predictor is undefined" is too strong. Instantaneous Hilbert phase in a fixed band does not require the segment to contain two cycles of the band's low edge. WHY A KERNEL SURVIVES: I verified the comparison sources directly. The bioRxiv preprint states "Bouts of NREM Stage 2 (N2) sleep data that lasted for at least 280 seconds" because "The minimum bout length of 280 seconds was chosen to ensure two full cycles of the lowest frequency of interest (0.0075 Hz)", and it uses the identical 0.01-0.04 Hz detection band the draft specifies — so this is the same method with a bout minimum nearly two to three times the draft's 150 s, relaxing to 100 s. The edge-effect concern is real in kind: order-2 filtfilt in this band has 99% impulse energy spanning about ±36 s, and median absolute phase error runs roughly 32-35° at a segment edge falling to 15-17° at 25 s inside, so a 100-150 s bout with onset near an edge carries material error. The draft does not say where in the bout onset must sit — the same specification gap as F94. DOWNGRADE: moderate to minor; frame it as "justify the bout minimum against the published 280-300 s and state the onset's minimum distance from the bout edge".
The problem
The draft's precision budget is expressed as a fraction of a 50 s cycle, and 50 s comes from rodent LC work (Osorio-Forero et al. 2025) and from 0.02 Hz in Lecci et al. The band actually used, 0.01-0.04 Hz, spans periods of 25-100 s — a fourfold range. If the true period in a given participant is 25 s, the draft's own quarter-cycle jitter tolerance is 6 s, not 12.5 s. The draft acknowledges an "empirical peak reported per participant" but never states what period range it expects, what happens when no peak is identifiable, or how the precision requirement is recomputed per participant. Relevant and uncited: a 2026 human paper co-authored by the very team member the draft invokes reports that the infraslow fluctuation of sigma power varies systematically across development, with "frequency, variability, and strength increasing from early to late adolescence" — evidence that the period is not a constant and has not been characterised in women aged 40-55, let alone in insomnia. Note the draft also introduces an undefined acronym, "IAF phase", in the sentence that defines the study's central predictor; IAF conventionally denotes individual alpha frequency, and a reviewer will read this as either a typo for ISF or a different quantity entirely.
In the draft — line 266-270 (also 315-318, 44-46)
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase.
Evidence
DRAFT.md line 269 ("0.01-0.04 Hz" = 25-100 s periods) against lines 317-318 ("a 50-s cycle", "a quarter of the cycle") and lines 44-46 ("the infraslow scale of roughly 50 seconds"), which derive from rodent LC data (Osorio-Forero et al. 2025, https://www.nature.com/articles/s41593-024-01822-0). Dimitriades, Osorio-Forero, Fattinger et al. (2026), Scientific Reports, https://www.nature.com/articles/s41598-026-58423-z: "The infraslow fluctuation of sigma power (ISFS)—the clustering of sleep spindles over 10–100 s"; "Results indicate that the ISFS is present across all ages, with frequency, variability, and strength increasing from early to late adolescence." I could retrieve only the abstract of this paper; its methods and any reliability figures are UNVERIFIED. DRAFT.md line 270 contains "the IAF phase" with no prior definition of IAF.
Fix
Make the period an estimated parameter with consequences, not a constant. Add: "The ISF period is estimated per participant from the spectral peak of the log-sigma envelope within 0.01-0.04 Hz; the jitter tolerance for that participant is a quarter of their own period, so the precision criterion is applied per participant rather than against a nominal 50 s. Participants without an identifiable peak are excluded from the phase analysis and reported." Cite Dimitriades et al. 2026 as evidence that ISF frequency varies across individuals and development, and name characterising it in midlife women as a contribution of this study. Fix "IAF" to "ISF".
What the refuter said
AGAINST: the headline is wrong. 50 s is not imported from rodents — Lecci reports a human peak of 0.019 ± 0.001 Hz (n=27), a ~53 s cycle, matching mice at 0.021 Hz, and the draft says as much at L96-97 ("present in both mice and humans"). The draft also already concedes participant variation ("empirical peak reported per participant", L268) and lists frontal expression and achievable precision as feasibility outcomes (L275-276). The passband-versus-50-s tolerance mismatch is F74's point, not a second finding. Two things survive, both small. First, "the IAF phase" at L269-270 is a genuine error: IAF is undefined in the document and conventionally denotes individual alpha frequency, in the sentence that defines the study's central predictor — trivial to fix, but it sits in the worst possible place. Second, the uncited Dimitriades et al. 2026 does exist (Scientific Reports s41598-026-58423-z, PMID 42310467) and reports the ISF's frequency and strength changing across development, so the period is not a constant and is uncharacterised in women aged 40-55; notably it comes from the applicants' own network (NIN, with Osorio-Forero), which makes the omission slightly awkward. Minor.
The problem
The primary outcome is per-flash and binary, so every press must be assigned to a flash or discarded, and every flash must be labelled pressed or not. That requires a matching window, and none is stated anywhere in the draft. The window is not a detail: a flash lasts one to five minutes, waking and orienting and locating a wristband button takes time that will vary with depth of sleep - which is the predictor - and the ISF cycle is 50 s. So a plausible latency spans more than a full cycle, and the latency itself is expected to correlate with phase. There is also no rule for presses with no nearby detected flash (false reports, or true flashes the detector missed), and no rule for two flashes inside one window. Each choice moves the primary result.
In the draft — line 216-218
On EEG nights participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration.
Evidence
Draft L66-69 and L216-218 describe the press; the Statistical analyses section (L280-302) treats registration as a per-flash binary with no matching rule; no latency window appears anywhere in the draft. Tsiartas et al. 2021 illustrates that even detector-to-reference matching required an explicit tolerance: "We time-aligned the features with the HF expert annotations for prediction and evaluation (+/-90 s matching window)" (/root/grantreview/Tsiartas2021.txt).
Fix
Pre-specify the matching rule and its sensitivity range: a press is attributed to the nearest detected flash whose onset falls within a stated window (e.g. onset to onset plus 300 s), with the primary analysis at one window and pre-registered sensitivity analyses at two others. Pre-specify handling of unmatched presses, multiply matched presses, and flashes with more than one press. Report the press-latency distribution and test whether latency varies with ISF phase; if it does, the binary outcome is contaminated and a latency-aware model is required. Also record and report the response-latency floor from an evening wake calibration (press on cue) so the sleep latencies can be interpreted.
What the refuter said
STRONGEST DEFENCE: L322-327 commits outcome definitions to a pre-registration before outcome data are inspected, which is where a matching window conventionally lives, and this is a strict subset of F70 — same quoted text, same absence, fewer sub-points, no new evidence. Reporting both double-counts one gap. WHY THE FACT SURVIVES: I confirmed by reading L216-218 and the entire Statistical analyses block (L278-311) that registration is treated as a per-flash binary with no matching window, no rule for unmatched presses, and no rule for two flashes inside one window. The finding's best contribution is the observation the parent findings understate: the press latency is expected to CORRELATE WITH THE PREDICTOR, because waking, orienting and locating a wristband button takes longer from deeper sleep — so the window choice is not merely arbitrary, it interacts with phase. Against a 50 s cycle, a plausible latency spans more than a full cycle. The Tsiartas parallel is apt ("±90 s matching window") — even detector-to-reference matching required an explicit stated tolerance. DOWNGRADE: moderate to minor, as a duplicate of F70; fold the latency-correlates-with-predictor point into F70 and drop this entry.
The problem
The Wassing criterion is quoted correctly, but it is not the AASM rule. AASM scores an arousal on an abrupt shift toward faster frequencies including alpha and theta as well as >16 Hz activity, excluding spindles, with a 3 s minimum and no 15 s ceiling — so '>16 Hz, 3-15 s' will systematically miss alpha- and theta-only arousals and cap long ones. More seriously, AASM arousal scoring in REM additionally requires a concurrent rise in submental EMG, and AASM staging of REM requires EOG and chin EMG. The ZMax records two frontal EEG channels, accelerometry and PPG — no EOG, no EMG. So 'scored in 30-s epochs (wake, REM, N1-N3) by an experienced scorer' (line 234-236) cannot be AASM-conformant staging, and the validation the draft relies on did not assess arousal detection from this device at all. Since the analysis excludes REM and the supporting outcome is cortical arousal, both the exclusion and the outcome rest on unvalidated scoring.
In the draft — line 236-237
Arousals follow standard criteria: transient \> 16 Hz activity lasting 3-15 s (Wassing et al., 2019).
Evidence
Wassing et al. (2019), Current Biology 29(14):2351-2358, STAR Methods (https://www.b-radlab.com/uploads/1/4/2/0/142020983/wassing_2019.pdf): "cortical arousals during sleep were indicated by transient high-frequency EEG activity (>16 Hz) lasting between 3 and 15 s", citing Bonnet et al. (1992), not the current AASM manual. On the AASM rule, StatPearls 'EEG Normal Sleep' (https://www.ncbi.nlm.nih.gov/books/NBK537023/): "EEG arousals appear as sudden shifts in frequency toward faster rhythms (theta, alpha, beta, but not sigma)". On the device, Esfahani et al. (bioRxiv 2023.08.18.553744): "two frontal EEG channels (F7-Fpz, F8-Fpz), a tri-axial accelerometer, and a PPG sensor", with no EOG or chin EMG in the Lite version; arousal scoring was not assessed, and the authors note "absence of EMG/EOG in Lite version" and "reliance on autoscoring (not manual visual inspection per AASM standard)" as limitations. UNVERIFIED verbatim: the AASM manual's exact arousal rule text (aasm.org and jcsm.aasm.org both blocked by robots.txt).
Fix
Drop 'standard criteria' and be explicit: 'Arousals are scored as transient >16 Hz activity lasting 3-15 s (Wassing et al., 2019). This is narrower than the AASM rule, which also counts alpha and theta shifts and, in REM, requires concurrent submental EMG; the headband carries no EOG or EMG channel, so staging and arousal scoring are EEG-only and are not AASM-conformant. We therefore report inter-rater agreement for staging and arousals on a double-scored subsample, and treat REM misclassification into NREM as a sensitivity analysis.'
What the refuter said
Verified independently: the AASM ISR page gives 'Arousals must last at least three seconds. Arousals must be preceded by at least 10 seconds of continuous sleep' and 'Arousals during R must also have an increase of chin EMG lasting at least one second'. Esfahani et al. confirms the ZMax Lite 'comprises two frontal EEG channels (F7-Fpz, F8-Fpz), a tri-axial accelerometer, and a PPG sensor' with no EOG or chin EMG, and confirms arousal detection was not assessed. So the facts hold. STRONGEST DEFENCE, and it substantially deflates the finding: the draft never claims AASM conformance. L234 says 'scored in 30-s epochs (wake, REM, N1-N3) by an experienced scorer', and L236-237 says 'standard criteria' citing Wassing et al. (2019) — a paper from the applicants' own Amsterdam UMC group, i.e. an in-house published criterion suited to limited-channel data. The finding attacks a conformance claim the document does not make. Also note the '>16 Hz' restriction is a stricter spindle exclusion than AASM's, not a looser one, which matters given the finding's own concern about an 11-16 Hz predictor. What survives is real but modest: the supporting arousal outcome and the REM exclusion both rest on scoring never validated for this device. Supporting outcome, not primary. Downgraded moderate to minor; duplicate of F119.
The problem
The 1.5 figure is a reported count, not a count of time-stamped in-night presses. The proposal's primary outcome is a different instrument: "a time-stamped button press made as soon as the participant thinks she is having a flash" (l.66-68), with the report-style measure demoted to "a complementary measure of recall and compliance" (l.68-69, l.218-219). The two can diverge in either direction: pressing requires waking enough to act, which will lose flashes that were consciously experienced but not acted on; conversely a woman may recall in the morning a flash she did not press for. So the 43% base rate underpinning every implicit expectation of outcome balance may not transfer to the press outcome, and the proposal never estimates press-based registration rates from any source. Relatedly, the morning diary is collected but has no stated analytic role beyond the word "complementary": it is not used to validate the press, to estimate press sensitivity, or to bound under-pressing, which is the obvious use and would cost nothing.
In the draft — line 40-41
Women have on average 3.5 objectively recorded flashes per night but only report 1.5 (De Zambotti et al., 2014).
Evidence
Draft l.40-41 (report-based 1.5/3.5) versus l.66-68 and l.216-218 (outcome is a prospective time-stamped press) and l.218-219 ("A morning diary estimate gives a complementary measure of recall and reporting behaviour") — no analysis links the two, and no press-based base rate is stated anywhere.
Fix
Distinguish the measures and give the diary a job: "Published registration rates (1.5 of 3.5 flashes; De Zambotti et al., 2014) rest on retrospective report and are not directly transferable to a prospective button press, which additionally requires sufficient arousal to act. We therefore treat the press-based registration rate as unknown, report it as a primary descriptive outcome, and use the morning diary count as an independent estimate of the number of flashes the participant was aware of, so that under-pressing can be quantified per participant (press count versus diary count versus detected count) and used as the compliance criterion for inclusion in the registration model."
What the refuter said
STRONGEST DEFENCE, MOSTLY SUCCESSFUL: the draft already knows the two instruments are different and says so — the press is "a prospective measure of conscious registration" (L217) and the diary is "a complementary measure of recall and reporting behaviour" (L218-219). Distinguishing them is the draft's design choice, not a conflation. More decisively, the finding's load-bearing claim is invented: it asserts the 43% figure underpins "every implicit expectation of outcome balance", but the draft states NO expected base rate for any outcome anywhere (I checked the whole file), so there is no expectation for the mismatch to corrupt. The 3.5-vs-1.5 pair is used as motivation in the Summary and Literature Review, not as a design parameter. WHY A KERNEL SURVIVES: no press-based registration rate is stated or sourced, and the morning diary — which is being collected anyway — is given no analytic role, when using it to bound under-pressing or estimate press sensitivity would cost nothing and is the obvious use. That is a modest, cheap improvement. DOWNGRADE: moderate to minor; and the no-base-rate half is already counted at major in F28.
The problem
Three problems. (1) It is called a prerequisite (l.181-182: "A prerequisite aim establishes whether onsets occur across the whole cycle or are themselves phase-structured") but nothing is contingent on its result: the draft states elsewhere that phase-structured occurrence "is not a reason to exclude such events; it is a separate testable outcome" (l.165-166). If no analysis decision depends on it, it is a third parallel aim, not a prerequisite, and calling it one implies a gating logic that is never stated. (2) The test is one clause. "Stage- and time-matched surrogates" specifies no test statistic (Rayleigh? omnibus circular test? a mixed model with phase as outcome?), no number of surrogates, no matching tolerance, and no criterion for concluding non-uniformity. It cannot be pre-registered as written. (3) It is not identifiable with the available measurements. The observable is the phase distribution of DETECTED onsets, which is the product of the true occurrence distribution and the detector's phase-dependent sensitivity — and the draft concedes phase-dependent sensitivity "cannot be assumed absent" (l.249-251). Without a reference standard whose sensitivity does not depend on ISF phase, phase-structured generation and phase-structured detection are indistinguishable, so the aim cannot be answered by this design as described.
In the draft — line 294-297
Whether onset is itself distributed across the ISF cycle is characterised first, against stage- and time-matched surrogates; phase-dependent occurrence would be a separate finding about flash expression, not conscious gating.
Evidence
Draft l.181-182 ("A prerequisite aim") against l.165-166 ("This is not a reason to exclude such events; it is a separate testable outcome") — no gating decision anywhere. Draft l.294-296 (surrogate test, no statistic or procedure). Draft l.249-251 ("It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent") makes the occurrence estimate non-identifiable without a phase-independent reference.
Fix
Either specify and qualify it, or relabel it. Suggested: "Whether flash onsets are themselves distributed non-uniformly across the ISF cycle is tested first, by comparing the observed circular distribution of onset phases to [N = 1000] surrogate onset sets drawn from the same participants' stable NREM matched on sleep stage and time within the night, using [named statistic] with the surrogate distribution as the null. Because detector sensitivity may itself depend on ISF phase, this test cannot separate phase-structured occurrence from phase-structured detection; the interpretation is therefore conditional on the sensitivity function estimated in the sternal-reference validation subsample. No analysis decision is contingent on this result, so it is reported as a parallel descriptive aim rather than a prerequisite."
What the refuter said
Strongest case against: two of three limbs are weak. (1) "Prerequisite" is plainly used in the sense the draft itself gives at L294-295 — "characterised first" — i.e. logically prior for interpretation. Reading an implied gating logic into it and then faulting the draft for not stating it is a word-choice nitpick. (2) Deferring the surrogate test's statistic is normal for a proposal that explicitly says the OSF registration will carry "outcome definitions, model specifications, inclusion criteria" (L322-327), and the draft labels this aim exploratory: "The primary registration model is confirmatory; all else is exploratory" (L310-311). A grant is not the pre-registration. (3) is correct and unaddressed: the observable is the phase distribution of *detected* onsets, which is the product of true occurrence and detection sensitivity, and the draft concedes at L249-251 that phase-dependent sensitivity "cannot be assumed absent", so stage- and time-matched surrogates do not identify occurrence. That is a genuine logical point, but it bears on an explicitly exploratory aim rather than the confirmatory test. Minor.
The problem
The lowest passband frequency corresponds to a 100 s period. The stable-bout requirement implied by the contingency section is three ISF cycles (l.345-346) — 150 s at 0.02 Hz — and the contingency relaxes it to two cycles, i.e. 100 s. Filtering and Hilbert-transforming a 100-150 s segment with a passband extending to 0.01 Hz leaves phase estimates dominated by edge effects across a large fraction of the segment: the analytic signal near both boundaries is unusable, and the usable interior may be a small central window. The draft states no edge-discard rule, no minimum distance from bout boundaries for an includable onset, and no treatment of bouts shorter than the filter's impulse response. Nor does it say whether filtering is applied per bout, per NREM period, or across the whole night with non-NREM interpolated (each choice has different artefacts, and the last would reintroduce non-NREM data into the phase estimate). Narrowing the band to the participant's empirical peak would mitigate this but the draft says the peak is only "reported", not used to define the filter.
In the draft — line 267-269
zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform
Evidence
Draft l.267-269 (0.01-0.04 Hz zero-phase band-pass plus Hilbert) against l.345-347 ("relaxing the stable-bout criterion from three ISF cycles to two"): 3 cycles at 0.02 Hz = 150 s, 2 cycles = 100 s, while the passband's low edge has a 100 s period. Arithmetic: 1/0.01 Hz = 100 s.
Fix
Specify the estimation geometry: "The sigma envelope is band-pass filtered within each stable NREM bout using a participant-specific band centred on that participant's empirical ISF peak ([peak] ± [width] Hz), rather than the full 0.01-0.04 Hz range, so that the filter's impulse response is short relative to the bout. Phase estimates within [E] s of a bout boundary are discarded, and a flash is analysable only if its onset lies at least [E] s inside the bout. Bouts shorter than [L] s are excluded. Filtering is performed per bout; no interpolation across non-NREM is used." Report how many events each of these rules removes.
What the refuter said
AGAINST: the premise that filtering happens per bout is the finding's own assumption, and it flags this itself ("Nor does it say whether filtering is applied per bout, per NREM period, or across the whole night"). The draft says only that "Analyses are restricted to stable (and non-artifactual) NREM" (L270) — a restriction on which *events* enter analysis, not on the filtering segment. Human NREM episodes run tens of minutes, so the standard implementation (filter the continuous sigma envelope over NREM periods, then select events sitting inside stable bouts) has no edge problem at all. Lecci did the analogous thing with continuous wavelet decomposition. So the arithmetic is right but the scenario is not established. What survives is a specification gap: no edge-discard rule, no minimum distance of an includable onset from a bout boundary, no statement of the filtering unit. Legitimate to raise as a question, but this is filter-parameter detail of the kind a 1,400-word methods section routinely omits, and it does not change validity. Minor.
Contradictions, undefined terms, and arguments that do not carry the weight put on them.
The problem
A pre-registration whose thresholds are described but not stated is not a pre-registration. Enumerating what is promised and left blank: (1) the "pre-registered threshold" for dispersion of inter-marker intervals (341); (2) "far more timing error" — unquantified (343); (3) "weak frontal ISF expression" — no definition, no threshold (344); (4) "Low event yield" — no triggering count (345); (5) "a pre-specified floor" for reporting compliance (348-349); (6) "artefact-free recording" — no artefact criterion (206); (7) "reliable onset timing" — no reliability criterion (206); (8) "stable (and non-artifactual) NREM" — "stable" defined only obliquely, 40 lines later, as three ISF cycles (270, 345-346); (9) "empirical peak reported per participant" for the ISF band — no rule if no peak is identifiable (269); (10) "the minimum detectable effect is reported" — no MDE (319-320); (11) "residual error is reported" for device synchronisation — no maximum acceptable value (230-231). Separately, the escape hatch in (2) is not what it claims: switching from phase to "pre-flash ISF amplitude" does not test "the same hypothesis". The draft's primary question is explicitly about phase — "whether the phase of the infraslow fluctuation at hot flash onset predicts whether that flash is consciously registered" (line 48-50) — and pre-flash amplitude is a different construct answering a different question (whether ISF strength, not ISF timing, gates registration). Presenting a fallback to a different hypothesis as robustness is the single sentence in this draft most likely to draw a hostile reviewer's pen.
In the draft — line 340-349
if it exceeds a pre-registered threshold, the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis; weak frontal ISF expression triggers the same rule. Low event yield is met by relaxing the stable-bout criterion from three ISF cycles to two, pre-specified as secondary… participants below a pre-specified floor are dropped from the registration model but retained for the arousal model.
Evidence
Internal: DRAFT.md lines 206, 230-231, 269, 270, 319-320, 340-349 all name a criterion without a value. Logical contradiction: line 48-50 defines the primary question as phase-dependent; line 342-344 substitutes amplitude while asserting it tests "the same hypothesis". Bial's assessment basis, /root/grantreview/BIAL_CONTEXT.md lines 42-44, is the objectives as stated in the application — objectives with unspecified decision rules cannot be assessed or later verified.
Fix
Put a number on each of the eleven. Concretely: the timing-precision gate as an interquartile range or SD in seconds (e.g. "if the SD of the inflection-to-temperature interval exceeds 15 s"); "weak frontal ISF expression" as a spectral criterion (e.g. "no identifiable spectral peak between 0.01 and 0.04 Hz in the log-sigma envelope in ≥50% of participants"); the low-yield trigger as a flash count (e.g. "fewer than 250 analysable flashes"); the compliance floor as a rate (e.g. "fewer than 30% of detected flashes registered on any night, or fewer than 2 presses per recorded night"); a maximum sync residual in seconds; a stated MDE. And reframe the amplitude fallback honestly: "If timing precision proves insufficient for phase, the confirmatory test is abandoned and pre-flash ISF amplitude is analysed as a distinct, secondary question about ISF strength rather than ISF timing." A stated kill criterion strengthens an application; a hypothesis substitution weakens it.
What the refuter said
Strongest case against the larger half: a grant narrative is not a pre-registration, and the draft says so - "Hypotheses, outcome definitions, model specifications, inclusion criteria and the sensitivity analyses above are registered on the Open Science Framework before outcome data are inspected" (L322-327). Thresholds for artefact rejection, onset reliability and ISF expression legitimately depend on pilot characteristics not yet in hand, and naming a decision rule without its numeric value is ordinary practice at proposal stage. So the eleven-item enumeration, which I checked item by item against L206, L230-231, L269-270, L319-320 and L340-349 and found accurate, is weaker than it looks. What is airtight and serious is the logical contradiction, which needs no external source. L48-50 defines the primary question as "whether the *phase* of the infraslow fluctuation at hot flash onset predicts whether that flash is consciously registered". L341-343 then says that on a threshold breach "the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing *the same hypothesis*". Amplitude is a different construct answering a different question - whether ISF strength, not ISF timing, gates registration - so the confirmatory hypothesis is not fixed, and the same rule is also triggered by "weak frontal ISF expression", where the substitute predictor is derived from the very signal that failed. Major, on that sentence.
The problem
The section names four risks, all analytic, all internal, all with the same shape (a threshold is crossed, an analysis is substituted). It omits every structural and existential risk. Missing: (1) the primary outcome variable may not exist in the data at all, because the button press is not in the host protocol; (2) half the sample receives a drug that suppresses the outcome event by ~75% and improves sleep, and a quarter receives a behavioural therapy that changes the predictor; (3) the host trial's recruitment may fall short of 222 or run past the Bial project window, and device participation is optional and revocable; (4) data access depends on a consortium proposal and publication may be embargoed behind the host trial's primary paper; (5) the flash detector has no ground truth in this sample and was built on three women with different hardware — its sensitivity here is unknown and its false positives dilute the effect; (6) the primary construct may be unmeasurable in principle at this timescale — conscious registration of an event is indexed by a motor act that itself requires and produces arousal, so the registration and arousal outcomes are not independent and the measurement perturbs both the EEG and the accelerometer features around onset; (7) the true effect may simply be small — the draft nowhere states what effect size it would consider meaningful; (8) frontal ISF may be poorly expressed or the ISF period may differ in midlife women, invalidating the 50 s working assumption. A risk section that names only risks with tidy analytic answers reads as advocacy rather than assessment.
In the draft — line 338
Insufficient onset-timing precision is the principal analytic risk.
Evidence
Internal: DRAFT.md lines 336-349 contain four risks, all analytic. External, for each omission: button press absent from host protocol and MHT suppressing flashes by 75% (see findings nested-mht-trial-contamination and button-press-absent-from-parent-protocol); "Participants are free to opt out of EEG and smartwatch measurements at any time" and "Target N: 222 participants (accounting for 10% dropout)" (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full); data access via "Data Sharing proposal form" (ibid.); detector N=3 with a bench sensor array (/root/grantreview/Tsiartas2021.txt). Self-inflicted circularity of the registration measure: draft lines 299-302 acknowledge that "a button press requires sufficient arousal" and that registered flashes are "largely a subset of aroused ones", but do not treat the press as a source of motion and arousal artefact in the very signals used to detect and time the flash.
Fix
Rewrite the section as a table with, for each risk: likelihood, impact, mitigation, and the decision point at which it is assessed. Add the eight above. Two deserve explicit mitigations in the design: (a) the button press as artefact — pre-specify that EEG arousal scoring and EDA change-point detection use windows that exclude the press-associated motion transient, and report press latency relative to detected onset as a descriptive outcome; (b) effect size — state the smallest effect worth detecting (e.g. an odds ratio of registration between opposing ISF phases of 1.5) so that a null result is interpretable rather than merely disappointing.
What the refuter said
Strongest case against: risk sections are conventionally analytic, several listed omissions are conditional on unproven premises (MHT contamination presupposes post-randomisation data, which the draft never claims - see F109/F181; the publication embargo is speculative), and item (7) restates F169's missing-MDE. Some padding, then. But the core is right and one item is sharper than anything else in this batch. I confirmed L336-349 contains exactly four risks, all analytic, all of the same shape - a threshold is crossed and an analysis is substituted - with no structural, access or measurement-existence risk anywhere. Item (5) is verified and material: the detector has no ground truth in this sample (no sternal skin conductance is collected), was built on three women with a different sensor array, and its in-sample sensitivity is therefore unknown. Item (6) is the strongest and is checkable against the draft alone: the press is a motor act whose motion enters the detector's own feature set at test time, so the draft's inference at L246-249 - "the detector is never trained on button presses... so it cannot be circular with respect to registration" - does not follow. Not trained on is not the same as not an input. That is a genuine, unflagged circularity between the primary outcome and the event-detection step. Major.
The problem
The proposal's mechanistic appeal is that infraslow noradrenergic activity does two jobs: it sets arousability and it drives the preoptic thermoeffector output that produces the flash. If both are true, the design cannot distinguish gating from common cause. The draft's answer - characterise occurrence phase first and call any structure "a separate finding about flash expression, not conscious gating" - does not work, for a reason the draft never states. Conditioning on flash occurrence, when occurrence is itself a function of ISF phase and of everything else that drives a flash (thermal load, threshold position, magnitude of central sympathetic drive), is conditioning on a common effect. That induces an association between phase and those other causes within the analysed sample: a flash occurring at a phase where noradrenergic drive is low must have been produced by stronger non-phase drive, and those other causes are exactly what determine arousal and reportability. The result is that a phase-registration coefficient is expected under pure common cause with no gating at all. The stronger the occurrence phase-locking, the worse this gets, and the smaller the residual phase variance available to the primary test. The draft treats the two analyses as independent when they are causally entangled, and the entanglement runs in the direction that produces a false positive.
In the draft — line 139-141
it receives noradrenergic input, raising the possibility that a single infraslow rhythm times arousability and thermoeffector drive together
Evidence
The draft's own primary hot-flash source raises the common-cause account and the draft never engages it: Gombert-Labedens et al. 2025 (/root/grantreview/GombertLabedens2025.txt, ~L975-985) - "Hot flashes may themselves trigger arousal from sleep but the strong overlap in timing between many (but not all) hot flash events and awakenings could also reflect a common mechanism within the central nervous system in response to estrogen withdrawal, involving central sympathetic activation or the KNDy neuron network." The two-job architecture is confirmed by Osorio-Forero et al. 2021 (/root/grantreview/OsorioForero2021.txt L269-284): LC phase was targeted "based on whether sigma power started to rise or decline, thereby targeting preferentially high or low arousability periods", and "NA levels rose rapidly before sigma power declined". Draft L295-297 asserts the separation: "phase-dependent occurrence would be a separate finding about flash expression, not conscious gating."
Fix
Say plainly in the Literature Review and in Statistical Analyses that occurrence and registration share a putative cause, and specify the identification strategy rather than asserting separation. Concretely: (a) pre-specify a magnitude- and stage-matched comparison as the confirmatory test, not a sensitivity analysis (see estimand-mismatch-magnitude); (b) pre-specify that if occurrence is phase-locked with resultant vector length above a stated threshold, the primary test is reported as non-identifiable and reframed descriptively; (c) add a discordance test with real inferential value - among flashes matched on peripheral magnitude, onset phase and time of night, does registration still differ by phase? (d) acknowledge that only an intervention that moves phase without moving thermoeffector drive (or vice versa) can settle this, and position the study as generating the constraint rather than the mechanism.
What the refuter said
AGAINST: the bias exists only if occurrence is phase-dependent, and the draft's prerequisite aim exists precisely to find that out first (L294-296) — so the design does contain the diagnostic. The draft also names the common-cause alternative explicitly (L162-169) rather than ignoring it, and expects occurrence to be roughly uniform (L353-354), in which case there is no collider and no bias. SURVIVES, because naming the alternative is not the same as handling it. Conditioning on flash occurrence, when occurrence depends on both phase and on other drivers (thermal load, threshold position, magnitude of central sympathetic drive) that themselves determine arousal and reportability, is conditioning on a common effect: within the analysed sample, phase becomes associated with those other drivers, and a nonzero phase-registration coefficient is expected with no gating of access whatsoever. The draft's answer — that phase-structured occurrence is "a separate finding about flash expression, not conscious gating" (L296-297) — asserts independence between two analyses that are causally entangled, and the entanglement runs toward a false positive. The mechanism is not speculative: the draft's own hot-flash review raises the common-cause account ("could also reflect a common mechanism within the central nervous system ... involving central sympathetic activation or the KNDy neuron network", GombertLabedens2025.txt), and OsorioForero2021.txt confirms the two-job architecture (LC phase targeted "based on whether sigma power started to rise or decline, thereby targeting preferentially high or low arousability periods"). Compounding it, the draft declines to adjust for magnitude on mediation grounds (L306-309), which is exactly the adjustment the collider account would require. Downgraded from fatal to major: the bias is conditional on the prerequisite result, so it is an unaddressed identification threat rather than a certainty.
The problem
Spontaneous non-flash electrodermal activity during sleep is concentrated in exactly the stages the study analyses and in the same half of the night, and is not random with respect to the predictor. Because the ISF is a brain-autonomic rhythm coupled to an infraslow cardiac oscillation, spontaneous EDA excursions may themselves be ISF phase-locked. A false-positive event has no true onset, so its measured phase is determined entirely by the EDA fluctuation that created it — which closes a circular path from sigma dynamics to apparent phase. The draft addresses only missed true events (sensitivity), which is the benign direction.
In the draft — line 249-251
It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent
Evidence
Sano, Picard & Stickgold 2014, Int J Psychophysiol: "More than 80% of the EDA peaks occurred in non-REM sleep, specifically during slow-wave sleep (SWS) and non-REM stage 2 sleep"; "EDA amplitude is higher in SWS than in other sleep stages"; longer storms cluster in the first half of the night; in home data "EDA levels were higher and the skin conductance peaks were larger and more frequent" at the wrist than the palm (https://www.media.mit.edu/publications/quantitative-analysis-of-wrist-electrodermal-activity-during-sleep). Lecci 2017 establishes the ISF is coordinated with an infraslow cardiac oscillation (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853).
Fix
Add a pre-specified negative-control analysis: apply the detector to nights or participants with no reported vasomotor symptoms (or to matched EDA excursions that fail the flash criteria) and test whether those pseudo-events show the same phase distribution. A positive result there invalidates the primary finding. Also add: "Phase-dependent specificity is assessed by testing whether non-flash electrodermal excursions in stable NREM show phase structure."
What the refuter said
AGAINST: magnitude is entirely unquantified, and the draft reports 95.6% specificity — if the false-positive population is small, a phase bias within it is second-order. The draft also characterises occurrence "against stage- and time-matched surrogates" (L295-296), which controls for exactly the stage and first-half-of-night clustering the Sano evidence describes. And "circular" is the wrong word: EDA and sigma share no measurement channel, so this is confounding of event timing by the predictor, not circularity. SURVIVES because the specific mechanism escapes the specific control. I verified Sano, Picard & Stickgold: "more than 80% of the EDA peaks occurred in non-REM sleep, specifically during slow-wave sleep (SWS) and non-REM stage 2 sleep", "EDA amplitude is higher in SWS than in other sleep stages", "Longer EDA storms were more likely to occur in the first two quarters of sleep", and at the wrist specifically "EDA levels were higher and the skin conductance peaks were larger and more frequent when measured on the wrist than when measured on the palm". Stage-and-time-matched surrogates do not control for phase-locking of spontaneous EDA to the ISF, which Lecci's verified sigma-heart-rate coupling makes plausible. The draft's stated guard is one-directional and I confirmed it: L249-251 names "phase-dependent sensitivity" only. The consequence is specific — phase-clustered false positives would appear in the prerequisite aim and be read, per L296-297, as "a separate finding about flash expression" rather than as a detector artefact. Moderate: a paragraph, not a redesign.
The problem
Phase gating and rhythm amplitude are distinct claims. "Does an event arriving at a particular moment in the cycle reach awareness?" is not answered by "is awareness more likely when the rhythm is strong?" The latter is a state or trait question, closer to the norepinephrine-amplitude/memory result the draft cites from Kjaerby, and it cannot address the title question of when the body reaches the mind. The fallback also does not escape the measurement problem it is invoked to solve: estimating the amplitude of a 0.01-0.04 Hz component still requires the same filter on the same short segments, and an amplitude estimated from ~3 cycles carries a large standard error. Reviewers reading the risk section will see the confirmatory hypothesis quietly replaced by a different one.
In the draft — line 341-343
the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis
Evidence
Internal: the primary aim (draft lines 176-180) is "whether the phase of the infraslow sigma fluctuation (ISF) at nocturnal hot flash onset predicts conscious registration"; the fallback predictor is a scalar amplitude with no timing content, so no phase-dependence claim can be tested by it. The draft itself frames amplitude-based inference as a separate literature: "Kjaerby et al. (2022) showed that the memory benefit of sleep depends on the oscillatory amplitude of norepinephrine rather than on its mean level" (lines 112-113).
Fix
Do not present the fallback as testing the same hypothesis. Either declare the amplitude analysis a distinct secondary aim with its own hypothesis, or choose a fallback that preserves the phase claim while tolerating jitter — a binary rising-versus-falling sigma classification at onset, which is exactly the operationalisation the team's own mouse work used and which degrades far more gracefully than continuous phase. Replace with: "...the primary predictor switches to a binary rising/falling sigma classification at onset (as in Osorio-Forero et al., 2021), which preserves the phase-gating hypothesis at coarser resolution; a pre-flash ISF amplitude analysis is reported as a separate secondary aim."
What the refuter said
AGAINST: at the level of the project's framing, both predictors interrogate the same rhythm, and having a pre-registered fallback at all is a virtue reviewers reward. The amplitude route is also not empty — Kjaerby's amplitude result gives it independent standing. SURVIVES on the specific words. The primary aim is "whether the phase of the infraslow sigma fluctuation (ISF) at nocturnal hot flash onset predicts conscious registration" (L177-179), and the title asks "When does the body reach the mind". A scalar amplitude carries no timing content, so it cannot answer a when-question; "testing the same hypothesis" (L342-343) is false as written, and the draft itself treats amplitude-based inference as a distinct literature at L111-113. The honest phrasing is "a related but weaker hypothesis about state rather than timing". Downgraded from major to moderate: this is one overstated clause in a contingency plan, the fallback remains worth having, and note that the fallback's *tolerance* claim is correct (see F92) — only its equivalence claim is not.
The problem
The draft's own listed publication identifies the 222-woman cohort as the Sleeping Through Menopause RCT: a four-arm trial randomising perimenopausal women with insomnia (ISI ≥10, Greene Climacteric ≥13) to menopausal hormone therapy, online CBT/circadian therapy for insomnia, both, or control, with ZMax and EmbracePlus recordings at multiple timepoints. Three claims in the draft are inconsistent with that: (a) participants "take no hormone-altering medication" — the trial excludes prior MHT users but randomises half of them TO MHT, which suppresses vasomotor symptoms; (b) "physically healthy" and "in healthy humans and without intervention" — the population is defined by clinically significant insomnia; (c) the draft never states which timepoint's nights it will use. If only baseline nights are eligible, that must be said, and it caps the design at four nights per participant with no possibility of adding more.
In the draft — line 200-201
They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms.
Evidence
van Baarzel et al. 2026, Sleeping Through Menopause protocol, medRxiv (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): "Expecting a drop-out of 10%, the total number of participants to be recruited is n=222"; arms "1. MHT, 2. CBCTi, 3. MHT + CBCTi and 4. control"; inclusion "Insomnia Severity Index ≥10", "Greene Climacteric Scale score ≥13"; "Z-max triple electrode EEG headbands (Hypnodyne)" worn "at home for four nights during each timepoint"; "Women are excluded from the study if they...already use MHT." Contradicts DRAFT.md line 370: "in healthy humans and without intervention."
Fix
State the host study by name and specify exactly which nights are used. Add to Participants: "Nights are drawn from the pre-randomisation baseline assessment of the Sleeping Through Menopause trial (van Baarzel et al., 2026); no participant is receiving MHT or CBT-I at the time of the recordings used here." Delete "physically healthy" (the cohort is defined by ISI ≥10) and delete "in healthy humans and without intervention" from Expected Outcomes, replacing with "observationally, without experimental stimulation." Note separately that a project nested in an MHT/CBT-I insomnia trial invites the Bial exclusion of "projects involving clinical or experimental models of human disease and therapy" — the framing must foreground the psychophysiological question, not the clinical host.
What the refuter said
STRONGEST DEFENCE, AND IT DEFEATS ALL THREE STATED CONTRADICTIONS: I fetched the protocol and confirmed the trial excludes participants who "already use MHT" and excludes hormonal contraceptive users, so at BASELINE "take no hormone-altering medication" is literally true; four ZMax nights is exactly one timepoint's block ("four consecutive weeknights per timepoint", T0/T1/T2), and T0 precedes randomisation, so "without intervention" holds too. "Physically healthy" is not contradicted by insomnia or climacteric symptoms — the draft itself says participants report sleep disturbances (L197-198) — and the trial excludes the actual physical comorbidities. So (a), (b) and largely (c) dissolve, and a FATAL verdict resting on three dissolved contradictions cannot stand. WHY IT SURVIVES ON WHAT IS LEFT: the identification of the cohort is airtight (222; ZMax four nights; EmbracePlus seven days; perimenopausal; Amsterdam UMC METC; the protocol listed among the applicants' own publications), and grep confirms the draft's body text never mentions the trial, its four arms, its randomisation, its ISI≥10 / GCS≥13 insomnia inclusion, or which timepoint's nights will be used. A proposal that closes by advertising "healthy humans and without intervention" (L369-370) while silently drawing its entire event stream from a four-arm MHT/CBT-I insomnia RCT has a disclosure problem, and given Bial's exclusion of "projects involving clinical or experimental models of human disease and therapy", an assessor is entitled to ask. One sentence naming the trial and the timepoint fixes it. DOWNGRADE: fatal to moderate.
The problem
The draft frames the objective/subjective flash gap as either a measurement artefact or evidence of phase gating, and dismisses the first. It omits a third published account with direct evidence: that perception of flashes tracks sleep fragmentation, i.e. flashes are reported when the sleeper is already awake or arousing, and awareness reflects arousal-dependent encoding rather than a gating mechanism. Bianchi et al. found that objective flashes were NOT associated with increased stage transitions while SELF-REPORTED flashes were, and concluded that sleep disruption increases awareness of and memory for flashes. Since the draft's own analysis expects registered flashes to be "largely a subset of aroused ones", this alternative predicts the same primary result the draft would claim as evidence for phase gating. Without a design feature that separates them, a positive finding is not diagnostic.
In the draft — line 152-154
This discrepancy has been treated as a measurement problem rather than as evidence about how endogenous signals gain access to awareness.
Evidence
Bianchi et al. 2016, J Clin Sleep Med 12(7):1003-1009, https://doi.org/10.5664/jcsm.5936: "Perception of HF, but not objective HF, is linked to increased sleep-stage transitions, suggesting that sleep disruption increases awareness of and memory for nighttime HF events"; "objective HF were not associated with sleep disruption as measured by increased transitions to wake or N1, but self-reported nocturnal HF correlated with an increase from pre- to post-leuprolide in the rate of transitions to wake (p = 0.01), and to N1 (p = 0.008)." Contrast with DRAFT.md lines 300-302: "Because a button press requires sufficient arousal, registered flashes are expected to be largely a subset of aroused ones."
Fix
Name and confront this account in the literature review and give it a pre-registered discriminating test. Add: "A third account holds that report reflects arousal-dependent encoding rather than gating: perceived, but not objectively detected, flashes track sleep fragmentation (Bianchi et al., 2016). Because our registration outcome requires a motor response, this account and phase gating make overlapping predictions. We therefore pre-specify the arousal-conditioned model — registration among aroused flashes only — as the discriminating test, and report the phase distribution of arousals independent of flashes as a nuisance baseline." The arousal-conditioned model is already in the draft (line 302) but is presented as an afterthought; it should be the discriminating analysis.
What the refuter said
Strongest case against: the diagnosticity claim is refuted twice by the draft. First, the draft does contain the design feature the finding says is missing — it models registration among aroused flashes (L301-302) and states at L54-56 that comparing the arousal and registration models "indicates whether phase modulates the cortical response, with registration following, or additionally modulates access to reportable awareness", which is precisely the contrast that separates arousal-driven from access-driven accounts. Second, Bianchi's mechanism is about *recall* of retrospectively self-reported flashes ("sleep disruption increases awareness of and memory for nighttime HF events"), whereas the draft's outcome is a prospective time-stamped press, with the morning diary kept separately as "a complementary measure of recall and reporting behaviour" (L217-219). A prospective press does not require memory, so the arousal-dependent-encoding account does not straightforwardly predict the same primary result. What survives, and I verified it verbatim from the JCSM abstract: "Objective HF were not associated with sleep disruption ... but self-reported nocturnal HF correlated with an increase ... in the rate of transitions to wake (p = 0.01), and to N1 (p = 0.008)." That makes the draft's L152-154 ("treated as a measurement problem") inaccurate — the discrepancy has already been treated as an awareness phenomenon — and Bianchi should be cited and dispatched. Moderate.
The problem
Phase and amplitude are close to orthogonal descriptions of this rhythm, and the mechanism the draft invokes is phase-based. The same sigma amplitude occurs twice per cycle, once on the rising limb and once on the falling limb, and the source literature makes arousability a function of which limb you are on: Osorio-Forero et al. manipulated arousability by targeting rising versus declining sigma, and noradrenaline is high on the falling limb and low on the rising limb at matched sigma levels. So an amplitude predictor averages over the two states the hypothesis is about, and a null on amplitude is compatible with a strong phase effect. Amplitude also carries a second meaning - ISF strength, which Lecci et al. linked to overnight memory - so an amplitude result would be confounded with a trait/night-level variable rather than an event-level state. The contingency is the study's only insurance against its acknowledged principal risk, and it does not cover the risk.
In the draft — line 341-343
the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis
Evidence
/root/grantreview/OsorioForero2021.txt L269-284: phases were targeted "based on whether sigma power started to rise or decline, thereby targeting preferentially high or low arousability periods"; "NA levels rose rapidly before sigma power declined ... NA had already declined when sigma levels started rising, producing a non-symmetrical U-shaped time course". Logical contradiction with draft L341-343: a predictor that is symmetric across the rising and falling limbs cannot test a hypothesis whose content is the difference between those limbs.
Fix
Replace the amplitude fallback with a coarse but phase-preserving fallback. Two options, both timing-tolerant. (a) Sign of the sigma derivative in a window around onset (rising versus declining), a binary predictor that survives several seconds of onset jitter far better than continuous phase and is the variable the rodent work actually manipulated. (b) A cycle-position bin analysis with four bins and onset-error-weighted assignment, or a measurement-error model that propagates the empirical onset-error distribution into the phase estimate rather than discarding phase. Keep amplitude as a genuinely separate, clearly labelled secondary question about ISF strength, and delete the claim that it tests the same hypothesis.
What the refuter said
AGAINST: the finding may be attacking the wrong construct. "Pre-flash ISF amplitude" most naturally reads as the Hilbert instantaneous amplitude of the 0.02 Hz oscillation — modulation depth, in the Kjaerby sense — not the sigma level, and modulation depth is not a within-cycle sinusoid, so "the same amplitude occurs twice per cycle, once on each limb" does not apply to that reading. The draft never says which it means, so the finding's argument holds only under one of two readings. A pre-specified amplitude fallback is also a legitimate, publishable analysis, not worthless insurance. SURVIVES on the conclusion, which holds under either reading. Verified in OsorioForero2021.txt: phases were targeted "based on whether sigma power started to rise or decline, thereby targeting preferentially high or low arousability periods", with "NA levels rose rapidly before sigma power declined ... NA had already declined when sigma levels started rising". Arousability is a function of which limb you are on. An amplitude predictor is symmetric across the limbs (level reading) or is a bout-level trait quantity (modulation-depth reading); neither carries information about WHEN in the cycle the flash arrived, which is the hypothesis at L177-179. So the draft's "testing the same hypothesis" (L341-343) is false. Moderate rather than major: the fallback is a legitimate secondary analysis mislabelled as equivalent, and the correction is to describe it as testing a different, coarser question.
The problem
A bias statistic describes agreement in the average level of sigma power. The study needs fidelity of the *within-night temporal dynamics* of the sigma envelope at 0.02 Hz, and specifically the accuracy of instantaneous phase. A device can reproduce mean band power exactly while distorting or adding sub-minute fluctuations — and the draft never states the units of the 0.005 figure, so the reader cannot even judge its magnitude. The two cited facts (headband mean bias; infraslow sigma modulation exists in humans) do not compose into the claim that this headband can recover infraslow sigma phase from two frontal derivations. The draft half-concedes this in the following sentence ("Frontal expression and achievable phase precision are themselves feasibility outcomes") — which means the load-bearing measurement is admittedly unvalidated, and the contingency at L338-344 becomes the likely path rather than a fallback.
In the draft — line 270-275
Feasibility is supported: sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023), and the infraslow modulation of sigma is established in humans
Evidence
Logical non-sequitur, stated on the draft's own terms: a mean-bias statistic quantifies agreement in level, not in temporal dynamics or phase, so it cannot license inference about 0.02 Hz phase estimation. The draft concedes the gap two sentences later at L275-276: "Frontal expression and achievable phase precision are themselves feasibility outcomes." The relevant human evidence for infraslow sigma modulation cited (Lecci et al. 2017; Lazar et al. 2019) is from polysomnography or full-montage EEG, not from this headband. UNVERIFIED: I did not obtain Esfahani et al. (2023), so I cannot confirm the 0.005 ± 0.012 figure, its units, or whether the paper reports any measure of temporal-dynamics agreement.
Fix
Replace the feasibility claim with a concrete validation step: "Sigma-band power from this headband agrees closely with polysomnography in mean level (Esfahani et al., 2023), but agreement in level does not establish fidelity of the sigma envelope at 0.02 Hz. We therefore quantify infraslow sigma expression and phase agreement between headband and polysomnography in a simultaneous-recording subsample of n participants, and report the phase-error distribution; the primary predictor is fixed only after that error is known." This makes the amplitude contingency an informed choice rather than a rescue.
What the refuter said
Strongest case against: the draft's sentence is a soft support claim ("Feasibility is supported"), and it concedes the residual gap two sentences later - "Frontal expression and achievable phase precision are themselves feasibility outcomes" (L275-276). Accurate sigma-band measurement is a necessary precondition for envelope work even if it is not sufficient, so offering it as partial support is not fabrication. Nonetheless the inferential gap is real and I can state it airtightly. I fetched https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full and confirmed the figure: ZMax mean bias "+0.005 +- 0.012 (unitless; relative bandpower)" for sigma, with sigma defined as 13-16 Hz. A mean relative-bandpower bias quantifies agreement in level. The analysis needs instantaneous phase of the 0.01-0.04 Hz sigma envelope - a device can match mean power exactly while distorting or adding sub-minute fluctuation, and Esfahani reports no envelope- or dynamics-level agreement metric. The finding's secondary point about units is now resolved (unitless relative bandpower), which slightly weakens it. Moderate, not major: the draft does not actually claim the bias validates phase, and it names the missing validation as a feasibility outcome.
The problem
Phase and amplitude answer different questions. The primary hypothesis is that *when in the infraslow cycle* a flash arrives determines whether it is registered — a within-cycle gating claim. Pre-flash amplitude asks whether registration is more likely when the oscillation is *strong*, which is a between-window or between-participant claim about oscillation quality, entirely compatible with no phase gating at all and equally compatible with confounding by sleep depth, spindle rate or time of night. Calling it 'the same hypothesis' misdescribes the contingency and would let a null phase result be reported as a positive finding about gating. Under Bial's rules the Specific Aims as submitted are what the project is assessed against (BIAL_CONTEXT.md line 42-44), so a predictor swap that changes the question is not a minor contingency.
In the draft — line 341-343
the primary predictor switches from ISF phase to pre-flash ISF amplitude, which tolerates far more timing error while testing the same hypothesis
Evidence
Logical contradiction internal to the draft: the primary aim is stated at lines 178-180 as testing "whether the phase of the infraslow sigma fluctuation (ISF) at nocturnal hot flash onset predicts conscious registration", and the confirmatory test at lines 291-292 is "likelihood-ratio test of the joint sin and cos term" — a test of preferred phase. An amplitude predictor enters no sin/cos term and cannot test preferred phase.
Fix
Rewrite as: 'If onset precision is insufficient, the phase hypothesis is not testable and we report that as the feasibility result. The pre-registered secondary predictor is then pre-flash infraslow amplitude, which asks a related but distinct question — whether registration is more likely when the rhythm is strongly expressed — and we will report it as such, not as a test of phase gating.'
What the refuter said
Strongest case against: the swap is disclosed, pre-registered, and triggered only when the phase hypothesis has become untestable (L338-343, conditional on pre-outcome dispersion exceeding a pre-registered threshold), which is orthodox pre-registration practice rather than a bait-and-switch; and the finding overstates in calling pre-flash amplitude "a between-window or between-participant claim" — the Hilbert envelope is within-participant and time-varying, so it can enter the same GLMM as a within-cluster predictor. But the core is an airtight logical contradiction and it stands. The confirmatory test is specified at L291-292 as a "likelihood-ratio test of the joint sin and cos term", i.e. a 2-df test for a preferred phase; an amplitude regressor enters no sin/cos term and cannot test preferred phase. "Testing the same hypothesis" (L342-343) is therefore false: within-cycle timing and oscillation strength are different claims, and an amplitude effect is fully compatible with no phase gating and with confounding by sleep depth or spindle rate. The reviewers' own contingency simulation reaches the same conclusion independently ("the BOUT-LEVEL STRENGTH of the ISF modulation ... tests a different hypothesis"). The defect is the sentence's claim of equivalence, not the existence of a fallback — which is why moderate, not major.
The problem
Decompose the primary quantity: P(registered | phase) = P(aroused | phase) x P(registered | aroused, phase). The draft's stated novelty is that infraslow phase gates "access to reportable awareness" over and above arousability (l.53-56, l.179-181). But phase-dependent arousability is precisely what Lecci et al. (2017) already established (l.100-103). If the first factor carries the whole effect, the confirmatory result is a replication of known phase-dependent arousability using a spontaneous rather than delivered stimulus — informative, but not the advertised claim. The only term that isolates the advertised claim is the second factor, P(registered | aroused, phase), i.e. "registration additionally modelled among aroused flashes" — and that model is explicitly non-confirmatory (l.310-311: "The primary registration model is confirmatory; all else is exploratory"). Worse, that model conditions on arousal, a post-onset variable on the causal path from phase to registration — the exact adjustment the draft forbids two paragraphs later on principled grounds for magnitude (l.306-308). So the proposal simultaneously (a) makes confirmatory a test that cannot demonstrate its central claim, and (b) makes exploratory a test whose logic it elsewhere declares invalid. There is no stated formal comparison of the two models either: "Comparing the two indicates whether phase modulates the cortical response, with registration following, or additionally modulates access to reportable awareness" (l.53-56) names no estimand, no test statistic, and no decision rule; informally comparing two sin/cos coefficient pairs fitted to nested outcomes is not a test of the additional-modulation hypothesis.
In the draft — line 299-302
Because a button press requires sufficient arousal, registered flashes are expected to be largely a subset of aroused ones; the joint distribution is reported and registration additionally modelled among aroused flashes.
Evidence
Draft l.299-302 ("registered flashes are expected to be largely a subset of aroused ones") versus l.310-311 ("The primary registration model is confirmatory; all else is exploratory") versus l.306-308 ("if ISF phase modulates the size of the autonomic response, adjusting for it would bias the total effect" — the same mediator objection applies verbatim to conditioning on arousal). Numerically, from the draft's own figures: 1.5/3.5 = 42.9% of flashes registered (l.150-152) and 80.2% disturb sleep (l.149-151), so registration among disturbing flashes is roughly 1.5/(3.5x0.802) = 53%; the registration outcome is close to a relabelled arousal outcome plus a coin flip.
Fix
Make the arousal-conditioned analysis the confirmatory test of the central claim, and state the estimand formally. Concretely: pre-register a single joint model with a formal decomposition — a mediation analysis with arousal as mediator, reporting the natural direct effect of phase on registration (phase -> registration not through arousal) as the confirmatory quantity and the indirect effect (phase -> arousal -> registration) as the replication quantity, with the assumptions each requires stated. Replace l.299-302 with: "Because a button press requires sufficient arousal, the scientifically distinctive quantity is the effect of phase on registration that does not operate through arousal. The confirmatory test is therefore the natural direct effect of ISF phase on registration in a mediation model with cortical arousal as mediator; the indirect (arousal-mediated) path is reported as a replication of phase-dependent arousability. A total-effect model on registration is reported alongside." If mediation assumptions are judged untenable, say so and make the primary test P(registered | aroused, phase), pre-specifying the expected number of aroused flashes.
What the refuter said
STRONGEST DEFENCE, PARTLY SUCCESSFUL: the mediator sub-argument is wrong. The draft does NOT condition on arousal in its primary model; it fits registration among ALL detected flashes unadjusted (the total effect, L280-283, L306-309) and reports the aroused-only model as an ADDITIONAL analysis explicitly labelled exploratory (L299-302, L310-311). Total-effect primary plus mediator-conditional secondary is textbook mediation practice, not a contradiction, so "the proposal simultaneously makes exploratory a test whose logic it elsewhere declares invalid" is a false dilemma. "Indistinguishable from a replication" is also over-read: under an arousal-only truth the primary registration test rejects at only 0.077 (OR=2) rising to 0.194 (OR=5) against α=0.05 — mildly inflated, not "firing". WHY THE CORE SURVIVES: the decomposition P(reg|phase) = P(aroused|phase)·P(reg|aroused,phase) is exact, and the draft's own words concede the near-nesting ("registered flashes are expected to be largely a subset of aroused ones", L299-302). The advertised novelty is additional modulation of "access to reportable awareness" (L53-56, L179-181), and the only term isolating it is the second factor — which the draft demotes to exploratory (L310-311) while making confirmatory the term that does not isolate the claim. The disambiguating test carries roughly 8% of the arousal test's information, so it is also the least powered of the three. And the Summary's comparison sentence (L54-56) names no estimand, statistic or decision rule. That is a genuine priority-of-inference defect: the confirmatory result, if positive, will not by itself establish what the proposal advertises. DOWNGRADE: fatal to moderate — the design is not broken, its confirmatory/exploratory labels are mis-assigned and the disambiguator is underpowered.
The problem
The two sentences contradict each other, and the first is the wrong argument. Circularity here does not require the detector to have seen button presses. It requires only that detection probability depend on the predictor (ISF phase) and on something that also drives the outcome (autonomic response magnitude, arousal proneness). The draft concedes exactly that in the following sentence. Since the primary model is fitted only to detected flashes (l.283: "Both are fitted to every detected flash"; l.298: "Conditional on a flash occurring"), detection is a collider on the paths phase -> magnitude -> detection and (unmeasured arousability/reporting propensity) -> magnitude -> detection. Conditioning on a collider induces association between phase and reporting propensity among detected flashes even when no causal phase -> registration effect exists, and the sign of the induced bias is not predictable from the information given. This is a threat to the primary result of the same class as the effect being sought, and it is not addressed. The proposed remedy does not remove it: "a pre-specified analysis repeats detection from skin conductance and temperature alone" (l.251-252). (a) Electrodermal activity and skin temperature are themselves sympathetically driven; the draft's own literature review states the sigma ISF "is coordinated with an infraslow cardiac rhythm, indexing a brain-autonomic rather than a purely cortical state" (l.103-105) and that it tracks LC activity (l.107), so removing the cardiac feature removes one pathway of phase-dependent sensitivity, not the phenomenon. (b) Dropping features voids the cited validation: the >90%/95.6% operating point (l.244-245) belongs to the multi-sensor configuration, so the fallback detector has unknown sensitivity and specificity. (c) A sensitivity analysis that yields a different estimate demonstrates that bias exists; it neither bounds nor corrects it. No detection-probability model, no bias bounds, and no phase-independent gold standard are proposed.
In the draft — line 247-251
the detector is never trained on button presses, which are themselves an outcome, so it cannot be circular with respect to registration. It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent
Evidence
Internal contradiction within one paragraph: draft l.247-248 ("so it cannot be circular with respect to registration") versus l.249-251 ("It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent"). Draft l.103-105 supplies the mechanism by which the fallback signals are also phase-coupled. Draft l.283 and l.298 establish that the analysis conditions on detection.
Fix
Delete the claim of non-circularity and replace the argument. Suggested replacement for l.246-254: "Labelling rests on objective physiology rather than self-report, so the detector cannot inherit the outcome by training. It can, however, inherit the predictor: because detection depends on autonomic response magnitude and the infraslow rhythm is brain-autonomic, detection probability may itself be phase-dependent. Because the primary model conditions on detection, this makes detection a collider and can induce a phase-registration association in the absence of any causal effect. We therefore (i) estimate detector sensitivity as a function of ISF phase in a validation subsample recorded with concurrent sternal skin conductance as a phase-independent reference standard, (ii) carry that sensitivity function into a quantitative bias analysis giving bounds on the primary estimate, and (iii) report the analysis repeated with a skin-conductance-and-temperature detector, whose own sensitivity and specificity we re-establish in the validation subsample." Without a phase-independent reference standard in some subsample, this bias cannot be bounded and the proposal should say so.
What the refuter said
Strongest case against: the alleged self-contradiction is a misreading of precision as inconsistency. L247-248 claims no circularity *with respect to registration* (the detector never sees the outcome) — true, and correctly argued. L249-251 raises phase-dependent detection *sensitivity* — a different problem, which the draft names unprompted and then pre-specifies an analysis for. Naming two distinct threats is not contradicting yourself. Also, the demanded remedy is not obtainable: no phase-independent gold standard for flash onset exists in an ambulatory home study (the field standard is sternal skin conductance at 30-s resolution), so "no phase-independent gold standard is proposed" asks for something unavailable. Two sub-points do survive and are verified. First, dropping the PPG features voids the quoted operating point: Tsiartas reports the >90%/95.6% figure for the full multi-sensor set, notes skin-conductance features alone "explain most of the variance (~65%)", and reports SC+ as "+10.7% in sensitivity" over SC — so an SC+T detector has unknown sensitivity and specificity, and the draft carries the multi-sensor number over as if it applied. Second, a sensitivity analysis that changes the estimate demonstrates bias without bounding it. Moderate.
The problem
Three problems. (1) The rule cited is correct for estimating a total effect, but the total effect is not the quantity of scientific interest here, by the draft's own account: "The study therefore separates gating of peripheral flash expression from gating of cortical and conscious registration" (l.75-76) and the aims section promises to separate "gating of the cortical response from gating of registration" (l.180-181). The total effect of phase on registration deliberately sums the peripheral route (phase -> larger sympathetic response -> more salient event -> registered) and the central route (phase -> gating of access -> registered). A positive total effect driven entirely by the peripheral route would satisfy the confirmatory test while contradicting the proposal's headline interpretation that an infraslow rhythm "sets when an internal event becomes an experience" (l.366-367). The analysis that answers the stated aim — the direct effect of phase holding magnitude fixed — is the one relegated to "sensitivity analysis only". (2) Magnitude is not only a mediator; because detection depends on magnitude and the sample is restricted to detected flashes, magnitude has already been conditioned on (truncated at the detector's threshold) before any model is fitted. The claim that the primary model "does not adjust for magnitude" is therefore not achievable by omitting the covariate: the total effect in the source population is not recovered by this design regardless. (3) The draft applies the mediator rule inconsistently: it refuses to adjust for magnitude on mediator grounds while conditioning on arousal — equally a post-onset mediator, and by the draft's own reasoning nearly a necessary condition for the outcome — in the analysis at l.301-302.
In the draft — line 306-310
Flash magnitude is a mediator, not a confounder: if ISF phase modulates the size of the autonomic response, adjusting for it would bias the total effect of phase on registration. The primary model therefore does not adjust for magnitude; a magnitude-adjusted model is a sensitivity analysis only.
Evidence
Draft l.306-310 against draft l.75-76 ("separates gating of peripheral flash expression from gating of cortical and conscious registration") and l.180-181 ("separating gating of the cortical response from gating of registration"). Conditioning on magnitude via detection follows from l.283 ("Both are fitted to every detected flash") plus l.242-245 (detection is a classifier on autonomic magnitude features). Inconsistent application: l.301-302 ("registration additionally modelled among aroused flashes").
Fix
State the estimand explicitly and align it with the aim. Replace l.305-311 with: "Flash magnitude lies on the causal path from phase to registration, so a magnitude-adjusted model does not estimate the total effect of phase. Because the aim is to separate peripheral expression from central gating, both quantities are pre-specified: the total effect of phase (unadjusted) and the direct effect of phase holding magnitude at its observed distribution, estimated in a single mediation model with magnitude as mediator. The direct effect is the confirmatory quantity for central gating; the total effect is reported alongside. Because the analysis sample is restricted to detected flashes, magnitude is already truncated at the detector's operating threshold, so neither quantity generalises to undetected flashes; the resulting restriction is stated as a limitation and probed in the bias analysis above."
What the refuter said
Strongest case against, taken point by point. Point (3) is simply wrong: L301-302 says registration is "additionally modelled among aroused flashes" - an additional, secondary model, the same status the draft assigns the magnitude-adjusted model ("a sensitivity analysis only", L309-310). There is no inconsistency; both mediator-conditioned analyses are secondary. Point (1) is an estimand preference dressed as an error. Pre-registering the well-identified total effect as confirmatory and the assumption-heavy direct effect (which requires no unmeasured mediator-outcome confounding) as sensitivity is defensible, indeed conservative, practice. And the draft's peripheral/central separation is delivered by a different device than the finding assumes: "Occurrence and consequence are analysed separately" (L294) - the prerequisite aim tests phase-structured expression, the primary tests registration among occurring flashes, and the arousal-vs-registration comparison separates cortical response from access (L180-181). Magnitude adjustment was never the separation mechanism. Point (2) survives and is airtight on the draft's own text: detection is a classifier on autonomic magnitude features (L241-245) and every model is "fitted to every detected flash" (L282-283), so the sample is already truncated on the mediator before any covariate decision. The draft's stated justification - that omitting the covariate preserves the total effect - therefore does not hold as written. That is a real internal-logic defect and an interpretive risk (a positive total effect could be entirely peripheral while being written up as "an infraslow rhythm sets when an internal event becomes an experience", L366-367), but with two of three prongs failing, major is not sustainable.
The problem
Section 'Study design' justifies four nights on the grounds that power depends on event count; the Summary and the Power section assert the opposite, that 'the binding constraint is not event count' (line 79-80) and 'The binding constraint is measurement precision rather than event count' (line 315). Both cannot be true, and the simulation shows both constraints bind simultaneously and multiplicatively: required N scales as 1/lambda^2, so precision sets the multiplier on an event count that is already an order of magnitude short.
In the draft — line 191-194
The unit of analysis is the flash, not the participant: power depends on event count, hence multiple nights.
Evidence
Required analysable flashes for 80% power at OR=2, press rate 0.15 (/root/grantreview/sim/08_yield_v2.py, 02_power_vs_N.py): 1,258 at zero jitter, 1,867 at 5 s, 14,837 at 12.5 s. Central yield 347. Neither constraint is slack at any jitter level: even at perfect timing the central yield gives power 0.19.
Fix
Delete the 'not event count' framing in the Summary and Power sections and say instead: "Two constraints bind multiplicatively. Power scales with the number of analysable flashes and with the square of the timing attenuation factor, so the required event count is the zero-jitter requirement divided by that factor squared. Both are reported below."
What the refuter said
STRONGEST DEFENCE, PARTLY SUCCESSFUL: "the binding constraint is X rather than Y" is a comparative claim about which constraint dominates, not a denial that Y matters. On that reading L191-192 ("power depends on event count, hence multiple nights") and L315 ("The binding constraint is measurement precision rather than event count") are perfectly compatible — both bind, precision binds harder. "Both cannot be true" is therefore an over-reading of ordinary English, and the finding's headline is the weakest available framing. WHY IT SURVIVES ON THE SUBSTANCE: the draft does not stop at the comparative claim. L317-318 asserts "power ceasing to improve with additional events once jitter exceeds a quarter of the cycle", which is false — I verified analytically that jitter attenuates the first-harmonic amplitude by λ = exp(−σ_φ²/2) but leaves the 2-df noncentrality n·p̄(1−p̄)(λA)²/2 strictly increasing in n, so required N inflates by 1/λ² (11.8x at a quarter cycle) and power never plateaus. And the reviewers' power code (sim/common.py, 02_power_vs_N.py) implements exactly this standard wrapped-normal attenuation and 2-df noncentrality, so I can vouch for it. The multiplicative point is also right in direction: 137 x 4 x 3.5 ≈ 1,900 detected events before ANY exclusion, against ~1,000 needed at zero jitter for OR=2 at a 20% base rate — no slack at all once the exclusion cascade runs. Caveat kept: the "central yield 347 / power 0.19" figures are the reviewer's own yield model, not the draft's. DOWNGRADE: major to moderate — largely a duplicate of F28's stronger version, with an over-read headline.
The problem
The title asks when the body reaches the mind; the measure is whether a woman wakes up enough to press a button and attributes the awakening to a flash. The press is a conjunction of at least five things: interoceptive signal reaching cortex, arousal sufficient for volitional action, correct causal attribution of the bodily state to a flash, motor execution, and compliance. Only the first is the construct of interest, and it is the one component the design cannot isolate. Take the frameworks in turn. Global neuronal workspace makes reportability constitutive of access, but it also holds that long-range ignition is what NREM sleep suppresses; on GNW the press indexes the state transition to wake, so a phase effect is re-described as phase modulating arousal threshold, which is the already-hypothesised finding and says nothing new about access. Higher-order theories would accept the press as evidence of a higher-order representation but insist that first-order interoceptive representation can be present without it, so a null on the press is uninformative; HOT would also demand a graded confidence or metacognitive measure, and the draft has only a binary press. Recurrent processing theory (Lamme) predicts that local recurrence in interoceptive cortex suffices for experience without report, so on RPT most flashes are felt and never pressed, and the press-versus-no-press contrast tracks access and attention, not the arrival of the body in the mind. Integrated information theory would look at post-onset cortical dynamics, not a motor act. So no framework predicts the proposed effect as an effect on consciousness; each reinterprets it as arousal-threshold modulation. Finally, the no-report critique is the field's standard objection to exactly this design, and the draft's own cited paper answers it with heartbeat-evoked potentials - a no-report interoceptive measure - which the draft does not adopt.
In the draft — line 66-67
Conscious registration is indexed by a time-stamped button press made as soon as the participant thinks she is having a flash
Evidence
Draft L216-218: "participants press a wristband button as soon as they think they are having a flash; this time-stamped press is a prospective measure of conscious registration" - "thinks she is having a flash" is a causal attribution, not a detection. The no-report critique: Tsuchiya, Wilke, Frassle & Lamme, "No-Report Paradigms: Extracting the True Neural Correlates of Consciousness", Trends in Cognitive Sciences 2015, https://www.cell.com/trends/cognitive-sciences/abstract/S1364-6613(15)00252-1 . A no-report interoceptive measure exists in sleep and is cited by the draft itself: Cataldi et al. 2026, Current Biology, https://www.cell.com/current-biology/fulltext/S0960-9822(26)00889-4 - heartbeat-evoked potentials "were preserved across REM microstates and were enhanced relative to wakefulness" while auditory responses declined. Covert task-relevant processing without behavioural report during NREM sleep is established: Kouider et al., "Inducing Task-Relevant Responses to Speech in the Sleeping Brain", Current Biology 2014, https://www.cell.com/current-biology/fulltext/S0960-9822(14)00994-4 . The draft's own text concedes the arousal dependence at L299-300: "Because a button press requires sufficient arousal".
Fix
Do two things. (1) Stop calling the press an index of conscious access. Rename the primary outcome "prospective self-report of a flash" and state that it is a compound of arousal, attribution and response, with the theoretical caveat spelled out in one sentence: "Because report during sleep requires a transition to wake, this measure cannot distinguish gating of access from gating of arousal; the arousal model is included precisely to bound that." (2) Add a no-report electrophysiological outcome to the same events - an event-related response time-locked to the electrodermal inflection, and a flash-locked heartbeat-evoked-potential change, following Cataldi et al. 2026. That gives a measure of the signal reaching cortex that does not require waking, converts the press from the sole outcome into one rung of a three-rung ladder (cortical response, arousal, report), and is the only version of this project that can honestly claim to address interoceptive access rather than arousability.
What the refuter said
Strongest case against: this is largely a philosophical over-read of a draft that has already made the moves it demands. The draft commits explicitly to reportability rather than phenomenality — "how a signal arising in the body reaches reportable conscious experience" (L175-176) — so the global-workspace objection is not a rebuttal but the draft's own frame; it concedes the arousal conjunction in the finding's own terms at L299-300 and then builds exactly the decomposition that separates them (L301-302 registration among aroused flashes; L54-56 comparing the arousal and registration models). The recurrent-processing objection that felt-but-unreported events exist would invalidate the entire access-consciousness literature, not this design. And the proposed no-report alternative is not available here: heartbeat evoked potentials need ECG and dense EEG, not two frontal derivations and wrist PPG at home, and they index cardiac rather than thermal interoception — a limitation the draft itself flags at L125-127 ("whether findings for cardiac signals generalise to thermal interoception is a further open question"). The four-theory survey is rhetorical: the draft nowhere claims to adjudicate between theories. What survives is one sharp observation: "thinks she is having a flash" is a causal attribution, and no false-press or misattribution analysis is proposed — which is F126's territory. Minor.
The problem
The claim is repeated at line 283 ('Both are fitted to every detected flash') but contradicted by the eligibility chain: analyses are 'restricted to stable (and non-artifactual) NREM' (line 269-270), which the contingencies define as bouts of three infraslow cycles (line 344-345), and the yield calculation applies exclusions for 'EEG night, NREM, stable NREM, estimation window, artefact-free recording, reliable onset timing' (lines 204-206). The literature makes this restriction bite hard rather than incidentally: roughly half to seven-tenths of nocturnal flashes coincide with arousal or awakening, about one in ten occurs when the woman is already awake, and flashes cluster in wake and N1 while being less likely in slow-wave sleep. The eligible set is therefore biased toward the quiet flashes precisely as the draft intends — but that is a selection on a variable downstream of the outcome, and the draft nowhere states an expected surviving event count.
In the draft — line 52-54
Every detected flash enters both analyses, including the quiet ones passing without arousal or report, events an experimenter-delivered stimulus can never provide.
Evidence
Internal contradiction between DRAFT.md lines 52-54 / 283 and lines 204-206 / 269-270 / 344-345. On the base rates: Baker et al. 2019 (Sleep 42(11):zsz175) reports 51.1% with arousal/awakening, 28.6% in undisturbed sleep, and "HFs were less likely to wake a woman in rapid-eye-movement and slow-wave sleep"; Gombert-Labedens et al. 2025 (/root/grantreview/GombertLabedens2025.txt lines 921-924) reports for the n=34 study "70% of hot flashes were associated with an arousal from sleep" and "the remaining hot flashes occurred when the participant had already been awake for at least one minute".
Fix
Replace 'Every detected flash' with 'Every flash detected within an eligible stable NREM bout', and add the arithmetic: give the expected number of eligible flashes per participant after each exclusion in the sequential yield chain, with the base rates and their sources, and state the minimum eligible event count below which the study is not run.
What the refuter said
Strongest case against: read in context there is no contradiction. L282-283 runs "Both are fitted to every detected flash, so events passing without arousal or report contribute" — the trailing clause states the point of the sentence, which is that nothing is excluded *for lacking arousal or a press*, contrasting with designs that count only reported or arousing flashes. That is fully compatible with technical exclusions, and the draft itself enumerates them openly two pages earlier (L204-206: "EEG night, NREM, stable NREM, estimation window, artefact-free recording, reliable onset timing"). A reviewer is not misled; the phrasing is loose, not contradictory. The base rates are verified — Bianchi et al. 2016 (JCSM abstract): "Most occurred during wake (51.0%) and stage N1 (18.8%)"; GombertLabedens2025 L921-924: "70% of hot flashes were associated with an arousal from sleep" — so the NREM requirement does bite hard. But the consequence (no stated surviving event count) is F114's and F96's finding, and the selection-on-a-variable-downstream-of-the-outcome claim inherits F18's refuted premise about fragility. Minor.
The problem
Two sentences earlier the draft cites human evidence that involves no stimulus at all: "in humans the fragility phase is associated with spontaneous arousals and transitions to lighter sleep" (l.102-103), repeated in the summary as "in humans, fragility is marked by absent sleep spindles and spontaneous arousals" (l.29-30), and it cites a memory-retention association (l.105-106) and an oscillation-amplitude/memory result (l.111-113) that likewise involve no delivered stimulus. So the literature does not rest entirely on delivered stimuli, and — more damagingly for the novelty claim — spontaneous arousals are endogenous events already shown to be phase-structured. The genuinely novel elements are narrower: an interoceptive, thermoregulatory endogenous event, and conscious report as the outcome. The over-claim is repeated at l.31-32 and l.363-364. A related over-claim appears in the summary: "Every detected flash enters both analyses, including the quiet ones passing without arousal or report, events an experimenter-delivered stimulus can never provide" (l.51-54). This is false: sub-threshold tones that produce neither arousal nor report are trivially deliverable and are standard in arousal-threshold designs. The distinguishing property of the flash is endogeneity, not the existence of quiet events. Finally, if fragility is partly identified by spontaneous arousals (l.29-30), then "a flash-related arousal occurring in fragility" is partly definitional rather than predictive, which compounds the non-causal filtering problem.
In the draft — line 115-117
This literature rests entirely on stimuli delivered by an experimenter, a design that deliberately randomises stimulus timing against brain state.
Evidence
Draft l.115-117 ("rests entirely on stimuli delivered by an experimenter") against l.102-103 and l.29-30 (spontaneous arousals in humans) and l.105-106, l.111-113 (memory findings involving no stimulus). Draft l.51-54 ("events an experimenter-delivered stimulus can never provide") is contradicted by the existence of sub-threshold stimuli, which the arousal-threshold literature the draft cites necessarily uses to establish a threshold.
Fix
Narrow the gap statement so it survives contact with a specialist reviewer: "The stimulus-response arm of this literature rests on experimenter-delivered stimuli whose timing is randomised against brain state. Endogenous events are not absent from it — spontaneous arousals and spindle-poor substates are themselves phase-structured — but no endogenous event has been used to ask whether infraslow phase gates access to reportable awareness, and no interoceptive event of any kind has been tested within NREM sleep." And replace l.52-54 with: "Every detected flash enters both analyses, including the quiet ones passing without arousal or report; unlike a delivered stimulus, these arrive on the body's own schedule, so their timing relative to the infraslow cycle is set by the organism rather than the experimenter."
What the refuter said
Strongest case against: read charitably, "this literature" at L115-117 means the literature on state-dependent gating *of delivered signals*, and the draft's narrower restatement at L363-364 ("State-dependent gating has so far been shown only for stimuli an experimenter chose to deliver") is accurate. Spontaneous arousals are not signals being gated - they are the arousal itself - so citing them at L102-103 is not self-contradiction on the gating claim. That defence rescues L115-117 as an imprecision rather than an error. It does not rescue L51-54, which is false as written and which I checked in context: "including the quiet ones passing without arousal or report, events an experimenter-delivered stimulus can never provide". Sub-threshold stimuli producing neither arousal nor report are not merely possible but necessary to the arousal-threshold designs the draft builds on - you cannot establish a threshold without them. The distinguishing property of the flash is endogeneity, and the draft has that argument available; "can never provide" is a rhetorical over-reach in the most-read section. The third prong is too vaguely stated to confirm, but it gestures at something sharper that the draft does not address: L267-269 obtains phase by "zero-phase band-pass filtering... and the Hilbert transform", so the phase estimate at onset depends on post-onset sigma, which an arousal suppresses - an outcome-to-predictor leak in the supporting model. Moderate: two verified over-claims plus a narrowed novelty claim, none of which sinks the design.
Whether each source exists, and whether it says what the draft says it says.
The problem
The two papers cited as feasibility support both report the opposite of what the design needs. The ISF has never been demonstrated at frontal derivations, and the only montage ZMax provides is frontal (F7-Fpz, F8-Fpz). The draft's own hedge concedes this, which makes the confirmatory primary hypothesis conditional on an untested measurement assumption.
In the draft — line 273-276
the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019). Frontal expression and achievable phase precision are themselves feasibility outcomes.
Evidence
Lecci 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853): the 0.02 Hz oscillation was "maximum over parietal derivations for power in both the sigma and the FSP band and declined toward anterior central and frontal areas"; frontal midline weaker than central/parietal (Friedman across midline electrodes, fast-spindle band, P = 3.5e-8); core human analysis at "C3 and C4 electrode sites". Lazar 2019 abstract (https://research-portal.uea.ac.uk/en/publications/infraslow-oscillations-in-human-sleep-spindle-activity): "These ISO are most prominent in the high sigma band and over the centro-parieto-occipital regions"; "ISO in sleep spindles are most prominent in the centro-parieto-occipital regions". ZMax montage from Esfahani 2023: "two frontal EEG channels (F7-Fpz, F8-Fpz)" (https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full).
Fix
Do not cite Lecci and Lazar as frontal feasibility support; state plainly that they show the opposite. Either (a) add a within-subject validation arm with a central/parietal derivation (a third ZMax channel or a small concurrent PSG subsample) and make the confirmatory aim conditional on passing it, or (b) demote frontal ISF phase to an exploratory aim. Replace the sentence with: "The infraslow sigma fluctuation is established in humans, but is maximal centro-parietally and declines toward frontal derivations (Lecci et al., 2017; Lazar et al., 2019); it has not been demonstrated at the F7/F8-Fpz montage used here. A validation subsample with concurrent central derivations therefore gates the primary analysis."
What the refuter said
Strongest case against: the finding's central assertion - "The ISF has never been demonstrated at frontal derivations" - is false, and I verified that from both papers it cites. Lecci et al. 2017 extended the analysis to nine sites "F3, FZ, F4, C3, CZ, C4, P3, PZ, and P4" and reported that although amplitude declined from posterior to anterior, "the oscillatory signature persisted across the anteroposterior axis". Lazar et al. 2019 did not fail to find frontopolar ISO; they measured and ranked it - "Frontopolar region had significantly (adjusted P < 0.05) lower integrated EIP compared to all other brain regions except the temporal brain region". So the signal is attenuated where ZMax records, not absent, and a fatal grade is not supportable. The substantive core is nonetheless verified and important: both human sources put the ISF maximum well away from the recording site (Lecci: "maximum over parietal derivations... declined toward anterior central and frontal areas"), the ZMax montage is F7-Fpz/F8-Fpz (Esfahani, confirmed), and the draft cites these two papers as feasibility support without disclosing the gradient. F133 states the same finding accurately and with a third independent source, so this one is the weaker duplicate.
The problem
Two errors compound. First, the operative variable in the cited work is the sign of the sigma derivative (rising versus declining), not the level (high versus low). Osorio-Forero et al. 2021 targeted "high or low arousability periods" by whether sigma was rising or declining, and showed noradrenaline peaks BEFORE sigma declines and has "already declined when sigma levels started rising". So the high-noradrenaline, high-arousability window sits at and just after the sigma peak, not in the trough. The draft's mapping of fragility onto low sigma power is a substantive misreading of its own team's paper. Second, the human evidence points the other way. Dimitriades, Osorio-Forero, Fattinger et al. (2026), with the draft's own rodent collaborator as second author, characterised the human infraslow sigma fluctuation in 154 people and found arousal markers concentrated at the spindle-rich peak. The draft's Expected Outcomes commit to the opposite direction. A reviewer who knows this paper - and the second author is on this team - will read the prediction as evidence the proposers have not read the human literature on their own predictor. Note the sin/cos parameterisation is direction-agnostic and survives; the mechanistic narrative and the fallback (see amplitude-fallback) do not.
In the draft — line 24-27
alternating roughly every 25 s between a continuity phase of high sigma power and a fragility phase of low sigma power. During continuity the sleeper is comparatively insulated from sensory input; during fragility the same stimulus more often produces an arousal (Lecci et al., 2017).
Evidence
Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853) states of the human data: "the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained" - so the draft's human arousability premise is not established by the paper it cites; the same source gives the sigma band as 10-15 Hz, not the draft's 11-16 Hz. Osorio-Forero et al. 2021 (/root/grantreview/OsorioForero2021.txt L269-284): "NA levels rose rapidly before sigma power declined ... Across animals, NA had already declined when sigma levels started rising, producing a non-symmetrical U-shaped time course." Dimitriades et al. 2026, Sci Rep, doi 10.1038/s41598-026-58423-z, abstract: "electrophysiological markers of arousal and memory reactivation are organized within the spindle-rich ISFS peak"; preprint full text (https://www.biorxiv.org/content/10.1101/2024.11.06.620875v1): "The highest percentage of microarousals occurred around the peak of the ISFS ... they clustered around the peak of the ISFS in all age groups", and their summary of the rodent result: "around the peak of the ISFS and in the subsequent decreasing phase of the ISFS, external tones were more likely to induce wakefulness or microarousals". Draft L355-357 predicts the reverse: "flashes arriving in fragility (low sigma power) are predicted to be more likely registered, and followed by an arousal, than those in continuity."
Fix
Rewrite the phase definition in terms of the cycle position used by the source work - rising versus declining sigma, with the peak and early decline as the high-arousability window - and drop the level-based "high sigma / low sigma" gloss. Cite Dimitriades et al. 2026 and state the human prediction it implies: registration and arousal should be more probable for flashes arriving near the ISF peak and early decline. Replace L355-357 with: "Following the human evidence that microarousals cluster around the peak of the infraslow sigma fluctuation (Dimitriades et al., 2026) and the rodent finding that noradrenaline rises before sigma declines (Osorio-Forero et al., 2021), we predict that flashes arriving near the sigma peak and on the declining limb are more likely to be registered than those arriving on the rising limb." Correct the sigma band to the cited value and reconcile it with the arousal band. Also delete or qualify the unsupported human claim at L101-102.
What the refuter said
AGAINST: the finding concedes the sin/cos parameterisation survives, and the draft pre-empts the direction question explicitly at L358-361 ("A reciprocal pattern is equally informative ... a direction predicted rather than defining"), so the primary test is not invalidated. "Fatal" is therefore not sustainable. But the construct error is real and I verified it against three primary sources. Lecci et al. 2017 defines fragility by the DERIVATIVE, not the level: "the progression into the declining sigma power period thus reflects the entry into a period of sleep fragility", and its worked examples put the aroused case at the peak — "In a wake-up trial ... sigma power was at its maximum before noise onset, such that noise exposure fell within a phase of declining power. In contrast, for a sleep-through trial ... sigma power had just exited the trough." OsorioForero2021.txt agrees on the physiology: "NA levels rose rapidly before sigma power declined ... NA had already declined when sigma levels started rising, producing a non-symmetrical U-shaped time course." And the human data are now published and point the other way: Dimitriades et al. 2026, Sci Rep s41598-026-58423-z, abstract verbatim — "electrophysiological markers of arousal and memory reactivation are organized within the spindle-rich ISFS peak" (N=154), with the preprint adding "The highest percentage of microarousals occurred around the peak of the ISFS". Alejandro Osorio-Forero is second of fourteen authors, at the Netherlands Institute for Neuroscience in Amsterdam. So L24-27, L98-99 and the L355-357 prediction misstate the team's own two papers and predict against its own human result. Major: fix the mechanism narrative, and note this also voids the F189 fallback.
The problem
Lecci et al. (2017) is cited four times (L11, L27, L95, L274) and is the empirical foundation of the entire rationale — it supplies the sigma-ISF phenomenon, the continuity/fragility distinction and the arousability effect. It does not appear in the REFERENCES section at all. The same defect affects Osorio-Forero et al. (2025), cited in the Summary (L60) and as a substantive result in the Literature Review (L108-111) but listed only under PREVIOUS OWN PUBLICATIONS, not in REFERENCES. An assessor who tries to look up the founding citation of the proposal and cannot find it will draw a conclusion about the care taken with everything else.
In the draft — line 22-27
That rhythm is an infraslow fluctuation of sigma-band (11-16 Hz) power at approximately 0.02 Hz ... during fragility the same stimulus more often produces an arousal (Lecci et al., 2017).
Evidence
Grep of /root/grantreview/DRAFT.md for 'Lecci' returns exactly four hits, all in body text: lines 11, 27, 95, 274. The REFERENCES block (L375-445) contains Baker, Cataldi, De Zambotti, Esfahani, Fernandez & Lüthi, Freedman & Roehrs, Freeman & Sherif, Gombert-Labedens, Kjaerby, Lazar, Osorio-Forero (2021, twice), Rothhaas, Sikder, Tsiartas, Wassing — and no Lecci entry. 'Osorio-Forero ... (2025)' appears only at L461-464, under PREVIOUS OWN PUBLICATIONS.
Fix
Add the Lecci et al. (2017) entry to REFERENCES, and add Osorio-Forero et al. (2025) there as well (keeping it in PREVIOUS OWN PUBLICATIONS only if a team member is an author). Then run a mechanical check that every in-text citation has a matching reference entry and vice versa before submitting.
What the refuter said
STRONGEST DEFENCE: a missing bibliography entry is copy-editing, changes no science, and a Scientific Board member in this field knows Lecci et al. 2017 on sight. WHY IT SURVIVES: grep is exact — "Lecci" appears at DRAFT.md lines 11, 27, 95, 274 (four hits, all body text) and nowhere in REFERENCES (L375-445). I read the reference block: Baker, Cataldi, De Zambotti, Esfahani, Fernandez & Lüthi, Freedman & Roehrs, Freeman & Sherif, Gombert-Labedens, Kjaerby, Lazar, Osorio-Forero (x2), Rothhaas, Sikder, Tsiartas, Wassing. No Lecci. "Osorio-Forero ... (2025)" appears at L60 and L108 and only in PREVIOUS OWN PUBLICATIONS (L461-464). The finding's line numbers and count are exactly right, unlike the duplicate F40. The reference is genuinely load-bearing: L274 cites it for the claim that infraslow sigma modulation is "established in humans", which is the feasibility premise for the whole ISF-extraction section. Moderate is the correct severity — it is unfixable-by-the-reader but costless to fix.
The problem
Four separate mismatches, none disclosed. (1) Hardware: Tsiartas used a custom array of discrete sensors hard-wired into a Compumedics PSG (SC via Grove 101020052 at 64 Hz on the anterior wrist, T at 16 Hz anterior, PPG 512 Hz and motion 1024 Hz dorsal). EmbracePlus is a single integrated device; nothing in the draft or the literature establishes that the decision tree transfers across sensor hardware, sampling rate, electrode geometry and placement. (2) The source calls the sensors "consumer-grade", not research-grade. (3) Population: three women, natural menopause, i.e. postmenopausal, not perimenopausal. (4) Setting: "sound-attenuated and temperature-controlled bedrooms", not home. The paper claims only "initial feasibility".
In the draft — line 65-66
Flashes are detected from multi-sensor features of a research-grade wristband (Tsiartas et al., 2021).
Evidence
Tsiartas 2021 (/root/grantreview/Tsiartas2021.txt): "Signals from a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700 - 1024 Hz; T sensor: TI-TMP36GT9Z - 16 Hz) were collected from each participant's wrist ... and integrated with the Compumedics recording system"; "Three women (Age, mean ± SD: 55.6 ± 0.6 y)"; "had undergone natural menopause"; "Women slept in sound-attenuated and temperature-controlled bedrooms"; "The current study shows initial feasibility". EmbracePlus sensor set: "Ventral electrodermal activity (EDA) sensor", 4-channel PPG, digital temperature, accelerometer/gyroscope (https://www.empatica.com/en-eu/embraceplus/).
Fix
Delete "research-grade". State that the algorithm must be re-trained and re-validated on EmbracePlus data against concurrent sternal skin conductance in perimenopausal women at home, and budget that as a work package with a pass/fail criterion. Replace with: "Candidate flashes are detected from EmbracePlus multi-sensor features using an algorithm adapted from Tsiartas et al. (2021), which was developed on discrete consumer-grade sensors in three postmenopausal women in a temperature-controlled laboratory; transfer to this device, population and setting is validated in work package 1 against concurrent sternal skin conductance."
What the refuter said
All four mismatches verified verbatim in Tsiartas2021.txt, which I read in full. Hardware: 'Signals from a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700 - 1024 Hz; T sensor: TI-TMP36GT9Z - 16 Hz) ... integrated with the Compumedics recording system' (L59-74) — a wired array across dorsal and anterior wrist, not a wristband. Wording: 'consumer-grade' appears at L14 and L59. Population: 'Three women (Age, mean ± SD: 55.6 ± 0.6 y)' who 'had undergone natural menopause' (L29-40) — postmenopausal, against a perimenopausal target sample. Setting: 'Women slept in sound-attenuated and temperature-controlled bedrooms' (L99-101), against a home study. And 'The current study shows initial feasibility' (L235). So L65-66's 'research-grade wristband (Tsiartas et al., 2021)' misdescribes the source on both words. STRONGEST DEFENCE, which is why I hold at moderate rather than major: the EmbracePlus that this study actually uses is itself a legitimately research/medical-grade integrated wristband, so the adjective is not false of the study's device — what is wrong is attaching it to Tsiartas. And the sensitivity/specificity figures the draft quotes are accurate. Near-duplicate of F137, which carries the substantive temporal prong; downgraded fatal to moderate.
The problem
The cited paper built an automatic cross-correlation method and then rejected it in favour of manual visual alignment because manual was more reliable; it reports no quantitative residual error in seconds; it used the Empatica E4, not EmbracePlus; and it recorded one night per participant, so it establishes nothing about drift stability or non-linearity across a night. It also states the devices lack a reliable internal clock. The draft's fixed-offset-plus-linear-drift model is therefore unsupported by its own citation, and taps at only the two ends of the night cannot detect mid-night non-linearity, sample-rate jitter or dropped samples — the three failure modes that matter against a 50-s cycle.
In the draft — line 227-230
EEG and wristband are synchronised by cross-correlating accelerometer and PPG signals (Sikder et al., 2026), with a fixed offset and linear drift estimated per recording; tap markers validate this and residual error is reported.
Evidence
Sikder et al. 2026, SLEEP Advances 7(1):zpaf094 (https://academic.oup.com/sleepadvances/article/7/1/zpaf094/8405700): "we later opted to manually synchronise each participant's recordings through visual inspection of their corresponding raw tri-axial accelerometer data"; an automatic cross-correlation method was developed but "the outcome of the manual process was more reliable, which has been included in the final dataset"; devices "lacked a reliable internal clock, making straightforward time-based synchronisation unreliable"; wristband was the "Empatica E4"; N=130 healthy participants aged 18-39, one night each.
Fix
Stop citing Sikder as validation of cross-correlation. Add mid-night synchronisation markers (a scheduled vibration-prompted tap, or taps at every wake episode) so drift can be estimated with more than two anchors, report the residual error distribution from the tap markers as a primary feasibility number with a pre-registered ceiling, and fold that error into the jitter budget explicitly alongside onset uncertainty. Note that a 2 s synchronisation SD alone costs 3% attenuation; combined in quadrature with onset jitter it is not negligible.
What the refuter said
AGAINST: the draft does not lean on the citation for accuracy. L227-230 commits to independent validation and disclosure — "tap markers validate this and residual error is reported" — so no inherited precision figure is being claimed, and the tap markers are an addition beyond Sikder. Citing a paper for the general idea of accelerometer-based alignment is defensible even if that paper's final dataset used a different implementation. SURVIVES on the facts, all four of which I verified in Sikder et al. 2026 (SLEEP Advances 7(1):zpaf094): the automatic cross-correlation method was developed and then rejected — "since the outcome of the manual process was more reliable, this information has been included in the final dataset", with alignment done "by visual inspection" of tri-axial accelerometer data; no quantitative residual error in seconds is reported anywhere; the wristband was the Empatica E4 (EDA at 4 Hz), not EmbracePlus; N=130 with one night each; and the devices had "inaccuracies in some devices' internal clocks, necessitating manual synchronisation". So the draft's "fixed offset and linear drift estimated per recording" has no support in its own citation, and a one-night dataset can establish nothing about drift linearity across a night. Two taps at the ends cannot detect mid-night non-linearity, sample-rate jitter or dropped samples. Moderate rather than major: the draft's commitment to measure and report residual error is a real mitigation, and the fix is to cite the method honestly and add mid-night markers.
The problem
Two errors. Factually, gamma has a smaller bias than sigma in the same list, so "smallest bias of any band" is false. Substantively, the claim is irrelevant to what the study needs: a bias in mean relative band power is a static offset, and a constant offset in log sigma is removed by the band-pass filter entirely. What matters is whether the 0.02 Hz modulation of the sigma envelope survives at frontal derivations with wearable noise — a temporal-fidelity question the validation paper does not address. The same paper also reports lower spindle amplitude on ZMax than PSG, which cuts the wrong way. The validation sample is healthy 21-28 year olds, not perimenopausal women, and the paper contains no data on human manual scoring of ZMax against PSG — which is the scoring approach the draft plans.
In the draft — line 271-273
sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023)
Evidence
Esfahani 2023 (https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full): "ZMax demonstrated the bias of -0.055 ± 0.085, 0.045 ± 0.044, 0.016 ± 0.027, 0.005 ± 0.012, -0.008 ± 0.023, +0.002 ± 0.002" for delta, theta, alpha, sigma, beta, gamma — gamma smallest; relative bandpower, not absolute; lower spindle amplitude on ZMax "due to ... montage configuration"; validation cohorts ~95 healthy participants aged 21-28; on human scoring only "We showed the feasibility of human scoring of ZMax data, even in real-time, elsewhere", with no agreement statistics reported.
Fix
Replace with a claim about the quantity actually used: "Relative sigma-band power on this headband has low static bias against PSG (0.005 ± 0.012; Esfahani et al., 2023), but temporal fidelity of the sigma envelope at infraslow frequencies has not been validated for this device, nor in this age group; a concurrent-PSG subsample quantifies the coherence of the frontal log sigma envelope with a central derivation in the 0.01-0.04 Hz band." Report per-stage agreement for your own manual scoring of ZMax against PSG in that subsample.
What the refuter said
Strongest case against: it duplicates F42 on the factual error, and the draft's next sentences already concede the substantive point (L274-276, "Frontal expression and achievable phase precision are themselves feasibility outcomes"). But this is the better-argued twin and its extra limb is correct and sharper than F42's: a bias in *mean relative* bandpower is a static offset, and a constant offset in log sigma is removed entirely by a 0.01-0.04 Hz band-pass — so the cited statistic is not merely insufficient for the draft's purpose, it is irrelevant to it. Everything is verified at source (biorxiv 2023.08.18.553744): the bias list with gamma at +0.002 +/- 0.002 against sigma at 0.005 +/- 0.012, making "smallest bias of any band" false; ZMax fluctuations "typically attenuated when compared to the PSG", which cuts against the draft; no infraslow analysis; no arousal scoring validation; and on manual scoring only a feasibility assertion with no agreement statistic reported. One correction to the finding: the validation is not solely 21-28 year olds — datasets 1-4 are (mean ages 21-28), but dataset 5 is IBD patients around 50, used as proof-of-concept with degraded autoscoring. Moderate.
The problem
In Lecci et al. the acoustic stimulation protocol was applied only to mice; the human arm involved spontaneous sleep recordings and a memory task, with no stimuli delivered during sleep. The Summary's framing — "whether a sound wakes you depends not only on how loud it is, but on the moment at which it arrives" (lines 21-24) and "the sleeper is comparatively insulated" — reads as a human finding. The Literature Review gets the species attribution right (lines 100-103), so the Summary contradicts the body of the same document. Since the Summary is what most Scientific Board members will read closely, the overstatement sits in the worst possible place. It also weakens the proposal's own novelty argument: the human phase-arousability link is less established than the draft implies, which makes the interoceptive extension a two-step rather than one-step inference.
In the draft — line 26-27
during fragility the same stimulus more often produces an arousal (Lecci et al., 2017).
Evidence
Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853): the noise protocol applied exclusively to mice implanted for polysomnography; no stimuli were delivered to the human participants during sleep recordings, whose data concerned the correlation between 0.02 Hz oscillation strength and overnight memory consolidation. Internal inconsistency with draft lines 100-103: "In mice, sensory stimulation delivered during fragility more often produced arousal ... in humans the fragility phase is associated with spontaneous arousals".
Fix
Fix the Summary to match the body. Replace with: "In mice, identical stimulation delivered during fragility more often produces an arousal than during continuity; in humans, the fragility phase is marked by spontaneous arousals and transitions to lighter sleep, but stimulus-locked arousability across the cycle has not been tested (Lecci et al., 2017)." And note that Lecci et al. 2017 is cited four times in the body but is absent from the reference list.
What the refuter said
STRONGEST DEFENCE, PARTLY SUCCESSFUL: the Summary is more careful than the finding allows. L26-30 reads "during fragility the same stimulus more often produces an arousal (Lecci et al., 2017). In mice, noradrenergic locus coeruleus activity rises and dips with this phase (Osorio-Forero et al., 2021); in humans, fragility is marked by absent sleep spindles and spontaneous arousals (Fernandez & Lüthi, 2020)." The very next sentence separates species and attributes to humans only spindle absence and SPONTANEOUS arousals — not stimulus-driven arousability. So the draft never says "in humans the same stimulus more often produces an arousal". WHY IT SURVIVES: I fetched Lecci et al. 2017 (PMC5298853) and the species split is starker than the draft implies. Acoustic stimulation was delivered ONLY to mice ("A white noise stimulus of 90-dB SPL yielded an arousal success rate of 38.7 ± 8.6%"); the human arms were observational polysomnography (n=27) plus a memory cohort (n=24) with no stimuli during sleep; and the paper explicitly states "the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained." That undercuts two draft claims beyond the one quoted: the framing at L19-24 ("whether a sound wakes you depends ... on the moment at which it arrives"), which reads as a human fact, and the premise at L31-32 and L115-117 that "this literature rests entirely on stimuli delivered by an experimenter" — true of the rodent work only. The proposal's extension is therefore a two-step inference (mouse stimulus arousability → human stimulus arousability → human interoceptive access), and the Summary is where that gets compressed to one. Severity stays moderate.
The problem
Three overstatements in one passage. (1) The draft says "establishing that it is functionally consequential" on the strength of a single correlation in 24 subjects at r = 0.45, p = 0.027 — that establishes nothing; it is suggestive. (2) The human sample was 27 healthy young males aged 22.5, with no acoustic stimulation experiment at all, and the authors said so explicitly. The draft's population is perimenopausal women with clinically significant insomnia — a group in whom neither ISF strength nor its arousal relationship has been characterised, and in whom the one clinical population studied so far showed reduced ISF presence and strength. (3) The clause "in humans the fragility phase is associated with spontaneous arousals" is the claim most directly contradicted by later human data (see finding human-arousals-at-sigma-peak).
In the draft — line 100-106
in mice, sensory stimulation delivered during fragility more often produced arousal than identical stimulation during continuity; in humans the fragility phase is associated with spontaneous arousals and transitions to lighter sleep. The oscillation is coordinated with an infraslow cardiac rhythm, indexing a brain-autonomic rather than a purely cortical state, and its strength predicts overnight memory retention, establishing that it is functionally consequential.
Evidence
Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853): human core dataset "27 healthy male subjects (age 22.5 ± 0.49 years)", memory extension n=24; "Recall correlated with the individual peak of 0.02-Hz oscillations in the fast spindle band during all-night non-REM sleep (r = 0.45, P = 0.027)"; "No acoustic stimulation was performed on humans"; "the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained." — Reduced ISF in a clinical population: Dimitriades et al., bioRxiv 2025.04.23.650209: "The presence and strength of the ISFS were reduced in both COS and EOS groups compared to controls."
Fix
Downgrade "establishing" to "suggesting", state the sample, and add ISF expression in this population as an explicit unknown. Rewrite: "...and its strength correlated with overnight memory retention in a small sample (r = 0.45, P = 0.027, n = 24), suggesting functional consequence. The human evidence rests on 27 young men without any stimulation experiment; Lecci et al. noted that the oscillation's role for human arousability remained to be established. ISF expression has not been characterised in perimenopausal women, and is reduced in at least one clinical population (Dimitriades et al., 2025), so we report ISF presence and strength in this cohort as a primary descriptive outcome." This is more honest and it also converts a weakness into a stated contribution.
What the refuter said
AGAINST: limb (3) duplicates F45/F183 rather than adding evidence, and limb (2) is partly pre-empted — the draft explicitly designates frontal ISF expression and phase precision as feasibility outcomes (L274-276) and builds a contingency for weak frontal ISF (L343), so it is not blind to the population question. On limb (1), the paragraph does not rest on Lecci alone: Kjaerby et al. 2022 is cited two sentences later (L111-113) for the functional consequence of NE oscillation amplitude. SURVIVES. Verified in Lecci 2017: the human core dataset is "27 healthy men (22.5 ± 0.49 years of age)"; "No acoustic stimulation was performed on humans"; the memory result is "r = 0.45, P = 0.027; n = 24". A single r=0.45 at p=0.027 in 24 young men does not warrant the draft's "establishing that it is functionally consequential" (L105-106) — Kjaerby's converging evidence is rodent, so the human claim rests on that one correlation. And the population gap is real and unstated: the draft's cohort is perimenopausal women with ISI>=10 (verified in the parent protocol), a group in which frontal ISF expression is unmeasured, while Dimitriades et al. 2025 (bioRxiv 2025.04.23.650209, verified) found "features of the ISFS, namely its presence and strength, are reduced in central-parietal regions in schizophrenia compared to healthy controls" — the one clinical population studied. Moderate: soften "establishing", and state the generalisation as an inference.
The problem
The number is accurately quoted — Esfahani et al. do report 0.005 ± 0.012 for sigma, the smallest absolute bias across bands. But it is close to the wrong statistic for this application, and the surrounding evidence is unfavourable in three ways the draft omits. (i) Bias is not the relevant property. A constant per-subject offset in log sigma power is irrelevant to the *phase* of the band-passed log-sigma envelope; what matters is epoch-level reliability, i.e. how faithfully the headband tracks within-night fluctuation. Esfahani et al. report no epoch-level sigma correlation at all. An independent 2026 PSG-concurrent validation does: frontal sigma r=0.512 in N2 and r=0.487 in N3 before calibration, with spindle density under-detected by roughly half (1.01 vs 1.62 spindles/min). Attenuation of that order directly reduces ISF envelope amplitude and degrades phase estimation — exactly the quantity the primary test depends on. (ii) Population. Esfahani's samples are healthy young adults (mean age 21-28); the same 2026 validation explicitly warns that "in clinical populations with severely disrupted sleep architecture, including insomnia disorder… the available N2 epoch count may be insufficient for reliable calibration". The cohort here is, by inclusion criterion, ISI≥10. Frontal ISF expression in perimenopausal women with insomnia is unmeasured, and the draft concedes this in passing ("Frontal expression and achievable phase precision are themselves feasibility outcomes", line 274-276) while still calling feasibility "supported". (iii) The same paper's usable-recording rates (55-68%) are the yield evidence the draft omits — see the yield finding.
In the draft — line 271-276
Feasibility is supported: sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023), and the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019).
Evidence
Esfahani et al. 2023, https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full.pdf: "ZMax demonstrated the bias of −0.055 ± 0.085, 0.045 ± 0.044, 0.016 ± 0.027, 0.005 ± 0.012, −0.008 ± 0.023" (delta-beta; sigma smallest) — claim accurate; sample "mean age 21.1 ± 3.4 years" (Dataset 1), "22 females, mean age 21.59 years ± 3.78" (Dataset 2); no epoch-level sigma correlation reported. Parry & Briganti 2026, https://www.medrxiv.org/content/10.64898/2026.06.01.26354593v1.full: frontal sigma "r=0.512 (N2), r=0.487 (N3)" pre-calibration, "systematic underestimation: −0.74 log units across sigma band", "per-subject calibration is essential before cross-subject comparison", spindle density "mean 1.01 spindles/min vs PSG F3 at 1.62", and "In clinical populations with severely disrupted sleep architecture, including insomnia disorder… the available N2 epoch count may be insufficient for reliable calibration".
Fix
Rewrite the feasibility sentence to argue the right property and disclose the caveats: "Absolute sigma power from this headband is minimally biased against PSG (0.005 ± 0.012; Esfahani et al., 2023), and because ISF phase is derived from the band-passed log-sigma envelope, a per-subject offset is not a threat. The relevant risk is epoch-level attenuation: frontal ZMax sigma correlates with PSG at r≈0.51 in N2 (Parry & Briganti, 2026), which would reduce ISF envelope amplitude. Both validations used young healthy adults, and frontal ISF expression in perimenopausal women with insomnia is unmeasured; establishing it, with a pre-specified adequacy criterion, is the first work package." Add the insomnia-fragmentation caveat where the stable-NREM criterion is defined.
What the refuter said
AGAINST, and one limb is genuinely mis-stated. The finding presents r=0.512 (N2) as an EPOCH-level correlation and builds its whole argument on within-night tracking — but Parry & Briganti report SUBJECT-level Spearman correlations, which say nothing about within-night fluctuation and so cannot carry the inference. That paper is also net favourable after correction: post-calibration N3 sigma reaches r=0.752 and "calibrated Zmax spectral features show excellent within-subject stability". The draft also pre-empts partly, conceding at L274-276 that "Frontal expression and achievable phase precision are themselves feasibility outcomes". SURVIVES on the two points that matter, both verified. (i) Bias is close to the wrong statistic: a constant per-subject offset in log sigma power is removed entirely by the band-pass filter, so it is irrelevant to ISF phase — and I confirmed Esfahani et al. report NO epoch-level sigma correlation anywhere, while accurately quoting their "0.005 ± 0.012" for sigma (their sigma band being 13-16 Hz, incidentally, not the draft's 11-16). (ii) Population: Esfahani's datasets are healthy young adults (21.1 ± 3.4; 21.59 ± 3.78; 27.82 ± 6.07), and Parry & Briganti 2026 (verified to exist, medRxiv 2026.06.01.26354593) warn that "In clinical populations with severely disrupted sleep architecture, including insomnia disorder (reduced N2 continuity) ... the available N2 epoch count may be insufficient for reliable calibration" — against a cohort defined by ISI>=10. Moderate: cite the right statistic and the right population.
The problem
Three separate errors in one clause. (i) The source repeatedly and deliberately describes its sensors as consumer-grade, and frames the whole exercise as testing whether consumer hardware can do the job — "research-grade" is the opposite of the paper's claim. (ii) It was not a wristband: sensors were placed on the wrist and cabled into a Compumedics PSG amplifier, an arrangement that cannot be used in unattended home recording. (iii) The draft will actually use the EmbracePlus, a different device with different sensors, sampling rates and signal chain, and offers no bridging validation. Empatica's published EmbracePlus specification lists EDA range 0.01-100 µS with resolution ~900 pS and accuracy "N/A" and does not state an EDA sampling rate at all, while Tsiartas' skin-conductance sensor ran at 64 Hz. The gold-standard flash criterion is a 2 µS rise within 30 s; a device whose manufacturer publishes no EDA accuracy figure cannot be assumed to resolve it.
In the draft — line 64-66 (with 214-217)
Flashes are detected from multi-sensor features of a research-grade wristband (Tsiartas et al., 2021).
Evidence
Tsiartas2021.txt L59-72: "Signals from a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; 3-axis motion sensor: NXP-FXOS8700 - 1024 Hz; T sensor: TI-TMP36GT9Z – 16 Hz) were collected from each participant's wrist ... and integrated with the Compumedics recording system, using a multi-channel output card". Abstract L13-15: "a novel hot flash (HF) classification algorithm based on multi-sensor features integration using commercial wearable sensors"; L59: "No current solution exists within the consumer space for measuring HFs". EmbracePlus specification: EDA "0.01 μSiemens – 100 μSiemens", resolution "1 digit ~ 900 pSiemens", accuracy "N/A", no EDA sampling rate given (https://ifelldh.tec.mx/sites/g/files/vgjovo1101/files/Empatica%20Embraceplus.pdf).
Fix
Rewrite as: "Flashes are detected by reimplementing the multi-sensor feature set of Tsiartas et al. (2021), developed on wrist-placed consumer-grade sensors wired to a laboratory amplifier. Because the EmbracePlus differs in sensor hardware, sampling rate and signal conditioning, we first re-derive and re-validate the classifier on EmbracePlus data against simultaneous sternal skin conductance in a laboratory subsample (n = ...), and report EmbracePlus EDA sampling rate, resolution and noise floor relative to the 2 µS/30 s criterion." State the actual EmbracePlus EDA sampling rate and accuracy figures in the methods; if the manufacturer does not publish them, say so and measure them.
What the refuter said
AGAINST: two of the three sub-errors are weak. (a) "Research-grade wristband" at L64-66 plausibly describes the EmbracePlus — a device Empatica markets as medical-grade — with the Tsiartas citation attaching to the detection method, not the hardware; and the Methods (L212-215) correctly name ZMax and EmbracePlus. (b) The EDA-spec argument is a non-sequitur: I fetched the spec sheet and confirmed range 0.01-100 uS, resolution ~900 pSiemens, accuracy "N/A", no sampling rate — but absolute accuracy is irrelevant to detecting a *rise*, 900 pS resolves a 2 uS excursion to ~2,200 steps, and the 2 uS/30 s criterion is a STERNAL rule the draft never claims to apply at the wrist. SURVIVES on (ii) and (iii), which are the substance. Tsiartas2021.txt is unambiguous: "Signals from a customized array of consumer-grade commercially available sensors (PPG: S/F SEN-11574 - 512 Hz; SC sensor: Grove 101020052 - 64 Hz; ... T sensor: TI-TMP36GT9Z - 16 Hz) were collected from each participant's wrist ... and integrated with the Compumedics recording system"; "No current solution exists within the consumer space for measuring HFs"; all recordings "took place at the SRI International Human Sleep Research Laboratory" in "sound-attenuated and temperature-controlled bedrooms" with concurrent sternal SC. The draft transfers that rig's ">90% sensitivity at 95.6% specificity" (L242-246) to unattended home EmbracePlus recording with no bridging validation and no caveat. Severity moderate, not major: the wording is a one-word Summary slip, but the unhedged import of lab performance figures onto a different device is a real, catchable gap.
The problem
Three errors in one sentence. (i) 'Smallest bias of any band' is false — gamma's bias is +0.002 ± 0.002, smaller in both magnitude and dispersion than sigma's 0.005 ± 0.012. (ii) The quantity is a bias in *relative* bandpower, a dimensionless proportion; the draft states it bare, so a reader cannot tell it is not an absolute-power or a phase measure. Sigma is a low-power band, so a small proportional bias is close to arithmetically guaranteed and carries little information. (iii) Most seriously, the inference does not follow. A small bias in epoch-averaged relative bandpower says nothing about whether the *infraslow modulation* of sigma — a ~0.02 Hz, ~50 s rhythm whose instantaneous phase the draft must estimate — survives this device. The validation reports only epoch-level bandpower and event detection; it contains no time-resolved spectral analysis, no infraslow analysis, and no arousal validation at all.
In the draft — line 271-273
sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023)
Evidence
Esfahani et al., bioRxiv 10.1101/2023.08.18.553744 (https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full.pdf): "Overall, with respect to PSG relative bandpower, ZMax demonstrated the bias of -0.055 ± 0.085, 0.045 ± 0.044, 0.016 ± 0.027, 0.005 ± 0.012, -0.008 ± 0.023, +0.002 ± 0.002, and 1.13e-10 ± 2.22e-10 for delta, theta, alpha, sigma, beta, gamma, and absolute overall power, respectively". Bands defined as "SO (0.5–1 Hz), delta (1–4 Hz), theta (4–8 Hz), alpha (8–12 Hz), and sigma (13–16 Hz)". No infraslow or time-resolved power analysis is reported; arousal scoring is explicitly not assessed.
Fix
Replace with: 'The headband's relative sigma-band power tracks polysomnography closely (bias 0.005 ± 0.012 in relative bandpower; Esfahani et al., 2023), but that validation is epoch-level and says nothing about the infraslow envelope. Recovering the 0.02 Hz modulation of frontal sigma from this device is therefore itself a feasibility outcome, and we report the per-participant infraslow peak, its amplitude relative to a phase-shuffled surrogate, and the split-half reliability of estimated phase before any outcome analysis.'
What the refuter said
Strongest case against: the draft's very next sentences already concede the substance. L274-276: "Frontal expression and achievable phase precision are themselves feasibility outcomes" — i.e. the draft does not claim the infraslow modulation is proven measurable on this headband; it says measuring it is part of what the project establishes. That blunts limb (iii). But limb (i) is a plain factual error and I verified it verbatim at the source (biorxiv 2023.08.18.553744): "ZMax demonstrated the bias of -0.055 +/- 0.085, 0.045 +/- 0.044, 0.016 +/- 0.027, 0.005 +/- 0.012, -0.008 +/- 0.023, +0.002 +/- 0.002 ... for delta, theta, alpha, sigma, beta, gamma". Gamma is smaller in both magnitude and dispersion, so "the smallest bias of any band" (L271-273) is false. Limb (ii) also checks: the quantity is relative bandpower, and the paper reports no infraslow or dedicated time-resolved power validation and no arousal scoring validation. One further mismatch the finding missed and that strengthens it: Esfahani defines sigma as 13-16 Hz, the draft as 11-16 Hz (L23), so the borrowed statistic is not even for the draft's band. Verifiably false superlative on a feasibility claim; moderate rather than major given the L274-276 hedge.
The problem
The cited paper does three things differently. (i) It cross-correlated accelerometer data only; PPG was not used for alignment. (ii) The alignment it performed was ZMax-to-PSG, not ZMax-to-wristband. (iii) It explicitly rejected the automated cross-correlation in favour of manual visual inspection because the manual result was more reliable, and it reports no residual synchronisation error or drift estimate at all. The draft is thus citing, as an established method, a procedure the source tried and discarded — and the whole phase analysis depends on this alignment, since a 50 s cycle allows no more than a few seconds of inter-device error. The wristband is also a different device generation (Empatica E4 in Sikder, EmbracePlus in the draft).
In the draft — line 227-230
EEG and wristband are synchronised by cross-correlating accelerometer and PPG signals (Sikder et al., 2026), with a fixed offset and linear drift estimated per recording; tap markers validate this and residual error is reported.
Evidence
Sikder et al. (2026), SLEEP Advances 7(1), zpaf094, DOI 10.1093/sleepadvances/zpaf094 (https://academic.oup.com/sleepadvances/article/7/1/zpaf094/8405700): "we later opted to manually synchronise each participant's recordings through visual inspection of their corresponding raw tri-axial accelerometer data"; of the automated approach, "cross-correlations were computed over nonnegative lags, and the lag with the maximum correlation was used to identify the best-aligned window" but "the outcome of the manual process was more reliable, so this information has been included in the final dataset". Devices: Zmax, Empatica E4, ActivPAL, Somnoscreen/Mentalab; 130 participants, one night each, at home; 100/130 usable for the synchronised PSG+Zmax subset. No residual synchronisation error is reported.
Fix
Rewrite as: 'EEG and wristband are aligned on their tri-axial accelerometer traces, following the manual visual-inspection procedure that Sikder et al. (2026) found more reliable than automated cross-correlation, with automated cross-correlation used as a first pass and a fixed offset plus linear drift fitted per recording. The five-tap markers at lights-off and on waking give an independent check; we report the tap-to-tap residual for every recording and exclude nights whose residual exceeds a pre-registered fraction of the infraslow cycle.' Note that no published residual-error figure exists for this device pair, so the tap markers are the only evidence, not a validation.
What the refuter said
AGAINST: claim (ii) is wrong. I fetched the paper and the accelerometer alignment covered the wristband too — the synchronisation dashboard aligns "(a) Zmax, (b) Empatica, and (c) Activpal recordings based on accelerometer data", not just ZMax-to-PSG. And the draft's method is independently sensible: both devices in the draft record PPG (ZMax has "accelerometry and PPG", L213-214; EmbracePlus has PPG, L214-215), so cross-correlating heart-rate series between them is not a mistake, and the draft does not rest on Sikder for validation — it adds tap markers and commits that "residual error is reported" (L229-230). SURVIVES on claims (i) and (iii), which are the load-bearing ones. Verified: PPG was not used for alignment, only accelerometer data; the automated cross-correlation ("cross-correlations were computed over nonnegative lags, and the lag with the maximum correlation was used") was set aside because "the outcome of the manual process was more reliable"; and no residual synchronisation error or drift estimate is reported anywhere. So the draft cites, as established method, a procedure the source tried and preferred not to use, and attributes to it a signal it did not use. One point the finding missed strengthens it: Sikder's participants "were instructed to perform five jumps before bedtime and upon waking" — the exact analogue of the draft's five-tap marker — and that event-based scheme was also abandoned. Downgraded from major to moderate: a citation misattribution on a method that is defensible on its own terms, not a broken method.
The problem
The stimulus-arousal result is a mouse result. Lecci et al. tested arousability with white noise in mice only, and their discussion states in terms that the human case is unresolved. The Summary presents it without species qualification, so a reader takes the human gating of external signals as established — which is the premise the whole proposal builds on. The Literature Review is more careful about the mouse stimulation (line 100-102) but then asserts 'in humans the fragility phase is associated with spontaneous arousals and transitions to lighter sleep', which is also not something Lecci reports; the human dataset was analysed for memory consolidation, not for arousals. The related claim at line 29-30, 'in humans, fragility is marked by absent sleep spindles and spontaneous arousals (Fernandez & Lüthi, 2020)', cites a review rather than a primary human observation, and a review by the same senior author is unlikely to establish what the primary paper says remains to be ascertained.
In the draft — line 25-27
During continuity the sleeper is comparatively insulated from sensory input; during fragility the same stimulus more often produces an arousal (Lecci et al., 2017).
Evidence
Lecci et al. 2017 (PMC5298853): mouse stimulation — "A white noise stimulus of 90-dB sound pressure level (SPL) yielded an arousal success rate of 38.7 ± 8.6%"; "wake-ups and sleep-throughs occur during declining and rising sigma power levels, respectively". On humans — "Although a protective function of sleep spindles for arousals is well established, the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained", and "We caution here against a simple transfer of approaches between species." UNVERIFIED: Fernandez & Lüthi (2020), Physiological Reviews 100(2):805-868, could not be accessed (journals.physiology.org returns 403 for both full text and PDF), so the exact wording of the human fragility claim in that review is unconfirmed.
Fix
In the Summary, add the species: 'In mice, the same stimulus more often produces an arousal during fragility than during continuity (Lecci et al., 2017); whether the human 0.02 Hz oscillation gates arousability is, in those authors' own words, still to be ascertained.' Then make that gap part of the pitch rather than papering over it. Replace the Fernandez & Lüthi attribution with a primary human source, or state that no human stimulation study of this rhythm exists.
What the refuter said
AGAINST: the finding partly answers itself — it concedes the Literature Review marks the species correctly at L100-102 ("In mice, sensory stimulation delivered during fragility more often produced arousal..."), and the Summary does mark species for the LC claim at L27-28 ("In mice, noradrenergic locus coeruleus activity..."). It also concedes UNVERIFIED on Fernandez & Lüthi 2020, which I could not reach either. SURVIVES. I verified against Lecci et al. 2017 (PMC5298853) directly: "No acoustic stimulation was performed on humans" (mice only, n=10); the human sample is "27 healthy men (22.5 ± 0.49 years of age)"; and the authors state "Although a protective function of sleep spindles for arousals is well established, the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained" and "We caution here against a simple transfer of approaches between species." The draft's Summary sentence at L25-27 — "During continuity the sleeper is comparatively insulated from sensory input; during fragility the same stimulus more often produces an arousal (Lecci et al., 2017)" — carries no species marker, sits in the section a Scientific Board member reads first, and states as established precisely what the cited authors flag as unascertained. Moderate, not major: the draft demonstrates elsewhere that it knows the distinction, so this is a drafting correction, not a design flaw. Note it compounds with F183, where the human evidence now points the other way.
The problem
Lecci et al. used 10-15 Hz, not 11-16 Hz. The same wrong band appears in the Summary (line 23). This matters beyond bibliographic tidiness: the device validation the draft leans on defines sigma as 13-16 Hz, Osorio-Forero et al. 2021 used 10-15 Hz, and the draft's own arousal criterion is '> 16 Hz'. Four different band edges are in play and the draft never states which it will use, so the pre-registration cannot fix the predictor. The 11-16 Hz upper edge also abuts the arousal criterion, inviting spectral leakage between the predictor and the outcome.
In the draft — line 96-97
Lecci et al. (2017) described an oscillation of sigma-band (11-16 Hz) power at approximately 0.02 Hz, present in both mice and humans during NREM sleep.
Evidence
Lecci et al. 2017 (PMC5298853): "sleep spindles are electroencephalographic hallmarks of non–rapid eye movement (non-REM) sleep in the sigma (10 to 15 Hz) power range" and "the sigma (10 to 15 Hz) power band". Esfahani et al. (bioRxiv 2023.08.18.553744): "sigma (13–16 Hz)". /root/grantreview/OsorioForero2021.txt (lines 120, 149, 1125, 1141, 1144): sigma given as "10–15 Hz" throughout, with the infraslow filter at "0.01 – 0.04 Hz" (line 1147) — the same band the draft uses.
Fix
Correct to '(10-15 Hz)' where Lecci is cited, and add one sentence in ISF extraction fixing the analysis band: 'Sigma is defined as 10-15 Hz, following Lecci et al. (2017) and Osorio-Forero et al. (2021), which leaves a 1 Hz gap to the >16 Hz arousal criterion.'
What the refuter said
Strongest case against: a band-edge slip in a literature review is bibliographic housekeeping, and 11-16 vs 10-15 Hz overlaps heavily, so no substantive conclusion changes. The claimed spectral-leakage risk is also weak: the arousal criterion the draft cites ("transient > 16 Hz activity lasting 3-15 s", L236-237, Wassing et al. 2019) is a visual scoring rule, not a spectral filter applied to the same time series, so predictor and outcome are not two adjacent passbands. Even so, the factual error is verified three ways and sits in a load-bearing sentence. Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853/): "10 to 15 Hz" sigma throughout, with the refined human analysis on "a 2-Hz band around the fast spindle peak (FSP; 13.16 +- 0.12 Hz)". /root/grantreview/OsorioForero2021.txt L120 and L149: "Summed sigma (10-15 Hz) and delta (1.5-4 Hz)" and "the sigma (10-15 Hz) and the delta (1.5-4 Hz)". Esfahani 2023: sigma = 13-16 Hz. So the band the draft attributes to Lecci is wrong, the error is repeated in the Summary (L23), and the draft never states which band it will use for its own predictor - so the pre-registration cannot fix the predictor without a further decision. Moderate: verifiable, repeated, and touches the definition of the primary predictor, but changes no conclusion.
The problem
The proposal inherits onset timing from a pipeline whose timing was validated only to +/-90 s, on a classifier that emits one decision every 15 s, against a gold standard defined on 30-s windows. Against a 50-s cycle, none of these three resolutions is adequate, and the last is a property of the field's reference standard, so the required precision cannot be demonstrated with any existing method.
In the draft — line 242-244
Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier
Evidence
Tsiartas2021.txt: "We time-aligned the features with the HF expert annotations for prediction and evaluation (+-90 s matching window)"; "a Decision Tree classifier, which makes a decision every 15 s". Gombert-Labedens et al. 2025 (GombertLabedens2025.txt, p.~110): "The most widely accepted rule for the detection of a hot flash is based on an observed rapid rise of at least 2 microSiemens in sternal skin conductance within a 30-second period". Attenuation implied by each resolution alone against T=50 s (/root/grantreview/sim/01_attenuation.py, 06_filter_edge.py): 15-s frame quantisation (uniform, SD 4.3 s) -> lambda = 0.86; 30-s gold-standard window (uniform) -> lambda = |sinc(pi*30/50)| = 0.505; +-90 s validation tolerance -> lambda indistinguishable from 0. A lambda of 0.505 alone inflates the required N by 3.9x.
Fix
Add to Onset timing: "The reference standard for hot-flash detection is defined on a 30-s window and the cited classifier was validated only within a +-90 s matching tolerance, so neither can establish onset precision at the resolution this design needs. Onset precision is therefore treated as an unvalidated quantity and established de novo in a laboratory sub-sample against simultaneous high-rate sternal skin conductance with change-point back-dating, with the resulting error distribution (bias and SD) reported before the primary analysis is run."
What the refuter said
Strongest case against: the premise misreads the draft. The finding says the proposal "inherits onset timing from a pipeline whose timing was validated only to +/-90 s". It does not. L259-262 states the opposite: "Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing." Tsiartas supplies event *labels*; onset time is computed independently on the continuous electrodermal signal at sampling resolution, so neither the 15-s decision frame nor the +/-90 s matching window propagates as timing error. All three quotes are verified at source (Tsiartas: "+-90 s matching window", "makes a decision every 15 s", gold standard "sudden increases (2 uS/30s) in sternal skin conductance"), but they characterise the classifier's evaluation protocol, not the draft's onset estimator. The draft also names this as "the principal analytic risk" with a pre-outcome quantitative check and a pre-registered fallback (L338-343). What survives is the finding's last clause, and it is the important one: because no gold standard finer than 30 s exists in this field, the change-point estimate cannot be validated for *accuracy* at all, and the draft's proxy — "Intervals between this inflection and the accompanying temperature and pulse-rate changes" (L261-263) — measures agreement among peripheral markers, i.e. precision, not accuracy against truth. Moderate.
The problem
Four problems. (1) The AASM rule is an abrupt shift in EEG frequency including alpha, theta and/or frequencies above 16 Hz, explicitly excluding spindles; the draft's version keeps only the >16 Hz clause and drops the spindle exclusion — in a study whose predictor is 11-16 Hz sigma power, that exclusion is not optional. (2) AASM requires at least 10 s of preceding continuous sleep, which the draft omits and which interacts with the stable-bout criterion. (3) AASM requires a concurrent chin EMG increase of at least 1 s to score an arousal in REM; ZMax has no EMG, so REM arousals cannot be scored to standard. (4) The 15 s upper bound is not an AASM criterion. Frontal-only recording with no EOG makes this worse: eye movements and frontalis EMG produce exactly the high-frequency transients that mimic arousals, and a hot flash produces movement — so the supporting outcome is measured by a channel the event itself contaminates.
In the draft — line 236-237
Arousals follow standard criteria: transient \> 16 Hz activity lasting 3-15 s (Wassing et al., 2019).
Evidence
AASM ISR scoring guidance (https://isr.aasm.org/helpv5/ScoringArousalsA.html): "Arousals must last at least three seconds. Arousals must be preceded by at least 10 seconds of continuous sleep"; "Arousals during R must also have an increase of chin EMG lasting at least one second". Esfahani 2023 on ZMax: "The limited number of EEG channels and the absence of EMG (and occasionally EOG) signals ... pose a challenge for sleep scoring" (https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1.full).
Fix
State the full criterion including alpha and theta, the spindle exclusion and the 10 s preceding-sleep rule, and state plainly that AASM REM arousal criteria cannot be met without chin EMG — then justify the NREM restriction on that basis rather than on convenience. Add an explicit artefact-rejection rule for movement and muscle transients in the peri-flash window, and report inter-rater agreement for arousal scoring separately from stage scoring.
What the refuter said
Two of four prongs verified, one is backwards, one is wrong-ish. Verified from the AASM ISR page: 'Arousals must be preceded by at least 10 seconds of continuous sleep' and 'Arousals during R must also have an increase of chin EMG lasting at least one second'; verified from Esfahani that the ZMax Lite has neither EOG nor chin EMG, so REM arousals cannot be scored to that rule. Prong (1) is inverted, and inverted on the finding's own logic: it argues that dropping the spindle exclusion matters because the predictor is 11-16 Hz sigma — but the draft's criterion restricts to >16 Hz, which excludes the sigma band by construction and is therefore a STRICTER spindle exclusion than AASM's, not a dropped one. Prong (4) is true (15 s is not AASM) but the ceiling is the conventional arousal/awakening boundary in this literature. As with F53, the draft never claims AASM conformance: L236-237 says 'standard criteria' and cites Wassing et al. (2019), the applicants' own group's published criterion. The closing observation — frontal-only recording with no EOG means eye movements and frontalis EMG mimic arousals, and a flash produces movement, so the outcome channel is contaminated by the event — is reasonable and I could not verify it either way. Duplicate of F53; downgraded major to minor.
The problem
The draft asserts as established that treating flashes does not resolve the sleep complaint. Its own cited review states the opposite. Gombert-Labedens et al. 2025 — the source the draft leans on for flash physiology — says the flash-sleep link "is further supported by studies showing that the treatment of hot flashes...with hormone therapy reduces sleep disturbances." A reviewer who checks the citation will find the draft contradicting it. The claim is nonetheless defensible on other evidence the draft does not cite: an OASIS/NIRVANA pooled mediation analysis found that a majority of elinzanetant's sleep benefit is NOT mediated by vasomotor-symptom reduction, and Matthews et al. found menopause-associated NREM beta (hyperarousal) increases independent of self-reported hot flashes. The problem is the overstated framing plus the missing support, not the underlying idea.
In the draft — line 371-373
why women report far fewer nocturnal hot flashes than they have, and why treating the flashes does not reliably resolve the sleep complaint that accompanies them.
Evidence
Gombert-Labedens et al. 2025 (source file, line ~903): "The strong link between hot flashes and sleep disturbance is further supported by studies showing that the treatment of hot flashes (see below) with hormone therapy reduces sleep disturbances [210–213]." — Supporting evidence not cited: elinzanetant pooled mediation analysis (OASIS-1/2/3 + NIRVANA, n=1,345; presented SLEEP 2026, Maki et al., reported at https://www.patientcareonline.com/view/elinzanetant-improves-sleep-independent-of-hot-flash-reduction-in-menopause): natural direct effect on PROMIS Sleep Disturbance −2.67 (95% CI −3.28 to −2.07) of a total effect −4.92, "The NDE represented 54.3% of the total treatment effect." — Matthews KA et al. 2021, Sleep 44(11):zsab139, https://doi.org/10.1093/sleep/zsab139: "Statistical controls for self-reported hot flashes did not explain findings."
Fix
Soften the claim and cite the evidence that actually supports it. Replace with: "...why women report far fewer nocturnal hot flashes than they have, and why sleep improvement under vasomotor-symptom treatment is only partly attributable to flash suppression — roughly half of elinzanetant's sleep benefit is not mediated by reduced flashes (OASIS/NIRVANA pooled mediation analysis), and menopause-associated NREM hyperarousal is independent of self-reported flashes (Matthews et al., 2021)." Note the mediation analysis is conference-presented, not yet published — verify its publication status before citing, or cite Matthews alone.
What the refuter said
AGAINST: "contradicted" is an over-reading. The draft says treating flashes "does not reliably resolve the sleep complaint" (L372-373); the review says hormone therapy "reduces sleep disturbances" (verified verbatim in GombertLabedens2025.txt, L~903). Reduce and resolve are not the same predicate — a treatment can lower a symptom score without reliably clearing the complaint, so the two statements are logically compatible and no reviewer holding both would find a contradiction. The finding also concedes the claim is defensible on other evidence. What survives is thinner than the title: an uncited assertion of clinical fact in a closing rhetorical paragraph, sitting in visible tension with the review the draft leans on for flash physiology, when better support (mediation analyses showing sleep benefit not mediated by VMS reduction; Matthews 2021 on NREM beta independent of self-reported flashes) exists and is not cited. Worth a citation; not a defect that moves a decision. Minor.
The problem
The REM suppression half is well supported. The reversal half is presented as a demonstrated fact and is then used to justify a covariate ("half of the night") in the confirmatory model. But the review the draft relies on describes this as one finding among conflicting ones, and specifically records that de Zambotti et al. 2014 — cited elsewhere in the draft — found no first-half/second-half difference. Building a covariate into a confirmatory model on a contested finding is defensible; asserting the finding as settled is not, and a reviewer holding both papers will notice.
In the draft — line 143-147
Freedman and Roehrs (2006) showed that hot flashes are suppressed during REM sleep, when thermoregulatory effector responses are largely inactive, and that the flash-arousal ordering reverses across the night: flashes precede arousals in the first half, arousals precede flashes in the second.
Evidence
Gombert-Labedens et al. 2025 (source file, lines 911-918): "One study found that awakenings were more likely to occur before, than after, a hot flash [214], and another found that hot flashes occurred before an awakening only in the first half of the night [215], whereas others found that the majority of hot flashes coincide with awakenings [179,194,216,217] with no differences in the first and second part of the night [194]." Reference [215] is Freedman & Roehrs 2006; [194] is de Zambotti et al. 2014, cited by the draft at lines 40-41 and 149-152. REM suppression is supported: "hot flashes are less likely in REM than in non-REM sleep [179,194,215]" (line ~974).
Fix
Hedge the reversal and keep the covariate. Rewrite: "Freedman and Roehrs (2006) showed that hot flashes are suppressed during REM sleep, when thermoregulatory effector responses are largely inactive, and reported that flashes precede arousals in the first half of the night and follow them in the second; other studies find flashes and awakenings largely coincident with no first/second-half difference (de Zambotti et al., 2014). We therefore include half of the night as a covariate rather than as a hypothesis." Mirror the same hedge at line 291.
What the refuter said
AGAINST: the draft attributes the reversal to a specific study — "Freedman and Roehrs (2006) showed ..." (L143-147) — which is accurate reporting of what that study found, not a claim that the literature is settled. And where the finding actually bears weight, in the model, the draft is correctly hedged: the covariate is justified because the relationship "may reverse across the night" (L290). Building a coarse covariate on a contested finding is conservative, and the finding concedes as much. Survives only as a completeness point, and its evidence does check out: GombertLabedens2025.txt L911-918 records the mixed picture, and the reference numbers verify — [215] is "Freedman RR, Roehrs TA. Effects of REM sleep and ..." and [194] is "de Zambotti M, Colrain IM, Javitz HS, et al.", so the null the finding cites is indeed from a paper the draft itself cites twice (L40-41, L149-152). Not flagging that conflict in a 6,000-character review is a modest omission. Minor, down from moderate.
The problem
The 3.5-per-night figure and the ~20% undisturbed figure are both verifiable and attributable to de Zambotti et al. 2014. The "only report 1.5" figure is not. The Gombert-Labedens review summarises exactly this study and gives 3.5 flashes per night, n=34, 70% associated with arousal, 20% without disturbance — but no reported-flash count. I could not reach the de Zambotti full text through any available route (ScienceDirect, PMC and PubMed are all blocked by robots or proxy policy from this session), so I state this as UNVERIFIED rather than wrong. The 3.5-versus-1.5 contrast carries a lot of rhetorical weight in the Summary and would be an embarrassing thing to have mis-transcribed.
In the draft — line 40-41
Women have on average 3.5 objectively recorded flashes per night but only report 1.5 (De Zambotti et al., 2014).
Evidence
Gombert-Labedens et al. 2025 (source file, lines 917-922): "One of those studies found that perimenopausal individuals (n = 34) had an average of 3.5 hot flashes per night and that 70% of hot flashes were associated with an arousal from sleep, with only a minority (20%) occurring without disturbance to sleep; the remaining hot flashes occurred when the participant had already been awake for at least one minute [179,194]." No reported/subjective count appears. UNVERIFIED: de Zambotti et al. 2014, Fertil Steril 102(6):1708-1715, full text not retrievable from this session.
Fix
Check the 1.5 figure against the de Zambotti et al. 2014 full text and, if it is a morning-diary count rather than a per-flash report count, say which. If it cannot be sourced, replace the sentence with the figures that are verifiable: "Of 3.5 objectively recorded flashes per night, roughly 70% are accompanied by an arousal and about 20% pass without any disturbance to sleep (De Zambotti et al., 2014)." The 19.8% figure at line 149 is consistent with the review's "20%" and can stand.
What the refuter said
AGAINST: an unverifiable claim cannot be CONFIRMED, and by the same token there is no demonstrated error here — only an unchecked number. The affirmative evidence is thin: the de Zambotti abstract lists subjective frequency as a measure, so a per-night self-report count almost certainly exists in the full text, and 1.5 is entirely plausible for it. Gombert-Labedens summarising a study without repeating one of its numbers is weak evidence of anything. WHAT I CONFIRMED, mirroring the finding: the Europe PMC record for DOI 10.1016/j.fertnstert.2014.08.016 gives "Women had an average of 3.5 (95% confidence interval: 2.8-4.2, range = 1-9) objective hot flashes per night" and "A total of 69.4% of hot flashes were associated with an awakening", n=34 — and neither 1.5 nor 19.8% appears in the abstract. GombertLabedens2025.txt reports the same study as "(n = 34) had an average of 3.5 hot flashes per night and that 70% of hot flashes were associated with an arousal from sleep, with only a minority (20%) occurring without disturbance to sleep" with no reported-flash count. Full text unreachable from here (ScienceDirect robots-disallowed, PubMed and PMC captcha-gated), so UNVERIFIED is the honest label. Minor: a pre-submission check on a sentence that carries rhetorical weight in the Summary, not a defect.
The problem
Two problems. First, precision: what Osorio-Forero et al. measured was thalamic noradrenaline by fiber photometry, inversely correlated with sigma power; LC spiking was manipulated optogenetically, not recorded, and the paper's own claim is hedged as "suggesting". "Rises and dips with this phase" reads as in-phase and states as measurement what the paper infers. Second, and far more damaging: the correct sign means noradrenaline is HIGH in the low-sigma fragility phase. The draft's own cited review reports that raising brain noradrenaline pharmacologically provokes hot flashes and lowering it ameliorates them. Combining the two sources the draft already cites yields a specific prediction that flash *occurrence* clusters in fragility for purely thermoregulatory reasons. The draft treats phase-structured occurrence as a hypothetical alternative to be checked ("flash expression itself may be phase-structured"); it is in fact the mechanistically predicted outcome, and it makes the prerequisite aim likely to fire and the primary conditional analysis likely to be confounded.
In the draft — line 27-29 (with 138-141, 162-169)
In mice, noradrenergic locus coeruleus activity rises and dips with this phase (Osorio-Forero et al., 2021)
Evidence
OsorioForero2021.txt L32-33 (Highlights): "Thalamic NA fluctuates over ~50 s and is anticorrelated to sleep spindles". L267-269: "NA signals fluctuated in a manner inversely correlated with sigma power (with a time lag <0.5 s ...)". L274-277: "NA levels rose rapidly before sigma power declined, consistent with the suppressant effects of LC activity. Conversely, NA levels declined as sigma power was rising." L82 (hedge): "suggesting that both fore- and hindbrain-projecting LC neurons show coordinated infraslow activity variations in natural NREMS". The noradrenergic trigger for flashes: GombertLabedens2025.txt L874-882 (left column): "Freedman and colleagues [189] showed that yohimbine (an alpha-2 adrenergic antagonist that increases brain norepinephrine levels) provoked a hot flash, and that clonidine (an alpha-2 adrenergic agonist that reduces brain norepinephrine) ameliorated them, leading them to hypothesize that elevated levels of norepinephrine in the brain could narrow the inter-threshold zone in symptomatic individuals".
Fix
State the sign and own the prediction. Replace with: "In mice, thalamic noradrenaline released from the locus coeruleus fluctuates on the same ~50 s timescale in antiphase to sigma power, peaking as sigma declines (Osorio-Forero et al., 2021)." Then, in the aims, upgrade the occurrence question from prerequisite to co-primary and state the directional prediction explicitly: because brain noradrenaline both peaks in fragility and pharmacologically provokes flashes (Gombert-Labedens et al., 2025), flash occurrence is *predicted* to concentrate in fragility, and the registration analysis must be interpretable under that condition — which requires the design change in the collider finding below.
What the refuter said
All source quotes verified in OsorioForero2021.txt: L32 'Thalamic NA fluctuates over ~50 s and is anticorrelated to sleep spindles'; L267-269 'NA signals fluctuated in a manner inversely correlated with sigma power'; L274-277 the rise-before-decline sequence; L82 the 'suggesting' hedge. The yohimbine/clonidine passage verified verbatim in GombertLabedens2025.txt L874-882. So the facts are right. The finding still overreaches on both prongs. Prong 1: 'rises and dips with this phase' asserts covariation and specifies no sign, so there is no misstatement to correct; and the 2021 paper's own abstract and its own section heading ('LC activity fluctuates on an infraslow timescale during NREMS') make the LC claim the draft paraphrases, so paraphrasing a paper's headline claim is standard practice, not a citation error. Prong 2 is the substantive half and its key inferential step fails: selection on the exposure alone (conditioning on flashes that occurred, when occurrence depends on phase) does not bias the phase-registration association — that requires selection on a collider, which the finding never establishes. And the draft pre-empts phase-structured occurrence three times: L162-169 ('a separate testable outcome'), L181-182 (prerequisite aim), L294-297 (characterised first against stage- and time-matched surrogates). What survives is a sharpening note: the draft's own two sources jointly predict a direction for occurrence, and the draft could state it. Not decision-changing.
The problem
The review contains nothing about the conscious registration of nocturnal flashes. What it reports is (a) that 20-29% of objectively detected nocturnal flashes occur without disturbing sleep, and (b) that flashes occurring during sleep may be under-reported retrospectively in the morning. Neither is a statement about conscious registration: "no sleep disturbance" is an EEG-defined absence of arousal, and "under-reported in the morning" is an absence of recall. Conflating undisturbed sleep with unregistered experience is precisely the inferential step the proposed study is meant to test, so presenting it as an established citation begs the question the project asks. A reviewer who checks the reference will find the proposal's founding observation unsupported by the source given for it.
In the draft — line 37-39
Interestingly enough, just as with external signals, they are sometimes consciously registered, and other times leave no trace (Gombert-Labedens et al., 2025).
Evidence
GombertLabedens2025.txt L917-923 (right column): "with only a minority (20%) occurring without disturbance to sleep"; L925-928: "29% of hot flashes occurred in undisturbed sleep"; L858-861 (left column): "hot flashes that occur during sleep may be under-reported retrospectively in the morning [194,195]". No passage in the review addresses whether a flash is consciously registered at the time it occurs.
Fix
Separate what is cited from what is inferred: "Between 20 and 29% of objectively detected nocturnal flashes occur without any disturbance to sleep, and flashes occurring during sleep are under-reported in the morning (Gombert-Labedens et al., 2025). Whether the undisturbed and unreported flashes were registered at the time and forgotten, or never registered at all, is unknown — and is the question this project asks."
What the refuter said
I searched GombertLabedens2025.txt independently and the finding is right that no passage addresses whether a nocturnal flash is consciously registered at the time it occurs. The three cited passages are verified verbatim: 'with only a minority (20%) occurring without disturbance to sleep' and '29% of hot flashes occurred in undisturbed sleep' (both in the L917-928 block) and 'hot flashes that occur during sleep may be under-reported retrospectively in the morning [194,195]' (L858-861). STRONGEST DEFENCE, which carries real weight: 'sometimes consciously registered, and other times leave no trace' is a loose gloss on material the review does contain — undisturbed-sleep flashes plus retrospective under-reporting — and the very next sentence (L40-41) supplies report-gap evidence that is about registration/reporting. So the phenomenon is real and cited; what is mis-cited is the conceptual framing. The finding's own point still lands: the draft states as established the identity between 'no EEG disturbance' and 'not registered' that the project exists to test, and it does so in the sentence that motivates the whole proposal. A reviewer who checks the reference finds a gloss, not the claim. Downgraded from moderate to minor: it is a framing/attribution imprecision in one sentence, not a load-bearing empirical error, and the draft's design (prospective button press) correctly does not depend on it.
The problem
Two errors and an overstatement. (i) The cited review devotes a paragraph to ruling that "thermoneutral zone" denotes the zone between ambient-temperature thresholds while the zone between core-temperature thresholds — the one hypothesised to narrow in symptomatic women — is the inter-threshold zone. The draft uses the term the review specifically sets aside. A thermophysiology reviewer will read this as not having read the source. (ii) The projection target is incomplete: the review names both the median preoptic nucleus and the medial preoptic area. (iii) The narrowing is presented as fact where the review reports mixed evidence, and states that the functional relevance of KNDy innervation of the MnPO for normal thermoregulation has never been established.
In the draft — line 135-137
Oestrogen withdrawal hyperactivates arcuate KNDy neurons, which signal through the neurokinin 3 receptor to the median preoptic nucleus and narrow the thermoneutral zone (Gombert-Labedens et al., 2025).
Evidence
GombertLabedens2025.txt L598-603 (right column): "That zone is often referred to as the inter-threshold zone, but in some papers, it was called the TNZ. For clarity, here we refer to the zone between ambient temperature thresholds as the TNZ, and the zone between core temperature thresholds as the inter-threshold zone." L676-678: "KNDy neurons project from the ARC to key thermoregulatory areas within the POA, the median preoptic nucleus (MnPO) and the medial preoptic area". L744-748 (right column): "the functional relevance of KNDy innervation in the MnPO for normal thermoregulation has not, to our knowledge, been elucidated." Contested evidence: L805-820 records a study finding "only 51% of the hot flashes detected ... were preceded by any increase in gastrointestinal temperature" and "similar temperature thresholds for the onset of sweating ... between symptomatic postmenopausal and young, premenopausal individuals [180], challenging the idea that thresholds for sweating are lower in individuals symptomatic for hot flashes". The abstract (L82-83) itself flags "the mixed research findings about thresholds for sweating in symptomatic individuals".
Fix
Rewrite as: "Oestrogen withdrawal hyperactivates arcuate KNDy neurons, which signal through the neurokinin 3 receptor to glutamatergic neurons of the median preoptic nucleus and medial preoptic area in the central heat-defence pathway; a narrowed inter-threshold zone is the leading account of hot flash generation, though the evidence on sweating thresholds in symptomatic women is mixed (Gombert-Labedens et al., 2025)." Note that the draft's anatomy is otherwise correct — the median preoptic nucleus is the right structure, so only the terminology and the certainty need fixing.
What the refuter said
STRONGEST DEFENCE, PARTLY SUCCESSFUL: (ii) is wrong. The review states "The NKB signaling pathway in the MnPO plays a key role in menopausal hot flashes [109,137]", so naming only the median preoptic nucleus for the NK3R signalling step is the review's own emphasis, not an incomplete quotation — that sub-point is refuted. On (i) and (iii) the defence is weaker but real: "narrowed thermoneutral zone" is long-standing field usage (Freedman), so the draft is following convention while the review is making a local stipulation; and one background sentence in a 6,000-character literature review is not where causal hedging usually lives. WHY IT PARTLY SURVIVES: I verified the stipulation verbatim in GombertLabedens2025.txt: "That zone is often referred to as the inter-threshold zone, but in some papers, it was called the TNZ. For clarity, here we refer to the zone between ambient temperature thresholds as the TNZ, and the zone between core temperature thresholds as the inter-threshold zone." The draft cites that exact source for a sentence using the term the source sets aside — a thermophysiology reviewer will notice. I also verified the hedges: "the functional relevance of KNDy innervation in the MnPO for normal thermoregulation has not, to our knowledge, been elucidated"; and the challenge study ("only 51% of the hot flashes detected ... were preceded by any increase in gastrointestinal temperature"; "similar temperature thresholds for the onset of sweating ... challenging the idea that thresholds for sweating are lower"). DOWNGRADE: moderate to minor. It is a one-sentence terminology and hedging fix in background text; it changes nothing about the design, the measurement, or the inference.
The problem
The uncited sentence at L107 and the summary claim at L27-29 both assert LC activity where what exists is thalamic noradrenaline measured by fiber photometry plus causal optogenetic manipulation; the paper's own claim about LC neuronal activity is explicitly hedged as "suggesting". Recordings were confined to the mouse light phase (ZT1-9) in the first hours of rest, and the sigma band used was 10-15 Hz, not the 11-16 Hz the draft uses throughout. None of this undermines the mechanism, but a proposal whose team includes these authors should describe their own work precisely, and the band mismatch should be resolved since the draft's exposure is defined by the sigma envelope.
In the draft — line 107 (with 27-29, 58-60)
In rodents this fluctuation tracks locus coeruleus (LC) activity.
Evidence
OsorioForero2021.txt L82: "suggesting that both fore- and hindbrain-projecting LC neurons show coordinated infraslow activity variations in natural NREMS" (hedged). L120: "Summed sigma (10-15 Hz) and delta (1.5-4 Hz) power dynamics were derived from the time-frequency distributions." L143: "(ZT0-ZT12), which is their preferred resting phase"; L1008: "All optogenetic manipulation took place during the first 20 min of each hour between ZT1 and ZT9."
Fix
At L107 write "In mice, thalamic noradrenaline released from the LC fluctuates in antiphase to this rhythm, and timed optogenetic LC manipulation suppresses, locks or entrains it (Osorio-Forero et al., 2021)" — and add the citation, which is currently missing from that sentence. State the sigma band used in this study and reconcile it with the sources: if 11-16 Hz is chosen after Lecci et al., say so and report the primary analysis with 10-15 Hz as a sensitivity check.
What the refuter said
Every sub-claim verified in OsorioForero2021.txt: the hedge at L82; sigma defined as 10-15 Hz at L120 and L143-146; 'ZT0-ZT12, which is their preferred resting phase' at L143; 'All optogenetic manipulation took place during the first 20 min of each hour between ZT1 and ZT9' at L1008. I also confirmed by grep that the 2021 paper measured thalamic NA by fiber photometry and manipulated LC optogenetically — no LC unit recordings (the paper itself notes 'LC unit data from NREMS are scarce', L421). STRONGEST DEFENCE, and it disposes of most of the finding: L107 is an uncited topic sentence introducing Osorio-Forero et al. (2025) in the very next sentence, and the 2025 paper did record LC activity directly, so L107 is not a misattribution to the 2021 paper at all; and the 2021 paper's own title, abstract and section heading make the LC claim the draft paraphrases. Species-specific sigma bands are also normal. What survives, and is why I confirm rather than refute, is the band point — strengthened by a check the finding did not do: I fetched Lecci et al. 2017 and its HUMAN sigma band is also '10 to 15 Hz', with the infraslow peak at '0.019 ± 0.001 Hz'. So the draft's 11-16 Hz (L23, L96) matches neither of the two sources it cites for it, in a study whose exposure is the sigma envelope. Minor, as filed, but a real and checkable inconsistency.
The problem
Lecci et al. (2017) is the single load-bearing reference of the whole proposal — it supplies the infraslow sigma oscillation, the continuity/fragility framing and the memory claim — and it is cited at DRAFT.md lines 11, 27, 96, 106(implicitly) and 274, yet it is absent from the REFERENCES section (lines 375-445). Osorio-Forero et al. (2025) is cited at lines 60 and 108 but appears nowhere in REFERENCES either; it is listed only under PREVIOUS OWN PUBLICATIONS (lines 461-464), which is a CV section, not a bibliography. A reviewer cannot resolve either citation from the reference list.
In the draft — line 274
established in humans (Lecci et al., 2017; Lazar et al., 2019).
Evidence
REFERENCES (DRAFT.md 375-445) contains, in order: Baker 2019, Cataldi 2026, De Zambotti 2014, Esfahani 2023, Fernandez & Lüthi 2020, Freedman & Roehrs 2006, Freeman & Sherif 2007, Gombert-Labedens 2025, Kjaerby 2022, Lazar 2019, Osorio-Forero 2021, Osorio-Forero 2021 (duplicate), Rothhaas & Chung 2021, Sikder 2026, Tsiartas 2021, Wassing 2019. No Lecci entry; no Osorio-Forero 2025 entry. The correct citation is: Lecci, S., Fernandez, L. M. J., Weber, F. D., Cardis, R., Chatton, J.-Y., Born, J., & Lüthi, A. (2017). Coordinated infraslow neural and cardiac oscillations mark fragility and offline periods in mammalian sleep. Science Advances, 3(2), e1602026. https://doi.org/10.1126/sciadv.1602026 (see also the erratum, https://www.science.org/doi/10.1126/sciadv.aaq0565).
Fix
Add the Lecci 2017 entry as given in the evidence field, and add Osorio-Forero, A., Foustoukos, G., et al. (2025). Infraslow noradrenergic locus coeruleus activity fluctuations are gatekeepers of the NREM-REM sleep cycle. Nature Neuroscience, 28(1), 84-96. https://doi.org/10.1038/s41593-024-01822-0 to REFERENCES, keeping the PREVIOUS OWN PUBLICATIONS listing separate.
What the refuter said
STRONGEST DEFENCE: this is F10 restated at four times the severity, with sloppier evidence — it says Lecci is "cited five times" at "lines 11, 27, 96, 106(implicitly) and 274", but grep returns exactly four hits at 11, 27, 95 and 274; line 96 is off by one and "106 (implicitly)" is padding a count. Calling a missing bibliography entry FATAL is indefensible: it costs one line to fix, changes no measurement, no analysis and no inference, and cannot alter the study's validity. WHY THE FACT SURVIVES: the absence itself is real (verified by grep, same as F10), and I did verify the supplied replacement citation is genuine rather than fabricated — Lecci, Fernandez, Weber, Cardis, Chatton, Born & Lüthi (2017), "Coordinated infraslow neural and cardiac oscillations mark fragility and offline periods in mammalian sleep", Science Advances 3(2):e1602026, DOI 10.1126/sciadv.1602026, confirmed at PMC5298853. That is the one thing this finding adds over F10. DOWNGRADE: fatal to minor — duplicate of F10 with an inflated severity and two small citation-location errors. Report F10, drop this one.
The problem
'That nucleus' refers back to the median preoptic nucleus (MnPO), the KNDy target named in the preceding sentence. The cited review does discuss noradrenergic inhibition of preoptic sleep-promoting neurons, but the evidence it presents is for the ventrolateral preoptic area (VLPO) and for POA neurons defined by their projections to TMN and lateral hypothalamus — not the MnPO. Swapping VLPO evidence for an MnPO claim is the load-bearing step in the proposal's mechanistic story: it is what licenses 'a single infraslow rhythm times arousability and thermoeffector drive together'. The review also contains no discussion of KNDy neurons, neurokinin 3 receptor, oestrogen, menopause or hot flashes, so it cannot support the convergence argument as framed.
In the draft — line 138-141
That nucleus is also where autonomic thermoregulation and sleep-wake control converge, and it receives noradrenergic input, raising the possibility that a single infraslow rhythm times arousability and thermoeffector drive together (Rothhaas & Chung, 2021).
Evidence
Rothhaas & Chung (2021), Frontiers in Neuroscience 15:664781 (https://www.frontiersin.org/articles/10.3389/fnins.2021.664781/full): "LTS cells are GABAergic and are inhibited by wake-promoting substances such as noradrenaline (NA) and acetylcholine"; "POA neurons innervating histaminergic neurons in the TMN and Hcrt/Orx expressing neurons in the LH were potently inhibited by NA and 5-HT"; retrograde tracing "into the VLPO" found labelled neurons in "the TMN, raphe nuclei, ventrolateral medulla, and LC". Targeted search of the same text found no mention of noradrenergic input specifically to the median preoptic nucleus, and no mention of KNDy neurons, neurokinin 3 receptor, oestrogen, menopause or hot flashes.
Fix
Be precise about what is known and what is conjectured: 'The preoptic area, where autonomic thermoregulation and sleep-wake control converge, receives noradrenergic input from the locus coeruleus, and its sleep-promoting neurons are inhibited by noradrenaline (Rothhaas & Chung, 2021). Whether the median preoptic nucleus specifically — the KNDy target — carries that input is not established, so the possibility that one infraslow noradrenergic rhythm times arousability and thermoeffector drive together is a conjecture this study does not test.'
What the refuter said
AGAINST: the draft hedges explicitly — "raising the possibility that a single infraslow rhythm times arousability and thermoeffector drive together" (L139-141). A stated possibility is not a claim, and nothing in the analysis plan depends on it: the primary test is a direction-agnostic sin/cos contrast that needs no mechanistic premise. Calling this "the load-bearing step in the proposal's mechanistic story" overstates its function. MnPO noradrenergic innervation is also, as a matter of general neuroanatomy, uncontroversial — so the claim is probably true, merely attached to the wrong source. SURVIVES on the attribution. I checked Rothhaas & Chung 2021 (Front Neurosci 15:664781) directly: it contains no sentence pairing MnPO with noradrenaline/norepinephrine; the NA-inhibition evidence is VLPO-specific ("Two-thirds of neurons in the VLPO are LTS cells"; "LTS cells are GABAergic and are inhibited by wake-promoting substances such as noradrenaline (NA) and acetylcholine"); and it contains no mention of KNDy neurons, neurokinin 3 receptor, oestrogen, menopause or hot flashes. One half of the draft's sentence IS supported — the review does describe MnPO/MPO neurons whose "reactivation promoted sleep together with hypothermia", i.e. the thermoregulation/sleep convergence. Minor: swap in a primary anatomical source, or soften to VLPO/POA.
The problem
The amplitude claim is the paper's title and is well supported. The added contrast 'rather than on its mean level' asserts a comparison the abstract does not report making: nothing in it contrasts oscillatory amplitude against mean noradrenaline level as competing predictors. It is a small overstatement, but it is in a sentence whose whole rhetorical work is to establish that oscillatory structure per se — not tonic level — is what matters functionally, which is the analogy the proposal leans on for phase.
In the draft — line 111-113
Kjaerby et al. (2022) showed that the memory benefit of sleep depends on the oscillatory amplitude of norepinephrine rather than on its mean level.
Evidence
Kjaerby et al. (2022), Nature Neuroscience 25(8):1059-1070, DOI 10.1038/s41593-022-01102-9, abstract: "The amplitude of NE oscillations is crucial for shaping sleep micro-architecture related to memory performance: prolonged descent of NE promotes spindle-enriched intermediate state and REM sleep but also associates with awakenings, whereas shorter NE descents uphold NREM sleep and micro-arousals. Thus, the NE oscillatory amplitude may be a target for improving sleep in sleep disorders." No mean-level comparison is reported. Species: mice.
Fix
Trim to what the paper shows: 'in mice, the memory-relevant micro-architecture of sleep depends on the oscillatory amplitude of noradrenaline, with micro-arousals riding on the peaks and spindles on the descending phase (Kjaerby et al., 2022)'. The peak/descent detail is stronger for the proposal's argument than the mean-level contrast, and it is actually in the source.
What the refuter said
Strongest case against: "rather than on its mean level" is a fair gloss on the paper's thesis - if the operative variable is oscillatory amplitude, then it is not the mean, and the paper's own closing line ("the NE oscillatory amplitude may be a target for improving sleep") treats amplitude as the mechanism. It is one subordinate clause in a literature review and would not move any reviewer. That said, the finding is factually right and I verified it. I fetched https://www.nature.com/articles/s41593-022-01102-9 and read the abstract in full: "The amplitude of NE oscillations is crucial for shaping sleep micro-architecture related to memory performance: prolonged descent of NE promotes spindle-enriched intermediate state and REM sleep but also associates with awakenings, whereas shorter NE descents uphold NREM sleep and micro-arousals." No comparison of oscillatory amplitude against tonic or mean NE as competing predictors is reported anywhere. The draft therefore asserts a negative claim its source does not test, in the sentence whose rhetorical job is to license the phase analogy. Real but trivial - minor, exactly as the finding graded it.
The problem
Two distinct problems. First, the 3.5-versus-1.5 comparison is between an event-level physiological count and an aggregate recall count; it does not license the inference that ~43% of individual flashes are registered in real time, which is the base rate the primary outcome depends on and the quantity that most strongly drives power. Second, I could not verify 1.5 in any accessible source. The draft's own cited review, which reports the 3.5 figure, does not report 1.5, and warns explicitly about the retrospective nature of morning reports.
In the draft — line 39-41
Women have on average 3.5 objectively recorded flashes per night but only report 1.5 (De Zambotti et al., 2014).
Evidence
de Zambotti et al. 2014 abstract, retrieved verbatim (Semantic Scholar, DOI 10.1016/j.fertnstert.2014.08.016): "Women had an average of 3.5 (95%CI:2.8-4.2, range=1-9) objective hot flashes per night. 69.4% of hot flashes were associated with an awakening." The abstract lists subjective measures as "Subjective (frequency and bother)" and reports only correlations, no count of 1.5 - so the figure is UNVERIFIED, not disproved. Gombert-Labedens et al. 2025 (GombertLabedens2025.txt, p.~111) reports the same study's breakdown as "(n = 34) had an average of 3.5 hot flashes per night and that 70% of hot flashes were associated with an arousal from sleep, with only a minority (20%) occurring without disturbance to sleep" and cautions (p.~110) that "hot flashes that occur during sleep may be under-reported retrospectively in the morning". Power consequence (/root/grantreview/sim/08_yield_v2.py): required N for 80% power at OR=2 and 5 s jitter is 971 flashes at a 43% registration rate but 1,867 at 15% and 5,012 at 5%.
Fix
Rewrite as: "Women have on average 3.5 objectively recorded nocturnal flashes per night (de Zambotti et al., 2014), of which 69.4% are associated with an awakening and about a fifth occur without disturbing sleep. Retrospective morning counts are lower, but a morning count is not a per-event measure and cannot be used to estimate the per-event probability of prospective registration. That probability is the base rate of the primary outcome; it is estimated from the pilot button-press data and pre-registered, and power is reported across the range 0.05 to 0.30." Verify the 1.5 figure in the full text of de Zambotti et al. 2014 or drop it.
What the refuter said
AGAINST: the draft pre-empts the measurement half of this squarely. L215-219 separates the two constructs the finding says are conflated — "this time-stamped press is a prospective measure of conscious registration. A morning diary estimate gives a complementary measure of recall and reporting behaviour" — so the design never relies on morning recall for the primary outcome. And the finding's claim that 43% "is the base rate the primary outcome depends on and the quantity that most strongly drives power" attributes to the draft a number the draft never states: it gives no base rate anywhere (that is F87/F165's finding, not evidence for this one). SURVIVES as a framing overstatement. I verified the de Zambotti 2014 abstract via Europe PMC: "Women had an average of 3.5 (95% confidence interval: 2.8-4.2, range = 1-9) objective hot flashes per night" and "A total of 69.4% of hot flashes were associated with an awakening", n=34; neither 1.5 nor 19.8% appears in the abstract. GombertLabedens2025.txt summarises the same study (3.5/night, 70% with arousal, 20% undisturbed) with no reported-flash count. So L39-41 sets an event-level physiological count against an aggregate recall count, and L152-154 then builds the proposal's motivation on it ("This discrepancy has been treated as a measurement problem rather than as evidence about how endogenous signals gain access to awareness") without noting that part of the gap is plainly retrospective under-reporting. Minor: one hedging clause in the Summary, since the design already handles it.
Work that supports, conflicts with, or pre-empts the premise and is not cited.
The problem
The draft's operational definition — fragility = LOW sigma power — is not what Lecci et al. 2017 reported, and every subsequent HUMAN dataset places arousal-associated events at or just after the sigma peak, i.e. in what the draft calls continuity. (1) Lecci defined the phases by the DERIVATIVE, not the level: continuity = rising sigma, fragility = DECLINING sigma. (2) Dimitriades et al. found human microarousals concentrated at the sigma peak in all four age groups. (3) Carro-Domínguez et al. found human heart rate highest at the sigma peak, explicitly opposite to mice. If the arousal-prone phase in humans is the peak or the falling flank rather than the trough, the draft's pre-registered prediction ("flashes arriving in fragility (low sigma power) are predicted to be more likely registered") points the wrong way, and the Expected Outcomes section commits the proposal to it.
In the draft — line 22-27 (restated 96-103, 356-359)
That rhythm is an infraslow fluctuation of sigma-band (11-16 Hz) power at approximately 0.02 Hz, alternating roughly every 25 s between a continuity phase of high sigma power and a fragility phase of low sigma power. During continuity the sleeper is comparatively insulated from sensory input; during fragility the same stimulus more often produces an arousal (Lecci et al., 2017).
Evidence
Lecci et al. 2017, Sci Adv 3:e1602026 (https://www.science.org/doi/10.1126/sciadv.1602026): "wake-ups and sleep-throughs occur during declining and rising sigma power levels, respectively." — Dimitriades et al. (bioRxiv 2024.11.06.620875; published Sci Rep 2026, https://doi.org/10.1038/s41598-026-58423-z): "The highest percentage of microarousals occurred around the peak of the ISFS (corresponding to phase bin 6 and/or 7) compared to phase bin(s) occurring in the negative half wave of the ISFS for all age groups" (children p<0.0001; early adolescents p=0.028; late adolescents p<0.0001; young adults p<0.0001). — Carro-Domínguez et al. 2025, Nat Commun 16:2070 (https://doi.org/10.1038/s41467-025-57289-5): "HR was significantly higher during the peak phase of sigma power than during the rise or fall (p<=0.022)" and "This is in line with research in humans but, interestingly, not with findings in mice for which heart rate was low during high sigma power."
Fix
Stop equating fragility with low sigma power. Rewrite the premise as: continuity = rising sigma, fragility = declining sigma (Lecci's definition), and state that in humans the phase distribution of spontaneous arousals is currently reported to peak near the sigma maximum, so the sign of the effect is genuinely open. Replace the Expected Outcomes sentence with a directionally agnostic one, e.g.: "Because human microarousals have been reported to cluster near the sigma peak (Dimitriades et al., 2026) whereas rodent arousability is highest on the declining flank (Lecci et al., 2017), we pre-register the two-degree-of-freedom sin/cos test of any preferred phase as confirmatory and treat the direction as an empirical question." Note that the sin/cos model already supports this — only the prose commits to a direction.
What the refuter said
Strongest case against, and it is substantial: the confirmatory test is direction-agnostic by construction — "Circular phase enters each model as sin(phase) and cos(phase), giving a two-degrees-of-freedom test for any preferred phase without fixing the fragility-continuity boundary in advance" (L285-287) — and Expected Outcomes explicitly declines to commit: "A reciprocal pattern is equally informative ... a direction predicted rather than defining" (L357-361). So an inverted human direction does not invalidate the study, and "the Expected Outcomes section commits the proposal to it" is wrong. Fatal falls. What is confirmed and would be caught by any reviewer in this field: (1) the operational definition does not match its own cited source. Lecci (PMC5298853) states "offline periods correspond to raising, whereas fragility periods correspond to declining portions of the 0.02-Hz oscillation" and "wake-ups and sleep-throughs occur during declining and rising sigma power levels, respectively" — a derivative-based split, rotated ~90 degrees (~12.5 s) from the draft's level-based split at L22-25, which matters given the draft's own claim that a quarter-cycle of error destroys the effect. (2) Lecci itself says "the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained", contradicting the draft's L99-103 presentation of human fragility-arousability as established. (3) Both human datasets invert it: Dimitriades, "The highest percentage of microarousals occurred around the peak of the ISFS ... for all age groups"; Carro-Dominguez, "HR was significantly higher during the peak phase of sigma power than during the rise or fall (p<=0.022)" and "not with findings in mice".
The problem
The draft's entire primary predictor presupposes that the sigma envelope contains a genuine narrowband oscillation whose instantaneous phase is meaningful. Chen et al. 2025 (PNAS, n=1,025) directly tested that and found spindle timing is dominated by short-term refractory dynamics, with the ~50 s infraslow component a minor contributor. Band-pass filtering a point-process-driven envelope at 0.01-0.04 Hz and Hilbert-transforming it will return a phase estimate regardless of whether an oscillation is present — the filter manufactures one. This is the reviewer objection most likely to sink the proposal, and the draft neither cites the paper nor offers a surrogate-data control that would answer it. Note also that Lazar et al. 2019 never estimated phase at all, and Bergel et al. 2026 characterise the infraslow rhythm as "a slow modulation of multiple frequency bands, not a discrete oscillation itself."
In the draft — line 267-270
Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase.
Evidence
Chen S, He M, Brown RE, Eden UT, Prerau MJ (2025). Individualized temporal patterns drive human sleep spindle timing. PNAS 122(2):e2405276121, https://doi.org/10.1073/pnas.2405276121. Abstract: "Short-term (<15 s) temporal patterns of past spindle history are the main determinant of spindle timing, accounting for over 70% of the statistical deviance—surpassing the contribution of factors such as cortical up/down-state (slow oscillation phase), sleep depth, and long-term history (15 to 90 s, including ~50 s infraslow activity)"; "Fingerprint-like timing patterns, characterized by a refractory period followed by a period of increased spindle activity, which are highly individualized yet consistent night-to-night." — Bergel et al. 2026, Nat Neurosci 29:543-550, https://doi.org/10.1038/s41593-025-02159-y.
Fix
Add a pre-registered surrogate/null control to the ISF extraction section and cite Chen et al. 2025 explicitly. Concretely: for each participant, generate surrogate sigma envelopes by shuffling inter-spindle intervals while preserving spindle count and the individual refractory-period distribution (the control Lazar et al. 2019 used), and require the observed ISF spectral peak to exceed the surrogate distribution before phase is treated as meaningful. Add one sentence to the literature review: "Whether the infraslow sigma fluctuation is a continuous oscillation or an emergent consequence of individualised spindle refractoriness is contested (Chen et al., 2025); we therefore validate ISF presence against interval-shuffled surrogates per participant before any phase analysis."
What the refuter said
AGAINST: Chen et al. is about what predicts individual *spindle event* timing, not about whether a narrowband sigma-power modulation exists. Lecci settles the latter in humans with a genuine spectral peak — I verified 0.019 ± 0.001 Hz, n=27, at 0.001 Hz resolution — and Lazar 2019 replicates it independently. The finding's own second citation cuts against it: Bergel et al. is titled "Sleep-dependent infraslow rhythms are evolutionarily conserved across reptiles and mammals" (Nat Neurosci s41593-025-02159-y), which supports rather than undermines the draft's premise, and I could not verify its quoted "not a discrete oscillation itself" wording (UNVERIFIED). "The filter manufactures one" is a general worry about envelope analyses, not a result Chen establishes. SURVIVES as a literature gap. The paper is real and large — n=1,025 from the National Sleep Research Resource, short-term (<15 s) spindle history accounting for "over 70 percent" of timing variability, "vastly outweighing longer-term infraslow activity (~50 seconds) and slow oscillation phase" (PNAS 2025, 10.1073/pnas.2405276121, PMID 39772740). It is the most prominent recent challenge to the informativeness of the draft's exact predictor, it is uncited, and no surrogate-data control is offered against it. Downgraded from fatal to moderate: it demands a paragraph and a control analysis, not a redesign.
The problem
The draft's mechanistic account is entirely rodent, and it does not tell the reviewer that Lecci et al. themselves flagged the human case as unsettled, nor that two human studies have since addressed it — one of which documents a species reversal in exactly the brain-autonomic coupling the draft treats as unified. The cross-species inference is defensible but only if stated as an inference and anchored to the human data that now exist: pupillometry as an LC proxy across the infraslow sigma cycle (n=17, with auditory probes), and human VLF heart-rate variability phase-locked to the same rhythm and predicting spindle expression and overnight memory (n=28). By contrast, human LC MRI work measures neuromelanin contrast or awake activity, not infraslow dynamics during sleep — so pupillometry and VLF-HRV are the only available human handles.
In the draft — line 27-30
In mice, noradrenergic locus coeruleus activity rises and dips with this phase (Osorio-Forero et al., 2021); in humans, fragility is marked by absent sleep spindles and spontaneous arousals (Fernandez & Lüthi, 2020).
Evidence
Lecci et al. 2017 (https://pmc.ncbi.nlm.nih.gov/articles/PMC5298853): "the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained"; human sample was 27 healthy males, age 22.5 ± 0.49 y, with no acoustic stimulation experiment. — Carro-Domínguez M, Huwiler S, Oberlin S, et al. 2025, Nat Commun 16:2070, https://doi.org/10.1038/s41467-025-57289-5: "pupil size was smallest during the rise of sigma power and reached its peak when sigma power was falling"; species caveat quoted in finding human-arousals-at-sigma-peak. — Jacobsen et al. 2026, eLife reviewed preprint, https://doi.org/10.7554/eLife.110252.2 (mice n=7-10, humans n=28): "human sleepers show the same pattern: stronger VLF-HR fluctuations during NREM correspond to increased spindle expression and better overnight memory retention."
Fix
Add a short paragraph to the literature review and change the contingency plan. Text: "Human evidence for the noradrenergic account has recently emerged: pupil size, a non-invasive LC proxy, tracks the infraslow sigma cycle, being smallest as sigma rises and largest as it falls (Carro-Domínguez et al., 2025), and very-low-frequency heart-rate variability is phase-locked to the same rhythm in both mice and humans, predicting spindle expression and overnight retention (Jacobsen et al., 2026). Both studies also document species differences in cardiac coupling, so the rodent mechanism motivates rather than establishes the human case." Then use it: both study devices record PPG, so VLF-HRV gives an independent, peripherally derived estimate of infraslow phase — a far better contingency than the amplitude fallback (see finding amplitude-fallback-invalid).
What the refuter said
AGAINST: "entirely rodent" overstates. The draft cites Lecci 2017 partly for its human dataset and Lazar et al. 2019 for human infraslow sigma (L273-274), and a 6,000-character literature review cannot be exhaustive. SURVIVES, and the sources check out. Lecci 2017 verbatim: "the role of the 0.02-Hz oscillation for arousability in humans will need to be ascertained", with 27 healthy men and no acoustic stimulation. Carro-Domínguez et al. 2025, Nat Commun 16 (Lüthi among the authors), N=17 with auditory probes: "Pupil size was smallest during the rise of sigma power and reached its peak when sigma power was falling", and it reports a species divergence in the cardiac coupling the draft treats as unified. Jacobsen et al. 2026, eLife reviewed preprint 110252v2 (mice n=7-10; humans n=28, mean age 20.5): "stronger VLF-HR fluctuations during NREM correspond to increased spindle expression and better overnight memory retention." Two human handles on infraslow LC dynamics during sleep now exist, and one of them bears directly on the direction of the draft's central prediction — the pupil (LC proxy) peaks on the FALLING limb, not in the trough. With ~1,140 characters of Literature Review headroom available, this is addable. Moderate rather than major: an incomplete literature, not a broken design.
The problem
DRAFT.md contains zero occurrences of "CAP", "cyclic alternating pattern", "periodic limb movement", "PLM", "apnea", "Manconi", "Ferri" or "Silvani" (verified by grep over the full file). This omission is doubly damaging. First, a published commentary on Lecci et al. 2017 argues the 0.02 Hz sigma fluctuation may be a subharmonic of, or phase-locked to, the cyclic alternating pattern — a decades-old framework for infraslow NREM microstructure with its own arousability literature. A reviewer from that community will read the draft as reinventing CAP without acknowledgement. Second, and more importantly, the CAP literature already contains the closest analogue of the draft's own hypothesis: periodic limb movements — endogenous, frequent, objectively timed — synchronise with CAP phase A. That is powerful precedent support for the draft's premise that endogenous events are gated by infraslow NREM structure, and the draft cites none of it.
In the draft — line 44-46
Whether one also operates within NREM sleep, at the infraslow scale of roughly 50 seconds, has never been tested.
Evidence
Manconi M, Silvani A, Ferri R (2017). Commentary: Coordinated infraslow neural and cardiac oscillations mark fragility and offline periods in mammalian sleep. Front Physiol 8:847, https://doi.org/10.3389/fphys.2017.00847 — argues sigma-ISO "might thus be akin to a subharmonic of CAP, with a possible phase-locked synchronization such that one sigma-ISO cycle contains two CAP cycles", and "Understanding whether the sigma-ISO is indeed synchronized with CAP, and, if so, whether synchronization occurs with CAP phase A or B, would pave the way to develop better sleep quality indexes." — Mogavero MP, DelRosso LM, Lanza G, Bruni O, Ferini Strambi L, Ferri R (2025). J Sleep Res 34(2):e14265, https://doi.org/10.1111/jsr.14265: "The pronounced tendency of PLMS to synchronize with the A phase of CAP." — Original CAP/PLMS gate-control study: Clinical Neurophysiology 1996, PubMed 8858493 (abstract not retrievable through the available proxies; UNVERIFIED beyond title and journal). Grep confirmation: no CAP/PLM/apnea terms anywhere in DRAFT.md.
Fix
Add two or three sentences to the literature review; ~1,145 characters of headroom exist against the 6,000-character limit. Suggested: "Infraslow structuring of NREM is not a new claim: the cyclic alternating pattern describes a comparable alternation, and the sigma fluctuation may be synchronised with it (Manconi et al., 2017). That literature already supplies the closest precedent for the present hypothesis — periodic limb movements, endogenous and objectively timed, cluster in CAP phase A (Mogavero et al., 2025). What has not been tested is whether infraslow phase governs whether such an event is consciously registered, as opposed to whether it occurs." This converts the draft's largest citation gap into its strongest precedent, and pre-empts a hostile reviewer from that field.
What the refuter said
STRONGEST DEFENCE: the draft's novelty claim is narrow and survives intact — "Whether one also operates within NREM sleep, at the infraslow scale of roughly 50 seconds, has never been tested" (L44-46) is about a gate on the conscious registration of hot flashes, which the CAP literature has not tested. Missing related literature is the most common and least consequential grant criticism, and the design is unaffected. WHY IT SURVIVES: the grep is decisive (zero occurrences of CAP, cyclic alternating, periodic limb, PLM, apnea, Manconi, Ferri or Silvani in DRAFT.md) and I verified both key sources. Manconi, Silvani & Ferri (2017), Front Physiol 8:847, is a commentary written specifically about the draft's founding paper, and I confirmed verbatim: "The sigma-ISO might thus be akin to a subharmonic of CAP, with a possible phase-locked synchronization such that one sigma-ISO cycle contains two CAP cycles" and "whether synchronization occurs with CAP phase A or B, would pave the way to develop better sleep quality indexes". Mogavero et al. (2025) J Sleep Res 34(2):e14265 and the 1996 Clin Neurophysiol CAP/PLMS gate-control paper (PubMed 8858493) both exist. The second half of the finding is the valuable half and is framed correctly as SUPPORT, not attack: PLMS-CAP synchronisation is the closest existing analogue of the draft's own hypothesis — endogenous, frequent, objectively timed events gated by infraslow NREM microstructure — and citing it would strengthen the proposal's plausibility. The literature review has ~1,145 characters of headroom under its 6,000 limit. DOWNGRADE: major to moderate — a reviewer-community exposure and a missed opportunity, not a defect in the science.
The problem
The speculation is built on noradrenaline specifically, but acetylcholine, serotonin, dopamine, histamine and noradrenaline all show synchronised infraslow oscillations at ~0.02 Hz in NREM, with peak coherence at that frequency and no difference in peak frequency between them. Any of them, or a shared upstream driver, could time thermoeffector drive; nothing in the draft's argument selects noradrenaline. Worse, the direction is unfavourable as stated: the very review the draft cites reports that noradrenaline INHIBITS POA sleep-active neurons via α2 receptors, and states the anatomical overlap the argument requires is unknown. If high-NA (fragility) inhibits warm-sensitive/sleep-active POA neurons that drive heat-loss effectors, high NA would suppress rather than promote flash expression — the opposite of what the draft's continuity/fragility prediction implies.
In the draft — line 138-141
That nucleus is also where autonomic thermoregulation and sleep-wake control converge, and it receives noradrenergic input, raising the possibility that a single infraslow rhythm times arousability and thermoeffector drive together (Rothhaas & Chung, 2021).
Evidence
Kjaerby C, Radovanovic T, Kovács ER, et al. Coordinated infraslow cortical oscillations of neuromodulators during NREM sleep. iScience 2025, doi:10.1016/j.isci.2025.114554 (https://pmc.ncbi.nlm.nih.gov/articles/PMC12860995/): "All five neuromodulators examined, acetylcholine, serotonin, dopamine, histamine, and NE, exhibited synchronized infraslow cortical oscillations during NREM sleep"; "peak coherence at ∼0.02 Hz with no difference in peak frequency." — Rothhaas & Chung 2021, Front Neurosci 15:664781 (https://www.frontiersin.org/journals/neuroscience/articles/10.3389/fnins.2021.664781/full): "Electrical stimulation of the LC inhibited the activity in 50% of sleep-active neurons through α2-adrenoreceptors"; "the extent of anatomical overlap between warm-sensitive and sleep-promoting POA neurons is still largely unknown."
Fix
Downgrade the mechanism from noradrenergic to neuromodulatory and hedge the overlap. Rewrite as: "The median preoptic region receives dense monoaminergic input and partially overlaps sleep-regulatory populations, though the extent of that overlap is unresolved (Rothhaas & Chung, 2021). Because acetylcholine, serotonin, dopamine, histamine and noradrenaline all oscillate synchronously at ~0.02 Hz in NREM (Kjaerby et al., 2025), a shared infraslow neuromodulatory rhythm could in principle time arousability and thermoeffector drive together; the sign of any such coupling is not predictable a priori, since noradrenaline inhibits POA sleep-active neurons." This is more honest and it also justifies the draft's own reciprocal-direction hedge at lines 359-361.
What the refuter said
Strongest case against: the headline reasoning partly cuts the draft's way. Both sources are verified verbatim — Kjaerby et al., iScience 2025: "Coherence peaked at ~0.02 Hz with no difference in peak frequency between the neuromodulators highlighting the synchronized nature of these oscillations in the infraslow range"; Rothhaas & Chung: "Electrical stimulation of the LC inhibited the activity in 50% of sleep-active neurons through alpha2-adrenoreceptors" and "The extent of anatomical overlap between warm-sensitive and sleep-promoting POA neurons is still largely unknown." But the draft's actual claim is that "a single infraslow rhythm times arousability and thermoeffector drive together" — and five neuromodulators oscillating coherently at ~0.02 Hz *supports* that; it removes noradrenergic specificity, which the draft does not assert, writing only "it receives noradrenergic input, raising the possibility". The directional argument also needs an unstated extra step: that flash effectors are driven by POA sleep-active/warm-sensitive neurons, whereas the draft's stated mechanism is KNDy-to-MnPO narrowing of the thermoneutral zone (L135-137). And phase-structured *expression* is already flagged as a separate testable outcome (L162-169). What survives: the draft states convergence at the MnPO as fact where its own cited review calls the overlap unknown, and Kjaerby is the obvious citation to add. Moderate.
The problem
The power section is entirely about jitter and says nothing about the plausible magnitude of the effect being sought, so the minimum detectable effect has no benchmark to be judged against. The relevant benchmark exists: in wakefulness, conscious detection of near-threshold stimuli is modulated by cardiac and respiratory phase, and the effects are small — circular concentration R = 0.34 for hits against uniform, inter-beat interval differences of ≤5.2 ms. Those are the effect sizes a hostile reviewer will assume for a phase-gating effect on conscious registration, and they are small enough that the draft's power argument needs to engage with them. This literature is also the intellectual precedent for the draft's core framing and its absence makes the proposal look unaware of the field it is joining.
In the draft — line 315-320
The binding constraint is measurement precision rather than event count: onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle.
Evidence
Grund M, Al E, Pabst M, Dabbagh A, Stephani T, Nierhaus T, Gaebler M, Villringer A (2022). Respiration, Heartbeat, and Conscious Tactile Perception. J Neurosci 42(4):643-656, https://doi.org/10.1523/JNEUROSCI.0592-21.2021: "Tactile detection rate was highest during the first quadrant after expiration onset"; "tactile detection rate was higher during diastole than systole and newly specified its minimum at 250–300 ms after the R-peak"; reported R = 0.34 for hits versus uniform distribution (p = 0.007), ΔIBI ≤ 5.2 ms. Awake participants only — "No mention of sleep, anesthesia, or altered consciousness." Related: Al et al., PNAS 2020, https://doi.org/10.1073/pnas.1915629117.
Fix
Anchor the power simulation to a published effect size and cite the precedent. Add to Power and precision: "Effect sizes for phase gating of conscious detection in wakefulness are modest (circular R ≈ 0.34 for cardiac and respiratory phase modulation of near-threshold tactile detection; Grund et al., 2022); simulations therefore report the minimum detectable effect against that benchmark as well as against the observed jitter." Add one clause to the literature review noting that phase of a bodily rhythm already predicts conscious detection in wakefulness, so the question here is whether an infraslow brain rhythm does the same for an endogenous event in sleep.
What the refuter said
Strongest case against: this is a missing-citation complaint, and a proposal cannot cite everything; the draft does engage adjacent interoceptive work (Cataldi 2026 on heartbeat evoked potentials, L120-127). But the omission is verified and it is the wrong one to make. Grep of DRAFT.md finds no cardiac-phase, respiratory-phase, systole/diastole or near-threshold-detection citation anywhere; the only related term in the entire document is "heartbeat evoked potentials" at L121. Meanwhile the effect-size anchors are real and I verified them at source (J Neurosci 42(4):643): "Tactile detection rate was highest during the first quadrant after expiration onset"; "tactile detection rate was higher during diastole than systole and newly specified its minimum at 250-300 ms after the R-peak"; R = 0.34, p = 0.007 for hits versus uniform; inter-beat interval differences of 5.2 ms and 4.4 ms; awake participants. Two consequences make this decision-relevant rather than pedantic. The draft's power section names jitter as the binding constraint but states no expected effect magnitude at all, so its minimum detectable effect has nothing to be judged against — and this literature supplies the benchmark, at small values. And phase-gating of conscious detection by a bodily rhythm is the direct precedent for the draft's own framing, in exactly the interoception-and-awareness territory this funder backs.
The problem
The draft's claim to unique standing rests on the rodent work, but the same named collaborator is a co-author on the largest human characterisation of the infraslow sigma fluctuation to date (n=154, ages 8-26), published in Scientific Reports in 2026 — and that paper is nowhere in the draft. Two consequences. First, it weakens the team-strength argument: the team's human ISF credentials are stronger than the draft claims. Second, and more awkwardly, the omitted paper's headline result — arousal markers organised within the spindle-rich ISFS peak — is the one most in tension with the draft's directional prediction. A reviewer who notices that the applicant's own collaborator published the contradicting human result, and that the draft neither cites it nor addresses it, will draw an unflattering conclusion.
In the draft — line 58-61
Our team is uniquely placed to run this test: it includes the researchers who characterised the infraslow noradrenergic mechanism in rodents (Osorio-Forero et al., 2021, 2025)
Evidence
Dimitriades ME, Osorio-Forero A, Fattinger S, et al. The infraslow fluctuation of sigma power during sleep and its links to markers of arousal and memory reactivation across development. Sci Rep 2026, https://doi.org/10.1038/s41598-026-58423-z (PubMed 42310467): N=154, ages 8-26; "Electrophysiological markers of arousal and memory reactivation are organized within the spindle-rich ISFS peak"; preprint bioRxiv 2024.11.06.620875 gives the microarousal phase statistics. Grep of DRAFT.md returns no occurrence of "Dimitriades" or of this title.
Fix
Cite it, use it for the team-strength claim, and address its direction head-on. Rewrite the team sentence to include human ISF work: "...the researchers who characterised the infraslow noradrenergic mechanism in rodents (Osorio-Forero et al., 2021, 2025) and its developmental characterisation in humans (Dimitriades et al., 2026)." Then add to the literature review: "In the largest human characterisation to date, markers of arousal and memory reactivation were organised around the spindle-rich ISFS peak rather than the trough (Dimitriades et al., 2026), so the human phase-arousal mapping differs from the rodent one and the direction of any gating effect must be treated as open." This turns the omission into the proposal's justification for a two-sided test.
What the refuter said
STRONGEST DEFENCE, PARTLY SUCCESSFUL: the draft pre-empts the directional objection explicitly. L357-361 states "A reciprocal pattern is equally informative: because exteroceptive and interoceptive processing are not weighted uniformly across vigilance states, registration may instead be more probable during continuity, a direction predicted rather than defining." A 2-df sin/cos test with no pre-fixed boundary (L285-287) is direction-agnostic by construction, so a contrary human result does not threaten the design. And the Dimitriades sample is 154 participants aged 8-26, not perimenopausal women. WHY IT SURVIVES: I verified the paper exists and the result is as stated — Dimitriades ME, Osorio-Forero A, Fattinger S, et al., Sci Rep 2026, DOI 10.1038/s41598-026-58423-z, N=154, ages 8-26, abstract: "electrophysiological markers of arousal and memory reactivation are organized within the spindle-rich ISFS peak"; the preprint is more explicit still: "The highest percentage of microarousals occurred around the peak of the ISFS ... compared to phase bin(s) occurring in the negative half wave", consistently across all age groups. Grep confirms "Dimitriades" appears nowhere in DRAFT.md. The tension is not with the draft's hedged PREDICTION but with its stated PREMISE at L29-30 and L102-103 ("in humans, fragility is marked by absent sleep spindles and spontaneous arousals"; "in humans the fragility phase is associated with spontaneous arousals") — which the largest human characterisation to date, co-authored by the collaborator whose credentials the draft invokes at L58-61, points the other way on. That is exactly the paper an assessor will find, and the draft neither cites it nor addresses it. Severity stays moderate.
The problem
Gombert-Labedens et al. (2025) — which the draft cites twice as authoritative — assesses the state of automated flash detection and places Tsiartas 2021 firmly in the "research is ongoing to develop" category, in a passage stating that such techniques are mostly in the development phase and have not been applied to track flashes even in clinical trials. Fiona C. Baker is a co-author of both papers, so this is the detector's own authorship team characterising it as immature. A reviewer who reads both cited sources will find the draft treating as a settled measurement instrument the thing its own cited review calls unproven. The review's future-directions table reinforces this: it lists validating an objective flash definition and advancing wearable detection as open needs.
In the draft — line 242-244
Detection follows Tsiartas et al. (2021): wrist skin conductance, temperature, pulse-rate and motion features in a decision-tree classifier
Evidence
GombertLabedens2025.txt L901-908 (left column): "these techniques are mostly in the development phase and have not been applied to track hot flashes in clinical trials. Research is also ongoing to develop algorithms that can detect the onset of hot flashes based on a rise in skin conductance [204] or other characteristics of the physiological hot flash that reflect the thermoregulatory response [205]." Reference [205] is the cited detector — L2055: "[205] Tsiartas A, Baker FC, Smith D, et al. A novel hot-flash classification algorithm via multi-sensor features integration." See also L1388-1392: "there is a need to advance the detection of objective hot flashes that can be developed and applied in wearable and non-wearable devices."
Fix
Acknowledge the status openly and turn it into a strength rather than let a reviewer find it: "Automated wearable flash detection is still in development (Gombert-Labedens et al., 2025); we therefore treat detector validation on our hardware as a stated deliverable of this project rather than a premise, and report subject-wise performance, false-alarm rate and onset-timing error as feasibility outcomes alongside the primary test."
What the refuter said
STRONGEST DEFENCE: the draft reports Tsiartas's own published metrics accurately, notes that sleep-vs-wake performance was "reported separately" (true — Fig. 4), and never claims the detector is validated on EmbracePlus or in perimenopausal women. Reviews routinely describe every method as needing more work; that is not a disqualification. WHY IT SURVIVES: I verified the passage in GombertLabedens2025.txt: "these techniques are mostly in the development phase and have not been applied to track hot flashes in clinical trials. Research is also ongoing to develop algorithms that can detect the onset of hot flashes based on a rise in skin conductance [204] or other characteristics of the physiological hot flash that reflect the thermoregulatory response [205]." I confirmed [205] in the reference list: "Tsiartas A, Baker FC, Smith D, et al. A novel hot-flash classification algorithm via multi-sensor features integration." I also confirmed the future-directions row "Validating an objective definition of hot flashes": "there is a need to advance the detection of objective hot flashes ... Validated measures need to be developed that can be applied in wearable and non-wearable devices." And Tsiartas's own conclusion says "initial feasibility". Fiona C. Baker is indeed an author of both. So the draft cites one source as authoritative for its mechanism while treating as settled the instrument that same source classes as unproven. That asymmetry is a legitimate reviewer catch. DOWNGRADE: major to moderate — it is a disclosure/framing defect layered on the substantive detector-transfer problem already captured at higher severity in F160.
The problem
The literature is mixed and the draft states one side as fact, then builds a model covariate on it. De Zambotti et al. (2014) — which the draft cites twice for its headline prevalence numbers — found no first-versus-second-half difference, and a further study found a similar pattern in both halves. Using a contested finding to justify a covariate is defensible; presenting it as settled while citing, elsewhere in the same document, a paper that contradicts it is not, and a reviewer who reads both citations will notice. Worse, the review reports a study titled "Lack of sleep disturbance from menopausal hot flashes" finding awakenings more likely to occur *before* than after a flash. If arousal frequently precedes the flash, the causal arrow in the supporting aim is reversed and cortical arousal becomes a cause of detection rather than a consequence of the flash — a threat to the central logic that the draft never addresses.
In the draft — line 143-147 (with 288-291)
Freedman and Roehrs (2006) showed that hot flashes are suppressed during REM sleep, when thermoregulatory effector responses are largely inactive, and that the flash-arousal ordering reverses across the night: flashes precede arousals in the first half, arousals precede flashes in the second.
Evidence
GombertLabedens2025.txt L909-916 (right column): "One study found that awakenings were more likely to occur before, than after, a hot flash [214], and another found that hot flashes occurred before an awakening only in the first half of the night [215], whereas others found that the majority of hot flashes coincide with awakenings [179,194,216,217] with no differences in the first and second part of the night [194]." Reference [194] is de Zambotti 2014 (L1994-1998); [215] is Freedman & Roehrs 2006 (L2034-2037); [214] is L2030-2033: "[214] Freedman RR, Roehrs TA. Lack of sleep disturbance from menopausal hot flashes. Fertil Steril. 2004;82(1):138-144." And L962: "with a similar pattern in both halves of the night [220]".
Fix
State the disagreement and keep the covariate as a hedge, not a consequence: "Reports of the flash-arousal ordering conflict: one study found flashes preceding awakenings only in the first half of the night, others found no half-of-night difference or awakenings preceding flashes (Gombert-Labedens et al., 2025). We include half of the night as a covariate because of that disagreement, not because a reversal is established." Separately, add an explicit analysis addressing the reverse arrow: report the sign and distribution of the arousal-onset-to-flash-onset interval, and pre-specify exclusion of flashes whose onset follows an arousal by less than a fixed interval from the arousal outcome model.
What the refuter said
Strongest case against: the draft's statement is accurate to its cited source. I fetched Freedman & Roehrs 2006 (https://pubmed.ncbi.nlm.nih.gov/16837879/): first half "Most hot flashes preceded arousals and awakenings", second half "these patterns reversed". So the draft is not misciting. And the draft hedges in the methods section where the covariate is actually justified: "the flash-arousal relationship *may* reverse across the night" (L289-291). Including a half-of-night covariate is harmless whether or not the effect is real, so on the finding's stated axis ("using a contested finding to justify a covariate") there is no defect at all. What survives, and is worth reporting, is the second half of the finding, which I verified verbatim in GombertLabedens2025.txt (L909-916): "One study found that awakenings were more likely to occur before, than after, a hot flash [214]... whereas others found that the majority of hot flashes coincide with awakenings [179,194,216,217] with no differences in the first and second part of the night [194]", with [214] = Freedman & Roehrs 2004 and [194] = de Zambotti 2014, a paper the draft cites twice. If arousal frequently precedes the flash, the supporting aim's causal direction inverts and cortical arousal becomes a cause of detectability rather than a consequence - the draft nowhere addresses this. Downgraded to moderate: the covariate complaint is a non-issue, the causal-direction threat is real but is a discussion-level omission the applicants can answer.
The problem
The draft states the reversal as fact and builds a covariate on it. Gombert-Labedens 2025, which the draft cites elsewhere as authoritative, reports the finding as one of several mutually inconsistent results, including a study that found the same ordering in both halves of the night. Presenting a contested finding as established, while citing a review that documents the disagreement, is the kind of thing a hostile reviewer checks first.
In the draft — line 146-149
the flash-arousal ordering reverses across the night: flashes precede arousals in the first half, arousals precede flashes in the second.
Evidence
Gombert-Labedens 2025 (/root/grantreview/GombertLabedens2025.txt ~line 967-999): "One study found that awakenings were more likely to occur before, than after, a hot flash [214], and another found that hot flashes occurred before an awakening only in the first half of the night [215], whereas others found that the majority of hot flashes coincide with awakenings [179,194,216,217] with no differences in the first and second part of the night [194]"; and "the majority of hot flashes (80%) occurred before or concurrently with an awakening, with a similar pattern in both halves of the night [220]".
Fix
Rewrite as: "Reports on the temporal ordering of flashes and arousals are inconsistent: Freedman and Roehrs (2006) found flashes preceding arousals in the first half of the night and the reverse in the second, while others report no difference between halves (reviewed in Gombert-Labedens et al., 2025). Half of the night is therefore included as a covariate on empirical rather than mechanistic grounds." Keep the covariate; drop the false certainty.
What the refuter said
AGAINST: the draft hedges exactly where it matters. The Methods say "half of the night - the last because the flash-arousal relationship MAY reverse across the night (Freedman & Roehrs, 2006)" (L289-291), and including a half-of-night covariate is correct whether or not the reversal is real, so no analytic decision rides on the claim. The reversal is also a genuine reported finding in Freedman & Roehrs 2006, which is the source the draft attributes it to — not to Gombert-Labedens. SURVIVES for the Literature Review sentence. I confirmed the passage in GombertLabedens2025.txt: "found that awakenings were more likely to occur before, than after, a hot flash [214], and another found that hot flashes occurred before an awakening only in the first half of the night [215], whereas others found that the majority of hot flashes coincide with awakenings [179,194,216,217] with no differences in the first and second part of the night [194]", plus "with a similar pattern in both halves of the night [220]". So L145-149 states as settled fact ("the flash-arousal ordering reverses across the night") something the review the draft treats as authoritative documents as contested. Severity cut from moderate to minor: it is a one-word fix ("has been reported to reverse"), and the consequence is nil because the covariate is warranted either way.
The problem
The draft reaches for Rothhaas & Chung 2021 to argue that a noradrenergic rhythm could time thermoeffector drive — a review that never mentions hot flashes. Meanwhile the pharmacological evidence that hot flashes are noradrenergically triggered is on the same page of Gombert-Labedens et al. 2025 that the draft cites for KNDy signalling: an α2 antagonist that raises brain noradrenaline PROVOKES a flash, an α2 agonist that lowers it ameliorates flashes, and the resulting hypothesis is that elevated brain noradrenaline narrows the inter-threshold zone. That is a far stronger and more specific link between noradrenergic tone and flash initiation than anything in the draft, and it directly addresses the question of whether flashes are centrally initiated in a way that could be phase-locked.
In the draft — line 134-137
Oestrogen withdrawal hyperactivates arcuate KNDy neurons, which signal through the neurokinin 3 receptor to the median preoptic nucleus and narrow the thermoneutral zone (Gombert-Labedens et al., 2025).
Evidence
Gombert-Labedens et al. 2025 (source file /root/grantreview/GombertLabedens2025.txt, lines 873-882): "Freedman and colleagues [189] showed that yohimbine (an alpha-2 adrenergic antagonist that increases brain norepinephrine levels) provoked a hot flash, and that clonidine (an alpha-2 adrenergic agonist that reduces brain norepinephrine) ameliorated them, leading them to hypothesize that elevated levels of norepinephrine in the brain could narrow the inter-threshold zone in symptomatic individuals [18,190], or could be a separate or related trigger for hot flashes [182]." Also (lines 965-972): "the strong overlap in timing between many (but not all) hot flash events and awakenings could also reflect a common mechanism within the central nervous system in response to estrogen withdrawal, involving central sympathetic activation or the KNDy neuron network."
Fix
Use it. Insert after the KNDy sentence: "Pharmacological evidence already implicates central noradrenergic tone in flash initiation: an α2 antagonist that raises brain noradrenaline provokes flashes and an α2 agonist that lowers it suppresses them, supporting the hypothesis that elevated brain noradrenaline narrows the inter-threshold zone (Freedman, reviewed in Gombert-Labedens et al., 2025). If flash threshold is noradrenergically set and NREM noradrenaline fluctuates on a ~50 s cycle, flash timing and arousability may share a driver." Trace and cite the Freedman primary sources directly rather than via the review; the review's reference [189] gives the pointer.
What the refuter said
Strongest case against: this is not a defect at all. It reports no error, no misquote and no unsupported claim - it says a stronger available argument went unused, in a section under a 6,000-character cap. Grant reviews that report unexercised options as findings dilute the ones that matter. The quotation is genuine: I read GombertLabedens2025.txt at the passage and confirmed verbatim "Freedman and colleagues [189] showed that yohimbine (an alpha-2 adrenergic antagonist that increases brain norepinephrine levels) provoked a hot flash, and that clonidine (an alpha-2 adrenergic agonist that reduces brain norepinephrine) ameliorated them, leading them to hypothesize that elevated levels of norepinephrine in the brain could narrow the inter-threshold zone in symptomatic individuals", and the second quote about a "common mechanism within the central nervous system" involving "central sympathetic activation or the KNDy neuron network" also checks out. So the observation is accurate and the suggested strengthening is genuinely better than the Rothhaas & Chung bridge. But the claim that Rothhaas & Chung "never mentions hot flashes" is UNVERIFIED - I did not obtain that paper. PLAUSIBLE at minor: worth passing on as a constructive note, not as a finding.
The cohort, the yield, the cost, and what a grant application needs and this lacks.
The problem
The study's only EEG is the ZMax headband, whose derivations are F7-Fpz and F8-Fpz — bilateral frontopolar, referenced to Fpz. Two independent human studies report that the infraslow sigma fluctuation is maximal over centro-parietal cortex and minimal frontopolarly. Lazar et al. 2019 — cited by the draft as feasibility support — is the very paper reporting that frontopolar electrodes have the lowest infraslow power of any region. Dimitriades et al. independently report the same topographic dissociation, and explicitly warn that spindle DENSITY is maximal frontally while the ISFS is not. A small mean bias in average sigma bandpower (Esfahani) is a statement about static power, not about the fidelity of the sigma ENVELOPE at 0.01-0.04 Hz, which is the quantity the whole analysis depends on. The two cited feasibility supports therefore do not support feasibility, and one of them argues against it.
In the draft — line 270-276
Feasibility is supported: sigma-band power from this headband shows the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023), and the infraslow modulation of sigma is established in humans (Lecci et al., 2017; Lazar et al., 2019).
Evidence
Lazar, Dijk & Lazar 2019, J Neurosci Methods 316:22-34 (https://pmc.ncbi.nlm.nih.gov/articles/PMC6390176/): "Frontopolar region had significantly lower integrated EIP compared to all other brain regions"; "ISO over the centro-parieto-occipital region features higher intensities and lower frequencies." — Dimitriades et al., bioRxiv 2025.04.23.650209 (https://doi.org/10.1101/2025.04.23.650209): "the presence and strength of the ISFS peak in central-parietal regions, while sleep spindle density is maximal in frontal regions"; group reductions were "particularly in central-parietal electrodes." — Esfahani et al. 2023 (https://www.biorxiv.org/content/10.1101/2023.08.18.553744v1): ZMax uses "two frontal channels (F7-Fpz, F8-Fpz)" and "consistently underestimated the number of detected...spindle counts" (919.97 ± 357.53 vs 1008.7 ± 255.11).
Fix
This is the single largest feasibility risk and the draft currently conceals it behind a mis-aimed citation. Either (a) add at least one centro-parietal derivation — a second wearable, or a subset of nights with a portable montage extending to Cz/Pz — or (b) reframe frontal ISF expression from a "feasibility outcome" to a hard pre-registered gate with an explicit stopping rule, and cite Lazar 2019 and Dimitriades 2025 honestly as the reason. Delete "Feasibility is supported" and replace with: "Frontal ISF expression is the principal feasibility risk: the ISF is strongest centro-parietally and weakest frontopolarly (Lazar et al., 2019; Dimitriades et al., 2025), while the ZMax provides only F7-Fpz and F8-Fpz. We therefore quantify frontal ISF presence and strength per participant against a pre-registered threshold before any outcome analysis."
What the refuter said
Strongest case against: the draft never asserts frontal ISF is established - it names "Frontal expression and achievable phase precision" as feasibility outcomes (L275-276) - and attenuation is not absence, so the primary hypothesis is degraded rather than untestable. That is what takes this off fatal. Everything else verified, from three independent sources, and this finding states it accurately ("weakest", not "never demonstrated"). Lazar et al. 2019 (https://pmc.ncbi.nlm.nih.gov/articles/PMC6390176/): "Frontopolar region had significantly (adjusted P < 0.05) lower integrated EIP compared to all other brain regions except the temporal brain region", ranking parietal > occipital > central > frontal > temporal > frontopolar. Lecci et al. 2017: maximum over parietal, "declined toward anterior central and frontal areas". Dimitriades et al. (https://www.biorxiv.org/content/10.1101/2025.04.23.650209v1.full): "the topography of the strength of the ISFS closely aligns with the topography of the presence of the ISFS, with maximal values over central-parietal regions (specifically C3, Cz, C4, P3, Pz, P4)", while "sleep spindle density is maximal in frontal regions in our population" - so frontal spindle abundance is not evidence of frontal ISFS, which forecloses the obvious rebuttal. Esfahani's spindle undercount also verified verbatim (919.97 +- 357.53 vs 1008.7 +- 255.11). The draft cites two of these papers as feasibility support while one of them ranks its recording site last. Major.
The problem
The cited own-publication is not an observational cohort. van Baarzel et al. (2026) is a four-arm 2×2 RCT: CBCTi alone, MHT alone, CBCTi+MHT, control. Two of four arms receive "estradiol patches (Systen, 50mcg) and progesterone tablet (Utrogestan, 200mg)" for 15 weeks. The draft's assertion that participants "take no hormone-altering medication" is true only of the trial's *entry* criteria ("Already use MHT" and "Use hormonal contraceptives" are exclusions), not of the sample during the measurement windows. Roughly half the participants are on transdermal estradiol at T1 (week 8) and T2 (week 15). This is not a documentable co-medication to be adjusted for — it is the study drug, and it suppresses the primary event. The draft compounds this by claiming the work is done "in healthy humans and without intervention" (line 369-370), which is false as written. The CBCTi arms are equally damaging in a different way: sleep-restriction and circadian therapy change NREM continuity, spindle expression and arousal thresholds, i.e. they perturb the *predictor* (ISF sigma phase) as well as the outcome. In an unblinded trial ("the participants and study team members will not be blinded"), expectancy also acts directly on the self-report outcome.
In the draft — line 197-201 (also 61-64, 369-370)
Participants are perimenopausal women (irregular cycles and climacteric symptoms) reporting sleep disturbances, from an ongoing cohort of 222 women already being recorded. They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms.
Evidence
medRxiv protocol, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: four conditions (CBCTi, MHT, CBCTi+MHT, control); "MHT consists of estradiol patches (Systen, 50mcg) and progesterone tablet (Utrogestan, 200mg)… over the course of 15 weeks"; exclusions include "Already use MHT" and "Use hormonal contraceptives"; "Due to the nature of the interventions, the participants and study team members will not be blinded to the condition a participant is allocated to." Effect size of the intervention on the outcome event, from the draft's own cited review: Gombert-Labedens et al. 2025, /root/grantreview/GombertLabedens2025.txt lines 1312-1317: "In placebo controlled, randomized trials, MHT has been shown to reduce the frequency and severity of hot flashes, for example, reducing the frequency of symptoms by 75%"; line 1307: "The onset of efficacy for MHT is two to three weeks"; lines 1317-1318: "MHT has additional benefits including improvements in mood, sleep".
Fix
Restrict the analysis explicitly and exclusively to the pre-treatment measurement block, and say so in one sentence: "Analyses use the baseline (T0) measurement week only, before any study medication takes effect; the 40 EEG nights recorded at weeks 8 and 15 under randomised MHT or CBCTi are excluded from the confirmatory model." Replace lines 199-200 with: "Participants are free of hormone-altering medication at study entry (hormonal contraceptive use and prior MHT are exclusion criteria of the host trial); the analysis uses only the pre-randomisation measurement block, so no participant contributes data while on study medication." Delete "and without intervention" from line 370, or replace with "without any experimental manipulation of the events studied". Add a named sensitivity analysis using T1/T2 nights with arm as a covariate, framed as exploratory only.
What the refuter said
AGAINST: as with F38, the entry-criterion reading of L199 is available (the protocol excludes women who "already use MHT" and hormonal contraceptives), a T0-only analysis would be pharmacologically clean given MHT's two-to-three-week onset, and the CBCTi argument is speculative — the proposal's contrast is within-subject and within-night, so a therapy that shifts overall NREM continuity mostly shifts a nuisance level absorbed by participant and night random intercepts. SURVIVES, and this is the best-evidenced finding in the batch. Verified from the protocol: "MHT consists of estradiol patches (Systen, 50mcg) and progesterone tablet (Utrogestan, 200mg)" over 15 weeks, four arms, recordings at T0/T1/T2, and "the participants and study team members will not be blinded". Verified from the draft's own cited review (GombertLabedens2025.txt L1312-1317): MHT reduces flash frequency "by 75%". So the intervention removes the outcome event, not merely adds noise, and L369-370's "without intervention" is false as written. One correction to the finding: I checked the protocol's sequence and allocation precedes the T0 week ("they sign informed consent after which they will be allocated"), so even baseline is post-randomisation and unblinded — expectancy reaches the self-report outcome from the start, though the drug effect does not. Downgraded from fatal to major: fixable by naming the timepoint, at the cost of two-thirds of the recording blocks.
The problem
The host protocol contains no button press, event marker, or any nocturnal participant-triggered annotation of hot flashes; it describes only questionnaires (GCS vasomotor items, HFRDIS), a consensus sleep diary, and "objective occurrence of nocturnal hot flashes… estimated from a multisensory approach". Nor does it contain the five-tap synchronisation marker. Both are additions to the host protocol. Two consequences the draft never addresses. (i) Ethics: adding a participant-facing nocturnal task to an approved interventional protocol requires a substantial amendment to METC AUMC approval NL87156.018.24 and re-consent; the draft's ethics paragraph asserts existing approval and leaves the number as a placeholder. (ii) Yield: the button press can only be collected from participants recruited *after* the amendment is approved. If the cohort is, as the draft says, "already being recorded", the primary outcome is missing for everyone already run — and the draft never states how many that is. A confirmatory registration model with no registration variable is not a study.
In the draft — line 215-218
On EEG nights participants press a wristband button as soon as they think she is having a flash; this time-stamped press is a prospective measure of conscious registration.
Evidence
medRxiv protocol full text (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full) contains no reference to a button press, event marker, event button, tap marker, or any participant-triggered marking of hot flashes; hot flash self-report is not described at all beyond questionnaires and the consensus sleep diary. Protocol wearable text: "participants will also be given a smartwatch (Embraceplus) that can measure among others, heartrate, electrodermal activity, and physical activity for seven days at each timepoint". Ethics: "Medical Research Ethics Committee of Amsterdam University Medical Center… NL87156.018.24"; trial NCT06306404. Bial regulation, /root/grantreview/BIAL_CONTEXT.md line 35: "Human subjects / sensitive data: documentary evidence of submission to the competent Ethics committee, plus a signed PI declaration."
Fix
State the amendment status explicitly and give the numbers, e.g.: "The event button and tap synchronisation are additions to the host protocol, submitted to METC AUMC as amendment [n] to NL87156.018.24 on [date] and approved on [date]. They apply to participants enrolled from [date]; N participants have completed the baseline measurement week under the amended protocol to date, and N remain to be recruited." If the amendment is not yet approved, say so and make the button press a prospective collection that Bial is being asked to fund. If it cannot be added, demote conscious registration to the morning diary count and rewrite the primary aim accordingly — but then say clearly that the primary outcome is a next-morning recall estimate, not a prospective within-night report.
What the refuter said
AGAINST: the inference is not proven. A 10-page protocol paper is not the METC dossier; EmbracePlus carries a tag/event button as standard, so a marking instruction adds no hardware and could plausibly have been approved from the outset; and the draft's Procedure section is written prospectively ("Participants complete online surveys ... then an instruction session on device handling, self-report and synchronisation", L210-212), which reads as describing what the funded project will do rather than what has already been done. So "does not exist in data already being recorded" is an inference, not a verified fact. But the confirmed core is serious and is readable off the draft alone. I verified the parent protocol enumerates hot-flash measurement objectively and only — "The objective occurrence of nocturnal hot flashes will be estimated from a multisensory approach (skin conductance, skin temperature, heart rate and movement) assessed with the EmbracePlus smartwatch" — with no button press, event marker, tap marker or morning hot-flash diary anywhere, and with device participation optional and revocable ("Participants are free to opt out of EEG and smartwatch measurements at any time"). Against that, the draft asserts a cohort "already being recorded" (L198) whose confirmatory outcome variable (L215-218) appears nowhere in the parent design, never states how many participants are already run, never says whether an amendment is needed, and leaves the ethics reference as "[number]" (L331-332). Major: a confirmatory study must state that its primary outcome will exist.
The problem
The 222-woman cohort is the team's own 'Sleeping Through Menopause' trial, listed under PREVIOUS OWN PUBLICATIONS (DRAFT.md 449-454). That trial randomises exactly 222 women to menopausal hormone therapy (MHT), so roughly half receive oestradiol plus progesterone — the most effective known suppressor of vasomotor symptoms. The recruitment section's claim that participants take no hormone-altering medication is therefore false for two of the four arms at the T1 and T2 recording timepoints, and the draft nowhere states that it will use only the pre-randomisation T0 recordings. Because the outcome is the count and consequence of nocturnal hot flashes, MHT does not merely add noise: it removes the events the study is powered on. Either the analysis is restricted to T0 (in which case the event yield is one four-night block per woman, not the multi-timepoint total the design implies) or it is confounded by an active anti-flash intervention. The draft resolves neither.
In the draft — line 196-199
Participants are perimenopausal women (irregular cycles and climacteric symptoms) reporting sleep disturbances, from an ongoing cohort of 222 women already being recorded. They are physically healthy and take no hormone-altering medication
Evidence
Protocol, medRxiv 10.64898/2026.07.22.26358657 (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): "222 women will be randomly allocated to one of four conditions in a two-by-two treatment design"; MHT arm receives "estradiol patches (Systen, 50mcg)" "twice per week over the course of 15 weeks" plus "progesterone tablet (Utrogestan, 200mg)". Recording occurs at "baseline (T0, week 1), post-treatment (T1, week 8), and follow-up (T2, week 15)", with ZMax "at home for four nights during each timepoint" and EmbracePlus "seven days at each timepoint T0, T1 and T2". Exclusion is only that women must not "already use MHT" at entry — i.e. before randomisation.
Fix
State explicitly which timepoint(s) are used. If T0 only, say so and recompute the expected yield on a single four-night block per participant; replace 'take no hormone-altering medication' with 'are hormone-therapy-naive at the baseline recording, which is the only timepoint analysed'. If post-randomisation nights are used, add treatment arm as a design variable, pre-specify an arm-stratified analysis, and drop the 'no hormone-altering medication' claim entirely.
What the refuter said
AGAINST: "take no hormone-altering medication" (L199) is defensible as an entry criterion — the protocol's exclusions do include "current MHT use" and "hormonal contraceptives" — and the draft does flag awareness of the issue ("other medication is documented, with attention to agents suppressing vasomotor symptoms", L200-201). If the analysis uses only the T0 block, the statement is true and the confound absent, since MHT's "onset of efficacy ... is two to three weeks" (GombertLabedens2025.txt L1307) and T0 is week 1. SURVIVES. I verified the protocol: four arms, "transdermal estradiol patches (Systen, 50mcg)" twice weekly plus "progesterone tablet (Utrogestan, 200mg)" over 15 weeks, with ZMax four nights and EmbracePlus seven days at each of T0 (wk 1), T1 (wk 8), T2 (wk 15). And GombertLabedens2025.txt L1312-1317 verifies the magnitude: MHT "reducing the frequency of symptoms by 75%". So for two of four arms at two of three blocks the sentence at L199 is false, and MHT removes the outcome events rather than adding noise. The draft nowhere restricts to T0, so the reader cannot tell whether the design is clean-but-one-third-the-size or three-times-the-size-and-confounded. Downgraded from fatal because the disclosure fix is one sentence and, if T0-only, the science is intact — but as written a checkable factual claim about the sample is wrong, which is major for a grant.
The problem
The 222-woman cohort is the Sleeping Through Menopause RCT, in which women are randomised 2x2 to transdermal estradiol plus progesterone, guided CBT/circadian therapy for insomnia, both, or control, and ZMax (four nights) plus EmbracePlus (seven days) are collected at three timepoints (T0 baseline, T1 8 weeks, T2 15 weeks). Half the cohort is therefore on hormone therapy at T1/T2, which both suppresses the events and plausibly alters sigma and arousal; CBT-I alters arousal and sleep continuity directly. The draft's exclusion of hormone-altering medication is incompatible with T1/T2 data and silently implies baseline-only, which cuts the available nights to one third of what "an ongoing cohort of 222 women" suggests. Treatment allocation is left unmentioned as a determinant of both predictor and outcome.
In the draft — line 199-202
They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms.
Evidence
van Baarzel et al. 2026 protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): "222 women will be randomly allocated to one of four conditions in a two-by-two treatment design"; arms are MHT alone, CBCTi alone, both, control; "estradiol patches (Systen, 50mcg)" plus "progesterone tablet (Utrogestan, 200mg)"; "Z-max triple electrode EEG headbands (Hypnodyne)" for "four nights during each timepoint"; EmbracePlus "for seven days at each timepoint"; timepoints T0/T1/T2.
Fix
State explicitly which timepoint(s) supply the analysed nights. If baseline only, say so and recompute the yield on that basis. If T1/T2 are used, treatment arm must enter every model as a fixed effect with a pre-specified arm-by-phase interaction check, and the "no hormone-altering medication" sentence must be deleted and replaced by a description of the randomisation and its expected effect on flash rate.
What the refuter said
Strongest case against, and half of it lands. I verified the protocol at https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full in every particular the finding asserts (four arms, Systen 50mcg estradiol plus Utrogestan 200mg, web-based CBCTi, target 222 "50 per arm after 10% attrition", ZMax "four nights during each timepoint", EmbracePlus "seven days at each timepoint", T0/T1/T2). But the alleged incompatibility is conditional, not actual. "Current MHT use" and "Hormonal contraceptive use" are both trial *exclusion criteria*, so the draft's statement that participants "take no hormone-altering medication" (L199-201) is TRUE at T0. And the draft describes exactly one measurement block - "seven consecutive days and nights... and on four of those nights also with an ambulatory EEG headband" (L188-191) - which matches one timepoint exactly. The draft is therefore internally coherent as a baseline analysis, and the finding's framing ("the draft claims X which is false") is an over-reading. What survives, and is worth reporting, is the disclosure gap named in the finding's own title: the timepoint is never stated, treatment allocation is never mentioned as a determinant of predictor or outcome, and any reviewer who follows the draft's own citation to the protocol will have to ask. If baseline-only, two-thirds of the collected nights are unavailable, which bears directly on the unquantified yield. Moderate, not fatal.
The problem
The draft is written two ways at once. The Participants section says the data are "already being recorded" — a secondary analysis. The Procedure section is written in the present tense as if the applicants run the protocol ("Participants complete online surveys… then an instruction session on device handling… Sleep EEG uses the ZMax headband"), which describes the host trial's procedure, funded by NWO grant NWA.1518.22.104 with co-funding from four Dutch health charities. A reviewer reading the Procedure section cannot tell whether they are being asked to pay for recruitment, devices, instruction sessions and a measurement week that are already paid for. That reads as either double-funding or as concealing that this is an analysis grant. The consequence is concrete: with no budget in the document, the reviewer cannot resolve it, and a Scientific Board member with no way to tell what the money buys has an easy reason to decline. Note that a secondary-analysis framing is entirely fundable and in fact a strength — it makes a €60k ask credible — but only if stated.
In the draft — line 197-198 (also 207-221, present tense throughout the Procedure section)
from an ongoing cohort of 222 women already being recorded
Evidence
Host trial funding, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: "Dutch Research Council (NWO) programme NWA ORC, project NWA.1518.22.104, co-funded by Dutch Heart Foundation, Diabetes Fund, Hersenstichting, and Dutch Digestive Health Fund." The host protocol already specifies the ZMax headband for four nights and the EmbracePlus for seven days at each of three timepoints. Draft line 197-198 vs lines 207-221: the same activity is described as pre-existing and as prospective.
Fix
Open the Study design section with one unambiguous sentence, e.g.: "Recruitment, devices and the measurement weeks are funded by NWO NWA.1518.22.104 (Sleeping Through Menopause, NCT06306404). Bial funding is requested for the components this project adds: the event-button amendment and its consumables, detector re-derivation and precision validation, manual sleep and arousal scoring of [N] nights with a double-scored subsample, and the analyst time for pre-registration, analysis and dissemination." Rewrite the Procedure section in the past or passive voice for anything already collected, and clearly mark the button press and tap markers as additions. Add an overlap statement naming the NWO grant.
What the refuter said
AGAINST: Bial requires a "[d]etailed description of the research project + schedule + budget" (BIAL_CONTEXT.md L32) as separate form content, so the absence of a budget from this narrative proves nothing about what a Scientific Board member can resolve — they will have the budget. And a Procedure section written in the present tense is the normal register for describing a protocol, whoever funds it. SURVIVES as a verified internal inconsistency of framing. L197-198 says the data are "already being recorded", while L207-221 describes recruitment, instruction sessions, device fitting and a measurement week in the present tense as project activity — and I verified that every one of those activities is already specified and funded in the host protocol (NWO NWA.1518.22.104 with four Dutch health charities, ZMax four nights and EmbracePlus seven days at each timepoint). The proposal never says which of the two things it is asking Bial to pay for. The finding's own closing observation is the useful one: a secondary-analysis framing is fundable and makes a modest ask credible, but only if stated. Downgraded from major to moderate because the budget form resolves the money question even where the narrative does not.
The problem
The host protocol runs the ZMax for four consecutive weeknights and the EmbracePlus for seven days at *each of three timepoints* — T0 (week 1), T1 (week 8) and T2 (week 15) — so each participant contributes up to twelve EEG nights and twenty-one wristband days, not four and seven. The draft describes one block and never says which. This is not pedantry: the timepoint choice determines whether the sample is on MHT, whether the CBCTi arms have been treated, and how many nights are available. It also determines whether the analysis is clean, and the draft's silence means a reviewer must assume the worst. There is a further unresolved point I could not settle: the host protocol's own text is ambiguous about whether allocation precedes the T0 measurement week — it states participants "sign informed consent after which they will be allocated to one of four conditions", with T0 measurements at week 1. If allocation precedes T0, then even the baseline block is post-randomisation and unblinded, so expectancy effects act on the self-report outcome from the start, though MHT's pharmacological effect (onset 2-3 weeks) would not yet be present.
In the draft — line 62-64 (also 189-192)
Each completes four nights of ambulatory frontal EEG alongside seven days and nights of wrist-worn autonomic monitoring at home.
Evidence
Host protocol, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: "ambulatory Z-max triple electrode EEG headbands (Hypnodyne) will be used… at home for four nights during each timepoint"; EmbracePlus "for seven days at each timepoint"; "T0 (Week 1): Baseline; T1 (Week 8): Post-intervention assessment; T2 (Week 15): Follow-up assessment"; "they sign informed consent after which they will be allocated to one of four conditions". UNVERIFIED: whether the T0 measurement week is completed before or after allocation — the protocol does not state the sequence explicitly.
Fix
State it in one sentence in Study design: "The confirmatory analysis uses the T0 (pre-treatment) measurement block only: four consecutive weeknights of frontal EEG within a seven-day wristband week, per participant. The T1 and T2 blocks are recorded under randomised MHT or CBCTi and are reserved for exploratory analysis with arm as a covariate." Confirm with the host-trial PI, and state, whether the T0 measurement week precedes or follows allocation, and if it follows, name unblinded expectancy as a limitation on the self-report outcome.
What the refuter said
AGAINST: "factor of three" assumes all three blocks are usable, which the MHT problem makes false — so this finding partly contradicts F38/F158, which argue that T1 and T2 are exactly the blocks one must not use. If the plan is T0-only, then "four nights ... seven days" (L62-64, L189-192) is simply correct, and the host protocol itself describes the measurement in per-timepoint terms. SURVIVES as the same core defect stated as a specification gap: the draft never names the timepoint, and that choice determines treatment status, blinding, and event yield simultaneously. I resolved the finding's own UNVERIFIED point against the draft: the protocol's flow is screening → consent → "after which they will be allocated to one of four conditions" → assessments "at week 1 (T0), at week 8 (T1) and at week 15 (T2)", so allocation precedes even the baseline week and every recording is post-randomisation and unblinded, though MHT's pharmacological effect (onset two to three weeks) is absent at T0. Moderate: real, verified, and fixed by the same sentence that fixes F38/F158, so it should be reported with them rather than separately.
The problem
Two issues. The placeholder is a submission defect: Bial requires documentary evidence of submission to the competent ethics committee plus a signed PI declaration, and no Agreement is signed without the approval document. More substantively, the sentence claims approval for "the study" when the approval that exists — METC AUMC NL87156.018.24 — is for the Sleeping Through Menopause randomised trial. It does not cover the nocturnal event button, the tap-marker procedure, or the ancillary use of the data for this question, none of which appear in the host protocol. Asserting blanket approval, then leaving the number blank, is the combination most likely to read as carelessness about human-subjects governance.
In the draft — line 331-333
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference \[number\]).
Evidence
DRAFT.md line 332 contains the literal string "reference \[number\]". Actual host approval: "Medical Research Ethics Committee of Amsterdam University Medical Center (METC AUMC)… Reference Number: NL87156.018.24" (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full). The host protocol contains no event button, no tap marker, and no ancillary-study provision. Bial requirement, /root/grantreview/BIAL_CONTEXT.md lines 35-36: "documentary evidence of submission to the competent Ethics committee, plus a signed PI declaration. Ethics approval must exist before the Agreement is signed — 'No Agreement will be issued and signed unless such document has been duly provided' (Art. 19(4))."
Fix
Replace with: "The host trial is approved by the Medical Research Ethics Committee of Amsterdam UMC (NL87156.018.24; NCT06306404). The event button, the tap synchronisation marker and the ancillary use of the recordings for this question are covered by amendment [n], submitted [date] / approved [date]; documentary evidence and the signed PI declaration are appended." If the amendment is not yet submitted, say when it will be, and note that the Bial start window (1 Jan – 31 Oct 2027) allows time for approval before the Agreement is signed.
What the refuter said
Strongest case against: the placeholder is a drafting artefact, not a submission defect — L332 is a working document, the number obviously gets filled in before the form is submitted, and no reviewer sees this draft. Reporting a blank bracket as a major finding inflates a typo. It is also possible that an approved amendment covering the button press already exists; I cannot verify either way, so the strongest form of the charge is unproven. What survives is the substantive half and it is verified on both sides. The string "reference \[number\]" is literally present at L332. The host approval is METC AUMC NL87156.018.24 for the Sleeping Through Menopause trial (confirmed from the protocol), and that protocol measures nocturnal flashes only algorithmically — "estimated from a multisensory approach ... with the EmbracePlus smartwatch" — with no button press and no event marker anywhere in it. So the draft's flat assertion that "The study is approved" covers measurements that are demonstrably not in the approved protocol, against a funder that will not sign without the document: "No Agreement will be issued and signed unless such document has been duly provided" (Art. 19(4)). Moderate.
The problem
The cohort is the Sleeping Through Menopause trial, a 2x2 factorial RCT in which participants are randomised to transdermal estradiol 50 mcg plus oral progesterone 200 mg, to web-based cognitive behavioural and circadian therapy for insomnia, to both, or to control, with n=222 and four ZMax nights plus seven EmbracePlus days per timepoint. Three consequences. (1) "take no hormone-altering medication" is false for up to three quarters of the sample at post-randomisation timepoints; estradiol and progesterone alter sigma/spindle activity and directly suppress vasomotor symptoms, so they perturb the predictor, the event rate and the outcome simultaneously. (2) If only pre-randomisation timepoints are usable, the draft must say so, and the event yield collapses to one four-night block per woman rather than the multi-timepoint yield the "seven days and nights" plus "four nights" description implies. (3) The draft nowhere discloses to the funder that its data come from a therapy trial, which interacts with the Bial exclusion (see bial-framing-clinical). The draft also says the 222 are "already being recorded" (L198) whereas 222 is the trial's recruitment target including 10% dropout; actual enrolment to date is UNVERIFIED.
In the draft — line 60-62, 200-201
has access to an ongoing cohort of 222 perimenopausal women with self-reported sleep disturbances ... They are physically healthy and take no hormone-altering medication
Evidence
van Baarzel et al. protocol, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full : "repeated measurement, randomised controlled trial (RCT) in perimenopausal women who experience insomnia"; interventions "transdermal estradiol patches (50mcg twice weekly) plus oral progesterone (200mg)" and "web-based guided cognitive behavioural and circadian therapy"; sample size 222 "accounting for 10% dropout; 50 per arm completing"; inclusion "Insomnia severity index (ISI) of 10 or higher" and "green climacteric scale (GCS) score of 13 or higher"; monitoring "ambulatory Z-max triple electrode EEG headbands" and "smartwatch (Embraceplus)". The draft cites this protocol itself under PREVIOUS OWN PUBLICATIONS (L449-454) without stating that it is the data source or that it is an intervention trial.
Fix
State explicitly which trial timepoints are used. Restrict the primary analysis to pre-randomisation baseline recordings and recompute the yield projection on that basis (the L204-206 sequential-exclusion projection must be redone). If post-randomisation nights are used, add treatment arm and time-on-treatment as pre-specified covariates, report the flash-rate and sigma-power differences by arm, and drop the "take no hormone-altering medication" sentence. Replace "an ongoing cohort of 222 perimenopausal women with self-reported sleep disturbances" with an accurate description: the number actually recorded to date, the parent trial, and the design.
What the refuter said
Strongest case against, and two of three prongs fail. I verified every protocol detail (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): four arms, Systen 50mcg plus Utrogestan 200mg, web-based CBCTi, ISI >=10 and Greene >=13, ZMax four nights and EmbracePlus seven days at each of T0/T1/T2, target 222 at 50 per arm after 10% attrition. But prong (1) is conditional, not actual: "Current MHT use" and "Hormonal contraceptive use" are trial exclusion criteria, so "take no hormone-altering medication" is true at T0, and the draft's single seven-day/four-night block (L188-191) matches one timepoint exactly - so "false for up to three quarters of the sample" describes data the draft never claims to use. Prong (3) also largely fails: the protocol IS cited, at L449-454, under PREVIOUS OWN PUBLICATIONS, with a title that names "a randomized controlled trial of menopausal hormone therapy and online guided cognitive behavioural and circadian therapy for insomnia in perimenopausal women suffering from insomnia". An assessor who reads the reference list is not being kept in the dark. Prong (2) survives and is the finding: the timepoint is never stated, allocation is never mentioned as a determinant of predictor or outcome, and "222 women already being recorded" (L197-198) overstates a recruitment target that includes 10% attrition, with actual enrolment unstated. Moderate, not fatal.
The problem
The draft correctly identifies precision as the binding constraint and then relies on a detector whose published operating characteristics are incompatible with it. Tsiartas et al. 2021 classifies in 15 s frames; its features are computed over windows of plus/minus 30 s, 120 s before and after, 250 s and 500 s; and performance is evaluated against expert annotations using a plus/minus 90 s matching window - i.e. an alignment tolerance of nearly four ISF cycles. The reported "over 90% sensitivity at 95.6% specificity" is therefore an event-presence figure at coarse tolerance, not evidence that onset can be timed to a fraction of 50 s, and it comes from three women and 27 flashes. The draft's back-dating step is a good idea but it is not part of the cited validation, so its achievable precision is UNVERIFIED and the study is proposing to establish it and to depend on it in the same project. That is acceptable only if the fallback is sound, and it is not (see amplitude-fallback).
In the draft — line 260-262
Because the ISF cycle is only about 50 s, onset uncertainty translates directly into phase uncertainty. Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing.
Evidence
/root/grantreview/Tsiartas2021.txt: "Three women (Age, mean +/- SD: 55.6 +/- 0.6 y)"; "A total of 27 physiological HFs were recorded"; "a Decision Tree classifier, which makes a decision every 15 s"; "We time-aligned the features with the HF expert annotations for prediction and evaluation (+/-90 s matching window)"; T feature uses "the prior and following 500 s"; PPG feature "averaged in two regions: 120 s before and after the window". Draft L242-247 cites this as "reaching over 90% sensitivity at 95.6% specificity". Additionally, the Tsiartas wrist skin-conductance channel was a wired lab sensor integrated into a PSG amplifier, not the EmbracePlus wrist EDA the draft will use (draft L217-218), so even the coarse figures do not transfer without revalidation.
Fix
Do not present the Tsiartas figures as support for onset timing. Add one sentence stating what they do and do not establish: "Tsiartas et al. (2021) establishes event-level detection at a plus/minus 90 s tolerance in three participants; onset timing at the precision this study requires is not established and is a primary feasibility outcome." Then make the precision result a formal gate: pre-specify a validation sub-study - a small in-lab night with simultaneous sternal skin conductance as the reference - that quantifies the distribution of back-dated onset error against the gold standard, and pre-commit that the phase analysis proceeds only if the standard deviation of that error is below a stated fraction of the cycle (the draft's own simulation implies a quarter-cycle, about 12 s, as the ceiling). Report the error distribution regardless.
What the refuter said
AGAINST: the draft does not use Tsiartas for onset timing and says so — detection is one step, timing is another: "Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing" (L259-261), with achievable precision named as a feasibility outcome (L275-276) and a pre-registered fallback if it fails. The finding concedes this ("The draft's back-dating step is a good idea"), which removes most of its force: coarse detection followed by fine re-timing on the raw EDA is a coherent architecture, and 15-s frames and a ±90 s matching window constrain the detector, not the change-point estimator. SURVIVES on the citation rather than the timing. Every quoted fact checks out in Tsiartas2021.txt — "Three women (Age, mean ± SD: 55.6 ± 0.6 y)", "A total of 27 physiological HFs", "a decision every 15 s", "±90 s matching window", the 500 s temperature window and 120 s PPG windows — and the wrist SC channel was a wired Grove sensor integrated into a Compumedics PSG amplifier, not the EmbracePlus of L214-215. The draft presents "over 90% sensitivity at 95.6% specificity" (L244) as a settled property with no caveat about n=3, lab conditions, or the different device. Moderate, down from major: the timing critique is pre-empted; the uncaveated transfer of three-participant lab figures to home wrist EDA is not.
The problem
The draft's entire measure of conscious registration is a new participant-facing procedure: a time-stamped wristband button press at perceived flash onset, five taps on the headband at lights-off and on waking, and a morning diary estimate of flash number. None of these appear in the approved parent protocol, which specifies objective detection from the wristband only. Claiming existing approval for a protocol that adds nocturnal self-report tasks is not sustainable, and Bial will not sign an Agreement without the ethics document (Art. 19(4), BIAL_CONTEXT.md line 36). The reference number is also left as a placeholder, so a reviewer cannot check what was actually approved.
In the draft — line 331-332
The study is approved by the Medical Ethics Review Committee of Amsterdam UMC (reference \[number\]).
Evidence
Same protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): hot flash measurement is "objective occurrence of nocturnal hot flashes will be estimated from a multisensory approach (skin conductance, skin temperature, heart rate and movement) assessed with the EmbracePlus smartwatch"; there is no mention of a hot-flash button press, no mention of tapping the headband for synchronisation, and no morning hot-flash diary. The protocol's approval is "Medical Research Ethics Committee of the Amsterdam University Medical Center (METC AUMC)" reference NL87156.018.24, trial registration NCT06306404.
Fix
State that an amendment to METC AUMC NL87156.018.24 covering the button-press, tap-marker and morning-diary procedures has been submitted (or is required), give the actual reference number, and say plainly that these procedures are additions to the parent trial rather than implying they are already approved.
What the refuter said
AGAINST: absence from a published protocol PAPER is not absence from the approved METC dossier. A protocol summary describing a 2x2 RCT would not be expected to enumerate an operational detail like "tap the headband five times at lights-off", and EmbracePlus ships with a tag/event button as standard, so an event-marking instruction could have been in the dossier from the outset. The chain "not in the medRxiv paper -> not approved -> the ethics claim is unsustainable" therefore has a real gap, and "fatal" is not supportable on this evidence. WHAT I DID VERIFY: the protocol paper enumerates hot-flash measurement explicitly and physiologically only — "The objective occurrence of nocturnal hot flashes will be estimated from a multisensory approach (skin conductance, skin temperature, heart rate and movement) assessed with the EmbracePlus smartwatch" — with no button press, event marker, tap marker or morning hot-flash diary anywhere; approval is METC AUMC NL87156.018.24, trial NCT06306404. Omitting a prospective self-report measure from a list of hot-flash measures is at least suggestive. And the confirmed core stands on the draft alone: L331-332 asserts approval while leaving "reference [number]" as a placeholder, so a reviewer cannot check what was approved, and BIAL_CONTEXT.md line 36 (verified in the regulation as Art. 19(4)) makes the ethics document a precondition of signature. PLAUSIBLE at moderate: fix the number and state whether an amendment is needed.
The problem
The draft sets its own dealbreaker at a quarter of a 50 s cycle, i.e. about 12.5 s of onset jitter. The method it cites for detection operates two orders of magnitude away from that: predictions are made in 15 s steps, the onset feature is defined as a ±2 minute region around the predicted onset, and events were matched to expert annotations within a ±90 s window. The gold-standard sternal criterion is itself a rise of at least 2 µS within a 30 s period. The draft's answer is to back-date onset by change-point detection on the electrodermal inflection, which is a reasonable idea, but it is an untested idea: no source in the proposal reports the achievable precision of wrist electrodermal onset localisation against any reference. The proposal therefore rests on an unevidenced assumption at precisely the point it identifies as binding.
In the draft — line 315-317
onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle
Evidence
/root/grantreview/Tsiartas2021.txt lines 84-86: "We time-aligned the features with the HF expert annotations for prediction and evaluation (±90 s matching window)" and the classifier "makes a decision every 15 s". Lines 118-119: "We designate the HF onset as the ±2 minutes around the HF predicted onset." Line 46: gold standard is "expert evaluation of sudden increases (2 uS/30s) in sternal skin conductance".
Fix
Make onset precision a hard, pre-data gate rather than a described intention. Add: 'Onset precision is unestablished for wrist electrodermal recording. In a calibration subsample recorded with concurrent sternal skin conductance we estimate the distribution of change-point-to-reference latency; the phase analysis proceeds only if the interquartile range is below 12.5 s (a quarter cycle). Note that the detection method we build on tolerates ±90 s and cannot itself supply onset timing at this resolution.' If the calibration subsample is not affordable, the phase hypothesis is not testable and the amplitude hypothesis should be primary from the outset.
What the refuter said
All source facts verified in Tsiartas2021.txt: '(±90 s matching window)' L82-84; 'makes a decision every 15 s' L89; 'We designate the HF onset as the ±2 minutes around the HF predicted onset' L120-121; 'expert evaluation of sudden increases (2 uS/30s) in sternal skin conductance' L46. Draft's own budget verified at L315-318. STRONGEST DEFENCE, and it forces a downgrade: the draft does not use Tsiartas's onset. L259-261 back-dates onset by change-point detection to the electrodermal inflection, so the ±90 s figure is Tsiartas's event-matching tolerance for classification scoring, not an error the draft inherits. The finding's headline arithmetic is also wrong — 90/12.5 is 7.2×, not 'forty times' and not 'two orders of magnitude' — which is exactly the kind of inflation that gets a reviewer's objection dismissed. What survives, and is genuinely decision-relevant: the draft identifies ~12.5 s as its own dealbreaker, proposes an untested method to meet it, and cites no source anywhere reporting achievable wrist-electrodermal onset localisation against any reference. The draft concedes this itself at L275-276 ('achievable phase precision are themselves feasibility outcomes'). Downgraded major to moderate; the core overlaps F161 and F137.
The problem
There is no budget, and no Euro figure anywhere in the document, against a hard €60,000 ceiling. Three constraints interact badly and none is acknowledged. (1) Art. 11(3) forbids overheads/indirect costs and any charge for use of Host Entity equipment, with an exception list covering only MRI, TMS, CT, MEG, PET, SPECT and fNIRS — ZMax headbands and EmbracePlus wristbands are not on it, so device-use costs are not recoverable. (2) The described work is labour-heavy: manual expert scoring of ~137 participants x 4 nights (~55,000 30-s epochs before any double-scoring), plus change-point detection development, device synchronisation, simulation-based power work and pre-registration. (3) The data collection itself is already funded by NWO project NWA.1518.22.104. So the application must simultaneously show that €60k buys something real and that it is not paying twice for collection. The draft addresses none of this, which leaves the assessor to guess — and guessing is not how you win a 19% competition.
In the draft — line 233-235
Recordings are scored in 30-s epochs (wake, REM, N1-N3) by an experienced scorer, with a double-scored subsample giving inter-rater agreement.
Evidence
€60,000 cap: "The maximum amount will be €60.000 per project" (https://www.fundacaobial.com/en-GB/news/bial-foundation-opens-new-call-for-grants-for-scientific-research-2026-2027); historical grants ran €5,000–€50,000 (https://www.25anosfundacaobial.com/en/grants-for-scientific-research/grants). Art. 11(3), verbatim from the Regulation PDF: "overheads/indirect costs or payments for the use of spaces or equipment belonging to the Host Entity or to the Research Centre where the Research Project is carried out are, under no circumstances, accepted or paid by Bial Foundation, except for costs associated with the use of neuroimaging equipment for magnetic resonance imaging (MRI), transcranial magnetic stimulation (TMS), computed tomography (CT), magnetoencephalography (MEG), positron emission tomography (PET), single-photon emission computed tomography (SPECT), and functional near-infrared spectroscopy (fNIRS)." Parent funding: "Dutch Research Council (NWO), Dutch Heart Foundation, Diabetes Fund, Hersenstichting, Dutch Digestive Health Fund; project NWA.1518.22.104" (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full). Grep of DRAFT.md: budget=0, Euro=0, €=0.
Fix
Build a Euro budget that is explicit about the division of labour with the parent grant: state that NWO funds recruitment and data acquisition, and that the Bial grant funds only the work specific to this question — e.g. named analyst/postdoc FTE, expert scoring time, the button-press hardware/firmware and amendment work, and pre-registration/open-data costs. Charge no overheads and no ZMax/EmbracePlus use fees. Add one sentence in the narrative making the complementarity explicit, so the assessor reads 'leveraged infrastructure' rather than 'double funding'.
What the refuter said
Strongest case against: this document is narrative content for the free-text boxes of an online form (its own headings mirror form fields, e.g. "LITERATURE REVIEW (MAX. 6000 CHARACTERS INCLUDING SPACES)"); the budget is a separate structured field in BF-GMS, so its absence from a prose draft is expected, not a defect. That kills the headline "No budget at all". The finding's labour arithmetic is also wrong by an order of magnitude: 137 women x 4 nights x ~8 h is ~526,000 30-s epochs, not "~55,000". What survives is substantive and verified. Cap: "The maximum amount will be €60.000 per project" (fundacaobial.com 2026/2027 call). Art. 11(3) verbatim from the Regulation PDF: "overheads/indirect costs or payments for the use of spaces or equipment belonging to the Host Entity ... are, under no circumstances, accepted or paid", exception list MRI/TMS/CT/MEG/PET/SPECT/fNIRS only — ZMax and EmbracePlus are host-owned equipment already in use, so their use cannot be charged. Parent funding verified from the Sleeping Through Menopause protocol: NWO "NWA. 1518.22.104", plus Dutch Heart Foundation, Diabetes Fund, Hersenstichting, Dutch Digestive Health Fund, with hot flashes measured by "the EmbracePlus smartwatch for seven days at each timepoint". Grep of DRAFT.md returns zero hits for budget, Euro, €, cost, NWO or funding. So the draft nowhere discloses that collection is already funded, nor what €60k buys against a labour-heavy plan whose device costs are non-recoverable. Real gap, moderate not major.
The problem
Art. 6 fixes a maximum duration of three years and requires the grant to commence between 1 January and 31 October 2027. The draft states no start date, no duration and no schedule of any kind — no work packages, no milestones, no data-freeze point. At the same time it repeatedly emphasises that recording is already under way. The assessor is therefore left to infer either that the project is a retrospective analysis of data collected before the funding period (which invites 'what is Bial paying for?'), or that accrual will still be running in 2027–2029 (in which case a schedule is essential and the yield projection depends on a recruitment rate the draft never gives). Both readings need a schedule to resolve; neither is addressed.
In the draft — line 197-198
from an ongoing cohort of 222 women already being recorded
Evidence
Regulation Art. 6(1-2), verbatim: "The Grants foreseen in this Regulation have a maximum duration of three (3) years, with no minimum duration defined." / "The Grants shall commence between 1 January and 31 October 2027." Grep of DRAFT.md: '2027' = 0 occurrences, no 'timeline', no project 'Schedule'. Parent trial registered May 2024, protocol v8.0 dated September 2025, 15-week per-participant follow-up (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full) — so accrual has been running for roughly two years already.
Fix
Add a Schedule section with a start date inside the window (e.g. 1 March 2027), a duration, and dated milestones: amendment in force, accrual completion or data freeze, pre-registration lodged, synchronisation and onset-precision validation, primary analysis, reporting. State the current accrual count and rate as of submission, and show that the remaining accrual plus analysis fits the funded period. If most collection will be complete before the grant starts, say so and frame the grant as funding the analysis programme — but frame it deliberately rather than leaving the assessor to notice.
What the refuter said
Strongest case against: as with F1, the schedule may be a separate BF-GMS field, so its absence from the narrative is not necessarily an eligibility failure; and Art. 6's window is a contract term the applicant satisfies at signature, not something a proposal must recite. Also, "already being recorded" is not in tension with a 2027 start per se - analysing an existing cohort is a normal funded activity. But the finding survives on a different and sharper axis than eligibility. I verified the grep (2027=0, timeline=0, milestone=0, "work package"=0) and the parent-trial facts at https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full (registered NCT06306404 on 13 May 2024; 15-week per-participant protocol T0/T1/T2; no recruitment start or end date stated). The draft therefore never tells the assessor what the grant buys or over what period: "an ongoing cohort of 222 women already being recorded" (L197-198) plus zero schedule leaves the assessor unable to distinguish funded data collection from a secondary analysis of data already in hand. Under a regulation whose Art. 6 fixes a 2027 commencement and whose assessment turns on the stated objectives, that is a real, decision-relevant gap. Downgraded from major because the omission is one of positioning rather than of a substantive component, and because it is cheaply fixed with two sentences.
The problem
The proposal's whole feasibility rests on onset-phase precision and never assembles the error terms. At least five contribute: (i) EEG-wristband synchronisation residual, from cross-correlation with "a fixed offset and linear drift estimated per recording" anchored by taps at lights-off and waking only (l.219-221, l.228-230) — two anchors per night, with the linear-drift assumption unverified mid-night, which is where the events are; (ii) change-point error in locating the electrodermal inflection (l.260-262); (iii) the lag between the true central initiation of the flash and its first peripheral expression, which no measurement in this design observes at all; (iv) sub-epoch uncertainty from 30 s staging (l.234) against a 25 s half-cycle; (v) error in the ISF phase estimate itself. No target for any of these is given and no total is computed. Second, the tolerance is misstated. A quarter of a 50 s cycle is 12.5 s, but the filter passband is 0.01-0.04 Hz with the "empirical peak reported per participant" (l.267-268): a participant whose peak sits at 0.04 Hz has a 25 s cycle and a quarter-cycle tolerance of 6.25 s. The proposal's stated precision requirement is therefore up to twice as lenient as its own method permits, and it never says what happens to participants at the fast end of the passband. The same ambiguity infects the bout criterion: "three ISF cycles" (l.345-346) is 75 s at 0.04 Hz and 300 s at 0.01 Hz — a fourfold difference in an inclusion rule that is never resolved. Third, the decision statistic does not measure the decision quantity: "the dispersion of intervals between peripheral markers" (l.338-339, l.262-263) measures inter-marker latency variability across EDA, temperature and pulse rate. That is a proxy for, not a measurement of, the error of the estimated onset against the true onset — the quantity the switching rule depends on.
In the draft — line 316-319
onset jitter translates directly into phase uncertainty against a 50-s cycle, and simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle
Evidence
Draft l.316-319 ("a quarter of the cycle" against "a 50-s cycle") versus l.267-268 ("zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant)"): 0.04 Hz gives a 25 s cycle, so quarter-cycle tolerance is 6.25 s, not 12.5 s. Draft l.345-346 ("three ISF cycles") is 75-300 s across the same passband. Draft l.219-221 and l.228-230 (two tap anchors per night, linear drift) versus l.262-263 ("Intervals between this inflection and the accompanying temperature and pulse-rate changes index achievable precision") — an index of marker agreement, not of onset error.
Fix
Add an error-budget paragraph with a number per term and a total, and express tolerances as fractions of each participant's own cycle: "The onset-phase error budget comprises device synchronisation residual (target < [X] s, verified against tap markers at both ends of each night and reported per recording), change-point localisation error of the electrodermal inflection (< [X] s, estimated against expert-marked onsets in a subsample), unobserved central-to-peripheral latency (bounded by the inter-marker interval distribution and treated as a limitation), and ISF phase-estimation error (< [X] rad, from simulation). Precision requirements are expressed per participant as a fraction of that participant's empirical ISF period; participants whose empirical peak exceeds [value] Hz face a proportionally stricter requirement and are handled by [rule]. Bout-length criteria are stated in cycles of the participant's empirical peak and in seconds." Add mid-night synchronisation anchors (e.g. a scheduled tap, or drift estimated on rolling windows) rather than relying on two per night.
What the refuter said
AGAINST: much of this asks the draft to specify what it has honestly declared unknown. L275-276 makes "achievable phase precision" a feasibility outcome, and the Risks section pre-registers a threshold and a switch rule. Term (iii) — the lag between central initiation and peripheral expression — is unobservable in any wearable design, so no proposal could budget it. Term (ii)'s complaint about using inter-marker dispersion is criticising the only estimator available: with no ground-truth onset, marker agreement is the only observable proxy, so "the statistic is not the quantity" is a limitation of the world, not a design error. What survives is real. The tolerance is stated against a 50-s cycle (L316-318) while the passband is 0.01-0.04 Hz (L268) — at 0.04 Hz the quarter-cycle tolerance is 6.25 s, and the draft never says what happens to fast-peak participants. The "three ISF cycles" bout rule (L344-345) is likewise 75-300 s across the same band, a fourfold ambiguity in an inclusion criterion. Both are verifiable internal inconsistencies in a proposal that names precision as its binding constraint. Moderate, down from major, and PLAUSIBLE rather than CONFIRMED because the underlying complaint (no assembled error budget) is partly answered by the draft's declaration that these are feasibility unknowns.
The problem
The Summary and Participants sections both present 222 as an existing recorded cohort ('an ongoing cohort of 222 perimenopausal women', 'already being recorded'), while five lines later the draft reveals that 65 have been screened. 222 is the parent trial's planned enrolment target (50 per arm x 4 arms plus 10% dropout), not a number of women with data in hand. The 137-woman analysis sample is therefore a projection resting on a projection: a 62% eligibility rate estimated from 65 screens, applied to an accrual target that is roughly 30% achieved. The proposal's central feasibility claim — 'our team has access to a cohort of 222' — is the strongest thing in the Summary and it is the least supported. A hostile reviewer who opens the cited protocol finds the 222 defined as a power target and will discount the whole feasibility argument.
In the draft — line 202-205
The analysis sample comprises those whose nocturnal flashes interfere with sleep: of the first 65 screened, 40 (62%, 95% CI 49-72) meet this criterion, projecting to roughly 137 women.
Evidence
Parent protocol (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full): "Target: 222 women (n=50 per arm in 2x2 design, accounting for 10% dropout)" — i.e. a sample-size calculation, not an enrolled cohort. Internal contradiction within DRAFT.md between L197-198 ("222 women already being recorded") and L203 ("of the first 65 screened").
Fix
Replace the claim with the actual state of accrual as of submission: 'Participants are drawn from an ongoing trial with a planned enrolment of 222 (NCT06306404); N enrolled and M with complete baseline recordings as of August 2026, accruing at approximately R per month.' Then present the 137 figure explicitly as a projection conditional on full accrual, and give the analysis sample that would be available under a realistic accrual scenario as well. Also report the yield chain numerically rather than only naming the exclusions (L205-207 lists the exclusions but gives no attrition figures).
What the refuter said
AGAINST: the draft is not hiding accrual — the very next sentence says "of the first 65 screened" (L203), which explicitly signals an in-progress cohort, so a reviewer is told the true state within the same paragraph. "Ongoing cohort of 222" can be read as "the ongoing study whose planned N is 222", and 137 is transparently labelled a projection. SURVIVES on the facts. The protocol states "Expecting a drop-out of 10%, the total number of participants to be recruited is n=222" and reports no enrolment figures at all — so 222 is a recruitment target, and "an ongoing cohort of 222 perimenopausal women" (L60-62) / "222 women already being recorded" (L197-198) overstates what exists. The Summary's version has no "65 screened" caveat next to it, so the strongest feasibility sentence in the proposal is the least supported. Downgraded from major to moderate: the arithmetic itself is sound (40/65=62%, Wilson CI 49-72, 62% of 222≈137), the projection is labelled as one, and the disclosure of 65 screens sits five lines away — this is an overstatement of maturity, not a concealed one.
The problem
The arithmetic here is correct (see notes), but three things around it are not resolved. (1) The cohort is described as "an ongoing cohort of 222 women already being recorded" (l.197-199) with no citation; the only plausible referent in the document is the applicants' own protocol, a "randomized controlled trial of menopausal hormone therapy and online guided cognitive behavioural and circadian therapy for insomnia in perimenopausal women" (l.449-454). If the analysis sample is drawn from that trial, the statement that participants "take no hormone-altering medication" is in direct conflict with the trial's intervention, and vasomotor symptom frequency — the exposure that generates every event in this study — would be actively suppressed in an arm of it. Either the cohort is a different one, or the eligibility statement needs to specify how randomised arm and treatment timing are handled. (2) The projection assumes all 222 will be screened and will be eligible at the observed rate; the interval is not propagated: 49.4-72.4% of 222 is 110-161 women, and the draft reports only "roughly 137". (3) If recording is already under way, some participants have already completed their nights, and the proposal never states whether the button-press measure and the EEG-plus-wristband night structure are part of the existing protocol or are additions requiring re-recording. This determines whether 137 is available or aspirational, and it also bears on whether the existing ethics approval covers the new measure.
In the draft — line 199-204
They are physically healthy and take no hormone-altering medication; other medication is documented, with attention to agents suppressing vasomotor symptoms. The analysis sample comprises those whose nocturnal flashes interfere with sleep: of the first 65 screened, 40 (62%, 95% CI 49-72) meet this criterion, projecting to roughly 137 women.
Evidence
Draft l.199-200 ("take no hormone-altering medication") and l.197-199 ("an ongoing cohort of 222 women already being recorded", uncited) against l.449-454 (own publication: "study protocol of a randomized controlled trial of menopausal hormone therapy and online guided cognitive behavioural and circadian therapy for insomnia in perimenopausal women suffering from insomnia"). Arithmetic: 40/65 = 61.5%; Wilson 95% CI 49.4-72.4%; 61.5% of 222 = 136.6; the interval maps to 110-161 women.
Fix
Name the cohort and reconcile the criteria: "Participants are drawn from the [named] cohort (van Baarzel et al., 2026), in which [state whether and how hormone therapy is administered]. Women receiving menopausal hormone therapy are [excluded / analysed separately / included with treatment arm and time since initiation as covariates], since these agents suppress vasomotor symptoms. Of the first 65 screened, 40 (62%, 95% Wilson CI 49-72) report nocturnal flashes that interfere with sleep, projecting to 137 women (interval 110-161). Of these, [N] have already completed recording under the present protocol, which includes the nocturnal button press; the remaining [M] are recorded prospectively." State explicitly whether the button press is an addition to the parent protocol and whether an ethics amendment is required.
What the refuter said
STRONGEST DEFENCE: sub-point (2) is a non-issue — "projecting to roughly 137 women" is explicitly hedged, and I independently confirm 40/65 = 61.5%, Wilson 95% CI 49.4-72.4%, 61.5% of 222 = 136.6, so the stated arithmetic is correct and honest. Sub-point (1)'s "direct conflict" also dissolves under the obvious reading: the parent trial EXCLUDES current MHT users and hormonal contraceptive users at entry (I fetched the protocol and confirmed both exclusions), so at baseline "take no hormone-altering medication" is TRUE; randomisation to MHT occurs after baseline, and four nights is exactly one timepoint's ZMax block. WHY IT SURVIVES ON SUB-POINT (3), WHICH IS THE REAL FINDING: the parent protocol contains neither the button press nor the tap-marker synchronisation (the fetch returned "No button-press marker mentioned" and no tap procedure), while the draft describes the cohort as "already being recorded" (L197-199). So the PRIMARY outcome measure is an addition to a study already in the field, and the draft never says whether nights recorded so far carry it. If it was added late, the analysable N for the confirmatory model is unknown and potentially far below 137 — and the draft nowhere states which timepoint's nights it uses (grep: no occurrence of baseline, timepoint, T0, randomis*, or the trial's name in the body text). DOWNGRADE: kept at moderate, with the note that the headline should be sub-point (3), not the 137 projection.
The problem
The contingency is well-motivated but numerically trivial relative to the gap it is meant to close, and the draft does not say what gain it expects. Relaxing the bout requirement changes only the bout-geometry factor.
In the draft — line 345-347
Low event yield is met by relaxing the stable-bout criterion from three ISF cycles to two, pre-specified as secondary.
Evidence
With exponentially distributed stable-NREM bout durations of mean mu, the fraction of stable NREM lying at least a seconds inside a bout is exp(-a/mu) one-sided (/root/grantreview/sim/08_yield_v2.py). Relaxing from 3 cycles (150 s) to 2 (100 s): mu=150 s, 0.368 -> 0.513 (+40% events); mu=240 s, 0.535 -> 0.659 (+23%); mu=420 s, 0.700 -> 0.788 (+13%). The shortfall to 80% power for OR=2 at 5 s jitter is 9-fold. Note also that mu is itself an assumption of this review, not a measured quantity: the draft gives no arousal index or NREM bout statistics for this cohort, and these are women selected for sleep disturbance, so short bouts are the expected case.
Fix
State the expected gain and add a contingency proportionate to the gap: "Relaxing the stable-bout criterion from three ISF cycles to two increases the analysable event count by 13-40% depending on the observed NREM bout-length distribution, which is reported from the pilot nights. Because this is small relative to the event count the confirmatory test requires, the substantive low-yield contingency is an increase in EEG nights per participant rather than a relaxation of the bout criterion."
What the refuter said
STRONGEST DEFENCE: almost everything numerical here is the reviewer's construction, not the draft's. The exponential bout-duration model, the mean bout length μ, and the "nine-fold shortfall" all come from an assumed yield cascade that the draft does not contain — the finding itself concedes "mu is itself an assumption of this review, not a measured quantity". And no grant proposal quantifies the expected gain from relaxing a secondary inclusion criterion; "pre-specified as secondary" is exactly the right way to handle a contingency, and pre-specifying it at all is better practice than most proposals manage. So the criticism reduces to "the draft does not state a number it had no obligation to state, and our own model says that number would be small". WHY A KERNEL SURVIVES: the geometric arithmetic is internally correct (one-sided exp(−a/μ): 0.368→0.513 at μ=150 s, 0.535→0.659 at μ=240 s, 0.700→0.788 at μ=420 s), and the observation that the draft's only yield contingency is small relative to any plausible shortfall is a fair note IF a shortfall exists — which depends on F28's finding that no yield or power figure is given at all. The point about short bouts in women selected for sleep disturbance is the most useful line here. Severity stays minor, correctly assigned.
The problem
On the draft's own projection of ~137 women × 4 nights, this is roughly 550 nights of manual 30-s epoch scoring plus arousal marking, plus a double-scored subsample — on the order of 550-1100 expert hours, before any analysis. If more than one measurement block is used, treble it. The host protocol says nothing about who stages the ZMax data or whether staging is manual at all, so this cannot be assumed to be already resourced. Against a €60,000 ceiling with no overheads recoverable, expert manual scoring at this volume plausibly consumes the entire grant, leaving nothing for detector validation, analysis, or dissemination. The draft's risk section does not mention it. This is the most likely place where the plan simply does not fit the money.
In the draft — line 233-235
Recordings are scored in 30-s epochs (wake, REM, N1-N3) by an experienced scorer, with a double-scored subsample giving inter-rater agreement.
Evidence
Draft lines 233-235 (manual scoring of all recordings, plus double-scored subsample) and lines 201-205 (~137 women), 190-192 ("on four of those nights also with an ambulatory EEG headband"). Bial cap and cost rules, https://www.fundacaobial.com/en-GB/grants/grants-programme-scientific-research: "Maximum funding per grant: €60,000"; /root/grantreview/BIAL_CONTEXT.md line 27: "Overheads and space/equipment-use costs NOT reimbursable". Host protocol is silent on staging: "No detail on staging procedures… It only notes 'sleep stages and sleep architecture' will be assessed" (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full).
Fix
Make scoring a costed, bounded line item. Either (a) state that staging is already performed by the host trial and cite that arrangement, so Bial funds only arousal marking on the subset of nights carrying analysable flashes; or (b) use validated automatic staging with a manually scored subsample for agreement, and budget only that subsample — e.g. "automatic staging (kappa 0.77 against manual in the Wearanize+ reference dataset) with 10% double-scored manually by an experienced scorer, agreement reported". Either way, state the number of nights scored, the hourly rate, and the euro total, and pre-specify that manual arousal marking is restricted to windows around detected flashes rather than whole nights.
What the refuter said
STRONGEST DEFENCE, AND IT UNDERMINES THE HEADLINE: "with no budget" is an inference about a document that has no budget section because the budget is a separate mandatory form field (the call requires a suggested payment plan and itemised cost categories submitted through the portal). The same is true of the schedule and CVs. So the absence proves nothing about whether scoring is costed. The cost arithmetic is also inflated: 550 nights at a realistic 30-60 min per night for staging plus arousal marking is roughly 275-550 expert hours, not 550-1100, which at Dutch technician cost is on the order of €14k-28k of a €60,000 cap (cap verified on the Bial call page: "€60,000 (sixty thousand euros)") — substantial, not "plausibly consumes the entire grant". And the parent trial itself assesses "sleep stages and sleep architecture", so staging is plausibly already resourced there; the finding's "this cannot be assumed to be already resourced" cuts equally against its own assumption. WHY A KERNEL SURVIVES: the commitment at L233-235 to manual 30-s epoch scoring of every night by an experienced scorer plus a double-scored subsample, across ~550 nights, is an unusual volume that most groups would auto-stage, it appears nowhere in the Risks section, and it is a fair question for an assessor. DOWNGRADE: major to minor — real question, unsound premise, overstated arithmetic.
The problem
The host protocol states that data ownership is "secured in a clinical trial agreement with participating centers" and that "If new questions arise… other parties can request to use the data upon filling a Data Sharing proposal form", with the trial's own data-availability statement restricting release to "after publication of findings" on request to the corresponding author. The draft offers no evidence that this project has been approved by the MenoPause Consortium, that the ancillary analysis is permitted, that a data-sharing proposal has been accepted, or that publication is not embargoed behind the trial's primary paper. If the applicants are consortium members, that is easy to demonstrate and should be. If they are not, "has access" is the load-bearing claim in the entire proposal and it is undocumented. Any funder committing three years of money to a secondary analysis will look for this and find nothing.
In the draft — line 60-62
and has access to an ongoing cohort of 222 perimenopausal women with self-reported sleep disturbances
Evidence
Host protocol, https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full: data ownership "secured in a clinical trial agreement with participating centers"; "other parties can request to use the data upon filling a Data Sharing proposal form"; data availability "After publication of findings" and "upon reasonable request to corresponding author". DRAFT.md offers no letter of support, no consortium approval, no data-sharing agreement, and names no host-trial investigator as a team member (though van Baarzel et al. appears under PREVIOUS OWN PUBLICATIONS at lines 449-454, implying overlap that is never stated).
Fix
State the relationship in one sentence and back it with a document: "This project is an approved ancillary analysis of NCT06306404; [name] is an investigator on the host trial and [name] chairs the consortium data-access committee. A letter of support and the approved data-sharing proposal are appended." Address publication sequencing explicitly — whether this analysis can be published before the host trial's primary outcome paper — because a three-year Bial project that cannot publish until the parent trial reports is a schedule risk the funder should be told about.
What the refuter said
The host-protocol language is verified: 'The data collected in this study will be available from the corresponding author upon reasonable request and after publication of findings', with other parties requesting via a Data Sharing proposal form. So the gate is real. But the finding's own parenthesis concedes what defeats it, and I confirmed it: PREVIOUS OWN PUBLICATIONS at L449-454 lists the host protocol — van Baarzel, Broekman, van Dijken, MenoPause Consortium, van Someren, Koning — as the applicant's own publication, and I verified from the medRxiv record that every one of those authors is Amsterdam UMC, the same institution named in the draft's ethics statement. A reviewer looking for provenance of the cohort claim finds it in the one place provenance is normally established. A letter of support from your own institution granting you access to your own trial's data is not standard practice and no funder would expect it. The data-sharing form and the publication embargo apply to third parties, which on the evidence the applicants are not. What survives is thin: the draft never states the relationship in prose, so the reader has to infer it from the publication list. Downgraded major to minor; the substance is really F167's non-disclosure point.
The problem
Bial's ceiling is €60,000, duration at most three years, overheads not reimbursable, and space/equipment-use costs not reimbursable except for a closed list of neuroimaging modalities (MRI, TMS, CT, MEG, PET, SPECT, fNIRS). Ambulatory EEG headbands and wearable wristbands are not on that list, so neither device purchase nor device-use costs are recoverable. Approval rates are roughly 19% (82 of 434 in the 2024 edition). The good news, which the draft never uses: because collection is NWO-funded, the incremental cost here is analyst and scorer time plus modest consumables — a genuinely good fit for the ceiling, and an unusually strong value proposition. The bad news is that the draft's present-tense Procedure section reads as if it is buying recruitment, devices, instruction sessions and measurement weeks, which would blow the ceiling several times over and would be partly non-reimbursable. Because there is no budget at all, the reviewer cannot see which reading is intended.
In the draft — line 210-215
Sleep EEG uses the ZMax headband (Hypnodyne): two frontal derivations at 256 Hz plus accelerometry and PPG. Autonomic activity uses the EmbracePlus (Empatica; wrist EDA, skin temperature, PPG, actigraphy).
Evidence
https://www.fundacaobial.com/en-GB/grants/grants-programme-scientific-research: "Maximum funding per grant: €60,000"; 2024 edition "434 research projects involving 1252 researchers were received, out of which 82 projects were approved for funding". /root/grantreview/BIAL_CONTEXT.md lines 23-27: "Maximum duration 3 years. Project must start between 1 Jan and 31 Oct 2027"; "Overheads and space/equipment-use costs NOT reimbursable, except specified neuroimaging equipment (MRI, TMS, CT, MEG, PET, SPECT, fNIRS)". Host trial funding: NWO NWA.1518.22.104 (https://www.medrxiv.org/content/10.64898/2026.07.22.26358657v1.full).
Fix
Add a budget table in euros within €60,000, containing only reimbursable, incremental items — e.g. postdoc or analyst FTE for detector re-derivation and analysis; scorer time for a bounded number of nights; hydrogel electrode patches and event-button consumables for the amended nights; open-access publication; one conference presentation — with no overhead line and no device purchase or device-use line. Then state the leverage explicitly in one sentence, because it is this application's strongest commercial argument: "Recruitment, devices and measurement are funded by NWO; Bial funding buys the analysis and the added measurement, so a €[x] grant delivers a question that would otherwise cost several hundred thousand euros to field."
What the refuter said
AGAINST: the central charge is an over-reading. A Methods section written in the present tense describes what happens to participants; it is standard scientific prose and no assessor infers a budget from it. Devices, recruitment and instruction sessions are already NWO-funded under the parent trial, so nothing in the draft actually requests unreimbursable equipment costs — the finding convicts the draft of an implication it does not carry. And the budget itself is a BF-GMS form field, not a narrative section, so its absence from a text draft proves nothing. The surrounding facts are all verified: fundacaobial.com gives "Approved applications shall benefit from grants up to a total amount of €60,000" and 434 applications / 82 approved in the 2024 edition (~19%); the regulation confirms overheads and space/equipment-use costs are not reimbursable except for the named neuroimaging modalities, none of which covers an EEG headband or a wristband; and the parent trial is funded by NWO NWA.1518.22.104. The one genuinely useful and confirmed observation is the positive one, which I checked by grep: DRAFT.md never mentions funding, NWO, cost or budget anywhere, so it never makes the strongest argument available to it — that collection is separately funded and the ask is analyst and scorer time, an unusually good fit for a €60k ceiling. Minor, and framed as an opportunity rather than a defect.
The problem
The sole evidence that the core measurement device works is a bioRxiv preprint posted in August 2023 and, as of August 2026, still with no peer-reviewed version findable — three years is long enough that a reviewer will read the absence as a signal. It is also cited without a DOI. Separately, that preprint's usable-data fractions undercut the draft's contingency: 'Data loss is mitigated by recording four nights rather than one' (line 346-347) assumes losses are independent across nights, but the reported failures are device- and participant-level (battery, hinge, fit, contact) and concentrated in the older clinical cohort — the population closest to the draft's. The yield chain lists 'artefact-free recording' as an exclusion but attaches no rate to it.
In the draft — line 392-394
Esfahani, M. J., Weber, F. D., Boon, M., Anthes, S., Almazova, T., Hal, M. V., \... & Dresler, M. (2023). Validation of the sleep EEG headband ZMax. bioRxiv, 2023-08.
Evidence
Preprint DOI 10.1101/2023.08.18.553744, posted August 2023; searches for a published version ("Esfahani ZMax headband validation published Journal of Sleep Research 2024 Dresler peer-reviewed") returned only the preprint, ResearchGate, Semantic Scholar and Sciety records, no journal version. Usable-data fractions from the preprint: "Dataset 1 (26% useful), Dataset 2 (68% useful), Dataset 3 (30% useful), Dataset 4 (63% useful)"; for the 512-patient clinical cohort, "280 (55%) were deemed entirely useful, 366 (71%) contained at least 75% useful data"; authors note "performance degradation in older/clinical populations". Sikder et al. 2026 independently reports 100/130 (77%) usable after requiring PSG plus Zmax plus successful synchronisation, in healthy 18-39-year-olds.
Fix
Cite the preprint as a preprint with its DOI (bioRxiv, https://doi.org/10.1101/2023.08.18.553744) and note that no peer-reviewed version has appeared. Replace the data-loss sentence with a quantified assumption: 'Published usable-data fractions for this headband range from 26% to 71% of recordings depending on cohort, with worse performance in older and clinical samples (Esfahani et al., 2023). We therefore assume 60% of nights yield analysable EEG and carry that through the yield chain; four nights per participant reduces but does not eliminate participant-level loss, since the dominant failure modes are fit and battery rather than night-specific.'
What the refuter said
Verified: the draft's reference at L392-394 gives no DOI; I searched for a peer-reviewed version and found only the preprint, ResearchGate and Semantic Scholar records, consistent with the finding. Fetching the preprint confirmed the sensor set and that arousal detection was not assessed. But two of the four usable-data figures do not check out: I read Donders 2018 26%, Donders 2022 68%, Stockholm 74%, Quantified Self 45%, plus the clinical cohort '55%' and '71%' — the finding reports '30%' and '63%' for datasets 3 and 4. The clinical-cohort numbers match; the dataset-level ones do not, so part of the evidence is wrong. The failure-mode attribution that the independence argument rests on (battery, hinge, fit, contact, concentrated in older cohorts) I could not confirm in the preprint; what the preprint does say about older populations is about the autoscorer, not data loss. STRONGEST DEFENCE: four nights genuinely does mitigate night-level loss (contact, battery, fit vary by night), so L346-347 is not the non-sequitur the finding implies. What survives, and is a fair reviewer objection: the sole validation for the core measurement device is a three-year-old unpublished preprint, cited without a DOI, and L271-276 leans on it for 'Feasibility is supported'. Minor, as filed.
Cheap to fix, and it governs how generously everything else is read.
The problem
Lecci et al. (2017) is the foundational citation of the entire proposal — cited at lines 11, 27, 95 and 274, carrying the sigma ISF, the continuity/fragility phases, the arousability asymmetry, the cardiac coupling and the memory link — and it does not appear in the REFERENCES list. The list runs Baker, Cataldi, De Zambotti, Esfahani, Fernandez & Lüthi, Freedman & Roehrs, Freeman & Sherif, Gombert-Labedens, Kjaerby, Lazar, Osorio-Forero (twice, same 2021 paper), Rothhaas, Sikder, Tsiartas, Wassing: no Lecci. Separately, Osorio-Forero et al. (2025) is cited in the text at lines 60 and 108-111 for a central mechanistic claim but appears only under PREVIOUS OWN PUBLICATIONS (l.461-464), not in REFERENCES. In the other direction, van Baarzel et al. (2026) — evidently the parent study of the 222-woman cohort — is never cited at the point where that cohort is described (l.197-199), so a reviewer cannot verify what the cohort is or what its protocol includes.
In the draft — line 274
established in humans (Lecci et al., 2017; Lazar et al., 2019)
Evidence
grep of DRAFT.md: "Lecci" appears at lines 11, 27, 95, 274 only — all in-text, none in the reference list (l.375-445). "Osorio-Forero et al. (2025)" cited at l.60 and l.108; the 2025 reference appears only at l.461-464 under PREVIOUS OWN PUBLICATIONS. van Baarzel et al. (2026) appears only at l.449-454, with no in-text citation.
Fix
Add the missing Lecci et al. (2017) reference entry and move Osorio-Forero et al. (2025) into REFERENCES (it can remain in the own-publications list as well). Cite van Baarzel et al. (2026) at l.197-199 where the cohort is introduced, so the parent protocol is identifiable.
What the refuter said
Strongest case against: this is proofreading, not science. The sources are real, correctly described in the text, and a missing bibliography entry changes nothing about the study's validity; it is a five-minute fix before submission. That is why major is too high. But every factual claim is verified by grep and stands. "Lecci" occurs at L11, L27, L95 and L274 only — all in body text — and nowhere in the reference list (L375-445), despite carrying the sigma ISF, the continuity/fragility dichotomy, the arousability asymmetry and the human-establishment claim at L274. Osorio-Forero et al. (2025) is cited in text at L60 and L108 for the central mechanistic claim but appears only at L461-464 under PREVIOUS OWN PUBLICATIONS, not in REFERENCES. van Baarzel et al. (2026) appears only at L449-454 with no in-text citation, so the cohort described at L197-199 is unverifiable to a reviewer — which is the concrete cost here, and it is the same gap F36 identifies. Add the independently confirmed defects in the same list (Osorio-Forero 2021 duplicated at L421-424 and L426-429, Wassing formatted as a heading at L445, stray asterisk at L433) and the list reads unproofread. Moderate.
The problem
IAF conventionally denotes individual alpha frequency. The quantity being extracted is the infraslow sigma fluctuation (ISF) phase, the proposal's central predictor. In the one paragraph that defines how the primary predictor is computed, naming a different construct will read to a sleep-EEG reviewer as either a slip or a confusion about what is being filtered.
In the draft — line 269-270
and the Hilbert transform to calculate the IAF phase.
Evidence
Draft line 268-270: "Instantaneous ISF phase is obtained from the log sigma envelope by zero-phase band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant) and the Hilbert transform to calculate the IAF phase." The abbreviation ISF is defined at line 178-179 and used consistently elsewhere; IAF appears nowhere else and is never defined.
Fix
Change to "... and the Hilbert transform, giving the instantaneous ISF phase at each sample."
What the refuter said
Strongest case against: it is a copy-editing slip with no capacity to mislead. The same sentence pins the quantity unambiguously — "the log sigma envelope", "band-pass filtering (0.01-0.04 Hz, empirical peak reported per participant)" — and individual alpha frequency lives at 8-12 Hz, orders of magnitude away from a 0.01-0.04 Hz filter. No sleep-EEG reader could actually mistake what is being computed, so it cannot change a funding decision. But the string is real and verified: grep shows "IAF" occurs at L269 only and is never defined, while "ISF" is used correctly at L178, L250, L258, L265, L267, L295, L306, L341, L343 and L344. Appearing in the one paragraph that defines the primary predictor, in a proposal whose reference list is also unproofread (F72), it reads as carelessness at exactly the wrong place. Correctly filed as minor already — no inflation to correct here.
The problem
Four mechanical defects that together signal a document assembled at speed. (1) Osorio-Forero et al. (2021) appears twice in REFERENCES, at L421-424 and again at L426-429, in slightly different author-list formats. (2) The Wassing et al. (2019) reference at L445 is prefixed with '# ', so it will render as a top-level section heading rather than a reference. (3) A stray '*' trails the Rothhaas & Chung DOI at L433. (4) In the ISF extraction section the text says 'the Hilbert transform to calculate the IAF phase' — IAF conventionally denotes individual alpha frequency, and the intended quantity is the ISF phase. Individually trivial; in the methods section of a competitive application, a reviewer reads (4) as either a typo or a confusion about what is being extracted, and cannot tell which.
In the draft — line 268-269
the Hilbert transform to calculate the IAF phase
Evidence
DRAFT.md L421-424 and L426-429 are the same Osorio-Forero et al. (2021) Current Biology 31(22):5009-5023 reference listed twice; L445 begins '# Wassing, R., Lakbila-Kamal, O., ...'; L433 ends 'https://doi.org/10.3389/fnins.2021.664781 \*'; L268-269 reads 'the Hilbert transform to calculate the IAF phase' while the surrounding section (L265-276) is throughout about ISF.
Fix
Delete the duplicate Osorio-Forero entry, strip the '# ' from the Wassing reference, remove the stray asterisk, and change 'IAF phase' to 'ISF phase'. Proofread the reference block once end to end — it is the cheapest credibility gain available in the remaining days.
What the refuter said
STRONGEST DEFENCE: four cosmetic artefacts of a docx-to-markdown conversion, none of which touches the science; reviewers of grant prose routinely ignore reference-list formatting. That defence fails only because item (4) sits in the sentence that defines the primary predictor. All four verified by direct inspection: L421-424 and L426-429 are the same Osorio-Forero et al. (2021) Current Biology 31(22):5009-5023 entry twice, in different author-list formats; L445 begins '# Wassing, R., Lakbila-Kamal, O., ...' and will render as a level-1 heading; L433 ends 'https://doi.org/10.3389/fnins.2021.664781 \*'; L269 reads 'the Hilbert transform to calculate the IAF phase' while the section (L265-276) and the whole document use ISF, and IAF is never defined. The finding cites the IAF line as 268-269 where it is on 269 — trivial slop, not a misquote. Held at minor, as filed. Worth flagging to the parent: this finding and its two duplicates (F83, F179) all miss a strictly worse presentation defect — Lecci et al. (2017) is cited four times in the body (L11, L27, L95, L274) and appears nowhere in the reference list, and it is the citation carrying the entire premise.
The problem
The Literature Review complies comfortably — 4,861 characters including spaces and its two subheadings, about 81% of the allowance, so there is headroom to add the citations and the eligibility framing recommended elsewhere. But the Literature Review is the only section carrying a stated limit in the draft, and the regulation states no page or character limits. The Summary runs 5,001 characters and the Research Plan 9,920; if the BF-GMS form imposes limits on those fields (which is common and which I could not verify, since the submission system requires authentication), text will be silently truncated at paste time. That is a self-inflicted failure mode with five days left.
In the draft — line 89
# LITERATURE REVIEW (MAX. 6000 CHARACTERS INCLUDING SPACES)
Evidence
Character counts computed with Python over /root/grantreview/DRAFT.md: Literature Review body lines 91-169 = 4,861 characters including spaces (unwrapping the hard line breaks leaves the count unchanged, since each newline becomes one space); including the limit heading itself, lines 89-169 = 4,922. Summary body lines 8-87 = 5,001 characters. Research Plan and Methods, lines 173-373 = 9,920 characters. Regulation: "No page limits stated in the regulation for application sections" (/root/grantreview/BIAL_CONTEXT.md line 37).
Fix
Log into the BF-GMS form today and record the character limit on every field, then trim to fit deliberately rather than discovering truncation on submission night. Spend the ~1,100 characters of Literature Review headroom on the missing Lecci citation and one sentence distinguishing this study's population and design from clinical menopause/insomnia research.
What the refuter said
AGAINST: this finding reports compliance, not a defect, and its central worry is explicitly unverifiable. Its parenthetical justification is also slightly wrong — unwrapping hard line breaks does NOT leave the count unchanged, because blank lines between paragraphs collapse (I get raw 5,001 vs unwrapped 4,995 for the Summary). SURVIVES: I reproduced every number exactly. Lines 91-169 = 4,861 characters raw (81% of the stated 6,000 limit); lines 89-169 = 4,922; Summary lines 8-87 = 5,001 raw / 4,995 collapsed; Research Plan lines 173-373 = 9,920 raw. I also confirmed against the regulation PDF that it states no character or page limits for application sections (only "The report must not exceed 10 pages" for the final report), so the Summary limit genuinely cannot be checked without portal access — the finding is right to say UNVERIFIED rather than guess. The Summary landing 1-5 characters either side of a round 5,000 is a real coincidence worth acting on, and the Bial call closes 31 August 2026 (verified on fundacaobial.com), i.e. five days from today. Keeping minor: a paste-truncation risk to check, not something an assessor scores.
The problem
Bial's regulation states that the objectives against which the project is assessed, and against which the final report is judged, are "those defined in the submitted application, namely in the 'Specific Aims' section", and that "These objectives are essential for the assessment of the application." The draft has no section with that name and, more importantly, nothing that functions as one: the "Research aims" paragraph is continuous prose containing a primary aim, a supporting aim and a prerequisite aim, none numbered, none stated as a measurable objective, none with a success criterion or a decision rule. There is nothing a Scientific Board member can score, and nothing a final report can be checked against. This is a form-compliance failure with substantive consequences.
In the draft — line 174-179
# Research aims\n\nThe project asks how a signal arising in the body reaches reportable conscious experience during sleep, and whether infraslow phase determines this.
Evidence
/root/grantreview/BIAL_CONTEXT.md lines 42-44: "The objectives to be achieved within the scope of the Research Project are those defined in the submitted application, namely in the 'Specific Aims' section. These objectives are essential for the assessment of the application." DRAFT.md contains no heading matching "Specific Aims"; the only aims content is the prose paragraph at lines 176-184.
Fix
Rename the section "Specific Aims" and rewrite as three numbered, falsifiable objectives, each with its test and its deliverable. For example — Aim 1 (prerequisite): characterise the distribution of detected nocturnal hot flash onsets across the ISF cycle in stable NREM, against stage- and time-matched surrogates; deliverable: per-participant ISF period and phase-occurrence distribution, N≥[x] flashes. Aim 2 (primary, confirmatory): test whether ISF phase at flash onset predicts prospective conscious registration (joint sin/cos likelihood-ratio test in a participant- and night-nested logistic GLMM); success criterion: a pre-registered minimum detectable effect of [x] achieved with onset jitter below [y] s. Aim 3 (supporting): test the same for cortical arousal and report the joint distribution of arousal and registration. Add a fourth, honest methodological aim: establish achievable onset-timing precision and frontal ISF expression in perimenopausal women using ambulatory EEG and a wrist multisensor — because that is genuinely novel and it is what the rest of the proposal depends on.
What the refuter said
Strongest case against: the consequential claim is false. "There is nothing a Scientific Board member can score, and nothing a final report can be checked against" does not survive reading L176-184, which states three identifiable, testable aims with named predictor and outcomes — phase of the ISF at onset predicting conscious registration; the same for cortical arousal; and whether onsets are themselves phase-structured — and the confirmatory test is specified precisely at L291-292 as a likelihood-ratio test of the joint sin and cos term. That is scoreable and checkable. Moreover this document is narrative content for the form's free-text boxes: its own headings mirror form fields ("LITERATURE REVIEW (MAX. 6000 CHARACTERS INCLUDING SPACES)", "RESEARCH PLAN AND METHODS"), so the absence of a box labelled "Specific Aims" from a prose draft is not evidence the field will be left empty. Fatal collapses entirely. What is verified and worth one line of advice: grep finds no "Specific Aims" heading in DRAFT.md, and the Regulation does name that section — "The objectives to be achieved ... are those defined in the submitted application, namely in the 'Specific Aims' section. These objectives are essential for the assessment" — so rename the heading and number the three aims. Minor.
The problem
Osorio-Forero et al. 2021 appears twice in the reference list, in two different formats (lines 421-424 and 426-429). The Rothhaas & Chung entry ends with a stray asterisk (line 433), an editing artefact. The Wassing et al. 2019 reference is formatted as a level-1 heading ("# Wassing, R., …", line 445), so it will render as a section title rather than a reference. And at line 270 the sentence that defines the study's central predictor uses an undefined acronym, "the IAF phase", where ISF is meant. Individually trivial; collectively they tell a reviewer the document was not read through before submission, which colours how generously the substantive gaps are read.
In the draft — line 421-429
Osorio-Forero, A., Cardis, R., Vantomme, G., Guillaume-Gentil, A., Katsioudi, G., Devenoges, C., Fernandez, L. M. J., & Lüthi, A. (2021). Noradrenergic circuit control of non-REM sleep substates. Current Biology, 31(22), 5009-5023.
Evidence
DRAFT.md lines 421-424 and 426-429: the same Osorio-Forero et al. (2021) Current Biology 31(22):5009-5023 entry, twice. Line 433: "https://doi.org/10.3389/fnins.2021.664781 \*". Line 445 begins "# Wassing, R., Lakbila-Kamal, O., …". Line 270: "the Hilbert transform to calculate the IAF phase" with no prior definition of IAF; ISF is defined at line 178.
Fix
Delete the duplicate Osorio-Forero 2021 entry, remove the stray asterisk, demote the Wassing entry from a heading to a normal reference, and change "the IAF phase" to "the instantaneous ISF phase". While there, check the Cataldi et al. (2026) and Sikder et al. (2026) entries against the published records, since 2026 citations cannot be verified from memory.
What the refuter said
All items verified by direct inspection: L421-424 and L426-429 are the same Osorio-Forero et al. (2021) Current Biology 31(22):5009-5023 entry in two formats; L433 ends 'https://doi.org/10.3389/fnins.2021.664781 \*'; L445 begins '# Wassing, R., Lakbila-Kamal, O., ...' and will render as a section heading; L269 uses 'IAF phase' undefined where ISF is meant, with ISF defined at L178. Held at minor, as filed. This is the third filing of the same finding (F11, F83, F179) — the parent should merge them into one; three separate entries for one set of typos inflates the apparent defect count. All three share a blind spot worth reporting instead: Lecci et al. (2017) is cited four times in the body (L11, L27, L95, L274), carries the entire premise of the proposal, and appears nowhere in the reference list — a strictly worse reference-list defect than a duplicate entry or a stray asterisk, and the one a reviewer trying to check the premise would hit first.
The problem
Three separate overclaims in one sentence, none cited. First, a phase effect on registration would not explain the report gap: the gap is a count over a whole night and a phase effect is a within-cycle probability, so at best it says some flashes are less likely to be noticed - which is the observation, not an explanation. Second, the premise that treating flashes does not reliably resolve the sleep complaint is contradicted by the draft's own hot-flash review, which reports that hormone therapy reduces sleep disturbance and that neurokinin-antagonist trials improved sleep quality and reduced awakenings. Third, and structurally worst, this sentence is a therapeutic claim in a proposal to a funder that excludes therapy. The proposal spends its last breath asserting something it cannot deliver, cannot support, and should not be claiming to this funder.
In the draft — line 370-373
It would also give a mechanistic account of a clinical puzzle - why women report far fewer nocturnal hot flashes than they have, and why treating the flashes does not reliably resolve the sleep complaint that accompanies them.
Evidence
Gombert-Labedens et al. 2025 (/root/grantreview/GombertLabedens2025.txt ~L903-905): "The strong link between hot flashes and sleep disturbance is further supported by studies showing that the treatment of hot flashes ... with hormone therapy reduces sleep disturbances". Same source ~L1424-1427: "Preliminary evidence from phase 2 trials of elinzanetant provided evidence of improvements in sleep quality and reductions in awakenings from sleep ... which was supported in Phase 3 trials". The draft offers no citation for its contrary claim (L370-373).
Fix
Delete the sentence and end on the mechanistic contribution. Replacement: "Establishing that an infraslow rhythm sets when an internal bodily event becomes reportable would connect a mechanism already known to govern arousability and overnight memory to the threshold of awareness, in healthy humans and without intervention." If a bodily-relevance sentence is wanted, keep it descriptive and non-therapeutic: "It would also supply a psychophysiological account of why objectively recorded nocturnal flashes differ so widely in whether they are noticed at all."
What the refuter said
Verified: L370-373 carries no citation, and the review's contrary evidence is real — 'the treatment of hot flashes ... with hormone therapy reduces sleep disturbances [210-213]' (GombertLabedens2025.txt L903-905) and 'Preliminary evidence from phase 2 trials of elinzanetant provided evidence of improvements in sleep quality and reductions in awakenings from sleep [285] ... which was supported in Phase 3 trials [282]' (L1424-1427). But each of the three prongs weakens on inspection. 'Does not reliably resolve' is not contradicted by 'reduces' and 'improvements' — reducing a complaint and reliably resolving it are different claims, and the draft chose the weaker verb. Prong 1 misreads 'a mechanistic account': explaining why some flashes go unregistered IS a mechanistic account of an aggregate report deficit, so the objection that a within-cycle probability cannot explain a nightly count is too strong. Prong 3 is a stretch — the sentence treats treatment as a puzzle to be explained, not as project scope, and Bial's exclusion is about what a project involves. What survives is worth one line to the applicant: an uncited empirical claim in the closing sentence, contrary to the balance of the draft's own cited review, is exactly the sentence a reviewer quotes back. Downgraded major to minor.
The problem
Four mechanical defects that together signal an unfinished document. (i) Osorio-Forero et al. 2021 appears twice in REFERENCES as near-identical entries (lines 421-424 and 426-429, differing only in whether the author list is elided) and a third time under PREVIOUS OWN PUBLICATIONS (lines 456-459). (ii) The Wassing entry begins with a markdown '#', so it renders as a section heading rather than a reference, orphaning it from the list. (iii) The Rothhaas & Chung entry ends with a stray asterisk after the DOI. (iv) In ISF extraction the text switches acronym mid-sentence: 'the Hilbert transform to calculate the IAF phase' where every other mention is ISF. DOI coverage is also inconsistent — present for Baker, Cataldi, De Zambotti, Fernandez & Lüthi, Lazar and Rothhaas, absent for Esfahani, Freedman & Roehrs, Freeman & Sherif, Gombert-Labedens, Kjaerby, Osorio-Forero 2021 and Tsiartas.
In the draft — line 445
# Wassing, R., Lakbila-Kamal, O., Ramautar, J. R., Stoffers, D., Schalkwijk, F., & Van Someren, E. J. (2019). Restless REM sleep impedes overnight amygdala adaptation. *Current biology*, *29*(14), 2351-2358.
Evidence
DRAFT.md 421-429: two Osorio-Forero et al. (2021) entries, both "Noradrenergic circuit control of non-REM sleep substates. Current Biology, 31(22), 5009-5023"; repeated again at 456-459. Line 433: "https://doi.org/10.3389/fnins.2021.664781 \*". Line 445 begins with '# '. Line 269: "the Hilbert transform to calculate the IAF phase".
Fix
Delete the duplicate Osorio-Forero 2021 entry from REFERENCES (keep one, with full author list); remove the '# ' from the Wassing entry so it sits in the list; delete the trailing asterisk after the Rothhaas & Chung DOI; change 'IAF phase' to 'ISF phase'; and add DOIs to the seven entries lacking them.
What the refuter said
STRONGEST DEFENCE: these are mechanical artefacts of a Word-to-Markdown conversion, invisible in the submitted online form, and no reviewer rejects a proposal over a stray asterisk. WHY IT SURVIVES: all four are real and I verified each by reading the file. (i) Two near-identical Osorio-Forero et al. 2021 entries at L421-424 and L426-429, both "Noradrenergic circuit control of non-REM sleep substates. Current Biology, 31(22), 5009-5023", plus a third at L456-459. (ii) L445 begins "# Wassing, R., ..." so the entry renders as a heading, not a reference. (iii) L433 ends "https://doi.org/10.3389/fnins.2021.664781 \*". (iv) L269 reads "the Hilbert transform to calculate the IAF phase" where every other mention is ISF — and IAF conventionally denotes individual alpha frequency, so in the one sentence that defines the primary predictor the acronym names the wrong quantity. That last one is the only item with any teeth: it sits in the operational definition of the predictor. Severity stays minor, correctly assigned by the finding.
The problem
A list of defects that individually are cosmetic but together signal an unproofed submission in the sections where precision matters most. (1) l.269: "IAF phase" — IAF is individual alpha frequency; the intended quantity is the ISF phase. This appears in the single sentence defining how the primary predictor is computed. (2) l.332: "reference [number]" — the ethics approval reference is a placeholder while the sentence asserts approval as fact. (3) l.421-429: Osorio-Forero et al. (2021) is listed twice in REFERENCES, once with the full author list and once truncated with an ellipsis; the same paper appears a third time at l.456-459 under previous publications, with "Luthi" unaccented there versus "Lüthi" in the reference list. (4) l.433: a stray "*" after the Rothhaas & Chung DOI. (5) l.445: the Wassing et al. reference is formatted as a level-1 heading ("# Wassing, R., ..."), so it will render as a section title. (6) l.272-273: "the smallest bias of any band against polysomnography (0.005 ± 0.012; Esfahani et al., 2023)" gives a bias with no units and no metric — uninterpretable as written, and Esfahani et al. is cited as a bioRxiv preprint for the key feasibility claim about the recording device. (7) l.237: an arousal scoring rule is attributed to a substantive research paper (Wassing et al., 2019) rather than to the scoring manual, and is stated incompletely (">16 Hz activity lasting 3-15 s" omits the alpha/theta shift and the preceding-stable-sleep requirement). (8) Escaped apostrophes survive from the conversion at l.33 and l.117 ("body\'s").
In the draft — line 269
the Hilbert transform to calculate the IAF phase
Evidence
Direct quotes at the lines given: l.269 ("IAF phase"), l.332 ("reference [number]"), l.421-424 and l.426-429 (duplicate Osorio-Forero 2021), l.433 (trailing asterisk), l.445 ("# Wassing, R., ..." as a heading), l.272-273 ("0.005 ± 0.012" unitless), l.237 ("transient > 16 Hz activity lasting 3-15 s (Wassing et al., 2019)"), l.33 and l.117 ("body\'s").
Fix
Fix each: "IAF" to "ISF"; insert the actual ethics reference number (required before any agreement can be signed under the Bial regulation); delete the duplicate Osorio-Forero 2021 entry and standardise "Lüthi"; delete the stray asterisk; demote the Wassing entry from a heading to body text; give the ZMax bias its units and metric ("mean bias in relative sigma power of 0.005 ± 0.012, in units of [x]") and note the preprint status; cite the AASM scoring manual for the arousal rule and state it in full ("an abrupt shift in EEG frequency including alpha, theta and/or frequencies greater than 16 Hz, but not spindles, lasting at least 3 s, with at least 10 s of stable sleep preceding"); repair the escaped apostrophes.
What the refuter said
Items 1-5 and 8 verified verbatim by direct inspection: L269 'IAF phase'; L332 'reference \[number\]'; L421-424 and L426-429 duplicate Osorio-Forero entries plus a third at L456-459 with 'Luthi' unaccented; L433 trailing '\*'; L445 '# Wassing, R., ...' as a heading; L34 and L119 'body\'s'. Item 6's quote is real (L272-273). Two items are padding and should be dropped before this reaches the applicant. Item 6 overstates: I fetched Esfahani and the 0.005 ± 0.012 figure is a dimensionless relative-bandpower bias (their sigma band being 13-16 Hz, incidentally, against the draft's 11-16), so 'no units' is true because the quantity has none — 'uninterpretable as written' is wrong; the fair complaint is that the metric is unnamed. Item 7 is not a defect at all: citing a scoring criterion to a research paper that used it, rather than to the scoring manual, is normal practice, and Wassing et al. is the applicants' own group. Held at minor. Third of three duplicate presentation findings (F11, F179), and like them it misses that Lecci et al. (2017) — four body citations, carrying the premise — is absent from the reference list entirely.
The problem
Art. 4(1) makes the 'Specific Aims' section the operative statement of what the grant holder is contracted to deliver, and BIAL_CONTEXT records it as 'essential for the assessment of the application'. The draft's equivalent section is headed 'Research aims' and is written as flowing prose, with the three aims (primary registration test, supporting arousal test, prerequisite occurrence test) embedded in sentences rather than enumerated. The same mismatch appears at the other end: the agreement's Clause Fourth binds the holder to the application's 'Expected outputs', while the draft's section is headed 'Expected outcomes' and contains discussion of significance rather than deliverables. These are cheap to fix and they are exactly the fields an assessor uses to score.
In the draft — line 173
# Research aims
Evidence
Regulation Art. 4(1), verbatim: "The objectives to be achieved within the scope of the Research Project are those defined in the submitted application, namely in the 'Specific Aims' section." Clause Fourth(1), verbatim: "The Grant Holder undertakes to execute the Research Project as described in the Application, including, but not limited to, the 'Expected outputs'." Grep of DRAFT.md: "Specific Aims" = 0, "Expected outputs" = 0. Draft headings at L173 ('Research aims') and L351 ('Expected outcomes').
Fix
Rename the section 'Specific Aims' and rewrite it as three numbered aims, each one sentence, each with its own testable outcome — Aim 1 (prerequisite): whether flash onsets are distributed uniformly across ISF phase; Aim 2 (primary/confirmatory): whether ISF phase at onset predicts conscious registration; Aim 3 (supporting): whether ISF phase at onset predicts cortical arousal. Add a short 'Expected outputs' list (pre-registration, dataset, N publications, methods deliverable) distinct from the significance narrative.
What the refuter said
AGAINST: the draft's ALL-CAPS headings ("LITERATURE REVIEW (MAX. 6000 CHARACTERS INCLUDING SPACES)", "RESEARCH PLAN AND METHODS", "REFERENCES", "PREVIOUS OWN PUBLICATIONS") are plainly BF-GMS form field labels, while "# Research aims" and "# Expected outcomes" are the author's own sub-headings inside a free-text field. An applicant cannot rename a form field, so "rename your heading to Specific Aims" may be attacking the wrong object — and the substance is present: L177-184 states three identifiable aims (primary registration test, supporting arousal test, prerequisite occurrence test) clearly enough to be assessed. SURVIVES on the facts. I fetched the regulation PDF and confirmed all three quoted clauses: Art. 4(1) "The objectives to be achieved within the scope of the Research Project are those defined in the submitted application, namely in the 'Specific Aims' section"; Clause Fourth(1) "...including, but not limited to, the 'Expected outputs'"; and additionally Clause Tenth(1)(a) "Failure to address the Specific Aims set forth in the Application and reflected in this Agreement" — so the named sections are contractually operative, not stylistic. DRAFT.md contains neither string, and the aims are prose rather than enumerated. Severity cut from moderate to minor: the assessable content exists, the fix is labelling and numbering, and no reviewer decision turns on it.
45 findings
These were raised and did not survive. They are here so nothing looks hidden, and so nobody spends the next five days fixing things that are not broken.
Strongest case against: DRAFT.md's headings ("Project title", "Summary", "LITERATURE REVIEW (MAX. 6000 CHARACTERS INCLUDING SPACES)", "RESEARCH PLAN AND METHODS", "REFERENCES", "PREVIOUS OWN PUBLICATIONS") map exactly onto the narrative free-text fields of the BF-GMS online form. CVs, the Host Entity declaration, the payment plan and the budget are separate form fields and uploads, not narrative prose. I verified the component list at https://www.fundacaobial.com/en-GB/grants/grants-programme-scientific-research: "CV for each team member - maximum 4 pages", "Host Entity declaration - signed commitment of institutional support", "Payment plan", each listed as a distinct application component alongside the research proposal. Deadline (31 August 2026) and cap ("up to a total amount of EUR60,000") both confirmed there. The grep results are real (host=0, budget=0, Euro=0, 2027=0, CV=0), but they establish only that a narrative document contains no form-field content, which is the expected state of a narrative document. Inferring from that that the other components do not exist is unwarranted; and the fact that a deadline is five days away is a fact about the calendar, not a defect in the draft. That the draft is "the narrative half of one" is exactly what was handed over for review. This is standard practice mistaken for an error. One item does survive: a project schedule/duration is arguably expected inside the research plan and is absent (grep: the only two "schedule" hits are the rhetorical "the body's own schedule" at L34 and L119; no duration, no milestones). That is carried by F7 and does not support a fatal eligibility finding.
Strongest case against: the finding's claim about the draft is false, and I checked it line by line. It says the interoception/conscious-access hook "is stated once in the title and once in the closing paragraph, and never connected to the awareness/consciousness literature". In fact it appears in the title; three times in the Summary (L12-15 "during sleep the brain does not only receive signals from the outside, but also monitors interoceptive signals"; L54-56 "modulates access to reportable awareness"; L85-87 "brings together sleep neuroscience, neuroendocrinology and consciousness science"); as an entire Literature Review paragraph on interoceptive vs exteroceptive processing (L118-127, incl. Cataldi 2026); in the Research aims (L175-177 "how a signal arising in the body reaches reportable conscious experience"); and in Expected Outcomes (L363-369 "a question about interoception and conscious access"). The draft_quote supplied (L266-269, the phase-extraction paragraph) has nothing to do with the stated problem. The portfolio evidence itself is verified — the 14th symposium programme contains no sleep-EEG, sigma or infraslow titles, and 432 applications / 80 approved / 52-18-30 field split checks out — but funder-portfolio character is context, not a defect in the draft. The only residual (cite the consciousness literature Bial reviewers know) is F150's finding, better evidenced there. Does not survive.
Strongest case against the finding, which I find decisive on both prongs. (i) Eligibility: BIAL_CONTEXT.md L10 excludes "Projects involving clinical or experimental models of human disease and therapy". That is a clause about what a project *does* - it bites on disease models and administered therapies, not on a motivational sentence in Expected Outcomes. This project administers nothing, recruits physically healthy women, and explicitly frames its population as being in "a normal stage of reproductive life" (L37-38) and its work as "in healthy humans and without intervention" (L369-370). Reading an impact sentence as triggering a design exclusion is an over-reading; virtually every funded psychophysiology proposal names a clinical relevance. Bial's "healthy human being" phrasing comes from a press release, not the regulation. (ii) The logical objection fails on its own terms. The draft says the work would "give a mechanistic account of a clinical puzzle", not test treatment response. If a large share of nocturnal flashes never reach registration, that is precisely a mechanism explaining why flash-directed treatment need not resolve the sleep complaint - the account requires no observed treatment. F85 makes the well-calibrated version of this concern at the right severity; graded major, this one is inflated.
AGAINST: this is a finding with no evidence, and it says so itself ("UNVERIFIED"). It reports that a bar *might* apply to *someone unnamed*. BIAL_CONTEXT.md L18 is real, but L20 records that the application requires a "CV of applicant and every team member" — the ongoing-grant bar is checked administratively from those CVs, not from the narrative document under review. A grant narrative that describes the team by capability rather than by name is standard, and no reviewer treats it as an eligibility defect. Does not survive. There is no verified fact about the draft here, only an unresolved fact about the world, and even if a team member did hold a Bial grant the defect would lie in the CV attachments, not in DRAFT.md. Reporting it invites the applicant to act on a possibility no one has established. Nitpick, and not even about the document.
STRONGEST DEFENCE, WHICH HOLDS: the draft does not use the Tsiartas classifier to time onsets. The very next subsection (DRAFT.md L256-263, "Onset timing") states: "Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing." Classification (which events are flashes) and onset estimation (when the EDA inflection occurs) are two separate procedures, and separating them is exactly the right architecture. The finding's central claim — "the detector cited as the enabling measurement is provably unfit for the enabling measurement" — attacks a role the draft never assigns it. The ±90 s matching window and 15 s frame grid are real (Tsiartas2021.txt: "±90 s matching window"; "makes a decision every 15 s"), as are the ±30/±120/250/500 s feature windows, but those bound the candidate-region search, not the change-point estimate. Tsiartas's own reference [4] (Forouzanfar et al., "Automatic detection of hot flash occurrence and TIMING from skin conductance activity") shows onset timing is a separate, existing capability in this lineage. The draft additionally makes achievable precision a pre-registered feasibility outcome (L261-263, L275-276) with a pre-specified switch to pre-flash ISF amplitude if dispersion exceeds threshold (L338-343). A "fatal" verdict on a point the draft pre-empts in the adjacent paragraph is the definition of an over-reading. Residual worth one line, not a finding: Tsiartas reports no onset-accuracy metric, so the draft's precision is unvalidated — which the draft itself says.
This is the textbook case of a finding that ignores the draft's pre-emption in the very next sentence. The finding quotes L246-249 and stops. L249-254 continues: 'It does use autonomic signals that may vary with ISF phase, so phase-dependent sensitivity cannot be assumed absent; a pre-specified analysis repeats detection from skin conductance and temperature alone, since flashes that arouse show a larger cardiac response (Baker et al., 2019) and a cardiac-weighted detector would preferentially find them.' The pre-specified fallback detector excludes BOTH the motion feature and the cardiac feature — i.e. exactly the two channels the finding's mechanism runs through. The draft names the bias, names its direction, and pre-specifies the sensitivity analysis. Second, the draft's use of 'circular' is narrow and correct: the label is not derived from the outcome. What the finding describes is differential detection sensitivity, which is a different thing and is what L249-254 addresses under its own name. Third, the mechanism as stated does not bias the phase effect: press-driven motion inflating detection makes registered flashes over-represented among detected flashes, which shifts the base rate, not the phase-registration association, unless phase also drives detection — again the thing the draft names. Fourth, the Shapley characterisation is inflated: for sleep-onset flashes Fig. 6 gives SC+ 58.3%, PPG 30.4%, M 7.9%, T 3.4% (verified in Tsiartas2021.txt L246-257), so motion is a distant third, barely above temperature. The ±30 s motion window quote is real (Tsiartas2021.txt L75-81) but load-bearing on nothing.
Strongest case against: the alleged internal contradiction does not exist, and the finding self-defeats. L306-309 says the model does not *adjust* for magnitude; L283/L298 say the model is fitted to detected flashes. These are different things: not conditioning on a mediator as a covariate is a modelling choice, whereas being unable to analyse events no instrument recorded is an unavoidable measurement limit — no design can include undetected flashes, so this cannot be a defect of this design specifically. The draft also pre-empts the substance twice: L249-254 concedes "phase-dependent sensitivity cannot be assumed absent" and pre-specifies re-detection from skin conductance and temperature alone, and L309-310 makes a magnitude-adjusted model a stated sensitivity analysis. Worse, the finding's own evidence undercuts its mechanism: it cites GombertLabedens2025 (verified at source, right column p.~110) "the amplitude of the skin conductance does not represent the self-reported severity of a hot flash [193]" to argue the magnitude-to-registration link is missing. But that link is exactly what makes magnitude a common cause of detection and registration; without it the collider bias it alleges largely vanishes. The residual estimand point (a total effect is compatible with phase-makes-flashes-bigger) is real but is addressed in kind by the arousal-versus-registration comparison at L54-56. Does not survive as stated.
Strongest case against, and it holds. Sub-claim (i) inverts the draft's own logic: retrospective forgetting IS a measurement problem (recall bias), so it falls inside the category the draft says the field has used, not outside it. The finding's own evidence proves the point - I read GombertLabedens2025.txt around L858-861 in context, and "hot flashes that occur during sleep may be under-reported retrospectively in the morning [194,195]" sits inside a passage about symptom-assessment instruments ("Daily symptom diaries yield detailed data with minimal risk of recall bias, however, they can be burdensome"). That is exactly "treated as a measurement problem". No false dichotomy. Sub-claim (ii) is backwards: a prospective, time-stamped press is the instrument that separates non-registration from forgetting, so the 3.5-vs-1.5 gap motivates the design rather than being dissolved by it. Sub-claim (iii) fails for the same reason - if a substantial share of flashes are never registered at the time, that is a direct partial answer to why morning reports fall short. On the numbers: I could not reach the de Zambotti 2014 abstract through PubMed (reCAPTCHA), but Europe PMC returned "3.5 hot flashes per night (95% CI: 2.8-4.2, range 1-9), with 69.4% associated with awakening" and no self-report count - so "1.5" remains UNVERIFIED. That said, the draft's "19.8%" (L149-151) is more precise than the review's rounded 20%, which indicates first-hand access to the paper rather than a laundered figure. Graded major, this is not sustainable.
AGAINST: the draft's clause claims only that sleep-onset performance is "reported separately" — and it is. Tsiartas2021.txt Analyses 2: "Evaluated the HF classification performance as a function of whether the HF onset occurred during wake or sleep", with Fig. 4 devoted to "System performance for hot flashes (HFs) onsets occurring during sleep vs wake". Point (i) is the reviewer's own misreading: "sleep-onset performance" in context of that paper plainly means performance for flashes whose onset occurred during sleep, not performance at the wake-sleep transition. The residual observations are true but do not land on the draft. Yes, no numbers appear in the text and Fig. 4 carries only axis labels — but the draft asserts no numbers for that condition, so nothing is misquoted. The alleged self-contradiction (L233-235 vs L236-240) is real for PPG only (Fig. 6: PPG 30.4% sleep vs 16.5% wake, against the second sentence's "greater ... for the PPG, SC, and M sets ... wake vs sleep") — but it concerns Shapley *contributions*, which the draft never cites. A citation-accuracy finding requires the draft to have said something the source does not support; here it did not. Minor at most, as a note that the sleep-condition figures are unlabelled.
Strongest case against, and it is decisive: the finding checks the draft against the wrong source. The draft cites Freedman & Roehrs (2006) at L41-44 and L143-145, not Gombert-Labedens. Freedman & Roehrs' own abstract (PubMed 16837879) reads: "In the second half of the night, rapid eye movement sleep suppresses hot flashes and associated arousals and awakenings." "Suppresses" is the cited source's own verb. The draft is therefore faithful to its citation, and the charge that it "overstates 'less likely in REM'" is an artefact of substituting a review's summary of three other studies for the primary paper the draft actually cites. The GombertLabedens quotes are real — I verified both, including "possibly due to the lower sensitivity of the thermoregulatory system" and "although this hypothesis remains to be tested" — but a review's hedge about other studies does not make a direct citation of a primary experimental result an overstatement. Residual, genuinely minor: the untested less-intense-in-REM alternative is worth a clause, and the draft's mechanism wording ("largely switched off") is firmer than the review's hedge. The draft also already uses Freedman & Roehrs correctly for the across-night reversal and carries half-of-night as a covariate (L289-291).
AGAINST: the quantitative core misattributes the numbers. The "+10.7% in sensitivity" is Tsiartas's comparison of the SC+ set against the bare SC set ("At 96.5% specificity, the SC+ set showed better HF classification performance than the SC set (+10.7% in sensitivity)") — a comparison between two skin-conductance-only feature sets, not between the draft's SC-plus-temperature fallback and the full multi-sensor detector. Likewise "below 75%" is stated for "the SC features only", i.e. the bare SC set. The finding's chain — temperature contributes little, therefore SC+T ≈ SC-only, therefore −11pp and <75% under corruption — requires SC ≈ SC+, which the source explicitly refutes by 10.7 points. Does not survive as stated. What remains is unquantified and already conceded by the draft: removing PPG costs some sensitivity, and electrodermal activity stays sympathetically coupled to the rhythm (the draft itself says "phase-dependent sensitivity cannot be assumed absent", L250-251). Minor.
The quotes are verbatim correct — GombertLabedens2025.txt: "sternal skin conductance has adequate sensitivity (0.69, proportion of hot flashes identified by a rise in sternal skin conductance that were accompanied by self-report) and specificity (0.97 ...)" and "Factors like body mass index, negative mood, and stress can influence the concordance ...". But the inference is wrong. 0.69 is a VALIDATION statistic computed against self-report, not a component of the criterion. The same source states the criterion itself, two paragraphs later, in purely physiological terms: "an observed rapid rise of at least 2 microSiemens in sternal skin conductance within a 30-second period". Tsiartas trained on exactly that — "Physiological HFs were recorded and scored (2 uS/30s rises in SC) by experienced scorers" — with zero self-report input. So there is no dependence for the wrist detector to "inherit by construction", and the draft's L246-249 defence ("the detector is never trained on button presses ... so it cannot be circular with respect to registration") is correct as stated. Worse for the finding: a 0.69 physiological/self-report divergence is not a measurement defect, it is the *phenomenon the draft proposes to study*. The residual — BMI/mood/stress as person-level modifiers of report probability — the finding itself concedes is absorbed by participant random intercepts. Nothing decision-relevant left.
The finding is filed fatal and its central factual premise is unverified. It asserts that the de Zambotti self-report comparator is 'a morning retrospective count'. Its only evidence is a general methodological remark in a review — 'hot flashes that occur during sleep may be under-reported retrospectively in the morning [194,195]' — which cites de Zambotti among others but says nothing about how de Zambotti's own comparator was collected. The finding also concedes it could not locate the 1.5 figure at all. So it asserts that an unlocated number is a morning recall measure. I could not verify either way: Europe PMC's full-text endpoint returned 429 on four attempts, PMC and ScienceDirect returned reCAPTCHA and robots.txt blocks. What I did retrieve (Europe PMC core record for DOI 10.1016/j.fertnstert.2014.08.016) confirms it was a laboratory PSG study, 34 women, 63 nights, 3.5 objective flashes/night, '69.4% of hot flashes were associated with an awakening' — and lab protocols in this group commonly use an in-night event marker, so the finding's premise is not even the default reading. Two further defects. The 'quiet concession' at L218-219 is the draft doing the right thing: it adopts a prospective time-stamped press as the registration measure and explicitly relegates the morning diary to 'a complementary measure of recall and reporting behaviour'. And L152-154 is misread — 'This discrepancy has been treated as a measurement problem rather than as evidence about how endogenous signals gain access to awareness' claims the gap may be informative, not that it is registration failure.
AGAINST: the headline is wrong on its own evidence. I fetched PubMed 16837879 and "suppresses" is Freedman & Roehrs's own word in their own conclusion: "In the second half of the night, rapid eye movement sleep suppresses hot flashes and associated arousals and awakenings." The draft's "flashes are suppressed during REM sleep" (L42-43) therefore tracks the cited source's wording, and the draft attributes it to that study rather than asserting a settled literature. Population differences (18 postmenopausal plus 12 cycling women, lab) are a caveat, not a misquote. Point (iii) also fails on mechanism. Ambient temperature varies between nights and slowly within a night; the primary predictor is instantaneous phase on a 50-s cycle. A variable that is near-constant across 50 s cannot confound a within-night phase contrast, and night-level variation is already absorbed by the draft's "random intercepts ... for night within participant" (L288-289). Bedroom temperature affects how many events you get, not which phase they land in. Nothing here would change a reviewer's decision or the study's validity.
Strongest case against, and it holds. "A majority of women experience them across the menopausal transition" is a cumulative, transition-wide claim, and it is amply supported by the draft's own 2025 source, which I read at GombertLabedens2025.txt L246-247: "In the USA, hot flashes occur in 50 to 82% of individuals who experience natural menopause", plus "The median duration for vasomotor symptoms is 7.4 years". The SWAN figure the finding leans on ("about 48% reporting hot flashes in the final year before last menstrual period", L275-281) is a point prevalence in a single year, derived from monthly yes/no reports - it is not a cumulative-across-the-transition rate and therefore does not contradict the draft's sentence. The finding's inference from a one-year point prevalence to "a majority is not established for the perimenopausal window" is a category error. The subsidiary complaint that the citation is nineteen years old is not a defect: a coarse prevalence claim does not decay, and Freeman & Sherif remains the standard systematic review for it. At most this is a stylistic preference for the newer citation, inside a 6,000-character section.
Strongest case against: no defect is demonstrated, and the finding says so itself — "the pages may be an assignment the applicant has from a proof rather than a published record — in which case it is fine". Everything checkable checks out: the article is real, and I confirmed the title, authors and journal independently (cell.com/current-biology/fulltext/S0960-9822(26)00889-4, and the CIBM institutional announcement). I could not verify the page range either: api.crossref.org is robots-disallowed from this session and the publisher page returns 403, so the finding's own inability to confirm is reproducible, but that is a limit on the checker, not an error in the draft. No grant reviewer verifies page ranges, and an August-2026 online-first article plausibly has proof pagination. Reporting an unverifiable-but-probably-correct page number as a finding dilutes a review; the "invites suspicion of the whole bibliography" rationale is F72's point, which stands on its own. Too trivial to report.
AGAINST: a pre-registered direction-agnostic omnibus test is more conservative than a directional one, not less, and the draft is explicit rather than sly about the two-sided expectation — "a direction predicted rather than defining" (L360-361). Pre-specifying an omnibus 2-df test and reporting which phase is preferred descriptively is standard, defensible practice; committing to a direction on a rodent-derived analogy would be worse. The finding's strongest claim also fails on the text. It asserts "a null is uninterpretable" and that the draft "provides no way to distinguish a true null from measurement failure" — but L319-320 says "Simulations use the observed counts and precision, and the minimum detectable effect is reported", which is precisely the machinery for interpreting a null, and L275-276 pre-registers achievable phase precision and frontal ISF expression as feasibility outcomes measured independently of the primary result. So a null is bounded against an MDE and against measured precision. What is left is that "any preferred phase" is a weaker hypothesis than the rhetoric implies — a framing complaint, minor.
This misreads the draft twice. First, "leave no trace" at L38-39 is inside a sentence about awareness — "they are sometimes consciously registered, and other times leave no trace" — so "trace" there means a conscious trace, not a polysomnographic one. The finding reads it as "no trace in sleep" and then convicts the draft of a claim it did not make. Second, the draft plainly assigns the two statistics to two different jobs at L149-152: 19.8% carries "occur without disturbing sleep", and 1.5-of-3.5 carries the report gap that the following sentence calls "This discrepancy". Nothing is being smuggled. The finding's own "much more interesting phenomenon" — flashes that disturb sleep but go unreported — is exactly the draft's target: L299-302 states registration is "expected to be largely a subset of aroused ones" and commits to modelling "registration additionally modelled among aroused flashes". So the draft is already pointed at the 37 percentage points. The power limb also fails: a base rate of 0.78 gives variance 0.172, 69% of the maximum 0.25 — not "badly imbalanced" in any material sense — and it contradicts F93, whose analytics (which I checked) show the arousal test is the MORE informative one. Finally, the draft's figures are conservative: GombertLabedens2025.txt reports a larger study (n=86) in which "29% of hot flashes occurred in undisturbed sleep". Minor residue only: "leave no trace" deserves one clarifying word.
Not a contradiction. Power in this design is a function of two arguments, event count and onset precision. 'Power depends on event count, hence multiple nights' (L191-193) and 'the binding constraint is not event count but the precision with which onset can be timed against it' (L79-80, L315-316) are simultaneously true of any such function in which one argument saturates — and the draft states the saturation explicitly in the sentence immediately following the second quote: 'simulation shows power ceasing to improve with additional events once jitter exceeds a quarter of the cycle' (L317-319). The draft supplies its own bridge; the finding quotes around it. The secondary argument — that resources should go to precision rather than nights — also fails on a fact the finding did not check: I fetched the host protocol and it already prescribes 'Z-Max triple electrode EEG headbands ... four nights per timepoint' and EmbracePlus seven days per timepoint for the whole cohort. The four nights are not a resource allocation this project is making, so there is no trade-off to justify. Nothing survives.
Strongest case against: the finding names its own escape hatch — "This may all be inherited from the parent cohort at no marginal cost, in which case saying so removes the objection" — and I verified that it is. The Sleeping Through Menopause protocol specifies that hot flashes "will be estimated from a multisensory approach ... assessed with the EmbracePlus smartwatch for seven days at each timepoint." The seven days and nights are the parent study's design, already running; this project adds four EEG nights and a button press. So there is no unjustified participant burden and no undirected collection attributable to this proposal, and the charge that it "cuts against the proposal's own claim to a focused design" collapses. The diary claim is also inaccurate: L218-219 gives it a stated role ("a complementary measure of recall and reporting behaviour"), echoed at L68-69. Residual: one clause disclosing that the seven nights are inherited would pre-empt the objection entirely, and the suggested cheap uses (habitual flash rate from non-EEG nights, diary as a bound on press under-reporting) are good ideas. Nice-to-have, not a defect.
Strongest case against, and it is decisive on the quantitative core. The whole arithmetic requires the classifier to score every 15-s frame of an 8-hour night ("an 8-h night is 1,920 frames"), and that premise is contradicted by the draft and unestablished in the source. The draft says "Each candidate event is classified as a hot flash or not" (L241) - a candidate-gated pipeline, not a whole-night sweep. And Tsiartas's own architecture presupposes upstream candidate generation: I read L~112-118 of Tsiartas2021.txt, where "SC feature set uses the HF onset output of a previously developed HF prediction algorithm [4], as a feature. We designate the HF onset as the +-2 minutes around the HF predicted onset." So candidates come from a prior sternal-SC threshold detector whose rate is far below one per frame, and the false-positive burden F91 computes is an unbounded over-estimate. The paper also never states whether the reported specificity is frame-level or region-level, so "the 95.6% specificity is a per-15-s-frame figure" is an inference presented as fact. The headline claim - "most 'detected flashes' entering the primary model would be spurious" - is therefore not supported, and the power table built on it is not usable. Everything sound in this finding (small N=3 base, non-subject-independent validation, unknown false-positive burden) is already carried by F19 and F73 with verified evidence. Reporting it separately would put a fabricated-looking event-level PPV of 0.04 in front of the applicants.
AGAINST: the technical core rests on reading "pre-flash ISF amplitude" as the *signal level* averaged in a pre-onset window, then simulating sinc attenuation of the first harmonic. That is not what amplitude means in this pipeline. The draft obtains phase from "the Hilbert transform" (L268-269); the companion quantity from the same analytic signal is the instantaneous *envelope magnitude*, which by construction varies on the timescale of the modulation envelope rather than within a cycle. An envelope is genuinely far more tolerant of onset jitter than a phase — jitter of 12.5 s moves phase by π/2 but moves the envelope hardly at all. So the draft's claim that amplitude "tolerates far more timing error" is correct under the standard reading, and the /root/grantreview/sim/06_filter_edge.py sinc numbers answer a different question. Branch (b) of the finding — that a bout-level modulation strength tests a different hypothesis — is valid, but that is F116's point, not this one. As a statistics finding about tolerance, this does not survive. Minor.
AGAINST: the premise is not in the draft. The finding claims a non-significant surrogate test "will be treated as licence for the primary analysis" — but the draft never makes the prerequisite a gate. L294-297: occurrence is "characterised first" and "phase-dependent occurrence would be a separate finding about flash expression, not conscious gating", with the primary model fitted "[c]onditional on a flash occurring" either way. Characterised, reported, not used as a permission slip. The finding invents the inferential move it then criticises. Its own numbers also cut the other way: it computes that if occurrence *is* phase-structured the primary test loses only modest efficiency (0.96 at ρ=0.2, 0.91 at 0.3), which is a point in the draft's favour. The residual true statement — that a Rayleigh test on a few hundred events detects only fairly strong clustering, and that jitter attenuates ρ by the same λ — is a limitation the draft does not claim otherwise about. Minor.
The arithmetic is correct and I reproduced it (40/65 = 61.5%, Wilson 95% CI 49.4-72.4% matching the draft's '49-72', 0.615 x 222 = 136.6). The finding is nonetheless not a defect. Quoting a point projection for a target sample size is universal practice in grant applications, and the draft states the interval in the same clause the point estimate appears in (L203), so any reviewer who wants the range has it in front of them. The projection is also interim by construction — 65 of 222 screened, with recruitment continuing — so the range will narrow without any action by the applicants. And the finding's own simulated yield spread (109 -> 71 events, 137 -> 347, 160 -> 1,249) is not attributable to the CI, as it admits: 'those cells also vary other assumptions'. A 1.5-fold swing in women producing an 18-fold swing in events is evidence about the exclusion cascade, not about CI propagation. Whatever force this has is entirely absorbed by F65, which asks for the yield range directly and asks for it against the right quantity. Too trivial to report separately.
AGAINST: the arithmetic is right (0.01 Hz → 100 s; three cycles at 0.02 Hz → 150 s) and the filter physics is right, but the scenario is assumed, not stated. The draft restricts *analyses* to stable NREM (L270); it does not say the band-pass is applied bout-by-bout. Filtering the continuous log sigma envelope over whole NREM episodes — tens of minutes, hundreds of ISF cycles — and then admitting only events that sit inside a stable bout has no transient problem and is the ordinary way this is done. The same disposes of the frequency-resolution point: a per-participant empirical peak is estimated over NREM episodes or the whole night, not inside a 150-s window, and Lecci's own human analysis used ≥120 s segments with 0.001 Hz wavelet resolution over far longer records. So the "filter transient exceeds the whole segment" conclusion follows only from an implementation the draft never proposes, and the sibling finding F80 concedes the ambiguity that this one treats as settled. Not fatal, not confirmed; worth one clarifying sentence from the applicants. Minor.
Every quoted fact is verbatim correct in Tsiartas2021.txt — "We time-aligned the features with the HF expert annotations for prediction and evaluation (±90 s matching window)"; "a Decision Tree classifier, which makes a decision every 15 s"; "the AUC difference between the last 250 s and the current window (±30 s)"; "the average differential between the prior and following 500 s"; "averaged in two regions: 120 s before and after the window"; "2 uS/30s rises in SC". The inference is nevertheless wrong, and it is wrong about the one thing the finding calls fatal. The draft does not take onset from the classifier. L259-263: "Onset is back-dated by change-point detection to the inflection at which the electrodermal rise begins, not to a threshold crossing. Intervals between this inflection and the accompanying temperature and pulse-rate changes index achievable precision." Detection answers "is this a flash?"; timing comes from a separate change-point analysis of the raw EDA trace. ±90 s is a validation MATCHING TOLERANCE, not a timing resolution, and a coarse 15 s region boundary does not bound the precision of a change point located inside that region. The finding dismisses this in one clause ("cannot recover timing the detector never had") without engaging it, and it ignores that the draft measures its own onset precision empirically (L261-263) and makes insufficient precision the pre-registered trigger for a design switch (L338-343). What remains is a genuine but empirical risk the draft already treats as a feasibility outcome. Minor.
THE DEFENCE HOLDS ON EVERY PRONG. (1) The claimed internal contradiction is not one: L52-54 ("Every detected flash enters both analyses, including the quiet ones passing without arousal or report") is a statement that quiet flashes contribute to the DENOMINATOR — which is precisely the design's real advantage over stimulus studies — and is fully compatible with L299-301 ("registered flashes are expected to be largely a subset of aroused ones"). Both sentences are true simultaneously. (2) "No contrast in the design identifies conscious access independently of cortical arousal" is false: registration among aroused flashes is exactly that contrast, it is specified in the draft (L301-302), and it behaves correctly — under an arousal-only truth it rejects at 0.046-0.061 against α=0.05, i.e. it does not spuriously fire. It is underpowered, not unidentified, and "underpowered" is a different (and lesser) finding than "unmeasurable by construction". (3) The draft states the press-requires-arousal limitation itself at L299-300 and responds to it with the joint distribution plus the aroused-only model, so this is a pre-empted objection. (4) The base-rate quote is accurate (GombertLabedens2025.txt: 3.5/night, 70% aroused, 20% undisturbed) but supports only the near-nesting the draft already concedes. What is left is the priority-of-inference point, which F64 makes properly and which I confirmed there at moderate. Calling this fatal on top of F64 double-counts one issue at the maximum severity twice.
THE DEFENCE HOLDS, AND THE FINDING HAS THE DIRECTION OF THE ERROR BACKWARDS. The algebra is correct — Var(M_i − M_j) is independent of the common offset E[d], so inter-marker intervals carry no information about bias — and the ordering claim is verified in GombertLabedens2025.txt ("an increase in heart rate that is detectable before the onset of sweating"). But the draft does not need accuracy; it needs PRECISION, and it says exactly that: "Intervals between this inflection and the accompanying temperature and pulse-rate changes index achievable precision" (L261-263). Its confirmatory test is a 2-df sin/cos likelihood-ratio test that deliberately does not fix the fragility-continuity boundary in advance (L285-287), so a CONSTANT onset bias merely rotates the estimated preferred phase and leaves the test statistic and its power untouched — only dispersion attenuates the effect. The draft's risk rule is therefore triggered by the correct quantity. Worse for the finding: because inter-marker dispersion sums true onset jitter AND the effector-ordering variance the finding itself identifies, it OVERSTATES onset jitter — so as a pre-registered veto threshold it can only be over-conservative, never permissive. "A proxy that cannot measure accuracy is being used to license or veto the confirmatory analysis" gets the safety direction exactly wrong. Residual worth one line, not a major finding: a systematic lag of ~12.5 s would rotate the preferred phase by 90° and could flip the fragility/continuity interpretation in Expected outcomes (L355-361), and the draft never estimates a systematic lag.
Every quote verified in GombertLabedens2025.txt — 'typically lasting between one and five minutes' (L757-759), the 2 microSiemens/30 s rule (L888-893), 'only 51%' preceded by a GI temperature rise 'on average, 0.03 C' (L803-814), 'there is no increase in Tcore ... in the 2-min period before a hot flash' (L791-792). The finding's conclusion does not follow from them, and two of them argue against it. If there is no Tcore rise in the two minutes before a flash, and only half of flashes show any prodromal temperature change at all averaging 0.03 C, then the effector onset is comparatively punctate, not diffuse — the evidence supports a defined onset. More decisively, the draft defines the estimand operationally and the finding does not engage with it: L259-261 sets onset at 'the inflection at which the electrodermal rise begins, not to a threshold crossing', located by change-point detection. That is a well-defined instant. Event duration is irrelevant to phase-at-onset — a tone has duration too, and the exogenous literature the draft extends estimates phase at stimulus onset, not over stimulus extent. The one real residual, whether the peripheral inflection is a fixed lag from the central event, is F89's point and is scored there. What is left is that the draft could argue the estimand's existence explicitly rather than assuming it.
THE DEFENCE HOLDS. The draft's position is coherent and standard, and the finding misreads which model is which. The PRIMARY model does not condition on arousal: it predicts registration among ALL detected flashes with no arousal term and no magnitude term (L280-283), which is the total effect of phase on registration — exactly what the mediator argument at L306-309 says it should be. The aroused-only model is an ADDITIONAL analysis ("registration additionally modelled among aroused flashes", L301-302) and the draft explicitly ring-fences it: "The primary registration model is confirmatory; all else is exploratory" (L310-311). Estimating a total effect as the confirmatory quantity while reporting a mediator-conditional decomposition as a labelled secondary analysis is textbook mediation practice, not a double standard. The finding's "either mediator conditioning is forbidden or the magnitude argument collapses" is a false dilemma: mediator conditioning is forbidden FOR THE CONFIRMATORY TOTAL EFFECT and permitted for a clearly-labelled exploratory decomposition, which is what the draft does with both magnitude (sensitivity analysis, L309-310) and arousal (additional model, L301-302). Symmetric treatment, not asymmetric. Residual worth one clause, not a major finding: the draft flags the mediator caveat for magnitude but not for the aroused-only model, so it should say the same thing twice.
AGAINST: both limbs fail. (1) The inclusion criterion is a person-level screen — women "whose nocturnal flashes interfere with sleep" — not an event-level filter, so quiet flashes within those women are retained. The draft's own cited base rates come from exactly such samples: GombertLabedens2025.txt reports "only a minority (20%) occurring without disturbance to sleep" and "29% of hot flashes occurred in undisturbed sleep", and the draft leans on the 1.5-of-3.5 reporting gap. So the contrast the model needs is present by the finding's own sources; the variance is not truncated away. (2) Universal ISI ≥10 is a constant of the sample, so it cannot confound a within-subject phase contrast — participant random intercepts absorb level differences, and sleep-state misperception would have to be *phase-dependent* to bias the primary coefficient, which is asserted nowhere. What survives is that the draft never discloses the insomnia inclusion criterion — a generalisability and candour point already carried by F2 and F56. As an eligibility/statistics finding in its own right it does not stand. Minor.
The finding's substantive prong is factually wrong about the draft. It claims the draft quotes the reporting gap 'while omitting the arousal rate', making the quiet-flash pool 'look larger than the source says it is'. The draft does the opposite: L149-151 states '19.8% of objectively detected nocturnal flashes occur without disturbing sleep (De Zambotti et al., 2014)' — the quiet pool given directly, and given more conservatively than the review's rounded 20%, and consistent with the complement of the arousal rate I verified independently from the Europe PMC record ('69.4% of hot flashes were associated with an awakening'). So the draft discloses exactly the number the finding says it suppresses, in the same paragraph. The remaining prong is an audit note rather than a finding: the 1.5 self-report figure is UNVERIFIED, and I could not reach the primary paper either (Europe PMC full-text 429 on four attempts, PMC reCAPTCHA, ScienceDirect robots.txt). Against that, the draft's use of the precise, non-obvious 19.8% figure — which does not appear in the review and must have come from the primary paper — is decent circumstantial evidence the applicants read the source, and hence that 1.5 is from it too. Nothing here would change a reviewer's decision.
The quote is verbatim right and the reading of it is wrong. I retrieved the Bianchi abstract: "AFTER STANDARDIZING TO SLEEP-STAGE TIME DISTRIBUTION, the majority of HF were recorded during wake (51.0%) and stage N1 (18.8%)." Standardising to stage-time distribution converts counts to per-unit-time RATES; 51.0% is therefore the share of the rate, not the share of events. Wake occupies a small fraction of the night, so the raw count of flashes in N2/N3 can be far larger than 30%. The finding's entire yield arithmetic ("~30% in N2/N3 gives ~4") treats a standardised rate as a raw proportion. Three further problems. The population is not the draft's: 28 healthy PREMENOPAUSAL volunteers given depot leuprolide, "a model of induced HF", two lab PSGs each — not natural perimenopause. The same abstract says "80% occurring just before or during the awakening", i.e. the flash ONSET typically precedes the awakening, so onsets are in NREM even when the epoch scores as wake — and onset phase is what the draft analyses. And the draft's own primary source contradicts the finding: de Zambotti 2014 (verified) gives 3.5 objective flashes per night with 69.4% associated with an awakening and ~20% in undisturbed sleep, i.e. ~90% during sleep. Minor residue: the stable-NREM cell is genuinely unquantified, but that is F165's finding, evidenced properly.
Strongest case against, and it is decisive. Türker et al. 2023 is real (https://www.nature.com/articles/s41593-023-01449-7, verified), but its paradigm is not transferable to this study. The facial responses are *categorisation responses to experimenter-delivered verbal stimuli* - participants "had to react by smiling or frowning to categorize them" - so the response is elicited and time-locked to a stimulus the experimenter controls. That is precisely the exogenous design this proposal exists to escape. Nothing in that literature establishes that a volitional covert facial response to a *spontaneous endogenous* event can be detected, or discriminated from spontaneous facial activity and arousal-related EMG, without a stimulus-locked window. It also requires facial EMG, which the ZMax (two frontal EEG derivations, accelerometer, PPG) does not provide, in an at-home ambulatory study. The serial-awakening alternative destroys the sleep continuity that is the object of study and yields sampled, not event-locked, reports. Meanwhile the mechanistic concern the finding raises - that the press requires arousal and collapses the two outcomes - is explicitly conceded and modelled by the draft at L299-302, which the finding itself acknowledges. Worth a sentence of related work at most; graded major it is a design preference presented as a defect.
Strongest case against, and it holds on three grounds. First, the headline evidence is species-shifted. I fetched the Bergel 2026 abstract (https://www.nature.com/articles/s41593-025-02159-y) and the coupling clause is explicitly restricted: "This rhythm is tightly coupled with eye movements, muscle tone, heart and breathing rate *in lizards*, with skin brightness in chameleons and with pulsatile changes in cerebrovascular volume... in bearded dragons and during non-rapid eye movement sleep in mice." The finding quotes it with the species qualifier removed, which is the same class of error it is auditing the draft for. Second, the timescales do not meet: respiration is ~0.2 Hz and respiratory modulation of spindle timing operates at breath scale, an order of magnitude above the 0.02 Hz predictor; the finding never specifies how a breath-scale rhythm produces a 0.02 Hz confound. Third, the general form of the objection - that a shared brain-autonomic rhythm could co-modulate sigma and flash occurrence - is explicitly pre-empted by the draft at L162-169, which names it as "one alternative account" and makes phase-structured expression "a separate testable outcome". The finding also concedes one of its three citations is title-level only and unverified. Minor at most.
AGAINST: the finding opens by verifying the citation is accurate, so there is no defect in the draft's use of Cataldi. What remains is a claim that the draft fails to stake out distinct territory — and that claim is false on the text. The differentiator is stated three times: "All of it rests on designs that randomise stimulus timing against brain state, precisely what an endogenous event does not do" (L31-32); "This literature rests entirely on stimuli delivered by an experimenter" (L115-117); "State-dependent gating has so far been shown only for stimuli an experimenter chose to deliver. Whether the same holds for signals the body generates itself is a question about interoception and conscious access, not about hearing" (L363-366). The draft also says outright that the NREM extension is "untested" (L124-125). The rest is speculation about a rival group's next paper and institutional adjacency. Proposals are not expected to name competitors, and no verified fact about the draft is at issue. The unverified volume/pages note is a reference-formatting matter, not this finding. Nothing to report.
Strongest case against, and it is conclusive: the draft does not make the claim being attacked. L30-32 reads "All of it rests on designs that randomise stimulus timing against brain state, precisely what an endogenous event does not do" — a statement about the *prior literature's designs*, not an assertion that flash timing is independent of brain state. The draft's position is the opposite of what the finding imputes: flash timing may well be state-structured, which is exactly why there is a prerequisite aim to test it (L181-182), why phase-structured occurrence is called "a separate testable outcome" rather than a nuisance (L165-166), and why occurrence and consequence are analysed separately (L294-297). The finding concedes the pre-emption in its own text — "That does not break the design — the prerequisite aim already tests occurrence phase" — which by the review's own standard makes it refuted. Its key evidence is also self-declared UNVERIFIED (Naghavi full text 403-blocked, with the accuracy figure and whether nocturnal flashes were included both unknown), leaving a press release as the sole support. The residual shared-antecedents point is real but is the selection concern already covered by F66.
Strongest case against, and the finding half-concedes it in its own first line ("Not an error - an omission"). I verified the paper exists and is as described (https://www.nature.com/articles/s41593-025-02159-y: "By recording brain activity in seven lizard species, humans, rats and pigeons, we demonstrate an infraslow brain rhythm during sleep in all species"), and the multiband characterisation too - the Extended Data show simultaneous fluctuation across delta, theta, alpha, sigma, beta and low gamma rather than a narrowband oscillation. So the facts are right. But the premise of the finding is not: the draft never mounts a defence of the ISF against a spindle-clustering-artefact critique, because it never needs to - the ISF's reality is not in dispute in the literature it cites, and the draft's predictor is the sigma envelope regardless of whether the underlying modulation is narrowband or broadband. "A relevant 2026 paper is uncited" is not a defect in a proposal, and the caveat it carries (report phase from more than sigma) is a nice-to-have that would weaken rather than strengthen a pre-registered single primary predictor. Nothing here changes a reviewer's decision.
THE DEFENCE HOLDS ON THE MAIN PRONG. The claim is evidenced in-document by exactly the mechanism the Bial form provides: Osorio-Forero et al. 2021 AND the 2025 Nature Neuroscience paper are both listed under PREVIOUS OWN PUBLICATIONS (L456-464), which is how an applicant substantiates "our team includes the researchers who characterised the infraslow noradrenergic mechanism in rodents" without naming individuals in a prose section. CVs are a separate mandatory form field (max 4 pages per team member per the call), so "no person is named anywhere in the document" describes the document's scope, not an omission — this file contains Summary, Literature Review, Research Plan, References and Own Publications, and no budget, host entity or CV section either. I also confirmed the underlying relationship is real: Osorio-Forero et al. 2025 is Nature Neuroscience 28:84-96, and van Someren is acknowledged in that paper for discussion and for reading earlier manuscript versions. The second prong also weakens on inspection: the van Baarzel protocol listed as an own publication is an Amsterdam UMC MenoPause Consortium trial with gynaecology and endocrinology co-authors, which covers the menopause-clinician competence, and the van Someren group covers human ambulatory sleep EEG. Residual worth one line, not a major finding: nothing evidences wearable-signal engineering competence to re-derive and validate the flash detector on EmbracePlus — which is genuinely the study's weakest link (see F160).
THE DEFENCE HOLDS ON THREE OF FOUR SUB-POINTS, AND THE FOURTH IS A DUPLICATE. I fetched the Bial call page and confirmed that "estimated number of publications and scientific dissemination activities" is collected as a distinct required application field ("Expected output indicators"), alongside CVs, the payment plan and the host-entity declaration. So (1) deliverables and (2) dissemination live in form fields that are simply not part of this document — whose sections are Summary, Literature Review, Research Plan and Methods, References and Previous Own Publications. Their absence here is scope, not omission, and the acknowledgement format is a reporting obligation that arises after award, not an application item. (3) Neither the regulation nor the call requires a data-management plan; the draft does state pseudonymisation at collection and institutional-server storage (L332-334), and the sub-study inherits its GDPR basis and retention regime from an approved parent protocol. Wanting more is a preference, not a defect. (4) Sample-size justification is real but is F28's finding verbatim, at higher severity and with better evidence. A finding built from four sub-points, three of which rest on mistaking a form field for a missing section and one of which duplicates another finding, should not be reported.
THE DEFENCE HOLDS. This asserts as the draft's position the very reading that F94 correctly identifies as merely one of two possible readings — and it is the less natural one. "Analyses are restricted to stable (and non-artifactual) NREM" (L270) sits inside the ISF EXTRACTION subsection, whose subject is estimating instantaneous phase; the stable-bout requirement exists so that phase at onset is estimable, which is a pre-onset property. The contingency "relaxing the stable-bout criterion from three ISF cycles to two" reads naturally as being about the bout containing the onset. Nothing in the draft requires the bout to survive PAST onset, so "a flash followed by a cortical arousal ... breaks stability and is excluded" is the finding's own construction, not the draft's rule. The claimed "logical contradiction" — "an inclusion criterion that requires uninterrupted NREM across the event window cannot coexist with an outcome that requires waking" — is only a contradiction under the premise being asserted, so it is circular. The de Zambotti 69.4% figure is genuine (verified from the abstract via Europe PMC) but it establishes only the stakes, not the reading. Report F94, which states the same issue honestly as an unresolved ambiguity, and drop this one; asserting the damaging reading as fact and grading it major is exactly the inflation this pass exists to catch.
THE DEFENCE HOLDS: the draft's move is deliberate, stated, and is the thesis rather than a conflation. L152-154 says explicitly "This discrepancy has been treated as a measurement problem rather than as evidence about how endogenous signals gain access to awareness" — the whole point is to stop treating the 3.5-vs-1.5 gap as a recall artefact and instead ask whether access is gated at the moment the event arrives. Measuring that prospectively rather than by recall is the correct instrument choice for that reframing, not a mismatch, and the draft explicitly retains the diary as a complementary recall measure (L218-219) rather than abandoning the recall construct. If in-the-moment access is phase-gated, unregistered events cannot be recalled, so the closing claim at L370-373 follows. The reactivity objection (pressing rehearses the event) is real in kind but second-order and applies to any prospective in-night report. On the 1.5 figure: I confirmed it is absent from the de Zambotti 2014 abstract (Europe PMC gives 34 women, 63 nights, 222 flashes, 3.5/night, 69.4% awakenings), but the abstract also omits the 19.8% figure that GombertLabedens2025.txt independently corroborates as ~20% from the same study — so absence from an abstract is no evidence against a full-text figure, and "UNVERIFIED" is not a defect. Residual worth one clause: state whether the 1.5 is a morning estimate. Not a major eligibility finding.
AGAINST: the draft makes the very distinction the finding says it conflates, and makes it twice. L156-161: "these literatures jointly establish that ... hot flash expression is gated by sleep state at the ultradian scale, and that many nocturnal flashes produce neither awakening nor report. Whether the conscious registration of an endogenous interoceptive event depends on the infraslow phase at which it arrives is untested." And L296-297: "phase-dependent occurrence would be a separate finding about flash expression, not conscious gating." That is verbatim the "honest statement" the finding demands. The sentence it objects to (L147-149) claims only that "state-dependent gating of an endogenous thermoregulatory event" is established at the ultradian scale — gating of the event, correctly scoped, immediately after the draft itself supplies the effector mechanism at L42-44. So the finding is refuted by a pre-emption it did not read forward to. The GombertLabedens observation that REM flashes may simply be less intense ("less likely to be associated with an awakening, although this hypothesis remains to be tested") is a genuine alternative to the effector account, but it does not create the conflation charged here. Minor.
The statistical premise is wrong and it contradicts a sibling finding. A binary outcome with base rate 0.78 has Bernoulli variance 0.172 — 69% of the theoretical maximum 0.25, against 0.245 for a 0.43 base rate. That is not "compressed information" in any material sense, and logistic regression is untroubled by it. F93, whose analytics I checked by hand ((1-p)/(1-qp) ~ 0.34 at p=0.694), shows the opposite of this finding: the arousal test is the MORE informative of the two, by roughly an order of magnitude, precisely because its effect on the logit scale is undiluted. Both cannot be right, and F93 is the one with a derivation. The "few informative cells" claim also fails on arithmetic: at ~78% aroused and ~43% registered, aroused-but-unregistered is ~35% and not-aroused ~22% — both large. And the draft does not claim quiet flashes are a majority; L52-54 claims only that they exist and cannot be obtained by stimulus delivery, which is true, and GombertLabedens2025.txt reports a larger study (n=86) with "29% of hot flashes occurred in undisturbed sleep", so the draft is conservative. The one accurate observation — that L315-320 states no base rate and no cell counts — is verified but is F87's and F165's finding, evidenced there with the exclusion cascade. Minor, and duplicative.
Duplicative on its substantive content and wrong on its distinctive content. The Host Entity, schedule and budget absences are verified (grep: zero 'host'; no budget or schedule section; L332 'reference \[number\]') but they are F5's finding, reported there with the team-composition argument that makes them matter. This entry's own two criticisms both fail. First, 'Research aims' versus 'Specific Aims' is a form-field label, not a document defect: the BF-GMS online form supplies the field, the draft supplies the prose that goes in it, and the regulation's language quoted from BIAL_CONTEXT.md concerns the objectives, not a heading string. Second, the structural complaint misdescribes the text. L177-182 reads 'The primary aim tests whether the phase of the infraslow sigma fluctuation (ISF) at nocturnal hot flash onset predicts conscious registration; a supporting aim asks the same of cortical arousal ... A prerequisite aim establishes whether onsets occur across the whole cycle' — the primary aim opens the second sentence and the three aims are ordered by priority, which is a sensible arrangement, not an aim 'embedded mid-paragraph'. Nothing here would change a reviewer's decision that F5 does not already cover.