Research

Publications

Peer-reviewed papers and preprints. For the most current list, see my Google Scholar.

  • 2026

    How Much of Speech Recognition Must Be Learned? A Parameter-Free Analysis of Lexical Decoding

    Po-Ting Lin

    Po-Ting Lin

    ResearchGate (preprint) Preprint

    End-to-end recognisers learn one function from acoustics to text, which makes it impossible to ask where the difficulty sits. Separating the acoustic–phonetic stage from the lexical one, we find that given correct phonemes and no word boundaries a decoder with zero trainable parameters recovers words at 5.66% WER on LibriSpeech dev-clean – within 1.34 points of the floor that out-of-vocabulary words and homophones impose on any decoder over this lexicon, every residual error falling into one of three interpretable causes. Sweeping phoneme accuracy under controlled corruption yields a calibration curve, WER = 9.3% + 2.11 × PER (R² = 0.997), validated to about a point by seven trained CTC heads: the relationship is linear, not amplifying. A frozen language model reordering phoneme-licensed candidates improves accuracy up to a broad optimum with no collapse at high weight, provided both scores are expressed on a common scale.

  • 2026

    Characterising Per-Pixel Rendering Difficulty by Transport Mechanism

    Po-Ting Lin

    Po-Ting Lin

    Computer Graphics Forum In Review

    Per-pixel rendering difficulty is conventionally characterised by one noisy scalar: the sample variance of a Monte Carlo estimator. We argue that it should instead be characterised through transport structure — a discrete description of how each contribution's energy reaches the sensor, deterministic under stated renderer conventions — with variance treated as a measurement whose reliability that structure helps predict. We make this concrete with a seven-class transport-mechanism descriptor assigned per contribution event, evaluate it on eleven scenes, and measure variance reliability on the seven first-party ones. The dominant label agrees 87–99.6% between 64 and 4096 samples per pixel, where quantile-binned variance agrees as little as 21%; its stability on unseen scenes is predicted from their pilots. Conditioning a pilot variance on the label improves equal-budget sample allocation wherever heavy-tailed buckets carry appreciable population, reduces to the incumbent where they are absent or negligible, and is neither reproduced by a random partition nor absorbed by a median-of-means estimator; where the pilot fails for other reasons, as on the classical ajar-door scene (6.8 dB below uniform), the label says so from the pilot alone.

  • 2026

    Structural Crystallization: A Unified Computational Framework for Memory Formation, Persistence, and Modification

    Po-Ting Lin

    Po-Ting Lin

    Research Square (preprint) Preprint

    Prevailing models of memory treat strength as a unitary quantity that increases with learning and decreases with forgetting or extinction. We argue that this conflation obscures a fundamental distinction: memory has separable dimensions of structural accumulation and representational fidelity, which are differentially modified by retrieval. Structural accumulation captures how much of a memory trace has been consolidated; representational fidelity captures how faithfully that trace preserves its original encoding. The two can change independently — a memory can retain its full structural extent while becoming progressively distorted, or dissolve while remaining perfectly faithful. This distinction, combined with a second distinction between structurally protective retrieval (externally guided, high-constraint) and structurally risky retrieval (internally generated, low-constraint), resolves several phenomena that have resisted unified explanation: why extinction is temporary but retrieval-extinction can produce lasting change, why stress-enhanced fear memories persist for orders of magnitude longer than ordinary memories, and why fear extinction is fragile while appetitive extinction is durable. We formalize these distinctions in a system of four coupled differential equations (the Structural Crystallization framework), calibrate it to two datasets, and show that it correctly predicts — without parameter adjustment — outcomes across five independent benchmarks where competing models (Rescorla–Wagner; latent cause) fail. Extensions to human declarative memory capture the testing effect crossover, while an explicit failure on the spacing effect reveals interpretable boundary conditions. The results demonstrate that separating accumulation from fidelity is not merely a modeling convenience but a theoretical necessity for any account of both the persistence and the modifiability of memory.

  • 2026

    Independent Samples, Correlated Variance A Learnable Cross-View Cue in Path-Traced Stereo Data

    Po-Ting Lin

    Po-Ting Lin

    arXiv (preprint) Preprint

    Path-traced synthetic stereo is a primary training substrate for disparity networks, and the pipelines that consume it assume Monte Carlo (MC) rendering noise is independent across the two views. The assumption is correct at the level it is stated—individual samples—but silent about the object a network actually sees. We show that the per-pixel MC variance fields, though built from independent samples, are strongly correlated once aligned by the ground-truth disparity, and that a network can learn to match with that correlation. Across 20 indoor scenes the warped correlation is 0.754 ± 0.016 against 0.360 unwarped; it replicates on a second renderer with a different sampler and sample budget (0.743 ± 0.044), and a seed-count analysis puts the population value near 0.85, making the measurement a lower bound. The effect is field-level, not sample-level: at the warp correspondence the signed per-seed residual correlation is -0.0004 while the residual envelope correlates at 0.38. A capacity-limited siamese probe given nothing but variance-field patches reaches 78.8% two-alternative forced-choice accuracy on held-out scenes and falls to chance once the alignment is destroyed; under the ordinary single-render condition, decorrelation still costs 2.40 percentage points in a difference-in-differences design, positive in all six held-out scenes. A real sensor's variance is fixed by its own signal rather than by transport difficulty, so across views it is redundant with intensity: the cue is specific to rendered data.

  • 2026

    CrystalASR: Hierarchical Phoneme-Grounded Speech Decoding

    Po-Ting Lin

    Po-Ting Lin

    Zenodo (preprint) Archived

    Archived by the author. A methodological flaw in this version undermines one of its central conclusions, so it should not be cited. It has been superseded by a later paper.

Curriculum Vitae

Choose a language

English PDF 中文 PDF