Skip to content

Commit c6b2c11

Browse files
nonmetalclaude
andcommitted
Update paper metadata: new title, Spellbrush-only authors, arXiv-preprint citation
Match the arXiv submission: title changed to "When Does Predictor-Based RL Align with Human Perception?", author line added (Park, Li — Spellbrush), EMNLP-2026 inproceedings citation replaced with an arXiv @misc entry pending the assigned identifier. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MUq8H7Bue5vjhcuKWSgZNs
1 parent 6c92dca commit c6b2c11

2 files changed

Lines changed: 21 additions & 16 deletions

File tree

README.md

Lines changed: 11 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -2,8 +2,9 @@
22

33
**Constrained perceptual GRPO for codec-based speech language models.**
44

5-
Companion code for *"When Do Perceptual Rewards Align Speech Language Models?
6-
Constrained GRPO for Subjective Style Control"* (EMNLP 2026).
5+
Companion code for *"When Does Predictor-Based RL Align with Human Perception?
6+
A Study of Subjective Rewards in Codec-Based Speech Language Models"*
7+
(Joonyong Park and Jerry Li, Spellbrush — arXiv preprint, 2026).
78

89
🔊 **[Audio demo →](https://sizigi.github.io/animeGRPO/)**
910

@@ -183,14 +184,17 @@ zone gate — the ablation in the paper is exactly "swap the axis, keep the gate
183184
## Citation
184185

185186
```bibtex
186-
@inproceedings{animegrpo2026,
187-
title = {When Do Perceptual Rewards Align Speech Language Models?
188-
Constrained {GRPO} for Subjective Style Control},
189-
booktitle = {Proceedings of EMNLP 2026},
190-
year = {2026}
187+
@misc{park2026predictorrl,
188+
title = {When Does Predictor-Based {RL} Align with Human Perception?
189+
A Study of Subjective Rewards in Codec-Based Speech Language Models},
190+
author = {Park, Joonyong and Li, Jerry},
191+
year = {2026},
192+
note = {arXiv preprint}
191193
}
192194
```
193195

196+
The arXiv identifier will be added to the entry once assigned.
197+
194198
## License
195199

196200
Code is Apache-2.0 ([LICENSE](LICENSE)), matching the verl fork it extends.

docs/index.html

Lines changed: 10 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -173,11 +173,11 @@
173173
<header>
174174
<div class="wrap">
175175
<p class="eyebrow">animeGRPO</p>
176-
<h1>When Do Perceptual Rewards Align Speech Language Models?
177-
<span class="sub">Constrained GRPO for subjective style control in codec speech LMs</span>
176+
<h1>When Does Predictor-Based RL Align with Human Perception?
177+
<span class="sub">A study of subjective rewards in codec-based speech language models</span>
178178
</h1>
179-
<!-- Author list intentionally left out until the paper is public. -->
180-
<p class="venue">EMNLP 2026 &nbsp;·&nbsp; LLaSA-1B-Multilingual &nbsp;·&nbsp; Japanese &amp; English</p>
179+
<p class="venue">Joonyong Park &nbsp;·&nbsp; Jerry Li &nbsp;&mdash;&nbsp; Spellbrush</p>
180+
<p class="venue" style="margin-top:6px">arXiv preprint &nbsp;·&nbsp; LLaSA-1B-Multilingual &nbsp;·&nbsp; Japanese &amp; English</p>
181181
<div class="links">
182182
<a class="btn" href="https://github.com/sizigi/animeGRPO">
183183
<svg viewBox="0 0 16 16"><path d="M8 0C3.58 0 0 3.58 0 8c0 3.54 2.29 6.53 5.47 7.59.4.07.55-.17.55-.38 0-.19-.01-.82-.01-1.49-2.01.37-2.53-.49-2.69-.94-.09-.23-.48-.94-.82-1.13-.28-.15-.68-.52-.01-.53.63-.01 1.08.58 1.23.82.72 1.21 1.87.87 2.33.66.07-.52.28-.87.51-1.07-1.78-.2-3.64-.89-3.64-3.95 0-.87.31-1.59.82-2.15-.08-.2-.36-1.02.08-2.12 0 0 .67-.21 2.2.82a7.4 7.4 0 0 1 2-.27c.68 0 1.36.09 2 .27 1.53-1.04 2.2-.82 2.2-.82.44 1.1.16 1.92.08 2.12.51.56.82 1.27.82 2.15 0 3.07-1.87 3.75-3.65 3.95.29.25.54.73.54 1.48 0 1.07-.01 1.93-.01 2.2 0 .21.15.46.55.38A8.01 8.01 0 0 0 16 8c0-4.42-3.58-8-8-8z"/></svg>
@@ -338,11 +338,12 @@ <h2>Listen</h2>
338338

339339
<footer class="wrap">
340340
<h2 style="margin-bottom:16px">Citation</h2>
341-
<pre class="cite">@inproceedings{animegrpo2026,
342-
title = {When Do Perceptual Rewards Align Speech Language Models?
343-
Constrained {GRPO} for Subjective Style Control},
344-
booktitle = {Proceedings of EMNLP 2026},
345-
year = {2026}
341+
<pre class="cite">@misc{park2026predictorrl,
342+
title = {When Does Predictor-Based {RL} Align with Human Perception?
343+
A Study of Subjective Rewards in Codec-Based Speech Language Models},
344+
author = {Park, Joonyong and Li, Jerry},
345+
year = {2026},
346+
note = {arXiv preprint}
346347
}</pre>
347348
<p>Code is Apache-2.0. Audio and score tables are released for research
348349
inspection and replication of the paper's results.</p>

0 commit comments

Comments
 (0)