|
173 | 173 | <header> |
174 | 174 | <div class="wrap"> |
175 | 175 | <p class="eyebrow">animeGRPO</p> |
176 | | - <h1>When Do Perceptual Rewards Align Speech Language Models? |
177 | | - <span class="sub">Constrained GRPO for subjective style control in codec speech LMs</span> |
| 176 | + <h1>When Does Predictor-Based RL Align with Human Perception? |
| 177 | + <span class="sub">A study of subjective rewards in codec-based speech language models</span> |
178 | 178 | </h1> |
179 | | - <!-- Author list intentionally left out until the paper is public. --> |
180 | | - <p class="venue">EMNLP 2026 · LLaSA-1B-Multilingual · Japanese & English</p> |
| 179 | + <p class="venue">Joonyong Park · Jerry Li — Spellbrush</p> |
| 180 | + <p class="venue" style="margin-top:6px">arXiv preprint · LLaSA-1B-Multilingual · Japanese & English</p> |
181 | 181 | <div class="links"> |
182 | 182 | <a class="btn" href="https://github.com/sizigi/animeGRPO"> |
183 | 183 | <svg viewBox="0 0 16 16"><path d="M8 0C3.58 0 0 3.58 0 8c0 3.54 2.29 6.53 5.47 7.59.4.07.55-.17.55-.38 0-.19-.01-.82-.01-1.49-2.01.37-2.53-.49-2.69-.94-.09-.23-.48-.94-.82-1.13-.28-.15-.68-.52-.01-.53.63-.01 1.08.58 1.23.82.72 1.21 1.87.87 2.33.66.07-.52.28-.87.51-1.07-1.78-.2-3.64-.89-3.64-3.95 0-.87.31-1.59.82-2.15-.08-.2-.36-1.02.08-2.12 0 0 .67-.21 2.2.82a7.4 7.4 0 0 1 2-.27c.68 0 1.36.09 2 .27 1.53-1.04 2.2-.82 2.2-.82.44 1.1.16 1.92.08 2.12.51.56.82 1.27.82 2.15 0 3.07-1.87 3.75-3.65 3.95.29.25.54.73.54 1.48 0 1.07-.01 1.93-.01 2.2 0 .21.15.46.55.38A8.01 8.01 0 0 0 16 8c0-4.42-3.58-8-8-8z"/></svg> |
@@ -338,11 +338,12 @@ <h2>Listen</h2> |
338 | 338 |
|
339 | 339 | <footer class="wrap"> |
340 | 340 | <h2 style="margin-bottom:16px">Citation</h2> |
341 | | - <pre class="cite">@inproceedings{animegrpo2026, |
342 | | - title = {When Do Perceptual Rewards Align Speech Language Models? |
343 | | - Constrained {GRPO} for Subjective Style Control}, |
344 | | - booktitle = {Proceedings of EMNLP 2026}, |
345 | | - year = {2026} |
| 341 | + <pre class="cite">@misc{park2026predictorrl, |
| 342 | + title = {When Does Predictor-Based {RL} Align with Human Perception? |
| 343 | + A Study of Subjective Rewards in Codec-Based Speech Language Models}, |
| 344 | + author = {Park, Joonyong and Li, Jerry}, |
| 345 | + year = {2026}, |
| 346 | + note = {arXiv preprint} |
346 | 347 | }</pre> |
347 | 348 | <p>Code is Apache-2.0. Audio and score tables are released for research |
348 | 349 | inspection and replication of the paper's results.</p> |
|
0 commit comments