# v0.5.4 Release Notes

> v0.5.3 -> v0.5.4

## Summary

Gives the router a voice layer with speech-to-text and text-to-speech agents implemented for OpenAI and Gemini behind a shared audio contract. Model discovery gains matching capability filters so callers can list what each direction actually supports.

<details>
<summary>翻譯</summary>

為 router 補上語音層，以共用的音訊契約提供語音轉文字與文字轉語音代理，並在 OpenAI 與 Gemini 上實作。模型探索同步新增對應的能力過濾，讓呼叫端能列出各方向實際支援的模型。

</details>

## Changes

### FEAT
- Add `STTAgent` and `TTSAgent` with OpenAI and Gemini implementations returning WAV audio (@pardnchiu) [7bf7640]
- Add `STTOnly` and `TTSOnly` model filters with speech capability detection (@pardnchiu) [7bf7640]

<details>
<summary>翻譯</summary>

- 新增 `STTAgent` 與 `TTSAgent`，並實作 OpenAI 與 Gemini 版本，統一輸出 WAV 音訊
- 新增 `STTOnly` 與 `TTSOnly` 模型過濾選項及語音能力判別

</details>

## Scope

- `core/` — FEAT (`audio.go`, `reasoning.go`)
- `core/openai/` — FEAT (`audio.go`, `models.go`)
- `core/gemini/` — FEAT (`audio.go`, `models.go`)

***

Generated by [SKILL](https://github.com/pardnchiu/skill-version-generate)
