fullseye

attention_softmax — LLMCORE attend op

使い方

一括の注意(参照実装)→ tokens。この族の答え合わせの基準。

(T, T) のスコア行列を実際に作ってから softmax する、いちばん素直な書き方。 attention_tiled / attention_linear / attention_grouped はこれと 一致することで正しさを示す。

Args: query: (T, d) float。 key: (S, d) float。 value: (S, dv) float。 mask: None / "causal" / "window" / (T, S) の bool 配列。 window: mask="window" のときの窓幅。 scale: スコアの倍率。None なら 1/√d。 Returns: tokens (T, dv) float: 注意で混ぜた列。

詳しい使い方ガイド

参考(サンプルデータ・文献)

実行できる例(この op を実際に呼ぶ検証済みサンプル)

型が繋がる次の op(tokens を入力に取れる)

rms_norm · rope_rotate · attention_scores · attention_apply · attention_tiled · attention_linear · attention_grouped · kv_cache_decode

同カテゴリ(attend)

attention_tiled · attention_linear · attention_grouped


Provenance: llmcore.py — LLMCORE operator registry. この per-op ノートは tools/opdocs.py md が自動生成(手編集しない)。

© 2026 Kazufumi Furuse — Fullseye operator documentation. Licensed under Apache-2.0.