Welvet examples

12. layers/mha — attention

Open original example ↗Recorded results · not a live execution

Part: III · Layers
Package: github.com/openfluke/welvet/layers/mha
Status: ok — ✅

When

Building or training a net that needs the layers/mha Op (also usable as a Parallel cam).

Where

import "github.com/openfluke/welvet/layers/mha"

cd 12-mha && source ../env.sh && go run .

Why

Transformers need multi-head attention with masks, RoPE/ALiBi, GQA/MQA, and cross-attn — without forking MatVec for every projection.

What

Q/K/V/O are Dense children. Presets: DecoderCausal, EncoderBidirectional, CrossAttention, … Attn/RoPE ALU is host today; on-device attn shaders are still open.

Sample output (captured)

true 256 <nil>

Live capture from go run ./cmd/runall on local ../../welvet (exit 0).

Source

Copied from the Welvet feature book examples (openfluke.github.io/welvet/examples/12-mha).

main.go

Download source ↓
package main

import (
	"fmt"

	"github.com/openfluke/welvet/core"
	"github.com/openfluke/welvet/layers/mha"
)

func main() {
	l, err := mha.New(mha.DecoderCausal(32, 4, 4))
	if err != nil {
		panic(err)
	}
	x := core.NewTensor[float32](1, 8, 32) // [batch, seq, dim]
	pre, post, err := mha.Forward(l, x)
	fmt.Println(pre != nil, len(post.Data), err)
}