Direct Memory Access Controller (DMA)¶
The DMA controller moves data between memory and memory-mapped peripherals without CPU involvement. It is a TileLink bus master with a configurable number of channels, a descriptor-chaining engine and per-peripheral request lines for flow control.
Features¶
1 to 8 channels sharing one transfer engine (round-robin arbitration per chunk)
TileLink UH master port of 32, 64 or 128 bits with naturally aligned bursts up to
burstBytesMemory-to-memory, memory-to-peripheral and peripheral-to-memory transfers
Incrementing or fixed source and destination addresses
8, 16 and 32 bit element widths with automatic byte-lane steering
Up to 16 request handshakes; a gated channel moves one element per handshake
Descriptor chaining: a channel follows
nextpointers through memory-resident descriptorsPer-channel done and error interrupts, abort, bus error and alignment error reporting
No cache coherency logic: buffers are managed by software (see below)
Transfer Engine¶
A transfer is executed as a sequence of chunks. Each chunk is one Get on the source
followed by one PutFullData on the destination; data is staged in an internal FIFO of
burstBytes bytes.
Burst mode applies when both addresses increment and no request line is enabled. The chunk is the largest power of two that satisfies the alignment of both addresses, the remaining length and
burstLimit. Unaligned buffers or odd lengths are handled by falling back to smaller chunks down to single bytes.widthis ignored.Element mode applies when either address is fixed or a request line is enabled. Every chunk moves exactly one element of
widthbytes.src,dstandlengthmust be multiples of the element size, otherwise the channel stops with an error.
Between chunks the engine re-arbitrates, so several channels progress concurrently and an abort takes effect at the next chunk boundary.
Request Lines¶
Each request input is a four-phase handshake (DmaHandshake) driven by a peripheral:
req is its “ready” indication, e.g. TX FIFO not full or RX FIFO not empty, and
ack is returned by the DMA. A channel with req_enable set is only scheduled while its
selected req is high and ack is low, and moves one element per grant:
The peripheral raises
reqwhile it can move one element.The DMA moves one element and waits for the bus response.
The DMA raises
ack; the peripheral holdsreqlow whileackis high.The DMA drops
ackonce it seesreqlow.reqthen follows the peripheral state again, which already reflects the completed access.
Both signals hold their level until the other side responds, so the handshake tolerates any
latency between the peripheral and the DMA. DmaHandshakeCc synchronizes one handshake
between a peripheral clock domain and the DMA clock domain. Unconnected req and ack
inputs default to low. Channels without req_enable run whenever they are busy.
FIFO-backed IPs expose these handshakes as a DmaRequest bundle (tx: write-side FIFO
can accept an element, rx: read-side FIFO holds one): UART, SPI controller, I2C
controller, PIO (rx only), Mailbox (one bundle per channel) and the AES accelerator
(tx follows the plaintext FIFO). PeripheralsComponent.getDmaRequests lists them for
platform wiring.
Descriptors¶
A descriptor is five consecutive 32-bit words in memory (4-byte aligned):
Offset |
Word |
Description |
|---|---|---|
0x00 |
config |
Same layout as the channel |
0x04 |
src |
Source address. |
0x08 |
dst |
Destination address. |
0x0C |
length |
Transfer length in bytes. |
0x10 |
next |
Address of the next descriptor, 0 terminates the chain. |
When a descriptor completes and its linked bit is set with next != 0, the engine
loads the descriptor at next into the channel registers and continues. Starting a channel
with length = 0, linked = 1 and next pointing at the first descriptor runs a
complete chain from memory; the channel registers always mirror the descriptor currently
in flight. A ring is formed by pointing the last descriptor at the first; use abort to
leave it.
Cache Coherency¶
The controller accesses memory directly and does not participate in CPU cache coherency.
On platforms with a data cache, software must either place DMA buffers in an uncached
region (e.g. an uncached alias of the memory) or use the Zicbom cache management
instructions: cbo.clean the source buffer before starting a memory-to-peripheral
transfer and cbo.inval the destination buffer after a peripheral-to-memory transfer
completes. Descriptors living in cacheable memory must be cleaned before the channel is
started.
Protocol¶
Write
config,src,dst,lengthand optionallynext.Enable the desired sources in
irq_mask.Write
1tocontrolto start. The channel reportsbusyuntil the descriptor (or chain) completes, an error occurs or the channel is aborted.On completion the
donepending bit is set whenirq_donewas set in the finishing descriptor; on failure theerrorpending bit and the channelerrorflag are set. Write1to the pending bit to clear it.Write
2tocontrolto abort a running channel. The channel goes idle after the current chunk without raising an interrupt.
Configuration¶
Available bus architectures:
APB3
TileLink
Wishbone
By default, all register buses are defined with 12 bit address and 32 bit data width. The
memory port is always TileLink (BusParameter.simple(addressWidth, dataWidth, burstBytes,
sourceWidth)). Element accesses and descriptor reads stay 8, 16 or 32 bits wide at their byte
position, so a wider memory port still reaches 32-bit peripherals through a width adapter.
Parameter¶
Name |
Type |
Description |
Default |
|---|---|---|---|
channels |
Int |
Number of channels. Must be between 1 and 8. |
2 |
requestLines |
Int |
Number of peripheral request inputs. Must be between 0 and 16. |
4 |
burstBytes |
Int |
Maximum transfer size on the memory port and FIFO depth in bytes. Power of two between 4 and 4096. |
64 |
addressWidth |
Int |
Memory port address width. Must be between 12 and 32. |
32 |
sourceWidth |
Int |
Memory port TileLink source width. Must be between 1 and 8. |
1 |
dataWidth |
Int |
Memory port data width: 32, 64 or 128. |
32 |
object Parameter {
def default() = Parameter()
def small() = Parameter(channels = 1, requestLines = 4, burstBytes = 16)
def medium() = Parameter(channels = 2, requestLines = 8, burstBytes = 64)
def large() = Parameter(channels = 8, requestLines = 16, burstBytes = 256)
}
Register Mapping¶
IP Identification:
The register map starts with an IP Identification block to provide all information about the underlying IP core to software drivers. This allows to provide backwards compatible drivers.
Address |
Bit |
Field |
Default |
Permission |
Description |
|---|---|---|---|---|---|
0x000 |
31 - 24 |
API |
0x0 |
Rx |
API version of the implemented IP Identification. |
23 - 16 |
Length |
0x8 |
Rx |
Length of the IP Identification block in Bytes. |
|
15 - 0 |
ID |
0x1A |
Rx |
IP value of this IP core. |
|
0x004 |
31 - 24 |
Major Version |
0x1 |
Rx |
Major number if this IP core. Version schema is major.minor.patch. |
23 - 16 |
Minor Version |
0x0 |
Rx |
Minor number if this IP core. Version schema is major.minor.patch. |
|
15 - 0 |
Patch Version |
0x0 |
Rx |
Patch number if this IP core. Version schema is major.minor.patch. |
Global Registers:
Address |
Bit |
Field |
Default |
Permission |
Description |
|---|---|---|---|---|---|
0x008 |
31 - 24 |
0 |
Rx |
Reserved. |
|
23 - 16 |
burstLog2 |
Rx |
log2 of |
||
15 - 8 |
requestLines |
Rx |
Number of request inputs. |
||
7 - 0 |
channels |
Rx |
Number of channels. |
||
0x00C |
2 x channels - 1 downto 0 |
irq_pending |
0 |
RW |
Pending interrupts, bit |
0x010 |
2 x channels - 1 downto 0 |
irq_mask |
0 |
RW |
Interrupt enable per pending bit. The |
0x014 |
channels - 1 downto 0 |
status |
0 |
Rx |
Busy flag per channel. |
Per-Channel Registers:
One set per channel. Channel 0 starts at 0x018; stride between channels is 0x20.
Address |
Bit |
Field |
Default |
Permission |
Description |
|---|---|---|---|---|---|
base + 0x00 |
2 |
request_active |
0 |
Rx |
Level of the selected |
1 |
error / abort |
0 |
RW |
Read: channel stopped with an error (bus error or misaligned element access); cleared by the next start. Write 1: abort the channel. |
|
0 |
busy / start |
0 |
RW |
Read: channel is running. Write 1: start the channel with the current registers. Ignored while busy. |
|
base + 0x04 |
31 - 20 |
0 |
RW |
Reserved. |
|
19 - 16 |
burst_limit |
0 |
RW |
log2 of the largest burst in burst mode; 0 selects |
|
15 - 11 |
0 |
RW |
Reserved. |
||
10 |
irq_done |
0 |
RW |
Raise the done interrupt when this descriptor completes. |
|
9 |
linked |
0 |
RW |
Fetch the descriptor at |
|
8 - 5 |
req_sel |
0 |
RW |
Index of the request line gating this channel. |
|
4 |
req_enable |
0 |
RW |
Gate transfers by the selected request line (element mode). |
|
3 - 2 |
width |
0 |
RW |
Element size in element mode: 0 = 8 bit, 1 = 16 bit, 2 = 32 bit, 3 = invalid. |
|
1 |
dst_inc |
0 |
RW |
Increment the destination address after each chunk. |
|
0 |
src_inc |
0 |
RW |
Increment the source address after each chunk. |
|
base + 0x08 |
addressWidth - 1 downto 0 |
src |
0 |
RW |
Source address. Advances during the transfer; read-only while busy. |
base + 0x0C |
addressWidth - 1 downto 0 |
dst |
0 |
RW |
Destination address. Advances during the transfer; read-only while busy. |
base + 0x10 |
31 - 0 |
length |
0 |
RW |
Remaining bytes. Counts down to 0; read-only while busy. |
base + 0x14 |
addressWidth - 1 downto 0 |
next |
0 |
RW |
Address of the next descriptor, 0 terminates the chain. Read-only while busy. |