Page MenuHomeFreeBSD

loongarch: add initial kernel support
Needs ReviewPublic

Authored by zhaoxiaoqiang007_gmail.com on Tue, Sep 22, 2:23 PM.
Tags
None
Referenced Files
F174745482: D59890.id188510.diff
Mon, Oct 5, 4:54 PM
Unknown Object (File)
Sun, Oct 4, 6:15 PM
Unknown Object (File)
Sat, Oct 3, 7:58 PM
Unknown Object (File)
Sat, Oct 3, 10:18 AM
Unknown Object (File)
Fri, Oct 2, 5:27 AM
Unknown Object (File)
Thu, Oct 1, 10:47 PM
Unknown Object (File)
Thu, Oct 1, 10:47 PM
Unknown Object (File)
Thu, Oct 1, 5:16 AM

Details

Summary

Add the LoongArch machine code. Including pmap, trap handling,
exception vectors, interrupt controllers, timer, SMP, bus_dma,
ptrace and DDB support, and FDT boot glue, plus small hooks in MI code.
Currently only QEMU is supported.

The port heavily references the existing riscv code.

LoongArch is a new RISC ISA from Loongson Technology, a bit like MIPS
or RISC-V. It defines four privilege levels, PLV0~PLV3: the kernel
runs at PLV0 and applications at PLV3. This port targets the 64-bit
variant (LA64): 32 GPRs, 32 FPRs, optional LSX/LASX vector extensions,
privileged-mode CSRs, a software-managed TLB with ASIDs, and DMWn
direct-mapped address windows used for the kernel direct map.

References:
https://github.com/loongson/LoongArch-Documentation/releases/latest/download/LoongArch-Vol1-v1.00-EN.pdf
https://github.com/loongson/LoongArch-Documentation/releases/latest/download/LoongArch-ELF-ABI-v1.00-EN.pdf
https://loongson.github.io/LoongArch-Documentation/
http://www.loongson.cn/
Signed-off-by: Xiaoqiang Zhao <zhaoxiaoqiang007@gmail.com>

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped
Build Status
Buildable 77168
Build 74051: arc lint + arc unit

Event Timeline

There are a very large number of changes, so older changes are hidden. Show Older Changes
sys/loongarch/loongarch/db_disasm.c
267

The decoder loop at line 278 relies on op->name != NULL to terminate:

for (op = ops; op->name NULL; op++) {
However, the table ends at line 267 with { "csrrd", ... } and has no sentinel { NULL, ... } entry. After processing csrrd, op++ moves past the and the next iteration reads out-of-bounds memory to check op->name. This is undefined behavior and may crash DDB or call a bogus function pointer through op->pr.

It is suggested to add a new line after this line:

	{ NULL,      0,          0,          NULL },

JIRL/LL/SC offsets are as *4 instead of *2.

sys/loongarch/loongarch/db_disasm.c
90–99

Based on the LoongArch manual's explanation of JIRL/LL/SC, these immediates are shifted left by 2 bits, the code here should be changed to *4.

JIRL: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#_jirl
LL: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#_ll_wd_sc_wd
SC: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#_sc_q

static void pr_r_i14s2(struct op *op, vm_offset_t loc, uint32_t insn)
{
	db_printf("%-9s\t%s, %s, %d", op->name,
	    rname[RD(insn)], rname[RJ(insn)], SI14(insn) * 4);
}
static void pr_rr_i16s2(struct op *op, vm_offset_t loc, uint32_t insn)
{
	db_printf("%-9s\t%s, %s, %d", op->name,
	    rname[RD(insn)], rname[RJ(insn)], SI16(insn) * 4);
}

addi.w skip logic is based on an incorrect premise, It is suggested to remove this code.

sys/loongarch/loongarch/db_disasm.c
281–289

addi.w and ld.b do not have the same opcode:

addi.w = 0x02800000,bits[31:22] = 0x00a
ld.b = 0x28000000,[31:22] = 0x0a0

This code is based on an incorrect assumption, and since ld.b comes before addi.w in ops[], it would already match and return first anyway. This skip logic is dead/misleading code.

invtlb is not in 3R register format, but rather op, rj, rk (an immediate plus two registers). Using pr_rrr would incorrectly display the op field as a register name.

invtlb is printed using pr_rr, which assumes a 3-register format rd, rj, rk. However, per LoongArch Vol1 section 4.2.4.7, invtlb has the format invtlb op, rj, rk where op` is a 5-bit immediate, not a register.

The current code interprets bits[4:0] as rd and prints it via rname[], so invtlb 0x0, a, a1 is displayed as invtlb zero, a0, a1, which is misleading.

Please add a dedicated printer that formats the first operand as an immediate: invtlb <op>, rj, rk.

invtlb: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#_invtlb

sys/loongarch/loongarch/db_disasm.c
27

Sync the same change:

#define	OP(insn)	(((insn) >> 0) & 0x1f)	/* invtlb op immediate */
108

suggested to add a new line after this line:

static void pr_invtlb(struct op *op, vm_offset_t loc, uint32_t insn)
{
	db_printf("%-9s\t%u, %s, %s", op->name,
	    OP(insn), rname[RJ(insn)], rname[RK(insn)]);
}
265

Suggested change to:

{ "invt",  0x06490000, 0xffff801f, pr_invtlb },
281–289

addi.w and ld.b do not have the same opcode:

addi.w = 0x02800000,bits[31:22] = 0x00a
ld.b = 0x28000000,[31:22] = 0x0a0

This code is based on an incorrect assumption, and since ld.b comes before addi.w in ops[], it would already match and return first anyway. This skip logic is dead/misleading code.

It is suggested to remove this code.

%u is used to print uint64_t fields.

sys/loongarch/loongarch/db_interface.c
116–117

%u is used to print uint64_t fields.

In db_show_mdpcpu():

db_printf("kernel    = %u\n", pc->pc_kernel_sp);
db_printf("asid_value   = %u\n", pc->pc_asid_value);

However, per machine/pcpu.h, both pc_kernel_sp and _asid_value are uint64_t. The %u format specifier expects unsigned int (32-bit), so on 64-bit builds this truncates the upper 32 bits of the value.

pcpu.h: https://reviews.freebsd.org/D59890

	uint64_t pc_kernel_sp;  /* kernel sp */ 	\
	uint64_t pc_asid_value;	/* ASID */ \

Please use %ju with auintmax_t` cast, suggested change to:

	db_printf("kernel_sp    = %ju\n", (uintmax_t)pc->pc_kernel_sp);
	db_printf("asid_value   = %ju\n", (uintmax_t)pc->pc_asid_value);

Please pass count through to db_stack_trace_cmd() and use it to bound the loop.

DDB passes a count argument to db_trace_thread() to limit the number
of stack frames printed.However, db_trace_thread() ignored the argument and
called db_stack_trace_cmd() with no limit, which used an unbounded
while (1) loop. This meant user-specified trace limits ignored
and the entire stack was always printed.

Add an int count parameter to db_stack_trace_cmd() and replace the
while (1) loop with a counted loop. A negative count is treated as
unlimited, preserving the previous behaviour for db_self().

sys/loongarch/loongarch/db_trace.c
54

Suggested change to:

db_stack_trace_cmd(struct thread *td, struct unwind_state *frame, int count)
62

Sync the same change:

	for (int depth = 0; count < 0 || depth < count; depth++) {
128

Sync the same change:

	db_stack_trace_cmd(thr, &frame, count);
140

Sync the same change:

	db_stack_trace_cmd(thr, &frame, -1);
sys/loongarch/loongarch/dump_machdep.c
50–55

This function is called during a crash dump to temporarily map chunk pages starting at physical address pa and return the virtual address via *va.

The function is not implemented here; it looks like it only does a single line of printing. Would it be possible to add a TODO or some explanation?

For reference, the x86(dumpsys_map_chunk) implementation maps each with pmap_kenter_temporary(trunc_page(a), i).

sys/loongarch/loongarch/eioic.h
12

The header is not self-contained.

eioic.h uses struct intr_irqsrc, device_t, and struct mtx, but does not include the headers that define them.

It currently compiles only because eioic.c and pch_pic.c already include <sys/interrupt.h>, <sys/types.h> (pulled in indirectly via <sys/bus.h>, etc.), and <sys/mutex.h> before including eioic.h.

#include <sys/types.h>		/* device_t */
#include <sys/cpuset.h>	/* cpuset_t */
#include <sys/mutex.h>		/* struct mtx */
#include <sys/intr.h>		/* struct intr_irqsrc */
24

struct eioic_softc defines struct resource *irq_res, but it is never read or written anywhere in eioic.c; it is suggested to remove it.

sys/loongarch/loongarch/elf_machdep.c
488–511

Should the 32-bit-related code be removed?

The full FP / LSX / LASX context is saved on every exception, which is a significant overhead on every exception.

The save_registers macro checks CSR_EUEN_FPEN / LSXEN / LASXEN at the exception entry and saves the context, which is a significant overhead on every exception.

Other FreeBSD ports typically use lazy saving: the context is saved only when the FPU/vector unit is first used, triggering an fpudis/lsxdis/lasxdis exception.

However, I see that this would touch many files and involve a large amount of code, so it can be handled later.

sys/loongarch/loongarch/exception.S
99

fix type

	/* save static register */
414

fix type

	 * LoongArch does.

Do not allow GDB to write $zero (r0) into the trapframe.

sys/loongarch/loongarch/gdb_machdep.c
87

gdb_cpu_setreg() currently accepts regnum == 0 because the regnum >= 0 check includes r0, and stores the value into gdb_frame->tf_regs[0].

LoongArch r0/$zero is architecturally hardwired to zero; any write to it is discarded by hardware. Writing through KGDB therefore not change the CPU register but does corrupt the trapframe's view of r0, which can mislead fill_regs(), get_mcontext(), or crash dumps.

	if (kdb_thread == curthread) {
		/* $zero (r0) is hardwired to zero by the architecture. */
		if (regnum == GDB_REG_ZERO)
			return;

Fix non-standard C99 anonymous unions/structs.

Modify struct pcb to remove the anonymous union, change it to a normal array member, and keep the named access macros.

sys/loongarch/include/pcb.h
43–70

Fix non-standard C99 anonymous unions/structs.

Modify struct pcb to remove the anonymous union, change it to a normal array member, and keep the named access macros.

	uint64_t	pcb_regs[32];	/* general purpose registers */
#define	pcb_a	pcb_regs[4]
#define	pcb_t		pcb_regs[12]
#define	pcb_s		pcb_regs[23]
#define	pcb_ra		pcb_regs[1]
#define	pcb_sp		pcb_regs[3]
	uint64_t	pcb_fregs[34];	/* floating point registers */
#define	pcb_fcsr0	pcb_fregs[32]
sys/loongarch/loongarch/exec_machdep.c
332

Sync the same change:

mcp->mc_fpregs.fp_fcsr = curpcb->pcb_fcsr0;
sys/loongarch/include/pci_cfgreg.h
12

This header is the machine-dependent header for the ACPI subsystem to access PCI configuration space, but here it is an empty file used as a placeholder. A TODO could be added to explain this.

sys/loongarch/loongarch/copyinout.S
163–164

That's reasonable

sys/loongarch/loongarch/exec_machdep.c
235

Right

297

Yes, will add validation in next version

449

yes

sys/loongarch/include/intr.h
146–147

will remove in next reversion

sys/loongarch/include/pcpu.h
48–50

fix type

	uint64_t pc_kernel_sp;  /* Kernel stack for user exceptions */ \
	uint64_t pc_asid_value;	/* Current ASID value */ \
	uint32_t pc_asid_mask;	/* ASID bit mask */ \
sys/kern/link_elf_obj.c
1004–1005

will add this two constant in next reversion

sys/loongarch/loongarch/elf_machdep.c
488–511

I think this should be kept, it's legal 32 bit relocation in 64 bit ELF, not 32 bit ABI.

cpu_desc[MAXCPU] is declared in pmap.h, which is not an appropriate place.

sys/loongarch/include/cpu.h
257

struct cpu_desc is defined in <machine/cpu.h>, and the global array is defined in identcpu.c. Putting the extern declaration in pmap.h introduces unnecessary dependencies; it should be moved into <machine/cpu.h>.

extern cpu_desc cpu_desc[MAXCPU];
sys/loongarch/include/pmap.h
182–183

Remove it here, and move extern struct cpu_desc into cpu.h.

This macro is defined in elf.h, but hardware capability bits are more appropriate in <machine/cpu.h>. Making pmap.h depend on an ELF header is not appropriate.

sys/loongarch/include/cpu.h
257

This macro is defined in elf.h, but hardware capability bits are more appropriate in <machine/cpu.h>. Making pmap.h depend on an ELF header is not appropriate.

#define	CPU_HWCAP_PTW	(1 << 13)	/* Page Table Walker extension */
sys/loongarch/include/pmap.h
51

suggested to remove this include header, as the reference has already been moved into cpu.h.

200

Sync the same change:

if (cpu_desc[cpu].cpuinfo.hwcap & CPU_HWCAP_PTW)
sys/loongarch/include/cpu.h
257

that's reasonable

sys/loongarch/loongarch/db_trace.c
54

seems a good improvement, but loongarch is sync with other arch now, so I will ignore this for now and maybe we can submit in a separate patchset later.

sys/loongarch/loongarch/dump_machdep.c
50–55

currently, it mirrors arm64 and riscv and just leave it blank

sys/loongarch/loongarch/eioic.h
12

will fix

sys/loongarch/loongarch/gdb_machdep.c
87

will fix this

TF_FCSR0/_FCC relies on anonymous union members.

Remove the anonymous union for the FP part and change it to separate fields.

sys/loongarch/include/frame.h
83–92

TF_FCSR0/_FCC relies on anonymous union members.

Remove the anonymous union for the FP part and change it to separate fields.

	uint64_t tf_fregs[32];	 FP registers $f0-$f31 */
	uint64_t tf_fcsr0;	/* FP control/status register 0 */
	uint64_t tf_fcc;	/* FP condition codes */
sys/loongarch/loongarch/genassym.c
173–174

Sync the same change:

ASSYM(TF_FCSR0, offsetof(struct trapframe, tf_fcsr0));
ASSYM(TF_FCC, offsetof(struct trapframe, tf_fcc));
sys/loongarch/loongarch/trap.c
571

Sync the same change:

uint64_t fcsr = frame->tf_fcsr0;

update to second reversion

  1. rebase main
  2. use platform KMOD_MIN/MAX_ADDRESS macro to limit module VA
  3. drop memguard related fixes
  4. remove anonymous union types
  5. misc code cleanup and optimization
  6. typos, errors and comments fixup

mvendorid/marchid/mimpid are RISC-V CSR names, and has_sstc/has_sscofpmf are RISC-V extensions. LoongArch does not have these, so they should not exist in LoongArch at all. LoongArch's CPU identity information should be read via CPUCFG-related instructions.

sys/loongarch/include/md_var.h
40–46

RISC-V-specific concepts left over directly from RISC-V, RISC-V uses the CSR names mvendorid/marchid/mimpid to read CPU identification. LoongArch does not have these CSRs; it reads PRID and Arch information via CPUCFG0/CPUCFG1. Yet the LoongArch version retains the RISC-V variable names.

https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#_cpucfg

https://docs.riscv.org/reference/isa/v20250508/priv/sstc.html

sys/loongarch/loongarch/identcpu.c
61–68

suggested to remove it as well.

Change the cpuhead list traversal to CPU_FOREACH_ISSET to iterate directly over the target cpuset, avoiding scanning all online CPUs, consistent with other architecture implementations, and remove the redundant pc_pending_ipis setting.

sys/loongarch/loongarch/intr_machdep.c
114

Sync the same change:

int cpu;
117–122

Change the cpuhead list traversal to CPU_FOREACH_ISSET to iterate directly over the target cpuset, avoiding scanning all online CPUs, consistent with other architecture implementations, and remove the redundant pc_pending_ipis setting.

	CPU_FOREACHSET(cpu, &cpus) {
		CTR3(KTR_SMP, "%s: cpu: %d, ipi: %x",
		    __func__, cpu, ipi);
		ipi_send(cpuid_to_pcpu[cpu (uint32_t)ipi);

cpu_flush_dcache() is an empty function, which would make callers believe the cache has been flushed.

sys/loongarch/loongarch/machdep.c
214–215

An empty implementation would make callers believe the cache has been flushed when it actually has not. In scenarios such as kexec loading a new kernel, writing data to a memory disk, or loading modules, the CPU may fetch stale instructions from the I-cache or read stale data from the D-cache, causing boot failures or data corruption.

A memory barrier can be implemented as a stopgap for now:

__asm __volatile("dbar 0" ::: "memory");
sys/loongarch/loongarch/mem.c
101–102

can be made more concise like this:

error = uiomove(PHYS_TO_DMAP(v), cnt uio);

FDT and ACPI code for starting an AP does the same thing, so they must send the same IPI. Changing both the original 2 (PREEMPT) and 1 (AST) to IPI_WAKEUP eliminates the inconsistency and avoids sending a preemption signal during early AP startup.

sys/loongarch/include/smp.h
40
/*
 * IPI used to wake a newly-started AP once its boot mailbox has been written.
 * Alias to IPI_AST because the AP is not yet in the scheduler and only needs
 * a lightweight door, not a preempt.
 */
#define	IPI_WAKEUP	IPI_AST
sys/loongarch/loongarch/mp_machdep.c
852

Sync the same change:

	ipi_cpu(cpuid IPI_WAKEUP);
909

Sync the same change:

ipi_write_action(coreid, IPI_WAKEUP);

When setting the CPU bitmap all_cpus, cpu_mp_start() incorrectly uses the hardware cpuid boot_cpu, whereas all_cpus should use the logical cpuid (0-based).

sys/loongarch/loongarch/mp_machdep.c
934

boot_cpu is the hardware cpuid, and all_cpus is a logical cpuid bitmap.

When setting the CPU bitmap all_cpus, cpu_mp_start() incorrectly uses the hardware cpuid boot_cpu, whereas all_cpus should use the logical cpuid (0-based).

so when the hardware boot CPU is not 0, the bit will be set in the wrong position, causing errors in IPI sending, CPU traversal, and similar logic.

Suggested change to:

	/* CPU 0 is always the boot CPUlogical id). */
	CPU_SET(0, &all_cpus);

ofwbus should only be added when the system has no ACPI RSDP (i.e. FDT boot).

sys/loongarch/loongarch/nexus.c
207

ofwbus is the root bus of the FDT device tree and should not appear on ACPI systems. Currently exus_attach() places ofwbus outside the if (loongarch_efi_acpi_rsdp() != 0) branch, causing ofwbus to be attached even during ACPI boot.

Note: this is only a minimal fix. The cleaner approach is to follow ARM64 and split FDT and ACPI into two separate sub-drivers, with ofwbus attached only in nexus_fdt_attach().

Suggested change to:

	if (loongarch_efi_acpi_rsdp() == 0)
		nexus_add_child(dev, 7, "ofbus", 0);
sys/loongarch/loongarch/ofw_machdep.c
28

This file does not use any macros from cdefs.h, suggested to remove it.

The version register read operation uses unnamed numbers.

sys/loongarch/loongarch/pch_pic.c
59

Define named constants and cite the manual:

#define	PCH_ID		0x00	/* INT_ID, identification reg 1 */
#define	PCH_PIC_INT_NUM		0x04	/* INT_ID2[23:16]: num sources - 1 */
#define	PCH_PIC_INT_NUM_SHIFT16
#define	PCH_PIC_INT_NUM_MASK	0xff

description-of-interrupt-related-registers: https://loongson.github.io/LoongArch-Documentation/Loongson-7A1000-usermanual-EN.html#description-of-interrupt-related-registers

316–317

Use the three defined constants PCH_PIC_INT_NUM, PCH_PIC_INT_NUM_SHIFT, and PCH_PIC_INT_NUM_MASK.

	/*
	 * Determine number of interrupt inputs from the identification
	 * register 2 (LS7A1000 manual §5.2: INT_ID2[23:16] + 1+	 */
	sc->vec_count = ((pch_pic_read(sc, PCH_PIC_INT_NUM) >>
	    PCH_PIC_INT_NUM_SHIFT) & PCH_PIC_INT_NUM_MASK) + 1;

pmap_enter() hardcodes PTE_CC and ignores m->md.pv_memattr.

When constructing the PTE, pmap_enter() hardcodes PTE_CC without reading m->md.pv_memattr, causing device/uncached pages to also be mapped as cacheable.

sys/loongarch/loongarch/pmap.c
349

Sync the same change:

static inline pt_entry_t pmap_memattr_bits(vm_memattr_t mode);
3561

Sync the same change:

	new_l3 = PTE_V | PTE_P | pmap_memattr_bits(pmap_page_get_memattr(m)) |
           PTE_NX | PTE_K;
3829

Sync the same change:

PTE_V | pmap_memattr_bits(pmap_page_get_memattr(m)) | PTE_P | PTE_NX);
4190

Sync the same change:

PTE_V | pmap_memattr_bits(pmap_page_get_memattr(m)) | PTE_P | PTE_NX;

ptrace_clear_single_step() only clears a software flag, td->td_md.md_ss_active, but does not write the hardware CSRs (LOONG_CSR_IB0CTRL and LOONGARCH_CSR_FWPS) to actually disable IB0.

If the target thread is running on the current CPU, IB0 is still armed and FWPS.SKIP is still set, so the next instruction fetch may still hit IB0 and trigger an extra single step, and only afterwards will la_watch_handler() clear IB0CTRL.

fetch-watchpoint-overall-status: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#fetch-watchpoint-overall-status

fetch-watchpoint-n-configuration: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#fetch-watchpoint-n-configuration

sys/loongarch/loongarch/ptrace_machdep.c
56

Sync the same change:

#include <machine/watch.h>
83

It is suggested to add after this line:

	if (td curthread)
		la_watch_load(td);

ll/sc only guarantees the atomicity of that read-modify-write, not ordering with surrounding memory accesses; a successful CAS should have at least acquire/release semantics.

ll.acq/sc.rel can be used on CPUs that support LLACQ_SCREL, otherwise, add a dbar 0 before the ll and after a successful sc.

dbar: https://loongson.github.io/LoongArch-Documentation/LoongArch-Vol1-EN.html#_dbar

sys/loongarch/loongarch/support.S
55

suggested to add after this line:

dbar	0			/* release: prior accesses complete before CAS */
60

suggested to add after this line:

dbar	0		/* acquire: subsequent accesses after CAS */
79

suggested to add after this line:

dbar	0			/* release: prior accesses complete before CAS */
84

suggested to add after this line:

dbar	0			/* acquire: subsequent accesses after CAS */

base_freq * cfm is computed as uint32_t, so the intermediate result may exceed 2^32-1 and overflow; it should be changed to ((uint64_t)base_freq * cfm) / cfd to perform 64-bit multiplication.

sys/loongarch/loongarch/timer.c
168

Suggested change to:

*freq = ((uint64_t)base_freq * cfm) / cfd;