Sources
Knowledge with a paper trail.
Shuka's corpus is eight published extension documents — about 850 pages — each license-verified before a single passage was indexed. If we can't name the page an answer came from, the answer doesn't ship.
The corpus
Indexed and answering.
Licenses were verified in the documents themselves or in their repository records, not assumed. The full table — with URLs, verification notes and page counts — is corpus/SOURCES.md in the repository. Because several licenses are non-commercial, the source PDFs are fetched by script rather than redistributed; the built index ships with attribution.
What we left out
The documents we found and did not index.
An honest corpus is defined as much by its exclusions.
Staged, awaiting clearance
NAERLS extension bulletins
Nigeria's national extension bulletins on maize, cassava, rice and tomato — the most locally-specific material there is — carry no license statement. They sit quarantined outside the index until cleared, and they are first in the roadmap.
Rejected
IITA vegetable IPM guide (2010)
Downloaded, then found to state "no reproduction without written permission of the publisher." Deleted. Excellent content is not a license.
Unverified
AfricaRice production handbooks
The best practical West African rice manuals name no license and parts of the PDFs carry broken font encoding. Held out of the index on both grounds.
Why this matters
Advice you can check is advice you can trust.
Every Shuka answer names its documents and pages, so an extension officer can open the manual and verify — or overrule. That is the relationship the tool is built for: it hands professionals a faster index into literature they already trust, not a rival authority.
8 manuals.
Every one cited.
~850 pages, embedded and searched entirely on the laptop.