You are a senior software architect, full-stack engineer, AI/RAG engineer, database engineer, cybersecurity engineer, and DevOps engineer. I want you to build a COMPLETE production-ready web application called: FAIDHA DIGITAL ARCHIVE & AI The purpose of this project is NOT primarily commercial. It is a long-term digital preservation, research, education, and knowledge project dedicated to preserving and making accessible authentic knowledge related to: - Shaykh Ibrahim Niasse - The Faidha Tijaniyya - Shaykh Ahmad al-Tijani - Tijaniyya history - Fayda scholars - Students and disciples of Shaykh Ibrahim - Faidha institutions and zawiyas - Historical events - Tafsir - Books and writings - Audio lectures - Historical documents - Biographies - Academic research - Translations - Contemporary Fayda activities The system must eventually become a highly intelligent AI research assistant grounded in a carefully curated Fayda knowledge archive. IMPORTANT: Do NOT build this as a simple chatbot. Do NOT build this as a simple PDF upload website. Build the digital archive and knowledge infrastructure first, with the architecture ready for advanced AI/RAG integration. ================================================== 1. CORE VISION ================================================== The final system should have two major layers: LAYER 1: FAYDA DIGITAL ARCHIVE LAYER 2: FAYDA AI Architecture: USER ↓ FAYDA AI ↓ KNOWLEDGE RETRIEVAL ↓ FAYDA DIGITAL ARCHIVE ↓ BOOKS / PDF / AUDIO / VIDEO / TRANSCRIPTS / TRANSLATIONS / HISTORICAL SOURCES ↓ EVIDENCE ↓ AI REASONING ↓ ANSWER + SOURCES The Fayda Archive is the source-of-truth layer. The AI should use its general intelligence for reasoning, language understanding, summarization, comparison, explanation, etc., but Fayda-specific historical claims should be grounded in the archive whenever evidence is available. ================================================== 2. NO HALLUCINATION POLICY This is one of the most important requirements. The AI must NEVER confidently invent: - quotations - dates - historical events - names - relationships - book references - page numbers - Arabic quotations - statements attributed to Shaykh Ibrahim - statements attributed to Shaykh Ahmad al-Tijani - statements attributed to any scholar If evidence is not found, the AI should clearly say: "I could not find sufficient evidence for this in the Fayda Archive." If sources disagree, the AI must say so. Example: "Source A reports X, while Source B reports Y." The AI must distinguish: 1. Primary source 2. Historical document 3. Contemporary testimony 4. Scholar/lecture 5. Academic research 6. Secondary source 7. User-submitted material Never treat a user claim as established fact automatically. ================================================== 3. SOURCE AUTHORITY SYSTEM Every material must have an authority tier. Tier 1: Primary works written by Shaykh Ibrahim Niasse or original historical material. Tier 2: Original historical documents / manuscripts / contemporary records. Tier 3: Works or testimony from direct students/contemporaries and established scholars. Tier 4: Academic research, books, dissertations and scholarly studies. Tier 5: Secondary/general material. Every material must also have: - verification_status - source_name - author - speaker - language - publication information - date - edition - volume - page information where available - provenance - notes - reviewer - verification date ================================================== 4. MATERIAL TYPES The archive must support: - Books - PDFs - Manuscripts - Scanned documents - Audio - Video - Lectures - Tafsir - Articles - Translations - Letters - Biographies - Historical documents - Photographs - Academic papers - Interviews - Other research material ================================================== 5. LIBRARY STRUCTURE Create major collections: A. Shaykh Ibrahim Niasse B. Shaykh Ahmad al-Tijani C. Fayda Tijaniyya D. Tafsir E. Kāshif al-Ilbās F. Diwan G. Fayda History H. Tijaniyya History I. Fayda Scholars J. Students and Disciples K. Zawiyas and Institutions L. Nigeria M. Senegal N. Ghana O. Niger P. Mauritania Q. Sudan R. Other Countries S. Academic Research T. Translations U. Lectures V. Historical Documents ================================================== 6. MATERIAL RELATIONSHIPS The system must allow materials to be linked. Example: TAFSIR WORK ├── Arabic PDF ├── Arabic Audio ├── Arabic Transcript └── English Translation These should be treated as representations/versions of the same underlying work. Another example: KĀSHIF AL-ILBĀS ├── Arabic Original ├── English Translation └── Other Translation The database must support: - original_material_id - translation_of - transcript_of - audio_of - video_of - related_work - related_person - related_event - related_place ================================================== 7. AUDIO SYSTEM Audio is extremely important. The archive must support large audio files. For every audio file store: - speaker - title - language - date - location - duration - source - description - related work - related event The system must be designed for: AUDIO ↓ TRANSCRIPTION ↓ TIMESTAMPS ↓ SEARCHABLE TEXT ↓ AI KNOWLEDGE Arabic audio must be supported. Do not assume English-only transcription. The architecture must be ready for: - Arabic - English - French - Hausa ================================================== 8. TAFSIR SPECIAL STRUCTURE The user already has complete Shaykh Ibrahim Niasse Tafsir PDFs and complete Tafsir audio recordings. The system must support linking: Arabic Tafsir PDF + Arabic Tafsir Audio + Arabic Transcript + English Translation The system should eventually be able to identify: Surah Ayah Volume Page Audio timestamp Example: Surah Al-Baqarah Ayah 255 Arabic PDF page 143 Audio timestamp 01:14:32 Arabic transcript English translation ================================================== 9. SEARCH Do NOT rely only on basic SQL keyword search. Build the architecture for two types of search: A. Keyword/full-text search B. Semantic/vector search The final AI should be able to understand questions rather than just match exact words. Example: User asks: "When did Shaykh Ibrahim first visit Kano?" The system should retrieve relevant historical passages even if the source uses different wording. ================================================== 10. AI/RAG ARCHITECTURE Prepare the application for Retrieval-Augmented Generation. Pipeline: USER QUESTION ↓ QUERY UNDERSTANDING ↓ SEARCH FAYDA ARCHIVE ↓ RETRIEVE RELEVANT CHUNKS ↓ RANK SOURCES ↓ CHECK AUTHORITY ↓ SEND EVIDENCE TO AI MODEL ↓ GENERATE ANSWER ↓ ATTACH CITATIONS ↓ RETURN ANSWER The AI must prioritize higher-authority sources. ================================================== 11. OPENAI INTEGRATION Design the application so OpenAI API can be integrated cleanly. Do NOT hard-code the AI provider throughout the application. Create an abstraction such as: AIProviderInterface This allows future support for: - OpenAI - local/open-source models - other providers OpenAI should be used for high-quality reasoning when appropriate. The architecture should support model routing. For example: Simple retrieval question → lower-cost model Complex historical comparison → stronger reasoning model Difficult Arabic Tafsir research → strongest available reasoning model Do not send every question to the most expensive model. ================================================== 12. AI SOURCE DISCIPLINE When answering a Fayda-specific question, the AI should provide: ANSWER SOURCE(S) AUTHOR WORK PAGE / VOLUME where available AUDIO TIMESTAMP where available LANGUAGE / ORIGINAL SOURCE where useful Example: Source: Kāshif al-Ilbās Author: Shaykh Ibrahim Niasse Page: 147 Evidence type: Primary source For audio: Source: Tafsir lecture Speaker: Shaykh Ibrahim Niasse Timestamp: 01:23:16–01:24:02 ================================================== 13. PUBLIC LIBRARY Create a beautiful public-facing library. Pages: Home Library Search Collections Books Audio Tafsir Lectures Scholars History Translations About Users should be able to: - Search - Browse - Filter - Open material pages - See metadata - See related materials - Read available text - Listen to permitted audio - See citations - Explore related scholars/events/works Do NOT automatically expose private/raw files for download. Allow administrators to control: - public - private - view only - downloadable ================================================== 14. ADMIN DASHBOARD Create a professional admin dashboard. Features: Dashboard statistics: - total books - total audio - total video - total documents - total transcripts - total verified materials - pending submissions - rejected materials - storage usage Admin actions: - Add material - Edit material - Delete/archive material - Upload file - Add metadata - Add transcript - Add translation - Link related materials - Verify material - Reject material - Change authority tier - Manage categories - Manage scholars - Manage events - Manage locations - Manage users/admins ================================================== 15. COMMUNITY SUBMISSION SYSTEM The project must eventually allow people around the world to contribute historical material. Example: A user asks: "Who was Shaykh X?" If the archive has insufficient information, the AI can say: "I do not currently have sufficient information about this person in the Fayda Archive." Then offer: "Do you have books, documents, recordings, photographs, or documentaries about this person and their relationship with Shaykh Ibrahim? You can submit them to the Fayda Archive for review." Create a submission system supporting: - PDF - audio - video - images - documents - external links - written historical information Every submission starts as: PENDING REVIEW Never automatically add it to trusted knowledge. Admin can: APPROVE REQUEST MORE EVIDENCE REJECT Only approved material enters trusted knowledge retrieval. ================================================== 16. KNOWLEDGE GRAPH Design the database so eventually we can build a Fayda Knowledge Graph. Entities: PERSON WORK BOOK AUDIO VIDEO EVENT PLACE SCHOLAR STUDENT ZAWIYA ORGANIZATION DATE COUNTRY SURAH AYAH Relationships: PERSON → STUDENT_OF → PERSON PERSON → TEACHER_OF → PERSON PERSON → MET → PERSON PERSON → ASSOCIATED_WITH → EVENT PERSON → AUTHORED → WORK WORK → TRANSLATION_OF → WORK AUDIO → DISCUSSES → WORK PERSON → ASSOCIATED_WITH → ZAWIYA EVENT → OCCURRED_AT → PLACE This will make future research extremely powerful. ================================================== 17. MEMORY SYSTEM The AI should have a distinction between: A. Conversation memory B. Verified archive knowledge C. Unverified user claims D. Pending submissions E. Approved new knowledge Never mix them. The AI can become more knowledgeable as verified material is added. Workflow: NEW INFORMATION ↓ SUBMISSION ↓ REVIEW ↓ EVIDENCE CHECK ↓ APPROVAL ↓ ARCHIVE ↓ INDEXING ↓ AI CAN RETRIEVE IT ================================================== 18. MULTILINGUAL SUPPORT The system must be Unicode-first. Support: Arabic English French Hausa The AI should eventually answer in the user's language. Arabic must be treated as a first-class language, not an afterthought. ================================================== 19. SECURITY Implement: - secure authentication - password hashing - CSRF protection - prepared SQL statements - input validation - output escaping - MIME validation - upload restrictions - filename randomization - upload size limits - rate limiting - admin authorization - session security - secure headers - audit logs - backup strategy - protection of private materials Do not expose: - database credentials - API keys - private storage paths - admin endpoints without authentication ================================================== 20. DATABASE Use MySQL/MariaDB. Design normalized tables for: admins users materials works authors speakers persons scholars categories collections languages translations transcripts audio_metadata video_metadata events places relationships submissions citations source_references embeddings chunks audit_logs settings Use indexes appropriately. Support UTF-8 / utf8mb4. ================================================== 21. FILE STORAGE Do not store huge files directly inside database BLOBs. Store files in filesystem/object storage. Database stores metadata and storage references. Architecture should eventually support: Local storage S3-compatible storage Cloud object storage ================================================== 22. PROCESSING PIPELINE Build a job-based processing architecture. When a PDF is uploaded: UPLOAD ↓ VALIDATE ↓ STORE ↓ EXTRACT TEXT ↓ OCR IF NEEDED ↓ CLEAN TEXT ↓ SEGMENT ↓ INDEX ↓ READY When audio is uploaded: UPLOAD ↓ VALIDATE ↓ STORE ↓ TRANSCRIBE ↓ TIMESTAMP ↓ CLEAN ↓ SEGMENT ↓ INDEX ↓ READY Do not make the web request wait for long transcription jobs. Use background jobs/queue architecture where possible. ================================================== 23. ADMIN PROCESSING STATUS Every material should show processing status: UPLOADED PROCESSING TEXT_EXTRACTED TRANSCRIBING TRANSCRIBED OCR_REQUIRED INDEXING READY FAILED Admin should see errors. ================================================== 24. CITATION ENGINE Build citations into the knowledge model from the beginning. A citation should be able to point to: Book page PDF page Volume Chapter Section Audio timestamp Video timestamp Document section Transcript segment This is essential for scholarly credibility. ================================================== 25. UI DESIGN Design should feel: - scholarly - premium - calm - trustworthy - modern - archival - easy to navigate Do NOT make it look like a children's Islamic website. Do NOT overload the interface with unnecessary decoration. The most important thing is usability and trust. Responsive: Desktop Tablet Mobile ================================================== 26. DEPLOYMENT The initial deployment target is: cPanel Apache PHP 8.1+ MySQL/MariaDB Provide: - complete source code - database schema - configuration example - .htaccess - installation instructions - environment configuration - storage setup - cron/job instructions - backup instructions The project must be packaged as a ZIP ready for deployment. ================================================== 27. CODE QUALITY Use clean architecture. Separate: Frontend Backend Database Authentication File storage Processing AI Search Admin API Do not put everything into one PHP file. Use reusable components. Document important code. Do not leave fake placeholder functions where real functionality is required. If a feature cannot be fully implemented without an external API, build the integration interface and clearly document what credentials/configuration are required. ================================================== 28. IMPORTANT EXISTING MATERIALS The initial archive will contain: 1. Complete Shaykh Ibrahim Niasse Tafsir PDFs in Arabic. 2. Complete audio recordings of Shaykh Ibrahim Niasse's Tafsir in Arabic. 3. Kāshif al-Ilbās Arabic PDF. 4. Kāshif al-Ilbās English translation PDF. 5. Additional Fayda/Tijaniyya books and documents. 6. More materials will be added later, including Diwan. The system must be designed to handle these from the beginning. ================================================== 29. FUTURE FEATURES Prepare architecture for: - AI chat - voice questions - Arabic voice search - semantic search - book comparison - scholar comparison - historical timeline - interactive map - knowledge graph visualization - quote finder - Tafsir verse explorer - audio search - transcript search - source comparison - multilingual translation - research mode - downloadable citations - researcher accounts - community submissions - moderation - API access Do not necessarily implement every future feature in version 1, but make the architecture extensible. ================================================== 30. VERY IMPORTANT DEVELOPMENT RULE Do not skip architecture. Before coding: 1. Design database schema. 2. Design folder structure. 3. Design API structure. 4. Design authentication. 5. Design file-processing architecture. 6. Design search architecture. 7. Design AI/RAG integration architecture. 8. Design citation system. 9. Design community verification workflow. 10. Then implement. Do not simplify the system into a basic CRUD application. ================================================== 31. DELIVERABLE I want a COMPLETE ZIP project. The ZIP must contain: - all source code - database schema - installation files - configuration examples - frontend - backend - admin dashboard - public library - upload system - search - authentication - metadata management - verification system - processing architecture - API foundation - AI integration foundation - documentation The project must be runnable after configuration. At the end provide: 1. Project structure 2. Installation instructions 3. Database setup 4. Admin setup 5. Upload instructions 6. AI API configuration 7. Storage configuration 8. Cron/background job configuration 9. Security checklist 10. Known limitations 11. Next recommended development phase IMPORTANT FINAL INSTRUCTION: Do not claim that AI knowledge has been created simply because files were uploaded. Knowledge ingestion must be an explicit pipeline: FILE → EXTRACTION/TRANSCRIPTION → CLEANING → CHUNKING → METADATA → INDEXING → VERIFICATION → RETRIEVAL → AI The system must be designed so that when thousands of books and hours of audio are eventually added, the architecture can continue to scale. Build this as the foundation of a serious, long-term Fayda knowledge preservation and research platform.