ما حُسم، وما بقي بحثيًا؟
هذه الصفحة تفصل قرارات النموذج الأولي عن النتائج البحثية. لا توجد هنا نتائج دراسات التحقق أو أرقام دقة جديدة؛ الموجود هو نطاق وبروتوكول مبدئي قابل للمراجعة الأخلاقية.
01متاح في النموذج الأوليالقسم المرجعي الأول
قسم تقنية المعلومات، كلية الحاسبات وتقنية المعلومات، جامعة الملك عبدالعزيز هو النطاق المرجعي الحالي للمقررات وربط البيانات التجريبي. هذا لا يعني موافقة مؤسسية أو إطلاقًا رسميًا.
02محسوم للنسخة العامةلا يوجد نموذج عربي توليدي داخل قرار الملاءمة الحالي
النسخة العامة تستخدم OCR عربيًا محليًا داخل المتصفح (Tesseract language data) لاستخراج النص، ثم قواعد/أنطولوجيا/graph traversal محكومة للربط والاستدلال. اختيار أي نموذج لغوي عربي مستقبلي يبقى قرارًا منفصلًا يتطلب تقييمًا وترخيصًا وحوكمة قبل أن يؤثر في المنتج.
03خطة — لا نتائجدراسة استرجاعية
الهدف: مقارنة مخرجات كامن على مجموعة تاريخية من السجلات المنزوعة الهوية بمرجع بشري محدد مسبقًا. قبل التنفيذ يجب تثبيت عينة الدراسة، معايير الإدراج، تعريف المرجع البشري، مخرجات القياس، وخطة التحليل قبل فتح البيانات.
04خطة — لا نتائجنسبة إكمال إيداع الدليل
الهدف: قياس نسبة المشاركين المدعوين الذين يكملون إيداعًا قابلًا للاعتماد ويصلون إلى أول ملف قدرات. يجب فصل هذا القياس عن أي موافقة بحثية إضافية وعدم جمع محتوى السجل في بيانات الاستخدام.
مبادئ القياس
- أهداف القياس المنشورة في صفحة المنهجية ليست نتائج.
- أي تقييم للدقة يجب أن يستخدم مرجعًا بشريًا موثقًا.
- تغطية الربط، نسبة الأحكام التي تحمل مرجعًا للدليل، ووقت الوصول لأول قيمة تُقاس كمؤشرات لجودة المنتج.
- لا تدخل السيكومتريات أو الخصائص الحساسة في الملاءمة قبل تحقق مستقل.
حالة الموافقة الأخلاقية
هذه الصفحة لا تدعي وجود موافقة أخلاقية أو IRB لدراسات التحقق. أي استخدام لسجلات حقيقية أو نشر نتائج يتطلب المسار الأخلاقي والنظامي المناسب قبل التنفيذ.
مقترح مؤرخ 30 سبتمبر 2026 — لا نتائجخطة التحقق خلال تسعة أشهر
يبدأ الشهر الأول بعد اعتماد البروتوكول والجهات المشاركة، وليس من تاريخ نشر الصفحة. التواريخ والمسؤوليات تحتاج تأكيد المشرف، ولا توجد شراكات أو عينات مجندة معلنة.
هدف الفصل هو منتج قابل للاستخدام وتجربة أولية في الأسابيع 10–12 بعد الاعتماد؛ بقية الأشهر لتعميق التحقق والمتابعة، والشهر العاشر للتسليم النهائي. لا تؤجل هذه الخارطة تسليم الفصل.
الشهر 1–2: اعتماد النطاق والبروتوكول
المشرف + الجهات المشاركة · تحديد الكليات والموافقات والمقاييس ومسؤوليات البيانات. هدف: مراجعة 20 مسارًا على الأقل بواسطة مختصين اثنين قبل اعتماد الربط.
الشهر 2–3: اختبارات استخدام تكوينية
فريق التجربة + المستشارون · هدف: 12 طالبًا سعوديًا على جولتين، مع العربية والجوال. تُقاس ست مهام ببيانات وهمية وتُعالج العوائق قبل التوسع.
الشهر 3–4: تجربة استكشافية محددة بنهاية الفصل
المشرف + الكليات الموافق عليها · هدف: 60 مشاركًا، 20 لكل كلية. تعرض النجاحات المستقلة والمساعدة والفشل، وفواصل عدم اليقين، مع فصل نتائج اللغة والجهاز.
الشهر 5–6: نموذج تحقق من المصدر
فريق الأدلة + جهة إصدار مخولة · هدف: 20 حالة اختبار اصطناعية تشمل التوقيع غير الصحيح والانتهاء والإلغاء. لا ترقية لمستوى مؤسسي قبل اجتياز الحالات السلبية ومراجعة أمنية.
الشهر 7–8: متابعة الاستخدام وإعادة التقييم
الفريق + المستشارون · متابعة اختيارية معتمدة للمشاركين، ومقارنة الإصلاحات بخط البداية على مجموعة اختبار منفصلة؛ لا ادعاء أثر سببي على التوظيف.
الشهر 9: قرار الاستمرار والتسعير
المشرف + مركز الإرشاد · هدف: 5 مقابلات مع أصحاب قرار شراء ومراجعة النتائج والقيود وتكلفة التشغيل. لا انتقال لخدمة مدفوعة تلقائيًا.
عتبات مقترحة قبل التجربة — ليست نتائج
- إتمام مستقل ≥80% لكل مهمة أساسية؛ الفشل والمساعدة في المقام، ولا يُستبعد الانسحاب من التقرير.
- فهم المصدر وحد الاستنتاج (2 من 2) لدى ≥80% من المحاولات المؤهلة؛ ترميز بمراجعين لعينة محددة مسبقًا.
- صفر ترقيات خاطئة للدليل في مجموعة اختبارات التحقق السلبية، وصفر عيوب خصوصية حرجة غير معالجة قبل تجربة سجلات حقيقية.
- يُجمّد التعريف والعينة وقواعد التوقف قبل الدراسة. هذه أهداف لاتخاذ قرار الاستمرار وليست برهانًا على الدقة أو العدالة أو تحسن التوظيف.
مهام الاستخدام: بناء ملف، تصحيح دليل، فهم توصية، اختيار خطوة، استعادة نسخة، والتصحيح والسحب والحذف. نشر النسب مع العدد وفواصل عدم اليقين؛ عدم تحقق الهدف يستدعي الإصلاح وإعادة القياس.
سرعة المراجعة وجودتها — أهداف قبل الدراسة
مقترح غير معاير؛ لا نتائج ميدانية منشورة. تُجمّد التعريفات والحالات والأهداف مع المشرف قبل جمع النتائج، وتُفصل اللغة والجهاز ودور المراجع.
زمن القبول أو الرفض بثقة
من عرض مهمة المراجعة والتوصية لأول مرة إلى تسجيل قبول أو رفض مع أعلى درجة ثقة من ثلاث درجات، وشرح صحيح للمصدر وحد الاستنتاج يراجعه مقيّم مستقل. هدف أولي: ≥80% من المهام المؤهلة تنتهي بقرار مستوفٍ، ووسيط الزمن ≤120 ثانية للمحاولات المستوفية. لا يصبح القرار صحيحًا بمجرد اختيار «واثق».
تُعرض أعداد الاعتراض والتأجيل والفشل والمساعدة والانقطاع والانسحاب منفصلة؛ لا تُحذف لتحسين النسبة. المؤقت المحلي اختياري ويبدأ عند تفعيله، ولذلك ليس بديلًا تلقائيًا لقياس أول تعرض في الدراسة. تُستبعد محاولاته المنقطعة من ملخص الزمن فقط مع إعلان عددها.
الأخطاء التي عبرت منظومة التحقق
في حزمة وهمية مجمّدة: عدد الحالات ذات العيب المعروف التي قُبلت بعد المراجعة ÷ جميع الحالات ذات العيب المعروف التي عُرضت للمراجعة. هدف أولي ≤5%، مع صفر قبول لترقية دليل زائفة أو عيب خصوصية حرج. تُعرض كذلك الأخطاء المؤكدة في عينة مستقلة من القرارات المقبولة مع مقامها؛ لا تُخلط مع معدل الحزمة المزروعة. الاعتراضات غير المفحوصة ليست أخطاء مؤكدة.
فحوص القواعد والمراجعة السلبية في المستودع تعمل الآن. دراسة قرارات الطلاب والمستشارين، والتحقق من مصدر جامعي، ومعالجة الاعتراضات مركزيًا لم تُنفذ بعد. لا تُحسب نتائج الاختبارات الهندسية نتائجَ دراسة بشرية.
مقاييس حلقة المهمة
أهداف قبل الدراسة، لا نتائج: نسبة المهام المبدوءة التي تصل إلى مخرج مرفق؛ نسبة المخرجات التي تستوفي المعايير عند مراجع مستقل؛ الفرق بين التقييم الذاتي وتقييم المراجع؛ وعدد المهام التي يقيّمها خبيرا الموارد البشرية والتخصص قبل اعتمادها (الهدف: المهام الست).
What is decided, and what remains research?
This page separates prototype decisions from research results. No validation study outcomes or new accuracy figures are presented here; this is a reviewable protocol outline only.
01Resolved for prototypeFirst reference department
The Department of Information Technology, Faculty of Computing & Information Technology, King Abdulaziz University is the current course namespace and pilot-adapter reference. This does not imply institutional approval or official launch.
02Resolved for public releaseNo generative Arabic model drives current Fit judgments
The public release uses in-browser Arabic OCR (Tesseract language data) for text extraction, followed by governed rules, ontology mappings, and graph traversal. Any future Arabic language model is a separate decision requiring evaluation, licensing, and governance before affecting product judgments.
03Plan — no resultsRetrospective study
Goal: compare Kamin outputs on a de-identified historical record set against a pre-defined human reference. Sample, inclusion rules, reference definition, outcomes, and analysis plan must be frozen before data access.
04Plan — no resultsEvidence-deposit completion rate
Goal: measure the share of invited participants who complete an approvable evidence deposit and reach a first capability profile. This metric must remain separate from optional research consent and must not collect transcript content in telemetry.
Measurement principles
- Published methodology Targets are not Results.
- Accuracy evaluation requires a documented human reference.
- Mapping coverage, evidence-cited judgment ratio, and time-to-first-value are product-health metrics.
- Psychometrics and sensitive traits do not enter Fit before independent validation.
Ethics status
This page does not claim ethics/IRB approval for these studies. Any use of real records or publication of results requires the appropriate ethical and regulatory pathway before execution.
Proposal dated 30 September 2026 — no resultsNine-month validation roadmap
Month 1 starts after protocol and institutional approval, not when this page is published. Dates and responsibilities require supervisor confirmation; no partnerships or recruited samples are announced.
The semester goal is a usable product and an initial study in weeks 10–12 after approval. Later months deepen validation and follow-up, with final handover in month ten. This roadmap does not postpone the semester deliverable.
Month 1–2: Scope and protocol approval
Supervisor + participating institutions · Confirm colleges, approvals, measures and data responsibilities. Target: two specialists review at least 20 pathways before mappings are accepted.
Month 2–3: Formative usability
UX team + advisors · Target: 12 Saudi students over two rounds, including Arabic and mobile. Measure six tasks using synthetic records and resolve blockers before scaling.
Month 3–4: Bounded end-of-semester pilot
Supervisor + approved colleges · Target: 60 participants, 20 per college. Report independent completion, assistance and failure with uncertainty, separated by language and device.
Month 5–6: Issuer-verification prototype
Evidence team + authorized issuer · Target: 20 synthetic test cases covering invalid signatures, expiry and revocation. No institutional evidence upgrades before negative cases and security review pass.
Month 7–8: Follow-up and reevaluation
Team + advisors · Approved optional follow-up, comparing improvements with the baseline on a separate test set; no causal employment-effect claims.
Month 9: Continuation and pricing decision
Supervisor + career center · Target: 5 interviews with purchasing decision-makers; review results, limitations and operating cost. No automatic transition to a paid service.
Proposed pre-pilot thresholds — not results
- Independent completion ≥80% per core task; failures and assisted attempts remain in the denominator, and withdrawals are reported.
- Source-and-limit comprehension (2 of 2) in ≥80% of eligible attempts; two reviewers code a prespecified subset.
- Zero false evidence upgrades in the negative verification suite and zero unresolved critical privacy defects before a real-record pilot.
- Freeze definitions, sample and stopping rules before the study. These are continuation targets, not proof of accuracy, fairness or improved employment.
Usability tasks: build a profile, correct evidence, understand a recommendation, choose a step, restore a backup, and correct, withdraw and delete data. Report counts and uncertainty alongside rates; a missed target requires remediation and remeasurement.
Review speed and quality — pre-study targets
Uncalibrated proposal; no published field results. Freeze definitions, cases and targets with the supervisor before collecting results; separate language, device and reviewer role.
Time to confident acceptance or rejection
From first presentation of the review task and recommendation to acceptance or rejection with the highest of three confidence levels and a correct source-and-limit explanation assessed independently. Initial target: ≥80% of eligible tasks end in a qualifying decision, with median time ≤120 seconds among qualifying attempts. Selecting “confident” alone does not make a decision correct.
Report contests, unresolved outcomes, failures, assistance, interruptions and withdrawals separately; do not remove them to improve completion rates. The optional local timer starts when explicitly enabled, so it is not automatically equivalent to study first-exposure timing. Interrupted local attempts are excluded only from time summaries, with counts disclosed.
Errors that escaped verification
In a frozen synthetic suite: known-defective cases accepted after review divided by all known-defective cases presented for review. Initial target ≤5%, with zero accepted false evidence upgrades or critical privacy defects. Also report confirmed errors in an independent sample of accepted decisions with its denominator; do not combine that rate with the seeded-suite rate. Untriaged contests are not confirmed errors.
Rule and review-boundary regression checks now run in the repository. Student/advisor decision studies, institutional-source verification and central contest triage have not been completed. Engineering test results are not human-study results.
Task-loop measures
Pre-study targets, not results: share of started tasks that reach an attached output; share of outputs meeting the rubric under an independent reviewer; gap between self-assessment and reviewer assessment; and the number of tasks assessed by the two HR and domain experts before adoption (target: all six).