Summary
Improve record grouping, ordering, and filtering in the annotation queue to enhance user experience when working with document extraction data. This phase leverages the new workspace-level schema configuration infrastructure to provide intelligent, schema-aware improvements.
Motivation
The current implementation has several issues that limit the effectiveness of multi-stage document extraction workflows:
- Records from the same document/reference appear scattered throughout the annotation queue
- Schemas are not displayed in a consistent, dependency-based order based on workspace configuration
- The UI only allows filtering by one record status at a time, limiting workflow flexibility
- These issues make it difficult for users to efficiently annotate related records and manage complex extraction workflows
Strategic Context
This phase is part of the larger workspace-level schema management architecture redesign. It focuses on leveraging workspace schema configuration to provide intelligent record organization and filtering, setting the foundation for the document-schema-fields table view.
Proposed Refactor
This phase focuses on three main improvements that work together:
1. Record Grouping by Reference (#87)
- Leverage Workspace Schema Configuration: Use schema configuration to understand document structure
- Integrate with SchemaService: Use the new SchemaService for intelligent grouping logic
- Backend Enhancement: Modify search API to support schema-aware reference grouping
- Frontend Integration: Update UI to display grouped records with schema context
2. Schema Topological Ordering (#88)
- Use Workspace Configuration: Leverage schema relationships defined in workspace config
- SchemaService Integration: Use SchemaService to compute topological ordering automatically
- Consistent Ordering: Ensure ordering is consistent across all application components
- Visual Indicators: Show schema relationships and dependencies in UI
3. Multi-Status Filtering (#89)
- Workspace-Aware Filtering: Use workspace schema configuration for intelligent status filtering
- Document-Centric Approach: Enable filtering that considers document-level status across schemas
- Enhanced UI: Create multi-select filtering controls with workflow awareness
- Cross-Dataset Support: Consider status across related datasets in multi-stage workflows
Dependencies
This phase builds on the workspace-level schema management infrastructure:
- Required: Workspace Schema Configuration Infrastructure
- Recommended: SchemaService implementation
- Future Enhancement: Document-centric APIs (for optimal implementation)
Implementation Strategy
Following the strategic principle of "make it run, make it right, make it fast":
- Make it Run: Implement basic functionality using workspace schema configuration
- Make it Right: Refine the implementation with proper schema service integration
- Make it Fast: Optimize performance with caching and efficient queries
Acceptance Criteria
Functional Requirements
Technical Requirements
Quality Requirements
Related Issues
Direct Components
Strategic Dependencies
- Workspace Schema Configuration Infrastructure (foundational requirement)
- SchemaService Implementation (enhances functionality)
- Document-centric API Layer (future enhancement)
Next Phases
Success Metrics
- User Experience: Reduced time to navigate between related records
- Workflow Efficiency: Improved annotation throughput for multi-schema documents
- System Performance: Maintained or improved response times
- Development Velocity: Clear foundation for subsequent phases
This is Phase 1 of the larger Extralit Document Extraction Data Architecture refactoring plan. This phase focuses on improving the user experience by leveraging workspace schema configuration without requiring significant changes to the underlying data model, while setting the foundation for more advanced capabilities in subsequent phases.
Summary
Improve record grouping, ordering, and filtering in the annotation queue to enhance user experience when working with document extraction data. This phase leverages the new workspace-level schema configuration infrastructure to provide intelligent, schema-aware improvements.
Motivation
The current implementation has several issues that limit the effectiveness of multi-stage document extraction workflows:
Strategic Context
This phase is part of the larger workspace-level schema management architecture redesign. It focuses on leveraging workspace schema configuration to provide intelligent record organization and filtering, setting the foundation for the document-schema-fields table view.
Proposed Refactor
This phase focuses on three main improvements that work together:
1. Record Grouping by Reference (#87)
2. Schema Topological Ordering (#88)
3. Multi-Status Filtering (#89)
Dependencies
This phase builds on the workspace-level schema management infrastructure:
Implementation Strategy
Following the strategic principle of "make it run, make it right, make it fast":
Acceptance Criteria
Functional Requirements
Technical Requirements
Quality Requirements
Related Issues
Direct Components
Strategic Dependencies
Next Phases
Success Metrics
This is Phase 1 of the larger Extralit Document Extraction Data Architecture refactoring plan. This phase focuses on improving the user experience by leveraging workspace schema configuration without requiring significant changes to the underlying data model, while setting the foundation for more advanced capabilities in subsequent phases.