Skip to content

[Refactor] Phase 1: Improve record grouping, ordering, and filtering #95

Description

@JonnyTran

Summary

Improve record grouping, ordering, and filtering in the annotation queue to enhance user experience when working with document extraction data. This phase leverages the new workspace-level schema configuration infrastructure to provide intelligent, schema-aware improvements.

Motivation

The current implementation has several issues that limit the effectiveness of multi-stage document extraction workflows:

  1. Records from the same document/reference appear scattered throughout the annotation queue
  2. Schemas are not displayed in a consistent, dependency-based order based on workspace configuration
  3. The UI only allows filtering by one record status at a time, limiting workflow flexibility
  4. These issues make it difficult for users to efficiently annotate related records and manage complex extraction workflows

Strategic Context

This phase is part of the larger workspace-level schema management architecture redesign. It focuses on leveraging workspace schema configuration to provide intelligent record organization and filtering, setting the foundation for the document-schema-fields table view.

Proposed Refactor

This phase focuses on three main improvements that work together:

1. Record Grouping by Reference (#87)

  • Leverage Workspace Schema Configuration: Use schema configuration to understand document structure
  • Integrate with SchemaService: Use the new SchemaService for intelligent grouping logic
  • Backend Enhancement: Modify search API to support schema-aware reference grouping
  • Frontend Integration: Update UI to display grouped records with schema context

2. Schema Topological Ordering (#88)

  • Use Workspace Configuration: Leverage schema relationships defined in workspace config
  • SchemaService Integration: Use SchemaService to compute topological ordering automatically
  • Consistent Ordering: Ensure ordering is consistent across all application components
  • Visual Indicators: Show schema relationships and dependencies in UI

3. Multi-Status Filtering (#89)

  • Workspace-Aware Filtering: Use workspace schema configuration for intelligent status filtering
  • Document-Centric Approach: Enable filtering that considers document-level status across schemas
  • Enhanced UI: Create multi-select filtering controls with workflow awareness
  • Cross-Dataset Support: Consider status across related datasets in multi-stage workflows

Dependencies

This phase builds on the workspace-level schema management infrastructure:

  • Required: Workspace Schema Configuration Infrastructure
  • Recommended: SchemaService implementation
  • Future Enhancement: Document-centric APIs (for optimal implementation)

Implementation Strategy

Following the strategic principle of "make it run, make it right, make it fast":

  1. Make it Run: Implement basic functionality using workspace schema configuration
  2. Make it Right: Refine the implementation with proper schema service integration
  3. Make it Fast: Optimize performance with caching and efficient queries

Acceptance Criteria

Functional Requirements

  • Records with the same reference are grouped together in the annotation queue using workspace schema configuration
  • Schema ordering is derived from workspace configuration and applied consistently
  • Users can filter records by any combination of available statuses with workspace awareness
  • UI clearly indicates schema-based record relationships and workflow context

Technical Requirements

  • Backend APIs leverage workspace schema configuration for intelligent operations
  • SchemaService is used for grouping, ordering, and filtering logic
  • Frontend components integrate with workspace schema configuration
  • Performance is maintained or improved with the new functionality

Quality Requirements

  • All existing functionality works correctly (sorting, search, etc.)
  • Integration tests verify proper grouping, ordering, and filtering behavior
  • Error handling provides meaningful feedback to users
  • The system maintains backward compatibility during transition

Related Issues

Direct Components

Strategic Dependencies

  • Workspace Schema Configuration Infrastructure (foundational requirement)
  • SchemaService Implementation (enhances functionality)
  • Document-centric API Layer (future enhancement)

Next Phases

Success Metrics

  • User Experience: Reduced time to navigate between related records
  • Workflow Efficiency: Improved annotation throughput for multi-schema documents
  • System Performance: Maintained or improved response times
  • Development Velocity: Clear foundation for subsequent phases

This is Phase 1 of the larger Extralit Document Extraction Data Architecture refactoring plan. This phase focuses on improving the user experience by leveraging workspace schema configuration without requiring significant changes to the underlying data model, while setting the foundation for more advanced capabilities in subsequent phases.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

refactorCode refactoring or technical debt improvements

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions