How a GovTech AI Startup Moved Millions of Files Off SharePoint Without Rebuilding Its Government AI Pipeline
A software startup builds private AI assistants for US local governments. It serves dozens of city and county clients, delivering purpose-built assistants for the functions a local government actually runs: the city manager's office, planning and community development, HR, legal. Each assistant is built entirely from that government's own records. Meeting minutes, staff reports, zoning codes, master plans, and internal policies become a closed-loop, retrieval-based knowledge base that staff query instead of digging through archives.
That proposition puts a serious file operation at the center of the business. Every engagement begins with ingesting a client's record library: scraping, OCR, conversion, and sanitization run on an on-premises Synology NAS, and the converted output lands in that client's own Azure Blob Storage containers, where Azure AI Search builds the indexes the retrieval system reads. An assistant's answers are only as trustworthy as the pipeline that carried the records into it. And between the local processing tier and the per-client cloud containers, everything passed through a staging layer.
Two Document Libraries, Millions of Files, and SharePoint's Real Ceiling
That staging layer was SharePoint: two document libraries holding client data for the whole customer base, millions of files and terabytes of data.
The estate had reached SharePoint's practical ceiling. Automated processes were throttled. The large libraries would no longer sync. The throttling reached beyond the libraries themselves: OneNote notebooks elsewhere in the tenant, with no connection to client data, were throttled too. The startup wanted a staging layer with no file ceiling, so ingestion could grow with the client roster.
For a company whose product is AI built from a government's complete records, the staging layer has to carry every file, every time.
The estate was also expensive and labor-intensive to keep alive. Additional SharePoint storage expense kept climbing, and the limit had to be raised twice inside six weeks as new client data arrived. Moving processed files out of SharePoint and into each customer's Blob container depended on hand-rolled Microsoft Graph API scripts running as scheduled tasks, three to four streams at a time, monitored by hand. The path could be run and controlled only by the engineers who built it.
Why It Couldn't Be Fixed in Place
The problem could not be configured away. SharePoint's file ceiling was architectural, and client growth kept pushing past it: the startup was ingesting new records by the terabyte, with new city and county clients closing continuously.
The obvious alternatives did not fit. A direct transfer from the NAS to Azure Blob would have meant another hand-built path to run and monitor. And the backend of the product could not simply be rebuilt around new storage: the Azure AI Search indexers that powered every client's assistant were bound to the existing per-customer Blob containers, and retying them across dozens of clients and their separate data lakes would have meant reworking the foundation of the product itself.
So whatever replaced SharePoint had a specific job. It had to carry millions of files with no practical ceiling. It had to reach the NAS, because the OCR and conversion tooling required files native on that machine. It had to deliver converted files into each client's own Blob container on a schedule, with no scripts to babysit. It had to keep every government's data in its own scoped folder tree. And it had to take over with nothing missed, while ingestion kept running.
The startup selected Files.com to be that orchestration layer.
Files.com Between the NAS and Every Client's Blob Container
Files.com became the layer that presents the whole estate as one set of governed folders and runs the movement between its parts on its own.
On the local side, the Files.com Agent runs as a Docker container directly on the Synology NAS, connecting over an outbound-only connection. The NAS stays, because the conversion path needs local files, and its two client libraries now appear as ordinary Files.com folders. Internal delivery and audit staff work on them in the browser at a scale SharePoint could never serve.
On the cloud side, every customer's containers are connected to Files.com as remote servers, a pair per client: one holding converted and indexed documents and one holding the unconverted source files the application displays and cites. Scheduled Files.com Syncs push converted output into the container each client's Azure AI Search indexer reads. The indexers never moved.
Access is scoped with folder-level permissions that mirror each government's own organization, keeping each client isolated while giving departments the appropriate reach. A city manager's scope covers everything for that client; HR and legal reach only their own folders.
A single team member at the startup owned the implementation through roughly weekly guided onboarding sessions with a Files.com onboarding architect, doing the configuration between sessions. The Agent was configured the day the team first attempted it; that team member then built a test agency to validate file movement in both directions before setting up the first Sync. The existing Azure Blob remote servers were connected during onboarding, and the pattern then expanded across customer containers. Initial onboarding ended with the team member self-sufficient on the configuration.
The cutover protected the data the whole way. Files were copied out of SharePoint rather than moved, so SharePoint stayed intact as a fallback while the two systems ran in parallel for a handful of months. One client's flow went live first, then the pattern fanned out and became the standard for every new client. Preserving the full folder structure with nothing missed was the stated bar, and it held.
Unattended Syncs Without SharePoint Storage Overhead
With Files.com in production, the startup replaced a SharePoint estate at its ceiling and a hand-run script path with a file layer that runs itself.
- After the initial full syncs carried the load, every client sync runs on schedule with nobody watching it. The Microsoft Graph scripts and scheduled-task jobs that only their builders could run and monitor are retired.
- The additional SharePoint storage expense—and the repeated limit increases—stopped when the libraries left SharePoint.
- Client records sit on a layer with no file ceiling, absorbing new records by the terabyte.
- Onboarding the next government client is now a repeat of the same pattern: a scoped folder tree, a pair of Blob connections, and a Sync. The file layer is no longer the constraint on that growth.
The Pipeline Stayed, the Layer Underneath It Changed
The larger lesson is what the startup did not do. It never rebuilt its AI pipeline to fix its file problem. The Azure AI Search indexers stayed bound to the storage they already read, the NAS kept doing the local work only it can do, and Files.com replaced the layer between them while old and new ran side by side until the new one had proved itself. The file layer feeding a production RAG system turned out to be swappable, and the startup swapped it under a business that never stopped ingesting.
Related Customer Stories
A Domain Registry Runs Self-Service Zone File Distribution for Vetted Outsiders on Files.com
The registry separated vetting and entitlement from account creation, giving hundreds of approved outsiders self-service access without putting them in its own identity systems.
Read The Story
A Database Software Company Gives Every Support Ticket Its Own HTTPS or SFTP Intake Route With Files.com
API-driven, write-only intake lets customers deliver diagnostics through their firewalls while the company keeps no standing credentials for external uploaders.
Read The Story
A Network Security Vendor Retires Box by Moving a Handful of Beta Users to Files.com
The workload was small, but absorbing it into the file-transfer environment already feeding Oracle ERP eliminated an entire external sharing surface.
Read The Story
Get The File Orchestration Platform Today
4,000+ organizations trust Files.com for mission-critical file operations. Start your free trial now and build your first flow in 60 seconds.
No credit card required • 7-day free trial • Live in minutes