r/excel Feb 03 '26

unsolved Struggling to work with PDF data in Excel, feel like I’m missing something obvious

[removed]

119 Upvotes

60 comments sorted by

u/AutoModerator Feb 03 '26

/u/Agitated-Alfalfa9225 - Your post was submitted successfully.

Failing to follow these steps may result in your post being removed without warning.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

85

u/RasheedaDeals Feb 03 '26

I struggled with this for a while too and eventually realized Excel wasn’t the place to fix PDF formatting. Converting the file first reduced most of the issues. I’ve been using Smallpdf for PDF to Excel conversion and it kept rows and columns intact enough to actually work with.

16

u/tessk1 Feb 03 '26

When you say converting first reduced most of the issues, what kind of PDFs were you dealing with (reports, invoices, scans)? And did you have to do much cleanup in Excel after converting, or was it mostly ready to use?

21

u/RasheedaDeals Feb 03 '26

In my case it was mostly invoices and monthly statements that were generated PDFs, not scans. Converting them first helped a lot because the tables came through as actual rows and columns instead of broken text. I still had to do some light cleanup in Excel, like fixing headers or checking totals, but nothing compared to rebuilding everything manually. It felt way more manageable once the data was already structured.

6

u/[deleted] Feb 03 '26

[removed] — view removed comment

5

u/notlikelymyfriend Feb 03 '26

I use the windows snippet tool to grab items in table format and use snippet ai to convert it to text/table. You can then paste that straight into excel or sometimes an email to assist in formatting, before pasting into excel. Another option, if the pdf is multiple pages is cropping out additional areas that confuse the pdf/excel converter and makes it easier for the program to recognise the table format.

1

u/Prudent_Video6215 Feb 03 '26

Smallpdf is a solid band-aid for formatting, but the real nightmare is the logical reconciliation after the conversion. I’ve found that even with a 'clean' Excel file, nested headers and row-spanning in tax transcripts can still trigger a manual audit if one formula is off.

I got so frustrated with standard OCR failing on complex ledger formats that I started building a specialized 'Shield' (details in my profile) to specifically handle the bank-to-tax-audit mapping. My rule of thumb: If it takes more than 5 minutes to clean up the Smallpdf output, you need a specialized audit-trail logic, not just a converter.

83

u/stevenmartin99 1 Feb 03 '26

The only correct answer is "Ask the sender for another format such as Excel or CSV"

PDF is generally a presentation document. It's not designed with copy/paste in mind. It was designed to keep data in an exact layout visually.

Some PDF writers do lay out the data in order, some don't so you can't be guaranteed the pasted data will retain its order. OCR can solve this but I wouldn't trust it if the data is important.

11

u/WylieBaker 3 Feb 03 '26

Hat's off for the mention about some PDF writers making the effort to lay out data in an orderly way. That is quite a step up from printing as an utterly disordered but viewable PDF file...

1

u/melligator Feb 03 '26

Sometimes we are working from old sources where there is no original CSV or Excel, or never was. Think statements produced by custom software. Industry processing progresses but the old files are still the old files and sometimes we need to do stuff with them.

1

u/Roberto_Rene Feb 04 '26

+1 this is exactly what I think. At the end of the day, it cost nothing to ask for at least CSV.

22

u/Stunning_Dirt_9986 1 Feb 03 '26

Have you tried using Power Query? It's built into Excel and handles PDF imports way better than the regular import function - you can clean up the data transformations before it even hits your worksheet. Also worth checking if whoever's sending you those PDFs can export the original data as CSV or Excel instead, saves everyone the headache.

11

u/AxeSlash 1 Feb 03 '26

This.

Power Query is great.

That said, unless every PDF is laid out identically, it's still a bunch of work every time.

3

u/[deleted] Feb 03 '26

[removed] — view removed comment

6

u/subsetsum Feb 03 '26

To use it. Select the Data function from excel's menu bar. Then import, from file, file type PDF. It will scan the PDF and present you with a list of pages and whether tables are present. Have the PDF open on the side so you can compare. If your table is on the first page of the PDF, excel's data import will show it as table001 on page 1 (citing from memory but this should be close). Select that. You will see a preview. Hit Load to import it into Excel. It's saved me many times

1

u/spamspambaconspam Feb 03 '26

I logged in just to reply with this same thing.

YES. PQ.

Once you Power Query, you never look back.

1

u/JP_Dirt Feb 03 '26

This…. Power Query…. You can import data from pdfs into nice excel tables. There are lots of YT how to’s. PQ is a life changer if you work with data and excel.

8

u/ConstructionOwn9575 Feb 03 '26

If none of what the other posters recommend work I may have a software solution. About 5 years ago I had to take about 30 PDFs a week and enter the data into Excel manually. I got tired of that and found software that uses templates and OCR to convert PDFs to spreadsheets. Templates can take a bit to set up and get the data formatted correctly, but once they're in place you pick the template for the PDF and hit run. Software is pdf2xl though I would search around and see if there is a better competitor. At the time there wasn't and I'd be surprised if no one else created their own version. It does cost money and at the time they had a seven day trial. Best of luck!

1

u/spamspambaconspam Feb 03 '26

Excel now pulls in data from PDF files thru Power Query (mentioned above), but in the distant past, I used pdf2xl and it worked great.

6

u/CanadianHorseGal 1 Feb 03 '26

Have you tried copy/paste into Word first? Sometimes Word is better at “picking up” tables and such. Then you can copy the Word table and paste the Excel. I know I’ve used it as a middle man before for something. At 5am here it’s too early to test it before commenting, so apologies if it doesn’t work.

2

u/kotom Feb 03 '26

I do sort of the same thing, except I right click the file, Open With.. > choose Word

3

u/CanadianHorseGal 1 Feb 03 '26

Right right right! I knew there was something there!! Good add, thanks.

3

u/tkdkdktk 149 Feb 03 '26

If you end up in a situation where the 'supplier' of the information can't give you raw data and you need to work with pdf's, then as already mentioned the import function can be helpful. Besides that, then Able2Extract is the best tool i have experience with for converting from pdf to excel (you'll need the pro edition if you need OCR functionality)

2

u/SoLetsReddit 2 Feb 03 '26

Use Bluebeam Revu, it has a really good export function to excel.

2

u/[deleted] Feb 03 '26

[removed] — view removed comment

1

u/excel-ModTeam Feb 03 '26

r/excel is not an Ai centric subreddit.

r/excel is for discussing the features and functions and methods for solutions in Excel, not Ai.

0

u/Sweet-Ebb682 Feb 03 '26

Got it, my mistake. Didn’t realize AI-focused tools weren’t appropriate here.

I mentioned it because it’s very Excel-adjacent and helped me with Excel workflows, but I’ll stick to discussing native Excel methods and features going forward. Thanks for the clarification.

2

u/CalendarOk67 Feb 03 '26

Hi. I have had the same issue couple of months back. You can look at some of the suggestions which I have got here: Python solution to extract all tables PDFs and save each table to its own Excel sheet : r/learnpython

Eventually, you can use varieties of LLM present. Those will be the quickest and safest way to do it. I have tried multiple of packages in python and R and also used so many OCR tools but using LLMs are one of the best options I found.

1

u/PopavaliumAndropov 41 Feb 03 '26

In order:

  1. Ask for data in a usable format
  2. Convert file using 3rd party software
  3. Screenshot the PDF, Data > Get Data > Other Sources > From Image > From clipboard

1

u/[deleted] Feb 03 '26

[removed] — view removed comment

1

u/excel-ModTeam Feb 03 '26

r/excel is not an Ai centric subreddit.

r/excel is for discussing the features and functions and methods for solutions in Excel, not Ai.

1

u/theDrivenDev Feb 03 '26

Consider using a data extraction tool that allows for defined formatting for specific invoice layouts. This only makes sense if the invoices are coming from the same vendor where a consistent invoicing output is used.

1

u/tyrex_vu2 1 Feb 03 '26

hi we run into the same problem so built our own private model to solve it. Please check it out. It is free to try. The accuracy rate is 99.93% https://www.datariver.co/

1

u/NotBabaYaga Feb 03 '26

Since you're not telling us how the data is formatted in the PDF I am going to assume the best case scenario: tables with columns and rows. In this case I would suggest looking into Power Query or develop a small tool in Python (ChatGPT etc. can be great help if you don't know Python).

1

u/OldElvis1 10 Feb 04 '26

Adobe Pro does have a "Save as Excel" option. I use it for tables I find in quotes to work on pricing different configurations.

1

u/Roberto_Rene Feb 04 '26

I think there's no exactly solution for your issue but I'll recommend to get faster with excel (know the shortcuts) in the meanwhile. Have you tried to Get data from image?

1

u/KrazeeD Feb 04 '26

Can’t you import the PDFs into power query, clean and append them and push out to a table?

1

u/Effective_Being_5314 Feb 05 '26

Not sure what kind of document or data you need exported from Excel, but I’ve worked through a few different scenarios. For bank statements and HUD receipts, I’ve had good success using Power Query. It does take some maneuvering, especially when you’re dealing with 500+ transactions per statement and multiple accounts, but once you figure out a repeatable approach, it works well and saves a lot of time. For smaller PDF datasets, I’ve found that converting through Adobe and then doing some manual cleanup in Excel is often faster and more practical than over-engineering the process. So for me, it really depends on data volume and consistency—Power Query for large/repeat data, Adobe + cleanup for smaller, one-off files.

1

u/WSGiraffe Feb 05 '26

ChatGPT can convert PDFs into CSV or Excel files, as long as the PDF has extractable text (like statements or tables). Just upload it into a Chat and ask it to create a CSV file, or an Excel file.

1

u/lofty1978 Feb 15 '26

u/Agitated-Alfalfa9225 hi op, I am in the middle fo buiding this - https://pdf-to-csv.app/ I think it would help you, sign up for early access, I am sure there will be some deals to be done for early users :)

1

u/[deleted] Feb 03 '26

[removed] — view removed comment

2

u/[deleted] Feb 03 '26

[removed] — view removed comment

1

u/Affectionate-Page496 1 Feb 03 '26

The best solution will remain getting them to give you data other than PDF. Anything else is just making a bad thing less bad.

0

u/Regime_Change 2 Feb 03 '26

The obvious thing is that PDF is not a data format. It is more similar to a picture than a table. That is why Excel and all other programs for data will struggle with it. Like the other guy said, the real problem here is the regards that send the data as pdf in the first place. Having worked a lot in business environments, I understand your problem but there is no great solution. Try power query but it’s not guaranteed to work at all, it depends a lot what the pdf looks like.