CSVREAD is returning a row of zeros for every other row

I have a matrix titled 'S.csv' that is essentially a compilation of "survival ratios" for different age groups, so most of the values are just below 1. The first few columns are years and "group numbers", so I am making sure to leave those out when I call CSVREAD on this matrix.
Here is what I am entering, I am reading the matrix between rows 0 to 47 and 3 to 101:
S = csvread('S.csv',0,3, [0, 3, 47, 101])
For some unknown reason, this returns a matrix where every other row is all zeros. The non-zero rows all have their correct respective values, and "S" is returned as the correct size (48x99). I really can't figure out why every other row is just zeros.
Has anyone ever seen this happen with the CSVREAD function? I can't find any documentation on this.

2 件のコメント

KAE
KAE 2020 年 1 月 22 日
編集済み: KAE 2020 年 2 月 5 日
Yes, I had this problem too due to poor formatting. I worked around it by generating a function using the Import Data button.
dpb
dpb 2020 年 1 月 22 日
@KAE -- read through the responses here carefully and you'll see the problem OP had is related to a badly formatted input file. In all likelihood your problem is, too. If that isn't enough of a hint to be able to find it in your case, post a new Question and include a sample of the file that errors, we can't diagnose what we can't see.

サインインしてコメントする。

回答 (3 件)

dpb
dpb 2015 年 11 月 14 日

2 投票

Your data were saved as a character string with a comma delimiter between fields within the strings. IOW, the first line looks like
"2014,1,1,0.995512232,0.9991681652,0.9996800384,0.9997750236,...
for each line which isn't a valid csv numeric format. You can get around (other than fixing the spreadsheet or other app that created it) by
>> x=textread('ben1.csv','','delimiter',',','whitespace','"');
>> whos x
Name Size Bytes Class Attributes
x 282x104 234624 double
>>
This uses the optional whitespace argument to tell textread to ignore the apostrophes in the file. textscan has the same facility.

8 件のコメント

per isakson
per isakson 2015 年 11 月 15 日
編集済み: per isakson 2015 年 11 月 15 日
... and an empty format string, which is a poorly documented option. It is used by TMW in the function, dlmread.
Ben Lockhart
Ben Lockhart 2015 年 11 月 15 日
Thank you so much. That fixed the problem.
One last question, how would I edit the boundaries to only read a certain number of rows/columns, like I did when I called CSVREAD?
per isakson
per isakson 2015 年 11 月 15 日
編集済み: per isakson 2015 年 11 月 15 日
There is no elegant way to do that. Read the entire file and pick the part you want or copy the solution used in dlmread. Search for the string "textscan". However, that is a waste of time.
per isakson
per isakson 2015 年 11 月 15 日
編集済み: per isakson 2015 年 11 月 15 日
If size matters, use a method similar to http://se.mathworks.com/matlabcentral/answers/57276-import-data-file-with-comma-decimal-point#answer_69283. In comma2point_overwrite replace
comma = uint8(',');
point = uint8('.');
by
comma = uint8('"');
point = uint8(' ');
and use the function to replace '"' by ' ' in a separate step. Then read the file with csvread
Walter Roberson
Walter Roberson 2015 年 11 月 15 日
If you add the 'Whitespace', '"' to the textscan() that I used in my Answer then the skipcols controls how many leading columns to ignore. If there are trailing columns that are being ignored a similar change could be made to add a repmat('%*f', 1, number_of_skipped_trailing_columns) to the fmt
dpb
dpb 2015 年 11 月 15 日
編集済み: dpb 2015 年 11 月 17 日
@Per re' "...an empty format string, which is a poorly documented option."
Indeed. There were another couple of threads just last couple days wherein this was explored in some depth. The "feature" is an absolute jewel in that it not only has the facility of letting one handle generic files such as OP's with an unknown number of fields per record and automagically return the correct shape, it consequently allows one to avoid the [proverbial-appendage-here]-ugly construct of repmat and/or a whole slew of field descriptors in the formatting string to explicitly match the record.
It does, of course, require the file be regular and all numeric.
I've submitted an enhancement request to TMW on both the patch to dlmread to get around the lack of a 'MultipleDelimsAsOne' option effective if the delimiter is given specifically in order to still be able to use the offset parameters as well as pointing out the deficiency in documentation regarding the behavior of textscan and textread with the empty format string.
I also continue to hammer on the fact that the textread abilities to accept a file name and to return a "regular" double array should be included in textscan or textread should be maintained and enhanced to be equivalent. Probably what should happen to it is that it becomes a wrapper to textscan as csvread is to dlmread (which also devolves to a textscan call in the end) altho the facility to use the file name in lieu of a file handle and still have a repetitive call on file to handle the case of the multi-segmented file would be handy as well, albeit a workaround exists.
ADDENDUM
In followup to the TMW support engineer responding to a query for more detail on the text[read|scan] behavior, I dug into the history of dlmread from V5.2 thru R11, R12, R14 and R2012b. It's an interesting history of evolution from an initial implementation as a doubly-nested loop using fgets and parsing every line character-by-character, to textread with an explicit format string (first incarnation inside a try...catch block that reverted to the V5.2 algo which was packaged as an internal subroutine so must have been pretty fragile at first :) ) to beginning with R12 the first introduction of the empty format string. This was carried over in R14 and later with the same idiom but using textscan instead.
Meanwhile, the documentation that initially consisted of a single example in textread that mentioned the empty format string but doesn't say anything about the shape slipped from release to release until there's only one seemingly unrelated example using the 'MissingValue' keyword in textread left that has the empty format string and no mention nor any example with textscan
Ben Lockhart
Ben Lockhart 2015 年 11 月 15 日
I was really overthinking that last question. The matrix has already been created by textread, so all I had to do was make a new matrix using subarrays by doing something like:
S=S(1:48,4:102);
dpb
dpb 2015 年 11 月 15 日
Yes, I had shown that in earlier response before we determined the problem in the formatting of the input file.
I'd still suggest fixing that in the file-generation process is step one...

サインインしてコメントする。

dpb
dpb 2015 年 11 月 14 日

0 投票

Don't have your file for testing, but I think the RANGE argument is incorrect. Try
S = csvread('S.csv',[0, 3, 47, 101]);
Of course, it's probably just as simple to read the whole array and then just eliminate what you don't want.
S=csvread('S.csv'); % read the file
S=S(1:48,4:102); % select the subarray desired (salt to suit ranges, of course)

2 件のコメント

Ben Lockhart
Ben Lockhart 2015 年 11 月 14 日
I'm unable to just read the matrix as a whole. It gives this error:
Error using dlmread (line 138) Mismatch between file and format string. Trouble reading 'Numeric' field from file (row number 1, field number 1)
dpb
dpb 2015 年 11 月 16 日
>> dlmread('ben.csv')
Error using dlmread (line 143)
Mismatch between file and format string.
Trouble reading number from file (row 1u, field 1u) ==>
"2014,1,1,0.995512232,0.9991681652,0.9996800384,0.9997750236,...
>>
BTW, if you had posted the full text of the error including the failed line that is echo'ed, we could have seen what the problem was then...which impresses that it is important to read ALL of the information pertaining to an error carefully...note there's the quote mark at the beginning of the line that is the culprit.

サインインしてコメントする。

Walter Roberson
Walter Roberson 2015 年 11 月 14 日

0 投票

Historically, csvread() and dlmread() could not be used for files that had any text in them at all, even in the header lines. That changed fairly recently, but it is possible that it still has bugs.
You should use textscan() instead:
skipcols = 3; %check whether this should be 2 or 3!
ncol = 99;
fmt = [repmat('%*s', 1, skipcols), repmat('%f', 1, ncol)];
fid = fopen('S.csv', 'rt');
datacell = textscan(fid, fmt, 'HeaderLines', 1, 'Delimiter', ',', 'CollectOutput', 1);
fclose(fid);
S = datacell{1};

3 件のコメント

Ben Lockhart
Ben Lockhart 2015 年 11 月 14 日
Same result except all the zeros are NaN's now.
Walter Roberson
Walter Roberson 2015 年 11 月 14 日
Could you attach the csv file for investigation?
Also, is it possible that the file was prepared on MS Windows but that you are reading it on OS-X or Linux?
Ben Lockhart
Ben Lockhart 2015 年 11 月 14 日
Sure! Here is the link. If you could take a look at why this happens with CSVREAD and/or suggest a workaround, I would be very grateful.

サインインしてコメントする。

カテゴリ

ヘルプ センター および File Exchange で Large Files and Big Data についてさらに検索

質問済み:

2015 年 11 月 14 日

編集済み:

KAE
2020 年 2 月 5 日

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by