CSVREAD is returning a row of zeros for every other row
古いコメントを表示
I have a matrix titled 'S.csv' that is essentially a compilation of "survival ratios" for different age groups, so most of the values are just below 1. The first few columns are years and "group numbers", so I am making sure to leave those out when I call CSVREAD on this matrix.
Here is what I am entering, I am reading the matrix between rows 0 to 47 and 3 to 101:
S = csvread('S.csv',0,3, [0, 3, 47, 101])
For some unknown reason, this returns a matrix where every other row is all zeros. The non-zero rows all have their correct respective values, and "S" is returned as the correct size (48x99). I really can't figure out why every other row is just zeros.
Has anyone ever seen this happen with the CSVREAD function? I can't find any documentation on this.
2 件のコメント
dpb
2020 年 1 月 22 日
@KAE -- read through the responses here carefully and you'll see the problem OP had is related to a badly formatted input file. In all likelihood your problem is, too. If that isn't enough of a hint to be able to find it in your case, post a new Question and include a sample of the file that errors, we can't diagnose what we can't see.
回答 (3 件)
dpb
2015 年 11 月 14 日
Your data were saved as a character string with a comma delimiter between fields within the strings. IOW, the first line looks like
"2014,1,1,0.995512232,0.9991681652,0.9996800384,0.9997750236,...
for each line which isn't a valid csv numeric format. You can get around (other than fixing the spreadsheet or other app that created it) by
>> x=textread('ben1.csv','','delimiter',',','whitespace','"');
>> whos x
Name Size Bytes Class Attributes
x 282x104 234624 double
>>
This uses the optional whitespace argument to tell textread to ignore the apostrophes in the file. textscan has the same facility.
8 件のコメント
per isakson
2015 年 11 月 15 日
編集済み: per isakson
2015 年 11 月 15 日
... and an empty format string, which is a poorly documented option. It is used by TMW in the function, dlmread.
Ben Lockhart
2015 年 11 月 15 日
per isakson
2015 年 11 月 15 日
編集済み: per isakson
2015 年 11 月 15 日
There is no elegant way to do that. Read the entire file and pick the part you want or copy the solution used in dlmread. Search for the string "textscan". However, that is a waste of time.
per isakson
2015 年 11 月 15 日
編集済み: per isakson
2015 年 11 月 15 日
If size matters, use a method similar to http://se.mathworks.com/matlabcentral/answers/57276-import-data-file-with-comma-decimal-point#answer_69283. In comma2point_overwrite replace
comma = uint8(',');
point = uint8('.');
by
comma = uint8('"');
point = uint8(' ');
and use the function to replace '"' by ' ' in a separate step. Then read the file with csvread
Walter Roberson
2015 年 11 月 15 日
If you add the 'Whitespace', '"' to the textscan() that I used in my Answer then the skipcols controls how many leading columns to ignore. If there are trailing columns that are being ignored a similar change could be made to add a repmat('%*f', 1, number_of_skipped_trailing_columns) to the fmt
@Per re' "...an empty format string, which is a poorly documented option."
Indeed. There were another couple of threads just last couple days wherein this was explored in some depth. The "feature" is an absolute jewel in that it not only has the facility of letting one handle generic files such as OP's with an unknown number of fields per record and automagically return the correct shape, it consequently allows one to avoid the [proverbial-appendage-here]-ugly construct of repmat and/or a whole slew of field descriptors in the formatting string to explicitly match the record.
It does, of course, require the file be regular and all numeric.
I've submitted an enhancement request to TMW on both the patch to dlmread to get around the lack of a 'MultipleDelimsAsOne' option effective if the delimiter is given specifically in order to still be able to use the offset parameters as well as pointing out the deficiency in documentation regarding the behavior of textscan and textread with the empty format string.
I also continue to hammer on the fact that the textread abilities to accept a file name and to return a "regular" double array should be included in textscan or textread should be maintained and enhanced to be equivalent. Probably what should happen to it is that it becomes a wrapper to textscan as csvread is to dlmread (which also devolves to a textscan call in the end) altho the facility to use the file name in lieu of a file handle and still have a repetitive call on file to handle the case of the multi-segmented file would be handy as well, albeit a workaround exists.
ADDENDUM
In followup to the TMW support engineer responding to a query for more detail on the text[read|scan] behavior, I dug into the history of dlmread from V5.2 thru R11, R12, R14 and R2012b. It's an interesting history of evolution from an initial implementation as a doubly-nested loop using fgets and parsing every line character-by-character, to textread with an explicit format string (first incarnation inside a try...catch block that reverted to the V5.2 algo which was packaged as an internal subroutine so must have been pretty fragile at first :) ) to beginning with R12 the first introduction of the empty format string. This was carried over in R14 and later with the same idiom but using textscan instead.
Meanwhile, the documentation that initially consisted of a single example in textread that mentioned the empty format string but doesn't say anything about the shape slipped from release to release until there's only one seemingly unrelated example using the 'MissingValue' keyword in textread left that has the empty format string and no mention nor any example with textscan
Ben Lockhart
2015 年 11 月 15 日
dpb
2015 年 11 月 15 日
Yes, I had shown that in earlier response before we determined the problem in the formatting of the input file.
I'd still suggest fixing that in the file-generation process is step one...
dpb
2015 年 11 月 14 日
Don't have your file for testing, but I think the RANGE argument is incorrect. Try
S = csvread('S.csv',[0, 3, 47, 101]);
Of course, it's probably just as simple to read the whole array and then just eliminate what you don't want.
S=csvread('S.csv'); % read the file
S=S(1:48,4:102); % select the subarray desired (salt to suit ranges, of course)
2 件のコメント
Ben Lockhart
2015 年 11 月 14 日
dpb
2015 年 11 月 16 日
>> dlmread('ben.csv')
Error using dlmread (line 143)
Mismatch between file and format string.
Trouble reading number from file (row 1u, field 1u) ==>
"2014,1,1,0.995512232,0.9991681652,0.9996800384,0.9997750236,...
>>
BTW, if you had posted the full text of the error including the failed line that is echo'ed, we could have seen what the problem was then...which impresses that it is important to read ALL of the information pertaining to an error carefully...note there's the quote mark at the beginning of the line that is the culprit.
Walter Roberson
2015 年 11 月 14 日
Historically, csvread() and dlmread() could not be used for files that had any text in them at all, even in the header lines. That changed fairly recently, but it is possible that it still has bugs.
You should use textscan() instead:
skipcols = 3; %check whether this should be 2 or 3!
ncol = 99;
fmt = [repmat('%*s', 1, skipcols), repmat('%f', 1, ncol)];
fid = fopen('S.csv', 'rt');
datacell = textscan(fid, fmt, 'HeaderLines', 1, 'Delimiter', ',', 'CollectOutput', 1);
fclose(fid);
S = datacell{1};
3 件のコメント
Ben Lockhart
2015 年 11 月 14 日
Walter Roberson
2015 年 11 月 14 日
Could you attach the csv file for investigation?
Also, is it possible that the file was prepared on MS Windows but that you are reading it on OS-X or Linux?
Ben Lockhart
2015 年 11 月 14 日
カテゴリ
ヘルプ センター および File Exchange で Large Files and Big Data についてさらに検索
Community Treasure Hunt
Find the treasures in MATLAB Central and discover how the community can help you!
Start Hunting!